Dan’s Editorial Checks/Sonnet 5.5/02 personal-benchmark-paragraph-first-draft
02 personal-benchmark-paragraph-first-draft
Sonnet 5.5Claude Codehigh effortrun 23 Sep 2026269,823 tokens
▸Instructions — what the model was asked
okay now we need the "what is a personal benchmark" graf
Answer
What is a personal benchmark?
A personal benchmark is a test built from your own work. Every time you correct an AI, that correction can become a case: the task, plus the yes-or-no checks that capture what you were looking for. Run a new model on your cases and you can see how often it clears your bar. Public benchmarks are like SAT scores. They can tell a 300 from a 1600, but they can't help you choose between models that all land around 1570. A personal benchmark is a reference check instead: it tells you which model is best for your work, whoever makes it, and when a new one is worth switching to.
Notes
Where it goes: the empty blocks under "What is a personal benchmark?", before "We believe that there is great human work to do with AI." It is about 105 words, so the full strategy doc stays under the 500-word cap (the current draft is roughly 300).
Choices I made
- The SAT / reference-check analogy comes from your pitch to Rick. It's the clearest thing you said about why generic benchmarks fall short. I used your "1,600 vs. 300 vs. 1570" version. The Rick transcript is an unverified Granola Chat extract, so check that the numbers match what you remember.
- The paragraph doesn't make the "replace you vs. work for you" contrast. Your closing lines right after it make that point, and this paragraph would repeat them.
- I used "cases" and "checks," not "KateBench." Both terms appear in your draft and the Rick demo. KateBench is a product name and reads oddly as a definition. Swap it in if you want the tie-in.
- The paragraph doesn't say "one click, automatically." Your bullet above it already promises that, and the definition should stand on its own. Every Checks is internal-only today, and the correction-to-case step is still an ambition. The editorial review flagged that the page should separate today from next.
Optional closing sentence if you want to mark that line: "Today we build these by hand for ourselves; we want it to happen as you work."
Shorter version (about 60 words):
A personal benchmark is a test built from your own work. Your corrections become cases, each with yes-or-no checks for what you were looking for. Public benchmarks are like SAT scores, which don't help when every model scores about the same. A personal benchmark is a reference check. It shows which model is best for your work, whoever makes it, and when to switch.
Assumptions: I assumed the paragraph goes in the "Personal benchmarks" section of the Journal draft, in your voice ("we") and matching the plain, short-sentence style there. I didn't touch the Notion page. This file is the only output.
Checks
3/8ai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md
Q1D01. First sentence answers the questionthis case
Judge's reasoning
Both offered versions open with "A personal benchmark is a test built from your own work," which restates 'benchmark' (test) and 'personal' (your own work) without naming AI/models or what the reader gets from it.
▸Rubric
D01. Does the first sentence answer what a personal benchmark is in plain, accurate terms? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “Does it answer the question accurately in a way I understand?” PASS if: The first sentence tells a new reader what a personal benchmark does for their own work or their choice of AI model. Simple language such as measuring how well an AI works for you can pass, even if a sharper sentence would say it helps identify the best model for your work. FAIL if: The first sentence only sets up a generic benchmark, offers an analogy that needs decoding, or restates 'personal' and 'benchmark' without saying what the person gets from it. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.
Q2D03. Each sentence follows the previous onethis case
Judge's reasoning
The 60-word version jumps from "Your corrections become cases, each with yes-or-no checks" straight to "Public benchmarks are like SAT scores," skipping the step where models are actually run on those cases, and then asserts "A personal benchmark is a reference check." with no bridge to the following sentence.
▸Rubric
D03. Does the paragraph develop its explanation in a coherent sequence, without missing connections between ideas? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “Each sentence of this paragraph needs to connect from one to the next, and that breaks the connection.” PASS if: Each sentence builds on an established idea or makes its change of direction understandable. The explanation supplies the connections a reader needs before relying on them. FAIL if: A sentence abruptly changes subjects, leaves a necessary explanation unfinished, or skips a step in the reasoning so the reader must invent the missing connection. Judge continuity throughout the paragraph, allowing ordinary reader knowledge and established document context. No particular topic order or transition words are required; an implicit connection is enough when clear. A change of topic, contrast, or analogy can pass when its relationship to the preceding idea is understandable. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.
Q3D04. Analogy clicks without algebrathis case
Judge's reasoning
Both versions drop the second half of the analogy unmapped — "A personal benchmark is a reference check instead" (and the bare "A personal benchmark is a reference check." in the short version) — never establishing the hiring frame or saying who gives the reference, so the reader must reconstruct the correspondence.
▸Rubric
D04. If the paragraph uses an analogy, does it clarify the concept immediately? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “I have to do a lot of algebra to understand it.” PASS if: There is no analogy, or the comparison promptly maps ordinary benchmarks and personal benchmarks to familiar things and helps the reader understand the distinction. It does not require reconstructing the relationship for oneself. FAIL if: It leads with a compressed line such as 'a reference check, not an SAT score' before the reader knows what corresponds to what, or delays the mapping so the analogy is forgotten before it is explained. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.
Q4D05. Make personal standards concretethis case
Judge's reasoning
Neither version gives a single example of an actual standard; they stop at abstractions like "the yes-or-no checks that capture what you were looking for" and "how often it clears your bar," with no sample check such as whether a headline makes the strongest claim.
▸Rubric
D05. Does the paragraph make a person's own standard concrete? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “The thing that's good about this is it is concrete. It gives concrete examples, which is really, really important for this.” PASS if: It gives at least one understandable example of a standard the person could set for their work, such as whether a headline is interesting or follows AP style, and makes clear that models can be compared against those standards. Other good standards count; Dan's examples are illustrative, not required wording. FAIL if: It merely lists tasks, corrections, taste, or judgment without showing what one actual standard would ask, or gives a detail that does not help the reader see how a person's standard applies across models. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.
Q5D07. Tests performance on your workthis case
Judge's reasoning
Both versions endorse the secondhand framing Dan rejected — "A personal benchmark is a reference check instead" — rather than presenting it as trying models on your own work like a work trial.
▸Rubric
D07. Does the paragraph frame a personal benchmark as a direct trial of a model on your work? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “personal benchmarks are more like a work trial than a reference check.” PASS if: It explains trying models on tasks representative of your work and judging their performance against your standards. A work trial, audition, test assignment, or accurate literal explanation can convey this; the phrase 'work trial' is not required. FAIL if: It endorses a reference check or other secondhand report of performance as the explanation of personal benchmarks, even if it also mentions tasks and checks, or never makes clear that the model's own performance on your work is tested. Explicitly rejecting the reference-check analogy is not a failure. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.
Q6D02. Explain the terms the reader needsthis case
Judge's reasoning
Both versions explain the jargon at first use — "that correction can become a case: the task, plus the yes-or-no checks that capture what you were looking for" — so 'case' and 'checks' are defined inline rather than assumed.
▸Rubric
D02. Can a new reader understand the terms and references needed to follow the paragraph when they appear? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “we haven't defined what checks are. So I'm lost.” PASS if: An unfamiliar term essential to the explanation is made understandable at first use, either by ordinary wording or an immediate explanation or example. Normal reader knowledge and context already established in the surrounding document count. FAIL if: The paragraph relies on an unexplained term or reference whose meaning the reader must reconstruct from the model's notes or unstated context. This extends the check beyond ‘checks’ to concepts such as ‘cases.’ It does not require a glossary, explaining familiar words, or repeating material already introduced. Do not fail an ordinary use of the verb ‘check’ when its object is clear. Assess an analogy's explanatory value separately under D04. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.
Q7D06. Claims fit the supplied contextthis case
Judge's reasoning
Claims track the source material — "yes-or-no" checks and running multiple models on a task come from the Rick demo, and "which model is best for your work, whoever makes it, and when a new one is worth switching to" mirrors the draft's own KateBench bullet; the hedged "can become a case" does not assert automatic capture.
▸Rubric
D06. Do the paragraph's claims about how the product works fit the supplied context? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Use the supplied task source material to assess the paragraph's claims. PASS if: Product claims and promises are supported by the source material, with proposed capabilities presented as intentions where that distinction matters. Clearly illustrative examples are allowed. FAIL if: The explanation depends on an invented or contradicted capability, or presents an aspiration as an established current capability. This assesses support, not whether the prose sounds exciting or cautious. Do not demand qualifications that do not change the claim's meaning. An aspiration about one-click capture, for example, should not become a factual claim that the product already captures every correction automatically. Return PASS or FAIL and one sentence of evidence identifying the relevant claim and its support or conflict in the supplied material. For a failure, identify the version and quote the relevant wording.