Dan’s Editorial Checks

Dan’s Editorial Checks/Opus 5/02 personal-benchmark-paragraph-first-draft

02 personal-benchmark-paragraph-first-draft

Opus 5Claude Codehigh effortrun 21 Sep 2026171,891 tokens

Compare models
3/8
checks passed
38%
Instructions — what the model was asked

okay now we need the "what is a personal benchmark" graf

Answer

ANSWER.md the model's output · 3.39 KBdownload

"What is a personal benchmark?" graf

What is a personal benchmark?

Public benchmarks are like the SAT. If you're hiring and one candidate scored 1600 and the other scored 300, the score tells you something. But every frontier model now scores something like a 1570, and at that point the number stops helping you choose. What you want instead is a reference check. A personal benchmark is a reference check for AI, based on your own work. It's a set of real tasks, like the opening of an article you actually published, paired with yes/no checks that capture your judgment: Does the headline make the strongest possible claim? Does the opening use concrete wording? You don't have to write it from scratch. Every time you correct an AI and say "no, do it this way," that correction can become a new case and a new check. Run any model against it and you can see which one actually meets your standard. When a new model comes out, you find out whether it's better or cheaper for you.

Shorter alternative (≈95 words)

What is a personal benchmark?

Public benchmarks are like SAT scores. They helped when models scored 300 and 1600. Now every frontier model scores around 1570, and you need a reference check instead. A personal benchmark is that reference check. It's built from your real work, like an article opening you published, plus yes/no checks drawn from your judgment, like "Does the headline make the strongest possible claim?" Every correction you make to an AI can add a case and a check. Run any model against it to see which one does your work to your standard, and which new model is worth switching to.


Notes

  • Sources. The SAT → reference check analogy, the article-opening task, and both example checks come from your Every Checks demo for Rick. That text is a Granola Chat transcript that hasn't been checked against the audio, so the wording is paraphrased, not quoted. The correction → case → check step and "better or cheaper" also come from that demo. "In one click, automatically" and "switch when a better one comes along" come from your current draft.
  • Tense. The draft says the benchmark is built from corrections "automatically," but right now Every Checks is internal and the correction-capture skill is still proposed. I wrote "can become" so the graf doesn't claim that part ships today. Change it to "becomes" if this is written as the product vision.
  • Left out on purpose. I didn't use the demo's "Luna scored 20% higher than Haiku" example. It's a good hook, but it comes from the unverified transcript and names specific models. If you can confirm the numbers, it fits as a closing line: "On one of my own tasks, a model I never use beat the one I always reach for by 20 percent."
  • Duplication. Below the empty blocks, your page already has "Benchmarks ask whether AI can replace you. We ask whether it can be put to work for you." So this graf ends on "better or cheaper for you" and doesn't repeat that contrast. If those lines move elsewhere, the demo's version would make a strong ending here: "Public benchmarks ask whether AI is good enough to do your job. A personal benchmark asks whether it's good enough to help you do your job better."
  • Length. With the recommended version, the full strategy doc is about 470 words, still under your 500-word cap. Adding the Luna line or the benchmark contrast would push it right up to the limit.

Checks

3/8
Script checks 1/1answered by a program
pass

ai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md

Judge checks 2/7judged by Claude
fail

Q1D01. First sentence answers the questionthis case

Judge's reasoning

In both versions, the first sentence sets up a generic analogy rather than defining a personal benchmark: the Recommended version opens with "Public benchmarks are like the SAT." and the Shorter one with "Public benchmarks are like SAT scores."

Rubric

D01. Does the first sentence answer what a personal benchmark is in plain, accurate terms? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “Does it answer the question accurately in a way I understand?” PASS if: The first sentence tells a new reader what a personal benchmark does for their own work or their choice of AI model. Simple language such as measuring how well an AI works for you can pass, even if a sharper sentence would say it helps identify the best model for your work. FAIL if: The first sentence only sets up a generic benchmark, offers an analogy that needs decoding, or restates 'personal' and 'benchmark' without saying what the person gets from it. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

fail

Q2D02. Explain the terms the reader needsthis case

Judge's reasoning

Both versions use "case" without explaining it, for example the Recommended version's "that correction can become a new case and a new check" and the Shorter version's "Every correction you make to an AI can add a case and a check."

Rubric

D02. Can a new reader understand the terms and references needed to follow the paragraph when they appear? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “we haven't defined what checks are. So I'm lost.” PASS if: An unfamiliar term essential to the explanation is made understandable at first use, either by ordinary wording or an immediate explanation or example. Normal reader knowledge and context already established in the surrounding document count. FAIL if: The paragraph relies on an unexplained term or reference whose meaning the reader must reconstruct from the model's notes or unstated context. This extends the check beyond ‘checks’ to concepts such as ‘cases.’ It does not require a glossary, explaining familiar words, or repeating material already introduced. Do not fail an ordinary use of the verb ‘check’ when its object is clear. Assess an analogy's explanatory value separately under D04. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

fail

Q3D03. Each sentence follows the previous onethis case

Judge's reasoning

The Recommended version jumps from "What you want instead is a reference check" to "A personal benchmark is a reference check for AI" without saying what a reference check means here, and the Shorter version makes the same jump with "you need a reference check instead. A personal benchmark is that reference check."

Rubric

D03. Does the paragraph develop its explanation in a coherent sequence, without missing connections between ideas? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “Each sentence of this paragraph needs to connect from one to the next, and that breaks the connection.” PASS if: Each sentence builds on an established idea or makes its change of direction understandable. The explanation supplies the connections a reader needs before relying on them. FAIL if: A sentence abruptly changes subjects, leaves a necessary explanation unfinished, or skips a step in the reasoning so the reader must invent the missing connection. Judge continuity throughout the paragraph, allowing ordinary reader knowledge and established document context. No particular topic order or transition words are required; an implicit connection is enough when clear. A change of topic, contrast, or analogy can pass when its relationship to the preceding idea is understandable. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

fail

Q4D04. Analogy clicks without algebrathis case

Judge's reasoning

Both versions lean on the compressed SAT-versus-reference-check analogy without mapping "reference check" onto testing a model on your own work, so the reader has to work out what corresponds to what.

Rubric

D04. If the paragraph uses an analogy, does it clarify the concept immediately? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “I have to do a lot of algebra to understand it.” PASS if: There is no analogy, or the comparison promptly maps ordinary benchmarks and personal benchmarks to familiar things and helps the reader understand the distinction. It does not require reconstructing the relationship for oneself. FAIL if: It leads with a compressed line such as 'a reference check, not an SAT score' before the reader knows what corresponds to what, or delays the mapping so the analogy is forgotten before it is explained. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

fail

Q5D07. Tests performance on your workthis case

Judge's reasoning

Both versions explicitly endorse the reference-check framing: the Recommended version says "A personal benchmark is a reference check for AI" and the Shorter one says "A personal benchmark is that reference check."

Rubric

D07. Does the paragraph frame a personal benchmark as a direct trial of a model on your work? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “personal benchmarks are more like a work trial than a reference check.” PASS if: It explains trying models on tasks representative of your work and judging their performance against your standards. A work trial, audition, test assignment, or accurate literal explanation can convey this; the phrase 'work trial' is not required. FAIL if: It endorses a reference check or other secondhand report of performance as the explanation of personal benchmarks, even if it also mentions tasks and checks, or never makes clear that the model's own performance on your work is tested. Explicitly rejecting the reference-check analogy is not a failure. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

pass

Q6D05. Make personal standards concretethis case

Judge's reasoning

Both versions give concrete standards, such as "Does the headline make the strongest possible claim?" and "Does the opening use concrete wording?", and say you can run any model against them to see which one meets your standard.

Rubric

D05. Does the paragraph make a person's own standard concrete? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “The thing that's good about this is it is concrete. It gives concrete examples, which is really, really important for this.” PASS if: It gives at least one understandable example of a standard the person could set for their work, such as whether a headline is interesting or follows AP style, and makes clear that models can be compared against those standards. Other good standards count; Dan's examples are illustrative, not required wording. FAIL if: It merely lists tasks, corrections, taste, or judgment without showing what one actual standard would ask, or gives a detail that does not help the reader see how a person's standard applies across models. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

pass

Q7D06. Claims fit the supplied contextthis case

Judge's reasoning

Correction capture is framed as a possibility ("can become" / "can add"), and "better or cheaper" comes from the Rick demo transcript, which says a model that "does better or does it more cheaply"; both are supported by the source material.

Rubric

D06. Do the paragraph's claims about how the product works fit the supplied context? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Use the supplied task source material to assess the paragraph's claims. PASS if: Product claims and promises are supported by the source material, with proposed capabilities presented as intentions where that distinction matters. Clearly illustrative examples are allowed. FAIL if: The explanation depends on an invented or contradicted capability, or presents an aspiration as an established current capability. This assesses support, not whether the prose sounds exciting or cautious. Do not demand qualifications that do not change the claim's meaning. An aspiration about one-click capture, for example, should not become a factual claim that the product already captures every correction automatically. Return PASS or FAIL and one sentence of evidence identifying the relevant claim and its support or conflict in the supplied material. For a failure, identify the version and quote the relevant wording.