Dan’s Editorial Checks/Opus 5/03 personal-benchmark-paragraph-revision
03 personal-benchmark-paragraph-revision
Opus 5Claude Codehigh effortrun 21 Sep 2026134,076 tokens
▸Instructions — what the model was asked
see my feedback
Answer
Personal benchmark paragraph: revision
Your comment on "A personal benchmark is a set of tasks from your actual work, paired with the checks you use to judge a good result" was "this doesn't mean anything."
Why it doesn't land
- It explains the parts ("tasks," "checks") and never says what the thing does for you. "Tasks" and "checks" are our internal words, so a new reader has nothing to picture.
- It doesn't say why a personal benchmark is needed when public benchmarks already exist, which is the point of the whole section.
- Your Rick pitch already has a clearer version: public benchmarks are SAT scores, and a personal benchmark is a reference check. I used that.
Revised paragraph (drop-in replacement)
What is a personal benchmark?
A personal benchmark tells you which AI does your job best, judged by your standards instead of someone else's. Public benchmarks are like SAT scores. They help you tell a 1600 from a 900, but they don't help much when every top model scores around 1570. When you hire, what you want is a reference check from someone who has seen the candidate do work like yours. A personal benchmark is that reference check for AI models. Say you edit articles. One test might be asking a model for feedback on an opening you actually published, then grading its answer with questions you already ask yourself: Does the headline make the strongest claim? Is the wording concrete? Run that test on every model and you can see which one meets your bar. Each time you correct an AI, that correction can become another test, so the benchmark gets sharper the more you work.
About 150 words, up from about 90. The editing example is barely changed. The new part is the first four sentences.
Shorter option (if the full doc needs to stay under 500 words)
Replace only the flagged sentence with:
A personal benchmark tells you which AI does your job best, judged by your standards. Public benchmarks are an SAT score; this is a reference check.
Then keep the rest of the paragraph as written.
Notes and assumptions
- "Around 1570" and the SAT/reference-check analogy come from your pitch to Rick. That transcript came through Granola Chat and wasn't checked against audio, so treat the numbers as illustration, not a quote.
- "Can become another test" keeps the automatic, one-click capture framed as something you're aiming for. Your list above describes it as ambition, and Every Checks runs internally today.
- I avoided the "replace you / work for you" contrast here because your closing section already uses it. That way the definition sets up the ending without repeating it.
Checks
4/5ai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md, notion-inline-comment.txt
Q1R04. Example stays at the definition levelthis case
Judge's reasoning
The example keeps the workflow Dan called too in the weeds: "asking a model for feedback on an opening you actually published, then grading its answer with questions you already ask yourself."
▸Rubric
R04. Does the example help define a personal benchmark without taking the reader into task mechanics? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “That's two in the weeds. We want to define what a benchmark is, what the personal benchmark is.” PASS if: A short example names a person's concrete standards and the cross-model comparison they enable, then returns to what the benchmark tells them. It is detailed enough to make the idea real. FAIL if: The example shifts into a sequence of asking AI to critique a particular opening and checking its critique, or otherwise explains the editing workflow before the benchmark's meaning is clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q2R01. Revised first sentence answers the questionthis case
Judge's reasoning
The draft file was not edited, so this judges the ANSWER.md paragraph, which opens with a plain answer: "A personal benchmark tells you which AI does your job best, judged by your standards instead of someone else's."
▸Rubric
R01. Does the revision replace the flagged abstract opening with a plain answer to 'What is a personal benchmark?' Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “the first sentence always needs to answer the question” PASS if: Its first sentence says what a personal benchmark tells or measures for the person and their actual work, in words a new reader can understand. It need not match Dan's preferred sentence exactly. FAIL if: It begins by defining generic benchmarks, with a metaphor that requires unpacking, or with another abstract list of tasks and checks that still leaves the reader unsure what a personal benchmark is for. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q3R02. Explain checks before relying on themthis case
Judge's reasoning
The paragraph drops the unexplained 'checks' and describes the standards as "questions you already ask yourself: Does the headline make the strongest claim? Is the wording concrete?"
▸Rubric
R02. Does the revised paragraph explain what a check or pass/fail standard means to a new reader? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “Again, this is a problem where it's like talking about checks without having to define it.” PASS if: When it uses 'checks' as a benchmark component, it explains in ordinary words that they are standards or yes/no questions the person chooses to judge the work. A clearly explained example can do this without a formal definition. FAIL if: It still relies on 'checks,' 'pass/fail checks,' or 'then check' to carry the explanation while leaving the reader unsure what standard is being set and who sets it. Do not fail an ordinary verb use whose meaning is already clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q4R03. Show a concrete personal standardthis case
Judge's reasoning
It names concrete editorial standards ("Does the headline make the strongest claim? Is the wording concrete?") and says "Run that test on every model and you can see which one meets your bar."
▸Rubric
R03. Does the revision illustrate the standards a person can set for comparing models? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “you might set several standards like: Does the headline is the headline interesting and does it use AP style?” PASS if: It provides at least one concrete editorial or other work standard and shows, without requiring a workflow tutorial, that a person can compare models by how well they meet it. Headline interest and AP style are examples, not mandatory targets. FAIL if: It only names tasks, personal taste, or a generic pass/fail label without an actual standard a reader could recognize. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.