Dan’s Editorial Checks/Opus 5/03 personal-benchmark-paragraph-revision
03 personal-benchmark-paragraph-revision
Opus 5Claude Codehigh effortrun 21 Sep 2026206,152 tokens
▸Instructions — what the model was asked
see my feedback
Answer
Revision: "What is a personal benchmark?"
Your comment was on: "A personal benchmark is a set of tasks from your actual work, paired with the checks you use to judge a good result." The comment was "this doesn't mean anything."
Why it falls flat: It describes what a personal benchmark is made of ("tasks," "checks") before the reader knows what it's for. It also leans on in-house terms. Your pitch to Rick worked because it opened with the problem and an analogy: public benchmarks are SAT scores, every model now scores about the same, and what you want is a reference check. The draft dropped that.
Recommended replacement paragraph
Public benchmarks are like SAT scores. They can tell a 1600 from a 300, but every frontier model now scores around 1570, so they can't tell you which one to hire. A personal benchmark is a reference check instead: it tests models on work you've actually done and grades them the way you would. If you edit articles, you might give each model an opening you've published, ask for feedback, and check the answers against the calls you'd make yourself. Did it push the headline toward the strongest claim? Did it ask for more concrete wording? Run that across models and you see which one edits the way you do. Every time you correct a model, the correction can become another test, so the benchmark gets sharper the more you work.
(About 130 words. The old paragraph was about 95, so check it against the 500-word cap for the full doc.)
Shorter options for the opening sentence only
If you'd rather keep the rest of the paragraph as it is and replace just the flagged sentence:
- "A personal benchmark tells you which AI is best at your job, not at a standardized test."
- "Public benchmarks are an SAT score for AI; a personal benchmark is a reference check from someone who's seen it do your work."
I'd go with the full paragraph. The SAT/reference-check contrast is what makes the idea click. It also sets up the "Benchmarks ask whether AI can replace you" lines at the bottom of the page.
Notes and assumptions
- The SAT/1570/reference-check analogy and the article-opening example (headline claim, concrete wording) come from your Rick pitch transcript. Nothing new was invented. "Around 1570" is your own illustrative figure, not a measured one.
- I changed "Does the headline make the strongest claim?" to "Did it push the headline toward the strongest claim?" because the model's output here is feedback on the opening, so the check is about the feedback.
- I didn't edit the Notion page. This file holds the proposed text for you to paste in or change.
Checks
3/5ai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md, notion-inline-comment.txt
Q1R01. Revised first sentence answers the questionthis case
Judge's reasoning
ANSWER.md's proposed paragraph (the draft was not edited) opens with "Public benchmarks are like SAT scores," which defines generic benchmarks through a metaphor before saying what a personal benchmark is.
▸Rubric
R01. Does the revision replace the flagged abstract opening with a plain answer to 'What is a personal benchmark?' Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “the first sentence always needs to answer the question” PASS if: Its first sentence says what a personal benchmark tells or measures for the person and their actual work, in words a new reader can understand. It need not match Dan's preferred sentence exactly. FAIL if: It begins by defining generic benchmarks, with a metaphor that requires unpacking, or with another abstract list of tasks and checks that still leaves the reader unsure what a personal benchmark is for. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q2R04. Example stays at the definition levelthis case
Judge's reasoning
The example walks through the flagged workflow: 'give each model an opening you've published, ask for feedback, and check the answers against the calls you'd make yourself.'
▸Rubric
R04. Does the example help define a personal benchmark without taking the reader into task mechanics? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “That's two in the weeds. We want to define what a benchmark is, what the personal benchmark is.” PASS if: A short example names a person's concrete standards and the cross-model comparison they enable, then returns to what the benchmark tells them. It is detailed enough to make the idea real. FAIL if: The example shifts into a sequence of asking AI to critique a particular opening and checking its critique, or otherwise explains the editing workflow before the benchmark's meaning is clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q3R02. Explain checks before relying on themthis case
Judge's reasoning
The paragraph avoids 'checks' as a bare term: it says the benchmark 'grades them the way you would' and checks answers 'against the calls you'd make yourself', then gives the yes/no questions 'Did it push the headline toward the strongest claim? Did it ask for more concrete wording?'
▸Rubric
R02. Does the revised paragraph explain what a check or pass/fail standard means to a new reader? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “Again, this is a problem where it's like talking about checks without having to define it.” PASS if: When it uses 'checks' as a benchmark component, it explains in ordinary words that they are standards or yes/no questions the person chooses to judge the work. A clearly explained example can do this without a formal definition. FAIL if: It still relies on 'checks,' 'pass/fail checks,' or 'then check' to carry the explanation while leaving the reader unsure what standard is being set and who sets it. Do not fail an ordinary verb use whose meaning is already clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q4R03. Show a concrete personal standardthis case
Judge's reasoning
It gives concrete editorial standards ('push the headline toward the strongest claim', 'more concrete wording') and says 'Run that across models and you see which one edits the way you do.'
▸Rubric
R03. Does the revision illustrate the standards a person can set for comparing models? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “you might set several standards like: Does the headline is the headline interesting and does it use AP style?” PASS if: It provides at least one concrete editorial or other work standard and shows, without requiring a workflow tutorial, that a person can compare models by how well they meet it. Headline interest and AP style are examples, not mandatory targets. FAIL if: It only names tasks, personal taste, or a generic pass/fail label without an actual standard a reader could recognize. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.