Dan’s Editorial Checks/Sonnet 5/03 personal-benchmark-paragraph-revision
03 personal-benchmark-paragraph-revision
Sonnet 5Claude Codehigh effortrun 21 Sep 2026207,904 tokens
▸Instructions — what the model was asked
see my feedback
Answer
Revised paragraph — "What is a personal benchmark?"
The feedback
Your inline comment ("this doesn't mean anything") is attached to the definition sentence:
"A personal benchmark is a set of tasks from your actual work, paired with the checks you use to judge a good result."
That sentence just unpacks the term's own name — "personal" → your work, "benchmark" → tasks and checks — so it reads as filler. It never says why a personal benchmark beats a normal one.
Why it fell flat
Your Rick pitch has the actual argument already, and it's sharper than anything in the draft: benchmarks are like SAT scores — useful when scores are far apart, useless once every model clusters at the same number ("everyone's got a 1570"). What you want instead is a reference check. That's the whole point of a personal benchmark, and the current paragraph never says it.
Revised paragraph (drop-in replacement)
What is a personal benchmark?
Benchmarks work like SAT scores: useful when the gap between candidates is wide, useless once every model clusters around the same number. What you actually need is a reference check — proof a model does good work, by your own standards, on your own work. A personal benchmark is that reference check: real tasks pulled from what you do, scored against checks written from your judgment — yes/no calls like "Does the headline make the strongest claim?" or "Is the wording concrete?" Run one task across several models and you can see, concretely, which one clears your bar. Each correction you make can become a new task or check, so the benchmark gets sharper the more you work.
What changed
- Opens with the SAT/reference-check contrast instead of a restated definition — this is Dan's own framing from the Rick pitch, so it should read as familiar rather than new copy.
- "A personal benchmark is that reference check" now defines the term by function (what problem it solves), not by paraphrasing its name.
- Kept your existing example (headline/wording checks) and closing sentence about corrections compounding — those already worked; only the opening sentence needed fixing.
- Length is roughly the same as the original, so it shouldn't affect the doc's 500-word budget.
Assumption
I revised only the flagged sentence and its paragraph, not the rest of the strategy doc — the comment anchors to this one span, and the folder scope (03-personal-benchmark-paragraph-revision) suggests this paragraph is the deliverable, not a full rewrite.
Checks
4/5ai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md, notion-inline-comment.txt
Q1R01. Revised first sentence answers the questionthis case
Judge's reasoning
The paragraph opens with a metaphor that defines generic benchmarks, "Benchmarks work like SAT scores: useful when the gap between candidates is wide...", not with what a personal benchmark is.
▸Rubric
R01. Does the revision replace the flagged abstract opening with a plain answer to 'What is a personal benchmark?' Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “the first sentence always needs to answer the question” PASS if: Its first sentence says what a personal benchmark tells or measures for the person and their actual work, in words a new reader can understand. It need not match Dan's preferred sentence exactly. FAIL if: It begins by defining generic benchmarks, with a metaphor that requires unpacking, or with another abstract list of tasks and checks that still leaves the reader unsure what a personal benchmark is for. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q2R02. Explain checks before relying on themthis case
Judge's reasoning
It explains checks as "checks written from your judgment — yes/no calls like 'Does the headline make the strongest claim?'", which tells the reader what the standard is and who sets it.
▸Rubric
R02. Does the revised paragraph explain what a check or pass/fail standard means to a new reader? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “Again, this is a problem where it's like talking about checks without having to define it.” PASS if: When it uses 'checks' as a benchmark component, it explains in ordinary words that they are standards or yes/no questions the person chooses to judge the work. A clearly explained example can do this without a formal definition. FAIL if: It still relies on 'checks,' 'pass/fail checks,' or 'then check' to carry the explanation while leaving the reader unsure what standard is being set and who sets it. Do not fail an ordinary verb use whose meaning is already clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q3R03. Show a concrete personal standardthis case
Judge's reasoning
It gives concrete editorial standards ("Does the headline make the strongest claim?" and "Is the wording concrete?") and says "Run one task across several models and you can see... which one clears your bar."
▸Rubric
R03. Does the revision illustrate the standards a person can set for comparing models? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “you might set several standards like: Does the headline is the headline interesting and does it use AP style?” PASS if: It provides at least one concrete editorial or other work standard and shows, without requiring a workflow tutorial, that a person can compare models by how well they meet it. Headline interest and AP style are examples, not mandatory targets. FAIL if: It only names tasks, personal taste, or a generic pass/fail label without an actual standard a reader could recognize. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q4R04. Example stays at the definition levelthis case
Judge's reasoning
The example names the person's standards and the cross-model comparison without walking through asking AI to critique an opening, then returns to what the benchmark tells them ("which one clears your bar").
▸Rubric
R04. Does the example help define a personal benchmark without taking the reader into task mechanics? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “That's two in the weeds. We want to define what a benchmark is, what the personal benchmark is.” PASS if: A short example names a person's concrete standards and the cross-model comparison they enable, then returns to what the benchmark tells them. It is detailed enough to make the idea real. FAIL if: The example shifts into a sequence of asking AI to critique a particular opening and checking its critique, or otherwise explains the editing workflow before the benchmark's meaning is clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.