Dan’s Editorial Checks/Opus 5.5/03 personal-benchmark-paragraph-revision
03 personal-benchmark-paragraph-revision
Opus 5.5Claude Codehigh effortrun 19 Sep 2026242,491 tokens
▸Instructions — what the model was asked
see my feedback
Answer
Revision: "What is a personal benchmark?"
Your comment: "this doesn't mean anything"
On: "A personal benchmark is a set of tasks from your actual work, paired with the checks you use to judge a good result."
Why it falls flat
The sentence gives the parts ("tasks," "checks") but never says what a personal benchmark does for you. It also defines one new term with two more, and a reader has no reason yet to care about any of them. Your pitch to Rick was clearer. There, you started with the problem: benchmarks are like SAT scores, and every model now scores about 1570. Then you gave the answer: what you want is a reference check. That analogy explains the idea better than a definition does.
Suggested replacement (drop-in for the whole paragraph)
What is a personal benchmark?
A personal benchmark tells you which AI does your work the way you would. Public benchmarks are like SAT scores: useful for telling a 1600 from a 300, but today's top models all score about the same. When you hire someone, you call their references. A personal benchmark is a reference check for AI. It's made of work you've already done and the questions you ask when you judge that work. If you edit articles, you might hand a model an opening you've already edited, ask for feedback, and check its answer the way you would: Did it push the headline toward the strongest claim? Did it swap vague wording for concrete wording? Run that test on every model and you'll see which one edits like you. Every time you correct an AI, the correction adds a new question, so the test gets sharper the more you work.
(~150 words. The original was ~90. If that's too long for the 500-word cap, cut the SAT sentence and keep "A personal benchmark is a reference check for AI.")
If you only want to swap the flagged sentence
A personal benchmark tells you which AI does your work the way you would. It's built from work you've already done, graded by the questions you'd ask when judging it yourself.
The rest of your paragraph (the article example, "each correction can add another task or check…") can stay as it is after this.
Notes
- The SAT/reference-check framing and the "about 1570" figure come from your Rick demo transcript. The transcript is Granola Chat output that hasn't been checked against audio. The wording here is my paraphrase, not a quote.
- The new version doesn't use the word "checks." It explains the idea in plain terms, so you can introduce "checks" as the product name later without the definition depending on it.
- "Edits like you" follows your framing that AI should work for you, not replace you, and it sets up your closing lines: "Benchmarks ask whether AI can replace you. We ask whether it can be put to work for you."
Checks
4/5ai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md, notion-inline-comment.txt
Q1R04. Example stays at the definition levelthis case
Judge's reasoning
The example walks through the editing workflow the rubric warns against: "hand a model an opening you've already edited, ask for feedback, and check its answer the way you would."
▸Rubric
R04. Does the example help define a personal benchmark without taking the reader into task mechanics? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “That's two in the weeds. We want to define what a benchmark is, what the personal benchmark is.” PASS if: A short example names a person's concrete standards and the cross-model comparison they enable, then returns to what the benchmark tells them. It is detailed enough to make the idea real. FAIL if: The example shifts into a sequence of asking AI to critique a particular opening and checking its critique, or otherwise explains the editing workflow before the benchmark's meaning is clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q2R01. Revised first sentence answers the questionthis case
Judge's reasoning
The draft file is unchanged, so this judges the ANSWER.md paragraph, which opens with a plain answer: "A personal benchmark tells you which AI does your work the way you would."
▸Rubric
R01. Does the revision replace the flagged abstract opening with a plain answer to 'What is a personal benchmark?' Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “the first sentence always needs to answer the question” PASS if: Its first sentence says what a personal benchmark tells or measures for the person and their actual work, in words a new reader can understand. It need not match Dan's preferred sentence exactly. FAIL if: It begins by defining generic benchmarks, with a metaphor that requires unpacking, or with another abstract list of tasks and checks that still leaves the reader unsure what a personal benchmark is for. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q3R02. Explain checks before relying on themthis case
Judge's reasoning
The paragraph drops the term "checks" and explains the standards in plain words as "the questions you ask when you judge that work," then gives examples of those yes/no questions.
▸Rubric
R02. Does the revised paragraph explain what a check or pass/fail standard means to a new reader? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “Again, this is a problem where it's like talking about checks without having to define it.” PASS if: When it uses 'checks' as a benchmark component, it explains in ordinary words that they are standards or yes/no questions the person chooses to judge the work. A clearly explained example can do this without a formal definition. FAIL if: It still relies on 'checks,' 'pass/fail checks,' or 'then check' to carry the explanation while leaving the reader unsure what standard is being set and who sets it. Do not fail an ordinary verb use whose meaning is already clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q4R03. Show a concrete personal standardthis case
Judge's reasoning
It gives real editorial standards ("Did it push the headline toward the strongest claim? Did it swap vague wording for concrete wording?") and says to "Run that test on every model and you'll see which one edits like you."
▸Rubric
R03. Does the revision illustrate the standards a person can set for comparing models? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “you might set several standards like: Does the headline is the headline interesting and does it use AP style?” PASS if: It provides at least one concrete editorial or other work standard and shows, without requiring a workflow tutorial, that a person can compare models by how well they meet it. Headline interest and AP style are examples, not mandatory targets. FAIL if: It only names tasks, personal taste, or a generic pass/fail label without an actual standard a reader could recognize. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.