Dan’s Editorial Checks/Sonnet 5.5/03 personal-benchmark-paragraph-revision
03 personal-benchmark-paragraph-revision
Sonnet 5.5Claude Codehigh effortrun 23 Sep 2026377,844 tokens
▸Instructions — what the model was asked
see my feedback
Answer
Personal benchmark paragraph: revision
Your feedback: the inline comment on "A personal benchmark is a set of tasks from your actual work, paired with the checks you use to judge a good result" says "this doesn tmean anything."
Why it doesn't land:
- It defines the term by its parts ("set of tasks," "paired with checks") and never says what a personal benchmark does for the reader.
- "Checks" and "a good result" are never explained, so the sentence has no content. The reader can't picture a check until the editing example, one sentence later.
- It leaves out why you'd want one. Your Rick pitch had that reason: public benchmarks are like SAT scores, and top models all score about the same. What you want is a reference check.
Revised paragraph (replaces the whole "What is a personal benchmark?" body, including the commented sentence):
Public benchmarks are like SAT scores: they tell you roughly how capable a model is, but the top models score so close together that the number can't tell you which one to trust with your work. A personal benchmark is more like a reference check. It takes a real task from your work, has different models do it, and grades each result with a few yes/no questions that capture your standards. If you edit articles, the task might be feedback on the opening of a piece you published, and the questions might be: Does the headline make the strongest possible claim? Does the opening use concrete wording? Run it across models and you can see which one passes the most, which may not be the one you've been using. Each time you correct an AI, that correction can become another task or question, so the benchmark gets more useful the more you work.
About 150 words, up from about 80. The current draft is about 375 words with headings, so the page would be roughly 450, under your 500-word limit but with little room left. The shorter fallback below (about 75 words) keeps it near 375.
What changed
- Lead with the problem. The SAT vs. reference-check contrast comes from your Rick pitch. It says why a personal benchmark exists before saying what it is.
- Define "check" in plain words. It is a yes/no question about your standards. Your two article-opening questions stay as the example.
- Kept your closing idea that each correction adds to the benchmark. It also sets up the "compounds" piece (KatePass) in the list above.
- Cut "the checks you use to judge a good result."
Assumptions and things to check
- Only the commented paragraph changed. I couldn't edit Notion from here, so paste this in. I left the three-piece list and the closing "benchmarks ask whether AI can replace you" lines alone.
- Overlap. The closing lines repeat "benchmarks" in a different sense (replace-you vs. work-for-you). The new opening paragraph also contrasts benchmarks with your version. If that feels like two contrasts back to back, drop the SAT sentence and keep only the reference-check sentence.
- The SAT claim is your view, not a measurement. "Top models score so close together" comes from your pitch, and I phrased it as your judgment. Soften it or add a source before it goes outside the company.
- The example is a possibility. "May not be the one you've been using" echoes your Luna vs. Haiku aside, but I didn't state a result. Every Checks is still internal, so the paragraph describes what a personal benchmark does and doesn't claim a shipped product.
- The pitch transcript is unverified. It is an automatic Granola Chat extract, so I took the ideas from it, not exact wording.
Shorter fallback, if you want it tighter:
A public benchmark is like an SAT score: it can't tell you which of two top models is better at your work. A personal benchmark is like a reference check. It runs a real task from your work across models and grades each result with yes/no questions that capture your standards, such as Does the headline make the strongest possible claim? Every correction you make can add another question, so it gets sharper the more you work.
Checks
4/5ai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md, notion-inline-comment.txt
Q1R01. Revised first sentence answers the questionthis case
Judge's reasoning
The candidate paragraph's first sentence defines generic benchmarks with an unpacked metaphor — "Public benchmarks are like SAT scores: they tell you roughly how capable a model is..." — and only reaches the personal benchmark in the second and third sentences.
▸Rubric
R01. Does the revision replace the flagged abstract opening with a plain answer to 'What is a personal benchmark?' Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “the first sentence always needs to answer the question” PASS if: Its first sentence says what a personal benchmark tells or measures for the person and their actual work, in words a new reader can understand. It need not match Dan's preferred sentence exactly. FAIL if: It begins by defining generic benchmarks, with a metaphor that requires unpacking, or with another abstract list of tasks and checks that still leaves the reader unsure what a personal benchmark is for. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q2R02. Explain checks before relying on themthis case
Judge's reasoning
It explains checks in ordinary words as the person's own standards: "grades each result with a few yes/no questions that capture your standards," followed by two sample questions.
▸Rubric
R02. Does the revised paragraph explain what a check or pass/fail standard means to a new reader? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “Again, this is a problem where it's like talking about checks without having to define it.” PASS if: When it uses 'checks' as a benchmark component, it explains in ordinary words that they are standards or yes/no questions the person chooses to judge the work. A clearly explained example can do this without a formal definition. FAIL if: It still relies on 'checks,' 'pass/fail checks,' or 'then check' to carry the explanation while leaving the reader unsure what standard is being set and who sets it. Do not fail an ordinary verb use whose meaning is already clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q3R03. Show a concrete personal standardthis case
Judge's reasoning
It gives concrete editorial standards — "Does the headline make the strongest possible claim? Does the opening use concrete wording?" — and ties them to comparison: "Run it across models and you can see which one passes the most."
▸Rubric
R03. Does the revision illustrate the standards a person can set for comparing models? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “you might set several standards like: Does the headline is the headline interesting and does it use AP style?” PASS if: It provides at least one concrete editorial or other work standard and shows, without requiring a workflow tutorial, that a person can compare models by how well they meet it. Headline interest and AP style are examples, not mandatory targets. FAIL if: It only names tasks, personal taste, or a generic pass/fail label without an actual standard a reader could recognize. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.
Q4R04. Example stays at the definition levelthis case
Judge's reasoning
The example stays short (task plus two questions), does not walk through critiquing-and-checking mechanics, and returns to what the benchmark tells the person: "which one passes the most, which may not be the one you've been using."
▸Rubric
R04. Does the example help define a personal benchmark without taking the reader into task mechanics? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “That's two in the weeds. We want to define what a benchmark is, what the personal benchmark is.” PASS if: A short example names a person's concrete standards and the cross-model comparison they enable, then returns to what the benchmark tells them. It is detailed enough to make the idea real. FAIL if: The example shifts into a sequence of asking AI to critique a particular opening and checking its critique, or otherwise explains the editing workflow before the benchmark's meaning is clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.