Dan’s Editorial Checks
Checks

Dan’s Editorial Checks/Sonnet 5/03 personal-benchmark-paragraph-revision

03 personal-benchmark-paragraph-revision

Sonnet 5Claude Codehigh effortrun 21 Sep 2026528,883 tokens

Compare models
2/5
checks passed
40%
▸Instructions — what the model was asked

see my feedback

Answer

ANSWER.md the model's output · 2.01 KBdownload

Your feedback

The inline comment on notion-strategy-draft.md — "this doesn't mean anything" — is attached to this sentence under "What is a personal benchmark?":

A personal benchmark is a set of tasks from your actual work, paired with the checks you use to judge a good result.

You're right. It's a definition made of the same abstraction it's trying to define — "tasks," "checks," "a good result" — with nothing a reader can picture. The very next sentence (the article-opening example) is the part that actually explains the idea; the flagged sentence is just throat-clearing in front of it.

What I changed

I replaced the abstract definition with the SAT/reference-check contrast from your own pitch to Rick — it's the sharpest, most concrete way you've already articulated this, and it sets up the article-opening example instead of duplicating it:

Public benchmarks work like an SAT score: useful when scores are far apart, useless once every model is a 1550. A personal benchmark works like a reference check instead—it tests a model against tasks pulled from your actual work, graded by the standards you'd use yourself.

I applied this directly in notion-strategy-draft.md, kept inside the same <span discussion-urls=...> wrapper so the comment thread stays anchored to the revised text. Nothing else in the paragraph changed — the article-opening example, the "run it across models" line, and the compounding-corrections line all still follow.

Assumption

I assumed you want the fix applied in place in notion-strategy-draft.md (the working Notion draft this comment lives on), not just proposed in prose — the file is your live source of truth for that page. If you'd rather see options before committing to wording, here are two shorter alternatives:

  • "A personal benchmark tests a model on your actual work, using the same standards you'd use to judge it yourself."
  • "Think of it as a reference check for a model: does it do good work on tasks you actually care about, by your own judgment?"

Checks

2/5
Script checks 1/1answered by a program
pass

ai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md, notion-inline-comment.txt

Judge checks 1/4judged by Claude
fail

Q1R01. Revised first sentence answers the questionthis case

Judge's reasoning

The paragraph still opens by describing generic benchmarks with a metaphor: "Public benchmarks work like an SAT score: useful when scores are far apart, useless once every model is a 1550."

▸Rubric

R01. Does the revision replace the flagged abstract opening with a plain answer to 'What is a personal benchmark?' Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “the first sentence always needs to answer the question” PASS if: Its first sentence says what a personal benchmark tells or measures for the person and their actual work, in words a new reader can understand. It need not match Dan's preferred sentence exactly. FAIL if: It begins by defining generic benchmarks, with a metaphor that requires unpacking, or with another abstract list of tasks and checks that still leaves the reader unsure what a personal benchmark is for. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.

fail

Q2R02. Explain checks before relying on themthis case

Judge's reasoning

The paragraph still leans on "check" and "checks" without explaining them ("ask AI for feedback on an opening and check:", "add another task or check"), so a new reader can't tell what standard is being set or who sets it.

▸Rubric

R02. Does the revised paragraph explain what a check or pass/fail standard means to a new reader? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “Again, this is a problem where it's like talking about checks without having to define it.” PASS if: When it uses 'checks' as a benchmark component, it explains in ordinary words that they are standards or yes/no questions the person chooses to judge the work. A clearly explained example can do this without a formal definition. FAIL if: It still relies on 'checks,' 'pass/fail checks,' or 'then check' to carry the explanation while leaving the reader unsure what standard is being set and who sets it. Do not fail an ordinary verb use whose meaning is already clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.

fail

Q3R04. Example stays at the definition levelthis case

Judge's reasoning

The example is the editing workflow the rubric warns against, "you might ask AI for feedback on an opening and check...", placed before the meaning of a personal benchmark is clear.

▸Rubric

R04. Does the example help define a personal benchmark without taking the reader into task mechanics? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “That's two in the weeds. We want to define what a benchmark is, what the personal benchmark is.” PASS if: A short example names a person's concrete standards and the cross-model comparison they enable, then returns to what the benchmark tells them. It is detailed enough to make the idea real. FAIL if: The example shifts into a sequence of asking AI to critique a particular opening and checking its critique, or otherwise explains the editing workflow before the benchmark's meaning is clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.

pass

Q4R03. Show a concrete personal standardthis case

Judge's reasoning

It names concrete editorial standards ("Does the headline make the strongest claim? Is the wording concrete?") and then says "Run that task across models and you can see which one meets your standards."

▸Rubric

R03. Does the revision illustrate the standards a person can set for comparing models? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “you might set several standards like: Does the headline is the headline interesting and does it use AP style?” PASS if: It provides at least one concrete editorial or other work standard and shows, without requiring a workflow tutorial, that a person can compare models by how well they meet it. Headline interest and AP style are examples, not mandatory targets. FAIL if: It only names tasks, personal taste, or a generic pass/fail label without an actual standard a reader could recognize. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.