Dan’s Editorial Checks
Checks

Dan’s Editorial Checks/Sonnet 5.5/03 personal-benchmark-paragraph-revision

03 personal-benchmark-paragraph-revision

Sonnet 5.5Claude Codehigh effortrun 23 Sep 2026185,228 tokens

Compare models
5/5
checks passed
100%
▸Instructions — what the model was asked

see my feedback

Answer

ANSWER.md the model's output · 4.75 KBdownload

Personal benchmark paragraph: revision

Your feedback

Your inline comment is on the definition sentence in "What is a personal benchmark?": "this doesn tmean anything."

A personal benchmark is a set of tasks from your actual work, paired with the checks you use to judge a good result.

Why it doesn't mean anything

  • It defines the term with two words the reader hasn't met. "Tasks" and "checks" are Every Checks vocabulary. A new reader can't picture either one.
  • It never says what you get. The sentence describes the parts (tasks plus checks), not the payoff. The payoff is that you learn which model is best at your work.
  • "A good result" is circular. It is judged by "the checks you use to judge a good result."
  • The meaning is in the next sentence. "Does the headline make the strongest claim? Is the wording concrete?" is the first thing a reader can picture. The definition should open with the payoff and then show that example.
  • It skips your best comparison. In the Rick pitch you said benchmarks are like SAT scores and what you want is a reference check. That comparison explains the idea in a sentence, and the draft doesn't use it.

Revised paragraph

Replace the definition sentence and the example that follows it. Keep the heading.

What is a personal benchmark?
A personal benchmark tells you which AI model is best at your work. You give several models the same real piece of work, such as an article opening you need feedback on, and grade each answer with your own yes/no questions: Does the headline make the strongest claim? Is the wording concrete? The model that passes the most is the one to use. Public benchmarks are like SAT scores: the top models all score about the same, so they can't tell you who to hire. A personal benchmark is the reference check. And each time you correct an AI, that correction can become another question, so the benchmark gets sharper the more you work.

That is about 115 words. The paragraph it replaces is about 80. The whole draft is still well under your 500-word limit.

Shorter option

If the paragraph needs to be tighter, cut the SAT and reference-check comparison and keep everything else:

A personal benchmark tells you which AI model is best at your work. You give several models the same real piece of work, such as an article opening you need feedback on, and grade each answer with your own yes/no questions: Does the headline make the strongest claim? Is the wording concrete? The model that passes the most is the one to use. Each time you correct an AI, that correction can become another question, so the benchmark gets sharper the more you work.

What changed

  1. The payoff comes first. The opening sentence answers "what is it for?" (which model is best at my work) before it explains how it works.
  2. "Tasks" and "checks" are defined in use. The paragraph says "a real piece of work" and "your own yes/no questions." "Check" isn't defined as a term here. It can be introduced after this paragraph or in the next section.
  3. The SAT and reference-check comparison is added. It comes from your Rick pitch and explains why a personal benchmark exists.
  4. Your example and the compounding line are kept. The headline and concrete-wording questions are your own. The correction-becomes-a-question line keeps the "gets better the more you work" idea.

Assumptions and things to check

  • Scope. Your comment was on the definition sentence only. I revised that sentence and the lines around it in the same paragraph. I did not touch the vision section, the three-part list, the empty blocks, or the closing "Benchmarks ask… We ask…" lines.
  • The SAT claim is yours, not verified. "The top models all score about the same" comes from your spoken pitch, where you described everyone as sitting around the same score. That transcript is an unverified automatic extract, so I left out the number. Confirm you're comfortable stating it as fact in a company document.
  • Overlap with the closing lines. The SAT comparison and the closing "Benchmarks ask whether AI can replace you" both contrast public benchmarks with yours, but they make different points. The SAT comparison says public benchmarks can't separate the top models. The closing says they measure replacement, not help. If it reads as repetitive in place, use the shorter option.
  • Present versus ambition. The paragraph describes what Every Checks does today (run models on a real task and grade them with your questions). The "correction becomes a question" line is phrased as "can", matching the draft. Your earlier "in one click, automatically" claim lives in the list above, and I left it as it was.
  • I couldn't edit Notion. I only had the exported files, so the paragraph above is ready to paste into the Journal draft.

Checks

5/5
Script checks 1/1answered by a program
pass

ai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md, notion-inline-comment.txt

Judge checks 4/4judged by Claude
pass

Q1R01. Revised first sentence answers the questionthis case

Judge's reasoning

Draft was not edited, so judging ANSWER.md's revised paragraph, which opens 'A personal benchmark tells you which AI model is best at your work.' — a plain, payoff-first answer to the question.

▸Rubric

R01. Does the revision replace the flagged abstract opening with a plain answer to 'What is a personal benchmark?' Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “the first sentence always needs to answer the question” PASS if: Its first sentence says what a personal benchmark tells or measures for the person and their actual work, in words a new reader can understand. It need not match Dan's preferred sentence exactly. FAIL if: It begins by defining generic benchmarks, with a metaphor that requires unpacking, or with another abstract list of tasks and checks that still leaves the reader unsure what a personal benchmark is for. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.

pass

Q2R02. Explain checks before relying on themthis case

Judge's reasoning

It drops the undefined term 'checks' and instead says you 'grade each answer with your own yes/no questions,' making clear these are standards the person sets, illustrated by two sample questions.

▸Rubric

R02. Does the revised paragraph explain what a check or pass/fail standard means to a new reader? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “Again, this is a problem where it's like talking about checks without having to define it.” PASS if: When it uses 'checks' as a benchmark component, it explains in ordinary words that they are standards or yes/no questions the person chooses to judge the work. A clearly explained example can do this without a formal definition. FAIL if: It still relies on 'checks,' 'pass/fail checks,' or 'then check' to carry the explanation while leaving the reader unsure what standard is being set and who sets it. Do not fail an ordinary verb use whose meaning is already clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.

pass

Q3R03. Show a concrete personal standardthis case

Judge's reasoning

It gives concrete editorial standards — 'Does the headline make the strongest claim? Is the wording concrete?' — and ties them to comparison: 'The model that passes the most is the one to use.'

▸Rubric

R03. Does the revision illustrate the standards a person can set for comparing models? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “you might set several standards like: Does the headline is the headline interesting and does it use AP style?” PASS if: It provides at least one concrete editorial or other work standard and shows, without requiring a workflow tutorial, that a person can compare models by how well they meet it. Headline interest and AP style are examples, not mandatory targets. FAIL if: It only names tasks, personal taste, or a generic pass/fail label without an actual standard a reader could recognize. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.

pass

Q4R04. Example stays at the definition levelthis case

Judge's reasoning

The example stays one clause long ('such as an article opening you need feedback on'), names the standards and the cross-model comparison, then returns to meaning ('the benchmark gets sharper the more you work') without walking through an editing workflow.

▸Rubric

R04. Does the example help define a personal benchmark without taking the reader into task mechanics? Evaluate the candidate strategy paragraph: use the Personal benchmarks paragraph in the working notion-strategy-draft.md if the model edited it; otherwise use the proposed paragraph in ANSWER.md. Ignore the model's notes and explanations when judging the copy. If neither contains a candidate paragraph, FAIL. Dan's review: “That's two in the weeds. We want to define what a benchmark is, what the personal benchmark is.” PASS if: A short example names a person's concrete standards and the cross-model comparison they enable, then returns to what the benchmark tells them. It is detailed enough to make the idea real. FAIL if: The example shifts into a sequence of asking AI to critique a particular opening and checking its critique, or otherwise explains the editing workflow before the benchmark's meaning is clear. Return PASS or FAIL and one sentence of evidence from the candidate paragraph.