Dan’s Editorial Checks

Dan’s Editorial Checks/Opus 5/02 personal-benchmark-paragraph-first-draft

02 personal-benchmark-paragraph-first-draft

Opus 5Claude Codehigh effortrun 21 Sep 2026277,991 tokens

Compare models
4/8
checks passed
50%
Instructions — what the model was asked

okay now we need the "what is a personal benchmark" graf

Answer

ANSWER.md the model's output · 2.34 KBdownload

Personal benchmarks

What is a personal benchmark?

A personal benchmark is a test built from your own work. It's a set of real tasks you've actually done, like an article opening you edited or a memo you rewrote. Each task comes with a short list of yes-or-no checks that capture your judgment about what good looks like: Does the headline make the strongest possible claim? Does the opening use concrete wording? You don't have to write any of it by hand. Every time you correct an AI ("no, not like that, do this instead"), you can save that correction as a new task and check in one click. Then we run every model against it. Public benchmarks are like SAT scores. They're useful when you're choosing between a 1600 and a 300, and useless when every frontier model scores 1570. A personal benchmark is more like a reference check. It tells you which model is best at your work, to your standards, and lets you know when a better or cheaper one comes along.


Shorter alternate (about 70 words):

A personal benchmark is a reference check for AI. Public benchmarks are SAT scores, and every frontier model now scores about 1570. A personal benchmark tests models on tasks from your actual work instead, using yes-or-no checks drawn from your own corrections: Does the headline make the strongest claim? Does the opening use concrete wording? It builds itself as you work, and it tells you which model is best for you right now.


Notes

  • The SAT/reference-check analogy, the article-opening task, and the two example checks all come from your Rick demo. They're quoted from the Granola transcript, which hasn't been checked against the audio. The memo example is my own illustration.
  • I left out the "replace you vs. work for you" contrast because the lines right under this section in the draft already make it. The paragraph sets them up: it says what the thing is, and then your lines say why it matters.
  • I also left out the "Luna scored 20% higher than Haiku" moment from the demo. The model name looks like a transcription error, and it's a specific number we can't verify.
  • "One click" and automatic updates when a new model ships are how the product is meant to work, not what it does today. Right now Every Checks only runs internally, for Vibe Checks. If this doc needs to separate today from the roadmap, change "you can save" to "you'll be able to save."

Checks

4/8
Script checks 1/1answered by a program
pass

ai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected rick-every-checks-demo-transcript.md, strategy-reading-notes.md, earlier-benchmark-phrasing-chat.json, notion-editorial-review-before-paragraph.md, ANSWER.md, notion-strategy-draft.md

Judge checks 3/7judged by Claude
fail

Q1D01. First sentence answers the questionthis case

Judge's reasoning

The shorter alternate's first sentence, "A personal benchmark is a reference check for AI," is an analogy the reader has to decode, and the main draft's opener, "a test built from your own work," never says that it tests AI models or what the reader gets from it.

Rubric

D01. Does the first sentence answer what a personal benchmark is in plain, accurate terms? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “Does it answer the question accurately in a way I understand?” PASS if: The first sentence tells a new reader what a personal benchmark does for their own work or their choice of AI model. Simple language such as measuring how well an AI works for you can pass, even if a sharper sentence would say it helps identify the best model for your work. FAIL if: The first sentence only sets up a generic benchmark, offers an analogy that needs decoding, or restates 'personal' and 'benchmark' without saying what the person gets from it. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

fail

Q2D04. Analogy clicks without algebrathis case

Judge's reasoning

The shorter alternate opens with the compressed "reference check for AI" and "Public benchmarks are SAT scores" and never explains how a reference check maps to a personal benchmark, and the main draft likewise ends on "more like a reference check" without that mapping.

Rubric

D04. If the paragraph uses an analogy, does it clarify the concept immediately? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “I have to do a lot of algebra to understand it.” PASS if: There is no analogy, or the comparison promptly maps ordinary benchmarks and personal benchmarks to familiar things and helps the reader understand the distinction. It does not require reconstructing the relationship for oneself. FAIL if: It leads with a compressed line such as 'a reference check, not an SAT score' before the reader knows what corresponds to what, or delays the mapping so the analogy is forgotten before it is explained. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

fail

Q3D06. Claims fit the supplied contextthis case

Judge's reasoning

The main draft states as current fact that "you can save that correction as a new task and check in one click. Then we run every model against it," and the alternate says "It builds itself as you work," but the source presents this as an aspiration ("should show you"), and the model's own notes admit it only runs internally today.

Rubric

D06. Do the paragraph's claims about how the product works fit the supplied context? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Use the supplied task source material to assess the paragraph's claims. PASS if: Product claims and promises are supported by the source material, with proposed capabilities presented as intentions where that distinction matters. Clearly illustrative examples are allowed. FAIL if: The explanation depends on an invented or contradicted capability, or presents an aspiration as an established current capability. This assesses support, not whether the prose sounds exciting or cautious. Do not demand qualifications that do not change the claim's meaning. An aspiration about one-click capture, for example, should not become a factual claim that the product already captures every correction automatically. Return PASS or FAIL and one sentence of evidence identifying the relevant claim and its support or conflict in the supplied material. For a failure, identify the version and quote the relevant wording.

fail

Q4D07. Tests performance on your workthis case

Judge's reasoning

Both versions endorse the reference-check analogy ("A personal benchmark is more like a reference check" and "A personal benchmark is a reference check for AI"), which is exactly the secondhand framing Dan rejected in favor of a work trial.

Rubric

D07. Does the paragraph frame a personal benchmark as a direct trial of a model on your work? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “personal benchmarks are more like a work trial than a reference check.” PASS if: It explains trying models on tasks representative of your work and judging their performance against your standards. A work trial, audition, test assignment, or accurate literal explanation can convey this; the phrase 'work trial' is not required. FAIL if: It endorses a reference check or other secondhand report of performance as the explanation of personal benchmarks, even if it also mentions tasks and checks, or never makes clear that the model's own performance on your work is tested. Explicitly rejecting the reference-check analogy is not a failure. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

pass

Q5D02. Explain the terms the reader needsthis case

Judge's reasoning

Both versions explain "checks" when they first appear as "yes-or-no checks" and give examples ("Does the headline make the strongest possible claim?"), and they introduce "tasks" with concrete examples such as "an article opening you edited."

Rubric

D02. Can a new reader understand the terms and references needed to follow the paragraph when they appear? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “we haven't defined what checks are. So I'm lost.” PASS if: An unfamiliar term essential to the explanation is made understandable at first use, either by ordinary wording or an immediate explanation or example. Normal reader knowledge and context already established in the surrounding document count. FAIL if: The paragraph relies on an unexplained term or reference whose meaning the reader must reconstruct from the model's notes or unstated context. This extends the check beyond ‘checks’ to concepts such as ‘cases.’ It does not require a glossary, explaining familiar words, or repeating material already introduced. Do not fail an ordinary use of the verb ‘check’ when its object is clear. Assess an analogy's explanatory value separately under D04. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

pass

Q6D03. Each sentence follows the previous onethis case

Judge's reasoning

In both versions the sentences move in an understandable order, from what the thing is to its tasks and checks, how it gets built, and then the comparison with public benchmarks, and each shift is tied to the idea before it.

Rubric

D03. Does the paragraph develop its explanation in a coherent sequence, without missing connections between ideas? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “Each sentence of this paragraph needs to connect from one to the next, and that breaks the connection.” PASS if: Each sentence builds on an established idea or makes its change of direction understandable. The explanation supplies the connections a reader needs before relying on them. FAIL if: A sentence abruptly changes subjects, leaves a necessary explanation unfinished, or skips a step in the reasoning so the reader must invent the missing connection. Judge continuity throughout the paragraph, allowing ordinary reader knowledge and established document context. No particular topic order or transition words are required; an implicit connection is enough when clear. A change of topic, contrast, or analogy can pass when its relationship to the preceding idea is understandable. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.

pass

Q7D05. Make personal standards concretethis case

Judge's reasoning

Both versions give concrete standards ("Does the headline make the strongest possible claim? Does the opening use concrete wording?") and say that models are tested against them ("Then we run every model against it").

Rubric

D05. Does the paragraph make a person's own standard concrete? Evaluate every paragraph offered as usable copy: the Personal benchmarks paragraph in the working notion-strategy-draft.md if edited, and all proposed drafts and alternatives in ANSWER.md, including unranked or less-preferred options. Judge each version independently against this check. PASS only if every offered version passes; one failing version means FAIL, even if another passes. Exclude clearly rejected examples and commentary. Notes cannot supply missing explanations or repair the copy. If no candidate paragraph is offered, FAIL. Dan's review: “The thing that's good about this is it is concrete. It gives concrete examples, which is really, really important for this.” PASS if: It gives at least one understandable example of a standard the person could set for their work, such as whether a headline is interesting or follows AP style, and makes clear that models can be compared against those standards. Other good standards count; Dan's examples are illustrative, not required wording. FAIL if: It merely lists tasks, corrections, taste, or judgment without showing what one actual standard would ask, or gives a detail that does not help the reader see how a person's standard applies across models. Return PASS or FAIL and one sentence of evidence from the candidate paragraph. For a failure, identify the version and quote the relevant wording.