# "What is a personal benchmark?" graf

## Recommended (≈165 words)

**What is a personal benchmark?**

Public benchmarks are like the SAT. If you're hiring and one candidate scored 1600 and the other scored 300, the score tells you something. But every frontier model now scores something like a 1570, and at that point the number stops helping you choose. What you want instead is a reference check. A personal benchmark is a reference check for AI, based on your own work. It's a set of real tasks, like the opening of an article you actually published, paired with yes/no checks that capture your judgment: Does the headline make the strongest possible claim? Does the opening use concrete wording? You don't have to write it from scratch. Every time you correct an AI and say "no, do it this way," that correction can become a new case and a new check. Run any model against it and you can see which one actually meets your standard. When a new model comes out, you find out whether it's better or cheaper *for you*.

## Shorter alternative (≈95 words)

**What is a personal benchmark?**

Public benchmarks are like SAT scores. They helped when models scored 300 and 1600. Now every frontier model scores around 1570, and you need a reference check instead. A personal benchmark is that reference check. It's built from your real work, like an article opening you published, plus yes/no checks drawn from your judgment, like "Does the headline make the strongest possible claim?" Every correction you make to an AI can add a case and a check. Run any model against it to see which one does your work to your standard, and which new model is worth switching to.

---

## Notes

- **Sources.** The SAT → reference check analogy, the article-opening task, and both example checks come from your Every Checks demo for Rick. That text is a Granola Chat transcript that hasn't been checked against the audio, so the wording is paraphrased, not quoted. The correction → case → check step and "better or cheaper" also come from that demo. "In one click, automatically" and "switch when a better one comes along" come from your current draft.
- **Tense.** The draft says the benchmark is built from corrections "automatically," but right now Every Checks is internal and the correction-capture skill is still proposed. I wrote "can become" so the graf doesn't claim that part ships today. Change it to "becomes" if this is written as the product vision.
- **Left out on purpose.** I didn't use the demo's "Luna scored 20% higher than Haiku" example. It's a good hook, but it comes from the unverified transcript and names specific models. If you can confirm the numbers, it fits as a closing line: *"On one of my own tasks, a model I never use beat the one I always reach for by 20 percent."*
- **Duplication.** Below the empty blocks, your page already has "Benchmarks ask whether AI can replace you. We ask whether it can be put to work for you." So this graf ends on "better or cheaper *for you*" and doesn't repeat that contrast. If those lines move elsewhere, the demo's version would make a strong ending here: *"Public benchmarks ask whether AI is good enough to do your job. A personal benchmark asks whether it's good enough to help you do your job better."*
- **Length.** With the recommended version, the full strategy doc is about 470 words, still under your 500-word cap. Adding the Luna line or the benchmark contrast would push it right up to the limit.
