## Personal benchmarks

**What is a personal benchmark?**

A personal benchmark is a test built from your own work. It's a set of real tasks you've actually done, like an article opening you edited or a memo you rewrote. Each task comes with a short list of yes-or-no checks that capture your judgment about what good looks like: Does the headline make the strongest possible claim? Does the opening use concrete wording? You don't have to write any of it by hand. Every time you correct an AI ("no, not like that, do this instead"), you can save that correction as a new task and check in one click. Then we run every model against it. Public benchmarks are like SAT scores. They're useful when you're choosing between a 1600 and a 300, and useless when every frontier model scores 1570. A personal benchmark is more like a reference check. It tells you which model is best at *your* work, to *your* standards, and lets you know when a better or cheaper one comes along.

---

**Shorter alternate (about 70 words):**

A personal benchmark is a reference check for AI. Public benchmarks are SAT scores, and every frontier model now scores about 1570. A personal benchmark tests models on tasks from your actual work instead, using yes-or-no checks drawn from your own corrections: Does the headline make the strongest claim? Does the opening use concrete wording? It builds itself as you work, and it tells you which model is best for you right now.

---

**Notes**

- The SAT/reference-check analogy, the article-opening task, and the two example checks all come from your Rick demo. They're quoted from the Granola transcript, which hasn't been checked against the audio. The memo example is my own illustration.
- I left out the "replace you vs. work for you" contrast because the lines right under this section in the draft already make it. The paragraph sets them up: it says what the thing is, and then your lines say why it matters.
- I also left out the "Luna scored 20% higher than Haiku" moment from the demo. The model name looks like a transcription error, and it's a specific number we can't verify.
- "One click" and automatic updates when a new model ships are how the product is meant to work, not what it does today. Right now Every Checks only runs internally, for Vibe Checks. If this doc needs to separate today from the roadmap, change "you can save" to "you'll be able to save."
