**What is a personal benchmark?**

A personal benchmark is a test built from your own work. It's a set of tasks you've actually done, each graded by simple yes-or-no checks that capture your judgment. One of mine is giving feedback on the opening of an article we published. The checks include: Does the headline make the strongest possible claim? Does the opening use concrete wording? Run any model on those tasks and you can see which one does your work to your standard. Public benchmarks are like SAT scores. They'll tell a 1600 from a 300, but the top models all score around 1570. A personal benchmark is a reference check. And you don't have to build it by hand. Every time you correct an AI, that correction can become a new case and a new check, so your benchmark grows as you work. When a new model comes out, you'll know whether it's better for you, or just as good for less.

---

**Shorter version (about 85 words), if you need the room:**

A personal benchmark is a test built from your own work: real tasks you've done, graded by yes-or-no checks that capture your judgment, like "Does the headline make the strongest possible claim?" Public benchmarks are SAT scores, and every top model now scores about the same. A personal benchmark is a reference check. It shows which model does *your* work to *your* standard. And it builds itself: every time you correct an AI, that correction becomes a new case and a new check.

---

**Notes**

- **Where it goes:** In the empty blocks under "What is a personal benchmark?" After it, your existing lines ("The industry is built on benchmarks. Benchmarks ask whether AI can replace you. We ask whether it can be put to work for you.") work as the payoff. The SAT line leads into them.
- **Where the material comes from:** It's all from your Rick demo: the SAT-versus-reference-check analogy, the article-opening task and its two checks, correction → case → check, and "better or cheaper" model updates. It adds no new numbers or customer claims. "1570" is your analogy, not a real benchmark score.
- **Now vs. later:** Every Checks with hand-written checks exists today (internally). Turning corrections into checks automatically is the next step, so the long version says a correction "*can* become" one. The short version's "builds itself" states it more strongly. Soften that if the doc needs to be strictly about what exists today.
- **Length:** The draft is about 300 words now. The long version brings it to about 460, still under your 500-word limit. The short version brings it to about 385.
- **Kate:** I used your article-opening example rather than KateBench. KatePass exists, but I couldn't confirm that a KateBench has been built. If you want to name KateBench here, swap in one of Kate's real copy-edit checks.
