# Revision: "What is a personal benchmark?"

**Your comment** was on: *"A personal benchmark is a set of tasks from your actual work, paired with the checks you use to judge a good result."* The comment was "this doesn't mean anything."

**Why it falls flat:** It describes what a personal benchmark is made of ("tasks," "checks") before the reader knows what it's for. It also leans on in-house terms. Your pitch to Rick worked because it opened with the problem and an analogy: public benchmarks are SAT scores, every model now scores about the same, and what you want is a reference check. The draft dropped that.

## Recommended replacement paragraph

> Public benchmarks are like SAT scores. They can tell a 1600 from a 300, but every frontier model now scores around 1570, so they can't tell you which one to hire. A personal benchmark is a reference check instead: it tests models on work you've actually done and grades them the way you would. If you edit articles, you might give each model an opening you've published, ask for feedback, and check the answers against the calls you'd make yourself. Did it push the headline toward the strongest claim? Did it ask for more concrete wording? Run that across models and you see which one edits the way you do. Every time you correct a model, the correction can become another test, so the benchmark gets sharper the more you work.

(About 130 words. The old paragraph was about 95, so check it against the 500-word cap for the full doc.)

## Shorter options for the opening sentence only

If you'd rather keep the rest of the paragraph as it is and replace just the flagged sentence:

1. **"A personal benchmark tells you which AI is best at *your* job, not at a standardized test."**
2. **"Public benchmarks are an SAT score for AI; a personal benchmark is a reference check from someone who's seen it do your work."**

I'd go with the full paragraph. The SAT/reference-check contrast is what makes the idea click. It also sets up the "Benchmarks ask whether AI can replace you" lines at the bottom of the page.

## Notes and assumptions

- The SAT/1570/reference-check analogy and the article-opening example (headline claim, concrete wording) come from your Rick pitch transcript. Nothing new was invented. "Around 1570" is your own illustrative figure, not a measured one.
- I changed "Does the headline make the strongest claim?" to "Did it push the headline toward the strongest claim?" because the model's output here is feedback on the opening, so the check is about the feedback.
- I didn't edit the Notion page. This file holds the proposed text for you to paste in or change.
