A personal benchmark is a test of how well an AI can do your work, judged by your standards. It pairs tasks from your actual work with checks drawn from your judgment and corrections. If you're an editor, a task might be improving an article opening, with checks like: Does the headline make the strongest claim? Does the opening use concrete language? Give different models the same task and score their answers against those checks, and you can see which best meets your standards. As new models come out, you can test whether they do the work better or meet the same bar for less money.
