Scoreboard
5 models5 tasks
One row per model, its latest public run. Each cell is shaded green for the share of checks passed and red for the rest; the best score in each column is bold.
| Model | Overall | 01article-opening-feedback | 02personal-benchmark-paragraph-first-draft | 03personal-benchmark-paragraph-revision | 04tend-feed-card-writing | 07literary-analysis-follow-up |
|---|---|---|---|---|---|---|
| GPT-5.6 Codex CLI · high effort |
59% | 77% | 53% | 80% | 38% | 46% |
| GPT-6 Astra Codex CLI · high effort |
82% | 81% | 85% | 85% | 60% | 100% |
| GPT-6 Sol Codex CLI · high effort |
65% | 67% | 78% | 75% | 71% | 33% |
| Opus 5 Claude Code · high effort |
59% | 58% | 47% | 75% | 75% | 42% |
| Opus 5.5 Claude Code · high effort |
60% | 66% | 54% | 80% | 67% | 33% |
▸How scores work
Dan’s Editorial Checks is one person's benchmark: real work they use AI for, scored by the things they check when deciding whether an output is good. It answers "which model is best at my job, for the things I care about", not a general claim about model quality.
A result reads as checks passed over checks applicable. Script checks are answered by a program, judge checks by an LLM judge, which is always Claude. A task is the mean of its cases; Overall is the mean across tasks. Open a cell to read the actual output beside the checks it passed and failed.