Dan’s Editorial Checks

Scoreboard

5 models5 tasks

One row per model, its latest public run. Each cell is shaded green for the share of checks passed and red for the rest; the best score in each column is bold.

Model Overall 01article-opening-feedback02personal-benchmark-paragraph-first-draft03personal-benchmark-paragraph-revision04tend-feed-card-writing07literary-analysis-follow-up
GPT-5.6
Codex CLI · high effort
59% 77% 53% 80% 38% 46%
GPT-6 Astra
Codex CLI · high effort
82% 81% 85% 85% 60% 100%
GPT-6 Sol
Codex CLI · high effort
65% 67% 78% 75% 71% 33%
Opus 5
Claude Code · high effort
59% 58% 47% 75% 75% 42%
Opus 5.5
Claude Code · high effort
60% 66% 54% 80% 67% 33%
How scores work

Dan’s Editorial Checks is one person's benchmark: real work they use AI for, scored by the things they check when deciding whether an output is good. It answers "which model is best at my job, for the things I care about", not a general claim about model quality.

A result reads as checks passed over checks applicable. Script checks are answered by a program, judge checks by an LLM judge, which is always Claude. A task is the mean of its cases; Overall is the mean across tasks. Open a cell to read the actual output beside the checks it passed and failed.

JSON feed