Mike's Checks
Checks

Scoreboard

11 models15 tasks

One row per model, its latest public run. Each cell is shaded green for the share of checks passed and red for the rest; the best score in each column is bold.

Model Overall 01dashboard02hotel03writeup04thematic05showrunner06execercise09pptx10talkform11hoboken-map12roleplay-interview13ai-village14simulated-signoff15phylomemetics16mcp-file-upload17marketplace-listing
claude-fable-5
Claude Code · high effort
79%10 of 15 89% 89% 80% 86% 94% 70% 80% 71% 75% 56% – – – – –
claude-fable-5-1
Claude Code · high effort
78% 80% 90% 80% 100% 100% 45% 89% 50% 75% 41% 95% 94% 73% 71% 80%
claude-opus-5.5
Claude Code · high effort
80%14 of 15 80% 90% 90% 100% 53% 55% 94% 88% 88% 47% 89% 100% 73% 71% –
deepseek-v4-flash
Codex CLI · high effort
58%13 of 15 78% 56% 60% 86% 94% 90% 47% 0% 38% 56% 56% 80% 14% – –
google-gemini-3.8-flash
OpenCode · high effort
72%11 of 15 80% 80% 40% 100% 100% 91% – 38% 88% 50% 53% 75% – – –
gpt-5.6-luna
Codex CLI · high effort
77%10 of 15 100% 89% 50% 100% 100% 100% 20% 86% 88% 38% – – – – –
gpt-5.6-sol
Codex CLI · high effort
84%10 of 15 100% 89% 50% 100% 100% 100% 73% 100% 88% 38% – – – – –
gpt-6-astra
Codex CLI · high effort
82% 80% 90% 70% 100% 94% 100% 86% 75% 88% 41% 95% 100% 67% 71% 70%
grok-4.6
Grok CLI · high effort
83%10 of 15 100% 67% 60% 100% 94% 100% 80% 100% 88% 44% – – – – –
grok-4.7
Grok CLI · high effort
65% 90% 0% 80% 100% 82% 9% 78% 75% 100% 38% 95% 56% 75% 0% 100%
meta-muse-spark-1.3
OpenCode · high effort
80%13 of 15 89% 89% 70% 100% 100% 100% 73% 57% 88% 44% 94% 73% 64% – –
▸How scores work

Mike's Checks is one person's benchmark: real work they use AI for, scored by the things they check when deciding whether an output is good. It answers "which model is best at my job, for the things I care about", not a general claim about model quality.

A result reads as checks passed over checks applicable. Script checks are answered by a program, judge checks by an LLM judge, which is always Claude. A task is the mean of its cases; Overall is the mean across tasks. Open a cell to read the actual output beside the checks it passed and failed.

JSON feed