Scoreboard
11 models15 tasks
One row per model, its latest public run. Each cell is shaded green for the share of checks passed and red for the rest; the best score in each column is bold.
▸How scores work
Mike's Checks is one person's benchmark: real work they use AI for, scored by the things they check when deciding whether an output is good. It answers "which model is best at my job, for the things I care about", not a general claim about model quality.
A result reads as checks passed over checks applicable. Script checks are answered by a program, judge checks by an LLM judge, which is always Claude. A task is the mean of its cases; Overall is the mean across tasks. Open a cell to read the actual output beside the checks it passed and failed.