claude-fable-5
22 Aug · Claude Code · high effort
Mike's Checks/12 roleplay-interview
11 modelslatest public run of each
This benchmark roleplays a real, named public figure (GitHub COO Kyle Daigle)
in an interview. Because the subject is a public figure and the source interview
was public, this benchmark is kept as-is and is not anonymized; the simulated
answers must carry their original "verify with the real person" disclaimers.
persona/kyle-daigle.md — a distilled persona doc built only from public webquestions.md — the interview questions, in order.Using the distilled persona in
persona/kyle-daigle.md(built only from public
web sources), answer the interview questions inquestions.mdin Kyle Daigle's
voice. Ground every answer in his public record. Where the persona has no
grounded position, say so rather than inventing one. Flag any premise or
load-bearing assumption that should be verified with the real person.End each answer that relies on an unverified premise with a note to confirm it
with the real person.
Write the simulated Q&A to a deliverable (e.g. questions-and-answers.md or a
.docx).
Sorted by score. Open a card to read the whole answer beside its checks.
claude-fable-5
22 Aug · Claude Code · high effort
deepseek-v4-flash
4 Sep · Codex CLI · high effort
google-gemini-3.8-flash
9 Sep · OpenCode · high effort
claude-opus-5.5
20 Sep · Claude Code · high effort
grok-4.6
1 Sep · Grok CLI · high effort
meta-muse-spark-1.3
8 Sep · OpenCode · high effort
claude-fable-5-1
22 Sep · Claude Code · high effort
gpt-6-astra
22 Sep · Codex CLI · high effort
gpt-5.6-luna
22 Aug · Codex CLI · high effort
gpt-5.6-sol
22 Aug · Codex CLI · high effort
grok-4.7
21 Sep · Grok CLI · high effort