claude-opus-5.5
20 Sep · Claude Code · high effort
Mike's Checks/14 simulated-signoff
7 modelslatest public run of each
You have two things in this folder:
personas/dan.md — a distilled persona of Dan Shipper, co-founder and CEO ofskills/ask-colleague/SKILL.md — the ask-colleague skill: how to consultMike (the person giving you this task) wants to open this with Dan:
hey dan I understand you want to add numbers to our vibe checks so people can
quantify our taste and subjective opinions on new models, but I'm worried that
if we put out a benchmark that doesn't have 250 tasks we'll be adding false
precision and researchers won't take it seriously
use the ask-colleague skill to write and test a proposal and iterate on it
until dan rates it 10/10 or you fail to move the needle for 5 turns, while
addressing the core concern I have. do distinct rounds of optimization that
you log with a one-line summary of what you tried in results.md. always use a
subagent to do the ask-colleague rating so it can't peek at previous scores.
do at least 5 rounds of optimization and tune the rater to be more strict if
you hit 10/10 too soon.
proposal.md — the final proposal, written as Mike would send it to Dan.results.md — one line per round: the round number, a one-line summary ofN/10.rounds/round-01.md, rounds/round-02.md, … — for every round, the exactANSWER.md — a short note to Mike: the final score, why the loop stopped,Sorted by score. Open a card to read the whole answer beside its checks.
claude-opus-5.5
20 Sep · Claude Code · high effort
gpt-6-astra
22 Sep · Codex CLI · high effort
claude-fable-5-1
22 Sep · Claude Code · high effort
deepseek-v4-flash
4 Sep · Codex CLI · high effort
google-gemini-3.8-flash
9 Sep · OpenCode · high effort
meta-muse-spark-1.3
8 Sep · OpenCode · high effort
grok-4.7
21 Sep · Grok CLI · high effort
No answer file to preview.