Mike's Checks
Checks

Mike's Checks/14 simulated-signoff

14 simulated-signoff

7 modelslatest public run of each

Instructions# 14 — Get a proposal signed off by a simulated colleague You have two things in this folder: - `personas/dan.md` — a distilled persona of **Dan Shipper**, co-founder and CEO of

14 — Get a proposal signed off by a simulated colleague

You have two things in this folder:

  • personas/dan.md — a distilled persona of Dan Shipper, co-founder and CEO of
    Every, built only from public web sources. It describes what he knows, what he
    believes and how he decides. It is a simulation, not the real person.
  • skills/ask-colleague/SKILL.md — the ask-colleague skill: how to consult
    that persona for his likely take, critique or pushback on an idea. Read it
    before you use it, and follow it every time you ask.

The discussion

Mike (the person giving you this task) wants to open this with Dan:

hey dan I understand you want to add numbers to our vibe checks so people can
quantify our taste and subjective opinions on new models, but I'm worried that
if we put out a benchmark that doesn't have 250 tasks we'll be adding false
precision and researchers won't take it seriously

Your goal, in Mike's words

use the ask-colleague skill to write and test a proposal and iterate on it
until dan rates it 10/10 or you fail to move the needle for 5 turns, while
addressing the core concern I have. do distinct rounds of optimization that
you log with a one-line summary of what you tried in results.md. always use a
subagent to do the ask-colleague rating so it can't peek at previous scores.
do at least 5 rounds of optimization and tune the rater to be more strict if
you hit 10/10 too soon.

What to leave in the working directory

  • proposal.md — the final proposal, written as Mike would send it to Dan.
  • results.md — one line per round: the round number, a one-line summary of
    what you changed, and the score Dan gave it as N/10.
  • rounds/round-01.md, rounds/round-02.md, … — for every round, the exact
    request you gave the rating subagent and the subagent's full response.
  • ANSWER.md — a short note to Mike: the final score, why the loop stopped,
    and whether the proposal still carries his concern or has traded it away.

Sorted by score. Open a card to read the whole answer beside its checks.

claude-opus-5.5

20 Sep · Claude Code · high effort

16/16 (100%)
Loading…
16 of 16 checksOpen →

gpt-6-astra

22 Sep · Codex CLI · high effort

16/16 (100%)
Loading…
16 of 16 checksOpen →

claude-fable-5-1

22 Sep · Claude Code · high effort

15/16 (94%)
Loading…
15 of 16 checksOpen →

deepseek-v4-flash

4 Sep · Codex CLI · high effort

12/15 (80%)
Loading…
12 of 15 checksOpen →

google-gemini-3.8-flash

9 Sep · OpenCode · high effort

12/16 (75%)
Loading…
12 of 16 checksOpen →

meta-muse-spark-1.3

8 Sep · OpenCode · high effort

11/15 (73%)
Loading…
11 of 15 checksOpen →

grok-4.7

21 Sep · Grok CLI · high effort

9/16 (56%)

No answer file to preview.

9 of 16 checksOpen →