# Note to Mike

**Final score: 8/10.** That's from simulated Dan, meaning a fresh subagent following the ask-colleague skill with `personas/dan.md`. On the rubric, 8/10 means "I'd sign off after small edits."

`proposal.md` is the round-10 version. Three rounds scored 8/10 (rounds 7, 8 and 10), and round 10 is the most complete of them. I shipped it exactly as it was rated, so the score applies to that text.

## Why the loop stopped

Five rounds in a row (8–12) failed to beat the best score, which is the stopping rule you set.

- The score went from 3 to 7 in round 2 and then stayed at 7 for four rounds.
- It reached 8 in round 7, once the proposal gave our taste itself a measuring instrument.
- After that it swung between 7 and 8.

Each round, the rater accepted the fixes and then raised new, smaller objections: how a "hit" is defined, the units on the card, reader pushback, what happens to LFG, EC Bench and KateBench.

The rater never hit 10/10, so I never needed to make it stricter. The same rater prompt was used for all 12 rounds. The score range is on the prompt, so the "strictness" is fixed.

My read is that a strict simulated Dan has a ceiling of about 8. More detail in the proposal just gives him more to push on.

## Does it still carry your concern?

**Mostly yes on false precision. Partly traded on the other two parts.**

**Kept:**
- Every number is shown with its range and its count.
- A gap is called "resolved" only when its range excludes zero. Otherwise the card says "not yet resolved."
- The task-count maths is stated openly: about 95 tasks per job to resolve a 10-point gap, and about 250 to claim a gap of around 6 points. Gaps that small aren't claimed until then.
- The first step is an audit of the gaps we've already published. If most turn out to be noise, you bring Dan a stricter threshold before any number goes out.

Your 250 survives as the bar for claiming small gaps.

**Traded:**
- **"No benchmark under 250 tasks."** Dan rejected this in every round. The proposal now publishes numbers below 250, with visible uncertainty. It no longer waits for 250.
- **"Researchers won't take it seriously."** This is now a secondary audience. The proposal leads with the practitioner cost (someone switching tools over a gap that's really noise) and says researcher credibility follows from that.

Simulated Dan pushed back on researchers as the audience in every round. However, the persona has no record at all of Dan on sample sizes or on researchers as an audience; its "Known gaps" section says so. His pushback on those two points is extrapolated from his other views. **Test those two points with the real Dan before you give them up.**

## Things to know before sending

- **Numbers you'd need to check or swap out.** The card numbers are marked illustrative. Several other figures are my estimates, not measured data:
  - the 35-point spread between tasks;
  - about $3 per agent run;
  - the task-pool sizes;
  - Katie and Kieran as rubric owners.
- **Worth folding into round 10.** Rounds 11 and 12 added fixes that didn't raise the score but are worth including before you send. See `rounds/round-11.md` and `rounds/round-12.md`:
  - never label a one-task benchmark like Senior Engineer "resolved";
  - put what the verdict rests on onto the card;
  - define a "hit" before the first launch.

  The first of these matters most for your concern.
- **Open objections if you want to keep iterating.** Round 12 left these, all in `rounds/round-12.md`:
  - Show taste scores as a preference rate with a range (e.g. "preferred in 61% of sessions, 48–73%") instead of "+9."
  - If the judge misses its accuracy bar, publish testers' own ratings rather than going qualitative-only.
  - Classify logged usage per job, and add blind day-90 re-ratings.
