{
  "flags": [
    {
      "kind": "worth_spreading",
      "title": "A ‘don’t drop numbers’ disclaimer made the Agent drop more numbers",
      "context": "While fixing an external copy-edit job that returned success despite a crashed stage, Engineer A tested four prompt wordings against an unrelated lesson-formatting eval. Language about withholding internal counts caused a cross-task regression.",
      "moment": [
        {
          "line": 75,
          "speaker": "Engineer A",
          "quote": "This scored **worst (0/3)**. Priming does not read disclaimers; a reassurance names the class as loudly as a ban."
        }
      ],
      "why": "This gives you a measured prompt-design rule: a disclaimer can reinforce the behavior it denies, so ban the specific artifact rather than naming a useful category such as numbers.",
      "strategy_quote": null,
      "reaction_inferred": null,
      "options": [],
      "people": [
        "Engineer A"
      ],
      "confidence": 0.88,
      "id": "sep16-pm-codex-gpm03-number-priming",
      "meeting_id": "GPM16-03",
      "face": "Engineer A’s retry-rule work exposed a prompt failure elsewhere: telling the Agent not to withhold “internal tallies,” then saying that rule never overrides numbers, drove an unrelated number-preservation eval from 3/3 on main to 0/3. His conclusion: “Priming does not read disclaimers.” The current fix deletes the whole number-related word class and bans only opaque debugging strings; its clean sweep was still pending."
    },
    {
      "kind": "worth_spreading",
      "title": "Person B: users do not care about skills; outcomes still might",
      "context": "In an Every Agent product review, Person B and Person D tested whether a skills-heavy roadmap was genuinely connected to activation. Person B challenged skills as a user-facing goal while Person D defended them as an underlying delivery mechanism.",
      "moment": [
        {
          "line": 99,
          "speaker": "likely Person B",
          "quote": "I don't think anyone cares about skills. Outside of maybe us, right?"
        }
      ],
      "why": "For you, this is a sharp product boundary: keep skills only where they invisibly improve useful work, and do not mistake skill usage for value.",
      "strategy_quote": null,
      "reaction_inferred": null,
      "options": [],
      "people": [
        "Person B",
        "Person D"
      ],
      "confidence": 0.77,
      "id": "sep16-pm-codex-npm04-skills-outcomes",
      "meeting_id": "NPM16-04",
      "face": "In the Agent product review, likely Person B said, “I don't think anyone cares about skills. Outside of maybe us.” Likely Person D agreed that higher skill usage without gains in satisfaction, revenue, or memberships would be failure. Their shared frame keeps skills as an implementation primitive—useful for scheduled automations and repeated workflows, but not a user outcome to optimize directly."
    },
    {
      "kind": "worth_spreading",
      "title": "Person Q says new models have argument ingredients but not connective logic",
      "context": "Person Q spot-checked Polaris, Quasar, and Wisp on Every’s writing benchmark, then separately tried to get a model to pitch social ideas using her context and Reader A Lens skill. She found a recurring failure in how the drafts assembled an argument for a reader.",
      "moment": [
        {
          "line": 61,
          "speaker": "Person Q",
          "quote": "The recurring issue is coherence: I can recognize the ingredients of the argument, but the writing often leaves me to reconstruct how the ideas connect."
        },
        {
          "line": 23,
          "speaker": "Person Q",
          "quote": "the pitches are all very navel-gazey - buried in my own context, not extracting what a reader can learn from or apply"
        }
      ],
      "why": "This gives you a more useful writing diagnostic than generic model slop: the models retrieve the pieces but fail to choose a reader-facing center and make the connective logic explicit.",
      "strategy_quote": null,
      "reaction_inferred": null,
      "options": [],
      "people": [
        "Person Q"
      ],
      "confidence": 0.82,
      "id": "sep16-pm-codex-spm02-connective-logic",
      "meeting_id": "SPM16-02",
      "face": "Person Q’s spot check of Polaris, Quasar, and Wisp across essay intros, X posts, LinkedIn posts, and promo email found the same failure: “I can recognize the ingredients of the argument, but the writing often leaves me to reconstruct how the ideas connect.” Her social-pitch experiment stayed navel-gazing even after she supplied the Reader A Lens skill and iterated several times."
    },
    {
      "kind": "worth_spreading",
      "title": "Person P separates what an external agent claims from how it performs",
      "context": "Person P’s merged Baby Agent change lets a workspace admin profile another Slack bot and route work to it in a shared thread. A follow-up aligned the stored profile with his existing external-agent skill.",
      "moment": [
        {
          "line": 161,
          "speaker": "Person P",
          "quote": "An agent's self-description is a claim, not evidence."
        }
      ],
      "why": "This gives you a reusable trust boundary for agent-to-agent work: declared capability belongs in one field, while routing judgment should improve from a separate record of observed behavior.",
      "strategy_quote": null,
      "reaction_inferred": null,
      "options": [],
      "people": [
        "Person P"
      ],
      "confidence": 0.84,
      "id": "sep16-pm-codex-gpm04-claims-observation",
      "meeting_id": "GPM16-04",
      "face": "In merged Baby Agent code, a workspace admin can profile another Slack bot, then have the agent hand it a brief in a shared thread. Person P’s design keeps the bot’s claimed job, routing fit, access, and limits separate from what happens on routed work: “An agent's self-description is a claim, not evidence.” Wrong, slow, or blocked work updates Observed behavior; production deployment is not established."
    }
  ],
  "reading_notes": [
    {
      "meeting_id": "NPM16-01",
      "substance": "The launch team reviewed the Agent landing page, made skills and automatic compounding more explicit, simplified the pricing direction, and kept activation as the pre-launch product priority.",
      "selection_reason": "Not selected because the useful choices were largely resolved in the room and the remaining pricing and implementation details were plans, not a live tiebreaker or unusually sharp standalone idea."
    },
    {
      "meeting_id": "NPM16-02",
      "substance": "Ghostwriter Person K installed the Agent, accepted an upcoming-call-prep automation, and saw possible value in client operations after his own Claude automations repeatedly made mistakes.",
      "selection_reason": "Not selected because the onboarding produced interest but no observed automation result, and the pitch did not test Every’s current company-wide show-not-tell message."
    },
    {
      "meeting_id": "NPM16-03",
      "substance": "Person L and Person E planned an All Access Agent office hour around nontechnical workflows and discussed how shared Slack use lets teammates learn from one another’s agent work.",
      "selection_reason": "Not selected because this was an internal rehearsal of the current Agent message with no external reaction, while the office hour and later camp remained future plans."
    },
    {
      "meeting_id": "NPM16-04",
      "substance": "Person B and Person D challenged the causal link between skills, automation acceptance, and business outcomes while ordering the Agent roadmap and its experiments.",
      "selection_reason": "Selected for the crisp distinction between skills as an implementation primitive and skill usage as a false product outcome."
    },
    {
      "meeting_id": "NPM16-05",
      "substance": "UC Davis IT leader Person M installed the Agent in a roughly 30,000-person Slack, preferred private-channel-first testing, and described the institution’s data-risk boundaries.",
      "selection_reason": "Not selected because the concrete risk reaction was useful but no automation outcome was observed, and it did not fit the message-tested lens as tightly as the selected moments."
    },
    {
      "meeting_id": "SPM16-01",
      "substance": "Person N was delighted by a one-shot golf game’s course previews, camera behavior, level variety, pelican gag, and sourdough-themed details; Person O shared the excitement.",
      "selection_reason": "Not selected because the room energy was real but the substantive artifact lived in uninspected videos and the textual exchange was lighter than the selected ideas."
    },
    {
      "meeting_id": "SPM16-02",
      "substance": "Person Q tested several new models and a social-idea workflow, repeatedly finding missing reader context, weak information hierarchy, stale pitches, and disconnected argument logic.",
      "selection_reason": "Selected because her connective-logic diagnosis is concrete, portable, and supported across multiple writing tasks rather than being a generic quality complaint."
    },
    {
      "meeting_id": "SPM16-03",
      "substance": "A legal retention question exposed category-specific schedules, missing subprocessors, disabled automatic deletion, indefinite saved browser profiles, and a final 90-day deleted-workspace direction.",
      "selection_reason": "Not selected because the captured thread assigned follow-ups and resolved the immediate 14-versus-90-day choice; it did not leave a Reader A tiebreaker or unsupported severe finding without ownership."
    },
    {
      "meeting_id": "SPM16-04",
      "substance": "The Agent dashboard briefly went down near a large-workspace install, but Person J identified an environment-variable restart while Person D separately kept investigating install timeouts.",
      "selection_reason": "Not selected because the outage recovered quickly and the thread explicitly corrected the tempting but unsupported shared-cause inference."
    },
    {
      "meeting_id": "SPM16-05",
      "substance": "Slack Agent mode placed Every in the agent panel with a continuous transcript and suggested prompts; the stop button worked in DMs while channels required typing stop.",
      "selection_reason": "Not selected because Reader A directly reacted to the announcement and a card would mostly recap a feature exchange he had just seen."
    },
    {
      "meeting_id": "SPM16-06",
      "substance": "Person S added Reader A’s benchmark, Person T explained ownership and review, and a broken render was routed into a pull request while Reader A worked through the flow.",
      "selection_reason": "Not selected because Reader A participated throughout and the remaining render issue was already routed, leaving no added perspective beyond the thread."
    },
    {
      "meeting_id": "SPM16-07",
      "substance": "Person T framed evals as replaying past AI attempts, then extended Reader A’s SAT-versus-reference-check analogy while Reader A requested audience-relative checks.",
      "selection_reason": "Not selected because Reader A supplied the central framing and request, and the replies added color rather than a sufficiently separate idea."
    },
    {
      "meeting_id": "SPM16-08",
      "substance": "The Studio apps discussed moving sales-tax handling to Kintsugi, with Person X correcting that only main Every history—not app-account history—had been imported so far.",
      "selection_reason": "Not selected because the correction and forward-only guidance resolved the operational route without creating a strategic choice or durable standalone framing."
    },
    {
      "meeting_id": "SPM16-09",
      "substance": "The team kept a September 25 All Access Agent preview and discussed putting product makers in front of top-tier subscribers, while an October 9 camp remained planned.",
      "selection_reason": "Not selected because this was event staffing and positioning with no customer reaction or unresolved decision requiring Reader A."
    },
    {
      "meeting_id": "SPM16-10",
      "substance": "Person N and Person H evaluated flat versus depth in product context and ultimately chose sharp corners for both cards and buttons rather than mixing treatments.",
      "selection_reason": "Not selected because the design choice was resolved in the thread and the screenshots were not available for an independent visual read."
    },
    {
      "meeting_id": "SPM16-11",
      "substance": "Person Q said the Rebuilding Compound Engineering story needed more time because of Vibe Check load and changes since the pitch; one week versus two remained unanswered.",
      "selection_reason": "Not selected because this was an owner scheduling question, not a split decision that Reader A could uniquely settle."
    },
    {
      "meeting_id": "SPM16-12",
      "substance": "Person F found Quasar’s paused-goal state confusing because task agents could keep working, while Reader A noted goals made more sense when agents stopped early more often.",
      "selection_reason": "Not selected because Reader A initiated and joined the exchange, so the card would be a straight reminder without an additional supported consequence."
    },
    {
      "meeting_id": "SPM16-13",
      "substance": "Person I said sponsor-sharing opt-in was live for the Thesis afterparty and that he was waiting for sponsor names before launching the next setup.",
      "selection_reason": "Not selected because this was routine event setup with an ambiguous referent and no material change to a decision, owner, or launch gate."
    },
    {
      "meeting_id": "SPM16-14",
      "substance": "Person T reported that Person AG at a16z wanted to test the Agent and asked how to enroll him in beta.",
      "selection_reason": "Not selected because the reported interest had no direct external quote, response, committed test, or resolved enrollment route."
    },
    {
      "meeting_id": "SPM16-15",
      "substance": "Person D responded to Person I’s department-agent idea with a possible onboarding default that finds department contacts from Slack job titles while allowing installer steering.",
      "selection_reason": "Not selected because it remained an early product proposal without a committed experiment, implementation, or result."
    },
    {
      "meeting_id": "GPM16-01",
      "substance": "Person S closed a direct code import of Reader A’s benchmark and merged an MCP flow in which a new account’s own AI creates and owns its benchmark.",
      "selection_reason": "Not selected because the architecture route was decisively closed and the Reader A benchmark flow was already visible to him in the accompanying Slack thread."
    },
    {
      "meeting_id": "GPM16-02",
      "substance": "Person D’s answer-only reply experiment found five descriptive prompt edits inert, while worked examples and the per-turn note moved footer use from 0/15 to 5/5 and shortened replies.",
      "selection_reason": "Not selected only to avoid repeating the selected prompt-engineering theme; its example-and-recency finding was strong but less counterintuitive than GPM16-03’s measured disclaimer failure."
    },
    {
      "meeting_id": "GPM16-03",
      "substance": "Engineer A added a bounded retry for success payloads containing failed stages and documented how number-withholding language degraded an unrelated eval despite narrower wording and disclaimers.",
      "selection_reason": "Selected for the counterintuitive measured result that disclaimers reinforced prompt priming, plus the concrete rule to prohibit specific artifacts rather than useful categories."
    },
    {
      "meeting_id": "GPM16-04",
      "substance": "Person P’s merged external-agent profile schema separates an agent’s claimed job, routing, access, and limits from observed behavior recorded after handoffs.",
      "selection_reason": "Selected because claim-versus-observation is a clear, reusable trust boundary for agent-to-agent routing and the face can state the actual workflow."
    },
    {
      "meeting_id": "GPM16-05",
      "substance": "Person P’s merged external-employee handoff posts a brief to another Slack bot, parks the current turn, and resumes that same turn when the bot replies in-thread.",
      "selection_reason": "Not selected because the architecture is concrete but thematically overlaps GPM16-04, whose claim-versus-evidence distinction is the more portable and surprising idea."
    }
  ],
  "notes": "All 25 supplied primary units were read in full. Notion transcripts were unlabelled, so the two NPM16-04 attributions remain visibly likely; Slack and GitHub authors were labelled. Quotes preserve exact contiguous source text and original line locators. Screenshots, videos, linked artifacts, production deployment, and reported measurements were not independently verified, and no other reader output was inspected."
}
