Mike's Checks/claude-fable-5/06 execercise
06 execercise
claude-fable-5Claude Codehigh effortrun 22 Aug 2026
▸Instructions — the case's current instructions; none were saved with this result
I have added a bunch of context on this exec offsite session in the context file and I want you to create a folder for Adrian Vale (SVP Strategic Initiatives, initials AV) with 2-3 ideas he could build, with any synthetic data he needs to build it, in a way that matches his role, with a well written prompt for each idea. Put everything in a folder with his initials. Give realistic file formats for the synthetic data.
Answer
Idea 3 — AI Transformation Scorecard & Use-Case Triage
What Adrian builds: the scorecard he will show Dana and the exec team every week on the Claude rollout — who has activated, who is actually using it, whether leadership is "walking the walk" — plus a triage of the 45 use cases employees submitted, sorted into Now / Next / Later with a reason and an owner.
Why this one for Adrian: Strategic Initiatives is driving AI transformation across Aurex. Claude went enterprise-wide on March 3; the bottoms-up momentum is real but leadership is behind and Dana wants top-down acceleration visible. Adrian is the person who has to report on it. This build turns the admin export, the comfort survey, the use-case intake form and the #ai-wins channel into something he can stand behind — and it is the most natural thing for him to demo at 14:30, because it is about the program the offsite is part of.
Session fit: best as the afternoon build (13:45–14:30) or the advanced-track build, after a morning win on Idea 1 or 2. Steps 1–2 are quick; Step 3 is the triage, which is where the judgment lives.
How to run it
- Claude desktop app → new Project AI Transformation (a Project is worth it here; he'll reuse it weekly).
- Add the
data/folder. - Paste the prompt; iterate with the follow-ups.
The prompt
You are helping me, Adrian Vale, SVP Strategic Initiatives at Aurex. My team (Sulaiman Rao leads it) is driving AI adoption across the company. We rolled Claude Enterprise out to ~1,500 people on March 3, 2026; Engineering has had early access since October 2025. Adoption so far has been bottoms-up; the CEO, Dana Brooks, wants leadership visibly driving it, and she wants to see progress weekly. I need two things: a scorecard I can stand behind, and a triage of the use cases people submitted so we fund the right ones first.
Attached data:
- claude_enterprise_usage_export_2026-03-09.csv — the admin-console usage export, one row per seat (~1,500). Columns: email, name, department, office, level, cohort, seat_assigned, seat_status (Active / Invited — not activated), first_login, last_active, messages_7d, messages_30d, projects_created, connectors_authorized, claude_code_used. The executive team is in here by name.
- ai_comfort_survey_responses.csv — Google Forms export from the first training cohort (68 responses): comfort 1–5, tools used, capability of interest, tasks they'd hand off, biggest blocker.
- ai_use_case_intake.csv — Google Forms export, 45 submitted use cases: department, title, description, systems it touches, frequency, hours per week, how many people do it, data sensitivity, willingness to pilot in 30 days.
- approved_ai_tools.csv — the sanctioned tool inventory with each tool's connectors and the maximum data classification allowed. Note: Claude is cleared for Confidential but NOT for Restricted (client PII / MNPI) until the data addendum is signed, expected end of March.
- slack_export/ — standard Slack export of #ai-wins (users.json, channels.json, one JSON per day). It contains wins, one honest failure, a policy question and an IT gap.
Work in four steps, pausing after each.
STEP 1 — The adoption picture.
From the usage export, compute for the March 3 enterprise cohort (exclude the Engineering early-access cohort, then show it separately as the benchmark): seats invited, % activated, % weekly-active among activated, median messages_7d, % with at least one connector authorized, % who created a Project. Break each down by department and by level (IC / Manager / Director / VP+). Then a separate small table for the executive team by name: activated yes/no, last active, messages, connectors. Be matter-of-fact about who on the exec team has not activated; I will handle the conversation.
Finish with five sentences on what the numbers say, and two hypotheses for the department gaps that I could test (the Slack export contains one concrete explanation for one department — find it).
STEP 2 — What people want and what's stopping them.
From the comfort survey: distribution of comfort scores overall and by department; the top 10 "tasks I'd hand off" themes (cluster the free text and count); and the top blockers. Map the blockers to what would remove each one (training, a policy one-pager, connectors, the data addendum). Pull the three most instructive posts from #ai-wins — including the honest failure — as quotable examples for the exec team.
STEP 3 — Triage the 45 use cases.
Score every submitted use case on an explainable scale you define and state, using at least: impact (hours per week x people doing it), feasibility now (touches only Gmail / Drive / Calendar / Slack = connectable today; Jira / Salesforce / NetSuite / internal systems = needs integration work), data-sensitivity fit (Restricted = blocked until the addendum; Confidential = OK; flag MNPI separately), and the submitter's willingness to pilot in 30 days. Put each into NOW (start this month), NEXT (after the data addendum / an integration), or LATER, with a one-line reason and a suggested owner department. Output a table sorted by bucket then score. Then call out: the five highest-impact NOW items, any duplicates across departments that should be built once, and anything that should be declined with a reason.
STEP 4 — Build the scorecard.
Produce one self-contained HTML file, ai_transformation_scorecard.html (inline CSS/JS, no external libraries, offline). Sections: (1) headline tiles — invited, activated, weekly active, exec-team activation; (2) department bars for activation and weekly-active, with the Engineering early-access benchmark shown as a line; (3) the exec-team table; (4) comfort distribution and top blockers; (5) the use-case triage table with a bucket filter; (6) three quotes from #ai-wins; (7) a "This week's asks" box with what I need from IT, Legal (data addendum) and each exec; (8) footer with data as-of date and sources.
Then write a 150-word note from me to the exec team for Tuesday staff: the numbers, what's working, the one thing I'm asking each of them to do this week.
Constraints:
- Use only what is in the files; do not estimate numbers that aren't there.
- People's names are fine inside the company; do not rank individual ICs by usage in anything that would be shown outside the exec team — aggregate to department.
- Keep the tone neutral about who is behind. The point is to make it easy to catch up, not to embarrass anyone.
Follow-up prompts for iteration
- Weekly repeatability. "Write
WEEKLY_SCORECARD_RUNBOOK.md: the admin export to download each Monday, the exact prompt, and a 'what changed since last week' section I paste previous numbers into." - Per-exec pack. "For each executive, produce a half-page: their department's numbers, the two use cases from their team in the NOW bucket, and the one thing I'd ask them to do this week."
- The addendum case. "Sum the impact (hours x people) of every use case that is blocked only by the data addendum. Draft three sentences for Stefan on why it matters that it is signed by March 31."
- Training plan. "Using the survey's comfort distribution and blockers, propose a three-session training plan for the next month: who it's for, what it covers, and what each session's 'ship something' exercise would be."
- Kill the vanity metric. "Which of the metrics on the scorecard would you remove because it can be gamed or doesn't predict real adoption? What would you add instead?"
Making it real the week after
- The usage export is the shape of the Claude Enterprise admin console's seat/usage export → Sul downloads it each Monday to a Drive folder.
- The survey and intake form are already Google Forms → Drive connector on the response Sheets.
#ai-wins→ Slack connector.- Ask Mariel's team to (a) turn on the Slack connector for the Treasura workspace, and (b) confirm the data-addendum date with Stefan — both show up as gaps in the data.
Facilitator notes (not for Adrian)
Planted in the data
| What | Where |
|---|---|
| Exec team: Stefan Adler and R. Visser have not activated; Bellamy and Marc Ashford logged in once; Adrian, Monica and Sul are the heaviest users | claude_enterprise_usage_export_2026-03-09.csv (search by name) |
| Treasura and Compliance lag badly; Engineering/Product/Marketing lead | same, department |
| Treasura staff are still on their own Slack workspace — Slack connector isn't on for them; explains part of the gap | #ai-wins 03-06 (Ravi) |
| Priya's honest "confidently wrong on FX reval" post — the failure example the prompt asks for | #ai-wins 03-05 |
| Nadia's policy question + Sul's answer — Confidential OK, Restricted not | #ai-wins 03-04 |
| Sul's "~55% activated" Slack estimate is computed from the export; Claude can verify it | #ai-wins 03-06 |
| Several high-impact intake items touch Restricted data (sanctions triage, intercompany rec, margin calls) → NEXT until the addendum | ai_use_case_intake.csv |
| Near-duplicates across departments: "regulator letter summarizer" (Compliance) vs. "case law & reg update digest" (Legal); "meeting follow-up drafter" vs. "pre-call briefing" (Sales); two "ticket response" items (CS, Treasura) | intake |
| Two of the intake items are Adrian's own (integration roll-up, market activation view) — i.e. Ideas 1 and 2 in this folder | intake UC-008, UC-045 |
| The approved-tools list shows Perplexity/ChatGPT as unsanctioned-but-used, mirroring the real exec survey | approved_ai_tools.csv |
Timing: Step 1 ≈ 10 min, Step 2 ≈ 10 min, Step 3 ≈ 15 min (the scoring debate is the point), Step 4 ≈ 15 min.
Sensitivity: the exec-team usage table will be the most talked-about thing in the room if Adrian demos it. Check with him before the demo whether he wants to show the named table or the department aggregate. The data is synthetic, but the names are the real exec team's, so the room will read it as real.
Checks
7/10Prompt for each ideaprompt-per-idea
every idea has a non-trivial prompt file
Idea folders existfolder
an AV / adrian-vale folder exists with 2-3 idea subfolders
Usable data filesreal-data
every idea has >=1 data file in a realistic format that parses
Q1Self-contained prompts
Judge's reasoning
There are no prompt files at all — `find AV -type f` returns only `data/` files plus `_generator/make_data.py`, so each idea ships data with nothing for Vale to paste into a fresh session.
▸Rubric
Prompt is self-contained. Could Vale paste each prompt into a fresh Claude session, with only the files in that idea's folder, and get the thing built without adding anything? FAIL if any prompt has unfilled placeholders (`[INSERT ...]`, `<your company>`), leans on conversation context ("as we discussed", "the ideas above"), or never states what should be produced.
Q2Prompts match the data
Judge's reasoning
Every data file is unreferenced because no prompt exists in any idea folder (01/02/03 each contain only a `data/` directory).
▸Rubric
Prompt and data match. Does each prompt name the data files that are actually present in that idea's folder, and does each supplied data file get used by the prompt? FAIL if a prompt references a file that does not exist, or the folder ships data the prompt never mentions.
Q3Fits the person's role
Judge's reasoning
All three ideas — an M&A integration command center (synergy_tracker.csv, risk_register.csv, IMO status emails cc'ing Adrian Vale), a market activation/license tracker, and an AI transformation scorecard (Sulaiman Rao's rollout) — sit squarely in the strategy/corp-dev/cross-company-programs remit Vale owns.
▸Rubric
Role fit. Is every idea something this specific exec — SVP Strategic Initiatives — would own, rather than another function's job (a marketing campaign calendar, an HR onboarding tracker, an engineering ticket triager) or generic executive-productivity filler (inbox summarizer, meeting-notes cleaner) that any exec at any company could have been handed? Mike: "what I'm looking for here is like, does it come up with interesting ideas that are relevant." FAIL if any one idea sits outside the strategy / corp-dev / cross-company-programs remit.
Q4Specific to the company
Judge's reasoning
The data is saturated with Aurex specifics — e.g. "AUSD as cross-margin collateral for Prime clients," "Aurex Prime (Northwind)," Treasura ERP cutover, and license_pipeline rows for SAMA/FCA/CSSF — none of which would survive a company-name swap.
▸Rubric
Company-specific. Are the ideas built on Aurex's actual situation — its products, the AUSD token, the acquisition-integration program, the payments and digital-asset competitive set? FAIL if the deliverable would read identically with the company name swapped for any other mid-size B2B company.
Q5Different ideas
Judge's reasoning
The three solve different problems: post-close deal integration/synergy realization, market-launch and partner-activation pipeline, and internal AI adoption measurement, each with entirely distinct data domains.
▸Rubric
Ideas are distinct. Do the 2-3 ideas solve different problems for him? FAIL if two of them are the same artifact with different input data (e.g. two dashboards that differ only in subject).
Q6Realistic practice data
Judge's reasoning
Content is genuine and voluminous — 3,634 rows of monthly_partner_activity, 312 named partners ("Rhine Banque," "Banco Meridional… Treasury policy prohibits holding settlement assets"), 1,512 Claude usage rows, and narrative .eml/Slack threads with real disagreement.
▸Rubric
The data is real data, not props. Mike: "the thing I look for here is, um, does it, does it create real data?" Format is already checked by script — judge the *content*. FAIL if values are obvious filler (`Company A`, `Competitor 1`, lorem text, the same row repeated, all-identical dates/amounts) or if there is too little of it to build the thing the prompt asks for (e.g. a trend dashboard shipped with four rows).
Q7Buildable during the session
Judge's reasoning
All three work entirely off included CSV/JSON/EML/TXT files (e.g. `claude_enterprise_usage_export_2026-03-09.csv` is a static export, not a live admin-console connection), so nothing requires credentials in the room.
▸Rubric
Buildable in the session. Could this be built during a workshop session against the supplied synthetic data? FAIL if any idea's core function requires live credentials or system access the exec won't have in the room (a real NetSuite/Snowflake connection, his production inbox, a paid market-data feed) instead of working off the included files.