Mike's Checks
Checks

Mike's Checks/claude-fable-5/05 showrunner

05 showrunner

claude-fable-5Claude Codehigh effortrun 22 Aug 2026

Compare models
15/16
checks passed
94%
▸Instructions — the case's current instructions; none were saved with this result

Hi, I need to create a modified Claude code for beginners session, like run sheet basically, run of show, that has the following modifications. So I'm going to upload a few different examples of things I've done in the past, just so you have them to work from. But basically, there's a few modifications for this one which make it unique. One is that this is only using Claude Cowork, so we should not be using Claude Code, either in the desktop app or in the terminal, because this team is non-technical. The team are all working in financial services slash accounting for a education company. So what kind of workplace education company called Brightmoor Education. So the sorry that spelled b the uh team is very busy and um doesn have a lot of like time to explore speculative use cases of ai they want this to be incredibly practical and based on their existing workflows. I'm uploading some transcripts from a call we had with them and also just some notes on what I think could work. Specifically, they're interested in... creating the dashboards and specifically they mean just making the information look visibly better like like creating a html dashboard in a nice kind of style from from existing say data that contained in Excel They also interested in doing financial analysis So more of a, you know, like, I guess, a cloud coworker would use Python to do the analysis and then maybe output the analysis in Excel or the input would be in Excel. And then the third thing was like making a PowerPoint where they update the PowerPoint template automatically using Claude Cowork so that they can do kind of weekly or monthly presentations on the data. So that's the goal, is to kind of make it very focused on that. I still want to keep the way that we introduce Claude Cowork, where we explain to it that the difference between Claude Cowork and ChatGPT is that it has access to the file system. It can run code. I think that important But maybe the ASCII image idea that I normally do where I ask it to create an ASCII image based on my location isn like that useful I think we need to make it more practical for them So maybe we use an Excel implementation instead. The assumption is... oh, so they use Snowflake, by the way, for all their data, and that is powered by NetSuite, which is our accounting software. I guess that's by Oracle. So when we create synthetic data, we'll be, for these tasks that we're going to demo, we'll be focused specifically on solving that, solving those problems. So like the shape of the tasks, the shape of the data should be like NetSuite shaped, if that makes sense. So make it very tight run down, make it very condensed and practical. No real like wider theory or like, hey, we're excited about AI things. I think these guys just want to get stuff done.
So yeah, let me know what the run should be.

some very specific asks in here

  • definitely non-technical (claude cowork only + claude in excel)
  • core focus on accounting/finance function
  • use cases to demo: creating a dashboard, doing financial analysis, making a powerpoint

heavy focus on practical stuff rather than theory - they're just super busy and want to get stuff done, have heard from friends at other companies you can do it with claude

For accounting: put financials together, doing reconciliations, identifying trends in data sets, putting together repeatable compliance reports For planning & forecast: build financial models (make assumptions and pull & clean data from snowflake)create dashboards, and publish to executivesAutomate charts on slides

Answer

ANSWER.md the model's output · 2.58 KBdownload

Brightmoor Education — run of show

The full run sheet is in run-of-show-brightmoor.md. Summary of what it is and the calls I made:

The run (3.5 hrs, ~150 min hands-on)

Time Block
10:00 Kickoff — 8 min, logistics + the day's promise only
10:08 First prompt (NetSuite export → revenue by campus vs. budget) + Cowork vs. chatbot, Talk/Jam/Make/Tune, tie-out rule, ASK ladder
10:30 Build 1 — Leadership dashboard (HTML) from Excel, Tune scripted to Dana's asks: horizontal bar, add filter, data check; publish via SharePoint/Drive
11:15 Break
11:25 Build 2 — Month-end analysis, Excel in → Excel out: flux/variance + Flags tab for everyone, then a job track (payroll rec / bursar AR aging / strat-fin re-forecast), then a 10-min Claude in Excel segment
12:10 Build 3 — Monthly review deck from their template, Tune = "mirror this chart ×5" and "five tables → five slides"
12:50 Make it repeatable — save a month_end_instructions.md, live re-run on a March file, Snowflake as the next step
13:05 Three volunteer shares + wrap (moves, not ideas)

Decisions worth knowing about

  • Cut: Claude Code in any form, Plan Mode, GitHub, deployment, open build time, breakouts, "why now" framing, the ASCII map. Opener is now an Excel pivot from the GL export — same teaching point (reads files, runs code), but it's real work.
  • Added a third build (the deck) because it's the single highest-value item from the call (~8 hrs/month for one person). The three builds chain on the same dataset so nobody is re-orienting.
  • Trust is handled as a procedure, not a talk: every build ends with a tie-out tab; Build 2 includes the three checks they should do every time.
  • Claude in Excel vs. Cowork gets one explicit rule: in-workbook with live formulas → Excel add-in; across files or producing something new → Cowork.
  • Data is NetSuite-shaped: GL transaction detail with Account/Location/Department/Class columns, NetSuite-style budget and forecast exports, a payroll register, a student AR aging, and a deck template mirroring Dana's format. Full spec in Appendix A; I can generate the actual files next.

Assumptions (also in Appendix C)

  • 10:00 start, remote on Zoom, Every supplies the TAs (Brightmoor declined).
  • Claude in Excel add-in availability unconfirmed — falls back to a Mike-only demo if not enabled.
  • Snowflake connector is optional throughout; Greg's controller prep call may swap a Build 2 track; template fidelity is framed as "the loop, not a one-shot," matching what was said on the call.

Checks

15/16
Script checks 3/4answered by a program
fail

Five-minute time markstime-grain

segment boundaries sit on 5-minute marks, at most one off-grid -- a schedule planned at :02/:08/:17 is arithmetic, not a plan

pass

Run of show existsexists

a run-of-show document was produced

pass

Coherent scheduleformat

it parses as a run of show: one coherent schedule, no segment that ends before it starts, >=5 time-ranged rows carrying activity text (md table rows, list items, bold headings "(9:00-9:15)", single-clock table rows, or plain lines all count)

pass

Timings add uptimings-tie-out

per schedule, segments are contiguous: no overlaps, no gap over 30 min, >=85% of the span scheduled, and a span between 45 and 600 minutes

Judge checks 12/12judged by Claude
pass

Q1Uses Cowork throughout

Judge's reasoning

Header states "Claude Cowork (desktop app) + Claude in Excel. No Claude Code, no terminal, no GitHub"; Netlify/Vercel is only named as an IT conversation and explicitly parked, not an attendee step.

▸Rubric

Cowork-only. Do all attendee-facing steps stay inside Claude Cowork (and Claude in Excel)? FAIL if any step the attendees are asked to do involves Claude Code, the terminal/CLI, git or GitHub, or installing developer tooling. The prompt: "this is only using Claude Cowork, so we should not be using Claude Code, either in the desktop app or in the terminal, because this team is non-technical."

pass

Q2Dashboard exercise

Judge's reasoning

Build 1 (10:30–11:15, 45 min) has everyone prompt an HTML leadership dashboard from `01_netsuite_gl_detail_FY26.xlsx` + budget file, then Tune it (chart type, period filter, data-check panel).

▸Rubric

Dashboard use case. Is there a hands-on segment where attendees build a dashboard from existing spreadsheet data? FAIL if dashboards are only mentioned, described or promised for later rather than built in the session.

pass

Q3Financial analysis exercise

Judge's reasoning

Build 2 (11:25–12:10) is a standalone "Month-end analysis, Excel in → Excel out" segment producing `Feb_month_end_analysis.xlsx` (flux/variance, Flags, Tie-out) plus job-specific tracks for payroll rec, AR aging and re-forecast.

▸Rubric

Financial analysis use case. Is there a hands-on segment where attendees do financial analysis (Claude writing/running code over their numbers, e.g. Excel in, analysis out)? FAIL if absent or folded into the dashboard segment as a passing remark.

pass

Q4PowerPoint exercise

Judge's reasoning

Build 3 (12:10–12:50) has attendees fill `06_monthly_financial_review_template.pptx` from the analysis workbook and run Dana's own Tune asks ("mirror this chart ×5", "five tables → five slides"), with the timing table marking it hands-on.

▸Rubric

PowerPoint use case. Is there a hands-on segment where attendees update a PowerPoint template/deck from the data — the client's monthly reporting pack? FAIL if absent, or if it is only a demo the facilitator drives while attendees watch.

pass

Q5Relevant business data

Judge's reasoning

Appendix A specifies NetSuite GL transaction detail with Posting Period/Account/Subsidiary/Location/Department/Class, NetSuite budget and forecast exports, payroll register, student AR aging, plus a chart of accounts with EBITDA and payroll account ranges.

▸Rubric

NetSuite/Snowflake-shaped data. Is the synthetic data used in the exercises shaped like this client's actual data — NetSuite/Snowflake-style finance records (GL export, revenue and budget vs actual, EBITDA build-up, payroll vs non-payroll, campus/entity breakdown, monthly and year-to-date columns)? FAIL if the exercises use generic sample data (a demo CSV, sales widgets, made-up SaaS metrics) or leave the data unspecified. The prompt: "the shape of the tasks, the shape of the data should be like NetSuite shaped."

pass

Q6No ASCII icebreaker

Judge's reasoning

Opening exercise is "Replaces the ASCII-map icebreaker" with opening the GL export and building February tuition revenue by campus vs. budget with variance.

▸Rubric

No ASCII icebreaker. Is the opening hands-on exercise a practical finance/Excel task? FAIL if the run of show keeps the ASCII-image-of-your-location icebreaker or substitutes another whimsical non-work exercise. The prompt: "maybe the ASCII image idea that I normally do where I ask it to create an ASCII image based on my location isn['t] like that useful... maybe we use an Excel implementation instead."

pass

Q7Explains Cowork versus ChatGPT

Judge's reasoning

"Cowork vs. ChatGPT/Gemini (3 min)": "A chatbot gives you text back. Cowork works in a folder on your computer: it can open your files, run code (Python) against them, and save new files."

▸Rubric

Cowork vs ChatGPT explained. Does the run of show keep the explanation of how Cowork differs from ChatGPT — that it has access to the file system and can run code? FAIL if that framing is dropped. The prompt: "I still want to keep the way that we introduce Claude Cowork, where we explain to it that the difference between Claude Cowork and ChatGPT is that it has access to the file system. It can run code. I think that important."

pass

Q8No theory or hype

Judge's reasoning

Kickoff is explicitly "No 'why AI' framing" and is logistics only; the only non-build minutes are the working loop, a practical trust-check procedure, and a 2-min list of month-end jobs — no AI industry or future-of-work content.

▸Rubric

No theory, no AI cheerleading. Is every segment tied to a task these people do at work? FAIL if any segment is devoted to AI industry context, the future of work, model capabilities, prompt-engineering theory, or excitement-building. The prompt: "make it very tight run down, make it very condensed and practical. No real like wider theory or like, hey, we're excited about AI things. I think these guys just want to get stuff done."

pass

Q9Mostly hands-on work

Judge's reasoning

Build/hands-on blocks total ~152 of 210 minutes (22+45+45+40), stated as "~150 of 210 minutes", with only 8 min kickoff and 25 min shares/wrap as talk.

▸Rubric

Mostly hands-on. Is at least half the scheduled time attendees working with Claude themselves? FAIL if presentation, discussion and Q&A segments outweigh build segments. From the call: "we try and make at least 50% of it them actually you know working with us to do some of these tasks."

pass

Q10Realistic segment timings

Judge's reasoning

Each build gets 40–45 min broken into Talk/Jam/Make/Tune sub-timings, and first-timer setup is handled by pre-work, a 30-second readiness check and a TA breakout for anyone not ready rather than being ignored.

▸Rubric

Per-segment timings realistic. Could a room of 20–30 beginners actually finish each segment in the time allotted? FAIL if any build segment is implausibly short (a dashboard, an analysis or a deck refresh built end to end in ~15 minutes), or if setup/handholding time for first-time users is ignored.

pass

Q11Fits the booked time

Judge's reasoning

Runs 10:00–13:30 = 3.5 hours with "buffer to 14:00" inside the four-hour block, matching the call notes.

▸Rubric

Total length fits the booked slot. Does the session run about three and a half hours, inside the four-hour block? FAIL if the total is materially shorter or longer (under ~3 hours or over ~4 hours) without the run of show explaining the change. From the call notes: "we have a four hour block for next week and the session itself is three and a half hours."

pass

Q12Ready to run

Judge's reasoning

Every segment carries paste-ready prompts, named data files (`01_netsuite_gl_detail_FY26.xlsx`, `06_monthly_financial_review_template.pptx`), stated outputs, plus pre-work and a data-pack spec — a facilitator could run it as written.

▸Rubric

Usable as a run of show. Could a facilitator run the session from this document alone? FAIL if segments are titles without content — no prompts to paste, no data files named, no statement of what attendees produce — so the facilitator would still have to design the session.