Mike's Checks/gpt-6-astra/05 showrunner
05 showrunner
gpt-6-astraCodex CLIhigh effortrun 22 Sep 2026164,605 tokens
▸Instructions — what the model was asked
Hi, I need to create a modified Claude code for beginners session, like run sheet basically, run of show, that has the following modifications. So I'm going to upload a few different examples of things I've done in the past, just so you have them to work from. But basically, there's a few modifications for this one which make it unique. One is that this is only using Claude Cowork, so we should not be using Claude Code, either in the desktop app or in the terminal, because this team is non-technical. The team are all working in financial services slash accounting for a education company. So what kind of workplace education company called Brightmoor Education. So the sorry that spelled b the uh team is very busy and um doesn have a lot of like time to explore speculative use cases of ai they want this to be incredibly practical and based on their existing workflows. I'm uploading some transcripts from a call we had with them and also just some notes on what I think could work. Specifically, they're interested in... creating the dashboards and specifically they mean just making the information look visibly better like like creating a html dashboard in a nice kind of style from from existing say data that contained in Excel They also interested in doing financial analysis So more of a, you know, like, I guess, a cloud coworker would use Python to do the analysis and then maybe output the analysis in Excel or the input would be in Excel. And then the third thing was like making a PowerPoint where they update the PowerPoint template automatically using Claude Cowork so that they can do kind of weekly or monthly presentations on the data. So that's the goal, is to kind of make it very focused on that. I still want to keep the way that we introduce Claude Cowork, where we explain to it that the difference between Claude Cowork and ChatGPT is that it has access to the file system. It can run code. I think that important But maybe the ASCII image idea that I normally do where I ask it to create an ASCII image based on my location isn like that useful I think we need to make it more practical for them So maybe we use an Excel implementation instead. The assumption is... oh, so they use Snowflake, by the way, for all their data, and that is powered by NetSuite, which is our accounting software. I guess that's by Oracle. So when we create synthetic data, we'll be, for these tasks that we're going to demo, we'll be focused specifically on solving that, solving those problems. So like the shape of the tasks, the shape of the data should be like NetSuite shaped, if that makes sense. So make it very tight run down, make it very condensed and practical. No real like wider theory or like, hey, we're excited about AI things. I think these guys just want to get stuff done.
So yeah, let me know what the run should be.
some very specific asks in here
- definitely non-technical (claude cowork only + claude in excel)
- core focus on accounting/finance function
- use cases to demo: creating a dashboard, doing financial analysis, making a powerpoint
heavy focus on practical stuff rather than theory - they're just super busy and want to get stuff done, have heard from friends at other companies you can do it with claude
For accounting: put financials together, doing reconciliations, identifying trends in data sets, putting together repeatable compliance reports For planning & forecast: build financial models (make assumptions and pull & clean data from snowflake)create dashboards, and publish to executivesAutomate charts on slides
Answer
Claude Cowork for Brightmoor Finance
Recommended session: turn a monthly finance export into a checked Excel analysis, a readable dashboard, and an updated leadership presentation. Everyone works on the same reporting cycle, using Cowork in Claude Desktop and Claude in Excel only.
Audience: 20–30 colleagues, mostly accounting, payroll and bursar teams, plus strategic finance. Assume beginner knowledge. Duration: four hours, including a ten-minute break; approximately 140 minutes of participant practice. The call mentions both three and four hours; this schedule assumes four, with a three-hour version below.
The practical promise: leave with three draft deliverables and a process for refreshing them next month. Use three representative slides from the recurring reporting pack; reproducing all 15 slides precisely is follow-on work.
Run of show
Times are elapsed from the session start. “Practice” includes participants reviewing and correcting their outputs.
| Time | Segment | Facilitator action | Participant action and finish line |
|---|---|---|---|
| 00:00–00:05 | Show the destination | Open the finished sample workbook, dashboard and three-slide monthly report. Explain that all three use the same numbers. | See exactly what they will make. |
| 00:05–00:25 | First win: make an Excel export usable | Give the two-minute Cowork introduction below. Show selecting the workshop folder and requesting a formatted workbook copy. Open that copy in Excel and demonstrate one question through Claude in Excel. 8 minutes demo / 12 practice. | Format the export, freeze headers, add filters, and ask which campus has the largest revenue variance. Find the supporting cells. Finish: a usable workbook and one traceable answer. |
| 00:25–01:20 | Build 1: financial analysis and reconciliation | In Cowork, assemble monthly and YTD actuals against budget, with payroll, non-payroll and an EBITDA bridge. Show an export-to-control reconciliation and investigate a seeded mismatch. In Claude in Excel, inspect a formula and change one forecast assumption. 20 minutes demo / 35 practice. | Produce the analysis, review exceptions, trace one variance to source rows and test a simple revenue/payroll scenario. Finish: checked totals, an exception list and an assumption-driven forecast. |
| 01:20–01:30 | Break | Leave the next prompt and checkpoint filenames on screen. | — |
| 01:30–02:15 | Build 2: executive dashboard | Ask Cowork to turn the checked workbook into one local HTML dashboard. Show how to open it in a browser. Demonstrate changing a chart to horizontal bars, then adding a campus filter. 15 minutes demo / 30 practice. | Build and tune their dashboard: revenue, EBITDA, budget variances and monthly trends. Compare filtered results with Excel. Finish: a readable dashboard with working campus and period filters. |
| 02:15–03:10 | Build 3: refresh the PowerPoint | Give Cowork the checked workbook and an existing .pptx template. Update three slides: revenue against budget, EBITDA bridge, and campus performance with commentary. Demonstrate reusing the template’s chart style for another campus. 20 minutes demo / 35 practice. |
Update the template, correct one formatting issue and inspect the chart data. Finish: a three-slide draft that matches the workbook and retains the supplied layout as closely as possible. |
| 03:10–03:45 | Repeat the workflow with next month’s data | Supply the next-month export. Reuse the saved instructions. Show the internal sharing route using the prepared example. 10 minutes demo / 25 practice. | Refresh the workbook, dashboard and one slide; catch a stale title or commentary sentence. Save a short refresh checklist. Finish: evidence that the process can be repeated. |
| 03:45–04:00 | Review and immediate next use | Take two brief walkthroughs and answer questions about the files people made. 10 minutes review / 5 practice. | Record one recurring task, its input file, desired output and reviewer. Leave with saved files and prompts. |
The introduction: two minutes, then do the task
“Today you’ll give Claude a folder, describe the finance task, and ask it to save the finished files. Cowork can work with files you grant it access to and run code behind the scenes to process data. You won’t write or read that code. Claude in Excel lets you work directly with the workbook—ask about cells, inspect calculations, and make changes.”
“If you normally use ChatGPT as a chat window, the useful change today is working through a task across files and producing deliverables. File analysis and code execution aren’t exclusive to Cowork; we’re learning this particular workflow.”
Keep the familiar Talk → Jam → Make → Tune sequence inside each exercise: describe the output, settle the calculation and layout rules, create the file, then request one specific improvement. Check the numbers before using them downstream. No separate theory block.
Prepare before the session
- Access: verify every attendee can open Cowork in Claude Desktop, select the workshop folder and save a file. Verify Claude in Excel is installed, enabled and signed in. Send a five-minute preflight task beforehand; avoid spending workshop time provisioning accounts.
- Materials: prepare the synthetic workbook, three-slide PowerPoint template, next-month export and completed checkpoints described in DEMO-PROMPTS.md. These are preparation requirements, not files supplied with this run sheet. Rehearse the actual template refresh, especially charts and fonts.
- Delivery: one lead facilitator and one organiser-provided helper for access issues. Do not rely on Brightmoor’s more experienced staff as teaching assistants; the call explicitly ruled that out. If an exercise stalls, use its prepared checkpoint and continue.
- Data route: explain NetSuite → Snowflake → Excel export → Cowork → reviewed outputs. Use synthetic exports for the workshop. If an approved Snowflake connection is already working, briefly show retrieving the same dataset through it; connector setup and SQL are outside this session.
- Sharing: rehearse the company-approved route for executives. Show sharing the reviewed deck/workbook through existing internal permissions. If internal HTML hosting is available, demonstrate it; otherwise use the local HTML and a PDF snapshot. A local HTML file is not an authenticated reporting service, and adding a password field does not secure its embedded data.
Make the checks part of the work
Use a small number of visible checks: revenue and expense totals agree to the supplied control report; campus totals sum to the overall result; YTD includes the intended periods; the EBITDA bridge uses the supplied account map; and the dashboard and deck match the checked workbook. Show both a revenue shortfall and an expense overrun so participants see how favourable/unfavourable labels differ.
Have Claude identify the largest numerical contributors to a variance and separate those observations from possible business explanations. Keep unresolved reconciliation items visible. For slides, inspect both the displayed number and its period label; last month’s commentary can survive a successful chart update.
Use a fixed recovery prompt when stuck: “Here is what I expected and what happened. Fix only this issue, preserve the checked numbers, and save a new version.” If it still fails, return to the last working checkpoint and reduce the task to one sheet, chart or slide.
If the session is three hours
Use: destination 5 min; first win 15; financial analysis/reconciliation 40; break 10; dashboard 35; PowerPoint 45; next-month refresh 20; wrap 10. Total: 180 minutes. Keep all three core outputs. Move the forecast scenario and compliance-report extension to take-home prompts, and update two slides live instead of three.
Basis and assumptions
Adapted from the Brightmoor call notes and previous workshop runs. The workbook, template, reporting definitions, exact session length and internal hosting route were not supplied. Use the synthetic brief below until real examples are available. Current product documentation could not be reached during preparation, so account access, Excel features and PowerPoint template behaviour must be verified in the facilitator rehearsal; no live connector or exact template-fidelity guarantee is assumed.
Checks
16/17No AI-tell wordingai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected run-of-show-nvalo.md, ANSWER.md, DEMO-PROMPTS.md, brightmoor-education-call-notes.md
Run of show existsexiststhis task
a run-of-show document was produced
Coherent scheduleformatthis task
it parses as a run of show: one coherent schedule, no segment that ends before it starts, >=5 time-ranged rows carrying activity text (md table rows, list items, bold headings "(9:00-9:15)", single-clock table rows, or plain lines all count)
Five-minute time markstime-grainthis task
segment boundaries sit on 5-minute marks, at most one off-grid -- a schedule planned at :02/:08/:17 is arithmetic, not a plan
Timings add uptimings-tie-outthis task
per schedule, segments are contiguous: no overlaps, no gap over 30 min, >=85% of the span scheduled, and a span between 45 and 600 minutes
Q1Explains Cowork versus ChatGPTthis task
Judge's reasoning
The intro undercuts the requested framing: 'File analysis and code execution aren't exclusive to Cowork; we're learning this particular workflow.' It does not present file-system access and running code as what sets Cowork apart from ChatGPT.
▸Rubric
Cowork vs ChatGPT explained. Does the run of show keep the explanation of how Cowork differs from ChatGPT — that it has access to the file system and can run code? FAIL if that framing is dropped. The prompt: "I still want to keep the way that we introduce Claude Cowork, where we explain to it that the difference between Claude Cowork and ChatGPT is that it has access to the file system. It can run code. I think that important."
Q2Uses Cowork throughoutthis task
Judge's reasoning
Everything attendees do stays in Cowork and Claude in Excel ('Cowork in Claude Desktop and Claude in Excel only'); there are no Claude Code, terminal, git or dev-tool steps.
▸Rubric
Cowork-only. Do all attendee-facing steps stay inside Claude Cowork (and Claude in Excel)? FAIL if any step the attendees are asked to do involves Claude Code, the terminal/CLI, git or GitHub, or installing developer tooling. The prompt: "this is only using Claude Cowork, so we should not be using Claude Code, either in the desktop app or in the terminal, because this team is non-technical."
Q3Dashboard exercisethis task
Judge's reasoning
Build 2 (01:30–02:15) has attendees build and tune a local HTML dashboard from the checked workbook, with 30 minutes of practice.
▸Rubric
Dashboard use case. Is there a hands-on segment where attendees build a dashboard from existing spreadsheet data? FAIL if dashboards are only mentioned, described or promised for later rather than built in the session.
Q4Financial analysis exercisethis task
Judge's reasoning
Build 1 (00:25–01:20) is a hands-on analysis: monthly/YTD budget vs actual, payroll/non-payroll, an EBITDA bridge, reconciliation and a forecast scenario, run by Cowork into monthly_analysis.xlsx.
▸Rubric
Financial analysis use case. Is there a hands-on segment where attendees do financial analysis (Claude writing/running code over their numbers, e.g. Excel in, analysis out)? FAIL if absent or folded into the dashboard segment as a passing remark.
Q5PowerPoint exercisethis task
Judge's reasoning
In Build 3, attendees refresh monthly_template.pptx themselves ('Update the template, correct one formatting issue and inspect the chart data'), with 35 minutes of practice.
▸Rubric
PowerPoint use case. Is there a hands-on segment where attendees update a PowerPoint template/deck from the data — the client's monthly reporting pack? FAIL if absent, or if it is only a demo the facilitator drives while attendees watch.
Q6Relevant business datathis task
Judge's reasoning
The data pack specifies a 'NetSuite-derived Snowflake export' with GL_Lines (transaction/line IDs, subsidiary, campus, account), an Account_Map with payroll/non-payroll and EBITDA treatment, Budget, Control_Totals and month/YTD definitions.
▸Rubric
NetSuite/Snowflake-shaped data. Is the synthetic data used in the exercises shaped like this client's actual data — NetSuite/Snowflake-style finance records (GL export, revenue and budget vs actual, EBITDA build-up, payroll vs non-payroll, campus/entity breakdown, monthly and year-to-date columns)? FAIL if the exercises use generic sample data (a demo CSV, sales widgets, made-up SaaS metrics) or leave the data unspecified. The prompt: "the shape of the tasks, the shape of the data should be like NetSuite shaped."
Q7No ASCII icebreakerthis task
Judge's reasoning
The opening exercise is an Excel task: format the GL export and ask Claude in Excel which campus has the largest revenue variance.
▸Rubric
No ASCII icebreaker. Is the opening hands-on exercise a practical finance/Excel task? FAIL if the run of show keeps the ASCII-image-of-your-location icebreaker or substitutes another whimsical non-work exercise. The prompt: "maybe the ASCII image idea that I normally do where I ask it to create an ASCII image based on my location isn['t] like that useful... maybe we use an Excel implementation instead."
Q8No theory or hypethis task
Judge's reasoning
Every segment is tied to a finance task; the document says 'No separate theory block' and has no AI-hype segment.
▸Rubric
No theory, no AI cheerleading. Is every segment tied to a task these people do at work? FAIL if any segment is devoted to AI industry context, the future of work, model capabilities, prompt-engineering theory, or excitement-building. The prompt: "make it very tight run down, make it very condensed and practical. No real like wider theory or like, hey, we're excited about AI things. I think these guys just want to get stuff done."
Q9Mostly hands-on workthis task
Judge's reasoning
About 140 of 240 minutes are participant practice, and each build segment is split so practice outweighs demo.
▸Rubric
Mostly hands-on. Is at least half the scheduled time attendees working with Claude themselves? FAIL if presentation, discussion and Q&A segments outweigh build segments. From the call: "we try and make at least 50% of it them actually you know working with us to do some of these tasks."
Q10Realistic segment timingsthis task
Judge's reasoning
Build segments get 45–55 minutes (30–35 minutes of practice) with prepared checkpoints, a helper and a preflight access check, which is plausible for beginners.
▸Rubric
Per-segment timings realistic. Could a room of 20–30 beginners actually finish each segment in the time allotted? FAIL if any build segment is implausibly short (a dashboard, an analysis or a deck refresh built end to end in ~15 minutes), or if setup/handholding time for first-time users is ignored.
Q11Fits the booked timethis task
Judge's reasoning
The run is 00:00–04:00. That is not over the ~4-hour limit and the document explains its assumption and gives a 3-hour variant, but it misreads the call (which says 3.5 hours) and uses up the whole buffer.
▸Rubric
Total length fits the booked slot. Does the session run about three and a half hours, inside the four-hour block? FAIL if the total is materially shorter or longer (under ~3 hours or over ~4 hours) without the run of show explaining the change. From the call notes: "we have a four hour block for next week and the session itself is three and a half hours."
Q12Ready to runthis task
Judge's reasoning
It gives paste-ready prompts, named files (finance_export.xlsx, monthly_template.pptx, next_month.xlsx), a finish line for each segment, checks and recovery steps.
▸Rubric
Usable as a run of show. Could a facilitator run the session from this document alone? FAIL if segments are titles without content — no prompts to paste, no data files named, no statement of what attendees produce — so the facilitator would still have to design the session.