Mike's Checks
Checks

Mike's Checks/grok-4.6/06 execercise

06 execercise

grok-4.6Grok CLIhigh effortrun 1 Sep 2026

Compare models
10/10
checks passed
100%
▸Instructions — the case's current instructions; none were saved with this result

I have added a bunch of context on this exec offsite session in the context file and I want you to create a folder for Adrian Vale (SVP Strategic Initiatives, initials AV) with 2-3 ideas he could build, with any synthetic data he needs to build it, in a way that matches his role, with a well written prompt for each idea. Put everything in a folder with his initials. Give realistic file formats for the synthetic data.

Answer

ANSWER.md the model's output · 3.37 KBdownload

Adrian Vale session pack

Everything for Adrian Vale (SVP Strategic Initiatives) is in AV/. Three build ideas, each with a copy-paste prompt and a synthetic data pack in realistic formats. Nothing in data/ is live Aurex information.

What to have him build

Folder Walk-out artifact Role fit
AV/01-ma-market-activation-dashboard/ Local HTML Activation OS dashboard Cross-company view of the 2025 acquisitions, eight new markets, and partner conversion from messaging to on-network settlement
AV/02-competitive-intelligence-brief/ Staff-ready weekly brief + reusable “bot” prompt The Sunday-night competitive read for Dana / Monica / staff (Clearnet, Brightbank Ledger, Parity, Cardinal, Beacon)
AV/03-exec-staff-preread/ Tuesday staff pre-read + talking points + Monday workflow SI’s weekly operating cadence — including the AI-transformation beat he is personally driving

If he can only ship one: 01 for demo impact, 03 if he wants something he will use at 08:30 the same morning. 02 is the fastest second build.

Each prompt is written in his voice, names the files to read, forbids fabricating extra deals/markets, and asks Claude to surface tracker-vs-narrative conflicts rather than smooth them over. That is deliberate: the weekly memos contradict the CSVs in a few places (Treasura RAG, the 40% on-network figure, Luxembourg as a “launch”).

Synthetic data formats

Format What it stands in for
.csv Smartsheet / Sheets / PMO exports (acquisitions, workstreams, synergies, markets, partner funnel, RAID, win/loss, KPIs, initiative portfolio)
.md Workstream lead updates, news clips, last week’s staff notes, data dictionaries, briefing template
.eml Internal Outlook/Gmail forwards (Corp Dev, Legal, Network, Treasura, People, Comms)
.json Slack export of #strategic-initiatives (plus a private leads channel)
.ics Calendar slice for Miami week and the deadlines the pre-read has to respect

Connectors are optional. Every idea runs off local files so a Gmail/Slack permission failure cannot block the session.

Assumptions (called, not hidden)

  • Hope Carver and Sulaiman Rao sit on Adrian’s SI team (they were on the 26 Feb planning call).
  • Integration PMO is dotted-line into SI; deal origination stays with Erik Jansen.
  • Internal headcount (~1,500) is used, not the third-party 925 figure.
  • “R. Visser, Treasura” is treated as Rik Visser for email realism.
  • Named workstream leads below the executive team (Daniela Ruiz, James Okonkwo, Priya Nair, Lena Vogt, etc.), all synergy dollars, partner-level volumes, and win/loss rows are invented.
  • Data freeze is Friday 6 March 2026; the session is Tuesday 10 March 2026.
  • Two company facts from the context pack are intentionally left in tension inside the data, because they are his actual job: “eight new markets” vs. licenses without GTM, and “~40% on-network” vs. a denominator Partnerships cannot currently defend.

What this pack does not include

A finished dashboard, brief, or pre-read. The point of the session is that he ships it. Prompts specify the output filenames so a trainer can tell, in the room, whether v1 landed.

Checks

10/10
Script checks 3/3answered by a program
pass

Idea folders existfolder

an AV / adrian-vale folder exists with 2-3 idea subfolders

pass

Usable data filesreal-data

every idea has >=1 data file in a realistic format that parses

pass

Prompt for each ideaprompt-per-idea

every idea has a non-trivial prompt file

Judge checks 7/7judged by Claude
pass

Q1Fits the person's role

Judge's reasoning

All three ideas sit in strategy/corp-dev remit — M&A integration + market activation dashboard, competitive intelligence brief, and a cross-company staff pre-read built from 'the SI-owned and SI-shepherded programs' (initiative_portfolio.csv), not generic exec productivity.

▸Rubric

Role fit. Is every idea something this specific exec — SVP Strategic Initiatives — would own, rather than another function's job (a marketing campaign calendar, an HR onboarding tracker, an engineering ticket triager) or generic executive-productivity filler (inbox summarizer, meeting-notes cleaner) that any exec at any company could have been handed? Mike: "what I'm looking for here is like, does it come up with interesting ideas that are relevant." FAIL if any one idea sits outside the strategy / corp-dev / cross-company-programs remit.

pass

Q2Specific to the company

Judge's reasoning

Deliverables are saturated with Aurex specifics — 'AUSD collateral attach fees' synergy lines, the five 2025 deals (Northwind/Prime, Treasura, Tideway, Vaultline, Castellan), the messaging-vs-on-network partner funnel, and the payments/token competitive set (Clearnet, Beacon, Brightbank Ledger, Cardinal, Parity/PUSD).

▸Rubric

Company-specific. Are the ideas built on Aurex's actual situation — its products, the AUSD token, the acquisition-integration program, the payments and digital-asset competitive set? FAIL if the deliverable would read identically with the company name swapped for any other mid-size B2B company.

pass

Q3Different ideas

Judge's reasoning

Distinct problems and artifacts: an HTML program-operating dashboard, an external competitive brief plus reusable weekly briefing engine, and a 700-word decision-oriented staff pre-read with talking points.

▸Rubric

Ideas are distinct. Do the 2-3 ideas solve different problems for him? FAIL if two of them are the same artifact with different input data (e.g. two dashboards that differ only in subject).

pass

Q4Self-contained prompts

Judge's reasoning

Each PROMPT.md is a copy-paste block with full company framing, named output filenames ('a single-file HTML dashboard named aurex-activation-os.html'), explicit structure, and no placeholders or references to prior conversation.

▸Rubric

Prompt is self-contained. Could Vale paste each prompt into a fresh Claude session, with only the files in that idea's folder, and get the thing built without adding anything? FAIL if any prompt has unfilled placeholders (`[INSERT ...]`, `<your company>`), leans on conversation context ("as we discussed", "the ideas above"), or never states what should be produced.

pass

Q5Prompts match the data

Judge's reasoning

Every prompt enumerates its folder's files exactly and each shipped file is referenced — e.g. prompt 02 lists all six data files plus 'twelve items' in news_clips/ (12 present) and three emails (3 present).

▸Rubric

Prompt and data match. Does each prompt name the data files that are actually present in that idea's folder, and does each supplied data file get used by the prompt? FAIL if a prompt references a file that does not exist, or the folder ships data the prompt never mentions.

pass

Q6Realistic practice data

Judge's reasoning

Content is specific and internally consistent, not filler — e.g. raid_log R-002 'AUSD go-live date promised as Q2 sold as Q3 engineered' with a cost-of-delay note, 30 workstream rows, 24 KPI rows with confidence flags, and .eml/.ics/.json files with realistic headers and named actors.

▸Rubric

The data is real data, not props. Mike: "the thing I look for here is, um, does it, does it create real data?" Format is already checked by script — judge the *content*. FAIL if values are obvious filler (`Company A`, `Competitor 1`, lorem text, the same row repeated, all-identical dates/amounts) or if there is too little of it to build the thing the prompt asks for (e.g. a trend dashboard shipped with four rows).

pass

Q7Buildable during the session

Judge's reasoning

Each prompt states 'All processing happens from these local files. Do not try to call Slack, Gmail, or an API' / 'Local files only', and the README notes connectors are optional so 'permissions cannot block the build.'

▸Rubric

Buildable in the session. Could this be built during a workshop session against the supplied synthetic data? FAIL if any idea's core function requires live credentials or system access the exec won't have in the room (a real NetSuite/Snowflake connection, his production inbox, a paid market-data feed) instead of working off the included files.