Mike's Checks
Checks

Mike's Checks/deepseek-v4-flash/06 execercise

06 execercise

deepseek-v4-flashCodex CLIhigh effortrun 4 Sep 2026760,089 tokens

Compare models
9/10
checks passed
90%
▸Instructions — the case's current instructions; none were saved with this result

I have added a bunch of context on this exec offsite session in the context file and I want you to create a folder for Adrian Vale (SVP Strategic Initiatives, initials AV) with 2-3 ideas he could build, with any synthetic data he needs to build it, in a way that matches his role, with a well written prompt for each idea. Put everything in a folder with his initials. Give realistic file formats for the synthetic data.

Answer

README.md the model's output · 4.17 KBdownload

Adrian Vale (AV) — AI Executive Offsite Project Ideas

Role: SVP Strategic Initiatives, Aurex
Initials: AV

Adrian's role sits at the intersection of M&A integration, market expansion, regulatory strategy, and cross-functional program management. These three ideas are designed to save him 10-20 hours per week on reporting and synthesis, freeing him to focus on execution and decision-making.


Idea 1: M&A Integration Tracker & Status Generator

Folder: 01-ma-integration-tracker/

The problem: Adrian is juggling integration of 4 acquisitions across 10 workstreams each. Producing the weekly exec status report for Dana Brooks takes him 6-8 hours of manual data collection from Slack, email, and spreadsheets.

What it does: Reads structured CSV data on acquisition profiles, workstream progress, risks, and blockers, then generates a polished markdown exec briefing + a consolidated CSV for Google Sheets import. Uses 🟢/🟡/🔴 health indicators across all workstreams.

Synthetic data:

  • data/acquisition_profiles.csv — deal terms, team sizes, revenue synergies
  • data/integration_workstreams.csv — 40 workstream records (4 acquisitions × 10 workstreams) with RACI, status, dates, notes
  • data/risks_and_blockers.csv — 10 flagged risks with owners and mitigation plans
  • data/weekly_exec_summary_template.md — template used by exec team

Expected output: output/integration_status_report.md + output/consolidated_tracker.csv


Idea 2: Market Activation Playbook Generator

Folder: 02-market-playbook/

The problem: Aurex is launching in 3 new markets simultaneously (Brazil, Saudi Arabia, Australia). Each market requires a custom playbook covering licensing, partnerships, hiring, product localization, and risk. Adrian's team spends ~15 hours per market entry on playbook creation.

What it does: Ingests market research data (regulatory checklists, partner profiles, hiring plans) and generates structured activation playbooks per market plus a comparison summary table — cutting playbook creation from 15 hours to under 1 hour.

Synthetic data:

  • data/market_pipeline.csv — 3 target markets with TAM, rationale, competitor landscape
  • data/regulatory_checklist.csv — 11 license/permit requirements across 3 markets with cost & timeline estimates
  • data/local_partners.csv — 12 partner profiles (banks, tech vendors, law firms) with engagement recommendations
  • data/hiring_plan.csv — 18 planned hires across 6+ functions per market

Expected output: output/playbook_brazil.md, output/playbook_saudi_arabia.md, output/playbook_australia.md + output/market_activation_summary.md


Idea 3: Weekly Strategic Initiative Briefing Generator

Folder: 03-strategic-briefing/

The problem: Adrian manages a portfolio of 15 cross-functional strategic initiatives. Every Friday he manually compiles a briefing for Dana Brooks by reading Slack threads, emails, spreadsheets, and meeting notes. This takes 4-5 hours of synthesis work.

What it does: Reads a registry of 15 initiatives, weekly Slack updates from initiative owners, key metrics, and exec meeting notes — then generates a concise, decision-oriented exec briefing that Dana can read in 5 minutes on Monday morning.

Synthetic data:

  • data/initiatives_registry.csv — 15 initiatives with owners, status health, metrics, and week-over-week changes
  • data/weekly_updates_slack.csv — 10 raw Slack messages from 6 execs with initiative updates
  • data/key_metrics.csv — 12 leading indicators with targets and trends
  • data/exec_meeting_notes.md — detailed notes from weekly strategy review including decisions and action items

Expected output: output/exec_briefing_2026-09-07.md + output/weekly_snapshot.csv (appendable to running log)


How to use these with Claude

Each idea folder contains a prompt.md that can be copy-pasted into Claude Desktop or Claude Code. Claude will:

  1. Read the synthetic data files
  2. Follow the structured output requirements
  3. Generate the specified deliverables into the output/ folder

The synthetic data is designed to be immediately replaceable with real Aurex data once the workflow is proven.

Checks

9/10
Script checks 3/3answered by a program
pass

Idea folders existfolder

an AV / adrian-vale folder exists with 2-3 idea subfolders

pass

Usable data filesreal-data

every idea has >=1 data file in a realistic format that parses

pass

Prompt for each ideaprompt-per-idea

every idea has a non-trivial prompt file

Judge checks 6/7judged by Claude
fail

Q1Different ideas

Judge's reasoning

Ideas 1 and 3 are the same weekly exec status artifact with different input data — the README states nearly identical problems ('Producing the weekly exec status report for Dana Brooks takes him 6-8 hours of manual data collection from Slack, email, and spreadsheets' vs 'Every Friday he manually compiles a briefing for Dana Brooks by reading Slack threads, emails, spreadsheets, and meeting notes'), both output a 🟢/🟡/🔴 health-table markdown briefing plus a CSV for Dana on Monday morning, and idea 3's registry (SI-001..SI-004) simply re-states idea 1's four acquisitions at a higher altitude.

▸Rubric

Ideas are distinct. Do the 2-3 ideas solve different problems for him? FAIL if two of them are the same artifact with different input data (e.g. two dashboards that differ only in subject).

pass

Q2Fits the person's role

Judge's reasoning

All three ideas (M&A integration tracker across PayVector/BlockShield/ClearRoute/OmniTreasury, market activation playbooks for Brazil/KSA/Australia, cross-portfolio strategic initiative briefing) sit squarely in the strategy/corp-dev/cross-company-programs remit and echo Vale's survey ask for a 'living dashboard for market activation and M&A programs'.

▸Rubric

Role fit. Is every idea something this specific exec — SVP Strategic Initiatives — would own, rather than another function's job (a marketing campaign calendar, an HR onboarding tracker, an engineering ticket triager) or generic executive-productivity filler (inbox summarizer, meeting-notes cleaner) that any exec at any company could have been handed? Mike: "what I'm looking for here is like, does it come up with interesting ideas that are relevant." FAIL if any one idea sits outside the strategy / corp-dev / cross-company-programs remit.

pass

Q3Specific to the company

Judge's reasoning

Data is built on Aurex's actual situation — 'AUSD Circulation ($B)', 'aurex_product_assignment' mapping deals to Aurex Payments/Custody/Treasura, competitors 'Clearnet (existing stronghold); Beacon (entering)' lifted from the context, plus BCB/SAMA/APRA licensing and the real exec roster (Dana Brooks, Stefan Adler, R. Visser).

▸Rubric

Company-specific. Are the ideas built on Aurex's actual situation — its products, the AUSD token, the acquisition-integration program, the payments and digital-asset competitive set? FAIL if the deliverable would read identically with the company name swapped for any other mid-size B2B company.

pass

Q4Self-contained prompts

Judge's reasoning

Each prompt is a complete standalone block naming its role context, its input files, and its exact deliverables ('Create a consolidated status dashboard as a markdown file (`output/integration_status_report.md`)') with no unfilled placeholders or references to prior conversation.

▸Rubric

Prompt is self-contained. Could Vale paste each prompt into a fresh Claude session, with only the files in that idea's folder, and get the thing built without adding anything? FAIL if any prompt has unfilled placeholders (`[INSERT ...]`, `<your company>`), leans on conversation context ("as we discussed", "the ideas above"), or never states what should be produced.

pass

Q5Prompts match the data

Judge's reasoning

Every named file exists and every shipped file is named — e.g. idea 2 lists market_pipeline.csv, regulatory_checklist.csv, local_partners.csv, hiring_plan.csv and the folder contains exactly those four (only defect is a missing closing backtick on '`hiring_plan.csv — planned hires').

▸Rubric

Prompt and data match. Does each prompt name the data files that are actually present in that idea's folder, and does each supplied data file get used by the prompt? FAIL if a prompt references a file that does not exist, or the folder ships data the prompt never mentions.

pass

Q6Realistic practice data

Judge's reasoning

Data is specific and non-repeating with adequate volume — 40 workstream rows, 12 partner profiles naming Nubank, Matera, Mattos Filho and SABB, regulatory rows like 'PSP License (Pagamento Eletrônico),Banco Central do Brasil (BCB),6-8,350000', and varied risks such as 'FATF compliance gap identified — potential $5M penalty exposure'.

▸Rubric

The data is real data, not props. Mike: "the thing I look for here is, um, does it, does it create real data?" Format is already checked by script — judge the *content*. FAIL if values are obvious filler (`Company A`, `Competitor 1`, lorem text, the same row repeated, all-identical dates/amounts) or if there is too little of it to build the thing the prompt asks for (e.g. a trend dashboard shipped with four rows).

pass

Q7Buildable during the session

Judge's reasoning

All three run entirely off the included CSV/markdown files into an `output/` folder with no live system, credential, or paid-feed dependency.

▸Rubric

Buildable in the session. Could this be built during a workshop session against the supplied synthetic data? FAIL if any idea's core function requires live credentials or system access the exec won't have in the room (a real NetSuite/Snowflake connection, his production inbox, a paid market-data feed) instead of working off the included files.