Mike's Checks
Checks

Mike's Checks/gpt-5.6-sol/06 execercise

06 execercise

gpt-5.6-solCodex CLIhigh effortrun 22 Aug 2026

Compare models
10/10
checks passed
100%
▸Instructions — the case's current instructions; none were saved with this result

I have added a bunch of context on this exec offsite session in the context file and I want you to create a folder for Adrian Vale (SVP Strategic Initiatives, initials AV) with 2-3 ideas he could build, with any synthetic data he needs to build it, in a way that matches his role, with a well written prompt for each idea. Put everything in a folder with his initials. Give realistic file formats for the synthetic data.

Answer

PROMPT.md the model's output · 5.64 KBdownload

Build prompt: M&A Integration Control Tower

Copy everything below this line into Claude with the data folder attached.


You are my product builder and integration-management analyst. I am Adrian Vale, SVP of Strategic Initiatives at Aurex. Aurex acquired four capability businesses in 2025—Northwind (now Aurex Prime), Treasura, Tideway, and Vaultline—and must integrate products, customers, systems, risk controls, and teams while capturing the strategic rationale of each deal.

Build an executive-ready M&A Integration Control Tower from the synthetic files in data/. The tool should help an executive find exceptions and unblock decisions; it is not a generic project-management board.

Working method

Inspect all source files first. Report file purpose, row counts, date range, missing values, and join keys. State a short plan before implementation. Do not silently fill gaps or convert an unknown into a green status. Use the source files as the only source of truth.

Inputs

  • data/integration_workstreams.csv: one row per acquisition/workstream, including schedule, progress, status, dependency, blocker, decision, and synergy fields.
  • data/integration_milestones.csv: detailed milestones and acceptance criteria.
  • data/decision_log.md: synthetic notes for decisions already made or still open.

Display “SYNTHETIC DATA — FOR DEMO ONLY” prominently in the product and generated brief.

Definitions and calculations

Use 2026-03-06 as the as-of date.

  • A workstream is overdue if target_date is before the as-of date and percent_complete < 100.
  • A milestone is overdue if its due_date is before the as-of date and its status is not Complete.
  • Planned progress: linearly interpolate from 0% on baseline_date to 100% on target_date, capped between 0% and 100%.
  • Schedule variance equals percent_complete - planned_progress. Show percentage points, not percent.
  • A synergy target is counted once per acquisition + synergy type. Because it repeats across workstream rows, deduplicate it before summing. Sum realized value the same way using the maximum value per acquisition + synergy type, and document this rule.
  • Forecast synergy confidence:
    • High: realized at least 80% of target, or all related workstreams are Green;
    • Medium: realized 40–79% of target and no related workstream is Red;
    • Low: any related workstream is Red, or realized below 40% when the target is nonzero.
  • An item requires executive attention if any of these are true: Red status; overdue; schedule variance at or below -15 points; an open decision is present; blocker has remained open more than 14 days based on blocker_opened_date; or the next milestone is due within 14 days and the item is below 80% complete.

Product requirements

Create a self-contained local web app with as few dependencies as practical. If a local server is unavailable, provide a single-file HTML app that opens directly.

Include:

  • Portfolio KPI cards: overall completion, Red/Amber/Green counts, overdue milestones, open executive decisions, FY26 synergy target, synergy realized, and synergy confidence.
  • One acquisition summary card each for Aurex Prime, Treasura, Tideway, and Vaultline, with completion, schedule variance, RAG mix, top blocker, next milestone, and synergy progress.
  • A cross-deal heat map with acquisitions as columns and workstreams as rows. Every colored cell must also carry a text/icon label.
  • A milestone timeline with overdue and upcoming-30-day views.
  • A deduplicated synergy view by acquisition and synergy type, showing target, realized, gap, and confidence.
  • An executive attention queue, ordered by severity and due date, with the source field(s) that triggered attention.
  • A decision center separating open, decided, and superseded decisions. Link workstream decisions to the corresponding entry in the decision log when IDs match; flag unmatched IDs.
  • Filters for acquisition, workstream, executive sponsor, RAG, owner, status, and “attention required.”
  • A data-quality panel for duplicate keys, impossible dates, percentages outside 0–100, missing owners, unmatched decision IDs, and contradictory milestone/workstream states.

The design should feel like a serious operating review: compact, calm, legible, and optimized for scanning. Include metric definitions in an “About” panel. Do not create a false portfolio average by averaging acquisition percentages; calculate the overall figure as the average of the individual workstream completion values and label it clearly.

Deliverables

Create:

  1. the working app;
  2. README.md with exact run/open instructions, expected input schemas, calculation rules, and steps to refresh the dashboard weekly;
  3. WEEKLY_EXECUTIVE_MEMO.md, maximum 600 words, containing overall status, changes that matter, synergy outlook, top five exceptions, decisions required (with owner and date), and next-two-week milestones;
  4. VALIDATION.md with row counts, join results, synergy-deduplication checks, schedule-variance spot checks, and interaction tests.

Guardrails and quality bar

  • Treat RAG values and notes as reported management data, not objective truth.
  • Never infer legal, regulatory, accounting, security, or personnel approval.
  • Clearly separate a source-data fact, a deterministic calculation, and an AI recommendation.
  • Test every filter, link, sort, modal/drawer, reset state, and empty state. Confirm all displayed totals come from the input files.
  • If something cannot be completed, disclose it in the README with a concrete next step.

At the end, tell me what you built, how to open it, which three synthetic issues most need Adrian's attention, and what live exports should replace these files first.

Checks

10/10
Script checks 3/3answered by a program
pass

Idea folders existfolder

an AV / adrian-vale folder exists with 2-3 idea subfolders

pass

Usable data filesreal-data

every idea has >=1 data file in a realistic format that parses

pass

Prompt for each ideaprompt-per-idea

every idea has a non-trivial prompt file

Judge checks 7/7judged by Claude
pass

Q1Fits the person's role

Judge's reasoning

All three ideas sit in the strategy/corp-dev remit — a partner activation funnel ('move the right partners into active on-network settlement'), an 'M&A Integration Control Tower' for the Northwind/Treasura/Tideway/Vaultline integrations, and a 'Competitive Intelligence Copilot' — matching Vale's own survey asks and his role as exec sponsor on Go-to-market workstreams.

▸Rubric

Role fit. Is every idea something this specific exec — SVP Strategic Initiatives — would own, rather than another function's job (a marketing campaign calendar, an HR onboarding tracker, an engineering ticket triager) or generic executive-productivity filler (inbox summarizer, meeting-notes cleaner) that any exec at any company could have been handed? Mike: "what I'm looking for here is like, does it come up with interesting ideas that are relevant." FAIL if any one idea sits outside the strategy / corp-dev / cross-company-programs remit.

pass

Q2Specific to the company

Judge's reasoning

Everything is Aurex-specific: AUSD penetration weighted by settlement volume, Aurex Prime/Treasura/Tideway/Vaultline integrations with real exec sponsors (Monica Lang, Mariel Soto, Stefan Adler), and competitors like 'Parity — Direct AUSD competitor' and Brightbank's tokenized deposits.

▸Rubric

Company-specific. Are the ideas built on Aurex's actual situation — its products, the AUSD token, the acquisition-integration program, the payments and digital-asset competitive set? FAIL if the deliverable would read identically with the company name swapped for any other mid-size B2B company.

pass

Q3Different ideas

Judge's reasoning

External partner conversion (attention score, funnel stages, corridors), internal post-close integration (synergy dedup, decision center, RAG heat map), and market-signal triage (priority score, corroboration, threat matrix) solve three different problems with different data models and logic.

▸Rubric

Ideas are distinct. Do the 2-3 ideas solve different problems for him? FAIL if two of them are the same artifact with different input data (e.g. two dashboards that differ only in subject).

pass

Q4Self-contained prompts

Judge's reasoning

Each PROMPT.md starts 'Copy everything below this line into Claude with the `data` folder attached' and specifies role, inputs, exact formulas, product requirements, and named deliverables (e.g. 'EXECUTIVE_BRIEF.md, no more than 500 words') with no placeholders or references to prior conversation.

▸Rubric

Prompt is self-contained. Could Vale paste each prompt into a fresh Claude session, with only the files in that idea's folder, and get the thing built without adding anything? FAIL if any prompt has unfilled placeholders (`[INSERT ...]`, `<your company>`), leans on conversation context ("as we discussed", "the ideas above"), or never states what should be produced.

pass

Q5Prompts match the data

Judge's reasoning

Each prompt's Inputs section names exactly the files present — e.g. idea 2 lists integration_workstreams.csv, integration_milestones.csv, decision_log.md, which are the only three files in its data/ folder, and all are exercised by the required calculations.

▸Rubric

Prompt and data match. Does each prompt name the data files that are actually present in that idea's folder, and does each supplied data file get used by the prompt? FAIL if a prompt references a file that does not exist, or the folder ships data the prompt never mentions.

pass

Q6Realistic practice data

Judge's reasoning

Data is substantive and non-generic — 15 named partners with distinct corridors/volumes ('Pacific Horizon Remit,...Philippines-Japan,460,31.8,54'), 20 workstreams, 18 milestones, an 83-line decision log, 6 competitor profiles and 19 signals, plus deliberate edge cases (orphan PTN-099, duplicate URL SIG-014/015).

▸Rubric

The data is real data, not props. Mike: "the thing I look for here is, um, does it, does it create real data?" Format is already checked by script — judge the *content*. FAIL if values are obvious filler (`Company A`, `Competitor 1`, lorem text, the same row repeated, all-identical dates/amounts) or if there is too little of it to build the thing the prompt asks for (e.g. a trend dashboard shipped with four rows).

pass

Q7Buildable during the session

Judge's reasoning

Every prompt asks for 'a self-contained local web app' with a single-file index.html fallback computing from the attached files only, and instructs 'do not browse the web...or add remembered facts' — no live credentials or feeds needed.

▸Rubric

Buildable in the session. Could this be built during a workshop session against the supplied synthetic data? FAIL if any idea's core function requires live credentials or system access the exec won't have in the room (a real NetSuite/Snowflake connection, his production inbox, a paid market-data feed) instead of working off the included files.