Mike's Checks
Checks

Mike's Checks/gpt-6-astra/06 execercise

06 execercise

gpt-6-astraCodex CLIhigh effortrun 22 Sep 2026301,249 tokens

Compare models
11/11
checks passed
100%
▸Instructions — what the model was asked

I have added a bunch of context on this exec offsite session in the context file and I want you to create a folder for Adrian Vale (SVP Strategic Initiatives, initials AV) with 2-3 ideas he could build, with any synthetic data he needs to build it, in a way that matches his role, with a well written prompt for each idea. Put everything in a folder with his initials. Give realistic file formats for the synthetic data.

Answer

The model left no answer file or screenshots.

Checks

11/11
Script checks 4/4answered by a program
pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected context.md, AV/VALIDATION.md, AV/README.md, AV/ANSWER.md, AV/03_competitive_radar/DEMO_CHECKS.md, AV/03_competitive_radar/PROMPT.md, AV/03_competitive_radar/README.md, AV/03_competitive_radar/data/README.md

pass

Idea folders existfolderthis task

an AV / adrian-vale folder exists with 2-3 idea subfolders

pass

Usable data filesreal-datathis task

every idea has >=1 data file in a realistic format that parses

pass

Prompt for each ideaprompt-per-ideathis task

every idea has a non-trivial prompt file

Judge checks 7/7judged by Claude
pass

Q1Fits the person's rolethis task

Judge's reasoning

All three ideas fit his strategy and cross-company programs remit: '01_execution_dashboard' covers market activation and acquisition integration, '02_ai_pilot_tracker' supports his co-leadership of AI transformation with Sulaiman Rao, and '03_competitive_radar' is his competitive intelligence bot.

▸Rubric

Role fit. Is every idea something this specific exec — SVP Strategic Initiatives — would own, rather than another function's job (a marketing campaign calendar, an HR onboarding tracker, an engineering ticket triager) or generic executive-productivity filler (inbox summarizer, meeting-notes cleaner) that any exec at any company could have been handed? Mike: "what I'm looking for here is like, does it come up with interesting ideas that are relevant." FAIL if any one idea sits outside the strategy / corp-dev / cross-company-programs remit.

pass

Q2Specific to the companythis task

Judge's reasoning

The ideas are built on Aurex specifics, such as 'P02,UAE AUSD activation', 'P04,Treasura customer mapping', 'Vaultline custody controls' and Clearnet/Parity/Brightbank Ledger signals mapped to priorities like 'Increase institutional AUSD usage'.

▸Rubric

Company-specific. Are the ideas built on Aurex's actual situation — its products, the AUSD token, the acquisition-integration program, the payments and digital-asset competitive set? FAIL if the deliverable would read identically with the company name swapped for any other mid-size B2B company.

pass

Q3Different ideasthis task

Judge's reasoning

The three ideas solve different problems: a program RAG/dependency intervention dashboard, an AI pilot value tracker with an enablement-hours allocation scenario, and a deduplicated competitor event board with a weekly brief.

▸Rubric

Ideas are distinct. Do the 2-3 ideas solve different problems for him? FAIL if two of them are the same artifact with different input data (e.g. two dashboards that differ only in subject).

pass

Q4Self-contained promptsthis task

Judge's reasoning

Each prompt opens with 'I am Adrian Vale...', uses 'only the supplied synthetic files', and names its concrete outputs (e.g. 'weekly_interventions.md', 'program_assessments.csv', 'event_register.csv'), with no placeholders or references to earlier conversation.

▸Rubric

Prompt is self-contained. Could Vale paste each prompt into a fresh Claude session, with only the files in that idea's folder, and get the thing built without adding anything? FAIL if any prompt has unfilled placeholders (`[INSERT ...]`, `<your company>`), leans on conversation context ("as we discussed", "the ideas above"), or never states what should be produced.

pass

Q5Prompts match the datathis task

Judge's reasoning

Every supplied file is covered: prompts name files directly (capacity.csv, signals.csv, strategic_priorities.csv) or point to data/README.md, which defines each file, and each data type is referenced (Slack timestamps, dependencies, March 2/9 activation snapshots, emails, source documents).

▸Rubric

Prompt and data match. Does each prompt name the data files that are actually present in that idea's folder, and does each supplied data file get used by the prompt? FAIL if a prompt references a file that does not exist, or the folder ships data the prompt never mentions.

pass

Q6Realistic practice datathis task

Judge's reasoning

The data is realistic and varied, with no filler: e.g. 'UK,P01,...,12,2,1,1800000', observation rows with differing review minutes, and source texts such as 'Parity opens enterprise PUSD payout API... usage is unknown'; volume fits a weekly snapshot.

▸Rubric

The data is real data, not props. Mike: "the thing I look for here is, um, does it, does it create real data?" Format is already checked by script — judge the *content*. FAIL if values are obvious filler (`Company A`, `Competitor 1`, lorem text, the same row repeated, all-identical dates/amounts) or if there is too little of it to build the thing the prompt asks for (e.g. a trend dashboard shipped with four rows).

pass

Q7Buildable during the sessionthis task

Judge's reasoning

All prompts work off the embedded files ('Use only the supplied synthetic files', 'do not browse', 'no external libraries, network calls'), so no live system access is needed.

▸Rubric

Buildable in the session. Could this be built during a workshop session against the supplied synthetic data? FAIL if any idea's core function requires live credentials or system access the exec won't have in the room (a real NetSuite/Snowflake connection, his production inbox, a paid market-data feed) instead of working off the included files.