Mike's Checks
Checks

Mike's Checks/claude-fable-5-1/06 execercise

06 execercise

claude-fable-5-1Claude Codehigh effortrun 22 Sep 2026870,783 tokens

Compare models
5/11
checks passed
45%
▸Instructions — what the model was asked

I have added a bunch of context on this exec offsite session in the context file and I want you to create a folder for Adrian Vale (SVP Strategic Initiatives, initials AV) with 2-3 ideas he could build, with any synthetic data he needs to build it, in a way that matches his role, with a well written prompt for each idea. Put everything in a folder with his initials. Give realistic file formats for the synthetic data.

Answer

The model left no answer file or screenshots.

Checks

5/11
Script checks 2/4answered by a program
fail

Idea folders existfolderthis task

an AV / adrian-vale folder exists with 2-3 idea subfolders

fail

Prompt for each ideaprompt-per-ideathis task

every idea has a non-trivial prompt file

pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected context.md, AV/01_integration_and_market_activation_dashboard/data/steerco_2026-02-24_notes_and_actions.md, AV/01_integration_and_market_activation_dashboard/data/slack_export/channels.json, AV/01_integration_and_market_activation_dashboard/data/slack_export/users.json, AV/01_integration_and_market_activation_dashboard/data/slack_export/integration-pmo/2026-03-06.json, AV/01_integration_and_market_activation_dashboard/data/slack_export/integration-pmo/2026-03-05.json, AV/01_integration_and_market_activation_dashboard/data/slack_export/integration-pmo/2026-03-04.json

pass

Usable data filesreal-datathis task

every idea has >=1 data file in a realistic format that parses

Judge checks 3/7judged by Claude
fail

Q1Self-contained promptsthis task

Judge's reasoning

There are no prompt files anywhere in AV/; each idea folder holds only an empty 'Icon' file and a data/ directory.

▸Rubric

Prompt is self-contained. Could Vale paste each prompt into a fresh Claude session, with only the files in that idea's folder, and get the thing built without adding anything? FAIL if any prompt has unfilled placeholders (`[INSERT ...]`, `<your company>`), leans on conversation context ("as we discussed", "the ideas above"), or never states what should be produced.

fail

Q2Prompts match the datathis task

Judge's reasoning

No prompts exist to name any files, and the data/ folders for ideas 02 and 03 are empty apart from 'Icon'.

▸Rubric

Prompt and data match. Does each prompt name the data files that are actually present in that idea's folder, and does each supplied data file get used by the prompt? FAIL if a prompt references a file that does not exist, or the folder ships data the prompt never mentions.

fail

Q3Realistic practice datathis task

Judge's reasoning

Ideas 02 and 03 ship no data at all ('02_competitive_intelligence_bot_v2/data' holds only 'Icon'), so there is nothing to build them from.

▸Rubric

The data is real data, not props. Mike: "the thing I look for here is, um, does it, does it create real data?" Format is already checked by script — judge the *content*. FAIL if values are obvious filler (`Company A`, `Competitor 1`, lorem text, the same row repeated, all-identical dates/amounts) or if there is too little of it to build the thing the prompt asks for (e.g. a trend dashboard shipped with four rows).

fail

Q4Buildable during the sessionthis task

Judge's reasoning

With no prompts and no data for 02 and 03, the competitive-intelligence bot and the AI scorecard cannot be built from included files; only idea 01 has inputs.

▸Rubric

Buildable in the session. Could this be built during a workshop session against the supplied synthetic data? FAIL if any idea's core function requires live credentials or system access the exec won't have in the room (a real NetSuite/Snowflake connection, his production inbox, a paid market-data feed) instead of working off the included files.

pass

Q5Fits the person's rolethis task

Judge's reasoning

All three idea folders fit the SVP Strategic Initiatives remit: '01_integration_and_market_activation_dashboard', '02_competitive_intelligence_bot_v2', '03_ai_transformation_scorecard'.

▸Rubric

Role fit. Is every idea something this specific exec — SVP Strategic Initiatives — would own, rather than another function's job (a marketing campaign calendar, an HR onboarding tracker, an engineering ticket triager) or generic executive-productivity filler (inbox summarizer, meeting-notes cleaner) that any exec at any company could have been handed? Mike: "what I'm looking for here is like, does it come up with interesting ideas that are relevant." FAIL if any one idea sits outside the strategy / corp-dev / cross-company-programs remit.

pass

Q6Specific to the companythis task

Judge's reasoning

The data that exists is specific to Aurex, e.g. market_activation_tracker.csv has 'AUSD corridor UAE-India live' and 'MAS AUSD recognition', and there are integration-pmo Slack exports.

▸Rubric

Company-specific. Are the ideas built on Aurex's actual situation — its products, the AUSD token, the acquisition-integration program, the payments and digital-asset competitive set? FAIL if the deliverable would read identically with the company name swapped for any other mid-size B2B company.

pass

Q7Different ideasthis task

Judge's reasoning

The three ideas solve different problems: an integration and market-activation dashboard, a competitive-intelligence bot, and an AI transformation scorecard.

▸Rubric

Ideas are distinct. Do the 2-3 ideas solve different problems for him? FAIL if two of them are the same artifact with different input data (e.g. two dashboards that differ only in subject).