Mike's Checks
Checks

Mike's Checks/grok-4.7/06 execercise

06 execercise

grok-4.7Grok CLIhigh effortrun 21 Sep 2026

Compare models
1/11
checks passed
9%
▸Instructions — what the model was asked

I have added a bunch of context on this exec offsite session in the context file and I want you to create a folder for Adrian Vale (SVP Strategic Initiatives, initials AV) with 2-3 ideas he could build, with any synthetic data he needs to build it, in a way that matches his role, with a well written prompt for each idea. Put everything in a folder with his initials. Give realistic file formats for the synthetic data.

Answer

The model left no answer file or screenshots.

Checks

1/11
Script checks 1/4answered by a program
fail

Idea folders existfolderthis task

an AV / adrian-vale folder exists with 2-3 idea subfolders

fail

Usable data filesreal-datathis task

every idea has >=1 data file in a realistic format that parses

fail

Prompt for each ideaprompt-per-ideathis task

every idea has a non-trivial prompt file

pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected context.md

Judge checks 0/7judged by Claude
fail

Q1Fits the person's rolethis task

Judge's reasoning

The work directory has only context.md, prompt.md and an empty .claude folder, so there is no AV/ folder and no ideas to check for role fit.

▸Rubric

Role fit. Is every idea something this specific exec — SVP Strategic Initiatives — would own, rather than another function's job (a marketing campaign calendar, an HR onboarding tracker, an engineering ticket triager) or generic executive-productivity filler (inbox summarizer, meeting-notes cleaner) that any exec at any company could have been handed? Mike: "what I'm looking for here is like, does it come up with interesting ideas that are relevant." FAIL if any one idea sits outside the strategy / corp-dev / cross-company-programs remit.

fail

Q2Specific to the companythis task

Judge's reasoning

There are no ideas in the output, so nothing draws on Aurex's products, the AUSD token or the acquisition integrations.

▸Rubric

Company-specific. Are the ideas built on Aurex's actual situation — its products, the AUSD token, the acquisition-integration program, the payments and digital-asset competitive set? FAIL if the deliverable would read identically with the company name swapped for any other mid-size B2B company.

fail

Q3Different ideasthis task

Judge's reasoning

The output has no ideas at all, so there are no 2-3 distinct ideas.

▸Rubric

Ideas are distinct. Do the 2-3 ideas solve different problems for him? FAIL if two of them are the same artifact with different input data (e.g. two dashboards that differ only in subject).

fail

Q4Self-contained promptsthis task

Judge's reasoning

The output has no prompt files.

▸Rubric

Prompt is self-contained. Could Vale paste each prompt into a fresh Claude session, with only the files in that idea's folder, and get the thing built without adding anything? FAIL if any prompt has unfilled placeholders (`[INSERT ...]`, `<your company>`), leans on conversation context ("as we discussed", "the ideas above"), or never states what should be produced.

fail

Q5Prompts match the datathis task

Judge's reasoning

There are no prompts or data files to match against each other.

▸Rubric

Prompt and data match. Does each prompt name the data files that are actually present in that idea's folder, and does each supplied data file get used by the prompt? FAIL if a prompt references a file that does not exist, or the folder ships data the prompt never mentions.

fail

Q6Realistic practice datathis task

Judge's reasoning

The output has no synthetic data files.

▸Rubric

The data is real data, not props. Mike: "the thing I look for here is, um, does it, does it create real data?" Format is already checked by script — judge the *content*. FAIL if values are obvious filler (`Company A`, `Competitor 1`, lorem text, the same row repeated, all-identical dates/amounts) or if there is too little of it to build the thing the prompt asks for (e.g. a trend dashboard shipped with four rows).

fail

Q7Buildable during the sessionthis task

Judge's reasoning

With no ideas or data supplied, there is nothing to build in the session.

▸Rubric

Buildable in the session. Could this be built during a workshop session against the supplied synthetic data? FAIL if any idea's core function requires live credentials or system access the exec won't have in the room (a real NetSuite/Snowflake connection, his production inbox, a paid market-data feed) instead of working off the included files.

Notes

1

Harness: Cursor CLI 2026.09.15-d2fe57e with model grok-4.7-high (Cursor reports it as Grok 4.7 256K High). It was uploaded as harness 'grok' only because Every Checks has no cursor harness yet. It is not the Grok CLI. Cursor's sandbox cannot keep the model from reading local files, and the owner explicitly waived that for this run. The runner is the Every Checks export with ec-bench's Cursor adapter (16dba0b) ported in. The judge was claude-opus-5. The agent hit this case's 10m time limit and was stopped. Scoring used whatever it had written by then.

22 Sep 2026