Mike's Checks
Checks

Mike's Checks/grok-4.7/09 pptx/compound engineering

09 pptx · compound engineering

grok-4.7Grok CLIhigh effortrun 21 Sep 20265,465,586 tokens

Compare models
14/18
checks passed
78%
▸Instructions — what the model was asked

create a powerpoint in the every brand style on compound engineering, 30 min session https://every.to/guides/compound-engineering

A pptx skill is in skills/pptx/ — read skills/pptx/SKILL.md first. pptxgenjs is installed in the workspace, so require("pptxgenjs") works from here with no network. The brand guide is every-visual-style.md.

Answer

slide-01.png

Checks

14/18
Script checks 5/6answered by a program
fail

Brand colorspalettethis task

share of rendered pixels on Every's documented palette, and whether the deck stays inside one theme (needs Pillow)

pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected create-deck.js, compound-engineering.pptx, every-visual-style.md, skills/pptx/SKILL.md, skills/pptx/LICENSE.txt

pass

Deck opensopensthis task

a .pptx exists and parses (python-pptx when installed, stdlib zip+XML otherwise)

pass

Slide countslide-countthis task

8-20 slides -- sane for a 30-minute session

pass

No overlapping textno-overlapthis task

no two text-bearing shapes on a slide overlap materially (Mike's first question: "Is the writing overlapped?")

pass

Slides renderrenderthis task

every slide rendered to a PNG in <output-dir>/_screenshots/ via LibreOffice (soffice --convert-to pdf) + pdftoppm. FAIL when required render tools are unavailable.

Judge checks 9/12judged by Claude
fail

Q1Required design elementsthis task

Judge's reasoning

There are no classical illustrations anywhere. The swooshes are homemade PNG strokes rather than the brand's E/O/R SVGs, they appear on only 4 of 13 slides (1, 2, 10, 13), and slides 3–9 go seven slides in a row without one.

▸Rubric

The signature elements are there. Every decks carry organic brushstroke swooshes (the E / O / R SVG shapes, cropped at the slide edge, scaled large) and stippled classical illustrations — Greek and Roman architecture, figures, busts, sometimes with modern elements. The guide says to use the swooshes "liberally once every two or three slides." FAIL if the deck has neither swooshes nor classical illustration, or if such elements appear on fewer than roughly one slide in three. Mike's question on this eval: "does it use the, uh, SVGs?"

fail

Q2Enough substancethis task

Judge's reasoning

The loop is explained only in generic terms, and none of the guide's specifics appear: no plugin or its commands, no parallel reviewer agents, no time split between planning/review and coding, no adoption stages. Apart from the UTC date example on slide 10, the deck could have been written without reading the source.

▸Rubric

The content has real depth. Mike: "the main failing of the decks generated by AI is still the content is quite shallow." PASS when the deck carries the load-bearing substance of the selected case's brief and source material: its important mechanisms, distinctions, evidence, and consequences. FAIL if the deck could have been written without reading those materials, reduces the subject to generic advice, or makes unsupported claims.

fail

Q3Does the deck carry the guide's load-bearing compound-engineering substance?this case

Judge's reasoning

The deck covers the plan→work→review→compound loop and names AGENTS.md, a skill and a test, but it leaves out the guide's 80/20 shift (80% of effort on planning and review) and its staged adoption path; the only plan-weighting line is 'Time in the plan returns as fewer rounds of rework.'

▸Rubric

Does the deck carry the guide's load-bearing compound-engineering substance? PASS when it explains the plan → work → review → compound loop, makes the compound step the differentiator from ordinary AI-assisted work, and uses concrete artifacts such as CLAUDE.md, reusable agents, skills, or commands. It should also convey the guide's 80/20 shift and staged adoption path. FAIL when it substitutes generic AI productivity or prompting advice for those mechanics.

pass

Q4Every's color palettethis task

Judge's reasoning

The deck uses Every Black backgrounds, Every Blue accents, white text and orange accents throughout, and slide 3 is the Warm Cream light variant of the same Editorial theme. It never mixes in the Teal theme.

▸Rubric

The palette is Every's. Theme 1 (Editorial): background Every Black `#121212` (or Warm Cream `#F5F2ED` for light slides), Every Blue `#C0F0FB` as the signature accent, white text, with `#FA7B20` / `#E9731E` / `#C4400F` / `#349361` / `#1324CB` / `#7301CC` as accents. Theme 2 (Consulting): background Teal Dark `#0F5258` with Teal Medium `#4999A0` for depth, Warm White `#FFFEFB` text, `#BB7B19` / `#F8DE6E` / `#2E8D23` accents. FAIL if the deck's colors are visibly not these — corporate blue, generic slide-template gray, purple-black AI dark mode, neon gradients — or if it flips between the two themes across slides. Close-but-different shades are fine; a different palette is not.

pass

Q5Serif typographythis task

Judge's reasoning

Headlines and body text are serif (Georgia-style) on every slide. Sans-serif appears only in small labels like the slide 1 session tag, the 'Work' label on slide 8 and the every.to footer.

▸Rubric

Typography is serif. The brand is serif-first: Signifier, or its documented PowerPoint substitutes Georgia and Times New Roman. Sans-serif (Switzer / DM Sans) is allowed only for small labels — session indicators, footers, captions. FAIL if headlines or body text are set in a sans-serif face (Arial, Helvetica, Calibri, Inter). "Sans-serif for main content" is on the brand guide's explicit avoid list.

pass

Q6No banned decorationthis task

Judge's reasoning

There are no triangles, slashes, brackets, drop shadows, gradients, stock photos or rotated text. The only borders are the single thin outlines on slide 9's boxes and the small circles on slide 8, neither of which is banned.

▸Rubric

No forbidden decoration. The brand guide bans, in both themes: triangles, hexagons and other geometric accents; diagonal slashes and hard section dividers; corner brackets or L-shaped frames; drop shadows on boxes; heavy or double borders; stock photography; gradient backgrounds; rotated or outlined text. FAIL if any of these appears on any slide.

pass

Q7Consistent footerthis task

Judge's reasoning

Every slide, 1 through 13, has the EVERY wordmark on the left and the every.to guide URL on the right.

▸Rubric

The footer is consistent. Theme 1: EVERY logo left, `every.to` right. Theme 2: EVERY *Consulting* left, session indicator right. FAIL if content slides carry no footer, or if the footer appears on some and not others.

pass

Q8Interesting materialthis task

Judge's reasoning

The deck argues one point from start to finish: sessions forget (slide 4), complexity can compound for or against you (slide 5), so pay for each lesson once (slide 10) because the repo is the memory (slide 13). It chooses what to say rather than walking through the guide's headings.

▸Rubric

It picked interesting things to say. Mike's question on this eval: "does it pick out interesting things to write about." PASS when the deck makes a selective, coherent argument about what this audience should remember or do. FAIL if it merely inventories the supplied material in source order, restates headings, or includes facts without a point of view about why they matter.

pass

Q9One idea per slidethis task

Judge's reasoning

Each slide carries one message with plenty of space. The busiest ones, slides 7 and 11, still hold to a single idea set out as four short parts.

▸Rubric

One idea per slide. Mike: it fails "if it tries to fit too many ideas on one slide." The brand guide says the same thing — "clear focal point: one main message per slide," "spacious, not cluttered." FAIL if any single slide carries more than one idea, or is packed so densely that a room could not read it while the presenter talks.

pass

Q10Useful visualsthis task

Judge's reasoning

Most slides use visual structure rather than bullets: the three-card problem layout (slide 4), the side-by-side comparison (slide 5), the four-step loop cards (slide 6), the numbered circles (slide 8) and the storage grid (slide 11). Only slides 2, 12 and 13 are list-style.

▸Rubric

Real visual elements, not walls of bullets. Mike's question: "does it create good visual elements." Real means: a diagram of the loop, a staged progression, a 50/50 split with illustration, a quote in an organic blob, a comparison built from shapes — something that carries meaning visually. FAIL if the majority of content slides are a title plus a bulleted list.

pass

Q11Readable slidesthis task

Judge's reasoning

No text is clipped or overflowing, and contrast is readable on every slide. The only near-miss is slide 6's fourth card, which sits tight against the right edge but is not cut off.

▸Rubric

Every slide is legible as rendered. FAIL if any slide has text running off the edge, text clipped by its container, text over an image or shape at unreadable contrast, or headline and body colliding. (A separate script check already catches overlapping text boxes; this question covers what the geometry check cannot see — overflow, clipping and contrast in the actual render.)

pass

Q12Does the session make a deliberate choice about what matters most in the guide?this case

Judge's reasoning

The deck makes a deliberate argument: it opens on 'Why sessions forget', builds through 'Where a lesson has to live' to the close 'The repo is the memory', and ends with a concrete 'The next task' exercise rather than walking through the guide's headings in order.

▸Rubric

Does the session make a deliberate choice about what matters most in the guide? PASS when the 30-minute deck selects and connects the guide's most consequential ideas for the audience. FAIL when it is effectively a table of contents—one slide per source heading in source order, without a point of view about what the audience should retain or change.

Notes

1

Harness: Cursor CLI 2026.09.15-d2fe57e with model grok-4.7-high (Cursor reports it as Grok 4.7 256K High). It was uploaded as harness 'grok' only because Every Checks has no cursor harness yet. It is not the Grok CLI. Cursor's sandbox cannot keep the model from reading local files, and the owner explicitly waived that for this run. The runner is the Every Checks export with ec-bench's Cursor adapter (16dba0b) ported in. The judge was claude-opus-5.

22 Sep 2026