Mike's Checks/grok-4.7/09 pptx/compound engineering
09 pptx · compound engineering
grok-4.7Grok CLIhigh effortrun 21 Sep 20265,465,586 tokens
▸Instructions — what the model was asked
create a powerpoint in the every brand style on compound engineering, 30 min session https://every.to/guides/compound-engineering
A pptx skill is in skills/pptx/ — read skills/pptx/SKILL.md first. pptxgenjs is installed in the workspace, so require("pptxgenjs") works from here with no network. The brand guide is every-visual-style.md.
Answer
Checks
14/18Brand colorspalettethis task
share of rendered pixels on Every's documented palette, and whether the deck stays inside one theme (needs Pillow)
No AI-tell wordingai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected create-deck.js, compound-engineering.pptx, every-visual-style.md, skills/pptx/SKILL.md, skills/pptx/LICENSE.txt
Deck opensopensthis task
a .pptx exists and parses (python-pptx when installed, stdlib zip+XML otherwise)
Slide countslide-countthis task
8-20 slides -- sane for a 30-minute session
No overlapping textno-overlapthis task
no two text-bearing shapes on a slide overlap materially (Mike's first question: "Is the writing overlapped?")
Slides renderrenderthis task
every slide rendered to a PNG in <output-dir>/_screenshots/ via LibreOffice (soffice --convert-to pdf) + pdftoppm. FAIL when required render tools are unavailable.
Q1Required design elementsthis task
Judge's reasoning
There are no classical illustrations anywhere. The swooshes are homemade PNG strokes rather than the brand's E/O/R SVGs, they appear on only 4 of 13 slides (1, 2, 10, 13), and slides 3–9 go seven slides in a row without one.
▸Rubric
The signature elements are there. Every decks carry organic brushstroke swooshes (the E / O / R SVG shapes, cropped at the slide edge, scaled large) and stippled classical illustrations — Greek and Roman architecture, figures, busts, sometimes with modern elements. The guide says to use the swooshes "liberally once every two or three slides." FAIL if the deck has neither swooshes nor classical illustration, or if such elements appear on fewer than roughly one slide in three. Mike's question on this eval: "does it use the, uh, SVGs?"
Q2Enough substancethis task
Judge's reasoning
The loop is explained only in generic terms, and none of the guide's specifics appear: no plugin or its commands, no parallel reviewer agents, no time split between planning/review and coding, no adoption stages. Apart from the UTC date example on slide 10, the deck could have been written without reading the source.
▸Rubric
The content has real depth. Mike: "the main failing of the decks generated by AI is still the content is quite shallow." PASS when the deck carries the load-bearing substance of the selected case's brief and source material: its important mechanisms, distinctions, evidence, and consequences. FAIL if the deck could have been written without reading those materials, reduces the subject to generic advice, or makes unsupported claims.
Q3Does the deck carry the guide's load-bearing compound-engineering substance?this case
Judge's reasoning
The deck covers the plan→work→review→compound loop and names AGENTS.md, a skill and a test, but it leaves out the guide's 80/20 shift (80% of effort on planning and review) and its staged adoption path; the only plan-weighting line is 'Time in the plan returns as fewer rounds of rework.'
▸Rubric
Does the deck carry the guide's load-bearing compound-engineering substance? PASS when it explains the plan → work → review → compound loop, makes the compound step the differentiator from ordinary AI-assisted work, and uses concrete artifacts such as CLAUDE.md, reusable agents, skills, or commands. It should also convey the guide's 80/20 shift and staged adoption path. FAIL when it substitutes generic AI productivity or prompting advice for those mechanics.
Q4Every's color palettethis task
Judge's reasoning
The deck uses Every Black backgrounds, Every Blue accents, white text and orange accents throughout, and slide 3 is the Warm Cream light variant of the same Editorial theme. It never mixes in the Teal theme.
▸Rubric
The palette is Every's. Theme 1 (Editorial): background Every Black `#121212` (or Warm Cream `#F5F2ED` for light slides), Every Blue `#C0F0FB` as the signature accent, white text, with `#FA7B20` / `#E9731E` / `#C4400F` / `#349361` / `#1324CB` / `#7301CC` as accents. Theme 2 (Consulting): background Teal Dark `#0F5258` with Teal Medium `#4999A0` for depth, Warm White `#FFFEFB` text, `#BB7B19` / `#F8DE6E` / `#2E8D23` accents. FAIL if the deck's colors are visibly not these — corporate blue, generic slide-template gray, purple-black AI dark mode, neon gradients — or if it flips between the two themes across slides. Close-but-different shades are fine; a different palette is not.
Q5Serif typographythis task
Judge's reasoning
Headlines and body text are serif (Georgia-style) on every slide. Sans-serif appears only in small labels like the slide 1 session tag, the 'Work' label on slide 8 and the every.to footer.
▸Rubric
Typography is serif. The brand is serif-first: Signifier, or its documented PowerPoint substitutes Georgia and Times New Roman. Sans-serif (Switzer / DM Sans) is allowed only for small labels — session indicators, footers, captions. FAIL if headlines or body text are set in a sans-serif face (Arial, Helvetica, Calibri, Inter). "Sans-serif for main content" is on the brand guide's explicit avoid list.
Q6No banned decorationthis task
Judge's reasoning
There are no triangles, slashes, brackets, drop shadows, gradients, stock photos or rotated text. The only borders are the single thin outlines on slide 9's boxes and the small circles on slide 8, neither of which is banned.
▸Rubric
No forbidden decoration. The brand guide bans, in both themes: triangles, hexagons and other geometric accents; diagonal slashes and hard section dividers; corner brackets or L-shaped frames; drop shadows on boxes; heavy or double borders; stock photography; gradient backgrounds; rotated or outlined text. FAIL if any of these appears on any slide.
Q7Consistent footerthis task
Judge's reasoning
Every slide, 1 through 13, has the EVERY wordmark on the left and the every.to guide URL on the right.
▸Rubric
The footer is consistent. Theme 1: EVERY logo left, `every.to` right. Theme 2: EVERY *Consulting* left, session indicator right. FAIL if content slides carry no footer, or if the footer appears on some and not others.
Q8Interesting materialthis task
Judge's reasoning
The deck argues one point from start to finish: sessions forget (slide 4), complexity can compound for or against you (slide 5), so pay for each lesson once (slide 10) because the repo is the memory (slide 13). It chooses what to say rather than walking through the guide's headings.
▸Rubric
It picked interesting things to say. Mike's question on this eval: "does it pick out interesting things to write about." PASS when the deck makes a selective, coherent argument about what this audience should remember or do. FAIL if it merely inventories the supplied material in source order, restates headings, or includes facts without a point of view about why they matter.
Q9One idea per slidethis task
Judge's reasoning
Each slide carries one message with plenty of space. The busiest ones, slides 7 and 11, still hold to a single idea set out as four short parts.
▸Rubric
One idea per slide. Mike: it fails "if it tries to fit too many ideas on one slide." The brand guide says the same thing — "clear focal point: one main message per slide," "spacious, not cluttered." FAIL if any single slide carries more than one idea, or is packed so densely that a room could not read it while the presenter talks.
Q10Useful visualsthis task
Judge's reasoning
Most slides use visual structure rather than bullets: the three-card problem layout (slide 4), the side-by-side comparison (slide 5), the four-step loop cards (slide 6), the numbered circles (slide 8) and the storage grid (slide 11). Only slides 2, 12 and 13 are list-style.
▸Rubric
Real visual elements, not walls of bullets. Mike's question: "does it create good visual elements." Real means: a diagram of the loop, a staged progression, a 50/50 split with illustration, a quote in an organic blob, a comparison built from shapes — something that carries meaning visually. FAIL if the majority of content slides are a title plus a bulleted list.
Q11Readable slidesthis task
Judge's reasoning
No text is clipped or overflowing, and contrast is readable on every slide. The only near-miss is slide 6's fourth card, which sits tight against the right edge but is not cut off.
▸Rubric
Every slide is legible as rendered. FAIL if any slide has text running off the edge, text clipped by its container, text over an image or shape at unreadable contrast, or headline and body colliding. (A separate script check already catches overlapping text boxes; this question covers what the geometry check cannot see — overflow, clipping and contrast in the actual render.)
Q12Does the session make a deliberate choice about what matters most in the guide?this case
Judge's reasoning
The deck makes a deliberate argument: it opens on 'Why sessions forget', builds through 'Where a lesson has to live' to the close 'The repo is the memory', and ends with a concrete 'The next task' exercise rather than walking through the guide's headings in order.
▸Rubric
Does the session make a deliberate choice about what matters most in the guide? PASS when the 30-minute deck selects and connects the guide's most consequential ideas for the audience. FAIL when it is effectively a table of contents—one slide per source heading in source order, without a point of view about what the audience should retain or change.
Notes
1Harness: Cursor CLI 2026.09.15-d2fe57e with model grok-4.7-high (Cursor reports it as Grok 4.7 256K High). It was uploaded as harness 'grok' only because Every Checks has no cursor harness yet. It is not the Grok CLI. Cursor's sandbox cannot keep the model from reading local files, and the owner explicitly waived that for this run. The runner is the Every Checks export with ec-bench's Cursor adapter (16dba0b) ported in. The judge was claude-opus-5.
22 Sep 2026












