Mike's Checks
Checks

Mike's Checks/gpt-5.6-sol/09 pptx

09 pptx

gpt-5.6-solCodex CLIhigh effortrun 22 Aug 2026

Compare models
11/15
checks passed
73%
▸Instructions — the case's current instructions; none were saved with this result

create a powerpoint in the every brand style on compound engineering, 30 min session https://every.to/guides/compound-engineering

A pptx skill is in skills/pptx/ — read skills/pptx/SKILL.md first. pptxgenjs is installed in the workspace, so require("pptxgenjs") works from here with no network. The brand guide is every-visual-style.md.

Answer

_screenshots_slide-01.png

Checks

11/15
Script checks 3/5answered by a program
fail

No overlapping textno-overlap

no two text-bearing shapes on a slide overlap materially (Mike's first question: "Is the writing overlapped?")

fail

Brand colorspalette

share of rendered pixels on Every's documented palette, and whether the deck stays inside one theme (needs Pillow)

pass

Deck opensopens

a .pptx exists and parses (python-pptx when installed, stdlib zip+XML otherwise)

pass

Slide countslide-count

8-20 slides -- sane for a 30-minute session

pass

Slides renderrender

every slide rendered to a PNG in <output-dir>/_screenshots/ via LibreOffice (soffice --convert-to pdf) + pdftoppm. FAIL when required render tools are unavailable.

Judge checks 8/10judged by Claude
fail

Q1Enough substance

Judge's reasoning

The deck carries almost none of the guide's load-bearing substance: it renames the loop's fourth step 'Codify' rather than Compound (slide 5), and never mentions CLAUDE.md, reusable agents, skills or slash commands, the 80/20 planning-and-review split or the 50/50 rule, the beliefs to let go, or the staged 0→5 adoption ladder — AGENTS.md appears once in slide 6's speaker notes and nowhere on a slide, so every specific shown (the CSV-export example, the evidence ladder, the seven-day rollout) is invented rather than sourced.

▸Rubric

The content has real depth. Mike: "the main failing of the decks generated by AI is still the content is quite shallow." The deck must carry the load-bearing substance of the compound-engineering guide — for example the four-step loop (plan → work → review → **compound**) with the compound step named as the thing that separates it from ordinary AI-assisted work, the 80/20 split where planning and review take most of an engineer's time, the specific compounding artifacts (CLAUDE.md, reusable agents, skills, slash commands), the beliefs the philosophy asks you to drop, and the staged 0→5 path. FAIL if the deck could have been written without ever reading the guide — generic "AI makes engineers faster / here are some prompting tips" content, or claims about compound engineering that the guide does not support.

fail

Q2Readable slides

Judge's reasoning

On slide 9 the 'EVIDENCE > CONFIDENCE' pill sits on top of the fifth card and hides its '05' number and most of the word 'Judgment', leaving only '...nt' readable; on slide 4 the gold 'Code is becoming abundant.' overflows into the white text beneath it and the descenders of 'How does the team keep the learning?' collide with the italic caption line.

▸Rubric

Every slide is legible as rendered. FAIL if any slide has text running off the edge, text clipped by its container, text over an image or shape at unreadable contrast, or headline and body colliding. (A separate script check already catches overlapping text boxes; this question covers what the geometry check cannot see — overflow, clipping and contrast in the actual render.)

pass

Q3Every's color palette

Judge's reasoning

Consistent Theme 2 throughout — Teal Dark #0F5258 ground with Teal Medium depth swooshes, Warm White text, and gold/amber/green accents (slides 1–16); the extra coral and lavender appear only as Theme 2's sanctioned multi-color categorization pills, and the deck never flips to Theme 1.

▸Rubric

The palette is Every's. Theme 1 (Editorial): background Every Black `#121212` (or Warm Cream `#F5F2ED` for light slides), Every Blue `#C0F0FB` as the signature accent, white text, with `#FA7B20` / `#E9731E` / `#C4400F` / `#349361` / `#1324CB` / `#7301CC` as accents. Theme 2 (Consulting): background Teal Dark `#0F5258` with Teal Medium `#4999A0` for depth, Warm White `#FFFEFB` text, `#BB7B19` / `#F8DE6E` / `#2E8D23` accents. FAIL if the deck's colors are visibly not these — corporate blue, generic slide-template gray, purple-black AI dark mode, neon gradients — or if it flips between the two themes across slides. Close-but-different shades are fine; a different palette is not.

pass

Q4Serif typography

Judge's reasoning

All headlines and body are Georgia/Times New Roman serif (slides 3, 5, 9, 16); Arial appears only in eyebrow labels, pill text, axis captions and footers, all at 11pt or smaller.

▸Rubric

Typography is serif. The brand is serif-first: Signifier, or its documented PowerPoint substitutes Georgia and Times New Roman. Sans-serif (Switzer / DM Sans) is allowed only for small labels — session indicators, footers, captions. FAIL if headlines or body text are set in a sans-serif face (Arial, Helvetica, Calibri, Inter). "Sans-serif for main content" is on the brand guide's explicit avoid list.

pass

Q5No banned decoration

Judge's reasoning

No triangles or hexagons used as accents, no diagonal slashes, no corner brackets, no drop shadows (zero shadow declarations in the build script), no gradient backgrounds, no rotated or outlined text; borders are single-weight 0.7–3pt hairlines on rounded cards and pills.

▸Rubric

No forbidden decoration. The brand guide bans, in both themes: triangles, hexagons and other geometric accents; diagonal slashes and hard section dividers; corner brackets or L-shaped frames; drop shadows on boxes; heavy or double borders; stock photography; gradient backgrounds; rotated or outlined text. FAIL if any of these appears on any slide.

pass

Q6Required design elements

Judge's reasoning

An organic teal swoosh is layered on every slide via addAtmosphere and reads clearly in the renders (slides 1, 5, 8, 11, 15, 16), joined by classical line illustrations — a robed figure with laptop (1), a Greek temple (6), a scroll and cog (8), and a laurel wreath (16).

▸Rubric

The signature elements are there. Every decks carry organic brushstroke swooshes (the E / O / R SVG shapes, cropped at the slide edge, scaled large) and stippled classical illustrations — Greek and Roman architecture, figures, busts, sometimes with modern elements. The guide says to use the swooshes "liberally once every two or three slides." FAIL if the deck has neither swooshes nor classical illustration, or if such elements appear on fewer than roughly one slide in three. Mike's question on this eval: "does it use the, uh, SVGs?"

pass

Q7Consistent footer

Judge's reasoning

Every slide carries the Theme 2 footer — 'EVERY Consulting' bottom-left, session indicator plus page number bottom-right (slides 1–16); slide 16 swaps the right-hand text to 'compound engineering' but the structure is unbroken.

▸Rubric

The footer is consistent. Theme 1: EVERY logo left, `every.to` right. Theme 2: EVERY *Consulting* left, session indicator right. FAIL if content slides carry no footer, or if the footer appears on some and not others.

pass

Q8Interesting material

Judge's reasoning

It is not a table of contents — the deck builds its own argument arc from premise (slide 3) through the loop (5), a worked 'Add CSV export' example (12), a participant exercise (13), leading-indicator metrics (14) and a rollout (15), with per-slide time budgets rather than restated guide headings.

▸Rubric

It picked interesting things to say. Mike's question on this eval: "does it pick out interesting things to write about." A 30-minute session forces a choice about what matters. FAIL if the deck is a table of contents of the guide — one slide per heading, in the guide's order, with the headings restated and no point of view about what this audience should take away.

pass

Q9One idea per slide

Judge's reasoning

Each slide holds a single message with supporting structure — one comparison on slide 3, one diagram on 5, one 2x2 on 10, one timeline on 15 — and margins stay generous; slide 4 is the busiest but still argues one point (the bottleneck moved from generation to context, verification and memory).

▸Rubric

One idea per slide. Mike: it fails "if it tries to fit too many ideas on one slide." The brand guide says the same thing — "clear focal point: one main message per slide," "spacious, not cluttered." FAIL if any single slide carries more than one idea, or is packed so densely that a room could not read it while the presenter talks.

pass

Q10Useful visuals

Judge's reasoning

Nearly every content slide is a built visual rather than bullets — the four-node flywheel diagram (5), the 50/50 temple split (6), the prompt-vs-spec comparison (7), the staircase evidence ladder (9), the 2x2 codification grid (10), the four-step timeline (15) and the declining-effort chart (14).

▸Rubric

Real visual elements, not walls of bullets. Mike's question: "does it create good visual elements." Real means: a diagram of the loop, a staged progression, a 50/50 split with illustration, a quote in an organic blob, a comparison built from shapes — something that carries meaning visually. FAIL if the majority of content slides are a title plus a bulleted list.