Mike's Checks/gpt-6-astra/09 pptx/compound engineering
09 pptx · compound engineering
gpt-6-astraCodex CLIhigh effortrun 22 Sep 2026951,129 tokens
▸Instructions — what the model was asked
create a powerpoint in the every brand style on compound engineering, 30 min session https://every.to/guides/compound-engineering
A pptx skill is in skills/pptx/ — read skills/pptx/SKILL.md first. pptxgenjs is installed in the workspace, so require("pptxgenjs") works from here with no network. The brand guide is every-visual-style.md.
Answer
Checks
15/18Brand colorspalettethis task
share of rendered pixels on Every's documented palette, and whether the deck stays inside one theme (needs Pillow)
No AI-tell wordingai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected FACILITATOR.md, QA.md, build_deck.js, ANSWER.md, compound-engineering-every.pptx, every-visual-style.md, qa/slide-text.txt, qa/structure-check.txt
Deck opensopensthis task
a .pptx exists and parses (python-pptx when installed, stdlib zip+XML otherwise)
Slide countslide-countthis task
8-20 slides -- sane for a 30-minute session
No overlapping textno-overlapthis task
no two text-bearing shapes on a slide overlap materially (Mike's first question: "Is the writing overlapped?")
Slides renderrenderthis task
every slide rendered to a PNG in <output-dir>/_screenshots/ via LibreOffice (soffice --convert-to pdf) + pdftoppm. FAIL when required render tools are unavailable.
Q1Enough substancethis task
Judge's reasoning
Slide 18 admits 'Guide access was unavailable... independent synthesis'; the deck misses the guide's real substance (80/20 plan and review, parallel agents, plugin commands, adoption stages) and stays generic.
▸Rubric
The content has real depth. Mike: "the main failing of the decks generated by AI is still the content is quite shallow." PASS when the deck carries the load-bearing substance of the selected case's brief and source material: its important mechanisms, distinctions, evidence, and consequences. FAIL if the deck could have been written without reading those materials, reduces the subject to generic advice, or makes unsupported claims.
Q2Does the deck carry the guide's load-bearing compound-engineering substance?this case
Judge's reasoning
The deck does cover Plan → Work → Review → Compound and treats compounding as the differentiator, but it never mentions CLAUDE.md, reusable agents, skills or commands (lessons are only generic 'notes'), and it leaves out the 80/20 shift and the staged adoption path; slide 18 admits 'Guide access was unavailable. Framework: independent synthesis.'
▸Rubric
Does the deck carry the guide's load-bearing compound-engineering substance? PASS when it explains the plan → work → review → compound loop, makes the compound step the differentiator from ordinary AI-assisted work, and uses concrete artifacts such as CLAUDE.md, reusable agents, skills, or commands. It should also convey the guide's 80/20 shift and staged adoption path. FAIL when it substitutes generic AI productivity or prompting advice for those mechanics.
Q3Every's color palettethis task
Judge's reasoning
Theme 2 teal (#0F5258-ish) with warm-white text and gold #F8DE6E accents throughout; light slides 6/10/14 use warm white with teal text, which stays within the consulting palette.
▸Rubric
The palette is Every's. Theme 1 (Editorial): background Every Black `#121212` (or Warm Cream `#F5F2ED` for light slides), Every Blue `#C0F0FB` as the signature accent, white text, with `#FA7B20` / `#E9731E` / `#C4400F` / `#349361` / `#1324CB` / `#7301CC` as accents. Theme 2 (Consulting): background Teal Dark `#0F5258` with Teal Medium `#4999A0` for depth, Warm White `#FFFEFB` text, `#BB7B19` / `#F8DE6E` / `#2E8D23` accents. FAIL if the deck's colors are visibly not these — corporate blue, generic slide-template gray, purple-black AI dark mode, neon gradients — or if it flips between the two themes across slides. Close-but-different shades are fine; a different palette is not.
Q4Serif typographythis task
Judge's reasoning
Headlines and body are Georgia-style serif on every slide; sans only appears in pills, footers and small captions (e.g. slides 6, 7, 16).
▸Rubric
Typography is serif. The brand is serif-first: Signifier, or its documented PowerPoint substitutes Georgia and Times New Roman. Sans-serif (Switzer / DM Sans) is allowed only for small labels — session indicators, footers, captions. FAIL if headlines or body text are set in a sans-serif face (Arial, Helvetica, Calibri, Inter). "Sans-serif for main content" is on the brand guide's explicit avoid list.
Q5No banned decorationthis task
Judge's reasoning
No geometric accents, slashes, brackets, shadows or gradients; thin rounded cards and pill borders match the guide, and the temple pediment belongs to the classical illustration.
▸Rubric
No forbidden decoration. The brand guide bans, in both themes: triangles, hexagons and other geometric accents; diagonal slashes and hard section dividers; corner brackets or L-shaped frames; drop shadows on boxes; heavy or double borders; stock photography; gradient backgrounds; rotated or outlined text. FAIL if any of these appears on any slide.
Q6Required design elementsthis task
Judge's reasoning
A classical temple illustration appears on slides 1, 9 and 17, and faint edge-cropped brushstroke swooshes on slides 1, 4, 7, 13 and 16, so roughly 7 of 18 slides.
▸Rubric
The signature elements are there. Every decks carry organic brushstroke swooshes (the E / O / R SVG shapes, cropped at the slide edge, scaled large) and stippled classical illustrations — Greek and Roman architecture, figures, busts, sometimes with modern elements. The guide says to use the swooshes "liberally once every two or three slides." FAIL if the deck has neither swooshes nor classical illustration, or if such elements appear on fewer than roughly one slide in three. Mike's question on this eval: "does it use the, uh, SVGs?"
Q7Consistent footerthis task
Judge's reasoning
Every slide carries 'EVERY Consulting' at left and 'Compound engineering / NN' at right.
▸Rubric
The footer is consistent. Theme 1: EVERY logo left, `every.to` right. Theme 2: EVERY *Consulting* left, session indicator right. FAIL if content slides carry no footer, or if the footer appears on some and not others.
Q8Interesting materialthis task
Judge's reasoning
Makes a coherent argument that the output is a capability, not just code (slide 3), threads an invoice-export example through slides 6-11, and covers failure modes (slide 12) plus a human/AI split (slide 13).
▸Rubric
It picked interesting things to say. Mike's question on this eval: "does it pick out interesting things to write about." PASS when the deck makes a selective, coherent argument about what this audience should remember or do. FAIL if it merely inventories the supplied material in source order, restates headings, or includes facts without a point of view about why they matter.
Q9One idea per slidethis task
Judge's reasoning
Each slide carries one message with generous whitespace, e.g. slides 4, 9 and 15; no slide is overpacked.
▸Rubric
One idea per slide. Mike: it fails "if it tries to fit too many ideas on one slide." The brand guide says the same thing — "clear focal point: one main message per slide," "spacious, not cluttered." FAIL if any single slide carries more than one idea, or is packed so densely that a room could not read it while the presenter talks.
Q10Useful visualsthis task
Judge's reasoning
Visuals include a loop diagram (4), comparison cards (3, 13), numbered progressions (5, 11, 14), a 50/50 split with illustration (9, 17) and a note card (10), not bullet walls.
▸Rubric
Real visual elements, not walls of bullets. Mike's question: "does it create good visual elements." Real means: a diagram of the loop, a staged progression, a 50/50 split with illustration, a quote in an organic blob, a comparison built from shapes — something that carries meaning visually. FAIL if the majority of content slides are a title plus a bulleted list.
Q11Readable slidesthis task
Judge's reasoning
No overflow, clipping or low-contrast text in any render; the text in slides 11, 12 and 16 wraps cleanly inside its space.
▸Rubric
Every slide is legible as rendered. FAIL if any slide has text running off the edge, text clipped by its container, text over an image or shape at unreadable contrast, or headline and body colliding. (A separate script check already catches overlapping text boxes; this question covers what the geometry check cannot see — overflow, clipping and contrast in the actual render.)
Q12Does the session make a deliberate choice about what matters most in the guide?this case
Judge's reasoning
Rather than walking through source headings in order, the deck builds a clear point of view around one worked example (filtered invoice export), stresses that retrieval completes the cycle (slide 11), and ends with an exercise and a one-task-this-week rollout aimed at what the audience should retain.
▸Rubric
Does the session make a deliberate choice about what matters most in the guide? PASS when the 30-minute deck selects and connects the guide's most consequential ideas for the audience. FAIL when it is effectively a table of contents—one slide per source heading in source order, without a point of view about what the audience should retain or change.

















