Mike's Checks/gpt-6-astra/01 dashboard
01 dashboard
gpt-6-astraCodex CLIhigh effortrun 22 Sep 2026265,942 tokens
▸Instructions — what the model was asked
I want you to build a static html NPS dashboard and pick out interesting insights after doing analysis of the comments
Answer
Checks
8/10No AI-tell wordingai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
shared AI-tell phrases: the world of
Page loadsrendersthis task
the dashboard loads in a headless browser with no console error and paints at least 40 characters of text. Cannot pass without Playwright installed.
Correct NPSnps-correctthis task
the page says "NPS" or "net promoter", and the NPS computed from the source CSV appears on it within 0.15. Missing source data fails closed.
Score breakdowndistribution-shownthis task
"promoter", "passive" and "detractor" all appear, and all three groups carry either their raw count (within 0.01) or their share (within 0.2 of a percentage point). Missing source data fails closed.
Working controlsinteractivethis task
driving a real control -- text/search input, select, button or tab -- changes what the page displays. FAIL when the page could not be driven at all.
Q1Useful headlinethis task
Judge's reasoning
"A strong signal. A clearer opportunity." is a generic slogan that could sit on almost any positive NPS dashboard and says nothing specific about what this survey found.
▸Rubric
The headline earns its place. Look at `fold.png` and read the dashboard's own top-level title/headline — the line the page leads with, not the browser tab title and not a section heading. This asks two things of one line, and FAILs if either is missing. It must (a) say what this particular data showed, and (b) be worth reading. FAIL if the headline is a bare label that would sit unchanged on any NPS dashboard for any product — "NPS Dashboard", "Every NPS Results", "Customer Feedback Analysis", "Survey Results Overview" — or if it is a naked metric restated as a title. FAIL also if it reaches for interest at the cost of accuracy: a claim the data on the page does not support. PASS if someone who read only that line would learn something true and specific about this survey, and would want to look at the rest. The reviewer's words: *"is the headline a good concise summary of the content while being interesting?"*
Q2Distinctive designthis task
Judge's reasoning
The design is light, with cream, forest-green and terracotta colours; it has no purple-on-black dark-mode styling.
▸Rubric
AI-slop design. Look at `fold.png` and `full.png`. FAIL if the dashboard looks like generic AI-generated design — the reviewer's words: "we should definitely weigh it down if it's purple-black dark mode," and "at the very least, like it shouldn't look like a Vibe Slop, like purple dark mode." The specific tell is a purple-on-black dark-mode treatment (purple/violet/indigo gradients, glow, neon accents on a near-black background). PASS anything that reads as a considered design, light or dark, that isn't that.
Q3Readable typographythis task
Judge's reasoning
A chosen serif display face is paired with a clean sans; the text has good contrast, and the numbers and comments are easy to read.
▸Rubric
Typography. Look at `fold.png` and `full.png`. FAIL if the type is unconsidered or hard to read — the reviewer's words: "the typography. Like did it pick like good fonts? ... Is, is it like easy to read?" Concrete failure evidence: raw browser-default fonts, low-contrast text, cramped or overflowing/clipped text, sizes so small the numbers and comments are hard to read. PASS if fonts look deliberately chosen and the page is comfortable to read.
Q4Useful findingsthis task
Judge's reasoning
The insights name patterns from the comments that a reader could act on: practical-AI comments go with an NPS of +67 versus +36 for everyone else, readers see the bundle but use it less, and platform limits cause friction before signup.
▸Rubric
Decision-useful insight. Read the insights section(s) of the dashboard. FAIL if every stated insight is arithmetic restated — counts, percentages, averages, or "X% are promoters" phrased as a finding — with nothing that would help someone decide. The bar, in the reviewer's words: "would this actually be useful for me if I was ... making a decision about whether Every is a good product or not." PASS if at least one insight names a pattern in the comments (a theme, a segment, a tension between what promoters and detractors say) that a reader could act on and that is not derivable from the score distribution alone.
Q5Key facts firstthis task
Judge's reasoning
The NPS of +45.1 and the promoter, passive and detractor cards with counts and percentages are all visible in fold.png.
▸Rubric
Key information above the fold. Look at `fold.png` only. FAIL if the headline NPS number and the promoter/passive/detractor breakdown are not both visible without scrolling. The reviewer's words: "does it, does it like put the important information above the fold? ... That's the main thing for me."


