Mike's Checks
Checks

Mike's Checks/claude-fable-5-1/10 talkform

10 talkform

claude-fable-5-1Claude Codehigh effortrun 22 Sep 20263,718,262 tokens

Compare models
4/8
checks passed
50%
▸Instructions — what the model was asked

build a typeform competitor where the backend is identical to typeform with structured questions and answers, but the frontend is purely an AI voice interview, with the AI interviewer aware of what structured data it needs to fill in for the user, but asking in a more natural conversational way. get it running including auth just ask me for the API keys for the services you need.

Answer

builder.png

Checks

4/8
Script checks 3/5answered by a program
fail

Real API callreally-calls-apithis task

driving the app through its screens produces a real outbound OpenAI request, by any of the five evidence classes below. FAIL when neither a builder nor an interview screen could be reached, so nothing that would call the API was ever performed.

fail

Authentication existsauth-existsthis task

a signup or login screen carrying credential fields is reachable from the landing page or the usual routes. FAIL without a browser.

pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected README.md, ANSWER.md, package-lock.json, package.json, web/index.html, web/package.json, web/tsconfig.json, web/vite.config.ts

pass

App startsbootsthis task

the app installs, starts, and serves HTTP on a port

pass

Screenshots capturedscreenshotsthis task

all three of landing.png, builder.png and interview.png were captured. FAIL without a browser.

Judge checks 1/3judged by Claude
fail

Q1Clear landing page pitchthis task

Judge's reasoning

No landing page exists: '/' is an auth-gated dashboard, and the captured landing text and screenshot show only 'Cannot GET /', so there is no headline and no hook.

▸Rubric

Landing page hook. Read `_screenshots/landing-copy.txt` and look at `landing.png` / `landing-full.png`. The reviewer's bar: "does it do a good job with like landing page copy? ... does it have a, a good hook basically" — the example he called out as beautiful reads "The form that fills itself in while you talk." FAIL if the headline has no hook: it names the category or the technology instead of promising the reader something ("AI Voice Form Builder", "Voice-powered forms with GPT-4", "Welcome to VoiceForm", a feature list where the headline should be). PASS if the headline states, in the reader's terms, what changes for them. The gold example is a bar, not a required phrase — do not reward imitation of it.

fail

Q2Distinctive designthis task

Judge's reasoning

None of the screenshots rendered (every one says 'Cannot GET'), and the source styles the interview screen as a violet (#6d5dfc) glowing orb on a near-black radial gradient, which is the default purple AI look.

▸Rubric

Design is not slop. Look at `landing.png`, `landing-full.png`, `builder.png`, `interview.png`. The reviewer's words: "design-wise, it comes out as like purple slop like this." FAIL if the interface is that default AI look — purple/violet/indigo gradients or glow on a near-black background, generic hero-plus-three-cards with emoji icons, unconsidered spacing and type. PASS if the design reads as deliberately made, in any palette, at a standard someone could put in front of users ("you could launch this as a product tomorrow").

pass

Q3Complete user journeythis task

Judge's reasoning

The builder defines typed, required-flagged Typeform-style fields; the interviewer prompt is built from that form's fields and refs; and a record_answer tool validates each value with coerceAnswer and stores it against its field_ref, so every link in the chain exists in code.

▸Rubric

The loop is coherent. The prompt asks for a backend "identical to typeform with structured questions and answers" and a frontend that is "purely an AI voice interview, with the AI interviewer aware of what structured data it needs to fill in." Look at `builder.png` and `interview.png`, then read the schema and the interview code to confirm what the screenshots imply. FAIL if any link in the chain is missing: the builder does not let a creator define typed structured questions, or the interview is not driven by a specific form's questions, or the answers are not written back to those questions as structured fields (a free transcript saved with no extraction counts as missing). PASS if all three links exist and connect.