Mike's Checks
Checks

Mike's Checks/claude-opus-5.5/10 talkform

10 talkform

claude-opus-5.5Claude Codehigh effortrun 20 Sep 20265,923,110 tokens

Compare models
7/8
checks passed
88%
▸Instructions — what the model was asked

build a typeform competitor where the backend is identical to typeform with structured questions and answers, but the frontend is purely an AI voice interview, with the AI interviewer aware of what structured data it needs to fill in for the user, but asking in a more natural conversational way. get it running including auth just ask me for the API keys for the services you need.

Answer

auth-after.png

Checks

7/8
Script checks 4/5answered by a program
fail

Real API callreally-calls-apithis task

driving the app through its screens produces a real outbound OpenAI request, by any of the five evidence classes below. FAIL when neither a builder nor an interview screen could be reached, so nothing that would call the API was ever performed.

not exercised — never authenticated, so the builder and interview stayed out of reach; no route in its source sits next to a model call

pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected README.md, ANSWER.md, package-lock.json, package.json, test/llm.test.js, test/api.test.js, test/units.test.js, test/helpers.js

pass

App startsbootsthis task

the app installs, starts, and serves HTTP on a port

pass

Authentication existsauth-existsthis task

a signup or login screen carrying credential fields is reachable from the landing page or the usual routes. FAIL without a browser.

pass

Screenshots capturedscreenshotsthis task

all three of landing.png, builder.png and interview.png were captured. FAIL without a browser.

Judge checks 3/3judged by Claude
pass

Q1Clear landing page pitchthis task

Judge's reasoning

"Your form, as a conversation." states what changes for the reader rather than naming the category or the tech, with a supporting subhead about respondents just talking while fields fill in.

▸Rubric

Landing page hook. Read `_screenshots/landing-copy.txt` and look at `landing.png` / `landing-full.png`. The reviewer's bar: "does it do a good job with like landing page copy? ... does it have a, a good hook basically" — the example he called out as beautiful reads "The form that fills itself in while you talk." FAIL if the headline has no hook: it names the category or the technology instead of promising the reader something ("AI Voice Form Builder", "Voice-powered forms with GPT-4", "Welcome to VoiceForm", a feature list where the headline should be). PASS if the headline states, in the reader's terms, what changes for them. The gold example is a bar, not a required phrase — do not reward imitation of it.

pass

Q2Distinctive designthis task

Judge's reasoning

The visible screens use a deliberate warm off-white canvas with strong Swiss-style typography and a single restrained violet accent — no dark purple-gradient hero, emoji cards, or slop layout.

▸Rubric

Design is not slop. Look at `landing.png`, `landing-full.png`, `builder.png`, `interview.png`. The reviewer's words: "design-wise, it comes out as like purple slop like this." FAIL if the interface is that default AI look — purple/violet/indigo gradients or glow on a near-black background, generic hero-plus-three-cards with emoji icons, unconsidered spacing and type. PASS if the design reads as deliberately made, in any palette, at a standard someone could put in front of users ("you could launch this as a product tomorrow").

pass

Q3Complete user journeythis task

Judge's reasoning

Despite both walk screenshots landing on wrong routes (a 404 JSON and a nonexistent interview slug), the code shows a full typed field builder (public/assets/app.js), a per-form interviewer prompt built from that form's fields (server/interview.js), and record_answers/skip_field tools coercing and writing values back into typed Typeform answer objects on the response.

▸Rubric

The loop is coherent. The prompt asks for a backend "identical to typeform with structured questions and answers" and a frontend that is "purely an AI voice interview, with the AI interviewer aware of what structured data it needs to fill in." Look at `builder.png` and `interview.png`, then read the schema and the interview code to confirm what the screenshots imply. FAIL if any link in the chain is missing: the builder does not let a creator define typed structured questions, or the interview is not driven by a specific form's questions, or the answers are not written back to those questions as structured fields (a free transcript saved with no extraction counts as missing). PASS if all three links exist and connect.

Notes

1

Alright, for the Talkform one, I did a pretty good job. It's a little bit boring of a design. The auth kind of just looks like a Typeform thing. It didn't finish the build. There are actually some errors in the screenshots. It didn't run the app fully. I guess, or if it did, it didn't get the builder or the interview module working. So the only thing I'm really basing this off is a very plain boring landing page. It couldn't actually build out. It couldn't actually make the real API calls. So that's, even though that's only really one thing, it got wrong, that's the big thing that it got wrong. So yeah, I would say this is a fail on the building.

21 Sep 2026