Mike's Checks
Checks

Mike's Checks/claude-fable-5-1/11 hoboken-map

11 hoboken-map

claude-fable-5-1Claude Codehigh effortrun 22 Sep 20264,314,465 tokens

Compare models
6/8
checks passed
75%
▸Instructions — what the model was asked

11 — Neighborhood map + self-validating "hidden gem score"

This is a short multi-turn task. There are no data inputs — do your own
public web research to find real places. Treat the prompts below as a
conversation that arrives in order; the interesting one is the last.

Prompts (in order)

  1. My wife and kids and I just moved to north Hoboken — make me an interactive
    map of family recommendations (playgrounds, food, activities, day trips).
  2. Can you add a date-night category?
  3. What about a tapas place?
  4. Now work on a proprietary "hidden gem" score. Determine a way to measure
    it and validate that it matches expectations. Use public data where possible.

What "good" looks like

  • A single self-contained, interactive map artifact (index.html) a non-technical
    family member could open and use, with the categories above.
  • Real, plausibly-correct Hoboken-area places — not invented ones.
  • For prompt 4: a genuinely thought-through metric. Pre-register a methodology
    and your priors/expectations before collecting and scoring data, then test the
    computed scores against those locked priors and report how well they matched.
    Keep the methodology, the data, the scores, and the validation as inspectable
    files (e.g. a methodology writeup, a data file, a scoring script, a scores file).

Once produced, the artifact must open and work by double-clicking index.html —
no build step, no local server, no API keys, nothing to install. Loading a mapping
library and map tiles from a public CDN is expected and fine; "no API keys" means
no paid or credentialed services, not no network.

Answer

full.png

Checks

6/8
Script checks 3/5answered by a program
fail

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

shared AI-tell phrases: furthermore

fail

Reproducible scoringreproduciblethis task

the scores file is deleted in a scratch copy and the scoring script re-run: it has to finish inside 180s and write the scores back. Missing grader prerequisites fail.

pass

Map loadsrendersthis task

a visible map container of at least 300x200 with >= 5 markers, in a headless browser. Cannot pass without Playwright installed.

pass

Required categoriescategoriesthis task

all six categories the prompt asks for are named somewhere in the output: playgrounds, food, activities, day-trips, date-night, tapas.

pass

Required filesartifactsthis task

the run left all four of a methodology write-up, a data file, a scoring script and a scores file.

Judge checks 3/3judged by Claude
pass

Q1Easy for family to usethis task

Judge's reasoning

The map opens straight from index.html using Leaflet and OSM tiles from a CDN, and it shows labeled category checkboxes, a search box, and a list of named places with addresses, with no raw data visible.

▸Rubric

Usable by a non-technical family member. Look at `map.png` and `full.png`, and at how `index.html` is meant to be opened. The bar, in the reviewer's words: "It's a single self-contained interactive map artifact, uh, that someone could open and use basically by clicking on the HTML." FAIL if using this would require technical knowledge — a build step, a local server, a terminal command, editing a file, or an interface that shows raw data (JSON blobs, coordinate dumps, unlabeled controls) instead of named places and readable categories. PASS if double-clicking the file gives someone a map with labeled places and controls they could figure out unaided.

pass

Q2Criteria fixed before scoringthis task

Judge's reasoning

PRIORS.md sets out GEM/MID/LANDMARK expectations per place plus anchor tests and hash-locks them before data collection, and the validation report tests against them and reports failures (ρ=0.25, below the pre-registered 0.50 bar).

▸Rubric

Priors were locked before scoring. Read the methodology writeup and the scores/validation files. The prompt asked it to "**Pre-register** a methodology and your priors/expectations *before* collecting and scoring data, then test the computed scores against those locked priors and report how well they matched." FAIL if the priors read as written after the results were known — signals: no distinct set of priors stated separately from the findings; priors phrased in terms of the computed scores rather than as expectations about places or patterns; a "validation" section that restates the scores instead of describing a test the priors could have failed. PASS if there is a stated set of expectations that came first and a reported comparison against them, whether they held up or not.

pass

Q3Real placesthis task

Judge's reasoning

The sampled places are all real with plausible addresses and coordinates: Carlo's Bakery at 95 Washington, Lola's Tapas on 14th St, KikiLu Gelato (a hidden gem) on Garden St, Pier A Park, and Turtle Back Zoo in West Orange.

▸Rubric

Places are real. Pick five places from the map — spread across categories, including at least one the output calls a hidden gem — and check them against what you know about Hoboken and the surrounding area. The bar, in the reviewer's words: "real plausibly correct Hoboken area places, not invented ones." FAIL if any of the five is one you can positively identify as fabricated or misplaced — a business that does not exist, a name attached to the wrong town, or coordinates that put it far from where it actually is. If you cannot positively identify a problem with any of the five, PASS.