Mike's Checks
Checks

Mike's Checks/gpt-6-astra/11 hoboken-map

11 hoboken-map

gpt-6-astraCodex CLIhigh effortrun 22 Sep 20262,235,334 tokens

Compare models
7/8
checks passed
88%
▸Instructions — what the model was asked

11 — Neighborhood map + self-validating "hidden gem score"

This is a short multi-turn task. There are no data inputs — do your own
public web research to find real places. Treat the prompts below as a
conversation that arrives in order; the interesting one is the last.

Prompts (in order)

  1. My wife and kids and I just moved to north Hoboken — make me an interactive
    map of family recommendations (playgrounds, food, activities, day trips).
  2. Can you add a date-night category?
  3. What about a tapas place?
  4. Now work on a proprietary "hidden gem" score. Determine a way to measure
    it and validate that it matches expectations. Use public data where possible.

What "good" looks like

  • A single self-contained, interactive map artifact (index.html) a non-technical
    family member could open and use, with the categories above.
  • Real, plausibly-correct Hoboken-area places — not invented ones.
  • For prompt 4: a genuinely thought-through metric. Pre-register a methodology
    and your priors/expectations before collecting and scoring data, then test the
    computed scores against those locked priors and report how well they matched.
    Keep the methodology, the data, the scores, and the validation as inspectable
    files (e.g. a methodology writeup, a data file, a scoring script, a scores file).

Once produced, the artifact must open and work by double-clicking index.html —
no build step, no local server, no API keys, nothing to install. Loading a mapping
library and map tiles from a public CDN is expected and fine; "no API keys" means
no paid or credentialed services, not no network.

Answer

full.png

Checks

7/8
Script checks 4/5answered by a program
fail

Reproducible scoringreproduciblethis task

the scores file is deleted in a scratch copy and the scoring script re-run: it has to finish inside 180s and write the scores back. Missing grader prerequisites fail.

pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected index.html, scores.json, SOURCES.md, map-template.html, VALIDATION.md, places.json, METHODOLOGY.md, geocoding.json

pass

Map loadsrendersthis task

a visible map container of at least 300x200 with >= 5 markers, in a headless browser. Cannot pass without Playwright installed.

pass

Required categoriescategoriesthis task

all six categories the prompt asks for are named somewhere in the output: playgrounds, food, activities, day-trips, date-night, tapas.

pass

Required filesartifactsthis task

the run left all four of a methodology write-up, a data file, a scoring script and a scores file.

Judge checks 3/3judged by Claude
pass

Q1Easy for family to usethis task

Judge's reasoning

index.html is self-contained (data inlined, Leaflet/tiles from public CDN) and the screenshots show a working map with named places, readable category filters, search, and sort — double-clickable with no build step or server.

▸Rubric

Usable by a non-technical family member. Look at `map.png` and `full.png`, and at how `index.html` is meant to be opened. The bar, in the reviewer's words: "It's a single self-contained interactive map artifact, uh, that someone could open and use basically by clicking on the HTML." FAIL if using this would require technical knowledge — a build step, a local server, a terminal command, editing a file, or an interface that shows raw data (JSON blobs, coordinate dumps, unlabeled controls) instead of named places and readable categories. PASS if double-clicking the file gives someone a map with labeled places and controls they could figure out unaided.

pass

Q2Criteria fixed before scoringthis task

Judge's reasoning

preregistration.json/METHODOLOGY.md lock a timestamped, hashed formula plus three prior groups and six ordered pairwise predictions with a 4/6 target stated before data, and VALIDATION.md honestly reports the target was missed (0/1 evaluable passes, 5 unevaluable) rather than restating the scores.

▸Rubric

Priors were locked before scoring. Read the methodology writeup and the scores/validation files. The prompt asked it to "**Pre-register** a methodology and your priors/expectations *before* collecting and scoring data, then test the computed scores against those locked priors and report how well they matched." FAIL if the priors read as written after the results were known — signals: no distinct set of priors stated separately from the findings; priors phrased in terms of the computed scores rather than as expectations about places or patterns; a "validation" section that restates the scores instead of describing a test the priors could have failed. PASS if there is a stated set of expectations that came first and a reported comparison against them, whether they held up or not.

pass

Q3Real placesthis task

Judge's reasoning

Spot-checked ResilienCity Park (1201 Madison), Elysian Park, Fiore's House of Quality (414 Adams, the top-scoring 'hidden gem'), Lola's Tapas at 153 14th St, and Turtle Back Zoo in West Orange — all real, correctly-towned, and plausibly geocoded.

▸Rubric

Places are real. Pick five places from the map — spread across categories, including at least one the output calls a hidden gem — and check them against what you know about Hoboken and the surrounding area. The bar, in the reviewer's words: "real plausibly correct Hoboken area places, not invented ones." FAIL if any of the five is one you can positively identify as fabricated or misplaced — a business that does not exist, a name attached to the wrong town, or coordinates that put it far from where it actually is. If you cannot positively identify a problem with any of the five, PASS.