Mike's Checks/grok-4.7/11 hoboken-map
11 hoboken-map
grok-4.7Grok CLIhigh effortrun 21 Sep 20264,128,282 tokens
▸Instructions — what the model was asked
11 — Neighborhood map + self-validating "hidden gem score"
This is a short multi-turn task. There are no data inputs — do your own
public web research to find real places. Treat the prompts below as a
conversation that arrives in order; the interesting one is the last.
Prompts (in order)
- My wife and kids and I just moved to north Hoboken — make me an interactive
map of family recommendations (playgrounds, food, activities, day trips). - Can you add a date-night category?
- What about a tapas place?
- Now work on a proprietary "hidden gem" score. Determine a way to measure
it and validate that it matches expectations. Use public data where possible.
What "good" looks like
- A single self-contained, interactive map artifact (
index.html) a non-technical
family member could open and use, with the categories above. - Real, plausibly-correct Hoboken-area places — not invented ones.
- For prompt 4: a genuinely thought-through metric. Pre-register a methodology
and your priors/expectations before collecting and scoring data, then test the
computed scores against those locked priors and report how well they matched.
Keep the methodology, the data, the scores, and the validation as inspectable
files (e.g. a methodology writeup, a data file, a scoring script, a scores file).
Once produced, the artifact must open and work by double-clicking index.html —
no build step, no local server, no API keys, nothing to install. Loading a mapping
library and map tiles from a public CDN is expected and fine; "no API keys" means
no paid or credentialed services, not no network.
Answer
Checks
8/8No AI-tell wordingai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected index.html, scores.json, VALIDATION.md, METHODOLOGY.md, LOCK.txt, data.json, ANSWER.md, PRIORS.json
Map loadsrendersthis task
a visible map container of at least 300x200 with >= 5 markers, in a headless browser. Cannot pass without Playwright installed.
Required categoriescategoriesthis task
all six categories the prompt asks for are named somewhere in the output: playgrounds, food, activities, day-trips, date-night, tapas.
Required filesartifactsthis task
the run left all four of a methodology write-up, a data file, a scoring script and a scores file.
Reproducible scoringreproduciblethis task
the scores file is deleted in a scratch copy and the scoring script re-run: it has to finish inside 180s and write the scores back. Missing grader prerequisites fail.
Q1Easy for family to usethis task
Judge's reasoning
The map opens straight from index.html using Leaflet and OSM tiles, and it shows named places in a readable list, labeled category chips, a search box and sort buttons, all of which a family member could figure out without help.
▸Rubric
Usable by a non-technical family member. Look at `map.png` and `full.png`, and at how `index.html` is meant to be opened. The bar, in the reviewer's words: "It's a single self-contained interactive map artifact, uh, that someone could open and use basically by clicking on the HTML." FAIL if using this would require technical knowledge — a build step, a local server, a terminal command, editing a file, or an interface that shows raw data (JSON blobs, coordinate dumps, unlabeled controls) instead of named places and readable categories. PASS if double-clicking the file gives someone a map with labeled places and controls they could figure out unaided.
Q2Criteria fixed before scoringthis task
Judge's reasoning
METHODOLOGY.md, PRIORS.json and LOCK.txt set out fame-based tier priors and four pass/fail tests, with hashes recorded before collection, and VALIDATION.md reports honestly that the test failed (2 of 4 checks passed).
▸Rubric
Priors were locked before scoring. Read the methodology writeup and the scores/validation files. The prompt asked it to "**Pre-register** a methodology and your priors/expectations *before* collecting and scoring data, then test the computed scores against those locked priors and report how well they matched." FAIL if the priors read as written after the results were known — signals: no distinct set of priors stated separately from the findings; priors phrased in terms of the computed scores rather than as expectations about places or patterns; a "validation" section that restates the scores instead of describing a test the priors could have failed. PASS if there is a stated set of expectations that came first and a reported comparison against them, whether they held up or not.
Q3Real placesthis task
Judge's reasoning
Lola's Tapas Bar (153 14th St), Bin 14 (1314 Washington), Hoboken Cove Community Boathouse (a hidden gem), Carlo's Bakery (95 Washington) and Liberty Science Center are all real places with correct addresses and plausible coordinates.
▸Rubric
Places are real. Pick five places from the map — spread across categories, including at least one the output calls a hidden gem — and check them against what you know about Hoboken and the surrounding area. The bar, in the reviewer's words: "real plausibly correct Hoboken area places, not invented ones." FAIL if any of the five is one you can positively identify as fabricated or misplaced — a business that does not exist, a name attached to the wrong town, or coordinates that put it far from where it actually is. If you cannot positively identify a problem with any of the five, PASS.
Notes
1Harness: Cursor CLI 2026.09.15-d2fe57e with model grok-4.7-high (Cursor reports it as Grok 4.7 256K High). It was uploaded as harness 'grok' only because Every Checks has no cursor harness yet. It is not the Grok CLI. Cursor's sandbox cannot keep the model from reading local files, and the owner explicitly waived that for this run. The runner is the Every Checks export with ec-bench's Cursor adapter (16dba0b) ported in. The judge was claude-opus-5.
22 Sep 2026

