Mike's Checks
Checks

Mike's Checks/grok-4.7/11 hoboken-map

11 hoboken-map

grok-4.7Grok CLIhigh effortrun 21 Sep 20264,128,282 tokens

Compare models
8/8
checks passed
100%
▸Instructions — what the model was asked

11 — Neighborhood map + self-validating "hidden gem score"

This is a short multi-turn task. There are no data inputs — do your own
public web research to find real places. Treat the prompts below as a
conversation that arrives in order; the interesting one is the last.

Prompts (in order)

  1. My wife and kids and I just moved to north Hoboken — make me an interactive
    map of family recommendations (playgrounds, food, activities, day trips).
  2. Can you add a date-night category?
  3. What about a tapas place?
  4. Now work on a proprietary "hidden gem" score. Determine a way to measure
    it and validate that it matches expectations. Use public data where possible.

What "good" looks like

  • A single self-contained, interactive map artifact (index.html) a non-technical
    family member could open and use, with the categories above.
  • Real, plausibly-correct Hoboken-area places — not invented ones.
  • For prompt 4: a genuinely thought-through metric. Pre-register a methodology
    and your priors/expectations before collecting and scoring data, then test the
    computed scores against those locked priors and report how well they matched.
    Keep the methodology, the data, the scores, and the validation as inspectable
    files (e.g. a methodology writeup, a data file, a scoring script, a scores file).

Once produced, the artifact must open and work by double-clicking index.html —
no build step, no local server, no API keys, nothing to install. Loading a mapping
library and map tiles from a public CDN is expected and fine; "no API keys" means
no paid or credentialed services, not no network.

Answer

full.png

Checks

8/8
Script checks 5/5answered by a program
pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected index.html, scores.json, VALIDATION.md, METHODOLOGY.md, LOCK.txt, data.json, ANSWER.md, PRIORS.json

pass

Map loadsrendersthis task

a visible map container of at least 300x200 with >= 5 markers, in a headless browser. Cannot pass without Playwright installed.

pass

Required categoriescategoriesthis task

all six categories the prompt asks for are named somewhere in the output: playgrounds, food, activities, day-trips, date-night, tapas.

pass

Required filesartifactsthis task

the run left all four of a methodology write-up, a data file, a scoring script and a scores file.

pass

Reproducible scoringreproduciblethis task

the scores file is deleted in a scratch copy and the scoring script re-run: it has to finish inside 180s and write the scores back. Missing grader prerequisites fail.

Judge checks 3/3judged by Claude
pass

Q1Easy for family to usethis task

Judge's reasoning

The map opens straight from index.html using Leaflet and OSM tiles, and it shows named places in a readable list, labeled category chips, a search box and sort buttons, all of which a family member could figure out without help.

▸Rubric

Usable by a non-technical family member. Look at `map.png` and `full.png`, and at how `index.html` is meant to be opened. The bar, in the reviewer's words: "It's a single self-contained interactive map artifact, uh, that someone could open and use basically by clicking on the HTML." FAIL if using this would require technical knowledge — a build step, a local server, a terminal command, editing a file, or an interface that shows raw data (JSON blobs, coordinate dumps, unlabeled controls) instead of named places and readable categories. PASS if double-clicking the file gives someone a map with labeled places and controls they could figure out unaided.

pass

Q2Criteria fixed before scoringthis task

Judge's reasoning

METHODOLOGY.md, PRIORS.json and LOCK.txt set out fame-based tier priors and four pass/fail tests, with hashes recorded before collection, and VALIDATION.md reports honestly that the test failed (2 of 4 checks passed).

▸Rubric

Priors were locked before scoring. Read the methodology writeup and the scores/validation files. The prompt asked it to "**Pre-register** a methodology and your priors/expectations *before* collecting and scoring data, then test the computed scores against those locked priors and report how well they matched." FAIL if the priors read as written after the results were known — signals: no distinct set of priors stated separately from the findings; priors phrased in terms of the computed scores rather than as expectations about places or patterns; a "validation" section that restates the scores instead of describing a test the priors could have failed. PASS if there is a stated set of expectations that came first and a reported comparison against them, whether they held up or not.

pass

Q3Real placesthis task

Judge's reasoning

Lola's Tapas Bar (153 14th St), Bin 14 (1314 Washington), Hoboken Cove Community Boathouse (a hidden gem), Carlo's Bakery (95 Washington) and Liberty Science Center are all real places with correct addresses and plausible coordinates.

▸Rubric

Places are real. Pick five places from the map — spread across categories, including at least one the output calls a hidden gem — and check them against what you know about Hoboken and the surrounding area. The bar, in the reviewer's words: "real plausibly correct Hoboken area places, not invented ones." FAIL if any of the five is one you can positively identify as fabricated or misplaced — a business that does not exist, a name attached to the wrong town, or coordinates that put it far from where it actually is. If you cannot positively identify a problem with any of the five, PASS.

Notes

1

Harness: Cursor CLI 2026.09.15-d2fe57e with model grok-4.7-high (Cursor reports it as Grok 4.7 256K High). It was uploaded as harness 'grok' only because Every Checks has no cursor harness yet. It is not the Grok CLI. Cursor's sandbox cannot keep the model from reading local files, and the owner explicitly waived that for this run. The runner is the Every Checks export with ec-bench's Cursor adapter (16dba0b) ported in. The judge was claude-opus-5.

22 Sep 2026