Mike's Checks
Checks

Mike's Checks/grok-4.7/02 hotel

02 hotel

grok-4.7Grok CLIhigh effortrun 21 Sep 2026

Compare models
0/10
checks passed
0%
▸Instructions — what the model was asked

I'm delivering a claude code workshop in Miami. I am working out of a friend's office in South Beach (Miami Beach). The actual client workshop is in Wynwood. I arrive at 2pm on Thursday and immediately have calls, and the workshop is the next day in the morning. We will probably want to find somewhere to meet the client for drinks / food that night before the workshop.

Start your recommendation with these two lines, exactly in this format, before
anything else:

Hotel pick: <name of the hotel>
Restaurant pick: <name of the restaurant or bar>

Answer

The model left no answer file or screenshots.

Checks

0/10
Script checks 0/4answered by a program
fail

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

no readable deliverable text to inspect

fail

Recommendation existsexiststhis task

a recommendation document of >= 30 words was produced

fail

Length limitnot-verbosethis task

the recommendation documents come to <= 1,000 words in total. This is a trip plan, not an essay.

fail

Names both placesconcretethis task

both lines the prompt requires verbatim are present and name something -- "Hotel pick: X" and "Restaurant pick: Y" -- rather than being empty or a placeholder.

Judge checks 0/6judged by Claude
fail

Q1Reasonable hotel pricethis task

Judge's reasoning

No output document was produced (work dir contains only prompt.md), so no hotel is named and the rule makes Q1–Q3 FAIL.

▸Rubric

Hotel is not too expensive. FAIL if the recommended hotel is a luxury, five-star or marquee resort property (the Four Seasons / Ritz-Carlton / Faena tier), or if the output frames it as a splurge. Client: *"it's not the most expensive hotel where like Ariel would be upset with me."*

fail

Q2Decent hotelthis task

Judge's reasoning

No output was produced, so there is no hotel pick to judge.

▸Rubric

Hotel is not too cheap. FAIL if the recommended hotel is a hostel, motel, no-frills budget property, or is framed as the cheap option. Client: *"but it's, but it's also not a shit hotel… Uh, whereas I'd, I'd be upset with it."*

fail

Q3Hotel characterthis task

Judge's reasoning

No output was produced, so there is no hotel pick to judge.

▸Rubric

Hotel has taste. FAIL if the hotel pick is an interchangeable corporate chain with no character (the Hilton Garden Inn / Courtyard / Holiday Inn tier), even when the price band is right. The target: *"middle of the road, like a little bit, a little bit higher than middle… and also whether there's like some taste to the hotel."* CitizenM is **an example** of a hotel that clears this bar, not the required answer: *"it's not like amazing, like five-star hotel, but it's also — it's much trendier and nicer than like the average hotel."* Do not reward naming CitizenM and do not penalise not naming it. Judge the property that was picked.

fail

Q4Suitable client dinnerthis task

Judge's reasoning

No output was produced, so there is no 'Restaurant pick:' line or drinks/dinner venue.

▸Rubric

Venue has taste and is client-worthy. Judge the drinks/dinner pick **only** — this question and question 3 are deliberately separate, because a run can get the hotel right and the venue wrong, or the reverse. Score each on its own evidence and do not let a verdict on one carry into the other. FAIL if reading the drinks/dinner pick would prompt the reaction *"Okay, this is like fine, but it's like not like that impressive for like a client dinner"* — a chain restaurant, the hotel's own lobby bar as a default, a generic sports bar, or a pick with nothing about it that would land with a client. PASS if it is a place with real atmosphere or reputation, chosen for this dinner.

fail

Q5Fits the schedulethis task

Judge's reasoning

No output was produced, so there is no plan to check against the schedule.

▸Rubric

Logistics are coherent with the stated schedule. The schedule is fixed: arrives 2pm Thursday, calls immediately, works out of the South Beach office, drinks/food with the client Thursday evening, workshop in Wynwood Friday morning. FAIL only if the plan contradicts it — e.g. dinner booked inside the Thursday call block, hotel check-in scheduled after the client dinner, or a Friday plan that cannot reach the Wynwood workshop in the morning. Do **not** fail this for choosing South Beach over Wynwood or the reverse, and do not fail it over travel-time estimates.

fail

Q6No unnecessary extrasthis task

Judge's reasoning

No output was produced, so nothing reads like a recommendation a friend or assistant would give.

▸Rubric

It is not trying too hard. The bar, in the reviewer's words: it should read like *"something a friend or a human assistant would give you — just the important details"*. FAIL if the output pads with work nobody asked for: an assumptions list, Uber fare calculations, minute-by-minute travel math, weather, contingency trees, or research like room-noise from nearby construction. The canonical failure: *"It calculated the travel time… it like figured out how much dollars it would cost for an Uber… it did like basically like way, way more work than it needed to,"* and *"This hotel has construction, and so it might be noisy on the south side, so you wanna choose the room on the north side"* — *"in some ways that's cool, but it's also like way, way too much."*

Notes

1

Harness: Cursor CLI 2026.09.15-d2fe57e with model grok-4.7-high (Cursor reports it as Grok 4.7 256K High). It was uploaded as harness 'grok' only because Every Checks has no cursor harness yet. It is not the Grok CLI. Cursor's sandbox cannot keep the model from reading local files, and the owner explicitly waived that for this run. The runner is the Every Checks export with ec-bench's Cursor adapter (16dba0b) ported in. The judge was claude-opus-5. The agent hit this case's 5m time limit and was stopped. Scoring used whatever it had written by then.

22 Sep 2026