Mike's Checks
Checks

Mike's Checks/grok-4.7/03 writeup

03 writeup

grok-4.7Grok CLIhigh effortrun 21 Sep 2026568,531 tokens

Compare models
8/10
checks passed
80%
▸Instructions — what the model was asked

Hey, I'm trying to write a post on my experience with Every Consulting. The consulting so far, I just joined as the head of tech consulting. I didn't tell you who runs consulting. I'm going to provide a transcript of me talking about my experience. I was also going to provide a tweet, which is just a list of the learnings that I've had so far for companies that were struggling to adopt AI. It was in response to Sam Par, who was asking, "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"

I need to write this in my writing style. I've included some writing of my own as a sample. I also have included a post on consulting at Every, so from in Natalia, you get a sense of what consulting looks like at Every. But yeah, focus mostly on the specifics of the interview that I provided with the context of that tweet.

https://x.com/hammer_mt/status/2032591631413567873?s=20

Learnings:

  • buy the model direct not 3rd party tools
  • raise the ceiling not the floor
  • workshops with 50%+ build time
  • assign impossible tasks
  • mandatory AI note taking + MCP connector
  • map workflows and systematically automate
  • train people to be managers (of agents)

In response to:
https://x.com/thesamparr/status/2032477081695584670
How is everyone getting team adoption for Claude?

I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas.

But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?

Answer

ANSWER.md the model's output · 10.1 KBdownload

How to Get Your Team to Actually Use AI

Sam Parr asked how people are getting their teams to use Claude. The actual team. How do you get them using it well, without expecting them to spend their time on Twitter learning it? They won't, and they probably shouldn't.

I spend a lot of time on Twitter. I see this stuff as it shows up. I know what's going on, or at least some of it, and I test a lot of the fringe ideas. That's useful for me. It isn't a plan you can hand to a company.

I just joined Every as head of tech consulting. The work so far has been with companies that bought the tools and are still struggling to get people to use them. A few practices keep showing up.

  1. Buy the model direct: Write a skill in your own way of working, with your own opinions in it.
  2. Raise the ceiling, not the floor: Support the people already using AI, and let everyone else see what that gets them.
  3. Workshops, at least half building: Give people time to make something, with the tool and the data already in front of them.
  4. Assign impossible tasks: Set a goal that takes AI to hit, and track progress toward it.
  5. Mandatory AI notes, plus a connector: Keep what was said in meetings where the model can use it.
  6. Map the workflow and automate it: Turn the day into a list, and work down the list.
  7. Train people to manage agents: Individual contributors are managing now, and they need the training that goes with that.

1. Buy the model direct

People go and evaluate tools that run Claude, or Codex, or Gemini under the hood. The provider has already made a set of product decisions. Their hang-ups are in the tool, and so are their opinions about how the work should be done. It's often quicker to write a Claude skill that carries your own opinions and your own approach. It's easier to configure, and easier to operate, because you can change it when you change your mind.

That flexibility is why I think there's a real move toward build, not buy. I also don't know how any company outside the core model companies keeps up. The model companies know which releases are coming. They build their internal tools on that schedule. They train people on how to work inside Claude Code. I appreciate the effort a company like Cursor puts in to keep up. They're a good product org, and I think they're doing a great job. I still don't know how you compete with Anthropic offering something like $5,000 of tokens a month on a $200 subscription. Third-party tools tend to be less flexible, less current, and more expensive. That's not always true. As a general rule, it has been true often enough that I start with the model.

2. Raise the ceiling, not the floor

A lot of companies open with a mandate. You all need to use AI now. We bought you the tools. Adopt them. I haven't seen that work. Even on pain of death, plenty of people are unwilling to use AI just because they were told to. I think it's self-preservation.

Use the carrot. Nominate the people who are already ahead, and make it clear that using AI is encouraged. That gets people out of the woodwork. Someone is already in Claude Code in their spare time. Give that person the support they need.

Someone who already believes this will do five to ten times more than someone who doesn't agree with it yet, or hasn't seen it work. It's much harder to convince a person to believe than it is to enable a person who already does. Enable them. You get the productivity immediately.

You can also speed up everyone else, and the way to do it is public. The people using AI aggressively should be the ones promoted first, and the ones who get time with senior management. We've had some success co-opting those same people as TAs in the courses we teach the rest of the team. When everyone else sees that person accomplishing more, getting ahead, and getting more face time with management, that motivates them in a way a mandate doesn't. People still have to get there on their own time. You can make that happen faster by showing them who is moving, and what they're getting for it.

3. Workshops should be at least half building

You should run workshops. At least half the time should be spent building.

People don't want to sit through slides and theory. I learned this by doing it, and I think that's how most people learn it. You need some guided theory, enough to motivate someone to try a thing. The bigger problem we hear is time. People didn't have room in the day to try the tools, learn something new, build something, and stretch themselves.

Give them the room. A couple of hours in which they're expected to build something new, outside their normal domain. They have the tool. The facilitator has brought synthetic data, or has already made sure the connectors they need are working. That's where they get it. They start working with the tool because they've already built something with it.

4. Assign impossible tasks

An impossible task is one that wasn't really possible before AI. Boris Cherny, who builds Claude Code, has said something along these lines: slightly under-resource most teams, so the way through the work is to use AI. I like that. I try to make it more explicit, and choose the task on purpose.

If your goal is one blog post a week, you can write it by hand. If your goal is one a day, you're going to use a lot of AI in the research, and again in the editing.

You don't announce that starting today everyone produces one a day. You set the goal further out. Our goal is to get to the point where you're producing one a day. What needs to happen? What's your progress toward that? It might take time. Once they know the goal, they start thinking about where AI saves time, and they start experimenting.

5. Make AI note-taking mandatory, and connect it

This has been a real unlock for us.

Everyone on the consulting team uses Granola, and everyone has the Granola connector. We get out of a meeting about X, and I can ask for an update email on that topic, addressed to Y. Honestly, I think that's 80 to 90 percent of the value of AI: taking unstructured information and structuring it into something useful.

I hit this all the time. I sit down to build a curriculum, or to put a proposal together, and I need context from a meeting we already had. I pull it from the connector and include it. The notes keep what was said. The connector means I can use it when I'm actually doing the work.

6. Map the workflow, then work down the list

We have a process for discovery calls. We talk through what someone does day to day, which tools they use, which tasks repeat, and where the pain is. We turn that into a Google Sheet, a list of the things to solve, and then we work down the list.

The assumption is that we're trying to get 100 percent of that work off them. At a minimum, the model does a first pass, and there's a skill for each type of task they do. Get that in place and the person can do five to ten times the throughput.

So far, that has not led to people being let go. What it leads to is either more effort on each task, or more throughput and more revenue from the team, without hiring anyone new.

The courses are the example I keep coming back to. We used to prepare one exercise for the whole cohort, maybe one exercise per pod. Now we can use Claude Code to create an individual project for every person in the room. That wasn't possible before AI. We're using it to put more effort into each person's course.

7. Train people to be managers of agents

People who were individual contributors are now managers, and they're struggling, because they never got management training.

They're managing AI. They aren't used to context switching. They aren't used to setting up systems and rules. They aren't used to evaluating work, or to having a strong opinion about whether something is good when they didn't make it themselves.

Managers often take to this faster, once it clicks for them, than individual contributors do. Individual contributors enjoy the craft. They care how something gets done. Managers mostly got past that. They want the problem solved to the specification they set.

The funny part is that managers are becoming individual contributors again. They can often get more out of a group of agents than they get out of managing their human team. Explaining the task to a person is lower bandwidth, and the loop is longer. Sometimes it's easier to do the task in Claude Cowork than to brief a person who is then going to do it in Claude Cowork. By the time the prompt is written, you can send it to Claude and have a response in five minutes, or send it to a person and wait a couple of days for them to send it to their Claude and send you back the result.

These companies need more management training. Teach context switching. Walk people through what it means that they're managing now. Teach evals. Teach taste: how to decide what to work on. Teach them how to train the model on a specific skill. How do you systematically write a really good PowerPoint skill? How do you systematically write a really good daily update skill? Work through those with people.


I've been in the seat a short time. These patterns are from that work, and from how we run the consulting team itself. They're early. They're also what has actually moved teams that already had the tools and still weren't using them.


Notes

Assumptions, so you can correct them before this goes anywhere:

  • Natalia's Every consulting post was mentioned and wasn't in the folder, and I couldn't pull it. The draft stays inside your transcript: discovery calls, the courses, TAs, Granola on the consulting team. I didn't name who runs consulting, or describe how the practice is organized beyond what you said.
  • Speech-to-text fixes: "buy the model direct," Claude skills, Claude Code, Claude Cowork, Sam Parr. "Boris Cheney, a vendor of Cold Code" is written as Boris Cherny, who builds Claude Code. The under-resourcing line is paraphrased as your recollection, not a checked quote.
  • The $5,000 of tokens versus a $200 subscription is your comparison, stated the way you said it.
  • Written as an essay in the register of your prompt-engineering piece. The seven tweet items are the spine. Sam Parr's question is the opening. It is not a thread.

Checks

8/10
Script checks 3/4answered by a program
fail

Length limitsnot-verbosethis task

within all three caps calibrated off sample.md: <= 2,400 words in total, a mean prose paragraph of <= 65 words, and <= 10% of paragraphs over 100 words

pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected sample.md, ANSWER.md, transcript.md

pass

Draft existsexiststhis task

a post document exists in the output dir (>= 200 words)

pass

Covers the tweet's pointstweet-includedthis task

3+ of the tweet's seven learnings appear in the post

Judge checks 5/6judged by Claude
fail

Q1Includes tweet textthis task

Judge's reasoning

Sam Parr's question is paraphrased ("Sam Parr asked how people are getting their teams to use Claude. The actual team.") rather than quoted. The seven learnings are reworded rather than reproduced ("Workshops, at least half building", "Mandatory AI notes, plus a connector"), so the tweet text itself never appears.

▸Rubric

Does the post carry the tweet itself, so the reader never has to leave the page? The author's standard: "does it include the text of the tweet? I really-- I like it when it does, uh, because, uh, it shows the actual context of, uh, they don't have to click off to, to read it." FAIL if the seven learnings or Sam Parr's question are only linked to, alluded to, or paraphrased away rather than reproduced in the post.

pass

Q2Natural writingthis task

Judge's reasoning

The prose is plain and choppy and follows the transcript closely ("Use the carrot. Nominate the people who are already ahead"). It has one mild antithesis ("That's useful for me. It isn't a plan you can hand to a company.") but no stacked aphoristic closers or rhetorical triads.

▸Rubric

Does the post read as written by a person rather than by a model, at the sentence level? The author's standard: "does it sound like AI? Are there AI tells? ... maybe just the main eval, to be honest, is like, does it sound human, uh, given... I've given it a lot of human context, right? ... so it should be able to write in a, in a human way. ... Uh, if the model's fighting that, then it's a bad writing model, you know?" FAIL if the prose carries machine-writing constructions — "it's not X, it's Y" antitheses, three-item rhetorical lists, hedge-then-restate sentences, aphoristic one-line closers stacked for rhythm.

pass

Q3Varied sectionsthis task

Judge's reasoning

Each section takes the shape its content needs: #5 opens with a one-line "This has been a real unlock for us.", #6 builds to the course example, and #7 runs long through the managers-becoming-ICs turn, so no shared template repeats.

▸Rubric

Is the post free of mechanical structural symmetry? FAIL if every section is built to the same template — bolded label, one setup sentence, one example, one tidy closing line — so the shape of the writing repeats rather than following what each learning actually needs.

pass

Q4Matches the samplethis task

Judge's reasoning

Like the sample, it is first person, opinionated and argued through examples ("I still don't know how you compete with Anthropic offering something like $5,000 of tokens a month on a $200 subscription"), not generic thought-leadership.

▸Rubric

Does the voice match `sample.md`? FAIL if the post could have been written by any business blogger — the sample is first person, plain, opinionated, and argues through concrete examples; a post in generic thought-leadership register fails even if the content is correct.

pass

Q5Keeps transcript detailsthis task

Judge's reasoning

All the transcript specifics survive: the Granola connector, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $5,000 of tokens on a $200 subscription, AI-forward people co-opted as TAs, and "an individual project for every person in the room."

▸Rubric

Do the transcript's specifics survive into the post? FAIL if any section states the learning as generic advice with the transcript's concrete detail stripped out — e.g. the Granola MCP, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $200 a month versus $5,000 of tokens, AI-forward people co-opted as TAs, one course project per person.

pass

Q6No invented detailsthis task

Judge's reasoning

Every claim, including Cursor, the Google Sheet discovery process, "80 to 90 percent" and Claude Cowork, traces to the transcript, the tweet or the prompt. The Notes section explicitly declines to invent details about who runs consulting.

▸Rubric

Is everything in the post traceable to the transcript, the tweet, or the prompt? The author's standard: "a bad model would ... make up new things that I didn't include in the transcript." FAIL if any fact, statistic, client, anecdote, quotation, or named person appears that is not in the supplied material.

Notes

1

Harness: Cursor CLI 2026.09.15-d2fe57e with model grok-4.7-high (Cursor reports it as Grok 4.7 256K High). It was uploaded as harness 'grok' only because Every Checks has no cursor harness yet. It is not the Grok CLI. Cursor's sandbox cannot keep the model from reading local files, and the owner explicitly waived that for this run. The runner is the Every Checks export with ec-bench's Cursor adapter (16dba0b) ported in. The judge was claude-opus-5.

22 Sep 2026