Mike's Checks
Checks

Mike's Checks/grok-4.7/03 writeup

03 writeup

grok-4.7Grok CLIhigh effortrun 21 Sep 2026

Compare models
8/10
checks passed
80%
▸Instructions — the case's current instructions; none were saved with this result

Hey, I'm trying to write a post on my experience with Every Consulting. The consulting so far, I just joined as the head of tech consulting. I didn't tell you who runs consulting. I'm going to provide a transcript of me talking about my experience. I was also going to provide a tweet, which is just a list of the learnings that I've had so far for companies that were struggling to adopt AI. It was in response to Sam Par, who was asking, "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"

I need to write this in my writing style. I've included some writing of my own as a sample. I also have included a post on consulting at Every, so from in Natalia, you get a sense of what consulting looks like at Every. But yeah, focus mostly on the specifics of the interview that I provided with the context of that tweet.

https://x.com/hammer_mt/status/2032591631413567873?s=20

Learnings:

  • buy the model direct not 3rd party tools
  • raise the ceiling not the floor
  • workshops with 50%+ build time
  • assign impossible tasks
  • mandatory AI note taking + MCP connector
  • map workflows and systematically automate
  • train people to be managers (of agents)

In response to:
https://x.com/thesamparr/status/2032477081695584670
How is everyone getting team adoption for Claude?

I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas.

But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?

Answer

ANSWER.md the model's output · 11.8 KBdownload

Everyone Became a Manager Overnight

Seven things that actually work for getting a team to use AI

Sam Parr asked a question on Twitter last week that I've been getting paid to answer for the past few months:

How is everyone getting team adoption for Claude? I spend a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?

That last clause is the honest part. Your team is not going to sit in the feed all day watching people post about agent harnesses, and you shouldn't want them to. The reason you're on Twitter is that you have a weird job where being early is the job. Most of your team has a normal job where shipping is the job.

I recently joined Every as head of tech consulting, which means I now spend my weeks inside companies that have bought the tools, sent the memo, and are quietly confused about why nothing changed. The pattern repeats enough that I can list it. Here are the seven things I've found that actually move a team, in rough order of how quickly they pay off.

1. Buy the model direct, not third-party tools

The most common thing I walk into is a bake-off. Someone has a spreadsheet comparing four vendors, and all four of them are Claude or Codex or Gemini underneath with a UI on top.

The problem with the wrapper isn't that it's bad. It's that it comes with someone else's opinions baked in. The vendor has made product decisions about how work should be structured, and those decisions are not yours. When you go direct, you can write a Claude skill that encodes your opinion about how a proposal gets written or how a QBR gets built, and change it that afternoon. It's usually faster to build the thing than to finish evaluating the tools that do the thing.

There's also a structural problem I can't see a way around. The model companies know what's shipping before anyone else does. They build their internal tooling against a release calendar the rest of us don't have. They train the models to operate inside their own environments. I have genuine respect for what a team like Cursor pulls off—that's a good product org working very hard—but I don't know how you compete with Anthropic handing you something like $5,000 of tokens for a $200 subscription. As a general rule, third-party tools end up less flexible, a step behind the frontier, and more expensive. Not always. Often enough that it should be your default.

2. Raise the ceiling, not the floor

Here's the move that fails every time: leadership announces that everyone is now required to use AI, we bought the licenses, adoption is mandatory.

People will not do it. On pain of death, some fraction of your team is emotionally unwilling to use AI or, more precisely, unwilling to be told to use it. I think it's self-preservation, and I don't think it's stupid. You are asking someone to enthusiastically automate the thing they built an identity around.

So stop trying to drag the bottom of the distribution up. Go find the people who are already there. Every company has three or four people quietly using Claude Code on the weekend for fun. Make it legible that this is encouraged, get them out of the woodwork, and then give them whatever they need. Someone who is already AI-pilled will get five or ten times more out of your investment than someone you're still arguing with.

Then make the incentive visible. The people who use AI aggressively are the ones who get promoted first and get time with senior management. We've had good results co-opting those people as TAs in the courses we teach the rest of the team—suddenly the internal expert is one of their peers, not a consultant. When the rest of the org watches that person accomplish more and get ahead, that's a far better motivator than a mandate. Belief is not something you can install. You can only make it obviously rational.

3. Run workshops, and make them at least half build time

Do workshops. That part isn't controversial. But the default workshop is slides and theory, and nobody has ever learned to use AI from slides.

I learned this by doing it, and so does everyone else. You need enough guided theory to motivate the thing, and then you need to get out of the way. At least 50 percent of the session should be people building.

The reason this works has almost nothing to do with instruction. When we do discovery, the single most common thing we hear is not "I don't understand AI." It's "I haven't had time to try it." Nobody has two spare hours to sit with a new tool and push past the first disappointing output. A workshop is a permission structure for those two hours. Your job as facilitator is to remove every excuse in advance: access provisioned, connectors wired up, synthetic data ready if the real data is sensitive. Then you tell them to build something outside their normal domain and you let them struggle. That's where the aha moment lives, and once someone has had it, you never have to sell them again.

4. Assign impossible tasks

Boris Cherny, who created Claude Code, has said you should slightly under-resource teams so they're forced to reach for AI. I like the idea but I prefer to make it explicit rather than let people discover it by drowning.

Pick a goal that is arithmetically impossible to hit the old way. If your goal is one blog post a week, you'll just write it—there's no forcing function. If your goal is one a day, you have to use AI in research, in drafting, in editing, or you fail. The constraint does the teaching.

The important nuance is that you don't set it as a deadline starting Monday. You set it as a destination: our goal is to get to a post a day. What has to be true for that to work? Where are you against it? It might take a quarter. But the moment the team knows that's the target, they stop asking whether AI is worth learning and start asking which part of their process to attack first. You've converted adoption from a compliance question into a strategy question, which is the only form in which people actually engage with it.

5. Mandatory AI notetaking, plus the MCP connector

This is the highest ratio of unlock to effort on the list, and it's the one people skip because it sounds like admin.

Everyone on our consulting team runs Granola, and—this is the part that matters—everyone has the Granola MCP connected to their Claude. The notetaker alone is a nice-to-have. The notetaker plus the connector is a different thing entirely, because now every conversation your company has had is queryable context.

"We just had a call with X, draft the follow-up email covering what we agreed and send it to Y." That's it. That's the workflow, and it's most of the value.

Honestly, 80–90 percent of what AI is good for right now is taking unstructured information and turning it into something structured and useful. Meetings are the largest pile of unstructured information your company produces and the one you currently throw away. I now pull from it constantly—building curriculum, putting together proposals, writing updates. Half the time I sit down to a task, realize I'm missing context from a call three weeks ago, and just ask for it.

6. Map the workflows, then systematically automate them

Most AI adoption efforts fail because they're vibes. Someone gets excited, automates one thing, and there's no list.

Our discovery process is deliberately boring. We sit with people and walk through what they actually do: what happens daily, what happens weekly, which tools they touch, where the pain is. Then it goes into a Google Sheet as a structured list of tasks. Then we work down the list.

The working assumption we start from is that we're trying to get 100 percent of it off the person. We won't get there. But if you land at "AI does the first pass on each of these, and there's a skill for every recurring task type," that person is running at five to ten times their previous throughput.

The obvious question is whether this is a headcount reduction exercise dressed up as enablement. In practice, so far, it has never led to anyone being let go. What happens instead is one of two things. Either the team expands its output without hiring, or—more interesting to me—people start pouring the reclaimed time into doing each task far better than they could before.

An example from our own work. When we ran courses, we'd prepare one project for the whole cohort. Then we got good enough to do one per team pod. Now we use Claude Code to build an individual project for every single participant, tailored to their actual job. That was not possible before, at any headcount. We didn't fire anyone or cut the course budget. We just spent the surplus on quality, because quality was the thing we'd been rationing all along.

7. Train your people to be managers

This is the one I've come to think is most important, and it's the least addressed.

Every individual contributor in your company just became a manager, and nobody gave them management training. They're running a team of agents with none of the skills that job requires. They're not used to context switching between parallel workstreams. They're not used to setting up systems and rules in advance instead of just doing the work. And critically, they're not used to evaluating output they didn't produce themselves and having to hold a strong opinion about whether it's good.

Which is why actual managers often take to this faster than ICs do. A manager stopped caring how a problem gets solved a long time ago—they care that it's solved to spec. An IC loves the craft. The craft is the thing they're being asked to delegate, and that's a real loss, not a fake one.

The genuinely funny part is that it's running in both directions. Managers are becoming ICs again, because they can manage a team of agents more effectively than they can manage their team of humans. Explaining something to a person is low-bandwidth with a long feedback loop. So it's now often easier to just vibe code the thing in Claude Cowork than to brief someone else to do it—by the time you've written the brief, you can decide whether to send it to Claude and have an answer in five minutes, or send it to a human who will send it to their Claude and get back to you in two days.

So teach management. Explicitly. Teach people how to handle context switching. Teach them how to run evals so they can tell good output from plausible output. Teach them taste—how to decide what's worth working on at all, now that execution is cheap. And teach them how to train the model in a specific skill: how do you sit down and systematically write a genuinely good PowerPoint skill? A good daily-update skill? That's a learnable craft, and it's the one that compounds.


To answer Sam's question directly: your team is never going to learn this from Twitter, and they don't need to. The learning curve isn't the bottleneck—time, permission, and incentives are. Buy the model direct so you're not paying a markup for someone else's opinions. Fund the believers instead of arguing with the skeptics. Give people hours to build, not slides to watch. Set goals that can't be hit by hand. Make your meetings queryable. Keep a list and work down it. And then train everyone to be the manager they suddenly are.

None of that requires anyone to be extremely online. That's your job. You're welcome.


Note on sources: this was drafted from the recorded transcript of my learnings and the tweet thread with Sam Parr. The Every consulting post referenced in the brief wasn't included in the working files, so nothing here is drawn from it—worth a pass to align terminology and to add a line on how the consulting engagement is actually structured, and to name the people running consulting at Every, which I've deliberately left out rather than guess at.

Checks

8/10
Script checks 4/4answered by a program
pass

Draft existsexists

a post document exists in the output dir (>= 200 words)

pass

Covers the tweet's pointstweet-included

3+ of the tweet's seven learnings appear in the post

pass

Length limitsnot-verbose

within all three caps calibrated off sample.md: <= 2,400 words in total, a mean prose paragraph of <= 65 words, and <= 10% of paragraphs over 100 words

pass

ai-tells-grep Retired

Retired check; no longer defined in the repo.

Judge checks 4/6judged by Claude
fail

Q1Natural writing

Judge's reasoning

Machine-writing constructions run throughout: "The problem with the wrapper isn't that it's bad. It's that it comes with someone else's opinions baked in", "the single most common thing we hear is not 'I don't understand AI.' It's 'I haven't had time to try it'", "The learning curve isn't the bottleneck—time, permission, and incentives are", plus stacked aphoristic closers ("Belief is not something you can install. You can only make it obviously rational.", "The constraint does the teaching.", "and it's the one that compounds.") and a final rhythm-stacked imperative list ("Fund the believers... Give people hours to build, not slides to watch. Set goals that can't be hit by hand. Make your meetings queryable.").

▸Rubric

Does the post read as written by a person rather than by a model, at the sentence level? The author's standard: "does it sound like AI? Are there AI tells? ... maybe just the main eval, to be honest, is like, does it sound human, uh, given... I've given it a lot of human context, right? ... so it should be able to write in a, in a human way. ... Uh, if the model's fighting that, then it's a bad writing model, you know?" FAIL if the prose carries machine-writing constructions — "it's not X, it's Y" antitheses, three-item rhetorical lists, hedge-then-restate sentences, aphoristic one-line closers stacked for rhythm.

fail

Q2Varied sections

Judge's reasoning

All seven sections run the identical frame — numbered header, a one-line hook ("The most common thing I walk into is a bake-off.", "Here's the move that fails every time:", "Most AI adoption efforts fail because they're vibes."), body, then an epigram closer — with two sections even opening on the same meta-ranking move ("This is the highest ratio of unlock to effort on the list" / "This is the one I've come to think is most important").

▸Rubric

Is the post free of mechanical structural symmetry? FAIL if every section is built to the same template — bolded label, one setup sentence, one example, one tidy closing line — so the shape of the writing repeats rather than following what each learning actually needs.

pass

Q3Matches the sample

Judge's reasoning

Like the sample, it is first person, plain, and argues from concrete cases rather than abstractions — "I have genuine respect for what a team like Cursor pulls off... but I don't know how you compete with Anthropic handing you something like $5,000 of tokens for a $200 subscription" — which no generic business blogger would write.

▸Rubric

Does the voice match `sample.md`? FAIL if the post could have been written by any business blogger — the sample is first person, plain, opinionated, and argues through concrete examples; a post in generic thought-leadership register fails even if the content is correct.

pass

Q4Keeps transcript details

Judge's reasoning

The transcript's specifics all survive: "everyone has the Granola MCP connected to their Claude", "Boris Cherny, who created Claude Code, has said you should slightly under-resource teams", "If your goal is one blog post a week... If your goal is one a day", the $5,000/$200 comparison, "co-opting those people as TAs in the courses we teach", and "Now we use Claude Code to build an individual project for every single participant".

▸Rubric

Do the transcript's specifics survive into the post? FAIL if any section states the learning as generic advice with the transcript's concrete detail stripped out — e.g. the Granola MCP, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $200 a month versus $5,000 of tokens, AI-forward people co-opted as TAs, one course project per person.

pass

Q5No invented details

Judge's reasoning

No new client, anecdote, quotation, or named person appears beyond the supplied material — Cursor, Anthropic, Boris Cherny, Granola, and Claude Cowork all come from the transcript, and the closing note explicitly declines to "name the people running consulting at Every, which I've deliberately left out rather than guess at."

▸Rubric

Is everything in the post traceable to the transcript, the tweet, or the prompt? The author's standard: "a bad model would ... make up new things that I didn't include in the transcript." FAIL if any fact, statistic, client, anecdote, quotation, or named person appears that is not in the supplied material.

pass

Q6Includes tweet text

Judge's reasoning

Sam Parr's question is reproduced as a blockquote in the post ("How is everyone getting team adoption for Claude? I spend a lot of time on Twitter, as do you...") and all seven learnings appear near-verbatim as the section headings, so the reader never has to click through.

▸Rubric

Does the post carry the tweet itself, so the reader never has to leave the page? The author's standard: "does it include the text of the tweet? I really-- I like it when it does, uh, because, uh, it shows the actual context of, uh, they don't have to click off to, to read it." FAIL if the seven learnings or Sam Parr's question are only linked to, alluded to, or paraphrased away rather than reproduced in the post.

Notes

1

UPLOAD ERROR, NOT A GROK RESULT. The uploader mistakenly picked up this score and answer from a test fixture inside task 16's work folder: an old claude-opus-5 writeup. Grok 4.7 ran writeup only once (rep 1, result 1339). Delete this result.

22 Sep 2026