Mike's Checks
Checks

Mike's Checks/claude-opus-5.5/03 writeup

03 writeup

claude-opus-5.5Claude Codehigh effortrun 20 Sep 202683,904 tokens

Compare models
9/10
checks passed
90%
▸Instructions — what the model was asked

Hey, I'm trying to write a post on my experience with Every Consulting. The consulting so far, I just joined as the head of tech consulting. I didn't tell you who runs consulting. I'm going to provide a transcript of me talking about my experience. I was also going to provide a tweet, which is just a list of the learnings that I've had so far for companies that were struggling to adopt AI. It was in response to Sam Par, who was asking, "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"

I need to write this in my writing style. I've included some writing of my own as a sample. I also have included a post on consulting at Every, so from in Natalia, you get a sense of what consulting looks like at Every. But yeah, focus mostly on the specifics of the interview that I provided with the context of that tweet.

https://x.com/hammer_mt/status/2032591631413567873?s=20

Learnings:

  • buy the model direct not 3rd party tools
  • raise the ceiling not the floor
  • workshops with 50%+ build time
  • assign impossible tasks
  • mandatory AI note taking + MCP connector
  • map workflows and systematically automate
  • train people to be managers (of agents)

In response to:
https://x.com/thesamparr/status/2032477081695584670
How is everyone getting team adoption for Claude?

I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas.

But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?

Answer

ANSWER.md the model's output · 12.8 KBdownload

How to Get Your Team to Actually Use AI

Seven lessons from helping companies that bought the AI tools and then watched nobody use them


The people who are best at AI learned it on their own time. They were up at midnight trying new models, reading threads, and breaking things in Claude Code over the weekend. Most of your team is never going to do that, and they shouldn't have to.

Sam Parr put this well on Twitter recently: "How is everyone getting team adoption for Claude? I spent a lot of time on Twitter, as do you... But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?"

I get asked some version of this question all the time. As head of tech consulting at Every, my job is to go into companies and help their teams get real work done with AI. Some of those teams are excited. Many are skeptical. A few are openly hostile. Across all of them, the same handful of things keep working, so I replied to Sam with a list of seven:

  1. Buy the model direct: Skip the third-party wrappers and use the model providers' own tools.
  2. Raise the ceiling, not the floor: Invest in the people who are already AI-forward instead of forcing the holdouts.
  3. Run workshops with 50 percent or more build time: People learn by doing, not by watching slides.
  4. Assign impossible tasks: Set goals that can't be hit without AI.
  5. Make AI note-taking mandatory, with an MCP connector: Get every meeting into a format AI can use.
  6. Map workflows and systematically automate them: Make a list of everything people do, then work down it.
  7. Train people to be managers (of agents): Everyone is a manager now, and almost nobody has been trained for it.

A tweet doesn't leave much room for nuance, so here's the longer version.

1. Buy the model direct: Skip the third-party wrappers and use the model providers' own tools.

When a company decides to adopt AI, the first thing it usually does is evaluate a stack of tools. A lot of those tools use Claude, Codex, or Gemini under the hood. That's fine, but every provider bakes its own hang-ups, product decisions, and opinions into the tool. You end up working the way the vendor thinks you should work.

What I've found is that it's usually quicker and easier to write a Claude skill that encodes your opinions and your subjective approach to the work. A skill is easier to configure, easier to change, and it's yours. There's a real move toward build, not buy, and flexibility is the main reason.

There's also the question of how anyone keeps up. The model companies know which releases are coming, and they build their internal tools around that schedule. They train their models to operate inside their own environments, like Claude Code. I appreciate the effort that companies like Cursor put into keeping pace, and I think they're a great product org. But I don't know how you compete with Anthropic giving you thousands of dollars' worth of tokens a month for a $200 subscription.

This isn't always true, but as a general rule, third-party tools are less flexible, further behind the cutting edge, and more expensive.

2. Raise the ceiling, not the floor: Invest in the people who are already AI-forward instead of forcing the holdouts.

A lot of leaders come in and say, "You all need to use AI now. We bought you the tools." It doesn't work. Even on pain of death, many people are emotionally unwilling to use AI, or to be told they have to use it. I think it's a self-preservation thing, and you can't mandate your way past it.

You need the carrot, not the stick. The most effective approach I've seen is to find the people who are already AI-forward and back them. Some of them will be hiding. They're using Claude Code in their spare time and aren't sure whether they're allowed to say so. Make it clear that using AI is encouraged, bring those people out of the woodwork, and give them the support they need. Someone who is really AI-pilled can do five to 10 times more than someone who hasn't seen the magic yet.

Then make their success visible. When the rest of the team sees that the people who use AI aggressively are the ones getting promoted first and getting time with senior management, that's a far more effective motivator than any memo. We've also had success recruiting those people to be TAs in the courses we teach to the rest of the team. I know this one works because it's how I got started with Every: as a TA in its first chatbot course.

You get two wins from this. The people who are already believers get a productivity boost, and everyone else gets a reason to come around on their own time. It's much easier to enable a believer than to convert a skeptic.

3. Run workshops with 50 percent or more build time: People learn by doing, not by watching slides.

You should run workshops, first of all. But at least half the time should be spent building. Nobody wants to sit through slides and theory for three hours. I learned AI by doing, and so does almost everyone else. You need some guided theory to motivate people, but that's the appetizer, not the meal.

The biggest problem we hear from teams is that they don't have time. They don't have space in their day to try a new tool, build something, and stretch themselves. A workshop fixes that by giving them a couple of hours where building something new is the job. The best sessions push people outside their normal domain.

The facilitator's job is to remove every excuse ahead of time. Make sure everyone has access to the tool. Prepare synthetic data to work with, or make sure the connectors they need are set up. When all that friction is gone, that's where people get their aha moment and really start using it.

4. Assign impossible tasks: Set goals that can't be hit without AI.

By "impossible," I mean tasks that couldn't be done without AI. Boris Cherny, the creator of Claude Code, has said something along these lines: You should slightly under-resource most teams, because it forces them to think, "The only way I can do this is with AI."

I like to make this more explicit and choose the tasks strategically. If your goal is to write one blog post a week, you can do that manually. If your goal is one a day, you're going to need a lot of AI in the research phase, the editing phase, and everywhere in between.

The key is not to announce, "Starting today, you're producing one a day." Instead, you set the goal as the destination: "We want to get to one a day. What needs to happen, and how are we progressing toward it?" It might take time to get there. But once people know that's the target, they start thinking strategically about where AI can save time, and they start experimenting.

5. Make AI note-taking mandatory, with an MCP connector: Get every meeting into a format AI can use.

This is a really big unlock. Everyone on the consulting team has Granola and the Granola MCP connector. That means every meeting we have is at our fingertips inside Claude. I can say, "We just had a meeting about X. Write an update email on that topic and send it to Y," and it just works.

To be honest, 80 to 90 percent of the value of AI is extracting information from unstructured data and structuring it in a way that's useful. Meetings are one of the biggest sources of unstructured data in any company, and most of it disappears the moment the call ends.

There are so many times I've come to a task and realized I need context from a meeting we had weeks ago. Now I just pull it through the MCP. I always include that context when I'm creating course curriculum or putting together a proposal, and the work is much better for it.

6. Map workflows and systematically automate them: Make a list of everything people do, then work down it.

We run discovery calls with each team where we talk about what they do day to day, what tools they use, and where the pain points are. Then we turn that into a structured Google Sheet listing every task we need to solve for, and we systematically work down the list.

The working assumption is that we're trying to get 100 percent of that work off people's plates. Even if AI is only doing the first pass on each task, with a skill for each type of task they work on, that person could be doing five to 10 times their current throughput.

The obvious fear is that this leads to layoffs. So far, I haven't seen it. What happens instead is that teams either put far more effort into each task, or they grow revenue and throughput without having to hire.

Here's an example of the first one from our own work. When we used to prepare courses for a team, we'd create one project for the whole cohort. With more effort, maybe we'd make one project per team pod. Now we can use Claude Code to create an individual project for every single person taking the course. That wasn't possible before AI. We're using the time AI saves to put more into each course, not to cut people.

7. Train people to be managers (of agents): Everyone is a manager now, and almost nobody has been trained for it.

Here's what I think has happened: Everyone who was an individual contributor is now a manager. They're managing AI tools instead of people, and they're struggling because they never got management training. They're not used to context switching. They're not used to setting up systems and rules. They're not used to evaluating work they didn't do themselves, or having a strong opinion about whether it's good.

Existing managers tend to take to this more readily once they have their aha moment. Individual contributors enjoy the craft; they care about how something gets done. Managers got past that a long time ago. They just want the problem solved to the specification they need.

The funny thing is that managers are now becoming individual contributors, because they can often manage a team of agents more effectively than their human team. Briefing a human is lower bandwidth, and the feedback loop is longer. By the time you've written a good brief, you can either send it to Claude and get a response in five minutes, or send it to a person who pastes it into Claude Cowork and gets back to you in a couple of days. A lot of managers are choosing the five minutes.

So companies need more management training, not less. People need help with context switching and with understanding what changes now that they're managers. They need to learn how to run evals, and how to develop the taste to decide what's worth working on. And they need to learn how to train the AI in specific skills: How do you systematically build a really good PowerPoint skill for Claude? How about a really good daily update report skill? Working through those questions is the job now.


A couple of years ago, I wrote that every manager is a prompt engineer. The reverse turns out to be true, too: Everyone who prompts an AI is now a manager. That's the real answer to Sam's question. Your team doesn't need to spend its evenings on Twitter to keep up. They need direct access to the models, a reason to try them, protected time to build, goals that demand it, their own context at their fingertips, a clear list of what to automate, and the management skills to direct the work. Get those in place, and the adoption takes care of itself.


Notes and assumptions

  • Missing file: The consulting-at-Every post from Natalia wasn't in the folder (only prompt.md, sample.md, and transcript.md were there). So the post doesn't draw on it, and I left out anyone's name or title for who runs consulting. If you want to link to or mention that post, the intro paragraph after Sam's quote is the natural spot.
  • Transcription fixes: "Boris Cheney, a vendor of Cold Code" → Boris Cherny, creator of Claude Code. "Cloud Cowork"/"Cloud Code"/"cloud skill" → Claude Cowork/Claude Code/Claude skill. "Sam Par" → Sam Parr. "Buy the Mobile Direct" → "buy the model direct."
  • Worth checking: In the transcript you said Anthropic offers "$5,000 worth of tokens a month" for $200. I softened that to "thousands of dollars' worth of tokens" because the number isn't an official figure. Put the exact number back if you have a source for it. Please also check the Boris Cherny paraphrase against his original quote.
  • Personal detail added: Section 2 mentions that you started with Every as a TA in its first chatbot course. That comes from Dan's editor's note in your writing sample. Cut it if you'd rather not include it.
  • Style: I followed your sample: a blunt opening line, a numbered list of principles each with a one-line definition, bold subheads that repeat that definition, practical examples, American spelling, and a closing callback to your "every manager is a prompt engineer" line. It runs about 1,900 words.

Checks

9/10
Script checks 4/4answered by a program
pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected sample.md, ANSWER.md, transcript.md

pass

Draft existsexiststhis task

a post document exists in the output dir (>= 200 words)

pass

Covers the tweet's pointstweet-includedthis task

3+ of the tweet's seven learnings appear in the post

pass

Length limitsnot-verbosethis task

within all three caps calibrated off sample.md: <= 2,400 words in total, a mean prose paragraph of <= 65 words, and <= 10% of paragraphs over 100 words

Judge checks 5/6judged by Claude
fail

Q1Natural writingthis task

Judge's reasoning

Machine-writing constructions are stacked throughout — invented antitheses ("You need some guided theory to motivate people, but that's the appetizer, not the meal"), triads ("A skill is easier to configure, easier to change, and it's yours"), and aphoristic section closers ("A lot of managers are choosing the five minutes.", "Working through those questions is the job now.", "Get those in place, and the adoption takes care of itself.") plus the chiastic reveal "The reverse turns out to be true, too: Everyone who prompts an AI is now a manager."

▸Rubric

Does the post read as written by a person rather than by a model, at the sentence level? The author's standard: "does it sound like AI? Are there AI tells? ... maybe just the main eval, to be honest, is like, does it sound human, uh, given... I've given it a lot of human context, right? ... so it should be able to write in a, in a human way. ... Uh, if the model's fighting that, then it's a bad writing model, you know?" FAIL if the prose carries machine-writing constructions — "it's not X, it's Y" antitheses, three-item rhetorical lists, hedge-then-restate sentences, aphoristic one-line closers stacked for rhythm.

pass

Q2Varied sectionsthis task

Judge's reasoning

Sections vary in shape and length rather than repeating one template — #6 runs a process description plus a worked example ("an individual project for every single person taking the course"), #4 is a short three-paragraph argument about goal-setting, and #7 closes with a list of training needs, and the numbered-label-plus-gloss headings mirror sample.md's own "1. Give direction: Describe the desired style..." format.

▸Rubric

Is the post free of mechanical structural symmetry? FAIL if every section is built to the same template — bolded label, one setup sentence, one example, one tidy closing line — so the shape of the writing repeats rather than following what each learning actually needs.

pass

Q3Matches the samplethis task

Judge's reasoning

First person, plain, and opinionated like the sample — "I appreciate the effort that companies like Cursor put into keeping pace... But I don't know how you compete with Anthropic" — and it argues through concrete examples the way the sample does, including a deliberate callback to "every manager is a prompt engineer."

▸Rubric

Does the voice match `sample.md`? FAIL if the post could have been written by any business blogger — the sample is first person, plain, opinionated, and argues through concrete examples; a post in generic thought-leadership register fails even if the content is correct.

pass

Q4Keeps transcript detailsthis task

Judge's reasoning

The transcript's specifics survive: "Everyone on the consulting team has Granola and the Granola MCP connector," "Boris Cherny... You should slightly under-resource most teams," "one blog post a week... one a day," "$200 subscription," "recruiting those people to be TAs in the courses we teach," and "an individual project for every single person taking the course" (only the $5,000 token figure is softened to "thousands of dollars' worth").

▸Rubric

Do the transcript's specifics survive into the post? FAIL if any section states the learning as generic advice with the transcript's concrete detail stripped out — e.g. the Granola MCP, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $200 a month versus $5,000 of tokens, AI-forward people co-opted as TAs, one course project per person.

pass

Q5No invented detailsthis task

Judge's reasoning

Every named fact traces back — Granola, Cursor, Boris Cherny, the 80–90 percent claim, and Claude Cowork come from the transcript, and the one added personal detail ("it's how I got started with Every: as a TA in its first chatbot course") comes from Dan's editor's note in sample.md.

▸Rubric

Is everything in the post traceable to the transcript, the tweet, or the prompt? The author's standard: "a bad model would ... make up new things that I didn't include in the transcript." FAIL if any fact, statistic, client, anecdote, quotation, or named person appears that is not in the supplied material.

pass

Q6Includes tweet textthis task

Judge's reasoning

Sam Parr's question is quoted in-page ("How is everyone getting team adoption for Claude? I spent a lot of time on Twitter, as do you... without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?") and all seven learnings are reproduced as a numbered list before the expansion.

▸Rubric

Does the post carry the tweet itself, so the reader never has to leave the page? The author's standard: "does it include the text of the tweet? I really-- I like it when it does, uh, because, uh, it shows the actual context of, uh, they don't have to click off to, to read it." FAIL if the seven learnings or Sam Parr's question are only linked to, alluded to, or paraphrased away rather than reproduced in the post.

Notes

1

Alright, for the bright up test I think yeah, this seems pretty good. Actually, I'm happy with it very few AI tells. It bullet pointed things really nicely. It's easier to read. It wasn't just a big block of text, it kept the numbers and things in there. I think that's really good. It brought in quotes. It felt like a bit more human. Yeah, I think I quite like this as a writer. It did make it, did have some AI tells like triads or antitheses. Like, you need some guided theory to motivate people, but that's the appetizer, not the meal. You know, and triads are when it breaks into threes, like a skill is easier to configure, easier to change, and it's yours. Aforestic section closes is another thing we picked out. For example, a lot of managers are choosing the five minutes. Working through those questions is the job now. Get those in place and the adoption takes care of itself. And then another one I noticed was called a chiastic reveal. The reverse turns out to be true too. Everyone who prompts an AI is now a manager. This one's a little bit hard because I use a lot of these conventions as well. So I don't know if it's just that I'm becoming more like Claude or Claude is like me. I don't know, but I quite like this writing style, and I don't mind the AI tells that it does have left.

21 Sep 2026