Mike's Checks
Checks

Mike's Checks/claude-fable-5-1/03 writeup

03 writeup

claude-fable-5-1Claude Codehigh effortrun 22 Sep 2026176,777 tokens

Compare models
8/10
checks passed
80%
▸Instructions — what the model was asked

Hey, I'm trying to write a post on my experience with Every Consulting. The consulting so far, I just joined as the head of tech consulting. I didn't tell you who runs consulting. I'm going to provide a transcript of me talking about my experience. I was also going to provide a tweet, which is just a list of the learnings that I've had so far for companies that were struggling to adopt AI. It was in response to Sam Par, who was asking, "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"

I need to write this in my writing style. I've included some writing of my own as a sample. I also have included a post on consulting at Every, so from in Natalia, you get a sense of what consulting looks like at Every. But yeah, focus mostly on the specifics of the interview that I provided with the context of that tweet.

https://x.com/hammer_mt/status/2032591631413567873?s=20

Learnings:

  • buy the model direct not 3rd party tools
  • raise the ceiling not the floor
  • workshops with 50%+ build time
  • assign impossible tasks
  • mandatory AI note taking + MCP connector
  • map workflows and systematically automate
  • train people to be managers (of agents)

In response to:
https://x.com/thesamparr/status/2032477081695584670
How is everyone getting team adoption for Claude?

I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas.

But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?

Answer

ANSWER.md the model's output · 12 KBdownload

Raise the Ceiling, Not the Floor

Seven things I've learned about getting teams to actually use AI, from my first few months running tech consulting at Every

Sam Parr asked a question on X recently that I've heard, in one form or another, from almost every company I've walked into this year:

"How is everyone getting team adoption for Claude? I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta... But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?"

It's the right question. If you're reading this, you are by definition the person at your company who spends too much time on Twitter. You've seen the demos. You have the $200-a-month subscription. The problem is that you can't clone yourself, and the 40 people on your team have actual jobs to do that don't include "keep up with AI."

I recently joined Every as head of tech consulting, which means I now spend my days inside companies that bought the tools, told everyone to use them, and then watched nothing happen. I replied to Sam with a list of seven things that have actually moved the needle. Here they are in a bit more depth.

1. Buy the model direct, not third-party tools

The first thing most companies do when they decide to "adopt AI" is run a vendor evaluation. They line up a dozen tools that all use Claude or Codex or Gemini under the hood, and they compare features. This is almost always a waste of time.

Every one of those tools has the provider's opinions baked into it: their hang-ups, their product decisions, their view of how your work should be done. Nine times out of ten it's quicker and easier to write a Claude skill that encodes your opinions and your way of working, and point it at the model directly. It's easier to configure, easier to operate, and you can change it on a Tuesday afternoon without waiting for a roadmap.

I don't see how the third-party tools keep up in the long run. The model companies know what releases are coming. They build their internal tools around those releases. They train the model itself on how to operate inside their own environment. I genuinely appreciate the effort a company like Cursor puts in, and they're a good product org, but I also don't know how you compete with Anthropic effectively handing you $5,000 worth of tokens a month for a $200 subscription.

There are exceptions, but as a general rule, third-party tools are less flexible, less cutting-edge, and more expensive. It's a real shift toward build, not buy.

2. Raise the ceiling, not the floor

The default corporate rollout goes something like: "We bought you AI tools. You all need to use AI now." Then leadership is surprised when three months later the usage dashboard is flat.

This doesn't work because for a lot of people, being told they have to use AI triggers something closer to self-preservation than curiosity. Even on pain of death, some of your team are just not emotionally ready to be told a machine can do their job. You can't mandate the aha moment. People have to get there on their own time.

What you can do is use the carrot rather than the stick. Instead of trying to drag the bottom of the distribution up to some minimum standard, find the people who are already AI-forward and remove every obstacle in their way. Every company has a few people quietly using Claude Code in their spare time. Get them out of the woodwork. Make it visible that this is encouraged. Give them the budget and the access they need.

Someone who is genuinely AI-pilled will get five or ten times more done than someone who hasn't seen the magic yet, so you get the productivity boost immediately, without having to convince anyone of anything. It is much easier to enable a believer than to convert a skeptic.

Then make the incentives legible. The people who use AI aggressively should be the ones who get promoted first and who get face time with senior management. We've even had success co-opting these people as TAs in the courses we run for the rest of the team. When their colleagues see that the TA is accomplishing more, getting ahead in their career, and sitting in meetings with leadership, that's a far more effective motivator than any mandate.

3. Workshops should be at least 50 percent build time

You should run workshops. But nobody wants to sit through two hours of slides on how transformers work. I didn't learn any of this from theory; I learned it by doing, and so will your team.

The single biggest complaint we hear in discovery is not "I don't believe in AI." It's "I don't have time." People don't have room in their day to check out a new tool, learn something unfamiliar, build something with it, and stretch themselves. So the workshop's real job is to give them that time, with permission attached.

Our format is a bit of guided theory to motivate the exercise, then the majority of the session is building. Everyone is expected to make something new, ideally outside their normal domain. The facilitator's job beforehand is to make sure the tool is set up, the connectors they need are working, and there's synthetic data ready to go so nobody spends 45 minutes on access requests. That's where the aha moment happens: not when they're watching, but when they're two hours in and the thing they didn't think they could build is sitting in front of them.

4. Assign impossible tasks

By "impossible," I mean tasks that simply couldn't be done without AI.

Boris Cherny, who created Claude Code, has said something similar: slightly under-resource most teams, and they'll work out for themselves that the only way to hit the target is to use AI. I like to make it more explicit and choose the tasks strategically.

If your goal is one blog post a week, you can do that manually, and you will. If your goal is one blog post a day, you're going to have to use AI heavily in research, drafting, and editing, and you'll have to think about where it fits. The task forces the question.

The important nuance is how you set the goal. You don't say, "Starting today, you're producing one a day." You say, "Our goal is to get to the point where you can produce one a day. What needs to happen for that to be true, and what's your progress toward it?" It might take a while. But once people know that's the destination, they start thinking strategically about how to get there and start experimenting on their own.

5. Mandatory AI note-taking, plus an MCP connector

This is a surprisingly big unlock, and it's cheap. Everyone on our consulting team has Granola and the Granola MCP, which means every meeting we've ever had is available to Claude as context.

What that looks like in practice: "We just had a meeting on X. Write an update email on that topic and send it to Y." Done. Or I'm putting together a proposal or a curriculum, I realize I need something a client said three weeks ago, and rather than digging through my notes I just pull it from the MCP.

If I'm honest, 80 to 90 percent of the value of AI in a business context comes down to one thing: extracting information from unstructured data and structuring it in a way that's useful. Meetings are the largest pile of unstructured data in most companies, and until recently they evaporated the moment they ended. Make note-taking mandatory, connect it to the model, and you've given every prompt your team writes a memory.

6. Map workflows and systematically automate them

Most AI adoption is ad hoc. Someone tries something, it works or it doesn't, and nobody writes it down. We take a more boring, and much more effective, approach.

We run a discovery call with each person or team: what do you do day to day, what tools do you use, where are the pain points? We turn that into a structured Google Sheet listing every task that needs solving. Then we work down the list, systematically, building a skill for each type of task.

The working assumption is that we're trying to take 100 percent of the work off the human. We won't get there. But if every task on the list at least gets a good first pass from a skill, that person can do five or ten times the throughput they do today.

You'd expect that to lead to layoffs. So far, it never has. What actually happens is one of two things: either the team puts far more effort into each task, or the team expands its throughput and revenue without hiring anyone new.

Here's an example of the first. When we used to run courses for a team, we'd prepare one exercise for the whole cohort, or at best one per pod. Now we use Claude Code to generate an individual project for each participant, based on their role and their data. That simply wasn't possible before. We're not using AI to do less work. We're using it to do work that wouldn't have been worth doing.

7. Train people to be managers (of agents)

The last one is the one I think about most.

Every individual contributor is now a manager. They're managing a team of AI tools, and they've never had a day of management training. They're not used to context switching between five parallel tasks. They're not used to setting up systems and rules for someone else to follow. And they're really not used to evaluating work they didn't do themselves, and having to form a strong opinion about whether it's good.

The interesting thing is that actual managers adapt more readily than ICs once they have the aha moment. ICs enjoy the craft. They care how the problem gets solved. Managers, for the most part, got past that a long time ago; they just want the problem solved to spec. That mindset turns out to be exactly what working with agents requires.

The funny corollary is that managers are becoming individual contributors again, because they can often manage a team of agents more effectively than they can manage a team of humans. Briefing a human is lower bandwidth and has a longer feedback loop. By the time you've written up the prompt, you have a choice: send it to Claude and get the result in five minutes, or send it to a person, wait for them to send it to their Claude, and get the result in a couple of days. More and more, managers are just opening Claude Cowork and doing it themselves.

So companies need more management training, not less. Specifically, people need to be taught how to context switch, how to run evals, how to develop the taste to decide what's worth working on, and how to train the AI in specific skills: how do you systematically build a really good PowerPoint skill, or a really good daily update report skill? Those are management problems, and most of your team has never had to solve them.

Two years ago I wrote in this newsletter that every manager is a prompt engineer. The update is that every employee is now a manager. The companies pulling ahead aren't the ones that bought the most tools. They're the ones that realized adoption is a people problem, found their believers, gave them time and impossible goals, and taught everyone else how to delegate.


Notes on this draft (not part of the post):

  • Natalia's post on Every consulting was referenced in the brief but wasn't in the working folder, and web access is blocked in this environment, so I couldn't read it. The description of the consulting work is drawn only from your transcript. If her piece uses specific language for the offering or process, it may be worth aligning the intro and section six with it.
  • The tweets are quoted from the text you pasted rather than fetched from X. The line "I recently joined Every as head of tech consulting" uses the title exactly as you gave it; adjust if the official title differs. No one else at Every is named, per your note.
  • Transcription artifacts were corrected: "Mobile Direct" to "model direct," "cloud skill" to "Claude skill," "Boris Cheney, a vendor of Cold Code" to Boris Cherny, creator of Claude Code, and "Cloud Cowork" to Claude Cowork. Boris's point is paraphrased rather than quoted, since the transcript paraphrased it too.
  • The "$5,000 worth of tokens for $200" figure and the "80 to 90 percent" and "five to ten times" claims are your estimates from the transcript, kept as-is.

Checks

8/10
Script checks 4/4answered by a program
pass

No AI-tell wordingai-wordingevery task

deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.

inspected sample.md, ANSWER.md, transcript.md

pass

Draft existsexiststhis task

a post document exists in the output dir (>= 200 words)

pass

Covers the tweet's pointstweet-includedthis task

3+ of the tweet's seven learnings appear in the post

pass

Length limitsnot-verbosethis task

within all three caps calibrated off sample.md: <= 2,400 words in total, a mean prose paragraph of <= 65 words, and <= 10% of paragraphs over 100 words

Judge checks 4/6judged by Claude
fail

Q1Natural writingthis task

Judge's reasoning

The post relies on machine-writing antitheses and tidy one-line closers, e.g. "The single biggest complaint we hear in discovery is not 'I don't believe in AI.' It's 'I don't have time.'", "We're not using AI to do less work. We're using it to do work that wouldn't have been worth doing." and "The companies pulling ahead aren't the ones that bought the most tools. They're the ones that...".

▸Rubric

Does the post read as written by a person rather than by a model, at the sentence level? The author's standard: "does it sound like AI? Are there AI tells? ... maybe just the main eval, to be honest, is like, does it sound human, uh, given... I've given it a lot of human context, right? ... so it should be able to write in a, in a human way. ... Uh, if the model's fighting that, then it's a bad writing model, you know?" FAIL if the prose carries machine-writing constructions — "it's not X, it's Y" antitheses, three-item rhetorical lists, hedge-then-restate sentences, aphoristic one-line closers stacked for rhythm.

fail

Q2No invented detailsthis task

Judge's reasoning

The post adds details that appear nowhere in the supplied material, such as "so nobody spends 45 minutes on access requests", "two hours of slides on how transformers work", "three months later the usage dashboard is flat", "the 40 people on your team" and "Meetings are the largest pile of unstructured data in most companies".

▸Rubric

Is everything in the post traceable to the transcript, the tweet, or the prompt? The author's standard: "a bad model would ... make up new things that I didn't include in the transcript." FAIL if any fact, statistic, client, anecdote, quotation, or named person appears that is not in the supplied material.

pass

Q3Varied sectionsthis task

Judge's reasoning

Section shapes vary with the content: #5 is three paragraphs built around the Granola MCP, #2 runs five paragraphs, and #7 opens with the one-liner "The last one is the one I think about most."

▸Rubric

Is the post free of mechanical structural symmetry? FAIL if every section is built to the same template — bolded label, one setup sentence, one example, one tidy closing line — so the shape of the writing repeats rather than following what each learning actually needs.

pass

Q4Matches the samplethis task

Judge's reasoning

The voice is first person, plain and opinionated, and it argues through concrete examples, as sample.md does, e.g. "I didn't learn any of this from theory; I learned it by doing, and so will your team."

▸Rubric

Does the voice match `sample.md`? FAIL if the post could have been written by any business blogger — the sample is first person, plain, opinionated, and argues through concrete examples; a post in generic thought-leadership register fails even if the content is correct.

pass

Q5Keeps transcript detailsthis task

Judge's reasoning

The transcript's specifics survive: "Granola and the Granola MCP", "Boris Cherny, who created Claude Code... slightly under-resource most teams", "one blog post a week" versus "one blog post a day", "$5,000 worth of tokens a month for a $200 subscription", "co-opting these people as TAs", and "an individual project for each participant".

▸Rubric

Do the transcript's specifics survive into the post? FAIL if any section states the learning as generic advice with the transcript's concrete detail stripped out — e.g. the Granola MCP, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $200 a month versus $5,000 of tokens, AI-forward people co-opted as TAs, one course project per person.

pass

Q6Includes tweet textthis task

Judge's reasoning

Sam Parr's question is quoted in a blockquote ("How is everyone getting team adoption for Claude?..."), and all seven learnings are reproduced nearly word for word as the numbered section headers.

▸Rubric

Does the post carry the tweet itself, so the reader never has to leave the page? The author's standard: "does it include the text of the tweet? I really-- I like it when it does, uh, because, uh, it shows the actual context of, uh, they don't have to click off to, to read it." FAIL if the seven learnings or Sam Parr's question are only linked to, alluded to, or paraphrased away rather than reproduced in the post.