Mike's Checks/gpt-5.6-luna/03 writeup
03 writeup
gpt-5.6-lunaCodex CLIhigh effortrun 22 Aug 2026
▸Instructions — the case's current instructions; none were saved with this result
Hey, I'm trying to write a post on my experience with Every Consulting. The consulting so far, I just joined as the head of tech consulting. I didn't tell you who runs consulting. I'm going to provide a transcript of me talking about my experience. I was also going to provide a tweet, which is just a list of the learnings that I've had so far for companies that were struggling to adopt AI. It was in response to Sam Par, who was asking, "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"
I need to write this in my writing style. I've included some writing of my own as a sample. I also have included a post on consulting at Every, so from in Natalia, you get a sense of what consulting looks like at Every. But yeah, focus mostly on the specifics of the interview that I provided with the context of that tweet.
https://x.com/hammer_mt/status/2032591631413567873?s=20
Learnings:
- buy the model direct not 3rd party tools
- raise the ceiling not the floor
- workshops with 50%+ build time
- assign impossible tasks
- mandatory AI note taking + MCP connector
- map workflows and systematically automate
- train people to be managers (of agents)
In response to:
https://x.com/thesamparr/status/2032477081695584670
How is everyone getting team adoption for Claude?
I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas.
But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?
Answer
How to Get a Team to Actually Use AI
Sam Parr recently asked how companies are getting their teams to use Claude effectively without asking everyone to spend all their time on Twitter learning about AI.
I have spent a lot of time on Twitter, as has Sam. We see all the new models, tools, and fringe ideas as they appear. We are usually on top of what is happening—or at least close enough to feel like we are.
Most people on a team are not like that. They probably should not be. They have a job to do, and “keep up with AI” is not a particularly compelling addition to their job description.
I recently joined Every as the head of tech consulting, where we help companies figure out how to use AI in their actual workflows. Here are the things that have worked best so far.
Buy the model directly
The first thing I have noticed is that companies often start by evaluating tools that use Claude, Codex, or Gemini under the hood.
That can work, but the tool also comes with its own product decisions, opinions, and limitations. It is often quicker to create a Claude Code skill that reflects your company’s own way of working. You get more flexibility, and you can configure and operate it yourself.
There is also a basic incentive problem for third-party tools. The model companies know which releases are coming. They can build their internal tools around those releases and train people on the environments they control. I appreciate the effort that products like Cursor put into keeping up—they are clearly a good product organization—but I am not sure how many companies can compete with a model provider offering thousands of dollars of tokens for the price of a relatively small monthly subscription.
This is not an absolute rule. Some third-party products will be the right choice. But as a general principle, I would start with the model provider’s own tools. They tend to be more flexible, closer to the cutting edge, and, increasingly, less expensive.
Raise the ceiling, not the floor
The usual approach to AI adoption is to tell everyone that they need to use AI. Then the company buys a set of tools and expects adoption to follow.
It usually does not.
Even on pain of death, some people are emotionally unwilling to use AI when they are told that they have to. That response is understandable. It feels like a threat to the craft they have spent years developing, or to the value they believe they provide.
The better approach is to raise the ceiling rather than the floor.
Find the people who are already AI-forward. They may be using Claude Code in their spare time, or experimenting with tools nobody else on the team has heard of yet. Give them support, time, and access to senior management. Let them work on problems that would have been difficult without AI.
Someone who is already a believer can often produce five or ten times as much as someone who has not yet seen the magic. It is much easier to enable that person than to persuade someone else to become a believer from scratch.
We have also had success turning these people into teaching assistants for the courses we run with the rest of the team. This creates a useful loop: the AI-forward people get more responsibility and visibility, while everyone else gets to see a peer using the tools successfully.
When people see that the person using AI aggressively is the person getting ahead in their career, getting promoted, or getting more time with senior management, that is a much more effective motivator than a mandatory training email.
Make workshops mostly build time
You should run workshops, but at least half of the workshop should be build time.
People do not want to sit through slides about what AI might eventually do. And they do not need a complete theory of AI before they can start using it. The way I learned was by doing things and seeing what worked.
Some guided theory is useful. It gives people enough context to understand why the exercise matters. But the real obstacle we hear from people is that they have not had time to check out the tools. They do not have spare hours in the workday to learn something new, build something, and stretch beyond their normal responsibilities.
Give them those hours.
In a good workshop, people have access to the model, the connectors they need, and some synthetic data to work with. They are asked to build something that is slightly outside their normal domain. The facilitator is there to unblock them, not to perform the task for them.
That is where the aha moment tends to happen. Someone goes from hearing that AI can help with their work to watching it complete a meaningful part of the work in front of them.
Assign impossible tasks
An impossible task is one that would not have been possible to complete without AI—or at least would have required a much larger team.
Boris Cherny, one of the people behind Claude Code, has made a similar point: you should slightly under-resource most teams. That forces them to ask whether there is another way to get the work done.
I prefer to make this more explicit and strategic. Choose goals that require people to figure out how to use AI.
If your goal is to write one blog post a week, you can probably do that manually. If your goal is to write one a day, you will need to use AI in the research, drafting, editing, or distribution process.
You do not need to announce that everyone must produce one post a day starting tomorrow. Instead, make the goal clear: we want to get to the point where we can produce one a day. What would need to change to make that possible?
Once people know the goal, they start experimenting. They look for the bottlenecks. They work out which parts of the process can be delegated. Adoption becomes a means of achieving something they already want, rather than a separate initiative imposed on top of their work.
Make AI note-taking mandatory
One of the biggest unlocks for our consulting team has been mandatory AI note-taking, along with an MCP connector to those notes.
Everyone on the team uses Granola and the Granola MCP. That makes it much easier to go from a meeting to a useful piece of work:
We just had a meeting about this topic. Create an update email and send it to this person.
The model can use the context from the meeting instead of asking someone to reconstruct it from memory or search through a collection of notes.
This is probably 80 or 90 percent of the value of AI in many businesses: extracting information from unstructured data and turning it into something useful.
There are many tasks where I realize I need context from a meeting we had three weeks ago. If that context is available through an MCP connector, it can be pulled directly into a curriculum, a proposal, an update report, or a follow-up email.
The important part is not just taking notes. It is making the notes available at the moment when someone is trying to do the next piece of work.
Map the workflow and automate it systematically
The next step is to stop treating AI as a collection of isolated prompts and start mapping the workflow.
For our discovery calls, we ask people what they do day to day, which tools they use, what tasks they repeat, and where the pain points are. We turn that information into a structured Google Sheet with the things we need to solve for daily and weekly.
Then we work down the list.
The goal is to get as much of the work off the person as possible. Even if AI only does the first pass, and even if there is a separate skill for each type of task, the person can often achieve five or ten times the throughput they had before.
So far, this has not meant letting people go. It has usually meant one of two things: the team puts much more effort into each task, or it expands the revenue and throughput of the team without needing to hire more people.
For example, when we used to run courses, we might prepare one task for the whole cohort, or one task for each team pod. Now we can use Claude Code to create an individual project for every person taking part.
That is a strange but important outcome. AI is not just helping us automate the work. It is letting us invest more in each individual piece of work than was previously economical.
Train people to be managers of agents
The final shift is the most important one: people need to learn how to manage agents.
Many individual contributors are now operating like managers. They are delegating work to AI, checking the results, switching between multiple tasks, setting rules, and deciding whether the output is good enough. But they have not necessarily received any management training.
They may not be used to context switching. They may not know how to set up systems around recurring work. They may not have developed the taste required to evaluate something they did not personally make. And they may not know how to give the model a strong opinion about what good looks like.
Managers sometimes take to this more quickly than individual contributors. Individual contributors often care deeply about the craft and about how a problem is solved. Managers have usually spent more time focusing on whether the problem was solved to the required specification.
The irony is that managers are also becoming individual contributors again. It can be easier for a manager to use Claude Cowork to complete a task than to brief a human, wait for the work to come back, and then have that person use Claude Cowork themselves.
By the time you have written the prompt, you may be able to decide whether to send it to Claude and get a response in five minutes, or send it to a human and wait a couple of days for the response to come back from Claude.
That does not mean human teams are obsolete. It means the management interface is changing.
Companies need to teach people how to evaluate outputs, how to develop taste, how to decide what is worth working on, and how to turn a good process into a reusable skill. What does a good PowerPoint skill look like? What does a good daily update report skill look like? How do you give an agent enough context to perform reliably without checking every line of its work?
Those are management questions.
The companies that get AI adoption right will not be the ones that force everyone to spend more time learning about AI. They will be the ones that give their best users room to run, give everyone else a concrete reason to experiment, and redesign the work so that using AI is the most natural way to achieve the goal.
That is how you get a team to use Claude effectively without asking them to become AI people on Twitter.
Context note: This draft is based primarily on the supplied interview transcript and the seven learning points. The referenced Every consulting article was not included in the workspace, so I kept the description of Every Consulting limited to what appears in the transcript and prompt.
Checks
5/10Draft existsexists
a post document exists in the output dir (>= 200 words)
Covers the tweet's pointstweet-included
3+ of the tweet's seven learnings appear in the post
Length limitsnot-verbose
within all three caps calibrated off sample.md: <= 2,400 words in total, a mean prose paragraph of <= 65 words, and <= 10% of paragraphs over 100 words
ai-tells-grep Retired
Retired check; no longer defined in the repo.
Q1Natural writing
Judge's reasoning
Machine-writing tells throughout: the not-X-but-Y antithesis ("That does not mean human teams are obsolete. It means the management interface is changing."; "The important part is not just taking notes. It is making the notes available..."), tricolons ("They tend to be more flexible, closer to the cutting edge, and, increasingly, less expensive."), stacked one-line closers ("It usually does not.", "Give them those hours.", "Those are management questions."), and a total absence of contractions across ~1,800 words.
▸Rubric
Does the post read as written by a person rather than by a model, at the sentence level? The author's standard: "does it sound like AI? Are there AI tells? ... maybe just the main eval, to be honest, is like, does it sound human, uh, given... I've given it a lot of human context, right? ... so it should be able to write in a, in a human way. ... Uh, if the model's fighting that, then it's a bad writing model, you know?" FAIL if the prose carries machine-writing constructions — "it's not X, it's Y" antitheses, three-item rhetorical lists, hedge-then-restate sentences, aphoristic one-line closers stacked for rhythm.
Q2Varied sections
Judge's reasoning
All seven learnings get the identical shape — H2 label, setup, one transcript example, tidy summarizing closer ("...less expensive." / "...than a mandatory training email." / "...in front of them." / "...imposed on top of their work." / "...trying to do the next piece of work." / "That is a strange but important outcome..." / "Those are management questions.") — so the sections repeat a template rather than following what each point needs.
▸Rubric
Is the post free of mechanical structural symmetry? FAIL if every section is built to the same template — bolded label, one setup sentence, one example, one tidy closing line — so the shape of the writing repeats rather than following what each learning actually needs.
Q3Matches the sample
Judge's reasoning
Where the sample is loose and contraction-heavy with sharp turns ("Everybody you work with is, by definition, already at AGI—but they still need to be prompted."), this reads as hedged generic thought leadership: "This is not an absolute rule. Some third-party products will be the right choice." and the closer "The companies that get AI adoption right will not be the ones that force everyone to spend more time learning about AI."
▸Rubric
Does the voice match `sample.md`? FAIL if the post could have been written by any business blogger — the sample is first person, plain, opinionated, and argues through concrete examples; a post in generic thought-leadership register fails even if the content is correct.
Q4Keeps transcript details
Judge's reasoning
The buy-direct section strips the transcript's numbers into vagueness — "a model provider offering thousands of dollars of tokens for the price of a relatively small monthly subscription" instead of Anthropic's $5,000 of tokens versus a $200/month subscription.
▸Rubric
Do the transcript's specifics survive into the post? FAIL if any section states the learning as generic advice with the transcript's concrete detail stripped out — e.g. the Granola MCP, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $200 a month versus $5,000 of tokens, AI-forward people co-opted as TAs, one course project per person.
Q5Includes tweet text
Judge's reasoning
Sam Parr's question is only paraphrased ("Sam Parr recently asked how companies are getting their teams to use Claude effectively without asking everyone to spend all their time on Twitter") and the seven learnings are never reproduced as the tweet's text; the post's only blockquote is the Granola prompt example.
▸Rubric
Does the post carry the tweet itself, so the reader never has to leave the page? The author's standard: "does it include the text of the tweet? I really-- I like it when it does, uh, because, uh, it shows the actual context of, uh, they don't have to click off to, to read it." FAIL if the seven learnings or Sam Parr's question are only linked to, alluded to, or paraphrased away rather than reproduced in the post.
Q6No invented details
Judge's reasoning
Every named element traces back — Granola/Granola MCP, Cursor, Boris Cherny on under-resourcing, Claude Cowork, the Google Sheet, TAs in courses, one project per person — with no invented clients, stats, or quotations.
▸Rubric
Is everything in the post traceable to the transcript, the tweet, or the prompt? The author's standard: "a bad model would ... make up new things that I didn't include in the transcript." FAIL if any fact, statistic, client, anecdote, quotation, or named person appears that is not in the supplied material.