Mike's Checks/gpt-6-astra/03 writeup
03 writeup
gpt-6-astraCodex CLIhigh effortrun 22 Sep 202668,389 tokens
▸Instructions — what the model was asked
Hey, I'm trying to write a post on my experience with Every Consulting. The consulting so far, I just joined as the head of tech consulting. I didn't tell you who runs consulting. I'm going to provide a transcript of me talking about my experience. I was also going to provide a tweet, which is just a list of the learnings that I've had so far for companies that were struggling to adopt AI. It was in response to Sam Par, who was asking, "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"
I need to write this in my writing style. I've included some writing of my own as a sample. I also have included a post on consulting at Every, so from in Natalia, you get a sense of what consulting looks like at Every. But yeah, focus mostly on the specifics of the interview that I provided with the context of that tweet.
https://x.com/hammer_mt/status/2032591631413567873?s=20
Learnings:
- buy the model direct not 3rd party tools
- raise the ceiling not the floor
- workshops with 50%+ build time
- assign impossible tasks
- mandatory AI note taking + MCP connector
- map workflows and systematically automate
- train people to be managers (of agents)
In response to:
https://x.com/thesamparr/status/2032477081695584670
How is everyone getting team adoption for Claude?
I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas.
But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?
Answer
How to get your team to actually use AI
Buying everyone a Claude subscription is easy. Getting them to change how they work is much harder. You can give people access to the same tools you use, show them what’s possible, and still find that they go back to doing everything manually.
Sam Parr recently asked how people are getting their teams to adopt Claude. He spends time on Twitter keeping up with AI, but he doesn’t expect his employees to do the same. How do you get people to use these tools effectively without making following AI news part of their job?
I recently joined Every as head of tech consulting, and this is a problem we’re working on with companies. I replied with seven things I’ve learned so far. They cover which tools to buy, how to run training, and what happens when someone who has spent their career doing the work suddenly has to delegate it to an agent.
One of the biggest complaints we hear is that people haven’t had time to try the tools. They have a job to do, and learning a different way to do it is another task on an already full list. If you want them to experiment, you have to make room for it—and give them something worth experimenting on.
The first decision is what to give them access to. My default is to buy directly from the model provider. I see companies spending time evaluating products that use Claude or another model under the hood, when they could often get what they need by configuring the model’s own tools.
Every product comes with someone else’s opinions about how work should happen. Those opinions might be useful, but they can also get in your way. If you have a particular approach to research, writing, or reporting, it can be easier to put that approach into a Claude skill—a reusable set of instructions—than to make it fit another company’s product.
This has changed how I think about building versus buying. Creating your own tool used to imply a software project and someone to maintain it. Sometimes now it means writing down how you want a task done, giving Claude examples, and improving the instructions when the results fall short. For those tasks, building becomes a much more reasonable option.
I also think it’s difficult to compete with the companies building the models. They can develop their tools alongside the models that power them. There are third-party products I like, but I want a specific reason to add another one. Start with direct access, find out where it falls short, and then decide what else you need.
Once people have access, the next question is where to put your effort. I’ve found it more useful to raise the ceiling than the floor.
There are probably people in your company already using Claude Code in their spare time. Some may be waiting for permission to use it at work. Find those people and give them support. They already want to experiment; you can help them get further by giving them tools, time, and access to the problems that matter.
It’s much harder to convince someone who hasn’t seen the value yet. For some people, the resistance is emotional. They’ve spent years getting good at something, and now their employer is asking them to hand part of it over to AI. Telling them that everyone has to use it doesn’t resolve that.
We’ve had success bringing people who already use AI into our courses as teaching assistants. Their colleagues get to see someone they know accomplishing more with the tools. Those people also get more visibility. If using AI well helps someone get time with senior management or advance in their career, it gives everyone else a reason to pay attention. You get the immediate benefit of helping an enthusiast do more, and a practical example for everyone still deciding whether it’s worth learning.
You still need to train the rest of the team. But at least half of a workshop should be time spent building.
I learned to use AI by doing things with it. Some theory helps explain what to try, but watching someone else use Claude isn’t the same as getting it to do something useful yourself. The workshop gives people a couple of hours when experimenting is the work they’re supposed to be doing.
That time needs preparation. Make sure people have access to the tool before they start. Give them synthetic data to work with, or make sure the connectors they need are available. Choose a task that stretches them beyond what they would normally attempt. You want them to spend the session discovering what they can do, with a facilitator available when they get stuck.
The moment that matters is when someone builds something they didn’t think they could build. That gives them a reason to open the tool again after the workshop is over.
You can encourage the same kind of discovery in their regular work by assigning impossible tasks. By “impossible,” I mean something they couldn’t reasonably accomplish with their existing way of working.
If your goal is to write one blog post a week, you can probably do it manually. You might use AI to suggest a title, but there’s little pressure to rethink the process. If your goal is to get to one good post a day, you have to look more carefully at where the time goes. How much research can AI help with? What can it do during editing? Which decisions still need your attention?
The way you set that goal matters. “Starting tomorrow, produce five times as much” is a good way to make people miserable. I’d frame it as a target we’re working toward: What would need to change for this to become possible? What have we tried? How much closer are we getting?
That gives people a reason to experiment and time to develop a better process. It also gives the experiments a direction. They’re trying to solve a problem they understand, and they can judge whether the tool is helping.
One of the most useful changes we’ve made on the consulting team is much simpler: everyone uses Granola for AI meeting notes, with its MCP connector making those notes available to our AI tools.
So much of the context for our work lives in conversations. When I’m creating a curriculum or putting together a proposal, I’ll realize that something we discussed in a meeting is relevant. Having those notes available means I can pull that context into the task as I’m working on it.
An update email becomes a straightforward request: We just had a meeting about this project; use the discussion to draft an update for the people who need to know. The useful information is already there. I don’t have to reconstruct the meeting from memory before the model can help me.
I think a huge amount of AI’s value comes from this ability to take unstructured information and turn it into something useful. A conversation becomes a proposal, a set of requirements, or material for a course. Making AI note-taking and the connector a standard part of how the team works gives everyone a much better starting point.
From there, we can get systematic about the rest of the job. In discovery calls, we ask people to walk us through their day: what they do, which tools they use, which tasks repeat, and where they get stuck. We turn that into a structured Google Sheet and work down the list.
The ambition is to look at everything. Can AI do this task? Can it at least do a first pass? Can we create a skill for it so the person doesn’t have to explain the same requirements every time?
That last question matters if you want adoption to last. An impressive demonstration is useful, but a repeatable process gives someone a reason to keep using the tool. Each recurring task becomes another opportunity to capture how they like to work and make the next attempt easier.
So far, the result I’ve seen has been people putting more effort into each task, or teams increasing their output without needing to hire more people. Our own courses are a good example. We used to prepare one exercise for an entire cohort. With more time, we might prepare a different exercise for each team. Now we can use Claude Code to create an individual project for each participant.
We’re spending that capability on making the course better. Each person can work on something more relevant to them. That level of preparation would have been impractical before, and it’s a useful reminder that saving time on a task can change how much you’re willing to attempt.
The final lesson is that people need training in management. An individual contributor who starts delegating work to agents suddenly has to do things they may never have been taught: write a clear brief, provide context, set standards, and evaluate work they didn’t produce themselves.
In my guide to prompt engineering, I wrote that every manager is a prompt engineer. You give direction when you brief a designer or explain a task to an intern. Working with agents makes those skills relevant to people who have never managed anyone.
This can be uncomfortable if you enjoy the craft of doing the work yourself. You know how to make something good, but reviewing someone else’s attempt requires you to explain what’s wrong with it. “I wouldn’t have done it that way” is different from identifying whether it meets the brief. You need to develop judgment about what to delegate, what to check, and when the result is good enough.
There’s a funny reversal happening at the same time: managers can become individual contributors again. Once you’ve written a detailed brief, you can sometimes give it directly to an agent and get a result to review in minutes. For some tasks, that’s a much shorter feedback loop than sending the brief to someone who will then use their own AI tools to do the work.
That’s why I think companies need to teach more than prompting. People need practice managing several pieces of work, evaluating outputs, and turning their feedback into instructions the agent can use next time. Building a good PowerPoint skill or a reliable daily update skill means making your standards explicit enough that something else can follow them.
Your team doesn’t have to follow every new release on Twitter to learn those skills. Give them time to build, access to the context they need, and a useful problem that makes them stretch. Then help them turn what works into a process they can use again tomorrow.
Drafting note: I treated this as a first-person article for Every, using the supplied transcript and writing sample. The referenced Natalia post was not present, and web retrieval was unavailable, so I have not added details about who runs consulting. The tweet context and links come from your prompt. I omitted the transcript’s unverified pricing comparison and third-party attribution.
Checks
7/10No AI-tell wordingai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
inspected sample.md, ANSWER.md, transcript.md
Draft existsexiststhis task
a post document exists in the output dir (>= 200 words)
Covers the tweet's pointstweet-includedthis task
3+ of the tweet's seven learnings appear in the post
Length limitsnot-verbosethis task
within all three caps calibrated off sample.md: <= 2,400 words in total, a mean prose paragraph of <= 65 words, and <= 10% of paragraphs over 100 words
Q1Natural writingthis task
Judge's reasoning
The post ends on a stacked three-item rhetorical list, "Give them time to build, access to the context they need, and a useful problem that makes them stretch," and uses the same pattern earlier in "They cover which tools to buy, how to run training, and what happens when...".
▸Rubric
Does the post read as written by a person rather than by a model, at the sentence level? The author's standard: "does it sound like AI? Are there AI tells? ... maybe just the main eval, to be honest, is like, does it sound human, uh, given... I've given it a lot of human context, right? ... so it should be able to write in a, in a human way. ... Uh, if the model's fighting that, then it's a bad writing model, you know?" FAIL if the prose carries machine-writing constructions — "it's not X, it's Y" antitheses, three-item rhetorical lists, hedge-then-restate sentences, aphoristic one-line closers stacked for rhythm.
Q2Keeps transcript detailsthis task
Judge's reasoning
The drafting note admits "I omitted the transcript's unverified pricing comparison and third-party attribution," so the $200 a month versus $5,000 of tokens comparison and Boris Cherny on under-resourcing teams are both stripped out.
▸Rubric
Do the transcript's specifics survive into the post? FAIL if any section states the learning as generic advice with the transcript's concrete detail stripped out — e.g. the Granola MCP, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $200 a month versus $5,000 of tokens, AI-forward people co-opted as TAs, one course project per person.
Q3Includes tweet textthis task
Judge's reasoning
The seven learnings are only linked ("I [replied with seven things I've learned so far]") and Sam Parr's question is paraphrased ("Sam Parr recently [asked how people are getting their teams to adopt Claude]") rather than reproduced.
▸Rubric
Does the post carry the tweet itself, so the reader never has to leave the page? The author's standard: "does it include the text of the tweet? I really-- I like it when it does, uh, because, uh, it shows the actual context of, uh, they don't have to click off to, to read it." FAIL if the seven learnings or Sam Parr's question are only linked to, alluded to, or paraphrased away rather than reproduced in the post.
Q4Varied sectionsthis task
Judge's reasoning
There are no bolded labels or repeated section template; the learnings flow as continuous prose of different lengths, from the four paragraphs on buying direct to the long management section with its "funny reversal happening at the same time" aside.
▸Rubric
Is the post free of mechanical structural symmetry? FAIL if every section is built to the same template — bolded label, one setup sentence, one example, one tidy closing line — so the shape of the writing repeats rather than following what each learning actually needs.
Q5Matches the samplethis task
Judge's reasoning
The post is first person, plain and opinionated, and argues through concrete examples like the sample ("I learned to use AI by doing things with it"), and it even reuses the sample's "every manager is a prompt engineer" line in the author's own framing.
▸Rubric
Does the voice match `sample.md`? FAIL if the post could have been written by any business blogger — the sample is first person, plain, opinionated, and argues through concrete examples; a post in generic thought-leadership register fails even if the content is correct.
Q6No invented detailsthis task
Judge's reasoning
Every claim traces to the transcript, prompt or sample; even the guide link and the "every manager is a prompt engineer" line come from the supplied sample.md.
▸Rubric
Is everything in the post traceable to the transcript, the tweet, or the prompt? The author's standard: "a bad model would ... make up new things that I didn't include in the transcript." FAIL if any fact, statistic, client, anecdote, quotation, or named person appears that is not in the supplied material.