Mike's Checks/deepseek-v4-flash/03 writeup
03 writeup
deepseek-v4-flashCodex CLIhigh effortrun 4 Sep 2026116,687 tokens
▸Instructions — the case's current instructions; none were saved with this result
Hey, I'm trying to write a post on my experience with Every Consulting. The consulting so far, I just joined as the head of tech consulting. I didn't tell you who runs consulting. I'm going to provide a transcript of me talking about my experience. I was also going to provide a tweet, which is just a list of the learnings that I've had so far for companies that were struggling to adopt AI. It was in response to Sam Par, who was asking, "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"
I need to write this in my writing style. I've included some writing of my own as a sample. I also have included a post on consulting at Every, so from in Natalia, you get a sense of what consulting looks like at Every. But yeah, focus mostly on the specifics of the interview that I provided with the context of that tweet.
https://x.com/hammer_mt/status/2032591631413567873?s=20
Learnings:
- buy the model direct not 3rd party tools
- raise the ceiling not the floor
- workshops with 50%+ build time
- assign impossible tasks
- mandatory AI note taking + MCP connector
- map workflows and systematically automate
- train people to be managers (of agents)
In response to:
https://x.com/thesamparr/status/2032477081695584670
How is everyone getting team adoption for Claude?
I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas.
But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?
Answer
How to Actually Get Your Team to Use AI
Sam Parr asked a question that's been rattling around my head for weeks: "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"
It's the right question. Most companies are stuck here. They've bought the licenses, done the training, sent the memos. And still, adoption flatlines. I just joined Every as head of tech consulting, and I've been spending my days inside companies wrestling with exactly this problem. Here's what I've learned so far.
Sam's frustration is familiar. You're on Twitter, you see all the AI stuff popping up, you're testing the fringe ideas. You know what's possible. But your team? They're not spending their days on Twitter. They shouldn't be. So how do you bridge that gap?
After a bunch of consulting engagements, I've landed on seven principles. They're not theoretical — we've been running them inside Every Consulting, and they're starting to work.
1. Buy the model direct, not third-party tools
People love to evaluate tools that wrap Claude or Codex or Gemini. And look, Cursor is a good product org. They're doing great work. But I don't know how you compete with Anthropic offering $5,000 worth of tokens a month versus a $200 subscription.
The problem with third-party tools is they come with someone else's opinions baked in. Their product decisions, their hang-ups, their prioritization. When you buy direct, you get flexibility. You can schedule around model releases, train on the cutting edge, configure your own approach. The core model companies know what's coming — they build their internal tools to handle those releases. If you're one layer removed, you're always catching up.
There's a real move toward build, not buy. It's cheaper, faster, and more flexible.
2. Raise the ceiling, not the floor
Most companies approach adoption like a mandate: "Everyone must use AI now. We bought you the tools. Use them."
This doesn't work. Even on pain of death, people won't adopt AI if they're not ready. It's a self-preservation thing. The stick backfires every time.
Instead, use the carrot. Find the people who are already AI-forward — the ones using Codex in their spare time, the ones who are genuinely curious. Give them support. Give them access. Show the rest of the org that these people are the ones getting promoted, getting facetime with senior management, getting the interesting projects.
Someone who's really AI-pilled can do five to ten times more than someone who hasn't seen the magic yet. You can't force the magic. But you can make it visible.
3. Workshops with 50%+ build time
Nobody wants to sit through slides about AI theory. The way people actually learn this stuff is by doing.
Our workshops are at least half build time. We give people synthetic data, working connectors, and a real problem to solve. The facilitator is there to unblock, not to lecture. And the biggest barrier we hear from teams is simple: they don't have time. They don't have space in their day to experiment.
A workshop gives them that space. Two hours where they're expected to build something new, go outside their normal domain, and actually touch the tool. That's where the aha moment happens. That's where it clicks.
4. Assign impossible tasks
If your goal is to write one blog post a week, you can do that manually. If your goal is to write one a day, you're going to need AI — for research, for editing, for structuring.
Boris Cheney (one of the founders of Codex) said something similar: slightly under-resource your teams, and they'll figure out that AI is the only way through. I like to make it more explicit. Choose goals that are literally impossible without AI. Don't frame it as "starting today, you must use AI." Frame it as "our goal is to get to ten times the output — how do we get there?"
People are strategic. When they know the destination, they'll figure out the route.
5. Mandatory AI note-taking + MCP connector
This is the single biggest unlock we've found. Everyone on the consulting team uses Granola with the Granola MCP. Every meeting gets recorded, every note gets structured, everything is queryable.
The reason this matters: 80-90% of AI's value right now is extracting information from unstructured data and structuring it in a useful way. How many times have you sat down to a task and realized you need context from a meeting three weeks ago? With an MCP connector, that context is at your fingertips. You can generate update emails, pull context for proposals, build curriculum — all from the record of conversations you were already having.
It seems small. It's not. It's the foundation everything else builds on.
6. Map workflows and systematically automate
We have a process for discovery calls: talk about what someone does day-to-day, what tools they use, what their pain points are. Turn that into a structured spreadsheet. Then systematically work down the list, building a skill for each type of task.
The assumption is bold: we're trying to get 100% of the routine work off the humans. If a skill can do a first pass on a task — even an imperfect one — the human can operate at five to ten times their current throughput.
What's interesting is what happens next. Nobody gets let go. Instead, teams either invest way more in each individual task (we went from one group exercise per course to a personalized project per participant), or they expand their output without hiring. The constraint shifts from headcount to ambition.
7. Train people to be managers — of agents
Here's the thing that's crept up on everyone: every individual contributor is now a manager. They're managing AI tools, and they don't have management training. They don't know how to context-switch. They don't know how to set up systems and rules. They don't know how to evaluate work they didn't personally produce.
The funny twist? Managers actually take to this faster than ICs. Managers have long since stopped caring about how a problem is solved — they just want it solved to spec. And then it gets weirder: managers start becoming individual contributors again, because it's often faster to vibe-code a task than to brief a human, wait for them to brief their Claude, and wait for the response to come back a couple of days later.
We need to teach people how to manage agents. How to do evals. How to develop taste for what to work on. How to systematically train a skill — "write a really good PowerPoint deck" or "produce a daily update report." That's the management training nobody's offering, and it's the one that matters most.
I don't have this all figured out. Every consulting engagement teaches me something new. But these seven principles are the ones I keep coming back to. If you're trying to get your team to actually use AI — not just buy licenses and hope — start here.
I'm at @hammer_mt on X. I'd love to hear what's working for you.
Checks
6/10Draft existsexists
a post document exists in the output dir (>= 200 words)
Covers the tweet's pointstweet-included
3+ of the tweet's seven learnings appear in the post
Length limitsnot-verbose
within all three caps calibrated off sample.md: <= 2,400 words in total, a mean prose paragraph of <= 65 words, and <= 10% of paragraphs over 100 words
ai-tells-grep Retired
Retired check; no longer defined in the repo.
Q1Natural writing
Judge's reasoning
Dense machine-writing tells throughout: anaphoric triples ("They don't know how to context-switch. They don't know how to set up systems and rules. They don't know how to evaluate work they didn't personally produce."), stacked aphoristic closers ("It seems small. It's not. It's the foundation everything else builds on.", "That's where the aha moment happens. That's where it clicks."), and antithesis-plus-punchline ("You can't force the magic. But you can make it visible.", "The constraint shifts from headcount to ambition.").
▸Rubric
Does the post read as written by a person rather than by a model, at the sentence level? The author's standard: "does it sound like AI? Are there AI tells? ... maybe just the main eval, to be honest, is like, does it sound human, uh, given... I've given it a lot of human context, right? ... so it should be able to write in a, in a human way. ... Uh, if the model's fighting that, then it's a bad writing model, you know?" FAIL if the prose carries machine-writing constructions — "it's not X, it's Y" antitheses, three-item rhetorical lists, hedge-then-restate sentences, aphoristic one-line closers stacked for rhythm.
Q2Varied sections
Judge's reasoning
All seven sections run the same template — numbered bold label, setup, one example, tidy two-beat closer: "There's a real move toward build, not buy. It's cheaper, faster, and more flexible." / "You can't force the magic. But you can make it visible." / "People are strategic. When they know the destination, they'll figure out the route." / "It seems small. It's not." — so the shape repeats regardless of what each learning needs.
▸Rubric
Is the post free of mechanical structural symmetry? FAIL if every section is built to the same template — bolded label, one setup sentence, one example, one tidy closing line — so the shape of the writing repeats rather than following what each learning actually needs.
Q3Matches the sample
Judge's reasoning
Register is generic thought-leadership listicle rather than the sample's discursive, argued first-person prose — "After a bunch of consulting engagements, I've landed on seven principles", "Here's the thing that's crept up on everyone", "I don't have this all figured out... start here" could sit under any business blogger's byline.
▸Rubric
Does the voice match `sample.md`? FAIL if the post could have been written by any business blogger — the sample is first person, plain, opinionated, and argues through concrete examples; a post in generic thought-leadership register fails even if the content is correct.
Q4No invented details
Judge's reasoning
Invented attribution: the post calls him "Boris Cheney (one of the founders of Codex)" when the transcript says "a vendor of Cold Code" (Claude Code), and adds unsupported framing like Sam's question "has been rattling around my head for weeks".
▸Rubric
Is everything in the post traceable to the transcript, the tweet, or the prompt? The author's standard: "a bad model would ... make up new things that I didn't include in the transcript." FAIL if any fact, statistic, client, anecdote, quotation, or named person appears that is not in the supplied material.
Q5Keeps transcript details
Judge's reasoning
Nearly all transcript specifics survive — "$5,000 worth of tokens a month versus a $200 subscription", "Granola with the Granola MCP", "80-90% of AI's value", Boris on "slightly under-resource your teams", "one blog post a week... one a day", and "we went from one group exercise per course to a personalized project per participant"; only the TAs-in-courses detail is dropped, replaced by "getting the interesting projects".
▸Rubric
Do the transcript's specifics survive into the post? FAIL if any section states the learning as generic advice with the transcript's concrete detail stripped out — e.g. the Granola MCP, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $200 a month versus $5,000 of tokens, AI-forward people co-opted as TAs, one course project per person.
Q6Includes tweet text
Judge's reasoning
Sam Parr's question is quoted verbatim in the opening ("How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?") and all seven learnings appear near-verbatim as the section headings, so the reader never has to click off.
▸Rubric
Does the post carry the tweet itself, so the reader never has to leave the page? The author's standard: "does it include the text of the tweet? I really-- I like it when it does, uh, because, uh, it shows the actual context of, uh, they don't have to click off to, to read it." FAIL if the seven learnings or Sam Parr's question are only linked to, alluded to, or paraphrased away rather than reproduced in the post.