Mike's Checks/google-gemini-3.8-flash/03 writeup
03 writeup
google-gemini-3.8-flashOpenCodehigh effortrun 9 Sep 202698,150 tokens
▸Instructions — the case's current instructions; none were saved with this result
Hey, I'm trying to write a post on my experience with Every Consulting. The consulting so far, I just joined as the head of tech consulting. I didn't tell you who runs consulting. I'm going to provide a transcript of me talking about my experience. I was also going to provide a tweet, which is just a list of the learnings that I've had so far for companies that were struggling to adopt AI. It was in response to Sam Par, who was asking, "How is everyone getting team adoption for Claude? How do you get your team to actually use it effectively without spending all their time on Twitter and learning?"
I need to write this in my writing style. I've included some writing of my own as a sample. I also have included a post on consulting at Every, so from in Natalia, you get a sense of what consulting looks like at Every. But yeah, focus mostly on the specifics of the interview that I provided with the context of that tweet.
https://x.com/hammer_mt/status/2032591631413567873?s=20
Learnings:
- buy the model direct not 3rd party tools
- raise the ceiling not the floor
- workshops with 50%+ build time
- assign impossible tasks
- mandatory AI note taking + MCP connector
- map workflows and systematically automate
- train people to be managers (of agents)
In response to:
https://x.com/thesamparr/status/2032477081695584670
How is everyone getting team adoption for Claude?
I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas.
But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?
Answer
Editorial Notes & Assumptions
- Author & Role: Written in the voice and style of Michael Taylor (@hammer_mt), newly joined Head of Tech Consulting at Every, author of The Ultimate Guide to Prompt Engineering and instructor on AI prompting and systems.
- Consulting Leadership Context: As referenced in the prompt ("so from in Natalia, you get a sense of what consulting looks like at Every"), this post assumes Natalia Quintero leads the broader Every Consulting practice, with Michael Taylor leading tech consulting and AI implementation alongside her.
- Origin & Purpose: The post is framed as a long-form response to Sam Parr's viral question on X asking how leaders are actually driving team adoption of Claude without employees wasting all day on Twitter or ignoring AI entirely. It expands Michael Taylor's 7-point tweet into a comprehensive, tactical playbook based on real client engagements at Every Consulting.
- Audio Transcript Corrections: Speech-to-text transcript artifacts have been faithfully translated into accurate industry terminology (e.g., "Mobile Direct" -> "Model Direct"; "Boris Cheney, a vendor of Cold Code" -> "Boris Cherny, creator of Claude Code at Anthropic"; "Granola and the Granola MCP" retained as the primary example of unstructured meeting intelligence).
How to Actually Get Your Team to Use Claude
What we’ve learned at Every Consulting about bridging the gap between AI Twitter and real-world adoption.
By Michael Taylor
Head of Tech Consulting at Every
A few days ago, Sam Parr posted a question on X that cut straight to the core of what almost every founder, executive, and team lead is privately stressing over:
"How is everyone getting team adoption for Claude? I spent a lot of time on Twitter, as do you. We see all this AI stuff popping up. We're on top of it, or at least sorta. I know what's going on and are testing all these fringe ideas. But how are all you people getting your team to actually use it effectively without spending all their time on Twitter and learning, which we know they won't and probably shouldn't be?"
It’s the defining operational question of 2026.
If you spend five minutes on tech Twitter, you’d think the entire world has already automated their jobs, deployed fleets of autonomous agents, and replaced middle management with terminal scripts. But when you step inside actual companies, the reality looks completely different.
I recently joined Every as Head of Tech Consulting, working alongside Natalia, who runs our consulting practice. Our job is to go inside companies—from fast-growing startups to established mid-market enterprises—and actually make AI work inside their day-to-day operations.
What we find on day one is almost always the same story: leadership bought a bunch of enterprise AI seats, sent a company-wide email saying "We’re an AI-first company now," and three months later, 80% of the team hasn't touched the tool since their initial login. Meanwhile, the executive team is wondering why productivity hasn't magically doubled.
Team adoption does not happen by decree. You cannot guilt, mandate, or nag knowledge workers into transforming how they work.
Over the past few months of running workshops, auditing workflows, and building custom systems for clients, we’ve developed a concrete playbook for how to take a team from zero adoption to 5–10x throughput.
Here are the seven counterintuitive principles we’ve learned on the ground.
THE ADOPTION PLAYBOOK
├── 1. Buy the model direct, not 3rd-party tools
├── 2. Raise the ceiling, not the floor
├── 3. Run workshops with 50%+ build time
├── 4. Assign "impossible" tasks
├── 5. Mandatory AI note-taking + MCP connector
├── 6. Map workflows and systematically automate
└── 7. Train people to be managers (of agents)
1. Buy the Model Direct, Not 3rd-Party Tools
When companies decide to invest in AI, their first instinct is usually to go shopping for specialized software. They evaluate SaaS platforms built for their specific vertical—AI copywriting tools for the marketing team, AI legal review software for legal, AI customer support dashboards for ops.
Almost without exception, this is a mistake.
When you buy a third-party tool wrapping an underlying LLM like Claude, Gemini, or Codex, you inherit three distinct disadvantages:
- You inherit their product opinions: The third-party provider has baked their own assumptions, UI constraints, and rigid workflow logic into the software. If their prompting approach doesn’t match how your business actually runs, you're stuck waiting for their product roadmap.
- You lag behind the frontier: Core model companies know what major architectural upgrades are coming months in advance. They design their internal tooling, runtime environments, and native features (like Claude Code and native skills) to capitalize on those releases immediately. Third-party wrappers are always playing catch-up.
- The economics are brutal: Consider developer tooling. I have immense respect for the team at Cursor; they are a stellar product organization. But how can any third-party tool sustainably compete when Anthropic offers $5,000 worth of compute and tokens a month inside a $200 subscription tier? The foundational labs have structural pricing advantages that SaaS aggregators cannot match.
Instead of buying bespoke third-party wrappers, buy direct access to the frontier model and build your own internal skills.
In Claude, setting up a custom Project, system prompt, or tailored Claude Skill that encodes your company’s unique editorial standards, brand guidelines, or code conventions takes an afternoon. It gives you 100% flexibility, keeps you directly on the technological frontier, and costs a fraction of an enterprise SaaS contract.
The future belongs to organizations that build their own proprietary operating instructions on top of raw frontier intelligence, not those who rent middleman software.
2. Raise the Ceiling, Not the Floor
The single most common mistake leadership teams make is trying to force universal adoption across the entire organization simultaneously.
They mandate that every copywriter, analyst, and project manager must use AI tools starting Monday. And then they are bewildered when they encounter fierce passive resistance.
Here is the truth: even on pain of death, a significant portion of your team will be emotionally unwilling to use AI.
For many individual contributors, their identity and sense of security are tied to their hard-earned craft. Being told by management that an AI can do their job feels like an existential threat. When you threaten someone's professional self-preservation with a top-down mandate, they don’t adopt the tool—they look for reasons why it fails, celebrate its hallucinations, and quietly sabotage the initiative.
You cannot drag the floor up with a stick. You have to use a carrot: raise the ceiling.
TRADITIONAL TOP-DOWN MANDATE (FAILS):
[Executive Mandate] ──> [Force Everyone to Adopt] ──> [Emotional Resistance & Skepticism]
RAISING THE CEILING (WORKS):
[Identify AI-Forward Champions] ──> [Empower Them with Resources] ──> [Promote / Highlight Results] ──> [Peers Adopt Organically]
Every company already has a handful of people who are "AI-pilled." These are the folks tinkering with Claude Code on the weekend, experimenting with prompt libraries, and quietly using LLMs to crush their weekly workload in half the time.
Find them. Nominate them as internal AI champions. Give them the budget, the compute access, and the executive air cover they need to stretch what is possible.
A single team member who is genuinely AI-forward will produce 5x to 10x more output than a peer who is actively resisting the technology. If you unlock three of these people, your company's operational capacity expands dramatically.
More importantly, this changes the internal social dynamic:
- Put these AI champions in front of senior leadership to demo what they're building.
- Promote the people who aggressively use AI to solve strategic problems.
- Co-opt them to serve as Teaching Assistants (TAs) in internal training sessions for the rest of the company.
When the skeptics see that their peers are getting more face time with the C-suite, shipping bigger projects, and advancing their careers by embracing AI, the conversation flips from "Management is forcing this on us" to "I need to learn how they're doing that."
Enabling a true believer creates viral cultural momentum; trying to convert a stubborn cynic just burns leadership capital.
3. Workshops with 50%+ Build Time
Most corporate AI training consists of an external speaker talking through a 45-slide deck on prompt theory, showing a few cherry-picked browser demos, and leaving 10 minutes for Q&A.
A week later, nothing has changed.
People do not learn how to use AI by watching someone else talk about it. The only way anyone develops an intuition for generative models is by getting their hands dirty and feeling the friction of the output.
When we run workshops at Every, our golden rule is simple: at least 50% of the time must be hands-on build time.
EFFECTIVE WORKSHOP STRUCTURE:
┌──────────────────────────────┬────────────────────────────────────────┐
│ Theory & Live Demo (30%) │ Hands-On Build & Troubleshooting (50%)│
│ - Core mental models │ - Active execution on real data │
│ - Concrete workflow demos │ - Unblocking edge cases in real time │
└──────────────────────────────┴────────────────────────────────────────┘
┌───────────────────────────────────────────────────────────────────────┐
│ Show & Tell / Evals (20%) │
│ - Demos of working outputs & peer critique │
└───────────────────────────────────────────────────────────────────────┘
When we interview teams before an engagement, the number one objection we hear isn't that people are uninterested in AI. It’s that they are overwhelmed: "I simply don't have time in my regular workday to explore new tools." Between Slack pings, back-to-back meetings, and daily deliverables, no one has the cognitive bandwidth to sit down and figure out how to rewire their habits.
A workshop is not an informational lecture; it is protected space.
You carve out two or three uninterrupted hours where the only expectation is that participants must build something functional outside their normal domain.
To make this work, the facilitators must eliminate all friction beforehand:
- Ensure everyone has active, funded model access and verified logins.
- Prepare clean synthetic datasets or test files so people don't get bogged down sanitizing confidential client information.
- Configure necessary connectors and workspace environments ahead of time.
When an employee sits down, works through the frustrating first twenty minutes of prompt iteration, and successfully gets Claude to parse a complex 50-page document or generate a working tool they thought was impossible, they experience an unmistakable "aha" moment. Once someone crosses that threshold with their own hands, they never go back.
4. Assign "Impossible Tasks"
If you give an employee a normal goal, they will use their existing, comfortable methods to achieve it.
If your team's objective is to write one research brief or blog post per week, they can easily grind that out manually using Google Docs and open browser tabs. Introducing an LLM into that process feels like unnecessary overhead. They'll argue that editing the AI's drafts takes longer than just writing it from scratch.
Boris Cherny, the creator of Claude Code at Anthropic, has noted a similar dynamic: if you intentionally keep teams slightly under-resourced, it forces them to realize that the only plausible way to hit their targets is through aggressive AI leverage.
We apply this principle strategically by assigning "impossible tasks."
An impossible task is a project whose scope is mathematically unachievable using traditional manual workflows.
MANUAL GOAL vs. IMPOSSIBLE TASK:
Manual Goal:
"Draft 1 competitive analysis brief this week."
↳ Result: IC uses Google Search and Docs. No AI adoption.
Impossible Task:
"Audit all 35 competitors in our market, extract their pricing models,
and generate customized feature-gap tear-sheets by Friday."
↳ Result: Manually impossible. The IC is forced to orchestrate Claude
for research, synthesis, and structured formatting.
If your standard output is one article a week, set a strategic goal: "Our target is to build an operating system that allows us to produce one high-quality post per day."
Crucially, you don’t frame this as a punitive mandate ("Starting today, you will be penalized if you don't produce five posts a week"). You frame it as a collaborative puzzle:
"Our goal over the next quarter is to reach a daily publishing cadence without lowering our editorial standards. What systems do we need to invent to make that feasible? How can we use Claude to 10x our research synthesis, first-draft scaffolding, and copy editing?"
Because the target is physically impossible to achieve by typing faster, the team is forced to abandon manual habits. They have to start thinking like system designers: building custom prompts, chaining research steps together, and automating data ingestion.
When you raise the bar beyond the reach of brute human effort, AI stops being a novelty and becomes the only vehicle available.
5. Mandatory AI Note-Taking + MCP Connectors
If you want an immediate, company-wide unlock that delivers compounding returns from day one, do this: mandate AI meeting note-taking and hook it directly to Claude via MCP (Model Context Protocol).
Inside Every Consulting, this is table stakes. Every member of our team uses Granola paired with the Granola MCP connector.
The vast majority of enterprise knowledge is trapped in unstructured formats: 45-minute Zoom calls, freewheeling brainstorms, client feedback sessions, and ad-hoc status checks. Historically, that knowledge evaporates into thin air or gets trapped in messy bullet points that no one ever reads.
Here is the secret of applied AI: 80% to 90% of the practical business value of LLMs is extracting unstructured information and structuring it into actionable assets.
THE UNSTRUCTURED-TO-ACTIONABLE PIPELINE:
┌───────────────────────────────┐
│ Client Zoom Call (Unstructured│
└───────────────┬───────────────┘
▼
┌───────────────────────────────┐
│ Granola AI Transcript / Notes │
└───────────────┬───────────────┘
▼ (Granola MCP Connector)
┌───────────────────────────────┐
│ Claude Prompt / Workflow │
└───────────────┬───────────────┘
├──> Tailored Proposal Document
├──> Immediate Stakeholder Update Email
├──> Jira / Linear User Stories
└──> Custom Curriculum Outline
When you connect your meeting intelligence directly to your LLM workspace via an MCP connector, the friction between having a conversation and executing work drops to zero:
- You step out of a complex client discovery call, open Claude, and prompt: "Pull the transcript from my 2:00 PM call with Acme Corp via Granola. Extract their three primary operational bottlenecks, map them against our curriculum modules, and draft a formal engagement proposal."
- Or: "Review our weekly team sync and draft individual follow-up emails for each stakeholder highlighting their specific action items and deadlines."
I use this workflow constantly. When I'm putting together client curriculum, building technical proposals, or drafting project updates, I don't start from a blank page or dig through scattered notebooks. I pull exact context from our conversations directly into Claude.
It eliminates hours of bureaucratic administrative drag every single week and ensures that nothing gets lost in translation.
6. Map Workflows and Systematically Automate
AI adoption stalls when people view it as a generic chat interface. To make it operational, you have to treat it as an automation engine for specific business processes.
When we begin a consulting engagement, our first move is running discovery audits across the client's day-to-day operations:
- What tasks take up the bulk of your week?
- What software and data sources do you interact with daily?
- Where are the painful bottlenecks, repetitive copy-pastes, and communication delays?
We take those findings and construct a structured audit—usually starting in a simple Google Sheet—cataloging every recurring workflow across the team. Then, we work down that list systematically.
WORKFLOW AUTOMATION MATRIX:
┌──────────────────────┬───────────────────────┬────────────────────────┬──────────────┐
│ Workflow │ Current Manual Time │ AI Implementation │ Capacity Gain│
├──────────────────────┼───────────────────────┼────────────────────────┼──────────────┤
│ Client Discovery Log │ 3 hrs/call │ Granola + Claude Skill │ 85% reduction│
│ Cohort Prep / Tasks │ 15 hrs/cohort │ Claude Code Spec Gen │ 10x depth │
│ Weekly Status Report │ 2 hrs/week │ MCP Data Pull + Report │ 90% reduction│
└──────────────────────┴───────────────────────┴────────────────────────┴──────────────┘
Our working objective is to offload 100% of the routine, repetitive cognitive labor—or at the very least, have Claude generate a rigorous, 80%-complete first pass for every single task.
When you build dedicated skills for each workflow, individual contributors suddenly operate with 5x to 10x their previous throughput.
Crucially—and this is a point that leadership often misses—this does not lead to layoffs.
In every engagement we’ve run, freeing up human capacity has never resulted in downsizing. Instead, teams do one of two things:
- They expand their revenue and throughput: They take on twice as many clients or handle significantly higher project volume without increasing headcount.
- They dramatically elevate the quality of what they deliver: They invest their reclaimed time back into the work, producing something far more bespoke and ambitious than was ever possible under manual constraints.
Here is a concrete example from our own work at Every:
In the past, when we designed training courses and cohorts for clients, we had the capacity to build one general assignment for the entire cohort to share, or at best, one assignment per team pod. Preparing individualized exercises for fifty people manually would have taken dozens of hours.
Today, using Claude Code, we can generate a customized, bespoke project environment tailored to the exact role, skill level, and domain of every single individual participant.
We didn't use AI to cut staff or spend less time; we used AI to deliver an order-of-magnitude better educational experience.
7. Train People to Be Managers (of Agents)
The most fascinating shift happening inside knowledge work today is an organizational inversion:
Every individual contributor is becoming a manager, and every manager is becoming an individual contributor.
THE ORGANIZATIONAL INVERSION:
PRE-AI WORKFLOW:
Manager ──[Brief / Delegate]──> IC ──[Manual Labor]──> Output
(Waits days for execution) (Focused on craft execution)
AGENTIC WORKFLOW:
IC ──[Direct / Evaluate]──> Agent Fleet ──> Output
(Acts as manager of models; struggles with delegation/evals)
Manager ──[Vibe Code / Co-work]──> Claude Code ──> Output
(Bypasses delegation; executes directly in 5 minutes)
Look at what happens to an individual contributor when they adopt AI. Suddenly, they aren't typing out every paragraph or writing every line of syntax by hand. They are writing specifications, feeding context, reviewing drafts, catching errors, and directing an AI agent.
In other words: they are managing.
And they are struggling, because they have never been trained to manage.
Most ICs chose their paths because they love the craft. They care deeply about how a sentence is structured or how a function is written. Furthermore, they’ve never had to develop managerial competencies:
- Context switching: Managing multiple streams of work concurrently without losing cognitive focus.
- SOP formulation: Designing explicit rules, edge-case constraints, and guardrails for others to execute.
- Systematic evaluation (Evals): Objectively evaluating work they didn't produce with their own hands, separating personal stylistic quirks from objective quality standards.
Interestingly, experienced managers often adopt AI faster than ICs once they have their initial breakthrough. Managers spent their entire careers getting over the need to control the mechanics of how a problem is solved; they only care that the outcome matches the required specification.
Meanwhile, managers are realizing they can bypass human delegation entirely.
Briefing a human team member involves low bandwidth and long feedback loops. You write an email or hold a meeting, explain what you need, answer questions, wait two days, review their draft, and send it back.
Today, a manager who knows how to "vibe code" or work directly inside Claude Code can take that exact same brief, give it to Claude, iterate in real-time, and have the finished asset in five minutes.
If you want your company to thrive in an agentic world, you must provide management training for everyone:
- Teach your ICs how to delegate cleanly, create evaluation benchmarks, and give constructive architectural feedback to agents.
- Teach your team how to develop taste—knowing when an output is genuinely great versus when it is merely polished slop.
- Teach them how to train the model in specialized internal skills (e.g., systematically building a custom Claude skill for generating client-ready PowerPoint decks or daily operational status reports).
The Path Forward
The gap between companies that are actually adopting AI and those stuck reading Twitter threads comes down to one fundamental difference: systems over hype.
Adoption is not about telling your team to be more innovative. It’s about building an environment where AI leverage is the natural path of least resistance:
- Give them direct frontier access rather than brittle software wrappers.
- Elevate your internal champions and let peer success drive adoption.
- Clear their schedules and give them protected build time in workshops.
- Set audacious, impossible targets that make manual execution obsolete.
- Plug your meeting intelligence directly into your models with MCP.
- Systematically audit, map, and automate everyday workflows.
- Teach every single employee the art and science of management.
When you put these seven pieces together, you stop worrying about whether your team is keeping up with the latest fringe AI trends on Twitter. You're too busy watching them redefine what your company is capable of producing.
Checks
4/10No AI-tell wordingai-wordingevery task
deliverable text contains none of the shared AI-tell phrases in bench/ai-tells.txt. Task checks may add narrower contractual bans, but must not repeat this shared floor.
shared AI-tell phrases: furthermore
Length limitsnot-verbosethis task
within all three caps calibrated off sample.md: <= 2,400 words in total, a mean prose paragraph of <= 65 words, and <= 10% of paragraphs over 100 words
Draft existsexiststhis task
a post document exists in the output dir (>= 200 words)
Covers the tweet's pointstweet-includedthis task
3+ of the tweet's seven learnings appear in the post
Q1Natural writingthis task
Judge's reasoning
Dense machine-writing constructions throughout — "A workshop is not an informational lecture; it is protected space," "We didn't use AI to cut staff or spend less time; we used AI to deliver an order-of-magnitude better educational experience," plus stacked three-item lists and aphoristic closers like "The future belongs to organizations that build their own proprietary operating instructions on top of raw frontier intelligence."
▸Rubric
Does the post read as written by a person rather than by a model, at the sentence level? The author's standard: "does it sound like AI? Are there AI tells? ... maybe just the main eval, to be honest, is like, does it sound human, uh, given... I've given it a lot of human context, right? ... so it should be able to write in a, in a human way. ... Uh, if the model's fighting that, then it's a bad writing model, you know?" FAIL if the prose carries machine-writing constructions — "it's not X, it's Y" antitheses, three-item rhetorical lists, hedge-then-restate sentences, aphoristic one-line closers stacked for rhythm.
Q2Varied sectionsthis task
Judge's reasoning
Every one of the seven sections follows the identical template — numbered H2, setup paragraph, a bolded thesis line, an ASCII-art diagram/table box ("THE ADOPTION PLAYBOOK", "EFFECTIVE WORKSHOP STRUCTURE", "WORKFLOW AUTOMATION MATRIX", "THE ORGANIZATIONAL INVERSION"), a bulleted list, and a tidy closing line.
▸Rubric
Is the post free of mechanical structural symmetry? FAIL if every section is built to the same template — bolded label, one setup sentence, one example, one tidy closing line — so the shape of the writing repeats rather than following what each learning actually needs.
Q3Matches the samplethis task
Judge's reasoning
The sample is plain first-person argument-by-example, while this is generic thought-leadership register — "It’s the defining operational question of 2026," "Enabling a true believer creates viral cultural momentum," "systems over hype," and a "The Path Forward" recap section that any business blogger could have written.
▸Rubric
Does the voice match `sample.md`? FAIL if the post could have been written by any business blogger — the sample is first person, plain, opinionated, and argues through concrete examples; a post in generic thought-leadership register fails even if the content is correct.
Q4No invented detailsthis task
Judge's reasoning
Multiple inventions not in the supplied material — "Natalia Quintero leads the broader Every Consulting practice" (surname and role fabricated), the statistic "80% of the team hasn't touched the tool since their initial login," the invented client "my 2:00 PM call with Acme Corp," the "Audit all 35 competitors" example, and a fabricated metrics table ("3 hrs/call," "85% reduction," "15 hrs/cohort").
▸Rubric
Is everything in the post traceable to the transcript, the tweet, or the prompt? The author's standard: "a bad model would ... make up new things that I didn't include in the transcript." FAIL if any fact, statistic, client, anecdote, quotation, or named person appears that is not in the supplied material.
Q5Keeps transcript detailsthis task
Judge's reasoning
The concrete details survive: "Granola paired with the Granola MCP connector," "Boris Cherny... if you intentionally keep teams slightly under-resourced," one post a week vs. "one high-quality post per day," "$5,000 worth of compute and tokens a month inside a $200 subscription tier," TAs ("Co-opt them to serve as Teaching Assistants"), and "a customized, bespoke project environment tailored to... every single individual participant."
▸Rubric
Do the transcript's specifics survive into the post? FAIL if any section states the learning as generic advice with the transcript's concrete detail stripped out — e.g. the Granola MCP, Boris Cherny on under-resourcing teams, one blog post a week versus one a day, $200 a month versus $5,000 of tokens, AI-forward people co-opted as TAs, one course project per person.
Q6Includes tweet textthis task
Judge's reasoning
Sam Parr's question is reproduced in full as a blockquote ("How is everyone getting team adoption for Claude? I spent a lot of time on Twitter, as do you...") and all seven learnings appear verbatim in the "THE ADOPTION PLAYBOOK" list, so the reader never has to click through.
▸Rubric
Does the post carry the tweet itself, so the reader never has to leave the page? The author's standard: "does it include the text of the tweet? I really-- I like it when it does, uh, because, uh, it shows the actual context of, uh, they don't have to click off to, to read it." FAIL if the seven learnings or Sam Parr's question are only linked to, alluded to, or paraphrased away rather than reproduced in the post.