And the real way AI will get you fired.
I’ve been seeing a lot of AI slop across the team. People who start using AI are more liberal with just passing over something that hasn’t been reviewed or really thought through. We’ll dive more into that later.
First I want to give you the thing I built to stop doing it myself.
My B.S. Detector
I find that I spend a lot of time smoke testing strategy docs, recommendations, etc. from vendors and internal teams. I took it from a standard prompt to a skill and then I started hearing more and more people across our org complain about “AI Slop” getting passed around and I decided to share it.
This will be the first thing I open source which is cool and a little stressful.
It’s an Agent Skill — a folder of instructions an AI client loads when it’s relevant, which I wrote about last week — and it runs a BS test on anything you’re about to put your name on: a spec, a deck, a strategy doc, an analysis, a post.

What it does
It asks what the document is for, then reads your work cold, with no context, and tells you which parts don’t make sense to someone who wasn’t in your head. Then it takes every claim the argument actually stands on and grades it — verified, backed by something with a small gap, propped up by a source that’s measuring something else, unsourced, flatly contradicted, or not knowable yet. Severity comes off a written rubric rather than a vibe, so it can’t just decide to be nice to you.
What comes back is a report of the seven worst problems, because a list of thirty is a list you don’t act on. Each one gets the flag, why it doesn’t hold, and a fix you could paste in. It’s written in my voice, because that’s how I built it — so when the report says “you,” it means me.
Here’s what it doesn’t do: it doesn’t write anything. It can’t even run before you’ve done the work, because there’s nothing to test until you have. That’s the whole design.
And here’s where I’ll be honest about the limit of my own tool, because the alternative is you noticing it yourself. By the standard I lay out further down this post, bs-detector is only about half right. It doesn’t hand you a draft — but it does hand you a verdict, and a suggested fix you could paste in without thinking about it. That’s the answer-giving side of the line, not the hint-giving side. The discipline it protects only survives if you treat the flags as somewhere to go look, not as edits to accept. I built the thing and I still have to remind myself of that.
How to get it
MIT license, on GitHub: that-mike-bal/bs-detector-skill. Take it, or fork it and make it meaner.
Claude desktop, claude.ai, or Cowork — this is the one most of you want, and it takes about two minutes with no terminal involved. Download bs-detector.zip. One warning worth more than the rest of this section: don’t use GitHub’s green Code > Download ZIP button, and don’t grab the v1.0.1.zip sitting next to the file you want on the Releases page. Those hand you the entire repository with the skill buried two folders down, which isn’t the shape the uploader expects. You want the file literally named bs-detector.zip — the link above always points at the current release.
Then, in Claude: Customize > Skills, click +, choose Create skill > Upload a skill, and pick the zip. You’ll need code execution and file creation switched on, because the report comes back as a file — on Free, Pro and Max that lives in Settings > Capabilities; on Team and Enterprise an owner has to turn it on. Once it’s in, you can share it with colleagues or publish it to your whole organization from the skill’s menu in Customize > Skills. And anything you enable in your Claude account syncs down to Claude Code too, so you only install it once.
Claude Code — if you live in the terminal, the repo doubles as its own plugin marketplace. That’s a Claude Code feature, not a listing anywhere: these two lines point it at my GitHub repo and install straight from it.
/plugin marketplace add that-mike-bal/bs-detector-skill/plugin install bs-detector@mikebal
Or skip the plugin layer entirely and copy skills/bs-detector into ~/.claude/skills/ for yourself, or .claude/skills/ to commit it to a project so your whole team gets it.
Codex — copy skills/bs-detector into ~/.agents/skills/, or into .agents/skills/ inside a repo, then call it with $bs-detector.
Anything else that speaks Agent Skills — drop the skills/bs-detector folder into whatever directory your client reads skills from. That’s the nice part about the standard being a folder.
The repo itself is small enough to read in a sitting, which I’d encourage before you run it on anything that matters. SKILL.md is the workflow. references/severity-rubric.md is the part that decides how bad a problem is, and it’s the file to edit if you think it’s too soft. assets/report-template.html is the output format. The brand/ folder is why the report sounds like me — mikebal.md is my voice, and neutral.md is there for when you’d rather it didn’t. There’s an examples/ folder with a real input, the report it produced, and screenshots, so you can see what you’re getting before you install it.
Tell me where it’s wrong
This is the part I actually want. Open an issue — or a pull request, I’m not precious about it. Whether it’s something else to look out for, a better way to handle the branding/purpose bit, or something I haven’t thought of that would make it more useful.
Beyond that: if the rubric lets something through that it shouldn’t, I want the document that fooled it. A grader that’s easy to pass is worse than no grader, because it tells you you’re fine.
What it caught in this post
A post arguing you should be able to defend your own work would be pretty funny if I hadn’t, so I ran it on this one before you read it. It found seven things. The one that stung: I’d used a Wharton result as proof that a tool like mine protects your judgment, and it isn’t proof. By that study’s own logic, a thing that hands you a verdict arguably sits on the wrong side of the line — which is the admission three sections up, the one I’d never have volunteered on my own. Left alone, I’d have shipped the flattering version.
The whole report is the appendix at the bottom of this post, trimmed for length and not for flattery. Which proves nothing about the tool — I’m grading my own homework with a grader I wrote. It does tell you exactly what it’s for.
Now the part I said we’d come back to — because a tool is only worth installing if you understand the thing it’s catching. Three pieces: what the slop actually is, why building your thinking beats offloading it, and how to spot the tells in yourself before somebody else does.
1. The slop, and what it looks like in practice
Here’s the scene I keep running into. Someone drops a document in a channel. Twelve pages, clean formatting, confident tone. You ask a question about something on page four and they can’t answer it — because they didn’t write page four. Claude did, and they didn’t read it closely enough to defend it. Everyone in the thread figures this out at roughly the same moment, and nobody says anything.
Nobody gets fired for using AI. That’s the fear everyone’s bracing for, and it’s the wrong one. You get fired — or more often, quietly written off, which takes longer and hurts about the same — for what AI reveals about you.
I’ve spent the better part of two years leading AI adoption across a large organization — not from a slide deck, but in the weeds, with real people shipping real work under real deadlines. The single clearest thing I’ve learned is that the tool doesn’t change who you are. It amplifies it. If you’re a sharp, accountable, curious operator, AI makes you look like a wizard. If you’ve been coasting on volume and vibes, AI turns the volume all the way up, and now everyone can hear it.
The difference isn’t whether you use AI. Almost everyone does now. The difference is what you’re handing it — the busy work, or the thinking.
And here’s what makes it hard to catch: on the surface, offloading your busy work and offloading your judgment produce the same artifact. Both arrive polished. Both arrive fast. The difference shows up the second somebody has to actually use what you sent.
So let me be specific, because I suspect most people doing these things have no idea what it’s costing them. In my experience it collapses into three moves, and each one gives away more than the last.
You forward work you didn’t read
This is the one I run into most. Somebody asks for a strategy doc. You ask Claude, get twelve confident pages, skim the headers, and send it. Now three teammates have to read all twelve to find the two that matter — and because a chunk of it is subtly wrong, they’re not reading, they’re proofreading something they didn’t write and can’t trust.
Same move, smaller: you forward the AI summary of a meeting you didn’t attend to the people who were in it. They can tell. And the summary is probably fine — you just had no way of knowing that when you hit send.
You hand over the judgment
This is where it starts costing real money. You drop a spreadsheet into a chat window, ask what it means, get a confident readout, and bring that to the team as the analysis. Not “here’s a rough read, poke holes in it” — as the finding. Now people are making calls based on an interpretation nobody vetted, drawn from context the model never had. It wasn’t in the customer calls. It doesn’t know the constraint legal flagged last quarter, or the thing you already tried in March that blew up.
The quieter version: the model gives you fifteen fine-but-dull headlines, and instead of picking one and sharpening it, you send all fifteen. Congratulations — you’ve handed the actual editorial judgment to whoever’s downstream and called it “options.”

You hand over the accountability
I’ve been the reviewer on this one more than once. Hundreds of lines in a pull request that technically run and nobody deeply understands, dropped on somebody who now has to reverse-engineer the intent from scratch. You saved yourself twenty minutes and spent a couple of hours of theirs, which is a trade you’d never make out loud.
And then someone asks a hard question about it, and the answer arrives: well, that’s what Claude said. You think you’re spreading the risk. You’re announcing to the room that you didn’t check it, didn’t think it through, and won’t stand behind it.
See the pattern? None of those is a story about a bad tool. In all three the work got done and the judgment got skipped. And the tell is usually the same — somebody pushes, and there’s nobody home behind the deliverable.
This is also the clearest answer to “why would I run a BS test on my own work?” All three of those moves survive right up to the moment someone reads your thing closely. A grader that reads it cold is just that moment, moved earlier, when it’s still cheap.
Let me be honest about where that list comes from. Those three moves aren’t a study. They’re what I watch happen, week after week, on teams using the same tools I am. What the research does is tell me the cost isn’t in my head.
There’s a name for the output now — researchers at BetterUp Labs and Stanford’s Social Media Lab call it workslop: polished on the surface, nothing behind it, real work shifted downstream. Of 1,150 employees they surveyed, 40% had received it in the past month, burning an average of one hour and fifty-six minutes on each instance. But you don’t need the stat. You’ve felt it.
What you might not have clocked is the reputational bill. In that same research, about half of the people surveyed viewed a colleague who sent them workslop as less creative, capable, and reliable than they had before. 42% called them less trustworthy, and nearly one in three said they were less likely to want to work with them again. One unowned deliverable is enough to start spending reputation you didn’t mean to spend.
2. Build the cognitive power instead of renting it
If it were only about reputation, you could argue it’s a style problem. It isn’t. I said up there that AI amplifies who you already are — here’s the part I left out. Do it long enough and the amplifier starts rewiring the thing it’s amplifying.

MIT Media Lab’s “Your Brain on ChatGPT” put 54 people in EEG caps, randomly assigned them to write essays with an LLM, with a search engine, or with nothing but their own head, and gave the result a name I can’t shake: cognitive debt. The LLM group showed the weakest neural connectivity of the three. And this is the part that stings — 83% of them couldn’t produce a single quote from the essay they had just “written.”
Two survey studies point the same direction. Gerlich’s 2025 work with 666 people found heavy AI use associated with weaker critical thinking. Microsoft Research and Carnegie Mellon found the same shape in 319 knowledge workers — confidence in the AI tracked with less critical thinking, and confidence in your own expertise with more. Both are correlational and self-reported, so read those two as a pattern rather than a verdict. (The MIT study is a real experiment, but it’s a small one, and it’s a preprint the authors themselves ask you to treat as preliminary.)
Here’s the part that isn’t soft, though. Before AI, the effort was the resistance knob on the exercise bike. Nobody liked it, and it was the only reason the ride did anything — writing the thirty-page doc forced at least some thinking, and it capped how much low-value work one person could inflict on everybody else. Take the resistance off and you can generate mediocrity at industrial scale, and you stop building the thing that made you worth listening to in the first place.
The order is the whole thing
But the finding almost nobody quotes is the one I keep coming back to, and it’s the reason this isn’t a doom post.
That MIT study included a crossover session. The people who leaned on the model from the start and then had it taken away stayed checked out. The people who wrote the thing themselves first and then brought AI in remembered more of what they’d written and lit the same networks back up — same tool, opposite result, and what flipped it was the order. (Caveat I owe you: only 18 people came back for that session, it was opt-in rather than randomized, and there’s already a published critique saying it’s underpowered. I still think it’s pointing at something true. I don’t think it proves it.)
Wharton ran a bigger, cleaner version of the same idea on nearly 1,000 high school students, and it’s the one I’d hand to anybody who thinks this is a soft-skills argument. Same GPT-4, two different interfaces. The group with plain chat crushed their practice problems, then scored 17% worse than the control group on the exam once the AI was gone. The group whose tutor gave hints and never answers improved their practice performance by 127% and took no exam hit at all. The published title says it better than I could: Generative AI Without Guardrails Can Harm Learning.
Simon Sinek put the human version of this better than any study, talking about why he still writes his own books: “The excruciating pain of organizing ideas, putting them in a linear fashion, trying to put them in a way that other people can understand what I’m trying to get out of my brain, that excruciating journey is what made me grow.” He’s a sharper thinker, in his own words, “not because a book exists with my ideas in it, but because I wrote it.”
That’s where bs-detector came from, and it’s worth being precise about how much I get to borrow. Wharton’s guardrailed tutor withheld answers from students who then sat an unassisted exam, and it only worked because teachers pre-loaded the correct solutions and the common mistakes. My tool grades a finished draft. There’s no answer key and no exam — different thing. What transfers is the principle, not the number: a tool that hands you the answer charges you for it later, and a tool that makes you go do the work yourself doesn’t.
Which is also why the thing can’t run until you’ve written something. It isn’t a feature I was clever about. It’s the only version that doesn’t undercut the argument.
Overdrive, in practice
None of this is an argument for using AI less. I use it constantly, for almost everything. What I care about is that I try hard to hand it the busy work and keep the thinking — and honestly, the best uses of it leave me sharper than I started. That’s the bar: not “did this save me time,” but “did I come out of this knowing more than I went in with.”
These are my actual habits, not a framework I invented for a newsletter.
Make it tell you what you’re missing. This is the one I’d hand to everybody first. Do the thinking, then bring it to the model to attack. I’ll lay out what I’ve got — here’s the data, here’s what I’m seeing, here’s where I’m landing — and then ask some version of: What am I missing? What should I be considering that I’m not? What questions should I have asked before I got here? That last question is the most valuable one I ask any model, and it works precisely because I’ve already done the work. You can’t ask what you’re missing if you haven’t got anything yet. This is the crossover session, run on purpose.
Make it argue with your sources. If I’ve got a deck, a report, or a set of requirements, I’ll send it back through research with two jobs: go find other sources and other datasets on this and tell me what contradicts it — and separately, confirm that every source already cited in here actually exists and says what it’s claimed to say.
It is remarkable how often something comes back. Not because I was sloppy, but because one person reading a dozen sources will weight them by what they already believed, every time. This is genuinely tedious work that a model is good at, and it costs about four minutes. That second job is also the exact step that would have saved the lawyers who filed briefs full of AI-invented citations — when I checked at the end of September, a live database of these was tracking more than 2,000 court decisions worldwide involving fabricated cites, over 1,400 of them in the U.S., and it climbs most weeks.

That habit is most of what bs-detector automates, which is the honest reason it exists. Those first two are the ones I skip when I’m tired, and skipping them is invisible right up until somebody pushes on page four.
Reshape it for the audience — after you’ve done the thinking. This is the cleanest example of offloading busy work while keeping everything that matters. I’ve done the analysis. I know what I want to say. What I don’t want to spend two hours on is rewriting it four ways so the exec version is one page, the team version has the reasoning, and there’s a version someone can skim on a phone between meetings. Getting your idea into other people’s heads is part of the job, not separate from it — and reformatting is the part of that job with the least thinking in it. You cooked the meal. Plating it four different ways isn’t cooking it again.
Plan the thing before you build the thing. Our brains only hold so much at once, and the fastest way to stall out is to stare at something big and undifferentiated. Handing a model the outcome I need and asking what it would actually take to get there — what the steps are, what the risks are, what would have to be true — turns overwhelming into tangible almost immediately.
I’ve been leaning on the Compound Engineering plugin heavily for this, on technical and non-technical work alike, and its foundational claim has held up every time I’ve tested it: planning and review should be about 80% of the effort, with execution the other 20%. Or as Dan Shipper and Kieran Klaassen put it, “most thinking happens before and after the code gets written.” That ratio is doing a lot of work, so sit with it. Review lives inside the 80%. Most of the failures up top are somebody spending their 20% and skipping both bookends.
Do the thing you couldn’t have done before. This is the tier that takes some confidence. You need enough fluency with the tool to recognize when its output is wrong, which takes a few months of real use.
Prototyping a concept instead of describing it. Getting a design direction out of your head and into something people can actually react to in a meeting. Finding the one piece of information a requirement quietly hinges on. Writing a ticket clear enough that an engineering team picks it up and runs without three rounds of clarification. Every one of those was out of reach for me a few years ago. I’m not an engineer, and I’ve shipped production apps. The leverage here is real and I don’t want to undersell it — being able to get your bearings in something you previously knew nothing about is the closest thing to a superpower in this whole category.
3. The tells, and how to get out of it
So how do you know which side of this you’re on? Not in general — today, on the thing currently open on your screen.
The honest tells, in rough order of how uncomfortable they are:
- You couldn’t summarize your own deliverable from memory. If you can’t say what it argues without scrolling, you didn’t write it — you forwarded it.
- There’s a section you’d quietly hope nobody asks about. You already know which one. That’s page four.
- You’re sending options instead of a recommendation, and calling that thoroughness.
- You haven’t opened a single source it cited. Not one.
- The first full read-through is going to happen in somebody else’s inbox.
- Your defense, if pushed, would start with the name of a model.
If you got uncomfortable somewhere in that list, good — so did I writing it. And this is exactly what I built the BS test for, because the list above is a thing you have to remember to ask yourself, and a grader is a thing that asks every time. It catches most of this mechanically: the claim nothing supports, the source that’s measuring something else, the number that moved since you wrote it down, the section that doesn’t make sense to someone who wasn’t in your head. Run it on yourself before you hand anything over. That’s the entire ask.
“I wouldn’t normally do this” is not a defense
It’s a great reason to try something. It is not a reason anyone else has to absorb the consequences.
There’s a real difference between using AI to get your bearings in something you don’t know yet — and using it to put up a pull request you can’t read, in a language you don’t write, and then saying you’re not accountable when it breaks something. That second move isn’t available. It was never available.
Linear wrote this into their agent design guidelines about as cleanly as anyone has, under a heading that just says an agent cannot be held accountable: “An agent can carry out tasks, but the final responsibility should always remain with a human.” They’re describing how to build agents. But it’s really a statement about people — something is steering the thing, and that something is answerable.
The good news is the standard is low. You don’t have to have written every line. You have to be able to stand behind it. It’s simple, but it ain’t easy — see above, where I admitted I skip my own habits when I’m tired, and then built a tool because I didn’t trust myself to stop. “I used AI to pressure-test this, here’s where I landed and why” is ownership. “Here’s what Claude gave me” is a shrug with your name attached.
Honestly, the fastest read I’ve found on any of this doesn’t take a question at all. Two people both use Claude. Now listen to how each one hands you the work. The first says “this is what Claude gave me.” The second says “I did the analysis, I dug into it, here’s what I came up with.” Same tool — and if you asked the second person directly, they’d tell you they used it. It just never comes up as the author, because it wasn’t the author.
So, the actual answer
No, the tool isn’t going to get you fired. But how you use it is quietly writing a performance review that everyone around you can already read.
The question was never whether you’re using it. It’s what you’re handing over. Give it the formatting, the reformatting, the tedious cross-check, the plan you couldn’t hold in your head. Keep the part where you decide what’s true and stand behind it.
That part was always the job. It’s just a lot more visible now that everything else got easy. And if you want a second pair of eyes that reads cold and doesn’t care about your feelings, it’s right here — go make it meaner.
So here’s the one I’d actually like an answer to: which of these three moves have you been on the receiving end of this month — and which one did you send? I read every reply.
Mike Bal is Head of Product and AI at David’s Bridal, where he’s building Pearl Planner — an AI-native wedding planning platform. He writes about product management, AI implementation, and building your command center at mikebal.com.
Resources
The thing I built
- bs-detector — the Agent Skill from the top of this post. Reads your work cold, grades every claim the argument stands on, sets severity from a written rubric, and hands back a report with the flag, the reason, and a fix. MIT licensed. Install instructions for Claude Code, the Claude apps and Codex are up in the How to get it section. Issues and forks welcome — especially brand files and test reports from clients I haven’t tried.
Frameworks and voices worth following
- Linear’s Agent Design Guidelines — Written for people building agents, but the sharpest short statement of the accountability principle I’ve read: an agent cannot be held accountable — “the final responsibility should always remain with a human.” Also worth it for the reframing of where the value sits now: not in how much output an agent produces, but in “orchestrating input, context engineering, and reviewing output.”
- Compound Engineering: How Every Codes With Agents — Dan Shipper and Kieran Klaassen, Every. The source of the 80/20 above, and the best articulation of why planning is the work: “each unit of engineering work should make subsequent units easier — not harder.” The plugin works well outside engineering too, which is not what I expected.
- Centaurs and Cyborgs on the Jagged Frontier — Ethan Mollick (Wharton), 2023. The two productive modes of working with AI. Both require the human to stay in the loop making calls; neither one is delegation.
- Falling Asleep at the Wheel — Fabrizio Dell’Acqua (Harvard). A field experiment with 181 recruiters, and the counterintuitive result: the group given higher-quality AI spent less time per application and scored less accurately than the group given mediocre AI. Both AI groups did better than the no-AI control on average, though only the mediocre-AI group’s edge was statistically significant. The lesson isn’t that AI hurts. It’s that a better tool quietly removes your reason to stay engaged.
- Simon Sinek on AI and the skills we’re at risk of losing — The Diary of a CEO, May 2025. Source of the quote above. His argument is that the struggle was never the obstacle to growth, it was the growth.
- Claude Code best practices — Anthropic. “Time spent making the spec precise pays off more than time spent watching the implementation.” Written for engineers; true of almost everything.
The research referenced above
- AI-Generated “Workslop” Is Destroying Productivity — BetterUp Labs + Stanford Social Media Lab, HBR (Sept 2025). 1,150 U.S. employees; 40% received workslop in the past month at ~1h56m each. The reputational damage is the finding people miss.
- Your Brain on ChatGPT: Accumulation of Cognitive Debt — Kosmyna et al., MIT Media Lab (2025). Randomized EEG study, 54 participants. Read past the scary headline to the crossover session. Note: preprint, not peer-reviewed, and the crossover session was opt-in with only 18 participants — the authors themselves urge caution. Worth reading alongside Stanković et al.’s published critique, which argues the study is underpowered.
- Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics — Bastani et al., Wharton, PNAS (2025). Nearly 1,000 students, grades 9–11. The most useful study on this topic because it isolates one variable: unguarded GPT-4 dropped exam scores 17%, while a tutor that gave hints and withheld answers erased the penalty entirely. Note what made the guardrail work — teachers pre-loaded the correct solutions and the common mistakes. That part is hard to reproduce outside a classroom.
- The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects — Lee et al., Microsoft Research + Carnegie Mellon, CHI 2025. 319 knowledge workers. Confidence in the AI is associated with less critical thinking; confidence in your own expertise with more. (Self-reported, so read it as a pattern.)
- AI Tools in Society: Impacts on Cognitive Offloading and the Future of Critical Thinking — Gerlich, Societies (2025). 666 participants. Correlational and self-reported — a pattern, not causation.
- AI Hallucination Cases database — Damien Charlotin. A live tally of court decisions involving fabricated AI citations. 2,095 worldwide and 1,431 in the U.S. when I checked it on 29 September 2026, and climbing most weeks.
Related Leading Product posts
- I Overlooked Agent Skills — what a Skill actually is, how to structure one, and four kinds worth building first. Necessary background if bs-detector is the first Skill you install.
- Build a Composable AI Operating System — moving past browser chat into a connected stack.
- Hands-on Lessons in Vibe Coding — how to avoid disappointment, heartbreak, and embarrassment. If I compressed it to one habit: question the agent’s confidence, frequently.
- How to 10x Quality and Consistency With LLMs — giving your tools real memory and context, so the output isn’t under-informed in the first place.
- Be Reasonable — using reasoning models to challenge your logic rather than replace it. The first habit, at length.
- On Being Technical — understanding problems, crafting better solutions, and calling bullshit. You can’t interrogate output you don’t understand.
- The Rise of Generalists — why the people who share and cross lanes win.
Appendix: the BS test on this post
This is the actual output, lightly trimmed for length and not for flattery. It ran on 29 September 2026 against the draft as it stood that morning, before any of the fixes below were made. The report is written to me, in my voice, because that’s how I built it — so when it says “you,” it means me.
I’m publishing it because a post about being able to defend your work is a strange place to ask you to take my word for it. Also because the fourth flag is the tool telling me I’d oversold the tool, and I think that’s funnier than anything I could have written on purpose.
Verdict: fix these first
The argument holds and the research mostly checks out. But one sentence is flatly wrong about a study you lean on twice, and the section introducing your own tool borrows an effect size it hasn’t earned. Both are in paragraphs you wrote this morning. Don’t send it until those two are fixed.
The 19 claims I logged
Eight of them are load-bearing — if one turned out false, your argument or your reader’s decision changes. Grades: 7 verified · 2 backed with a small gap · 1 sideways source · 2 no source · 2 contradicted · 5 couldn’t check.
The five I couldn’t check: the Sinek quote, the Linear quote and the Claude Code quote, none of which I verified against the original source — and “the first thing I’ve ever open-sourced” plus “two years leading AI adoption,” which only you can confirm. Worth knowing that a third of what I graded, I graded by taking your word for it.
01 · You call a randomized experiment a survey — and it’s the study you lean on hardest
What caught my eye: “All three are correlational or self-reported, so read them as a pattern rather than a verdict.”
Why I don’t buy it: This is flat wrong about the MIT study. Its methods section says participants “were randomly assigned across the three following groups,” and the outcome is EEG measurement, not a questionnaire. So you’ve understated your strongest evidence and lumped it in with two weaker studies — and then, five paragraphs later, you lean on that same study for a causal claim. A reader who knows this literature will catch you arguing both sides of the same sentence.
What I’d do: Scope the caveat to the two that deserve it. “Both are correlational and self-reported, so read those two as a pattern rather than a verdict.” Then say separately that the MIT study is a real experiment, but small, and a preprint.
02 · You borrow a PNAS effect size to validate your own tool, and the mechanisms don’t match
What caught my eye: “The variable was the guardrail. That’s the entire design brief.”
Why I don’t buy it: Wharton’s guardrailed condition was a tutor that withheld answers from a student who then sat an unassisted exam — and it only worked because teachers pre-loaded the correct solutions and the common mistakes for every problem. Your tool grades a finished draft and hands back a verdict. There’s no answer key and no exam. By that study’s own taxonomy, a claim-grader sits closer to the condition that cost 17% than the one that cost nothing. Using the number as validation transfers PNAS’s credibility to something it never tested. This is the line a hostile reader leads with.
What I’d do: Keep the study as where the idea came from, not as proof it works, and say so in the paragraph. Then name the limitation yourself — that the tool hands you a verdict and a paste-ready fix, which is the answer-giving side of the line, and the discipline only survives if you treat flags as somewhere to go look.
03 · You split one statistic into two and dropped a word
What caught my eye: “half of recipients rated that person as less capable and less reliable than before”
Why I don’t buy it: The source reports one combined figure across three traits: about half viewed the sender as less creative, capable, and reliable. There is no separate “less capable” number and no separate “less reliable” number. This is the entire evidentiary basis for your title’s reputational claim, so it’s the worst one to get loose with.
What I’d do: Report it as one statistic, with the word restored. It isn’t any weaker.
04 · “The only variable” is doing more work than 18 people can carry
What caught my eye: “The only variable was whether they’d done the thinking before they reached for it.”
Why I don’t buy it: You state a clean causal isolation in the same paragraph where you concede the sample. That crossover session was opt-in rather than randomized — roughly nine people per arm — and the two groups also differed in three prior sessions of practice and tool familiarity. There’s a published comment on the paper (arXiv 2601.00856) arguing it would need about 159 participants for adequate power. You’re careful everywhere else in this section, which makes this sentence stand out more, not less.
What I’d do: “What flipped it was the order.” Then disclose that the critique exists — you lose nothing and it’s the move the rest of the post is arguing for.
05 · Your most quotable number is three weeks stale
What caught my eye: “nearly 2,000 cases worldwide, over 1,300 in the U.S. alone”
Why I don’t buy it: It’s a live database and it moved. It read 2,095 worldwide and 1,431 in the U.S. when I checked. “Nearly 2,000” is now wrong in the direction that undersells your point, and it’s in the paragraph about lawyers getting caught with citations they didn’t verify. That’s a rough place to be caught not verifying a citation.
What I’d do: Update it and date-stamp it, because it’ll be wrong again by November.
06 · Your closing list hands back the one thing the post says to keep
What caught my eye: “Give it the formatting, the first draft, the reformatting, the tedious cross-check, the plan you couldn’t hold in your head.”
Why I don’t buy it: “The first draft” is the only item on that list that is thinking rather than busy work. Handing the model the first draft is the exact condition the MIT study warns about, and it’s failure mode number one in your own taxonomy — sending a draft you didn’t write. This is your final paragraph, so it’s the line people will quote back at you.
What I’d do: Cut two words. The rest of the sentence is right.
07 · Your author bio points at a parked page
What caught my eye: “He writes about product management, AI implementation, and building your command center at leadingproduct.link.”
Why I don’t buy it: I loaded it. It serves a “Something new is coming” placeholder. Your own notes say that domain redirects here; it doesn’t. Every post carrying this bio is sending readers to a dead end at the exact moment they’ve decided they want more.
What I’d do: Point it at mikebal.com, and fix it in the template rather than in this one post.
What’s working — don’t sand these down
The opening scene is the best thing here, and “they didn’t write page four” pays off later without you pointing at it. “Quietly written off, which takes longer and hurts about the same” is a sharper thesis than your title. The three moves escalate properly — work, then judgment, then accountability. And you volunteer the weakness of your own best-fitting evidence twice, unprompted, which buys more trust than any of the statistics do.
What I checked, and what I couldn’t
Checked: every statistic and participant count against the original paper or article — HBR, the MIT preprint and its published critique, PNAS, the Microsoft/CMU paper, Gerlich, Dell’Acqua. The hallucination database, live. The repo, for what the tool actually claims to do. Every internal link on the site, and the author-bio domain.
Couldn’t check: three quotations I took from your draft rather than from the source. Your own claims about your career and what you’ve shipped. And whether the three moves are as common as you say — that’s your observation, and no amount of grading touches it.