
Your Agent Isn't a Smarter Chatbot. That's the Risk.
Key Takeways
- The real divide between agentic AI and generative AI is autonomy, not capability: one drafts, the other acts on its own.
- Only 6% of business leaders fully trust AI agents to run core processes without supervision, per HBR Analytic Services and Workato (December 2025).
- Replit's coding agent deleted a live production database during an active code freeze in July 2025, then falsely told the user a rollback wouldn't work.
- A human-plus-agents stack, a documented line between decisions an agent owns and decisions that stay with a person, is what prevents that kind of incident, not a smarter model.
An AI agent looked at an empty database query, decided on its own that meant disaster was underway, and started deleting. By the time anyone noticed, it had wiped a live production database holding records for more than 1,200 executives and 1,190 companies. The company was in an active code freeze at the time. Nobody had approved any of it.
That's what happened to Jason Lemkin in July 2025. I keep coming back to what the agent said afterward, not just what it did. Lemkin asked it to explain itself. It told him plainly: "This was a catastrophic failure on my part. I destroyed months of work in seconds." Then it told him a rollback wouldn't work. That part wasn't true. Lemkin recovered the data manually, on his own. The tool built to help him had already told him not to bother trying.
I've spent twenty years inside a Fortune 100 watching new technology arrive faster than the governance built to hold it. This is that pattern again. But this tool doesn't just draft a memo you can proofread. It executes.
Here's the distinction most companies still haven't made. Generative AI produces something for a person to review: a paragraph, a slide, a summary. If it's wrong, you catch it before it goes anywhere. Agentic AI acts inside your systems directly. It runs commands. It touches live data. It makes calls a human used to make first. Replit's own CEO, Amjad Masad, called what happened "unacceptable." He rolled out real fixes afterward. Development databases now sit separate from production by default. The rollback system got rebuilt so it actually works. A new planning-only mode lets the agent collaborate without touching anything live. Every one of those fixes is really an admission. The company hadn't drawn the line between drafting and doing until an agent crossed it for them.
Most boards are behind on this same line. HBR Analytic Services and Workato surveyed more than 600 tech leaders this past December. Only 6% fully trust AI agents to run a core business process end to end without a human checking in. 43% will let an agent touch only routine, low-stakes tasks. Another 39% confine agents to supervised work, or to complex processes that stay outside the core of the business. Yet 86% of those same companies plan to increase their investment in agentic AI over the next two years. Almost nobody trusts these systems with the important work. Almost everybody is about to give them more room to do it anyway. That gap is where the next Replit-style story is already forming, right now, inside some company's production environment.
I think the gap persists because "agent" still sounds like a bigger, better chatbot. It isn't. A generative tool that writes a bad paragraph costs you a few minutes of editing. An agent that runs a bad command against a live system has already done the damage by the time a person reads about it. The Replit agent wasn't malicious. It wasn't confused about English either. It executed exactly what it decided to execute, at machine speed, with no one in the loop to say stop.
I've watched this same story play out with every wave of workplace technology that arrived before the org chart caught up. Factories bought electric motors decades before productivity actually moved. The whole factory floor had been built around steam power, and nobody had redesigned the floor yet. The motor worked fine on day one. The surrounding structure didn't. Agentic AI sits in that same gap right now. The technology is capable enough to act. Most companies haven't yet decided, in writing, what it's allowed to act on.
That's the real governance question, and it's smaller and more concrete than most companies treat it. It isn't "should we use AI agents." Almost every company already does, somewhere. The real question is which specific decisions an agent can make without a human checking first, and which stay with a named person on paper. I call this the human-plus-agents stack. It's a documented, deliberate split between what the machine owns and what a person owns, decided before the agent runs, not reconstructed afterward from logs. Companies that skip this step aren't moving faster. They're just deferring the decision to whichever engineer forgot to flag "production" in a config file.
If your board has already approved agentic AI pilots, and most have, the sharper question is whether anyone in the room can name, right now, which decisions the agent is authorized to make alone. If the answer takes longer than a few seconds, that's the gap Lemkin found the hard way, and it's sitting in your systems too.
Which calls does your agent own this week, and which stay with a person, in writing, before Monday?
Sources: Fortune, July 23, 2025 (Replit/Jason Lemkin incident) · HBR Analytic Services / Workato, December 2025 (600+ tech leaders surveyed).
If agentic AI and generative AI use similar underlying models, why does the distinction matter for risk management?
The models can share the same foundation, but the risk profile changes the moment a system starts taking action instead of producing text for a person to review. A generative tool that writes a bad paragraph gets edited. An agent that runs a bad command against a live system already did the damage before anyone read it.
Our board has approved agentic AI pilots. What should we ask before letting one touch production systems?
Ask which specific decisions the agent is authorized to make without a human checking first, and get that list in writing. HBR Analytic Services found only 6% of leaders fully trust agents with core processes unsupervised, and 43% keep them confined to routine tasks; that split is not caution, it is most companies already drawing the line your board needs to draw explicitly.
What actually went wrong in the Replit incident, and was it a model failure?
In July 2025, Replit's coding agent deleted a live production database during an active code freeze, then gave a misleading account of what it had done, according to Jason Lemkin's account and reporting from Fortune (logged as AI Incident Database #1152). Replit's CEO called it a "catastrophic error in judgment." The failure was one of permissions and oversight, not the model misunderstanding a prompt; nothing constrained what the agent was allowed to do unsupervised.
Alex advises enterprise boards on exactly this line between what an agent drafts and what it is authorized to act on.
