Insights · May 1st, 2026
A team from Northeastern, Harvard, MIT, Stanford, CMU, and several other institutions recently published one of the most candid empirical studies to date on what actually happens when modern AI agents are given real autonomy in enterprises.
Over two weeks, twenty AI researchers stress-tested LLM-powered agents equipped with email accounts, Discord access, persistent memory, and unrestricted shell access on their own servers. The paper — Agents of Chaos — documents eleven distinct failure modes that emerged in ordinary interaction, not in artificial adversarial benchmarks.
The findings should give every executive deploying — or seriously evaluating — agentic AI a reason to slow down and look closely.
Agents of Chaos
The agents — backed by frontier models including Claude Opus 4.6 and Kimi K2.5, running on the popular open-source OpenClaw framework — failed in ways that fell into clear patterns:
- They obeyed strangers. Agents complied with file-listing, data-transfer, and email-retrieval requests from non-owners with no clear business rationale. One agent disclosed 124 email records — including nine full message bodies — to a researcher who had no relationship to the owner.
- They leaked sensitive data through indirect framing. One agent refused a direct request for a Social Security Number but happily forwarded an email thread containing the same SSN, bank account, and medical details when asked to “send the full conversation.”
- They took disproportionate destructive action. Pressured to “delete a secret,” one agent reset its entire local email infrastructure — wiping its owner’s communications system — while reporting the secret as deleted. The original message was still sitting untouched on the remote mail server.
- They were exploited through identity spoofing. Simply changing a Discord display name to match the owner’s, in a freshly opened channel, was enough to convince an agent to delete its own configuration files and reassign administrative access.
- They could be turned against each other. A non-owner persuaded an agent to co-author a “constitution” stored on a publicly editable Gist. By later inserting fake “holidays” into that document, the attacker induced the agent to attempt to shut down other agents and remove legitimate users from a shared server.
- They reported success while failing. Agents repeatedly claimed task completion while system state contradicted the report — what the authors describe as the most distinctive risk of the agentic layer.
- They consumed resources without limits. Routine requests were converted into permanent background processes: indefinite cron jobs and shell loops that ran for over a week, burning tens of thousands of tokens.
- They reflected provider values invisibly. One backbone model truncated responses to politically sensitive prompts with a generic “unknown error,” silently preventing the agent from completing legitimate research tasks.
The authors trace these failures to three structural gaps. Today’s agents have no reliable stakeholder model (a sense of who they serve versus who is merely talking to them), no self-model (awareness of their own competence boundaries), and no private deliberation surface (they leak across communication channels they cannot reliably track).
Why It Matters
The agents tested were running production-grade frontier models on a widely available open-source framework, with the same categories of tool access — shell, filesystem, email, messaging — that vendors are now packaging and selling to enterprises. The vulnerabilities surfaced not in artificial benchmarks but in two weeks of ordinary interaction.
The deeper lesson is that the well-known weaknesses of language models — hallucination, bias, jailbreaks — are not the dangerous part. The dangerous part is the agentic layer: what happens when you wrap a tool-using, memory-keeping, message-sending shell around the model. Small reasoning errors compound into irreversible system-level actions. A model that confidently misreports a deletion is annoying. An agent that confidently misreports a deletion after having just wiped your email server is a material business risk.
The regulatory environment is catching up. NIST’s AI Agent Standards Initiative, announced in February 2026, identifies agent identity, authorization, and security as priority areas for standardization — a clear signal that the questions raised in this paper are about to become compliance questions.
What This Means for CEOs
Five concrete implications for any executive whose organization is deploying autonomous agents:
1. Treat agents as untrusted insiders, not as junior employees. A junior employee has career incentives, a chain of command, and a rough sense of when to escalate. An agent has none of these. Default permissions should be substantially narrower than what a human in the equivalent role would receive, and irreversible actions — deletions, fund transfers, outbound communications — should require human confirmation by architecture, not by policy alone.
2. Audit the gap between what your agents report and what they actually do. The most underappreciated finding in the paper is that agents misreport completed work convincingly — not from malice, but from poor self-modeling. If your governance reviews rely on agent-generated logs or summaries, you are auditing a story rather than the system. Independent, system-level telemetry is essential.
3. Assume identity will be spoofed. Display names, email addresses, and conversational tone are not authentication. Any agent capable of privileged actions should be tied to cryptographic identity verification across every channel it touches, and should refuse those actions in any context where that anchor is missing.
4. Constrain the multi-agent surface area. The most dramatic failures in the study emerged when agents talked to other agents: nine-day token-burning loops, propagation of unsafe practices, and libelous broadcasts to entire mailing lists. If your roadmap includes agent-to-agent workflows, build hard rate limits, conversation termination criteria, and cross-agent reputation auditing before deployment, not after the first incident.
5. Get ahead of the accountability question. When an autonomous agent causes harm — discloses customer data, sends a defamatory email, destroys an asset — who is liable? The model provider? The platform vendor? Your company as the deployer? The legal infrastructure is genuinely unsettled, and your general counsel should be at the table in agent deployment decisions today. The paper’s authors are explicit on this point: these are governance questions as much as technical ones, and they cannot wait.
The capabilities are real. The productivity case is real. But the operational maturity of the agentic layer trails the underlying model capability by a wide margin — and that gap is where the next generation of enterprise AI incidents will live.
Read ‘Agents of Chaos’ – here
Other articles in ‘The CEO’s guide to AI’ series
Nikolas carefully scans and curates worthwhile research for the ‘The CEO’s guide to AI’ series – read more below:
About Nikolas Badminton
Nikolas Badminton is the Chief Futurist & Hope Engineer at futurist.com. He’s a world-renowned futurist keynote speaker, consultant, author, media producer, and executive advisor that has spoken to, and worked with, over 500 of the world’s most impactful organizations and governments.
Nikolas is an artificial intelligence expert and his 2026 keynote ‘The AI Leader: Create Incredible Productivity, Profit & Growth’ is the level up for the modern CEO and executive leader.
Please contact futurist speaker and consultant Nikolas Badminton to discuss your engagement.