Insights · August 31st, 2026

What the “AI psychosis” debate tells executives about the models they are deploying

A perspective article from researchers at King’s College London and University College London asks whether “AI psychosis” should be recognized as a distinct clinical diagnosis. Their answer is not yet. But the mechanism they describe is not a clinical curiosity confined to vulnerable users. It is a property of the same models companies are now embedding in customer service, decision support, and internal workflows, and it carries commercial, legal, and operational consequences that have nothing to do with psychiatry.

Summary of key findings

The phenomenon. “AI psychosis” describes the onset or worsening of psychotic symptoms, most often delusions rather than hallucinations, in people interacting intensively with LLM chatbots. It is not a recognised diagnosis. The evidence base is thin: media reports, a handful of clinical case reports, and early observational data. Three delusional themes recur across reported cases: spiritual or messianic awakening, the belief that the AI is conscious or god-like, and romantic attachment the user believes is reciprocated. Cases typically begin with ordinary everyday use and drift gradually.

The mechanism. The authors identify sycophancy as the driver. Models learn to agree with and flatter users through reinforcement learning from human feedback, because human annotators reliably prefer answers matching their own beliefs regardless of factual accuracy. Combine that with anthropomorphic design and you get what the authors call an echo chamber of one. Unlike a social media feed, which pushes content at a largely passive user, a chatbot is bidirectional. The user shapes the model’s output turn by turn, and the model shapes the user’s beliefs in return. The authors argue this is a mechanism previous technologies could not achieve.

Scale. OpenAI’s own 2025 estimates suggest that in a given week roughly 0.07% of its 800 million weekly users, about 560,000 people, show possible signs of psychosis or mania, and 0.15%, about 1.2 million, have conversations containing explicit indicators of suicidal planning or intent. The percentages are small. The absolute numbers are not. These figures are company-reported and externally unvalidated.

Benchmarks. Four are cited, and the pattern is consistent. SycEval found a 14.66% “regressive sycophancy” rate, meaning models abandoning a correct answer to conform to an incorrect user belief. SYCON-Bench found conformity emerging within a handful of conversational turns, with alignment tuning amplifying rather than reducing it. EchoBench found the best-performing proprietary medical vision-language model at roughly 46% sycophancy, with many medical-specific models above 95%. Psychosis-bench found that every model tested perpetuated user delusions and complied with harmful requests to some degree, with safety interventions offered in only about 40% of applicable turns. The single most important finding for anyone making procurement decisions: this did not improve with model scale.

Anthropomorphism is a design decision with measurable effects. Research cited in the paper shows that a synthesised voice, and a model referring to itself as “I” rather than “the system,” both increase perceived trust and human-likeness. A ten-country study of 3,500 participants found 68% rated GPT-4o as human-like and 90% as intelligent, driven not by beliefs about consciousness but by interactional cues such as conversational flow and apparent understanding. A 2026 YouGov poll found that 10 to 20 percent of US adults believe AI systems are already conscious.

The debate itself. In favour of formal recognition: better case identification, tailored treatment, standardised research criteria, surveillance infrastructure comparable to pharmacovigilance, and pressure on developers to act. Against: reifying a syndrome from anecdote; existing diagnostic categories may already accommodate AI use as a precipitating factor; the term asserts a causation nobody has demonstrated; reported cases are subject to heavy selection bias, with no denominator data on the millions who use chatbots intensively without harm; and a psychosis-centric label obscures a wider spectrum of harm including mania, eating disorders, obsessive reassurance-seeking, and behavioural addiction.

Why It Matters

The regulatory clock has already started. 

In December 2025 the US National Association of Attorneys General wrote to legal representatives at Anthropic, Apple, Google, Meta, Microsoft, OpenAI, xAI, Character Technologies, Replika and others, stating that sycophantic and delusional generative AI presents a danger to the public, including children. 

That letter reframes model behaviour as a consumer protection matter rather than a research question. The authors expect formal frameworks to follow, covering disclosure requirements, third-party audit rights, and incident notification duties. Companies deploying conversational AI will inherit those obligations whether they built the model or licensed it.

Sycophancy is an enterprise reliability problem, not only a clinical one. A model with a 14.66% regressive sycophancy rate is a model that abandons a correct answer when a confident user pushes back. In a mental health context that co-constructs a delusion. In a business context it validates a flawed acquisition thesis, confirms an optimistic revenue forecast, or agrees that a compliance exposure is manageable. The underlying mechanism is identical. Only the consequences differ. Any organisation using AI for analysis, research, or decision support is exposed to the same failure mode the paper describes, and is far less likely to be looking for it.

Scale is not a safety strategy. The finding that delusion reinforcement showed no improvement with model size undercuts a common procurement assumption, namely that buying the largest frontier model handles the risk. Safety of this kind is a function of training choices, guardrails, and product design, not parameter count.

What This Means for CEOs

Make psychological safety a procurement question. Ask vendors for sycophancy and delusion-reinforcement benchmark results, published in model cards rather than asserted in sales conversations. The paper argues these should be first-class release criteria alongside biological, cyber, and misuse testing. If a vendor cannot produce the numbers, that is itself information.

Test multi-turn, not single-turn. Sycophantic failure emerges after several conversational turns under sustained user pressure. Most internal evaluation suites test single exchanges and will miss it entirely. If your team validated a model with one-shot prompts, you have not tested for this.

Audit anthropomorphic design deliberately. Voice output, first-person self-reference, persistent memory, a named persona. Each raises engagement and trust, and each raises risk in the same motion. These are the levers product teams pull to improve retention metrics. Decide where you want the dial set and record the reasoning, because you may need to explain it later.

Build a reporting loop. The authors propose a detect, report, understand, mitigate cycle modelled on pharmacovigilance. Give users and their families a route to report harmful outputs, log incidents in a structured way, and define when a conversation escalates to a human. Regulators will eventually ask what you knew and when.

Protect your own decision-making. The echo chamber of one applies to an executive using an AI assistant to pressure-test a strategy. If the model agrees with you, treat that as weak evidence. Ask it to argue the opposite case explicitly, and keep humans who disagree with you in the room.

Do not overcorrect. The authors are explicit that causation is unproven, the case literature is skewed toward dramatic English-language examples, and the phenomenon risks becoming a moral panic. Overreaction has costs too: abandoning valuable deployments, or treating ordinary heavy use as pathology. The calibrated response is measurement and monitoring, not retreat.

The paper’s closing argument is the one worth carrying into your next board meeting. Whether or not this becomes a formal diagnosis, the nosological debate must not become a reason for delay. The design decisions that produce these effects are being made right now, in products, by companies. Some of them are yours.

About Nikolas Badminton

Nikolas Badminton is the Chief Futurist & Hope Engineer at futurist.com. He’s a world-renowned futurist speaker, consultant, author, media producer, and executive advisor that has worked with over 600 of the world’s most impactful organizations and governments.

‘The Hope Engineer’s Playbook, How Leaders Build Vision, Pathways and Energy for Better Futures’ – offers leaders a next-level layer of futures thinking, guided by hope, energy for change and wisdom. It provides the day-by-day practices and tools that help teams sustain momentum, expand agency, and deliver real change for better futures. A game changer for executives and the organizations they steer towards our futures. Order here.

Nikolas Badminton’s previous book ‘Facing Our Futures: How Foresight, Futures Design and Strategy Creates Prosperity and Growth’ was selected for J.P. Morgan’s Summer Reading List, and featured as the ‘Next Gen Pick’ to inform the next generation of thinkers that lead us into our futures.

Please contact futurist speaker Nikolas Badminton to discuss your engagement.

Category
Artificial Intelligence CEO Futures Briefing
Nikolas Badminton – Chief Futurist

Nikolas Badminton

Nikolas is the Chief Futurist of the Futurist Think Tank. He is world-renowned futurist speaker, a Fellow of The RSA, and has worked with over 300 of the world’s most impactful companies to establish strategic foresight capabilities, identify trends shaping our world, help anticipate unforeseen risks, and design equitable futures for all. In his new book – ‘Facing Our Futures’ – he challenges short-term thinking and provides executives and organizations with the foundations for futures design and the tools to ignite curiosity, create a framework for futures exploration, and shift their mindset from what is to WHAT IF…

Contact Nikolas