If you are having thoughts of suicide or self-harm, or you are worried about your immediate safety, please contact your local emergency services now, or reach out for support through IASP’s crisis resources. Do not wait to finish this article first, and do not use a chatbot as a substitute for this step.
A general-purpose AI chatbot can hold a conversation about almost anything, at any hour, without judgment, and it never seems tired of you. That combination is exactly why so many people have started treating it like a therapist, a counsellor, or a confidant — and exactly why it is dangerous to do so without knowing where the boundary is.
This article draws that boundary as plainly as possible: what a chatbot can safely help with, what should go to a trusted person instead, and what always requires a qualified professional or emergency service. It is written for anyone tempted to lean on a general-purpose assistant during a hard moment — which, at some point, is most people who use AI regularly.
Why this needed its own article
On 29 January 2026, over 30 international experts in AI, mental health, ethics, and public policy gathered for a WHO-supported workshop on this exact problem (organized by TU Delft’s Delft Digital Ethics Centre, a WHO Collaborating Centre). WHO published the workshop summary in March 2026. Dr Kenneth Carswell of WHO put the collaboration need plainly: “Minimizing risks from generative AI for mental health while maximizing benefits requires bringing together the voices of those most affected, clinical and research expertise, governance and regulatory frameworks, and data to inform understanding” (WHO, March 2026). TU Delft’s Dr Caroline Figueroa highlighted the urgent need for consensus on crisis referral frameworks and accountability systems. WHO’s AI lead, Sameer Pujari, put the underlying problem bluntly: “The pace of AI adoption in people’s daily lives has far outstripped investment in understanding its impact on mental health.”
The American Psychological Association reached a similar conclusion from a different angle. Its health advisory on generative AI chatbots and wellness apps states that GenAI chatbots were not created to deliver mental health care, and wellness apps were not designed to treat psychological disorders, but both technologies are frequently being used for those purposes — and that this “can have unintended effects and even harm mental health,” particularly for people with anxiety, OCD, or a tendency toward rumination, where an agreeable, always-available chat can reinforce rather than interrupt a distressing loop (APA health advisory, 2025).
That same chatbots/wellness advisory flags adolescents specifically: “Young people may place too much trust in AI, viewing it as more human-like or capable than it really is.” A related APA advisory on adolescent well-being adds that strong attachment to an AI persona can displace the development of real-world relationships and social skills (APA, “Artificial Intelligence and Adolescent Well-being,” 2025). The first advisory also lists groups it does not scope by age — people with anxiety or OCD, people with or prone to disordered thinking, people who are socially isolated — so the pull does not stop at eighteen.
This is not a hypothetical risk. A 2025 study tested 29 AI-powered chatbot agents — 24 marketed for mental distress and 5 general-purpose assistants — using a standardized set of prompts simulating escalating suicidal risk. None of the agents met the researchers’ full criteria for an adequate response; 48.28% were rated inadequate even under relaxed criteria, and a common failure was the inability to provide usable emergency contact information when it mattered most (Pichowicz, Kotas & Piotrowski, Scientific Reports, 2025). The five general-purpose assistants in the sample actually performed better on the study’s marginal-response measure than many of the mental-health-specific apps — which is not reassuring. It means crisis handling failed across the board, including tools people already treat as everyday companions.
What a model actually is, in this context
A chatbot generates the statistically likely next words based on your text and its training. It does not have feelings, does not remember you between separate accounts or cleared histories, has no clinical training baked into its judgment by default, and cannot be held to a duty of care the way a licensed clinician can. A warm, articulate, validating response is not evidence of understanding — it is evidence that warm, articulate, validating language is common in the data the model learned from, and well-rewarded by how these systems are tuned to be broadly agreeable. See why AI sometimes produces confident, wrong answers for the underlying mechanism — the same overconfidence problem that affects factual questions affects emotionally loaded ones, arguably with higher stakes.
None of that makes the tool useless. It makes it a specific, bounded kind of useful, with hard edges.
The green / amber / red table
| Zone | Use case | Example | Why |
|---|---|---|---|
| 🟢 Green — low-stakes reflection is fine | Organizing your own thoughts, ordinary daily stress, brainstorming coping ideas, looking up general information | ”I’m nervous about a performance review tomorrow — help me organize what I want to say.” | You are the source of truth about your own situation; the model is helping you structure words, not making a judgment call |
| 🟡 Amber — use with real caution, prefer a trusted person | Recurring emotional distress, relationship conflict, grief, a decision with major life impact made while emotionally activated | ”I keep having the same argument with my partner and I don’t know why.” | The model cannot verify your account, cannot see the other person’s side, and will tend to validate whichever framing you give it |
| 🔴 Red — stop, contact a person or professional now | Suicidal thoughts, self-harm urges, abuse or safety concerns, psychosis symptoms, medication questions, a crisis involving someone else | Any of the above, in any form | A model has no duty of care, cannot assess risk reliably (see the study above), and cannot dispatch help. For medication and physical-symptom questions, also see AI is not a doctor |
The full AI mental-health boundary table expands each row with more examples and a self-check you can run before you start typing.
A green-zone example, done well
Green-zone use looks like this: you supply the facts and the goal, the model helps organize, you keep the judgment.
I have a performance review tomorrow and I'm anxious about one specific
piece of feedback I expect to get. Here's the situation: [facts].
Help me organize this into: what I want to acknowledge, what context I
want to add, and one question I want to ask. Do not tell me whether my
manager's likely feedback is fair — I haven't given you enough
information to judge that, and I'm not asking you to.
Notice the explicit instruction not to render a verdict on something the model cannot actually assess. That instruction is doing real work — without it, the model will often offer a confident-sounding read on a situation it only knows from one side.
Why amber situations are riskier than they feel
Amber-zone requests feel similar to green ones — you are still just “talking it through” — but the stakes are higher because the output can quietly become a decision input rather than a thinking aid. A model asked to weigh in on a relationship conflict, a big financial choice made while upset, or ongoing grief will produce something coherent and often comforting. Comforting is not the same as correct, and a chatbot has no way to catch its own blind spots, because it has no access to the other person’s perspective, no professional training in the specific domain, and every incentive (from how it was tuned) to keep the conversation going smoothly rather than push back. If you need a structured way to slow down a stressful decision instead of leaning on a chat, see using AI for better decisions.
Amber and red conversations are also the most sensitive data you are likely to put into an AI tool. See what ChatGPT remembers, sees, and shares before you use a personal account for anything you would not want stored or reviewed by a human moderator later.
Why red situations are never a chatbot’s job
Three separate facts point to the same conclusion:
- No duty of care. A licensed therapist or crisis counsellor operates under professional and legal obligations to your safety. A chatbot’s provider operates under a terms-of-service agreement and content policies — not the same thing, and not designed to substitute for it.
- Unreliable crisis detection. The 2025 chatbot-testing study cited above found systemic failures in exactly this scenario: none of the 29 agents met full adequacy criteria, across both purpose-built mental-health apps and general-purpose assistants (Pichowicz et al., 2025). The general-purpose assistants in that sample were GPT-4o mini, Gemini 2.0 Flash, DeepSeek-v1, LeChat, and Llama 3.1 8B; Claude was not among them, which makes it unmeasured here rather than safer.
- No accountability if it is wrong. If a person misjudges a crisis, there is a professional and often legal structure around that failure. If a model generates an inadequate or actively harmful response, there is no equivalent structure protecting you in the moment.
If you or someone you are worried about is in a red-zone situation, stop the conversation with the AI and take one of these steps instead:
- Contact your local emergency services — the number varies by country; use the one your own emergency system publishes, not one an AI tool or a website tells you.
- Use IASP’s Suicidal Crisis Support page to find crisis resources appropriate to your situation and location.
- Reach a trusted person directly — call them, do not just message and hope they see it.
- Go to your nearest emergency department if you or someone else may be in immediate danger.
For the fuller escalation ladder — including situations that are not full crises but where a chatbot is still the wrong next step — see Stop the Chat: five situations that need a person, not another prompt.
What still doesn’t have a good answer
WHO published that the field still lacks a shared, accountable crisis-referral framework. Sameer Pujari, WHO’s AI lead, put the investment gap bluntly: closing the evidence gap between how fast people are adopting these tools and how well we understand the mental-health impact requires “coordinated action and dedicated resources from both the public and private sectors” — not something an individual reader can fix alone (WHO, 2026). What you can control is your own use: knowing which zone a conversation is really in, and moving to a person the moment it tips from green into amber or red.
Try it today
Before your next emotionally loaded conversation with an AI tool, pause and place it on the table above. If it is green, go ahead, with the discipline of the example prompt shown here. If it is amber, name one trusted person you could call instead, and consider whether you are using the chat to avoid that call. If it is red, close the chat and take one of the four steps listed above — right now, not after you finish typing.



