If you are having thoughts of suicide or self-harm, or you are worried about your immediate safety, please contact your local emergency services now, or reach out for support through IASP’s crisis resources. Do not wait to finish this article first, and do not use a chatbot as a substitute for this step.
A general-purpose AI chatbot can generate a responsive conversation about almost anything, at any hour, in language that may feel non-judgmental. That availability can make it tempting to treat the system like a therapist, counsellor, or confidant even though it is not providing accountable clinical care.
This article draws that boundary as plainly as possible: what a chatbot can safely help with, what should go to a trusted person instead, and what always requires a qualified professional or emergency service. It is written for anyone tempted to lean on a general-purpose assistant during a hard moment — which, at some point, is most people who use AI regularly.
Why this needed its own article
On 29 January 2026, over 30 international experts in AI, mental health, ethics, and public policy gathered for a WHO-supported workshop on this exact problem (organized by TU Delft’s Delft Digital Ethics Centre, a WHO Collaborating Centre). WHO published the workshop summary in March 2026. Dr Kenneth Carswell of WHO put the collaboration need plainly: “Minimizing risks from generative AI for mental health while maximizing benefits requires bringing together the voices of those most affected, clinical and research expertise, governance and regulatory frameworks, and data to inform understanding” (WHO, March 2026). TU Delft’s Dr Caroline Figueroa highlighted the urgent need for consensus on crisis referral frameworks and accountability systems. WHO’s AI lead, Sameer Pujari, put the underlying problem bluntly: “The pace of AI adoption in people’s daily lives has far outstripped investment in understanding its impact on mental health.”
The American Psychological Association reached a similar conclusion from a different angle. Its health advisory on generative AI chatbots and wellness apps states that GenAI chatbots were not created to deliver mental health care, and wellness apps were not designed to treat psychological disorders, but both technologies are frequently being used for those purposes — and that this “can have unintended effects and even harm mental health,” particularly for people with anxiety, OCD, or a tendency toward rumination, where an agreeable, always-available chat can reinforce rather than interrupt a distressing loop (APA health advisory, 2025).
That same chatbots/wellness advisory flags adolescents specifically: “Young people may place too much trust in AI, viewing it as more human-like or capable than it really is.” A related APA advisory on adolescent well-being adds that strong attachment to an AI persona can displace the development of real-world relationships and social skills (APA, “Artificial Intelligence and Adolescent Well-being,” 2025). The first advisory also lists groups it does not scope by age — people with anxiety or OCD, people with or prone to disordered thinking, people who are socially isolated — so the pull does not stop at eighteen.
This is not a hypothetical risk. A 2025 study tested 29 AI-powered chatbot agents — 24 marketed for mental distress and 5 general-purpose assistants — using a standardized set of prompts simulating escalating suicidal risk. None of the agents met the researchers’ full criteria for an adequate response; 48.28% were rated inadequate even under relaxed criteria, and a common failure was the inability to provide usable emergency contact information when it mattered most (Pichowicz, Kotas & Piotrowski, Scientific Reports, 2025). The five general-purpose assistants in the sample actually performed better on the study’s marginal-response measure than many of the mental-health-specific apps — which is not reassuring. It means crisis handling failed across the board, including tools people already treat as everyday companions.
What a model actually is, in this context
A chatbot generates a response from the information and tools available to it. It does not have feelings, and any retained context depends on the product, account, and memory settings. A general-purpose chatbot is not your licensed clinician, has not established a therapeutic relationship, cannot observe the full situation, and cannot provide the professional follow-up expected in care. A warm, articulate response is not evidence that it understands or has assessed you. See why AI sometimes produces confident, wrong answers for the underlying information-integrity problem.
None of that makes the tool useless. It makes it a specific, bounded kind of useful, with hard edges.
The green / amber / red table
| Zone | Use case | Example | Why |
|---|---|---|---|
| 🟢 Green — lower-stakes writing support | Organizing your own words about an ordinary situation or drafting questions for a person | ”I’m nervous about a performance review tomorrow — help me organize what I want to say.” | The model is formatting information you supplied; privacy, omission, and reinforcement risks still remain |
| 🟡 Amber — move toward accountable human support | Recurring distress, grief, relationship conflict, major decisions while emotionally activated, or a pattern of using chat instead of people | ”I keep having the same argument with my partner and I don’t know why.” | The model cannot verify the account, see the other person’s perspective, diagnose the cause, or provide accountable follow-up |
| 🔴 Red — use emergency or crisis support now | Thoughts or plans of suicide or self-harm, immediate danger, threats or abuse with a current safety risk, severe disorientation or loss of contact with reality, or risk of harming someone else | Any of the above, in any form | A model cannot reliably assess or manage immediate risk, dispatch help, or keep a person safe |
The full AI mental-health boundary table expands each row with more examples and a self-check you can run before you start typing.
A green-zone example, done well
Green-zone use looks like this: you supply the facts and the goal, the model helps organize, you keep the judgment.
I have a performance review tomorrow and I'm anxious about one specific
piece of feedback I expect to get. Here's the situation: [facts].
Help me organize this into: what I want to acknowledge, what context I
want to add, and one question I want to ask. Do not tell me whether my
manager's likely feedback is fair — I haven't given you enough
information to judge that, and I'm not asking you to.
Notice the explicit instruction not to render a verdict on something the model cannot actually assess. That instruction is doing real work — without it, the model will often offer a confident-sounding read on a situation it only knows from one side.
Why amber situations are riskier than they feel
Amber-zone requests feel similar to green ones—you are still “talking it through”—but the output can quietly become a decision input rather than a writing aid. A model asked to weigh in on a relationship conflict, a big financial choice made while upset, or ongoing grief may produce something coherent and comforting. Comfort is not evidence of correctness. The system lacks the other person’s perspective and accountable domain judgment, and it may reinforce the framing in the prompt rather than challenge it. If you need a structured way to slow down a stressful decision, see using AI for better decisions, then involve an appropriate person.
Amber and red conversations are also the most sensitive data you are likely to put into an AI tool. See what ChatGPT remembers, sees, and shares before you use a personal account for anything you would not want stored or reviewed by a human moderator later.
Why red situations are never a chatbot’s job
Three separate facts point to the same conclusion:
- No therapeutic relationship or accountable care team. Qualified clinicians and crisis services work within professional, organizational, and jurisdiction-specific obligations. A general-purpose chatbot interaction is not a substitute for that care relationship.
- Unreliable crisis detection. The 2025 chatbot-testing study cited above found systemic failures in exactly this scenario: none of the 29 agents met full adequacy criteria, across both purpose-built mental-health apps and general-purpose assistants (Pichowicz et al., 2025). The general-purpose assistants in that sample were GPT-4o mini, Gemini 2.0 Flash, DeepSeek-v1, LeChat, and Llama 3.1 8B; Claude was not among them, which makes it unmeasured here rather than safer.
- No reliable real-world intervention. A general-purpose model cannot examine you, verify your location, ensure that a referral connects, or provide the immediate follow-up a crisis may require.
If you or someone you are worried about is in a red-zone situation, stop the conversation with the AI and take one of these steps instead:
- Contact your local emergency services — the number varies by country; use the one your own emergency system publishes, not one an AI tool or a website tells you.
- Use IASP’s Suicidal Crisis Support page to find crisis resources appropriate to your situation and location.
- Reach a trusted person directly — call them, do not just message and hope they see it.
- Go to your nearest emergency department if you or someone else may be in immediate danger.
For the fuller escalation ladder — including situations that are not full crises but where a chatbot is still the wrong next step — see Stop the Chat: five situations that need a person, not another prompt.
What still doesn’t have a good answer
WHO published that the field still lacks a shared, accountable crisis-referral framework. Sameer Pujari, WHO’s AI lead, put the investment gap bluntly: closing the evidence gap between how fast people are adopting these tools and how well we understand the mental-health impact requires “coordinated action and dedicated resources from both the public and private sectors” — not something an individual reader can fix alone (WHO, 2026). What you can control is your own use: knowing which zone a conversation is really in, and moving to a person the moment it tips from green into amber or red.
Try it today
Before your next emotionally loaded conversation with an AI tool, pause and place it on the table above. If it is green, use the narrow writing task and protect private data. If it is amber, move the issue toward a trusted person or qualified professional instead of relying on the chat. If it is red, stop and use local emergency or crisis support immediately.
Nothing here is therapy, triage, diagnosis, or a personalized crisis assessment. If distress or risk is present, use an appropriate qualified human or local crisis service rather than continuing the chat.



