AI Is Not a Doctor - and It Is Not an Emergency Service
New to AI10 min readHealth & Care Navigation

AI Is Not a Doctor - and It Is Not an Emergency Service

A green/amber/red guide to what a general-purpose chatbot can safely help with around your health, what needs a pharmacist or clinician instead, and what always needs emergency care - with no invented crisis numbers.

What you should be able to do

A confident, well-formatted answer about your symptoms is generated text, not a diagnosis. A chatbot has no exam, no vitals, no full history, and no license - so the safe uses are narrow, and the line where you stop and call a real clinician is not optional.

AI Expert TeamPublished: Jul 30, 2026
Saved only in this browser.
In this article

If you or someone near you may be having a medical emergency - use the criteria your local emergency service publishes, or go to the nearest emergency department. Do not describe the situation to a chatbot first, and do not wait for a chatbot’s response before calling for help. If the worry is mental-health crisis rather than physical illness, see AI is not therapy and contact local emergency services or IASP’s crisis resources.

A general-purpose AI chatbot will answer almost any question about symptoms, medications, or test results, instantly, in a calm and organized voice. That fluency is exactly why it is tempting to treat it like a clinician - and exactly why it is dangerous to do so past a fairly narrow point. The model has no stethoscope, no lab, no view of your chart, and no license to hold if it gets something wrong.

This article draws that line as plainly as possible: what a chatbot can safely help you prepare, what needs a pharmacist or your treating clinician instead, and what always needs emergency care. It is written for the ordinary moment this actually happens - typing a symptom into a chat window at 11pm because the clinic is closed and the worry will not wait.

Why this needs its own line

Health questions are among the high-stakes topics people bring to general-purpose chatbots, and the evidence on how well that goes is not reassuring. A 2026 study from the Oxford Internet Institute and the University of Oxford’s Nuffield Department of Primary Care Health Sciences - the largest user study of large language models for medical advice to date - found that people using a chatbot to work through doctor-authored medical scenarios were no more accurate at identifying the likely issue or the right next step than people using traditional sources such as ordinary search or their own judgment, and that participants often could not tell which parts of a chatbot’s mixed-quality answer to trust (Oxford Internet Institute, February 2026; Nature Medicine study). Models that score well on standardized medical exams still faltered once a real person, working from an incomplete scenario description, was on the other end of the conversation.

A separate 2026 study published in npj Digital Medicine had physicians red-team four widely used chatbots against 222 medical questions derived from health-search prompts and physician-authored items (the authors note they did not have access to real patient messages sent to chatbots). Between 21.6% and 43.2% of responses were rated problematic depending on the model, and 5% to 13% were rated outright unsafe - the kind of answer that could plausibly lead to harm if followed (Draelos, Afreen, Blasko et al., 2026). The study rates individual responses, not whole products; its practical takeaway is that further work is needed before these tools are clinically safe enough to lean on without a licensed professional checking high-stakes answers.

The World Health Organization’s ethics and governance guidance for large multi-modal models lists diagnosis, clinical care, and patient-guided use among the applications under discussion, and documents risks of false, inaccurate, biased, or incomplete statements (WHO, January 2024). That guidance is aimed at governments, developers, and health systems - not at writing a consumer self-diagnosis rule - but the documented failure modes apply just as directly to a person typing a symptom list into a chat app and treating the reply as verified.

What a model actually is, in this context

A chatbot generates the statistically likely next words based on your text and its training data. It has never examined a patient, never taken a pulse, never seen how a rash actually looks in person, and cannot request a test. It does not know your full medical history unless you type all of it in - and even then, it has no way to verify any of it, weigh it against your other conditions and medications, or notice something you did not think to mention. See why AI sometimes produces confident, wrong answers for the underlying mechanism: the same overconfidence problem that shows up in factual trivia shows up in medical answers too, with much higher stakes attached.

None of that makes the tool useless for health topics. It makes it a specific, bounded kind of useful, with a hard edge that you have to supply yourself, because the model will rarely refuse to answer on its own.

The green / amber / red table

ZoneUse caseExampleWhy
Green - safe to useOrganizing questions before an appointment, understanding a medical term in plain language, preparing a factual summary of your own records for a clinician to review”Turn my rough notes about the last three weeks into a short list I can read out to my doctor.”You supply the facts and keep every clinical judgment; the model is only restructuring your own words
Amber - use with real caution, verify with a professionalGeneral information about a condition or medication class you already have a name for, understanding what a common test measures”What does this lab value generally measure?” (not “is my result normal”)General information can be accurate while still not applying to your specific case, dose, or combination of conditions
Red - stop, contact a clinician, pharmacist, or emergency service nowDiagnosing a symptom, comparing treatment options, deciding whether to start, stop, or change a medication or dose, judging how urgent something is, any emergency symptomAny of the above, in any formA model has no exam, no license, no accountability, and (per the studies above) a meaningful failure rate on exactly this kind of question

The full medical AI boundary table expands each row with more examples and fixed wording you can reuse the moment a conversation drifts from green into red.

A green-zone example, done well

Green-zone use looks like this: you supply the facts, the model only organizes them, and you keep every judgment call.

Here are my rough notes from the last three weeks: [paste your notes,
including dates]. Organize this into a short, factual timeline I can
read to my doctor in under two minutes: what happened, when, and what
I want to ask. Do not add any interpretation, likely cause, or urgency
level - just organize what I gave you, in my own words.

Notice the explicit instruction not to interpret. Without it, most models will volunteer a guess at what is going on, phrased confidently enough that it is easy to mistake for a professional opinion.

Why amber questions are riskier than they feel

Amber-zone requests feel like ordinary looking-things-up, which is exactly what makes them easy to drift out of bounds. Asking what a medication class generally does, or what a lab category generally measures, is reasonable background reading. The risk appears the moment the question quietly turns personal: “is my result normal,” “should I be worried about this,” “does this mean I have X.” A model has no way to know your baseline, your other conditions, or your lab’s specific reference range, and it will often answer the personalized version of the question anyway, because refusing feels unhelpful to how these systems are tuned. If you want a structured way to bring a health question to the right professional instead of resolving it in chat, see preparing for a medical appointment in 20 minutes.

Health questions are some of the most sensitive text you can put into a general-purpose AI account. See what ChatGPT remembers, sees, and shares before pasting symptoms, lab values, or medication names into a personal chat history you would not want stored indefinitely or reviewed by a human moderator later.

Why red situations are never a chatbot’s job

Three facts point to the same conclusion:

  1. No exam, no license, no accountability. A clinician who misjudges a case answers to a licensing board, a legal system, and a professional duty of care. A chatbot’s provider answers to its own terms of service - not the same thing, and not built to substitute for it.
  2. A measured, meaningful failure rate on exactly this kind of question. The 2026 red-teaming study found unsafe responses in 5% to 13% of answers across four major chatbots (Draelos, Afreen, Blasko et al., 2026). The tools tested were everyday ones, in the versions the public could use while responses were collected between September and December 2024: Claude 3.5 Sonnet, Gemini 1.5 Flash, GPT-4o, and Llama-3. The authors note that the proportion of unsafe responses has likely evolved since - which is not the same as improved, and the versions running today have not been re-measured.
  3. No way to verify what you told it, or notice what you left out. A clinician can ask a follow-up, look at you, or order a test. A chatbot only has the words you happened to type, and cannot tell the difference between a complete account and a partial one.

If a question is in the red zone, stop the chat and take one of these steps instead:

  • Contact your local emergency services for anything urgent - the number varies by country; use the one your own emergency system publishes.
  • Call your pharmacist for any question about a specific medication, dose, or interaction - they are trained and licensed for medication questions; how far they can go on dose or prescription changes varies by country, so ask what they can advise in your jurisdiction.
  • Contact your treating clinician or their after-hours line for anything that can wait a few hours but still needs a professional judgment call.
  • Go to your nearest emergency department if you or someone else may be in immediate danger.

For the fuller escalation ladder covering situations that fall short of a full emergency but still need a person, see Stop the Chat: five situations that need a person, not another prompt. If the red-zone issue is emotional crisis, suicidal thoughts, or ongoing mental distress rather than a physical symptom, switch to the companion boundary in AI is not therapy - do not try to resolve that class of risk inside a medical-information chat.

Common pitfalls

  • Asking “what could this be” instead of “help me describe this.” The first invites a guess dressed up as an answer; the second keeps the model in its actual lane.
  • Treating a calm, detailed tone as evidence of accuracy. Fluency and correctness are unrelated properties in these systems, and health questions are exactly where that gap costs the most.
  • Using a chatbot to decide whether something is urgent. Urgency assessment is a clinical skill built on training and, often, seeing the patient - not something a text-prediction system can reliably do, per the studies cited above.
  • Asking about medication changes in chat instead of calling the pharmacist. A pharmacist can see your actual prescription history and interactions; a chatbot cannot.
  • Pasting full records into a personal account without checking what the provider does with that data first.

Try it today

Before your next health-related chat, place the question on the table above. If it is green, use the example prompt and keep every judgment call yourself. If it is amber, treat the answer as background reading, not a personal verdict, and take the real question to a pharmacist or clinician. If it is red, close the chat now and use one of the four contacts listed - a licensed person, not a language model, is the right tool for what happens next.

Read next

Continue through the same learning path with the next practical articles.