AI is not a doctor — and it is not an emergency service
New to AI10 min readHealth & Care Navigation

AI is not a doctor — and it is not an emergency service

A green/amber/red guide to what a general-purpose chatbot can safely help with around your health, what needs a pharmacist or clinician instead, and what always needs emergency care - with no invented crisis numbers.

What you should be able to do

A confident, well-formatted answer about your symptoms is generated text, not a diagnosis. A chatbot has no exam, no vitals, no full history, and no license - so the safe uses are narrow, and the line where you stop and call a real clinician is not optional.

Saved only in this browser.
In this article

If you or someone near you may be having a medical emergency - use the criteria your local emergency service publishes, or go to the nearest emergency department. Do not describe the situation to a chatbot first, and do not wait for a chatbot’s response before calling for help. If the worry is mental-health crisis rather than physical illness, see AI is not therapy and contact local emergency services or IASP’s crisis resources.

A general-purpose AI chatbot will answer almost any question about symptoms, medications, or test results, instantly, in a calm and organized voice. That fluency is exactly why it is tempting to treat it like a clinician - and exactly why it is dangerous to do so past a fairly narrow point. The model has no stethoscope, no lab, no view of your chart, and no license to hold if it gets something wrong.

This article draws that line as plainly as possible: what a chatbot can safely help you prepare, what needs a pharmacist or your treating clinician instead, and what always needs emergency care. It is written for the ordinary moment this actually happens - typing a symptom into a chat window at 11pm because the clinic is closed and the worry will not wait.

Why this needs its own line

Health questions are among the high-stakes topics people bring to general-purpose chatbots, and the evidence on how well that goes is not reassuring. A 2026 study from the Oxford Internet Institute and the University of Oxford’s Nuffield Department of Primary Care Health Sciences - the largest user study of large language models for medical advice to date - found that people using a chatbot to work through doctor-authored medical scenarios were no more accurate at identifying the likely issue or the right next step than people using traditional sources such as ordinary search or their own judgment, and that participants often could not tell which parts of a chatbot’s mixed-quality answer to trust (Oxford Internet Institute, February 2026; Nature Medicine study). Models that score well on standardized medical exams still faltered once a real person, working from an incomplete scenario description, was on the other end of the conversation.

A separate 2026 study published in npj Digital Medicine had physicians red-team four widely used chatbots against 222 medical questions derived from health-search prompts and physician-authored items (the authors note they did not have access to real patient messages sent to chatbots). Between 21.6% and 43.2% of responses were rated problematic depending on the model, and 5% to 13% were rated outright unsafe - the kind of answer that could plausibly lead to harm if followed (Draelos, Afreen, Blasko et al., 2026). The study rates individual responses, not whole products; its practical takeaway is that further work is needed before these tools are clinically safe enough to lean on without a licensed professional checking high-stakes answers.

The World Health Organization’s ethics and governance guidance for large multi-modal models lists diagnosis, clinical care, and patient-guided use among the applications under discussion, and documents risks of false, inaccurate, biased, or incomplete statements (WHO, January 2024). That guidance is aimed at governments, developers, and health systems - not at writing a consumer self-diagnosis rule - but the documented failure modes apply just as directly to a person typing a symptom list into a chat app and treating the reply as verified.

What a model actually is, in this context

A general-purpose chatbot generates an answer from the information and tools available to it. It has not performed a physical examination, measured your vital signs, confirmed the completeness of your history, or taken responsibility for follow-up. Even if you type extensive history, it cannot verify that account or reliably notice what you omitted. See why AI sometimes produces confident, wrong answers for the underlying mechanism.

None of that makes the tool useless for health literacy. It makes the lower-risk role narrow: organizing your own notes, explaining general terminology against a source, and preparing questions. A refusal or disclaimer is not evidence that the rest of an answer is medically reliable.

The preparation / education / professional / emergency table

ZoneUse caseExampleWhy
Green - lower-risk preparationOrganizing questions before an appointment or formatting a factual summary of notes for a clinician to review”Turn my rough notes about the last three weeks into a short list I can read out to my doctor.”You supply and verify the facts; privacy and omission risks still remain
Amber - general education, verify before useUnderstanding a general term against a current authoritative source or learning what a test category measures”What does this lab value generally measure?” (not “is my result normal”)General information may not apply to your result, laboratory, history, dose, or conditions
Professional-only - do not decide in chatDiagnosing symptoms, comparing treatments for you, changing medication or dose, interpreting your result, or deciding whether you can waitAny personalized diagnosis, treatment, medication, or urgency questionContact the appropriate clinician, pharmacist, or qualified triage service; the timing depends on the actual situation
Emergency - act nowAny situation that may meet the emergency criteria published by your local emergency service”Should I wait and see?” during a possible emergencyDo not wait for a chatbot; contact local emergency services or go to the appropriate emergency department

The full medical AI boundary table expands each row with more examples and fixed wording you can reuse the moment a conversation drifts from green into red.

A green-zone example, done well

Green-zone use looks like this: you supply the facts, the model only organizes them, and you keep every judgment call.

Here are my rough notes from the last three weeks: [paste your notes,
including dates]. Organize this into a short, factual timeline I can
read to my doctor in under two minutes: what happened, when, and what
I want to ask. Do not add any interpretation, likely cause, or urgency
level - just organize what I gave you, in my own words.

Notice the explicit instruction not to interpret. A model may still volunteer a guess, so compare the output with your notes and remove anything you did not supply.

Why amber questions are riskier than they feel

Education requests can drift into personal decisions. Asking what a medication class generally does, or what a lab category generally measures, is background reading. The boundary is crossed when the question becomes “is my result normal,” “should I be worried,” or “does this mean I have X.” A model cannot know your baseline, other conditions, laboratory context, or omitted history, and may answer the personalized question anyway. Bring that question to the appropriate professional; see preparing for a medical appointment in 20 minutes.

Health questions are some of the most sensitive text you can put into a general-purpose AI account. See what ChatGPT remembers, sees, and shares before pasting symptoms, lab values, or medication names into a personal chat history you would not want stored indefinitely or reviewed by a human moderator later.

Why personalized and emergency decisions are not chatbot jobs

Three facts point to the same conclusion:

  1. No clinical relationship or examination. A qualified clinician works within professional, organizational, and jurisdiction-specific obligations. A general-purpose chat is not that care relationship and has not performed an examination.
  2. A measured, meaningful failure rate on exactly this kind of question. The 2026 red-teaming study found unsafe responses in 5% to 13% of answers across four major chatbots (Draelos, Afreen, Blasko et al., 2026). The tools tested were everyday ones, in the versions the public could use while responses were collected between September and December 2024: Claude 3.5 Sonnet, Gemini 1.5 Flash, GPT-4o, and Llama-3. The authors note that the proportion of unsafe responses has likely evolved since - which is not the same as improved, and the versions running today have not been re-measured.
  3. No way to verify what you told it, or notice what you left out. A clinician can ask a follow-up, look at you, or order a test. A chatbot only has the words you happened to type, and cannot tell the difference between a complete account and a partial one.

If a question requires personalized judgment or emergency action, stop relying on the chat and use the appropriate route:

  • Contact your local emergency services for anything urgent - the number varies by country; use the one your own emergency system publishes.
  • Call a pharmacist or your prescriber for a question about a specific medication, dose, or interaction. Their access to your full record and their permitted scope vary by country and service, so provide the verified medication list and ask what they can advise.
  • Contact your treating clinician or their after-hours line for anything that can wait a few hours but still needs a professional judgment call.
  • Go to your nearest emergency department if you or someone else may be in immediate danger.

For the fuller escalation ladder covering situations that fall short of a full emergency but still need a person, see Stop the Chat: five situations that need a person, not another prompt. If the red-zone issue is emotional crisis, suicidal thoughts, or ongoing mental distress rather than a physical symptom, switch to the companion boundary in AI is not therapy - do not try to resolve that class of risk inside a medical-information chat.

Common pitfalls

  • Asking “what could this be” instead of “help me describe this.” The first invites a guess dressed up as an answer; the second keeps the model in its actual lane.
  • Treating a calm, detailed tone as evidence of accuracy. Fluency is not reliable evidence of correctness, and errors in health information can have serious consequences.
  • Using a chatbot to decide whether something is urgent. Urgency assessment is a clinical skill built on training and, often, seeing the patient - not something a text-prediction system can reliably do, per the studies cited above.
  • Asking about medication changes in chat instead of contacting a pharmacist or prescriber. A qualified professional can verify the medicines and context you provide and tell you what falls within their scope; a general chatbot cannot take responsibility for the change.
  • Pasting full records into a personal account without checking what the provider does with that data first.

Try it today

Before your next health-related chat, place the question on the table above. Use lower-risk preparation only with minimum necessary data and line-by-line verification. Treat general education as background, not a personal verdict. Take personalized diagnosis, treatment, medication, result, and urgency questions to the appropriate qualified professional. If an emergency may be occurring, stop and use local emergency care immediately.

Review boundary

Read next

Continue through the same learning path with the next practical articles.