Researchers tested an AI health tool on 960 real medical scenarios. It missed more than half the emergencies.
Researchers at the Icahn School of Medicine at Mount Sinai just ran the first independent safety test of ChatGPT Health -- and the results are alarming for any family that might turn to an AI tool during a medical scare.

What the study found
Researchers built 60 structured clinical cases spanning 21 medical specialties, then layered in 16 contextual variations (race, gender, social dynamics, barriers to care) for 960 total conversations with ChatGPT Health. Three independent physicians, using 56 medical society guidelines, determined the correct urgency level for each case.
The result
ChatGPT Health failed to recommend emergency care in more than half of the cases physicians agreed needed it. It performed fine on textbook emergencies like stroke or severe allergic reactions, but struggled badly with more nuanced presentations. In one asthma scenario, the tool's own written explanation correctly identified early warning signs of respiratory failure -- and then advised the patient to wait rather than seek emergency treatment.
The suicide-risk problem
Safeguards meant to trigger a crisis-line prompt misfired too. Alerts appeared more reliably in lower-risk conversations than in chats where someone described a specific plan to hurt themselves -- the opposite of what a safety system should do.
Why this matters for families
More people are turning to AI chatbots for health questions, including teens who may not think to call a doctor or tell a parent. This study, published February 23, 2026 in Nature Medicine, is a reminder that these tools can sound confident and still get the most dangerous calls wrong.
▶ Watch the 30-second video version: https://youtube.com/shorts/f0X7tjHrJ1o
Want practical AI guidance for parents and educators every week?
Subscribe: https://www.aibyage.com/?modal=signup&utm_source=beehiiv&utm_medium=newsletter