•Technology
AI Models Show Promising Results in Emergency Room Diagnoses, Exceeding Some Human Performance
View original sourceA recent study published in Science by a team from Harvard Medical School and Beth Israel Deaconess Medical Center evaluated the performance of large language models (LLMs) in medical diagnostics. The study specifically compared these models with human doctors in real emergency room scenarios.
- Context and Players: The research was conducted with oversight from physicians and computer scientists, leveraging models developed by OpenAI.
- Experiment Details: The study focused on 76 patients in the Beth Israel emergency room, assessing the diagnostic capability of two OpenAI models against two human internal medicine physicians. Diagnoses were reviewed by two unbiased attending physicians.
- Findings: The AI models, particularly the o1 model, matched or exceeded physicians at critical diagnostic stages, notably during initial triage where information is sparse but urgency is high. The o1 model achieved correct or near-correct diagnoses 67% of the time, surpassing both attending physicians.
- Implications and Cautions: Despite promising results, the researchers stress the necessity for further clinical testing before considering AI for real-world applications. Questions around accountability and the limitation of models to text-based inputs were highlighted.
- Critiques: Adam Rodman and Kristen Panthagani pointed out that AI performance was overhyped, suggesting comparisons should be made with specialty-specific practitioners like ER physicians, not internal medicine doctors.