A new study examining medical artificial intelligence tools found the same flaw across four major platforms, including systems from OpenAI, Anthropic, Doximity, and OpenEvidence, according to Fortune. The finding raises concerns about the reliability of AI when applied to clinical medicine.
The study tested how each platform handled medical questions and found a consistent problem shared by all four. The researchers did not find that one system performed significantly better than the others on this particular issue. The flaw appeared to be widespread rather than unique to any single product.
AI tools for medicine have expanded rapidly. Physicians, nurses, and patients have increasingly turned to these systems for information about diagnoses, treatments, and drug interactions. Proponents argue that AI can help extend the reach of medical knowledge and reduce the burden on overburdened health systems.
But critics have raised concerns about accuracy and consistency. Medical information is high-stakes, and errors in how AI systems present clinical data can contribute to misdiagnosis or incorrect treatment decisions. The new study adds specific evidence to those concerns by identifying a shared flaw that cuts across competing products.
The platforms tested represent some of the most prominent names in both general-purpose AI and AI designed specifically for healthcare. OpenEvidence and Doximity are targeted at medical professionals, while OpenAI and Anthropic produce general large language models that have been widely adopted in healthcare settings.
Fortune reported the findings as part of ongoing coverage of AI's role in medicine. The study does not call for any of the platforms to be taken offline, but it does point to the need for rigorous, independent evaluation of medical AI tools before they are widely deployed in clinical settings. The specific nature of the flaw identified was described in the study but was not detailed further in available reporting.
