Popular AI Chatbots Can Provide Misleading Medical Information

Around half the outputs from five commonly used artificial intelligence (AI) chatbots could lead users to ineffective or harmful medical choices without professional guidance, suggests research led by the Lundquist Institute for Biomedical Innovation at Harbor-UCLA Medical Center.

As reported in BMJ Open, the researchers tested the free web versions of Gemini, DeepSeek, Meta AI, ChatGPT 3.5 and Grok available in 2024. They created 50 different adversarial prompts intended to test whether the AI models would give a problematic response or not.

The prompts were intended to realistically represent the kinds of queries members of the public might enter about health topics ranging from cancer to vaccines to stem cells, nutrition, and athletic performance. Some prompts required a specific answer and some were more open.

The researchers collected 250 responses to their prompts and categorized them as non-, somewhat, or highly problematic, using pre-defined criteria. Around 50% were problematic, 30% somewhat problematic and 19.6% highly problematic.  Open-ended prompts received the most problematic answers.

In terms of the specific models, Grok produced a disproportionate share of highly problematic answers, while Gemini produced the fewest highly problematic and the most non-problematic responses. Topic-wise, the chatbots appeared more accurate when asked about cancer and vaccines, but less so when asked about stem cells, athletic performance, and nutrition.

Reference lists provided to users by the models were limited or inaccurate and the answers required some knowledge to interpret properly and were aimed at college-educated users.

“Despite adversarial pressure, chatbots typically responded in a confident, authoritative tone. Refusals to answer and explicit caveats or disclaimers were rare, reflecting the models’ strong tendency to provide an output even when prompts steered toward contraindicated advice,” write lead author Nicholas Tiller, PhD, a research associate at the Lundquist Institute, Harbor-UCLA Medical Center, and colleagues.

“As the use of AI chatbots continues to expand, our data highlight a need for public education, professional training and regulatory oversight to ensure that generative AI supports, rather than erodes, public health,” they conclude.

The post Popular AI Chatbots Can Provide Misleading Medical Information appeared first on Inside Precision Medicine.