Nature Medicine, Published online: 02 July 2026; doi:10.1038/s41591-026-04500-9
Large language models achieve high scores on health application benchmarks, yet adversarial stress tests now reveal prevalent brittleness — shortcut reliance, fragile visual grounding and fabricated reasoning traces — which exposes substantial gaps between benchmark performance and the robustness evidence needed to support claims of readiness for medical decision-support and patient-facing applications.

