Background: Large language models (LLMs) have demonstrated expert-level performance on medical licensing examinations, but most benchmarks focus on final accuracy. Critical gaps remain in understanding model efficiency (latency), the efficacy of tiered “rescue” protocols for error correction, and the systematic correlation between performance and human-rated question difficulty. The German M2 examination, paired with the AMBOSS platform’s user data–driven difficulty ratings, provides an opportunity to map AI performance against human cognitive load. Objective: This study aimed to move beyond singular accuracy scores by (1) evaluating and comparing the baseline (tier 1; T1) accuracy and response latency of next-generation rapid-response LLMs, (2) analyzing the efficacy of a 2-tiered rescue (tier 2; T2) protocol in correcting initial errors, and (3) correlating model performance with the user data–driven AMBOSS difficulty rating. Methods: We evaluated 4 LLMs (Gemini 2.5 Flash, Gemini 2.5 Pro, GPT-5 Instant, and GPT-5 Thinking) on the complete 316-item German M2 (Fall 2024) medical examination, including all multimodal (image-based) questions. A zero-shot copy-paste prompting strategy was used, and outputs were evaluated against ground-truth answers using a strict exact-match criterion. A 2-tiered protocol was used: T1 (Gemini Flash and GPT-5 Instant) provided baseline responses. If incorrect, a T2 (Gemini Pro and GPT-5 Thinking) model was deployed as a “rescue.” Performance was analyzed using the McNemar test, the Wilcoxon signed-rank test, the Fisher exact test, and logistic regression. Results: Baseline (T1) accuracy was identical at 91.5% (289/316; 95% CI 87.85%‐94.06%) for both Gemini 2.5 Flash and GPT-5 Instant, with 27 errors each. However, Gemini Flash (mean 1.57, SD 1.06 s) was significantly faster than GPT-5 Instant (mean 2.07, SD 1.89 s; <.001). Additionally, GPT-5 Instant expended significantly more time on incorrect answers compared with correct ones (=.002), whereas Gemini Flash showed no such hesitation (=.81). The T2 rescue rate for GPT-5 Thinking (13/27, 48.2%; 95% CI 30.74%‐66.01%) was higher, though not statistically significant (=.41), than that for Gemini 2.5 Pro (9/27, 33.3%; 95% CI 18.64%‐52.18%). This rescue protocol elevated final accuracy to 94.3% (298/316; 95% CI 91.18%‐96.37%) for the Gemini system and 95.6% (302/316; 95% CI 92.70%‐97.34%) for the GPT-5 system (=.48). A strong, inverse relationship with difficulty was found: for every 1-point increase in difficulty, the odds of a correct T1 response decreased by 42.1% (odds ratio 0.579, 95% CI 0.425‐0.788; <.001) for Gemini Flash and 47.7% (odds ratio 0.523, 95% CI 0.379‐0.720; <.001) for GPT-5 Instant. This negative correlation persisted even after the rescue (=.01 and =.006, respectively). Conclusions: Expert-level LLM performance on the German M2 examination masks a critical vulnerability: a decrease in accuracy correlated with increased question difficulty. A 2-tiered “rescue” system is an effective strategy to mitigate these difficulty-based failures and achieve >95% accuracy.


