AI Model Predicts 348 Diseases from Electronic Health Record, Genetics

Research led by Harvard University shows an artificial intelligence algorithm can predict how likely it is that a given patient will develop 348 diseases based on data collected from electronic health records (EHRs) and genetic data.

As reported in Nature, the researchers built a new computational tool called ALADYNOULLI, a statistical machine learning model known as a Bayesian generative model, that jointly analyzes EHR data and genetics to model how disease risk evolves over an individual’s lifetime.

They applied it to over 683,000 participants across three independent biobanks, the UK Biobank, Mass General Brigham, and All of Us, covering 348 diseases and up to 52 years of follow-up.

ALADYNOULLI identified 21 disease signatures, which the researchers defined as clusters of conditions that tend to co-occur and evolve together over time, such as cardiovascular or metabolic disease. These were remarkably consistent across all three biobanks.

Crucially, the model also revealed biological subtypes within the same diagnosis: for example, early-onset and late-onset heart attacks showed distinct signature trajectories, suggesting different underlying mechanisms. Genetics also contributed to disease predictions, for example, in the case of cardiovascular disease the team found 23 cardiovascular variants missed by conventional single-disease analyses.

“We have painstakingly curated these signatures, which is a big differentiator. In contrast to deep learning approaches, which are typically ‘black boxes,’ our curated signatures capture the underlying biology in an interpretable way,” said co-lead author Giovanni Parmigiani, PhD, a Dana-Farber researcher and associate director of the Division of Population Sciences, in a press statement. “These signatures are then the drivers of the model’s ability to make predictions.”

For disease prediction, ALADYNOULLI substantially outperformed established clinical risk scores such as the Pooled Cohort Equation and PREVENT for cardiovascular disease, and the Gail model for breast cancer.

“There is a lot of useful information in a medical record, both over time and across different disease areas,” says Parmigiani. “That data would be difficult for a human to process in their head, but tractable for a machine learning model.”

Current medicine largely treats diseases in isolation and uses static, single-disease risk scores. ALADYNOULLI offers a unified, continuously updating health prediction for each patient that integrates their full diagnostic history and genetic predisposition simultaneously. This has major implications for earlier and more precise risk prediction, patient stratification in clinical trials, and genetic discovery.

“People are thinking about what their health is going to look like over the next few years, especially with increasing intervention options. This tool offers a path toward improving the prediction of future diseases so doctors and patients can take action to try to prevent them,” said co-senior author Alexander Gusev, PhD, a Dana-Farber scientist.

The team behind the model is now working to make the model broader and more accurate. They are also looking for opportunities to test the model in clinical practice and to help design better clinical trials.

The post AI Model Predicts 348 Diseases from Electronic Health Record, Genetics appeared first on Inside Precision Medicine.

Machine Learning Simplifies Accurate LDL Cholesterol Calculation

Researchers at Johns Hopkins University have developed a machine learning-based version of the widely used Martin-Hopkins equation that simplifies the calculation of low-density lipoprotein cholesterol (LDL-C) without compromising accuracy.

The new approach, which is published in JAMA Cardiology, could make it easier for laboratories to estimate LDL-C, improving treatment decisions for patients at risk of cardiovascular disease.

“We’ve optimized the calculation of LDL cholesterol and made this equation accessible and easier for all labs to implement,” said Seth Martin, MD, MHS, senior study author and director of the Advanced Lipid Disorders Program and Digital Health Lab at the Johns Hopkins Ciccarone Center for the Prevention of Cardiovascular Disease. “Our goal is to enable clinicians and patients to make better decisions about starting treatments that prevent heart attacks and strokes, and save lives.”

LDL-C is a major cause of atherosclerotic cardiovascular disease (ASCVD) and a primary treatment target. Current guidelines recommend using LDL-C cut offs, such as 70 mg/dL or 55 mg/dL (to convert to mmol/L multiply by 0.0259) in patients with ASCVD, to guide clinical lipid management.

The gold standard for measuring LDL-C concentration is preparative ultracentrifugation but this method is expensive and time-consuming. LDL-C concentrations are therefore usually estimated in routine practice.

One of the most accurate ways to estimate LDL-C concentration is the Martin-Hopkins method, which is recommended for clinical use in the U.S., Europe, and South America. However, implementation can be difficult because it requires users to look up an adjustable factor in a large table that is based on the patient’s triglyceride and non-high-density lipoprotein cholesterol levels.

“A lipid profile with low cholesterol and high triglycerides is the ultimate stress test of the LDL cholesterol calculation,” said Martin. He explains that a 5, 10 or 20 mg/dL difference, based on various equations, could change a person’s eligibility for treatment, such as with PCSK9 inhibitors, which have been shown to significantly lower LDL cholesterol levels. “It’s these types of on-the-cusp examples that benefit most from more accurate results,” he added.

To overcome this barrier and facilitate implementation, Martin and team used a transparent machine learning approach—multivariate adaptive regression splines—to create a simplified, formula-based LDL-C equation.

They trained and tested the tool on data from 4,939,528 adults and children (mean age, 56 years; 53% women) with complete lipid panel test results. These samples, which are representative of the U.S. population, had a median LDL cholesterol level of 114 mg/dL and came from the Very Large Database of Lipids.

The researchers report in JAMA Cardiology that the machine-learning version of the Martin-Hopkins equation estimated LDL-C concentrations that were similar to the original equation, with a minimal difference of 0.5 mg/dL.

Both Martin-Hopkins equations classified 90% of samples within the correct treatment category. Among other commonly used tools for LDL-C estimation, the Sampson-NIH equation correctly classified 86%, the modified Sampson-NIH equation classified 85%, and the Friedewald equation classified 83% in the correct category.

Importantly, said Martin, the investigators found that the Martin-Hopkins equations were the most accurate for classifying high-risk patients with lower ranges of LDL cholesterol levels.

When it came to assessing people who had triglyceride levels between 200 mg/dL and 399 mg/dL and LDL cholesterol levels less than 70 mg/dL, the Martin-Hopkins machine learning equation accurately classified 84% of high-risk samples, the original Martin-Hopkins equation classified 83%, the modified Sampson-NIH equation classified 72%, the Sampson-NIH equation classified 61%, and the Friedewald equation classified 40%.

Martin and co-authors conclude: “Given its high accuracy and straightforward implementation as a single line of code in laboratory information systems, [the Martin-Hopkins machine learning equation] is an alternative option to consider implementing in practice.”

The post Machine Learning Simplifies Accurate LDL Cholesterol Calculation appeared first on Inside Precision Medicine.

<![CDATA[Explore how dopamine D2 blockade shapes antipsychotic benefits and risks, revealing dosing pitfalls, polypharmacy harms, and why plasma level monitoring improves outcomes.]]>

What to know about the parasitic diarrhea outbreak

Get your daily dose of health and medicine every weekday with STAT’s free newsletter Morning Rounds. Sign up here.

Good morning. I dedicate the first item in today’s newsletter to a friend who is suffering from what she suspects to be cyclosporiasis. Feel free to forward this email to your friends who need information but are scared to wade through online discussions on the outbreak. 

Read the rest…

STAT+: FTC settles lawsuit with CVS Caremark over charges it manipulated insulin prices, impeded access

The Federal Trade Commission settled a lawsuit against CVS Caremark, one of the largest pharmacy benefit managers in the U.S., over allegations that the company artificially inflated the price of insulin and impeded access to the lifesaving diabetes treatment.

As part of the deal, which the agency maintained will save Americans up to $8.5 billion in out-of-pocket costs over 10 years, CVS Caremark, which is owned by CVS Health, must make several changes to its dealings with employers, health plans, and pharmacies. The FTC also estimated the deal will unlock up to $4.5 billion in further savings for patients through pharmacy counter rebates.

In its complaint, the FTC alleged that CVS Caremark — as well as Cigna’s Express Scripts and UnitedHealth’s Optum Rx — created a “perverse” system of rebates that favored insulin, which was then sold at higher list prices in order to “line their pockets” at the expense of patients who were forced to pay more for the medication.

Continue to STAT+ to read the full story…

STAT+: Sales from controversial U.S. drug discount program rose to $100 billion last year

Prescription medicines purchased in the U.S. under a controversial government discount program amounted to $100 billion in 2025, a 22.8% increase from the previous year, according to the Health Resources and Services Administration, which oversees the program.

Expensive medicines represented an increasing proportion of spending in the 340B Drug Discount Program, accounting for $61.9 billion, or nearly 62% of all prescription drugs purchased through the program. Nearly $8.9 billion was spent on Merck’s Keytruda immunotherapy treatment, followed by more than $4.47 billion on Biktarvy, an HIV medicine sold by Gilead Sciences.

The data mark a steady rise in sales under the 340B program, which requires drugmakers to offer discounts that are typically estimated to be 25% to 50% — but could be higher — off all outpatient drugs to hospitals and clinics that primarily serve lower-income patients. 

Continue to STAT+ to read the full story…