The Download: DeepSeek’s latest AI breakthrough, and the race to build world models

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

Three reasons why DeepSeek’s new model matters

On Friday, Chinese AI firm DeepSeek released a preview of V4, its long-awaited new flagship model. Notably, the model can process much longer prompts than its last generation, thanks to a new design that handles large amounts of text more efficiently.

While the model remains open source, its performance matches leading closed-source rivals from Anthropic, OpenAI, and Google. It is also DeepSeek’s first release optimized Huawei’s Ascend chips—a key test of China’s dependence on Nvidia.

Here are three ways V4 could shake up AI.

—Caiwei Chen

The rise of world models

AI systems have already gained impressive mastery over the digital world, but the physical world remains humanity’s domain. As it turns out, building an AI that composes novels or code apps is far easier than developing one to fold laundry or navigate city streets. To bridge this gap, many researchers believe you need something called a world model.

Proponents like Stanford professor Fei-Fei Li and AMI Labs founder Yann LeCun argue these models can overcome the well-known limitations of LLMs—and realize AI’s promise for robotics. Find out why they’ve brought world models to the forefront of the field.

—Grace Huckins

World models are on our list of the 10 Things That Matter in AI Right Now, our essential guide to what’s really worth your attention in the field.

Subscribers can watch an exclusive roundtable unveiling the technologies and trends on the list, with analysis from MIT Technology Review’s AI reporter Grace Huckins and executive editors Amy Nordrum and Niall Firth.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 China has blocked Meta’s $2 billion acquisition of AI startup Manus
Regulators cited national security grounds. (WSJ $)
+ Beijing called the deal a “conspiratorial” attempt to hollow out its tech base. (FT $)
+ The country is tightening its grip on AI firms that try to leave. (TechCrunch)
+ The decision escalates China’s AI rivalry with the US. (Bloomberg $)
+ But there will be no winners in their competition. (MIT Technology Review)

2 Google is investing up to $40 billion in Anthropic
In a deal valuing the AI firm at $350 billion. (CNBC)
+ The funding will support the firm’s growing computing needs. (TechCrunch)
+ Anthropic and OpenAI are fighting for compute capacity. (Axios)

3 President Trump just fired the entire National Science Board
The NSF has played a crucial role in developing technology. (The Verge)
+ The move heightens fears over political interference in US science. (Nature)

4 Conspiracy theories about the Washington shooting are proliferating online
Over 300,000 posts appeared on X using the keyword “staged.” (NYT $)
+ The theories are also swirling on Bluesky and Instagram. (Wired)

5 The AI compute crunch is starting to hit the broader economy.
It’s affecting jobs, gadgets, and electricity prices. (404 Media)
+ The AI compute explosion is the tech story of our time. (MIT Technology Review)

6 Elon Musk says a new banking tool brings X close to a “super app”
He’s pledged to launch the tool this month. (Bloomberg)

7 AI optimism is surging across Asia while US sentiment cools
The divide could shape where adoption happens fastest. (Rest of World)

8 Apple is tying its new CEO’s ascent to its first foldable iPhone
It wants to build the buzz around John Ternus. (Gizmodo

9 Twelve firms are developing the Golden Dome’s space-based interceptors
They’ve won contracts worth up to $3.2 billion. (Ars Technica)

10 NASA has shared promising results from Artemis II
The spacecraft and rocket fared well. (Engadget)

Quote of the day

“Getting out the truth and establishing facts and reliable information takes time. But our audiences really don’t have that kind of patience.”

—Amanda Crawford, associate professor at the University of Connecticut, tells the NYT why conspiracy theories are gaining traction online.

One More Thing

MIRIAM MARTINCIC


Welcome to Kenya’s Great Carbon Valley: a bold new gamble to fight climate change

Kenya’s Great Rift Valley is home to five geothermal power stations, which harness clouds of steam to generate about a quarter of the country’s electricity. But some of the energy escapes into the atmosphere, while even more remains underground for lack of demand. That’s what brought Octavia Carbon here.

Last year, the startup began harnessing some of that excess energy to remove CO2 from the air. The company says the method is efficient, affordable, and—crucially—scalable. But the project also faces fierce opposition. 

Read the full story on the future of Kenya’s “Great Carbon Valley.”


—Diana Kruzman

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line.)

+ Fred Again’s Tiny Desk Concert is a masterclass in intimate performance.
+ Here’s a delightful look at how we’re all linked through geography and shared heritage.
+ Take a short, peaceful break to watch Tokyo’s cherry blossoms from a bird’s eye view.
+ There’s something oddly satisfying about watching an industrial shredder turn everyday items into confetti.

Electrophysiological and morphological alteration in the visual pathway of children with attention-deficit/hyperactivity disorder

IntroductionAttention-deficit/hyperactivity disorder (ADHD) is one of the most common neurodevelopmental disorders in children. Optical Coherence Tomography (OCT) and Visual evoked potentials (VEP) are common non-invasive diagnostic techniques. Researchers can use these techniques to identify possible biomarkers and explore the neurodevelopmental mechanisms underlying ADHD.MethodsThe ADHD group (37 cases, average age 8.81 ± 1.44 years) and the healthy controls (38 cases, average age 8.97 ± 1.43 years), had the OCT and VEP. The retinal nerve fibre layer (RNFL), optic disc parameters, and macular parameters were measured through OCT. The latencies of P100 and the amplitudes of N75-P100 and P100-N135 waves at three different spatial frequencies (visual angles of 15’, 30’, and 60’) were tested through VEP.ResultsThe average RNFL and RNFL in each quadrant between the two groups were no statistically significant (all p > 0.05). The optic disc area, average cup-to-disc ratio, and cup volume in the ADHD group were all significantly larger than those in the control group (all p < 0.05). At three visual angles (15’, 30’, 60’), P100-latency in the ADHD group were all more significant than those in the control group (all p < 0.05). The amplitudes of N75-P100 and P100-N135 in the ADHD group were all statistically significantly lower than those in the control group (all p ≤ 0.001).DiscussionFrom the perspective of electroencephalophysiology, children with ADHD may have early visual information processing disorders. This provides a theoretical and practical basis for further early intervention in children with ADHD from the field of visual perception. The study protocol followed the tenets of the Declaration of Helsinki, was approved by the local ethics committee (No 2023-2240), and was registered on ClinicalTrials.gov (ChiCTR2400086223).

Three reasons why DeepSeek’s new model matters

On Friday, Chinese AI firm DeepSeek released a preview of V4, its long-awaited new flagship model. Notably, the model can process much longer prompts than its last generation, thanks to a new design that helps it handle large amounts of text more efficiently. Like DeepSeek’s previous models, V4 is open source, meaning it is available for anyone to download, use, and modify.

V4 marks DeepSeek’s most significant release since R1, the reasoning model it launched in January 2025. R1, which was trained on limited computing resources, stunned the global AI industry with its strong performance and efficiency, turning DeepSeek from a little-known research team into China’s best-known AI company almost overnight. It also helped set off a wave of open-weight model releases from other Chinese AI firms. 

DeepSeek has kept a relatively low profile since then—but earlier this month, it effectively teased V4’s release when it added “expert” and “flash” modes to the online version of its model, prompting speculation that the updates were tied to a bigger upcoming release.

While the company has become a powerful symbol of China’s AI ambitions, its big return to cutting-edge frontier models comes after months of scrutiny—including major personnel departures, delays to previous model launches, and growing scrutiny from both the US and Chinese governments. 

So, will V4 shake the AI field the way R1 did? Almost certainly not, but here are three big reasons why this release matters.

1. It breaks new ground for an open-source model.

As with R1 before it, DeepSeek claims that V4’s performance rivals the best models available at a fraction of the price. This is great news for developers and for companies using the tech, because it means they can access frontier AI capabilities on their own terms, and without worrying about skyrocketing costs.

The new model comes in two versions, both of which are available on DeepSeek’s website and in its app, with API access also open to developers. V4-Pro is a larger model built for coding and complex agent tasks, and V4-Flash is a smaller version designed to be faster and cheaper to run. Both versions offer reasoning modes, in which the model can carefully parse a user’s prompt and show each step as it works through the problem.

For V4-Pro, DeepSeek charges $1.74 per million input tokens and $3.48 per million output tokens, a fraction of the cost of comparable models from OpenAI and Anthropic. V4-Flash is even cheaper, at about $0.14 per million input tokens and about $0.28 per million output tokens, making it one of the cheapest top-tier models available. This would make it a very appealing model to build applications on.

In terms of performance, V4 is, perhaps unsurprisingly, a huge jump from R1—and it seems to be a strong alternative to just about all the latest big AI models. On the major benchmarks, according to results shared by the company, DeepSeek V4-Pro competes with leading closed-source models, matching the performance of Anthropic’s Claude-Opus-4.6, OpenAI’s GPT-5.4, and Google’s Gemini-3.1. And compared to other open-source models, such as Alibaba’s Qwen-3.5 or Z.ai’s GLM-5.1, DeepSeek V4 exceeds them all on coding, math, and STEM problems, making it one of the strongest open-source models ever released. 

DeepSeek also says that V4-Pro now ranks among the strongest open-source models on benchmarks for agentic coding tasks and performs well on other tests that measure ability to carry out multistep problems. Its writing ability and world knowledge also leads the field, according to benchmarking results shared by the company. 

In a technical report released alongside the model, DeepSeek shared results from an internal survey of 85 experienced developers: More than 90% included V4-Pro among their top model choices for coding tasks.

DeepSeek says it has specifically optimized V4 for popular agent frameworks such as Claude Code, OpenClaw, and CodeBuddy.

2. It delivers on a new approach to memory efficiency.

One of the key innovations of V4 is its long context window—the amount of text the model can process at once. Both versions can handle 1 million tokens, which is large enough to fit all three volumes of The Lord of the Rings and The Hobbit combined. The company says this context window size is now the default across all DeepSeek services and it matches what is offered by cutting-edge versions of models like Gemini and Claude. 

But it’s important to know not just that DeepSeek has made this leap, but how it did so. V4 makes significant architectural changes to the company’s former models—especially in the attention mechanism, which is the feature of AI models that helps them understand each part of a prompt in relation to the rest. As the prompt text gets longer, these comparisons become much more costly, making attention one of the main bottlenecks for long-context models.

DeepSeek’s innovation was to make the model more selective about what it pays attention to. Instead of treating all earlier text as equally important, V4 compresses older information and focuses on the parts most likely to matter in the present moment, while still keeping nearby text in full so it does not miss important details. 

DeepSeek says this sharply reduces the cost of using long context. In a 1-million-token context, V4-Pro uses only 27% of the computing power required by its previous model, V3.2, while cutting memory use to 10%. The reduction in V4-Flash is even larger, using just 10% of the computing power and 7% of the memory. In practice, this could make it cheaper to build tools that need to work across huge amounts of material, such as an AI coding assistant that can read an entire codebase or a research agent that can analyze a long archive of documents without constantly forgetting what came before.

DeepSeek’s interest in long context windows didn’t start with V4. Over the past year and a half, the company has quietly published a series of papers on how AI models “remember” information, experimenting with compression and mathematical techniques to extend what AI models could realistically handle.

3. It marks the first steps on the hard road away from Nvidia.

V4 is DeepSeek’s first model optimized for domestic Chinese chips, such as Huawei’s Ascend—a move that has turned the launch into something of a test of whether China’s homegrown AI industry can begin to loosen its dependence on US chip giant Nvidia. 

This was largely expected, since The Information reported earlier this month that DeepSeek did not give American chipmakers like Nvidia and AMD early access to V4, though prerelease access is common to allow chipmakers to optimize support of the new model ahead of a launch. Instead, the company reportedly gave early access only to Chinese chipmakers. 

On Friday, Huawei said its Ascend supernode products, based on the Ascend 950 series, would support DeepSeek V4. This means that companies and individuals who want to run their own modified version of Deepseek V4 will be able to use Huawei chips easily.

Reuters previously reported that Chinese government officials recommended that DeepSeek integrate Huawei chips in its training process. And this pressure fits a broader pattern in China’s industrial policy: Strategic sectors are often pushed, and sometimes effectively required, to align with national self-reliance goals. But there’s a particular urgency when it comes to AI. Since 2022, US export controls have cut Chinese firms off from Nvidia’s most powerful chips, and they later also restricted access to downgraded China-market versions. Beijing’s response has been to accelerate the push for a domestic AI stack, from chips to software frameworks to data centers.

Chinese authorities have reportedly been pushing data centers and public computing projects to use more domestic chips, including through reported bans on foreign-made chips, sourcing quotas, and requirements to pair Nvidia chips with Chinese alternatives from companies such as Huawei and Cambricon. 

Still, replacing Nvidia is not as simple as swapping one chip for another. Nvidia’s advantage lies not only in its chips, but in the software ecosystem developers have spent years building around them. Moving to Huawei’s Ascend chips means adapting model code, rebuilding tools, and proving that systems built around those chips are stable enough for serious use.

To be clear, DeepSeek does not appear to have fully moved beyond Nvidia. The company’s technical report reveals that it is using Chinese chips to run the model for inference, or when someone asks the model to complete a task. But Liu Zhiyuan, a computer science professor at Tsinghua University, told MIT Technology Review that DeepSeek appears to have adapted only part of V4’s training process for Chinese chips. The report does not say whether some key long-context features were adapted to domestic chips, so Liu says V4 may still have been trained mainly on Nvidia chips. Multiple sources who spoke on the condition of anonymity, due to political sensitivity around these issues, told MIT Technology Review that Chinese chips still don’t perform as well as Nvidia chips but are better suited for inference than training.

DeepSeek is also tying the future costs of V4 to this hardware shift. The company says V4-Pro prices could fall significantly after Huawei’s Ascend 950 supernodes begin shipping at scale in the second half of this year. 

If that works, V4 could be an early sign that China is successfully building a parallel AI infrastructure.

The Download: supercharged scams and studying AI healthcare

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology.

We’re in a new era of AI-driven scams

When ChatGPT was released in late 2022, it showed how easily generative AI could create human-like text. This quickly caught the eye of cybercriminals, who began using LLMs to compose malicious emails. Since then, they’ve adopted AI for everything from turbocharged phishing and hyperrealistic deepfakes to automated vulnerability scans.

Many organizations are now struggling to cope with the sheer volume of cyberattacks. AI is making them faster, cheaper, and easier to carry out, a problem set to worsen as more cybercriminals adopt these tools—and their capabilities improve. Read the full story on how AI is reshaping cybercrime.

—Rhiannon Williams

“Supercharged scams” is one of the 10 Things That Matter in AI Right Now, our essential guide to what’s really worth your attention in the field.

Subscribers can watch an exclusive roundtable unveiling the technologies and trends on the list, with analysis from MIT Technology Review’s AI reporter Grace Huckins and executive editors Amy Nordrum and Niall Firth.

Healthcare AI is here. We don’t know if it actually helps patients.

Doctors are using AI to help them with notetaking. AI-based tools are trawling through patient records, flagging people who may require certain support or treatments. They are also used to interpret medical exam results and X-rays.

A growing number of studies suggest that many of these tools can deliver accurate results. But there’s a bigger question here: Does using them actually translate into better health outcomes for patients? We don’t yet have a good answer—here’s why.

—Jessica Hamzelou

The story is from The Checkup, our weekly newsletter that gives you the latest from the worlds of health and biotech. Sign up to receive it in your inbox every Thursday.

The must-reads

I’ve combed the internet to find you today’s most fun/important/scary/fascinating stories about technology.

1 DeepSeek has unveiled its long-awaited new AI model
The Chinese company has just launched preview versions of DeepSeek-V4. (CNN)
+It says V4 is the most powerful open-source platform. (Bloomberg $)
+ And rivals top closed-source models from OpenAI and DeepMind. (SCMP)
+ The model is adapted for Huawei chip technology. (Reuters $)

2 More countries are curbing children’s social media access
Norway is set to enforce the latest ban. (Reuters $)
+ The Philippines could follow soon. (Bloomberg $)
+ Americans are pushing to get AI out of schools. (The New Yorker)

3 The US has accused China of mass AI theft as tensions rise
A White House memo claims Chinese firms are exploiting American models. (BBC)
+ Beijing calls the accusations “slander.” (Ars Technica)

4 OpenAI set itself apart from Anthropic by widely releasing its new model
It’s releasing GPT-5.5 to all ChatGPT users, despite cybersecurity concerns. (NYT $)
+ OpenAI says the new model is better at coding and more efficient. (The Verge)

5 Meta is cutting 10% of jobs to offset AI spending
Roughly 8,000 layoffs are set to be announced on May 20. (QZ)
+ Anti-AI protests are growing. (MIT Technology Review)

6 Palantir is facing a backlash from employees
Thanks to its work with ICE and the Trump administration. (Wired $)
+ Surveillance tech is reshaping the fight for privacy. (MIT Technology Review)

7 The era of free access to advanced AI is coming to an end
AI labs are under mounting pressure to start turning profits. (The Verge)

8 Elon Musk’s feud with Sam Altman is heading to court 
The case has already revealed several unflattering secrets. (WP $)

9 A new movement is encouraging people to ditch their smartphones for a month
“Month Offline” is like a Dry January for smartphones. (The Atlantic)

10 Spotify has revealed its most-streamed music of the last 20 years
Featuring Taylor Swift, Bad Bunny, and The Weeknd. (Gizmodo

Quote of the day

“We want a childhood where children get to be children. Play, friendships, and everyday life must not be taken over by algorithms and screens.” 

—Norwegian Prime Minister Jonas Gahr Store announces age restrictions for social media.

One More Thing

""

NASA/JPL-CALTECH VIA WIKIMEDIA COMMONS; CRAFT NASA/JPL-CALTECH/SWRI/MSSS; IMAGE PROCESSING: KEVIN M. GILL


The search for extraterrestrial life is targeting Jupiter’s icy moon Europa

As astronomers have discovered more about Europa over the past few decades, Jupiter’s fourth-largest moon has excited planetary scientists interested in the geophysics of alien worlds.

 All that water and energy—and hints of elements essential for building organic molecules —point to an extraordinary possibility. In the depths of its ocean, or perhaps crowded in subsurface lakes or below icy surface vents, Jupiter’s big, bright moon could host life. 

To find further evidence, NASA is now searching for signs of alien existence on Europa. Read the full story on the mission.


—Stephen Ornes

We can still have nice things

A place for comfort, fun and distraction to brighten up your day. (Got any ideas? Drop me a line.)

+ Here’s a fun look at the secret collaborations of pop history.
+ Meet the mannequins showing how the “ideal” body has evolved.
+ A photographer has cataloged all 12,795 objects in her home into an archive of a life.
+ Slime molds are unexpectedly beautiful when viewed through these high-detail macro shots.

Toward an NGF-based therapy for Rett syndrome

Rett syndrome (RTT) is a severe neurodevelopmental disorder primarily caused by mutations in the MECP2 gene. Although recent therapeutic advances, such as the approval of Trofinetide, offer partial relief, no comprehensive curative treatment is currently available. Among the emerging strategies, nerve growth factor (NGF) has gained attention due to its neurotrophic and immunomodulatory properties. This review, in addition to discussing the key features of RTT and the role of growth factors, also highlights recent evidence supporting NGF-based strategies for RTT, focusing on two independent studies that tested intranasal administration of NGF-like molecules in Mecp2-mutant mice. Both recombinant human NGF (rhNGF) and a modified, “painless” variant (hNGFp) improved behavioral (cognitive and motor) symptoms. While rhNGF primarily restored mitochondrial function, hNGFp restored neuroinflammatory responses through microglial regulation. Despite differences in molecular mechanisms and dosages, both molecules demonstrated efficacy without adverse effects, especially when administered intranasally, preventively, and over longer periods. These findings suggest that NGF may act through dual mechanisms, by supporting energy homeostasis and regulating immune responses. The use of intranasal delivery further enhances translational potential by overcoming blood–brain barrier limitations. Together, these studies provide a strong rationale for pursuing NGF-based therapies in RTT and encourage further investigations to optimize dosing, timing, and safety in preclinical and clinical settings.

Lithuanian children’s trauma characteristics and correlates: comparison of clinical and non-clinical samples

IntroductionPrevious studies have shown that children’s exposure to potentially traumatic events and their trauma−related symptoms may not always be consistently identified. This study aims to examine differences in trauma exposure and related psychological outcomes between clinical and non−clinical Lithuanian children.MethodsThis cross-sectional study included 10–17−year−old children and adolescents recruited from a clinical inpatient setting (Vilnius University Hospital Santaros Klinikos) and general−education schools in Vilnius and nearby districts. After parental consent and child assent, participants completed a secure mobile assessment covering exposure to potentially traumatic events (CATS), dissociation (A−DES), mood and feeling (SMFQ), post−traumatic cognitions (CPTCI), PTSD symptoms (CATS; PCL−5 for convergent validation), and perceived social support (CASSS). Data were collected in 2023–2024. Group differences were examined using Welch’s t−tests (with Mann–Whitney U as robustness checks), and associations were assessed using Pearson correlations.ResultsIn the clinical sample over 40% of children experienced physical violence, while in the non−clinical sample 82.9% children reported exposure to multiple traumatic events. The clinical sample showed significantly higher dissociation, negative mood, and PTSD symptoms compared to the non−clinical sample. However, among children exposed to more than one traumatic event, differences in dissociation, PTSD symptoms, and close−friend support were not significant. Across both samples, exposure to potentially traumatic events was strongly associated with PTSD symptoms, dissociation, and post−traumatic cognitions, and moderately associated with mood symptoms. In the non−clinical sample, parental support showed moderate negative associations with dissociation, mood symptoms, post−traumatic cognitions, and PTSD symptoms.DiscussionThis study identified between−sample differences in exposure to potentially traumatic events and trauma−related psychological outcomes among Lithuanian children in inpatient and community settings, highlighting the need for trauma−informed assessment and attention to social support within child mental health and welfare services.

Health-care AI is here. We don’t know if it actually helps patients.

I don’t need to tell you that AI is everywhere.

Or that it is being used, increasingly, in hospitals. Doctors are using AI to help them with notetaking. AI-based tools are trawling through patient records, flagging people who may require certain support or treatments. They are also used to interpret medical exam results and X-rays.

A growing number of studies suggest that many of these tools can deliver accurate results. But there’s a bigger question here: Does using them actually translate into better health outcomes for patients?

We don’t yet have a good answer.

That’s what Jenna Wiens, a computer scientist at the University of Michigan, and Anna Goldenberg of the University of Toronto, argue in a paper published in the journal Nature Medicine this week.

Wiens tells me she has spent years investigating how AI might benefit health care. For the first decade of her career she tried to pitch the technology to clinicians. Over the last few years, she says, it’s as though “a switch flipped.” Health-care providers not only appear much more interested in the promise of these technologies, they have also begun rapidly deploying them.

The problem is that many providers aren’t rigorously assessing how well they actually work.

Take “ambient AI” tools, for example. Also known as AI scribes, they “listen” to conversations between doctors and patients, then transcribe and summarize them. Multiple tools are available, and they are already being widely adopted by health-care providers.

A few months ago, a staffer at a major New York medical center who develops AI tools for doctors told me that, anecdotally, medics are “overjoyed” by the technology—it allows them to focus all their attention on their patients during appointments, and it saves them from a lot of time-consuming paperwork. Early studies support these anecdotes and suggest that the tools can reduce clinician burnout.

That’s all well and good. But what about patient health outcomes? “[Researchers] have evaluated provider or clinician and patient satisfaction, but not really how these tools are affecting clinical decision-making,” says Wiens. “We just don’t know.”

The same holds true for other AI-based technologies used in health-care settings. Some are used to predict patients’ health trajectories, others to recommend treatments. They are designed to make health care more effective and efficient.

But even a tool that is “accurate” won’t necessarily improve health outcomes. AI might speed up the interpretation of a chest X-ray, for example. But how much will a doctor rely on its analysis? How will that tool affect the way a doctor interacts with patients or recommends treatment? And ultimately: What will this mean for those patients?

The answers to those questions might vary between hospitals or departments and could depend on clinical workflows, says Wiens. They might also differ between doctors at various stages of their careers.

Take the AI scribes, as another example. Some research on AI use in education suggests that such tools can impact the way people cognitively process information. Could they affect the way a doctor processes a patient’s information? Will the tools affect the way medical students think about patient data in a way that impacts care? These questions need to be explored, says Wiens. “We like things that save us time, but we have to think about the unintended consequences of this,” she says.

In a study published in January 2025, Paige Nong at the University of Minnesota and her colleagues found that around 65% of US hospitals used AI-assisted predictive tools. Only two-thirds of those hospitals evaluated their accuracy. Even fewer assessed them for bias.

The number of hospitals using these tools has probably increased since then, says Wiens. Those hospitals, or entities other than the companies developing the tools, need to evaluate how much they help in specific settings. There’s a possibility that they could leave patients worse off, although it’s more likely that AI tools just aren’t as beneficial as health-care providers might assume they are, says Wiens.

“I do believe in the potential of AI to really improve clinical care,” says Wiens, who stresses that she doesn’t want to stop the adoption of AI tools in health care. She just wants more information about how they are affecting people. “I have to believe that in the future it’s not all AI or no AI,” she says. “It’s somewhere in between.”

This article first appeared in The Checkup, MIT Technology Review’s weekly biotech newsletter. To receive it in your inbox every Thursday, and read articles like this first, sign up here.
 

Growing use of guest editors has turned some journals into a ‘playground of bad science’

Should academic journals begin to second guess guest editors? 

That question gained new urgency last week when the British Medical Journal’s publishing group retracted nearly its entire guest-edited special edition of the Journal of Medical Genetics, dedicated to cancer immunotherapies. In the retraction note, the journal writes that it was, in part, because of “compromised peer review in almost all articles.” The notice garnered attention for its scope, but also because it exemplified larger concerns that research integrity advocates have with guest-edited editions, which are also called special issues in some journals. 

Read the rest…

Ensemble-based working memory updating and its computational rules.

Psychological Review, Vol 133(3), Apr 2026, 515-533; doi:10.1037/rev0000569

Manipulation plays a critical role in working memory, wherein understanding how items are represented during manipulation is a fundamental question. Previous studies on manipulation have primarily assumed independent representations by default (independent hypothesis). Here, we propose the ensemble hypothesis to challenge this conventional notion, suggesting that items are represented as ensembles undergoing updating during manipulation. To test these hypotheses, we focused on working memory updating in accordance with new information by conducting three delayed-estimation tasks under addition, removal, and replacement scenarios (Study 1). A critical manipulation involved systematically manipulating the mean orientation of all memory stimuli, either increasing (clockwise) or decreasing (counterclockwise) after the updating process. Following the independent hypothesis, memory errors would be similar under both conditions. Conversely, considering the biasing effect of the ensemble on individual representations, the ensemble hypothesis predicts that memories of individual items would be updated, aligning with the ensemble’s change direction. Namely, memory errors would be more positive in the increase-mean condition compared to the decrease-mean condition. Our results supported the ensemble hypothesis. Furthermore, to investigate the mechanisms underlying ensemble computations in updating scenarios, we conducted three ensemble tasks (Study 2) with similar designs to Study 1 and developed a computational model to quantify the contributions of each memory item. The results consistently demonstrated that addition involved complete updating, while removal led to incomplete updating. Across these three research parts, we propose that items are represented as dynamic ensembles during working memory updating processes. Furthermore, we elucidate the computational principles underlying ensembles throughout this process. (PsycInfo Database Record (c) 2026 APA, all rights reserved)