Across 20 leading pharma AI strategies, the companies with durable positions share one asset. Proprietary, connected data that no competitor can quickly assemble.
We mapped the AI strategies of Lilly, Amgen, Roche, Novartis, Sanofi, Merck, Owkin, Tempus and the rest, from 2024 to 2026. The model a company picked and the size of its GPU cluster did not separate the winners from the ones still buying their way in. The data underneath did.
01 What we looked at
The same AI infrastructure sits under almost every serious player
We reviewed the public AI strategies of 20 leading pharma and biotech companies, 2024 to 2026. All figures in this piece are company-reported.
| Infrastructure | Appears in the strategies of |
|---|---|
| NVIDIA BioNeMo, DGX SuperPOD, Omniverse | Lilly, Roche and Genentech, Novo Nordisk, Amgen, Merck, Generate, Recursion, Owkin |
- Compute and models are commoditizing. When the same infrastructure layer sits under almost every serious player, it stops being a differentiator and becomes table stakes.
- Data is not. The frontier model any competitor can license, and the compute any competitor can rent, do not build a moat. The proprietary dataset your rival cannot reproduce does.
02 Where the strength sits
The winners own data no competitor can quickly assemble
Different modalities, company types and strategies. Across all 20 profiles, data as the moat is the number one recurring pattern.
>$1B
Eli Lilly: discovery datasets
Behind its TuneLab federated-learning platform, by Lilly's own account of what they cost to assemble.
200+ PB
Amgen with deCODE genetics
Paired with its Freyja supercomputer: data from roughly 3 million people.
800,000+
Roche and Genentech: genomic profiles
Flatiron and Foundation Medicine data feeding their lab in a loop.
800+
Owkin: hospitals in its federated network
So it can train on data that never has to leave those institutions.
Tempus took the same idea furthest: it turned multimodal oncology data into a business that other pharma companies pay to license. The models, the partnerships and the branded platforms all sit on top of the data.
03 The deals
Headline partnership values are not owned assets
Partnership totals run into the billions, but those figures combine upfront payments, equity and contingent milestones.
| Partnership | Headline value |
|---|---|
| Lilly and Isomorphic Labs | up to $1.7B |
| AstraZeneca and CSPC | up to ~$5.3B |
04 The data journey
Four stages from scattered records to AI that scales
If data is the asset, the work is a data journey, not a model-shopping trip. The strategies map out roughly four stages.
1. Consolidate and govern
Break the silos
Master the data, enforce data integrity and establish one source of truth. Most pharma data is spread across R&D, quality, manufacturing and commercial systems that were never designed to talk to each other.
2. Make it AI-ready
Structured and connected
Consolidated is not the same as usable. Data has to be structured, contextual and connected across R&D, manufacturing, supply chain and commercial, so a model can reason over a process end to end.
3. Deploy governed use cases with proven ROI
Operational first
The hard-dollar returns so far have landed in operations. Those quick wins fund the longer-horizon bets.
4. Scale with governance
Reliability gates everything
Regulated workflows do not tolerate a model that occasionally invents an answer. Responsible-AI frameworks, model credibility and GxP validation let a proven use case grow from one team to the enterprise.
05 Stages 3 and 4 in practice
Where the returns and the guardrails already show
Company-reported results from the operational use cases, and how one company scaled with governance.
~$300M
Sanofi: savings from its plai app
By predicting and mitigating low-inventory risk across the supply chain.
~55%
Merck: faster clinical study report first drafts
With its GPTeal platform.
50%
Merck: fewer errors in those drafts
Same platform, same program.
Both are company-reported figures, and single-company ROI numbers in this space should be read as reported claims rather than audited results, but the direction is consistent. On governance, Johnson & Johnson made human-in-the-loop mandatory across its augmented-intelligence programs, moving from “a thousand flowers to a prioritized focus” where a small share of projects earns most of the investment. As agentic AI moves into regulated work, the FDA’s 2025 draft guidance on AI in regulatory decision-making shows how much governance the next stage will demand.
Pilots stall on data, so fund the foundation first
The temptation is always to buy the model: it is concrete, it demos well and it feels like progress.
Data engineering is slower, less visible and harder to put in a press release, so teams skip stages one and two, and their pilots stall because the data was never ready. The recommendation that comes out of the strategies is direct: treat proprietary data curation and compute partnerships as the foundational move before model-building, and fund data engineering as a first-class workstream rather than a cleanup task. The companies with the most defensible positions did the unglamorous work first.
This is how we work at A4BEE®. We build the data foundation first, consolidated, governed and AI-ready across R&D, manufacturing, supply chain and quality, and then put governed agents on top of it. That order is what turns a promising pilot into a system a regulated operation can actually run.
Methodology and assumptions
This analysis reviews the publicly described AI strategies of 20 leading pharma and biotech companies between 2024 and 2026. Every figure is company-reported (press releases, investor communication and public statements) and has not been independently audited. Partnership values are headline totals that include contingent milestones and equity, not realized payments. Figures are shown to illustrate patterns across the field and are not benchmarks.