We've Been Talking About External Data for Years. Here's What's Actually Changed.

At DataDrivenPharma East 2026, our “Inside the Science” panel brought together data scientists from BMS, GSK, Wave Life Sciences, and PsiThera to answer a deceptively simple question: what's new and what still breaks in external data-driven drug discovery?

Here are three things that stuck with us from this session:

1. The external data sources we rely on today were either non-existent or unconvincing a decade ago.

The panel opened with a deceptively simple question: what can you do with external data now that you couldn't do three years ago? The answers traced a longer arc than three years.

Take human genetics. One panelist recalled the human genome announcement and the biotech crash that followed within a year. All that data, no immediate way to exploit it. The inflection came quietly, over years, through GWAS studies growing in scale, UK Biobank crossing hundreds of thousands of patients, proteomics layered on top. The result: the ability to find rare loss-of-function variants pointing to real biology, not one-off signals. New targets, after 25 years of staring at the same list.

On the clinical side, real world data went through its own credibility journey. Nine years ago, one panelist had to explain to colleagues what it even was. The concept of studying patients outside controlled trial settings felt, at the time, like a contradiction in terms. Now it's the baseline assumption for any serious oncology program.

Liquid biopsy and circulating tumor DNA are the data type still mid-journey. Sensitivity has improved. Standardization is progressing. ctDNA is now a routine exploratory endpoint in oncology trials across the industry. But FDA acceptance as a primary endpoint hasn't arrived yet. The science is ahead of the regulatory framework, and the gap is where a lot of the field's attention is currently focused.

Credibility isn't granted, it accumulates. Scale, standardization, and time are what convert a promising data type into one you can actually build a program around.

2. Fit-for-purpose beats volume. Every time.

There's a pattern in how external data gets sold versus how it actually gets used.

Vendors lead with scale (dataset size, patient counts, breadth of coverage). Practitioners ask a different question: does this data speak to my specific scientific question?

One panelist put it plainly: "We often make a deal because of the volume, but then once you really try to focus on the question, it really becomes clear." The fit-for-purpose frame, defining the scientific question first and then evaluating which data serves it, is the discipline that separates useful external data from expensive noise.

Different providers have different strengths. The same data asset that's transformative for one program is irrelevant for another. Volume is a proxy metric. Specificity is the real one.

3. Harmonizing external data is hard. Harmonizing your internal data is harder.

Ask any data science team what's actually blocking them and you'll get some version of the same answer: the substrate isn't ready. Before you can build a model, run an analysis, or make a decision, someone has to wrangle the data into a shape that's usable. That work is still the majority of the job.

External data harmonization is difficult, but it has a forcing function. Vendors have a commercial incentive to deliver clean-enough data. If they don't, you don't renew. There are also contractual guardrails (you often can't directly compare datasets from competing providers), but at least the problem is bounded. You're evaluating units, assay conditions, methodology thresholds. It's technical work, and it's tractable.

Internal harmonization is a different beast. One panelist called it a psychological problem and "ten years of blood, sweat and tears". Within large organizations, data silos form not just because of technical debt but because of organizational structure, competing priorities, and legacy decisions that nobody wants to revisit. Getting the right people to care, and to keep caring, is much harder than writing the pipeline.

The panel was unanimous on one thing: the organizations that have made real progress on internal harmonization did it by having passionate people sitting at the center of the process and never letting it drop. It doesn't happen on its own.

There was, however, genuine optimism about what's changing. Agentic AI, the ability to pull and reason across data from disparate sources without requiring everything to be pre-unified, is starting to lower the activation energy. You still want harmonized data. But the bar for "good enough to work with" is shifting. For the first time in a while, the panelists felt like the trajectory was pointing in the right direction.

The Bottom Line

What connected all three conversations was a version of the same idea: more data, better models, and improved tooling are necessary. There’s still a lot of nuance to tackle on this topic, and we will further expand in future sessions and panels.

This panel was part of the Inside the Science track at DDP East 2026. DataDrivenPharma West 2026 is coming up October 15 - 16 in South San Francisco. Details and passes available on the conference page.

Previous
Previous

Personal Branding in Biotech and Pharma - panel readout from #ddpeast26

Next
Next

5 Ideas That Stuck With Us from #ddpeast26