Fifteen Challenges for Generative AI in Cell Biology

Even one of these would have far reaching effects - discuss the implications in-person at #ddpwest26


That is the title of a new perspective in Cell from Andrea Califano (Chair, Dept of Systems Biology, Columbia University) and fifteen co-authors, including Stephen Quake (Stanford), Emma Lundberg (Stanford), and John Tsang (Yale). The suggestion behind it is blunt: generative AI in biology is currently optimizing against the wrong scoreboard, and the field needs a public list of hard, biologically meaningful problems to organize itself around.

The precedent

In 1900, David Hilbert presented a list of 23 unsolved problems in math to the International Congress of Mathematicians in Paris. More than 125 years later, roughly nine are considered conclusively solved. Some were resolved in ways nobody expected, including a few that turned out to be unanswerable as posed. And even the ones that remain unsolved provided value by generating entire fields on the way to not being solved.

The new list

The paper's central claim is that AI's biggest wins in biology, predicting and designing protein structures, worked because those problems had two things going for them that cell biology at large does not. A protein is a string of letters read in order, which is exactly the shape of problem these models are built to handle. And decades of painstaking structural work had already produced a large, standardized library of correct answers to learn from.

Cells behavior is not linear, and there is not much data. LLMs train on trillions of words of text, while even the most ambitious biological profiling efforts (like the Human Immunome Project, or CellxGene) fall roughly a thousandfold short of that. Biology is also combinatorial: things work in groups, not pairs. Building a functional 60S ribosomal subunit takes up to 47 proteins binding in coordination, and studying any two of them tells you nothing about what the 47 do together.

This, the authors argue, is why single-cell foundation models like scGPT and Geneformer keep failing to meaningfully beat simple, decades-old statistical methods when tested on cell types or conditions they were not trained on. The problem is the shape of the biology, not the size of the dataset, so more data will not fix it. Their proposed remedy is to build what we already know about biology into the model itself, using existing maps of which molecules interact with to stop the model from searching everywhere at once.

The second complaint is about how progress gets measured. Models are graded on whether they can reproduce results we already have, not on whether their predictions later turn out to be right. The standard test, sorting cells into types, is one that plain linear regression handles just as well.

The authors offer a cautionary tale. In 1995 the Washington Post declared protein folding solved because a program had correctly predicted six structures out of seven. Those structures were already known. Getting it right on structures nobody had solved yet took another few decades and a Nobel Prize.

So: fifteen challenges across four levels, each with proposed proof-of-concept designs, success metrics, and the non-gen-AI baselines they should be measured against.

Level 1: Molecular interactions

1. Regulatory and signaling interactions. Complex regulatory logic: elucidating the transfer function that allows complex regulatory regions (e.g., a set of >10 super-enhancers) to regulate gene expression, including the formation of TF and cofactor condensates on chromatin.

2. Epigenetic interactions. Epigenetic logic: elucidating the chromatin architecture that determines the implementation of specific gene expression programs, including chromatin methylation state and histone marks based on genomic sequence and baseline transcriptomic and proteomic profiles.

3. Cell-cell interactions. Elucidating the molecular interactions between a ligand-secreting (source) and a receptor-expressing (sink) subpopulation, which determine the transcriptional state of the latter.

Level 2: Molecular function

4. Synthetic mechanisms. Predicting the optimal/minimal plasmid architecture and genomic insertion loci for both constitutive and inducible promoters that would prevent (1) epigenetic or functional promoter silencing following iPSC differentiation into specific lineages and (2) inducible promoter leakiness before induction.

5. Genome to function. Predicting the functional effects of specific de novo mutations, e.g., variants of unknown functional significance (VUFSs) in cancer and other human diseases, as identified by genetic profiles. 

6. Drug mechanism of action. Elucidating the proteome-wide drug mechanism of action and polypharmacology, including high-affinity, off-target, and indirect effector proteins. 

Level 3: Cellular and systems function

7. Genome to phenotype. Predicting the smallest genome sufficient to implement an independently living organism with a desired function, based on multi-omics data from thousands of different micro-organisms.

8. Cell state reprogramming. Predicting genetic or pharmacologic perturbations driving cell state transition, e.g., from an exhausted to a non-exhausted effector CD8+ T cell state, based on existing Perturb and Sequencing (Perturb-seq) assays, or multiome baseline profiles, and drug perturbation profiles. 

9. Logic biocircuit design. Predicting the minimal noise-tolerant genetic circuitry to implement a specific function. 

10. Co-culture and microenvironment. Predicting the minimum number of cell types and reagents to be included in an in vitro co-culture assay to support the viability of cell types that would otherwise not survive in isolation. 

Level 4: Translation

11. Biomarker identification. Identifying multi-omics biomarkers that predict response to an investigational agent based on multi-omics data.

12. Drug toxicity. Predicting dosage-dependent toxicity either in a specific organ (e.g., liver, kidney, heart, brain, blood, skin, etc.) or associated with systemic events (e.g., cytokine storm); toxicity would be predicted prospectively, based on pre-treatment multi-omics data, drug perturbation profiles, and drug structure. 

13. Drug efficacy. Predicting cell-specific drug sensitivity based on drug perturbation profiles, drug structure, and multi-omics data from the specific in vivo disease models.

14. Organismal responses. Predict an individual's immune setpoint and their response to vaccination, infection, and therapy from baseline multi-omics and available consortia.

15. Clinical trial outcomes. Predict the fraction of responders vs. non-responders as well as the mechanism of response. 

Hilbert's list worked because it was public, specific, and hard enough to organize a century of effort. The authors are betting the same structure can pull biological AI away from benchmarks that are already solved and toward the ones that are not. A solution to even 5 of these would be a very good outcome.


Dig deeper into AI problems that matter in pharma and biotech 

DataDrivenPharma West 2026 Days 1 and 2 will do a deep dive into where AI and Data Science are making an impact in Discovery and in Clinical Development. The stories will be told from the perspective of pharma, therapeutics, and diagnostics leaders.

Get #ddpwest26 passes on our website, and for discounts on 3+ non-commercial and 2+ commercial teams. Prices increase EOD August 31, so don’t wait!

See you at #ddpwest26!

Next
Next

Pharma's nine-figure data+AI deals: why not just write a service contract?