Lab-in-the-Loop, GenAI, and Why We're Finally Getting Somewhere

I don't think we'll ever get to a biology-agnostic, end-to-end lab-in-the-loop. But with the right use case and the right interface, we can get pretty damn close in specific sub-domains. And to be honest, this is actually the most optimistic I’ve ever been on the “AI will automate science” topic. 

Let me explain.

I first ran into the concept of lab-in-the-loop in 2018, working at a pharmaceutical manufacturing data management company, and then at a “cloud lab” company a couple years later. As a former organic chemist, the idea was immediately attractive. I knew from experience how unforgiving even the smallest nuance in science can be, and how a combination of increased scale, and in-line decision making could be huge.

But two problems kept it from really working: the use case, and the interface.

Use case was a limitation of the AI of the era. ML models needed very specific inputs and outputs, which meant you needed a narrow, well-defined problem to make them useful. They also needed a lot of training data in a very similar domain, which is hard to come by in R&D. The lack of generalization severely limited where you could actually create useful models.

Interface was the other half. You had to invest serious time in learning a specific tech stack, and even when you understood it well, bugs at the boundary between software and physical equipment were constant. That was fine for specialists in the technology, but a deal-breaker for scientists whose job was to do science. They didn't want to learn a whole new framework to do work they already knew how to do, especially when that framework was untested and may not stick around.

Both problems still exist. But GenAI and agentic systems make them a lot more surmountable. These models can be pre-trained on broad scientific domains and generalize within them. They can interface with humans in natural language, with much simpler UIs sitting on top. It's not a clean removal of either barrier, but it's a productive shift.

Think of a pizza making lab-in-the-loop. In the pre-GenAI version, you'd design a rigid automation system, outsource most of the ingredients, and end up with something that mostly just assembles. You may have seen one in an airport. The assembly is super good enough, but the "lab" part is where it falls apart. If you have feedback on the pizza, there's no good way to convey it, and no confidence the system understands you. Let’s instead imagine that the bot runs off of a pizza+italian cuisine foundation model and you can describe in plain language what you like and don't like, and what the bot can do about it. Now the system is genuinely useful. Not a full kitchen, but a real tool with real pizza-domain value.

This brings me to our sponsor, Amazon Web Services (AWS), and their recent release of Amazon Bio Discovery. I think it's a solid embodiment of what "constrained use case plus the right interface" looks like in practice. They smartly chose Antibody discovery as a use case, which is a great choice because it’s constrained enough, high-value enough, and amenable enough to today's technology to actually work. The interface is a model orchestration system combined with a catalog of 40+ AI biology models or user-imported ones. A scientist can take a target, choose a sequence, and design an antibody (or nanobody, bispecific, or another flavor). The process is AI-assisted, allowing the scientist to sanity-check logic and assumptions against literature and related work. Once they're happy with the constructs, which can run into the thousands, they rank them with built-in or custom models and send the most promising candidates for synthesis through integrated CROs (Ginkgo Bioworks, Twist Bioscience, A-Alpha Bio). The outputs are then screened, and the results are fed back in for the next cycle.

What I like about it is that the wet lab scientist gets their hands dry (get it?) with compute tools without needing to become a software engineer. And the data scientist can fine-tune the process with custom models, context, and data specific to their therapeutic area. There's an Memorial Sloan Kettering Cancer Center (MSK) example here, and you can try it yourself here. And get this, the class to understand the usage basics is only one hour.

AWS will give a 10-minute talk on Amazon Bio Discovery at DataDrivenPharma Forum at Genentech on May 27. The event is at 88% capacity and will sell out, so grab your pass soon.

If you want to keep this conversation going, we'll also have a talk and a roundtable on lab-in-the-loop at #ddpeast26 in Boston. Discounted and free passes are available this month for some pharma, therapeutics, and diagnostics company employees. Reply and ask me about it.

To come back to where I started: I don't think we'll ever get to a biology-agnostic, end-to-end lab-in-the-loop. But with the right use case and the right interface, we can get close enough in specific sub-domains that it actually changes how the work gets done. This take isn’t going to make headlines, but I think it’s most likely to actually play out.

Upcoming DataDrivenPharma Events

  • May 27: DataDrivenPharma Forum @Genentech (SSF)

    • Theme: AI’s Role in Drug Development Costs: Reverse or Reinforce? 

  • June 10: DataDrivenPharma East in Waltham, MA

    • Our Boston Area gathering of hands-on data science leaders across discovery, clinical, and strategy.

    • Some employees of pharma, therapeutics, or diagnostics companies can get discounted or Free passes, respond to this email to learn more

  • Online: Join our DataDrivenPharma Slack community here!

Previous
Previous

Is AI Actually Reducing the Cost of Drug Development?

Next
Next

The Context Switch