← All posts / Industry

Evolution as Training Data: Basecamp Research's $140M Series C Bets on Programmable Biology

London's Basecamp Research raised a $140M Series C led by S32 with NVIDIA and Anthropic's Anthology Fund on the cap table, valuing it at $800M. Its EDEN models, trained on 15 trillion DNA tokens from 30+ countries, have already designed an antibiotic that matched a last-resort drug in mice.

Evolution as Training Data: Basecamp Research's $140M Series C Bets on Programmable Biology

On Wednesday, Basecamp Research announced a $140 million Series C that values the London-based AI drug discovery company at $800 million. The round was led by S32, with NVIDIA, Anthropic’s Anthology Fund, the NATO Innovation Fund, and Redalpine participating, according to Reuters, GEN News, and The Decoder.

That investor list reads like a map of 2026’s AI power centers — a chipmaker, a frontier lab, a defense-oriented sovereign fund, and a deep-tech European VC — all converging on a company whose core asset is not a language model but a database of nature: genetic material gathered from rainforest soil, volcanic ground, hot springs, and the deep sea off Antarctica, paired with the chemical and ecological conditions of every sampling site.

What the money buys

Basecamp, founded in 2020, will use the capital for two things: training the next generation of its EDEN biological foundation models, and pushing its own therapy candidates toward clinical development. The first target is in vivo cell therapy — genetic treatments that reprogram a patient’s cells inside the body, rather than in a manufacturing plant.

The company’s pitch lands on a specific pain point. CAR-T cell therapies, in which a patient’s immune cells are extracted, genetically modified, multiplied, and reinfused to fight cancer, can cost hundreds of thousands of dollars per patient, largely because of that ex vivo manufacturing loop. If EDEN-designed insertion tools could perform the genetic modification directly inside the patient, the economics of the entire field change.

The data thesis: biology’s wall is a data wall

CTO Philip Lorenz frames the company’s strategy around a single statistic. Language models train on text corpora that Epoch AI caps at roughly five quadrillion tokens — an upper bound on all words ever written. Basecamp estimates the number of nucleotides on Earth at 10^37. “If you took a stack of cards of 10^37 poker cards, that stack would surround the observable universe a million times,” Lorenz told The Decoder. Nobody needs a dataset that size — but it frames how early biological AI remains.

Public genome databases can’t carry the load either. Basecamp’s own BaseData paper found that about 68% of the sequence volume in the NIH Sequence Read Archive comes from just five species; humans alone account for 54%. “If you were to train an LLM only on newspaper articles from 1975, it would be a really, really bad model,” Lorenz said in a Microsoft case study. “That is kind of where we are in biology.”

So Basecamp collects its own data. The network now spans more than 30 countries and all seven continents — Microsoft counts 208 participating organizations across 31 countries — and the dataset currently holds about 15 trillion tokens, where a token is a single DNA nucleotide. That already puts it in the same numerical range as the text datasets behind frontier chatbots. Over the next year and a half, the plan is to grow it roughly a hundredfold past one quadrillion tokens under the Trillion Gene Atlas initiative, announced in March together with Anthropic, NVIDIA, PacBio, and Ultima Genomics.

Proof points: an antibiotic and a gene-insertion toolkit

Two results anchor the claim that this data translates into medical value.

First, antibiotics. Basecamp prompts EDEN with a pathogen, and the model designs antimicrobial peptides to kill it. In the company’s (not yet peer-reviewed) EDEN paper, 97% of a curated set of tested candidates showed activity in the lab. More striking: a candidate called EDEN-7, generated directly by the model with no rounds of tweaking, performed about as well as a last-resort antibiotic in mice infected with multidrug-resistant bacteria. That work was done with Cesar de la Fuente’s lab at the University of Pennsylvania.

Second, gene insertion. Basecamp uses EDEN to design large serine recombinases — enzymes, originally forged in phage-bacteria arms races, that can rejoin DNA at chosen spots. Programmed to insert a healthy copy of a gene regardless of the patient’s specific mutation, they target a core problem of gene therapy for inherited disease. In the EDEN paper, half of the generated recombinases were active in human cells, and an earlier announcement reported that primary human T cells modified this way cleared more than 90% of tumor cells in the lab. The pipeline published on the company’s website lists four programs in lead optimization: in vivo CAR-T for blood cancer, in vivo CAR-T for an autoimmune disease, a liver gene therapy for phenylketonuria (PKU), and antimicrobial peptides against resistant pathogens.

Scaling curves, with caveats

Lorenz’s claims about data quality are the most technically interesting part of the story. Over the past six months, Basecamp trained identical architectures under identical conditions on its own data versus public data — and its own data produced a steeper scaling curve, with the gap widening as models grew. To reach a perplexity of 2, Basecamp’s data reportedly saves several hundred thousand to millions of GPU hours; for a perplexity of 1.5, the extrapolation points to billions of GPU hours saved. That last figure is a projection, not a measurement.

He is equally candid that benchmark scores don’t equal better molecules. When the team compared the StripedHyena architecture (the backbone of the Arc Institute’s Evo model) against Meta’s Llama, StripedHyena sometimes achieved lower perplexity — but Llama performed better on downstream biological tasks. Basecamp went with Llama. A third of the company’s GPUs are currently reserved for reinforcement learning experiments, using data from its Boston lab, where T-cell experiments yield hundreds to thousands of data points rather than billions of tokens, to steer broadly trained models toward clinical tasks.

The Anthropic connection

Selected EDEN capabilities are already available through Claude and Claude Science: researchers describe a task in plain language — designing antibiotics or picking vaccine targets — and Claude taps EDEN to suggest candidates. That integration helps explain Anthropic’s presence in the round. It also previews a distribution model in which frontier chatbots become the front door to specialized scientific models, with EDEN as an early proof of concept.

Open questions

The honest caveats are the story’s other half. None of Basecamp’s six programs has moved past lead optimization; everything rests on lab results and mouse data so far. Delivery remains the hard problem — Arc Institute co-founder Patrick Hsu has publicly named delivery, toxicity, and immune response as open challenges for insertion tools. And the benefit-sharing model, under which source countries receive licensing payments (52 beneficiaries across 19 countries by end of 2024; as little as nine months from sampling to first payment), has drawn criticism that roughly 1% of revenue is too small a share for the communities whose biodiversity drives the company’s value. Basecamp keeps every EDEN training token traceable to its geographic origin and the consent granted — no small claim — but “traceable” and “fairly compensated” are different bars.

The bet, in Lorenz’s own words: “Our ambition is to make biology programmable. You prompt on a disease and out comes a molecule that will address that.” With $140 million more in the bank, the next eighteen months — the hundredfold data expansion, the Trillion Gene Atlas, and the first attempts to move a program toward the clinic — will test whether evolution really can be treated as training data.