AlphaGo's Architect Returns: Thore Graepel Raises Millions for Metis Reasoning
AlphaGo co-creator Thore Graepel has left Google DeepMind to found Metis Reasoning, a startup betting that AlphaGo-style search-and-planning — not bigger LLMs — is the road to machines that can act in the physical world.
One of the minds behind AlphaGo is going independent — and he is betting against the idea that bigger language models alone will get us to real intelligence.
Bloomberg reported on September 24 that Thore Graepel, a co-creator of AlphaGo and until this summer a research leader at Google DeepMind, is raising tens of millions of dollars for a new startup called Metis Reasoning. The company’s target: artificial intelligence that can respond to unfamiliar problems and choose an action, with applications in robotics, science, and engineering. A follow-on tranche worth hundreds of millions of dollars at a higher valuation is already planned, though people familiar with the talks caution that both the structure and the size of the funding could still change.
The name is a thesis. In Greek mythology, Metis was the Titaness of wisdom and deep counsel — cunning intelligence applied to practical problems. That is precisely the lineage Graepel is drawing on.
The man behind the machines that reason
Graepel’s career reads like a pre-history of the current AI boom. Trained as a physicist in Hamburg, Imperial College London, and Berlin, he earned his PhD in machine learning in 2001. At Microsoft Research from 2003, he co-founded the Online Services and Advertising group, where his probabilistic reasoning systems shipped at enormous scale: TrueSkill, the rating and matchmaking system behind Xbox Live, and AdPredictor, the click-through prediction model behind Bing. Both rest on message passing over factor graphs — old-school Bayesian reasoning doing industrial work.
Then came DeepMind, where Graepel joined the team that built AlphaGo, the first program to defeat a professional player at full-sized Go — a feat most researchers believed was still a decade away when it happened in 2016. He went on to work on AlphaGo’s heirs: AlphaGo Zero, which learned entirely from self-play with no human game data; AlphaZero, which generalized the approach to chess and shogi; and MuZero, which mastered games without even being told the rules.
His personal site, updated for the new venture, describes the through-line: bringing “AlphaGo-style reasoning to frontier AI, so that machines can plan and act under uncertainty.” After DeepMind, Graepel led machine learning for cellular rejuvenation at Altos Labs, spent a stint on Google DeepMind’s Post-AGI team, and holds the Chair of Machine Learning at University College London. Now those threads — search, learned judgement, multi-agent systems, biology — converge at Metis.
The technical bet: search beats scale
Why does this matter beyond another well-funded lab launch? Because Metis Reasoning is making a pointed architectural wager at a moment when the industry’s default answer to every capability gap is “more parameters, more tokens.”
A large language model, at its core, predicts the next word from human-written text. Graepel’s project, as his site puts it, argues that “search and learned judgement care little where the uncertainty springs from — a game tree, a cell, an agent choosing what to do next.” The AlphaGo recipe was never about a giant network; it was about a policy that suggests moves, a value function that weighs them, and a tree that looks ahead. That combination — learned intuition fused with deliberate search — is what produced Move 37, the shoulder-hit on the fifth line against Lee Sedol that human commentators first called a mistake and then recognized as a masterpiece.
Metis Reasoning’s pitch is that this recipe transfers: to robots that must plan under physical uncertainty, to scientific problems where the “game tree” is a space of experiments, to engineering systems that must weigh options before committing. It is, in effect, the post-LLM thesis — systems that learn from their own actions and deliberate before acting, rather than interpolating from human text.
A DeepMind diaspora is being financed
Graepel’s exit is part of a broader pattern that investors have noticed. His departure from Alphabet follows that of David Silver, who led DeepMind’s reinforcement learning team and in April raised $1.1 billion at a $5.1 billion valuation for Ineffable Intelligence, with Sequoia and Lightspeed co-leading and Nvidia and Google participating. Yann LeCun’s AMI Labs raised $1.03 billion at a $3.5 billion pre-money valuation in March to build world models — systems that predict how an environment responds to an action. And Emulate, founded in August by three DeepMind veterans, is in talks to raise up to $700 million at a $3.7 billion valuation, the third lab to spin out of DeepMind’s London offices this year with nine-figure financing.
The financing structure itself is a sign of the times. Tranched rounds — raising capital in steps — are spreading among AI startups because new labs want a high headline valuation to signal confidence and attract talent, while early backers want in at a lower price before later investors mark it up. Metis Reasoning’s planned second tranche of hundreds of millions follows exactly this playbook.
What to watch
Metis Reasoning has not shipped a product, and the talks are described as preliminary. But the signals are worth tracking for anyone mapping where frontier AI goes next:
- Domain choice. Robotics, science, and engineering are domains where pure next-token prediction struggles and planning genuinely pays — and where physical and experimental feedback loops provide exactly the trial-and-error signal that reinforcement learning needs.
- Compute efficiency. If AlphaGo-style search genuinely composes with frontier-scale models, it could deliver capabilities at a fraction of the brute-force training cost — a proposition that becomes more attractive every quarter as pure-scale economics tighten.
- The talent magnet. A founder with AlphaGo, AlphaZero, and MuZero on his record, plus a UCL chair, will draw researchers who want to work on agency rather than chatbots. In this market, that is a moat.
- The verification question. Planning-based systems are, in principle, more inspectable than opaque monoliths — you can examine the tree. If safety regulators increasingly demand interpretable decision-making, that could become a commercial advantage, not just an ethical one.
The takeaway
For a decade, the AlphaGo lineage — search plus learned judgement — has lived mostly inside game engines and research demos while the industry poured capital into language models. Graepel’s Metis Reasoning is one of the clearest signals yet that serious money and serious pedigree are now converging on the idea that the next capability jump comes from teaching frontier models to deliberate — to plan, weigh, and act under uncertainty — rather than merely to predict.
Whether Metis becomes the next DeepMind or a footnote depends on execution nobody can see yet. But the bet itself marks a shift in the AI landscape’s center of gravity: from the chat window to the world.
Coverage based on reporting by Bloomberg (September 24, 2026), PYMNTS, and AI Weekly; biographical details from Thore Graepel’s personal site.
Sources
- [1] https://www.bloomberg.com/news/articles/2026-09-24/ex-deepmind-researcher-thore-graepel-raising-funds-for-ai-reasoning-startup
- [2] https://thoregraepel.github.io/
- [3] https://www.pymnts.com/news/artificial-intelligence/2026/deepmind-alumni-draw-investors-beyond-large-language-models/
- [4] https://aiweekly.co/alerts/alphago-co-creator-thore-graepel-raises-tens-of-millions-for-ai-reasoning