← All posts / Industry

No Pretraining Required: Tsinghua Professor's Naive AI Hits $1.4B Valuation on Open-Weight Bet

Seven months after founding, Jifeng Dai's stealth startup has raised $400M and is betting everything on mid-training and post-training over an existing Chinese open-weight base model.

No Pretraining Required: Tsinghua Professor's Naive AI Hits $1.4B Valuation on Open-Weight Bet

A seven-month-old startup with fewer than 100 employees, no product on the market, and a website consisting of exactly one sentence — “100x Intelligence for the pioneers” — is now valued at more than $1.4 billion. On September 18, 2026, The Information reported that Naive AI, the Beijing-based LLM startup founded in February by Tsinghua University associate professor Jifeng Dai, has closed three funding rounds totaling roughly $400 million, with its latest round valuing the company at about $1.42 billion (over 9.5 billion RMB, approaching the 10-billion-RMB “unicorn” threshold in Chinese market terms).

The round-by-round breakdown is unusually aggressive even by 2026 standards: $100 million, $180 million, and $120 million across three successive raises. Investors include Tencent, IDG Capital, MPCi, and HSG (the firm formerly known as Sequoia China). The speed of the stack — nearly half a billion dollars committed inside a single year to a company that has never shipped a model — says as much about the current state of the Chinese AI market as it does about Naive AI itself.

What Naive AI is actually building

The technical bet is what makes this story more than another funding headline. Naive AI is not training a foundation model from scratch. Instead of repeating the enormously expensive pretraining pass — the data collection, the tens of thousands of GPU-hours, the full pipeline from raw corpus to base model — the company takes an existing Chinese open-weight model as its starting point. Which model exactly remains undisclosed. From there, it restructures the model’s architecture and continues training with reinforcement learning and related post-training techniques to push performance on targeted tasks.

Dai’s wager is that mid-training and post-training are where the remaining headroom lives. His team’s thesis: even when the base model comes from someone else’s lab, there is still enormous performance to be extracted from architecture adjustments after pretraining, mid-training refinement, RL, and task-specific post-training. The company is also researching recursive self-improvement — AI models participating in upgrading their own capabilities — a direction currently being pursued head-on by OpenAI, Anthropic, Google, and Zhipu among the frontier labs.

This route is no longer an outlier. When Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, released its first LLM Inkling in July, it acknowledged that the architecture substantially followed DeepSeek’s open-weight V3 model released in late 2024. The Information also notes that Cursor, after its acquisition by SpaceX, built its own coding models on top of Zhipu’s models. As frontier base models become freely downloadable commodities, a new class of AI labs is choosing to skip the pretraining arms race entirely and concentrate capital on reasoning, agents, RL, and task-specific capability instead.

The founder is the product

For a company this stealthy, the core asset is Jifeng Dai’s track record. Dai earned his bachelor’s degree (2009) and Ph.D. (2014) in automation at Tsinghua, with Zhou Jie as his doctoral advisor. From 2014 to 2019 he worked in the vision group at Microsoft Research Asia, rising to Principal Researcher and Research Manager. From 2019 to 2022 he was Executive R&D Director at SenseTime Research. He joined Tsinghua’s Department of Electronic Engineering full-time in July 2022, where his research focuses on agentic AI and continual learning, aimed at AGI.

Two artifacts anchor his reputation in the open-source community. At Tsinghua he led development of InternVL, the open-weight multimodal model family that became a staple of academic research and the open-source ecosystem. And in 2025 he joined Shanda-backed AI startup MiroMind, where he contributed to MiroThinker — an open-source agent capable of multi-step web research and financial analysis that drew significant developer attention.

Then came the split. In January 2026 Dai left MiroMind over disagreements about team direction — the other co-founders wanted to relocate the team outside China. One month later he founded Naive AI in Beijing. The subtext of the funding story is that Chinese capital is betting on the founder who stayed.

Why the post-training thesis matters now

The timing of Naive AI’s rise tracks a structural shift in how the industry thinks about base models. DeepSeek, Moonshot, Zhipu, and Alibaba have largely stabilized China’s foundation-model landscape — the pretraining race among those players is mature, and their strongest models are downloadable. If the marginal gains in raw pretraining are shrinking while open-weight bases keep improving, then the scarce skill shifts from “burn compute on a corpus” to “extract capability from an existing base.”

That is precisely the skill set Dai has spent a decade accumulating: InternVL was an exercise in careful post-training over existing vision-language foundations, MiroThinker was an exercise in agentic capability engineering, and continual learning — his stated academic focus — is the study of how models improve after their initial training ends.

For the open-source ecosystem, a well-funded lab whose entire strategy is sophisticated post-training over open weights is good news twice over: it validates open-weight models as legitimate foundations rather than cheap substitutes, and its outputs — Naive AI plans to release its first model, also named Naive, as early as this month, with open weights and free re-use for downstream development — flow back into the commons.

The test arrives almost immediately

The Information reports that the first model could ship within weeks, released open-weight. A sub-100-person team, $400 million raised, and a valuation flirting with $1.42 billion will be judged against DeepSeek, Qwen, Kimi, and GLM the moment weights drop — models built by teams with far more compute and far longer track records at scale.

The bear case is straightforward: post-training gains compress as base models absorb the same techniques upstream, and a thin company on someone else’s foundation has no moat if the foundation provider iterates. The bull case: if mid-training and recursive self-improvement really are where the headroom is, Naive AI is one of the purest public bets on that thesis — and its first release, expected imminently, will be the earliest readable signal of whether the post-training-only lab is a durable category or a 2026 artifact.

Either way, the fact that investors pushed a no-product, one-sentence-website startup to a $1.4 billion valuation in seven months is a snapshot of where AI funding sentiment sits right now: pretraining is commoditizing, founders with proven open-source track records are the scarce resource, and the market is paying unicorn prices for a credible claim on whatever comes after the base model.