← All posts / Tools

A Crash Test for Every Pull Request: Raindrop Raises $50M to Stop AI Agents Failing in Production

Backed by CRV and researchers from OpenAI and Anthropic, Raindrop launched Simulations, which replays real production traffic against every agent change before it ships.

A Crash Test for Every Pull Request: Raindrop Raises $50M to Stop AI Agents Failing in Production

AI agents have quietly acquired a failure mode that traditional software never had: they do the wrong thing convincingly, at scale, until someone happens to notice. There is no stack trace. There is no 500 error. An agent simply answers with quiet confidence — a hallucinated fact, a misused tool, a subtle behavior drift introduced by a model upgrade — and keeps going. On September 17, 2026, San Francisco startup Raindrop announced a $50 million total funding milestone and a new product built specifically for this problem, and the round reads like a who’s-who endorsement of the “agent reliability” thesis.

What was announced

Raindrop closed a Series A led by CRV, bringing its total funding to $50 million. The round size itself was not disclosed, but the cap table was: Lightspeed Venture Partners and Y Combinator — both existing investors from the December 2025 seed round — participated, alongside Figma Ventures and Vercel Ventures. Most notably, lead researchers from OpenAI, Anthropic, and Thinking Machines invested personally. When the people building frontier models put their own money into a failure-detection startup, that is a signal worth reading.

Alongside the funding, Raindrop launched Simulations, now in research preview (early access is broadening, with general availability promised “over the coming month”). The company also confirmed a customer base that includes Vercel, Framer, Clay, and unnamed Fortune 100 enterprises across healthcare and logistics.

The product: see what your change will change

Raindrop’s existing platform is a production monitor for agents. It reads live agent trajectories — every run, every tool call, every decision — and flags what the company calls semantic anomalies: hallucinated answers, improper tool usage, agents stuck in loops, and behavior changes that arrive when a team upgrades the underlying model. Issues surface in traces before a user ever complains, and engineering teams can see what changed, when it started, and which users were affected.

Simulations extends that idea upstream to the moment before a change ships. The product runs on every pull request and answers a question traditional evaluation cannot: what will this change actually change?

The mechanism is deceptively simple to describe. Simulations replay real production traffic and a team’s existing test cases against a proposed change to the agent harness, then apply Raindrop’s anomaly detection to the results. Teams measure performance against their own anticipated test cases while simultaneously detecting the unanticipated behavior changes — the ones nobody wrote a test for.

The hard part, as the company explains it, is doing this in a harness-agnostic way. You cannot simply replay network calls or cached tool responses from an old trace. If a developer adds a brand-new tool, nothing exists for that tool in any historical trace. If a team migrates the entire harness to a new runtime, the old replays are useless. Raindrop’s answer is to actually simulate the world around the agent. The team’s mental model is Swiss cheese: every recorded tool call is a hole that reveals the actual state of the underlying world, and those holes become the raw material for reconstructing plausible environments to test against.

Why now: agents got long, and failures got expensive

Two numbers explain the urgency. Research from METR shows that the length of tasks frontier AI agents can complete autonomously with 50% reliability has roughly doubled every seven months since 2019. And in production today, a single agent run can last for days and involve thousands of tool calls.

“Agents now run for hours, call thousands of tools, and handle real money, real health data, and real customers,” said Zubin Koticha, Raindrop’s CEO. “When an agent fails, it does the wrong thing convincingly at scale until someone happens to notice.”

Longer runs multiply the blast radius of any single bad decision. An agent that loops, hallucinates, or misuses a tool at step 40 of a 4,000-step run does not just produce a wrong answer — it may have taken real actions in real systems along the way. Lightspeed partner Bucky Moore, who led the seed round, framed the stakes bluntly: as agents deploy into defense, healthcare, and financial services, bad behavior stops being an inconvenience and becomes a liability.

There is also a competitive-intelligence angle that Raindrop leans into openly. The company claims Simulations marks the first time companies outside the frontier labs get access to the same training and testing process those labs use internally. OpenAI recently published research on deployment simulation — regenerating responses to de-identified production conversations with a candidate model to predict misbehavior rates before release. Anthropic builds synthetic universes to train and stress-test its agents. Raindrop is effectively productizing that methodology for everyone else.

The founding story and the team

Raindrop was founded by Zubin Koticha (CEO), Ben Hylak (CTO), and Alexis Gauba (COO). Koticha and Gauba previously co-founded the cryptography startup Apache (secure computation, ZEXE), and Hylak built autonomous intelligent agents at Apple before a stint at an ML observability company. The nine-person San Francisco team carries unusually specific scars for the problem: engineers who built fraud transformer models at Robinhood, pioneered Pinterest’s recommender systems, did malicious anomaly detection at Square, and security engineers from Segment, Semgrep, and Socket.dev.

That background shows in how Raindrop frames the category. “Agents are fundamentally different from traditional software. They are highly capable, autonomous, and non-deterministic,” said Reid Christian, general partner at CRV. “Raindrop treats agent failure as a detection problem, the way a security company would.” It is the intrusion-detection playbook applied to your own software — assume failure is inevitable and non-obvious, and build the sensor grid to catch it.

Customer evidence suggests the framing lands. “If we’re having an issue like a build failure or agents stuck in a loop, we see that issue in Slack,” said Bani Singh, an AI engineer at Vercel.

A crowded lane with a distinctive bet

Raindrop is not alone in smelling opportunity. groundcover raised $100 million in July for AI-era observability. Scaled Cognition took $100 million from Khosla Ventures in June for reliable agents. Harvey acquired Guardrails AI earlier this month. The differentiator is where each pitch intervenes: design-time guardrails, runtime monitoring, or pre-deployment testing.

Raindrop’s bet is twofold. First, that agent reliability is fundamentally a detection problem rather than a design one — you cannot specify your way out of non-determinism. Second, that the highest-leverage intervention point is the pull request: the moment when a human still has hands on the keyboard and a bad change costs nothing to revert. If agents are going to run for days on end, the industry’s last cheap checkpoint is the PR — and Simulations wants to be the crash test that runs there.

The bet is not settled. Simulations sits in research preview, built hand-in-hand with Fortune 100 partners, and general availability is a month out. But with agent task horizons doubling every seven months, the window between “agent does something weird” and “agent does something catastrophic” is closing fast — and the market for catching it is heating up accordingly.

Raindrop says it is hiring across the board, with a focus on go-to-market and ML engineering.