← All posts / Tools

Test the Agent Before It Breaks: Raindrop Raises $50M and Ships Simulations

Raindrop's CRV-led Series A brings total funding to $50M. Its new Simulations product replays real production traffic against every pull request, applying anomaly detection to catch agent failures before they ship.

Test the Agent Before It Breaks: Raindrop Raises $50M and Ships Simulations

AI agents have quietly crossed a line. They no longer answer a question and stop — they run for hours, chain together thousands of tool calls, and touch real money, real health data, and real customers. When software of that kind fails, it does not crash with a stack trace. It does the wrong thing convincingly, at scale, until somebody happens to notice.

That gap between “broken” and “silently wrong” is the entire business of Raindrop, a San Francisco startup founded by Zubin Koticha, Ben Hylak, and Alexis Gauba. On September 17, 2026, the company announced a Series A led by CRV, bringing its total funding to $50 million. Alongside the round, it shipped a new product called Simulations — and the combination says a lot about where the AI infrastructure market is heading.

What Raindrop actually does

Raindrop’s core product is production monitoring for AI agents. It reads live agent traffic — every run, every tool call, every decision — and flags what it calls semantic anomalies: hallucinated answers, tool misuse, and the subtle behavior changes that arrive the moment someone swaps in a new model or upgrades a harness. When something goes wrong, engineering teams can see what changed, when it started, and which users were hit.

This is a deliberately different bet from most of the “agent reliability” space. As CRV general partner Reid Christian put it, “Raindrop treats agent failure as a detection problem, the way a security company would.” You do not try to prove the agent can never misbehave. You instrument it, watch it, and catch the misbehavior fast — the same posture security teams learned to take with software they could never fully verify.

The approach has attracted real customers: Vercel, Framer, and Clay are named users, alongside unnamed Fortune 100 enterprises in healthcare and logistics. “If we’re having an issue like a build failure or agents stuck in a loop, we see that issue in Slack,” said Bani Singh, an AI engineer at Vercel.

Simulations: move the checkpoint to the pull request

Monitoring catches failures after they reach production. Simulations, launched this week in research preview, tries to stop them before they ship. The mechanics are straightforward to describe: it replays real production traffic and a team’s existing test cases against a proposed change to an agent harness, then runs Raindrop’s anomaly detection over the results. It runs on every pull request, so an engineer can finally answer the question the company likes to pose — “what will my change change?”

The pitch is a direct swipe at conventional evaluation. Traditional eval suites depend on test cases written in advance, which means they mostly catch the failures a team already anticipated. The failure modes that actually burn companies in production are the ones nobody wrote a test for. By replaying live traffic rather than a curated benchmark, Simulations is designed to surface the unanticipated behavior changes — the regression that only shows up on the weird requests real users actually make.

The technically hard part, the company says, is doing this in a harness-agnostic way. You cannot simply replay cached tool responses from an old trace. If a developer adds a brand-new tool, nothing exists for it in any historical trace. If a team migrates the whole harness to a new runtime, the recorded world no longer matches. Raindrop’s answer is to simulate the environment around the agent, using each recorded tool call as a small window — “a hole in the Swiss cheese” — into the actual state of the underlying world at that moment.

The frontier-lab connection

Two details make this more than another observability startup story.

First, the cap table. Lightspeed Venture Partners and Y Combinator, already investors from Raindrop’s $15M seed in December 2025, joined the Series A — and so did lead researchers from OpenAI, Anthropic, and Thinking Machines, investing personally. When the people building frontier agents put their own money into a failure-detection company, that is a signal about how worried they are.

Second, the pedigree of the product itself. Raindrop explicitly frames Simulations as the first time companies outside the frontier labs get access to the training-and-testing process those labs use internally. OpenAI recently published research on deployment simulation — regenerating responses to de-identified production conversations with a candidate model to predict misbehavior rates before release. Anthropic builds synthetic universes to train and stress-test its agents. Raindrop is betting those techniques can be productized for everyone else, and it has been building Simulations hand in hand with its Fortune 100 partners before broadening early access this week.

Why now: the METR curve

The timing rests on an uncomfortable trend. Raindrop cites research from METR showing that the length of tasks agents can complete autonomously roughly doubles every seven months. A single agent run can now last days and involve thousands of tool calls. Every doubling makes hand-written test suites proportionally weaker and the blast radius of a silent failure proportionally larger.

The funding environment agrees that this is a category worth backing. groundcover raised $100M in July for AI-era observability; Scaled Cognition took $100M from Khosla in June for reliable agents; Harvey acquired Guardrails AI earlier this month. What differentiates the pitches is where they intervene — and Raindrop is placing its chip on two squares: detection rather than design, and the pull request rather than the incident review. Lightspeed partner Bucky Moore’s warning about where this ends up — bad agent behavior becoming catastrophic in high-stakes settings like defense — is the bear case that justifies the spend.

Simulations remains in preview, so the bet is not settled. But the direction is clear: as agents take on longer, riskier, less supervised work, the industry’s center of gravity is shifting from “make the model smarter” to “make the deployment provably survivable.” Raindrop’s $50M is the latest, and one of the more credible, wagers that the second problem is the bigger business.