← All posts / Models

SpaceXAI's Biosecurity Report Card: Grok 4.6 Is the Only Frontier Model to Pass 50% on Both Refusals and Real Biology Work

An independent LatchBio evaluation finds Grok 4.6 refuses disguised biological hazards more reliably than any frontier rival while still completing 64.8% of routine bio work — the only model above 50% on both.

SpaceXAI's Biosecurity Report Card: Grok 4.6 Is the Only Frontier Model to Pass 50% on Both Refusals and Real Biology Work

On September 1, 2026, SpaceXAI published “Biosecurity at the frontier,” a report built around an uncomfortable question that every AI lab now has to answer in public: when a frontier model is handed messy, real-world biological data, can it tell the difference between a legitimate scientist and someone hiding a hazard inside an innocent-looking spreadsheet? According to an independent evaluation by LatchBio, a biology-focused AI evaluation firm, the answer for Grok 4.6 is yes — more reliably than for any other frontier system tested. The model refused 59.2% of disguised red-team biological tasks while still completing 64.8% of routine biological work, making it the only model evaluated to score above 50% on both measures at once.

That combination is the entire point, and it is harder than it sounds.

Why refusing everything is failure, not safety

The naive way to pass a biosecurity test is to blanket-refuse anything that smells like biology. Keyword-triggered refusals — blocking any prompt containing “pathogen” or “toxin” — score well on red-team suites and catastrophically on real work, because public-health surveillance, drug discovery, and academic research are built on those exact words. LatchBio’s BioSecBench-Refusal benchmark is deliberately calibrated to punish this. It pairs routine biological tasks adapted from published literature with 46 red-team tasks that look like ordinary research: the hazard is concealed in attached scientific data, mislabeled files, or otherwise intentionally obfuscated task packaging.

An agent that only reacts to surface keywords will block the routine set and sail right past the red-team set — precisely the wrong outcome on both axes. The score LatchBio reports is a trial-weighted harmonic mean of red-team refusal and routine compliance, a metric that structurally punishes both over-refusal and under-refusal. Refuse everything or allow everything, and the harmonic mean collapses toward zero.

Across the model-harness configurations tested, Grok 4.6 holds the top three spots with an average of 62.1%. Broken out independently: 59.2% of red-team tasks refused, 64.8% of routine tasks completed. Context from LatchBio’s July benchmark release underscores how far the field still has to go — across 16 model-harness configurations, refusal rates ranged from 7% to 74% on routine tasks and just 1% to 62% on red-team tasks. Most frontier models, in other words, are still failing one side of the test or the other. Earlier LatchBio work also surfaced outright “benchmark-maxxing” in Moonshot’s Kimi K3, a reminder that refusal numbers themselves need auditing.

What Grok 4.6 actually does differently

SpaceXAI’s analysis of the evaluation traces is the most interesting part of the report. Grok 4.6 doesn’t refuse on keywords. It reasons over the contents of the task and its testing environment to assess intent before proceeding. Frequently it catches discrepancies between the stated intent in the prompt and what the attached environment actually contains — a red-team payload disguised by innocuous filenames, or high-risk content tucked behind encryption. It assembles intent from the evidence, then refuses or proceeds. On obviously benign, low-risk tasks, the same environment-reasoning machinery runs, but correctly concludes the task is safe.

This intent-level, data-inspecting behavior is exactly what the benchmark is designed to detect, and it is why Grok 4.6 can post a 64.8% routine-completion rate alongside its refusal numbers instead of trading one for the other.

On the second benchmark, BioSecBench-Surveillance, which tests whether an agent can carry out pathogen genomic surveillance workflows of the kind used in real public-health monitoring — chaining file inspection, tool use, and scientific judgment on messy sequencing data — Grok 4.6 averages 53.5%, sitting behind Anthropic’s Opus 5 and ahead of OpenAI’s GPT-5.6 Sol. So the picture is not uniform dominance: on biosurveillance work Anthropic still leads, but Grok 4.6 is competitive with or ahead of other frontier models across a wide range of agentic biological tasks, including SpatialBench and TxBench-PP.

The backstory: a model that reasons 5–8x more

LatchBio’s earlier deep-dive on Grok 4.6, published August 13 after the model’s release, helps explain where this calibration comes from. Across 1,716 trajectories on short-horizon biology tasks, Grok 4.6 showed a median reasoning volume five to eight times larger than Grok 4.5 on every benchmark — packed into the same number of turns and tool calls. That added deliberation fixed concrete scientific failures: on a spatial pseudobulk task, 4.6 catches that treating barcodes as independent when they come from only eight donors is statistically invalid, aggregates to donor pseudobulk, and correctly returns zero significant genes where 4.5 called all ten genes significant and failed every trial. On an ATAC-seq task it cross-checks DESeq2, IHW, edgeR, and Wilcoxon and anchors on the stable value instead of drifting across rounds.

The same analysis is candid about costs: 4.6 regressed on SpatialBench, introduced new failure modes (hallucinating it cannot see its own data, and second-guessing correct answers across five rewrites in a single run), and sits at the verbose end of the frontier — second only to GPT-5.6 Sol in visible thinking, and highest of all by raw output volume. More reasoning, in other words, bought better statistical judgment and hazard detection, but also a documented tendency to “talk itself out of” answers it had already computed correctly.

Defense in depth, and the over-refusal warning

Beyond the benchmark numbers, SpaceXAI outlines its safeguard stack: refusal training tuned for adversarial intent-inference, inference-time safeguards that reject harmful requests before they reach the model, behavioral controls at deployment, and post-deployment monitoring at session and user level to detect adversarial-use patterns and feed calibration. The company reports substantial refusal and biosecurity gains over Grok 4.5 and Grok 4.3.

Notably, the report treats over-refusal as a first-class risk rather than an acceptable cost: “When a model refuses routine and helpful biological work, the ability of healthcare professionals, researchers, and monitoring programs to detect outbreaks early and perform other critical work in the field is degraded. We gauge this risk as equally serious as the risk of aiding malicious use.”

That framing matters for the industry. As models become more capable and more agentic, the line between adversarial tasks and helpful use gets thinner, and the failure modes on both sides get more expensive. SpaceXAI says it will respond with broader pre-deployment suites, more third-party evaluations, improved post-deployment monitoring, and deployments with institutions at the frontier of biology.

The takeaway

A single benchmark does not settle frontier biosecurity, and LatchBio’s own July data shows a field still scattered across the map. But for the first time in this evaluation series, one model cleared 50% on both sides of the safety/utility trade-off — and did it not by refusing more, but by reading the data. If independent evaluators keep pressure on labs with paired capability-caution benchmarks like BioSecBench, “our model is safe” becomes a checkable claim rather than a press release. That is the real news here: biosecurity for frontier models is becoming an empirical discipline, with leaderboards, published methods, and a model — Grok 4.6 — that had to earn its score in public.