One Agentic Task, 10,000x the Footprint: Vals AI Puts Hard Numbers on AI's Energy Bill
Independent benchmarker Vals AI measured electricity, carbon, and water across 16 open-weight models and found long 'thinking' agentic tasks carry up to 10,000x the environmental impact of a simple query.
For years, the debate about AI’s environmental cost has run on vibes: data centers are big, GPUs are thirsty, therefore AI is bad for the planet. On September 3, independent benchmarking firm Vals AI replaced the vibes with data. In its first Environmental Impacts Report, the firm measured electricity consumption, carbon emissions, and water usage across 16 open-weight models running realistic agentic workloads — and the headline number is stark: a long, multi-step “thinking” task can carry roughly 10,000 times the environmental impact of a simple one-shot question.
What Vals actually measured
Most sustainability claims about AI focus on training — the one-time energy spike of building a frontier model. Vals looked at the other side of the ledger: inference at scale, doing real work. The team took token-usage data from its own Vals Index benchmark, a suite of long-horizon, multi-turn, agentic tasks drawn from finance, coding, healthcare, and law, and fed it through EcoLogits, an open-source library that estimates carbon, electricity, and water consumption from token counts plus a model’s total and active parameter counts.
That architecture requirement is why the study covers only open-weight models — 14 of the 16 are Chinese, a reflection of where open weights live today. Closed labs simply don’t disclose the internals needed to run the numbers. To partially close that gap, Vals shipped an interactive calculator that estimates footprints under user-specified assumptions. Its default preset is sobering: a hypothetical 6-trillion-parameter model with 30% Mixture-of-Experts activation would burn through 21.2 kWh of electricity on a single Vals Index task — a full day of average US home power — while emitting 11.5 kgCO₂e (about 50 miles of driving) and consuming 91.9 liters of water (a 10-minute shower).
The 10,000x gap, in household terms
Why do agentic tasks blow past simple queries by four orders of magnitude? Because agents don’t answer — they iterate. A model asked to build a web app plans, writes files, runs tools, reads errors, and retries, generating tens of thousands of output tokens where a chat reply might produce a few hundred. Vals engineer Omar Almatov put it plainly: “Footprints scale dramatically because these are no longer just single-shot questions.”
The report’s equivalences make the scale tangible. On some models, building a single web app consumes energy comparable to powering a home for 2.5 hours. Extrapolated across the full 2,157-task Vals Index, running the entire benchmark on Moonshot AI’s Kimi K3 — the per-task heavyweight — would emit as much carbon as a flight from San Francisco to Miami, use enough water to fill a fire truck’s tank, and consume on the order of 1.7 MWh of electricity, which the report frames as roughly a month of average US home energy use.
The full-index league table spreads across two orders of magnitude. At the efficient end, Ling 3.0 Flash 2607 completes the whole index on an estimated 23.8 kWh — one day of home electricity. MiniMax M2.7 follows at 34.8 kWh, with Mistral Medium 3.5 at 47.0 kWh. At the other extreme sit DeepSeek V4 Pro (459.6 kWh), DeepSeek V4 Pro 0813 (869 kWh), Kimi K3 (1.7 MWh), and Qwen3.8 Max (1.8 MWh) — tens of times the footprint of the leaders for tasks of comparable ambition.
Capability is cheap; footprint is not
The most commercially loaded finding is the disconnect between price and impact. Kimi K3 tops the open-weight Vals leaderboard at 57% accuracy, but runner-up DeepSeek V4 Flash sits only 4% behind at a fraction of the resource cost. Vals frames the gap between the two as the difference between charging your laptop once versus twenty times, or drinking a glass of water versus flushing a toilet.
Pricing, it turns out, is a terrible proxy. DeepSeek V4 Pro is nearly 15x cheaper per output token than Kimi K3 — and roughly 6x cheaper per completed task — yet has approximately the same environmental footprint. Discounts, market-share wars, and token-efficiency differences decouple what you pay from what the planet pays. A model can look like a bargain on an API price list while quietly being among the heaviest consumers of electricity and water in the fleet.
The pattern holds even at consumer scale. Using data from its upcoming Child Safety Benchmark, Vals estimated a 10-turn personal conversation: Kimi K3 carries roughly 3x the footprint of GLM-5.2 while using half as many tokens — enough energy to charge 14 phones, carbon equal to driving 0.2 miles, and three glasses of water. Parameter count, not verbosity, drove the bill.
Why this lands now
The report arrives amid genuine political friction. Vals cites polling showing seven in 10 Americans oppose data center construction in their locality, and cities including Minneapolis, Denver, and Seattle have enacted temporary restrictions on data center builds while they study the environmental outcomes. In March, Senator Markey and Representative Beyer introduced the Artificial Intelligence Environmental Impacts Act of 2026, which would mandate environmental and energy reporting for data centers, and California already requires large companies to disclose greenhouse gas emissions and climate-related financial risk.
Regulators, in other words, are going to need exactly the kind of measurement Vals just published — and the first credible dataset suggests the industry’s environmental problem is not the existential hum of data centers in general, but a specific, quantifiable design choice: how much thinking we ask models to do, and on what hardware.
The caveats that matter
The report is candid about its limits. Estimates rely on assumed hardware (the calculator presets 256 H100s at batch size 64) rather than measured facility data; input, cache, and prefill compute are excluded from the headline numbers; and closed-weight models — the GPTs, Claudes, and Geminis serving most consumer traffic — are entirely absent because labs don’t disclose architecture. The absolute figures are best read as well-grounded estimates, not meter readings.
But the direction is unambiguous. As models scale toward ever-larger parameter counts and agentic workflows become the default interface for software, per-task footprints grow with them. The report’s closing thought is the one that should stick with every builder picking a model this week: if capability continues to scale with size, the environmental footprint scales too — and unlike accuracy, nobody’s putting that number on the pricing page yet.