← All posts / Industry

Meta Kills Token-Count Performance Reviews Just as It Hands Employees the Hatch Agent

Meta walks back 'AI-driven impact' metrics after a lawsuit from workers on medical leave, while pushing its new autonomous Hatch agent to employees who aren't sure they trust it.

Meta Kills Token-Count Performance Reviews Just as It Hands Employees the Hatch Agent

For roughly a year, one number quietly shaped how thousands of Meta engineers were judged: how much AI they used. Internal dashboards tracked “AI adoption,” leaderboards ranked token consumption, and workers were sorted into labels like “AI Native,” “AI First,” and “AI Enabled.” This week, Meta told employees that era is over — performance reviews “will not use AI adoption dashboards or token counts to evaluate impact.”

The reversal, first reported by The Information and confirmed by Wired, lands at a strange moment. At the same time Meta is retiring the metrics, it is putting its most autonomous internal AI tool yet — an agent called Hatch — into employees’ hands. One hand takes away the measuring stick; the other hands over the machine.

What actually changed

The updated review guidance removes all references to AI usage as a measure of individual impact. The old framing, built around “AI-driven impact,” is replaced with language saying outcomes “can be supported by AI or other means.” Engineers across the company were told directly that AI-adoption dashboards and token counts will play no role in evaluating their work.

Meta spokesperson Tracy Clayton said the company evaluates employees on their contributions and that the “AI Native” labels were never used for performance evaluation — a claim several current employees dispute, given that the labels appeared in internal review materials for the past year.

The change didn’t come from nowhere. The company had already removed an internal token-use leaderboard in April after employees gamed it, and later had to ration employee AI access outright when usage ballooned. The leaderboard’s brief life illustrated the core problem: when you measure tokens, you get tokens. Employees reported colleagues “prompt-maxxing” — running AI tools continuously, on trivial tasks, purely to climb the rankings.

The lawsuit that shadowed the metrics

The timing is hard to separate from litigation. In July, 26 current and former Meta employees sued the company in federal court in Oakland, alleging that AI systems used in a May reduction-in-force — which cut roughly 8,000 jobs — disproportionately selected workers who had taken medical or parental leave or requested disability accommodations.

One plaintiff alleges her AI-adoption score dropped after leave because the system treated her time away as a performance gap rather than excluding it from the calculation. In other words, the “neutral” metric encoded a penalty for taking legally protected time off. Legal analysts have since held the case up as a textbook example of how proxy metrics — numbers that look objective — can launder bias into employment decisions.

Meta maintains that humans, not AI, made the layoff decisions. But with discovery looming, keeping an AI-usage score anywhere near the performance-review pipeline became a liability the company could no longer afford. Walking the metrics back costs little; defending them in court could cost much more.

Hatch asks for a different kind of trust

Meanwhile, Hatch — the agentic tool Meta has been testing internally for weeks on corporate devices — is being pushed across the company. Unlike a chatbot, Hatch can browse the web and control other applications: booking appointments, organizing calendars, acting on a user’s behalf across services. A public release is expected but has not been announced.

Early testers report a paradox the review-policy change only sharpens. Token usage keeps climbing anyway, because agents are hungry by design. And employees are hesitant to connect Hatch to personal email, calendars, and other accounts — not because of policy, but because an autonomous agent with write access to your digital life can make consequential mistakes.

That hesitation has recent precedent. Meta previously paused a project that used work devices to collect AI-training data after employees objected, an episode that eroded trust in the company’s AI tools generally. Some employees now use Hatch only for low-stakes personal logistics — booking appointments, organizing their week — keeping it at arm’s length from anything that matters.

There is also a quieter fear circulating in internal discussions: that productivity gains from Hatch could eventually justify smaller teams. The tool that Meta wants employees to embrace is, in the worst case, the same tool that makes some of them redundant. Removing token-count reviews does nothing to address that anxiety — if anything, it removes the one visible metric employees could point to as proof their AI usage “counted.”

The bigger picture: from adoption metrics to outcomes

Meta’s retreat is the most visible sign yet that the industry’s first phase of internal AI measurement is ending. In 2024 and 2025, companies across tech rushed to quantify AI adoption — dashboards, training quotas, usage scores — on the theory that what gets measured gets adopted. Two years in, the results are in: measured adoption produces theater. Employees run models on autopilot, leaderboards reward volume over value, and the metrics themselves become legal exposure when they interact with protected leave.

The stated philosophy now is outcomes-only: did the work get done, and was it good? AI is a means, not a grade. That is cleaner, fairer to employees on leave, and harder to game. It also shifts the burden to managers, who must now judge output quality without a proxy number to lean on — precisely the skill most organizations have spent the last two years atrophying in favor of dashboards.

For Meta specifically, the calculation is also competitive. Hatch only succeeds if employees trust it with real work and real accounts. Mandated adoption — or even the perception that usage is being scored — produces checkbox behavior, not the deep feedback an agent platform needs to improve. Voluntary use by genuinely motivated employees generates better training signal than any leaderboard ever did.

What to watch

Three things will tell us whether this reversal is substance or optics. First, whether the “AI Native” labels and dashboards actually disappear from internal tooling, or merely from review language. Second, how the Oakland lawsuit proceeds — if discovery surfaces documents tying adoption scores to layoff rankings, the walk-back will look prescient rather than principled. Third, Hatch’s public launch: if it ships with personal-account integration and employees opt in at scale, trust will have won. If it launches to the same hesitation reported internally, Meta will have learned that the scarcest resource in enterprise AI isn’t tokens — it’s belief.

The industry lesson generalizes beyond Meta. Any company scoring employees on AI usage today is building the same lawsuit and the same theater tomorrow. Measure outcomes. Let agents earn their tokens.