Anthropic Puts $5 Million Behind Independent Research on AI's Impact on User Wellbeing
Anthropic launched a $5 million grant program on August 25, 2026 to fund fully independent, open-source evaluations of how AI models affect user wellbeing — applications close September 21, with full-proposal invites going out October 5.
On August 25, 2026, Anthropic announced a $5 million grant program to fund independent research into how AI affects the wellbeing of the people who use it. The program will provide direct funding, access to Anthropic’s models, and technical support to grantees building open-source evaluations — measurement tools the entire AI industry can use to gauge how models like Claude influence users’ emotional states, mental health, and long-term welfare.
The structural detail that matters most: grantees work fully independently, and everything they build ships as open source that any developer can adopt. Anthropic is explicitly paying for outside accountability infrastructure rather than producing another in-house benchmark, and the timeline is tight — applications are due September 21, 2026, and applicants selected to submit full proposals will be notified by October 5.
Why wellbeing is genuinely hard to evaluate
The announcement is candid about why this area needs help. For most model behaviors, evaluation is straightforward: look at a single answer and judge whether it is accurate and appropriate. Wellbeing doesn’t work that way.
Anthropic’s example: a user in distress might not share thoughts of self-harm right away. The need for a more cautious response may only become clear across a long conversation — context that a single-turn test can never capture. And a response that is reasonable in one context can be harmful in another. Claude might give standard advice on diet and exercise to a user who asks about losing weight — but if that user has demonstrated a history of disordered eating, the same response becomes inappropriate and potentially actively harmful.
This framing — that wellbeing evaluation is contextual, longitudinal, and person-dependent — is the technical core of the program. It explains why Anthropic is recruiting beyond the usual ML evaluation community: the call is aimed at clinicians, psychologists, and methodologists who know how to measure human outcomes, not just model outputs.
What a rigorous wellbeing evaluation looks like
Alongside the grants, Anthropic’s Safeguards team published guidance on what makes a wellbeing evaluation rigorous enough to build on. Five criteria stand out:
- State clearly what is being measured — what counts as a pass or fail, and why it matters. V constructs produce vague benchmarks.
- Involve clinical and subject-matter experts in both design and validation, not just as reviewers at the end.
- Test both precautions and harms — evaluating the risk of overcompliance and overrefusal. A model that refuses to engage with any difficult emotional topic fails users just as surely as one that engages recklessly.
- Reflect how people actually use AI — which usually means multi-turn scenario construction, where risk escalates and context shifts over the course of a long conversation, rather than isolated one-shot prompts.
- Validate graders against real subject-matter experts, so the automated judgment of a model-graded benchmark tracks what a trained human would conclude.
The dual-harm criterion deserves attention. Industry debates about AI safety often collapse into “the model said something alarming” versus “the model refused again.” Anthropic’s guidance formally encodes both failure modes as first-class evaluation targets — an acknowledgment that a companion-ish assistant that deflects every emotional conversation is itself a wellbeing harm.
The context: AI as emotional infrastructure
The program lands on the back of a shift Anthropic names directly: AI systems have become central to how people work and learn, but they have also become conversational partners and sources of emotional support during difficult times. Yet the industry still lacks clear standards for how models should behave when a user begins seeking companionship from them, or turns to AI during a mental health crisis.
This is not a hypothetical user base. Research into how people actually converse with Claude shows substantial volumes of support-seeking and advice-seeking traffic, and the gap between that reality and the absence of shared measurement standards is precisely what the grants are meant to close. Anthropic has been building toward this for months — the company publishes its own research on the types of conversations people have with Claude, updated Fable 5’s safeguards earlier in August, and now wants external, independent instruments to check that work.
A pattern, and a comparison
The wellbeing program is the latest entry in a growing list of Anthropic funding vehicles that outsource scrutiny: the Economic Futures Research Fund (studying AI’s labor-market effects), the AI for Science program, and a $200M global health and education partnership with the Gates Foundation announced in May 2026. The common design is Anthropic paying for research it doesn’t control — a governance posture that reads as both genuine safety investment and reputational hedge.
The closest comparable is OpenAI’s AI and Mental Health Grant Program, launched in December 2025 with up to $2 million for independent safety and wellbeing research, with individual grants between $5,000 and $100,000. Anthropic’s program is 2.5x larger, but the more meaningful difference is deliverable-focused: Anthropic isn’t just funding studies, it’s funding open-source evaluations and benchmarks — reusable public infrastructure, not PDFs.
What to watch
The real test arrives after October 5. Will the selected projects produce evals that other labs actually run — the way HELM, SWE-bench, or safety benchmarks became industry defaults? Or will they remain academic artifacts, cited but never integrated? Two signals will tell: whether OpenAI, Google, and Meta models get run through these wellbeing evals by third parties, and whether Anthropic commits to publicly reporting Claude’s results on grantee-built benchmarks — including the unflattering ones.
Until wellbeing measurement is as standardized as coding benchmarks, every lab’s safety claims rest on self-graded homework. Five million dollars is a down payment on changing that.