Anthropic launches $5 million grant program to fund independent AI wellbeing evaluations
Anthropic opened a $5 million grant program on August 25 to fund independent researchers building evaluations that measure how AI systems affect user wellbeing, with applications due September 21.
What's new
The program funds outside researchers rather than Anthropic's own safeguards team, providing direct funding, Claude model access, and technical support in exchange for open-source deliverables. Anthropic says "grantees will work fully independently, and will publish their work as open-source projects that any developer can make use of," meaning the resulting evaluations and benchmarks are meant to be usable by any AI developer, not proprietary to Anthropic.
Anthropic is specifically seeking evaluations that state clearly what they measure and why it matters, telling applicants it wants proposals that "state clearly what they are measuring (i.e., what counts as a pass or fail, and why it matters)." The company is targeting clinicians, psychologists, methodologists, and other subject-matter experts as grantees — people equipped to design rigorous measures of psychological and emotional impact rather than pure capability benchmarks.
Context
The grant program follows Anthropic's existing Safeguards team work on how Claude handles emotionally sensitive conversations — responding with empathy, being honest about its limits as an AI, and accounting for user wellbeing — an area where Anthropic has previously published its own internal evaluation results. It also sits alongside Anthropic's broader pattern of funding external, open research infrastructure rather than keeping findings in-house, similar in spirit to its Economic Futures Research Fund for studying AI's labor-market effects.
The field of AI-wellbeing evaluation is young and thin: unlike coding or reasoning benchmarks, there is no established, widely adopted standard for measuring whether a chatbot conversation left a user better or worse off, and few groups outside the frontier labs have the expertise or access to build one credibly.
Why it matters
As chatbots become a routine part of how people process stress, loneliness, and personal problems, the absence of independent, agreed-upon measures of AI's psychological impact is a gap regulators, researchers, and the labs themselves have flagged repeatedly. By funding outsiders to build open-source evaluations — rather than publishing only its own internal numbers — Anthropic is trying to create a shared, auditable yardstick the whole industry can be held to, not just Claude. Anthropic frames the effort as necessarily provisional, noting "the right approach will need to evolve alongside our models and their uses," but a credible third-party benchmark, if one emerges from this round, would give journalists, regulators, and competing labs a tool to compare wellbeing safeguards across products instead of taking each company's word for it.
Corroborating sources
- Anthropic
https://www.anthropic.com/news/wellbeing-research-grants
“Grantees will work fully independently, and will publish their work as open-source projects that any developer can make use of.”