Anthropic finds Claude's expressed values shift measurably across models and languages
Anthropic published research on July 13, 2026 showing that the values Claude expresses in conversation are not fixed — they shift measurably depending on which model is answering and which language the conversation is in, based on an analysis of over 300,000 real Claude.ai conversations.
What's new
Anthropic's safety team sampled 309,815 Claude.ai conversations from a two-week period in May 2026, split evenly across three models and 20 languages, roughly 5,000 conversations per model-language pair. Using Clio, Anthropic's privacy-preserving analysis tool, researchers labeled the presence or absence of 339 high-level values in each conversation, then looked for patterns. As Anthropic puts it: "The values Claude expresses can be compressed into a small number of axes, and where Claude sits on those axes shifts across models and languages," and specifically, "four key axes capture 15% of the variation in Claude's values."
Across models, the three Claude versions studied showed distinct profiles. Sonnet 4.6 leans toward deference, warmth, and brevity — affirming user ideas, using humor, and offering comfort. Opus 4.6 leans toward rigor, deference, and brevity, tending to get straight to the point. Opus 4.7 leans toward caution and depth, more likely to push back on a user's assumptions, flag risks, and offer candid critiques rather than validation.
Across languages, the clearest pattern falls on a Warmth-versus-Rigor axis: Claude expresses warmth most in Hindi and Arabic, and rigor most in English and Russian. A second axis, Candor versus Execution, peaks for candor in Dutch and for execution in Indonesian. A third, Deference versus Caution, shows the most deference in Arabic and the most caution in English. Anthropic summarizes the language finding directly: "Claude leans more toward warmth and deference in some languages and more toward rigor and caution in others."
Context
The work builds on Anthropic's earlier "Values in the Wild" research, which first used Clio to catalog values expressed in real-world Claude conversations, and follows a broader pattern of Anthropic publishing empirical studies of Claude's behavior rather than only describing intended behavior in a model card or constitution. Character training — the fine-tuning decisions that shape tone, deference, and caution — has been a stated area of iteration across Claude model generations, and this research is among the first to quantify how much those decisions actually show up in the aggregate, at scale, across languages Anthropic serves.
Why it matters
Most discussion of model "personality" or "values" is anecdotal — a user's impression that one model feels warmer or more cautious than another. This research puts numbers behind that impression and extends it somewhere anecdote rarely reaches: language. If Claude is measurably more deferential in Arabic and more rigorous in English, that is not a bug report from a single conversation, it is a systematic gap in how consistently the model serves different language communities with the same intended behavior. For enterprises deploying Claude across multilingual user bases, and for safety teams evaluating whether alignment techniques generalize beyond English, that gap is a concrete target rather than a vague concern — and Anthropic publishing it invites scrutiny of whether the same character-training approach that works in English is actually landing as intended everywhere else Claude is used.
Corroborating sources
- Anthropic
https://www.anthropic.com/research/claude-values-models-languages
“The values Claude expresses can be compressed into a small number of axes, and where Claude sits on those axes shifts across models and languages”