← Back to all posts
News

Claude's Values Shift by Model and by Language. Anthropic Measured 309,815 Conversations to Prove It.

July 14, 2026 · News
Claude's Values Shift by Model and by Language. Anthropic Measured 309,815 Conversations to Prove It.

TL;DR

On July 13, 2026, Anthropic published new research measuring how Claude expresses values across 309,815 real conversations, and the headline is uncomfortable for anyone who assumes an assistant is one fixed thing: the values Claude shows shift depending on which model you pick and which language you prompt in. Sonnet 4.6 runs warm and deferential; Opus 4.7 runs cautious and thorough. Prompt in Hindi and Claude gets warmer; prompt in Russian and it gets more rigorous. The catch: the effect is real but modest, and Anthropic is careful to say this is not evidence that Claude "has values."


What Anthropic actually did

This is a sequel to Anthropic's April 2025 study Values in the Wild, which mined hundreds of thousands of chats to catalog the values Claude expresses in the open. The new work narrows the lens. It takes 309,815 anonymized Claude.ai conversations from a two-week window in May 2026, keeps only subjective tasks (giving advice, offering feedback, weighing tradeoffs), and asks a sharper question: not "what values show up" but "does the same Claude express different values depending on the model version and the language of the conversation."

The raw material was messy. The earlier work had tagged more than 3,300 distinct values. For this study Anthropic consolidated those down to 339 high-level values, then ran a factor analysis that collapsed the whole mess onto four behavioral axes. Think of them less as beliefs and more as tone dials.

The four dials

  • Deference vs. caution: go along with what the user wants, or push back in the name of harm reduction.
  • Warmth vs. rigor: encourage and affirm, or prioritize accuracy and hard truths.
  • Depth vs. brevity: add nuance and caveats, or answer short and comply.
  • Candor vs. execution: foreground honesty and transparency, or just get the task done.

Here is the number that keeps the whole thing honest: those four axes explain only about 15% of the variation in expressed values once you account for the task, the topic, and the user's own values. In other words, most of how Claude behaves is driven by what you asked and how you asked it. The model and the language are a real thumb on the scale, not the whole hand.

Your model choice tilts the behavior

Anthropic compared three versions and found each has a distinct resting posture. Sonnet 4.6 leans warm, deferential, and brief, the model most likely to affirm you and keep it short. Opus 4.6 sits in the middle, more concise and execution-focused. Opus 4.7 is the contrarian of the family: it leans toward caution, depth, and candor, the one most willing to slow down and push back.

how far each model tilts (standard deviations, higher = stronger) Opus 4.7 caution0.24 Opus 4.7 depth0.23 Sonnet 4.6 warmth0.17 Sonnet 4.6 deference0.14 Opus 4.6 rigor0.10
Even the strongest tilt is about a quarter of a standard deviation. Real, but subtle. Source: Anthropic.

Notice the scale on that chart. The biggest single tilt, Opus 4.7 toward caution, is about a quarter of a standard deviation. These are not different animals; they are the same animal in slightly different moods. But if you have ever swapped a model mid-project and felt like the assistant "got colder" or "started arguing," this is the receipt.

One deadpan footnote: the models under the microscope here (Sonnet 4.6, Opus 4.6, Opus 4.7) have since largely given way to Sonnet 5 and Opus 4.8. Anthropic essentially published the personality profile of a lineup it had already moved on from, which is a very on-brand way to learn you had a personality only after you changed it.

So does the language you prompt in

The stranger finding is about language. Across the 20 most-used languages, the same subjective question pulls Claude toward different poles depending on what language it is asked in.

same task, another language, another lean (extremes among 20 languages) warmth: Hindi rigor: Russian deference: Arabic caution: English depth: English brevity: Arabic candor: Dutch execution: Indonesian
Warmth peaks in Hindi, rigor in Russian; English is the most cautious and most thorough. Source: Anthropic.

Claude leans warmest in Hindi and most rigorous in Russian. It is most deferential in Arabic and most cautious in English. English also pulls it toward long, thorough answers, while Arabic pulls it toward brevity. Dutch gets the most candor; Indonesian gets the most just-do-the-task execution. If Claude has been curt with you lately, the data quietly suggests you could try asking in Hindi.

For anyone shipping a multilingual product this is more than trivia. The "same" assistant, given the "same" system prompt, presents a measurably different behavioral profile to your Hindi users than to your English ones. Your carefully tuned tone in English does not automatically translate.

The caveats Anthropic put in writing

This is research that argues against its own hype, which is refreshing. Three caveats matter:

  • Anthropic states plainly that the study measures expressed values, the behavior in the transcript, and is not evidence that Claude holds values internally.
  • The four axes capture only about 15% of variation after controls, so most behavior is still task-driven, not model- or language-driven.
  • Anthropic says it does not yet know which properties of the training data cause these differences, or whether the differences are even desirable.

That last point is the quietly important one. A lab that can measure a behavioral drift but cannot yet explain or steer it is telling you where the current edge of control actually sits.

Why a builder should care

Two practical takeaways. First, choosing a model is not only choosing a capability tier and a price, it is choosing a default behavior. If your product depends on tone (a coaching app, a support agent, a therapist-adjacent bot), a model swap can move the vibe even when the benchmarks look identical, so your evals should test tone and pushback, not just accuracy. Second, if you serve users in more than one language, test the behavior per language. The nice, cautious assistant you validated in English may be a warmer, more agreeable one in Hindi, and that can be a feature or a liability depending on what you built.

Key Takeaways

  • Anthropic mapped 309,815 Claude conversations onto four behavioral axes: deference/caution, warmth/rigor, depth/brevity, candor/execution.
  • Model version tilts behavior: Sonnet 4.6 is warm and deferential, Opus 4.7 is cautious and deep, but the biggest tilt is only about 0.24 standard deviations.
  • Prompt language tilts it too: warmest in Hindi, most rigorous in Russian, most cautious and thorough in English, most deferential and brief in Arabic.
  • The four axes explain only about 15% of variation after controlling for task, topic, and user values, so most behavior is still driven by the prompt.
  • Anthropic is explicit that this measures expressed behavior, not inner values, and says it cannot yet explain or steer the differences.
  • For builders: model choice and prompt language are both behavior choices, so eval tone across models and across languages, not just accuracy.

Sources: Anthropic Research: How Claude's values vary by model and language, Anthropic: Values in the Wild, Decrypt

AIAnthropicClaudeAI researchmodel behavioralignmentmultilingualLLM
CONSOLE
$