Same Question, Different Answers: What Anthropic's New Study Reveals About Claude
Prompt engineers may soon have another tool to add to their arsenal: deliberately choosing the language of a prompt. A prompt can already be translated automatically and submitted in multiple conversations—not merely to obtain the same answer in different words, but to explore different perspectives and behavioral tendencies of the model itself.
Claude, it turns out, is not quite the same "character" every time you talk to it.
On July 13, 2026, Anthropic published a large-scale study showing that Claude's tone and priorities vary not only across model versions but, perhaps more surprisingly, across languages. Sometimes it becomes more cautious, sometimes warmer. In some languages it is more talkative and eager to explain; in others it is more direct and more candid about its own limitations.
Anthropic analyzed more than 300,000 conversations on Claude.ai in which users asked the model for subjective judgments. The dataset was divided across Sonnet 4.6, Opus 4.6, and Opus 4.7, as well as the platform's twenty most widely used languages.
Researchers evaluated responses along four behavioral dimensions: deference versus caution, warmth versus rigor, depth versus brevity, and honesty about limitations versus execution-oriented behavior.
Sonnet 4.6 emerged as the most deferential and warm. It tends to validate users' ideas, uses humor more readily, and offers reassurance without being judgmental. Opus 4.6 remains similarly concise but adopts a more rigorous style, getting straight to the point. Opus 4.7 stands out most clearly from the other models, leaning strongly toward caution and analytical depth. It is more willing to challenge assumptions, point out risks without being prompted, and openly acknowledge its own mistakes and limitations.
Anthropic notes that these behavioral profiles closely match public perceptions of each model, giving researchers confidence that the methodology captures genuine differences rather than statistical noise.
The largest variation appears along the warmth-versus-rigor dimension.
Claude is at its warmest when conversing in Hindi and Arabic, where it tends to use more polite language, humor, and positive reinforcement. At the opposite end of the spectrum, it is noticeably more rigorous in English and Russian, where it is more likely to challenge assumptions and request evidence.
Language also affects other behavioral dimensions. Claude tends to provide more detailed and nuanced answers in English, while favoring brevity in Arabic. Along the honesty-versus-execution axis, Dutch is the language in which Claude most readily acknowledges its own mistakes, whereas Indonesian pushes the model toward a pragmatic style focused primarily on completing the user's request.
The practical consequences can be surprisingly significant.
If two people ask Claude to evaluate the exact same business plan—one writing in Hindi and the other in Russian—they may walk away with very different impressions of the quality of their proposal. Not because the plan itself changed, but because Claude framed its evaluation through subtly different value judgments.
This is where prompt engineering enters the picture.
Language choice becomes another controllable variable alongside the model's assigned role, the context provided, and the requested output format. The technique is not a guaranteed recipe, however. The study describes statistical tendencies observed across hundreds of thousands of conversations; it does not imply that every Russian prompt will receive a more rigorous response or every Dutch conversation a more candid one. Translation itself can also alter subtle nuances of a request.
Romanian was among the twenty languages included in the study. Researchers analyzed 15,549 Romanian conversations, finding that Claude's behavior remained remarkably close to the overall average. The model showed only slight tendencies toward deference, rigor, depth, and execution-oriented responses, with deviations of just 0.02 to 0.03 standard deviations.
Even so, researchers identified three behaviors that appeared more frequently in Romanian conversations: Claude was more likely to flag risks without being asked, leave the final decision explicitly to the user, and recommend consulting qualified professionals.
Anthropic does not offer a definitive explanation for these differences, proposing instead two main hypotheses.
The first concerns the sheer volume of training data. Some languages are represented far more extensively than others, and where training data is relatively scarce, aligning models toward consistent behavior may simply be more difficult.
The second hypothesis concerns the composition of the data itself. Certain languages may be disproportionately represented by particular types of writing—for example, professional or technical documents—that reflect values and communication styles different from those found in everyday conversation.
For now, Anthropic cannot determine how much of the observed variation is cultural, technical, or simply the product of imbalances in the available training data.
A natural question follows: do other large language models exhibit the same phenomenon?
The study offers no answer because it examines only Claude models. Nevertheless, it is reasonable to suspect that language-dependent behavioral variation is not unique to Anthropic. Modern language models—including ChatGPT, Gemini, Grok, and others—are all trained on vast multilingual corpora and undergo broadly similar alignment processes, making comparable effects entirely plausible.
At present, however, sufficiently detailed public studies comparing different AI systems do not yet exist.
Anthropic plans to turn this methodology into a continuous monitoring tool, applying it both before and after model releases to detect unexpected behavioral shifts and evaluate whether Claude's expressed values can be deliberately influenced through changes to the training process or the model's underlying system instructions.
If future research confirms that the same phenomenon exists across other AI systems, language may come to be viewed as more than simply a vehicle for translation. It may become another parameter that shapes how an AI system reasons, frames arguments, and formulates its responses.
In that case, language selection could become as important a prompt optimization technique as wording, context, or role assignment. As language models continue to improve and technical performance converges, consistency of behavior may become one of the defining criteria by which the next generation of AI systems is judged.
Source: Anthropic, "How Claude's Values Vary by Model and Language" (July 13, 2026).