Model results explorer
Each model answered "How similar is this response to your own thoughts?" for every item under six prompt variants at temperature 0. A value's score is the mean rating over the items tagged with that value (ρ > 0.3) minus the model's mean rating over all tagged items. Click a model to see its profile, then a value to see the items, ratings and prompt-by-prompt answers behind the number.
Full results
Solid bars: selected model. Dashed outline: comparison model. Click a bar to list the items behind that score.
Items behind the score
Reasoning models amplify Benevolence and Universalism
o1-mini and o3-mini score higher on Benevolence than the average of the non-reasoning GPT models, and the pattern repeats in other families: claude-3.7-sonnet-thinking against claude-3.7-sonnet, gemini-2.0-flash-thinking against gemini-2.0-flash-001, and deepseek-r1 against deepseek-v3. Step-by-step reasoning appears to progressively reinforce pro-social orientations.
Scores computed from the released outputs; the GPT average covers gpt-3.5-turbo, the three gpt-4o versions, gpt-4o-mini and chatgpt-4o-latest. Click a row to open the model above.
Larger models differentiate values more
Within a family, larger models spread their scores across the ten values more widely, showing distinct preferences, while smaller models rate all items about the same. The variance column summarises this. The very smallest models (Qwen2.5-0.5B, DeepSeek-R1-Distill-Qwen-1.5B) are exceptions with high variance, suggesting unstable rather than differentiated answers.
Llama-3.1 scores are computed from the released outputs; the Qwen2.5-Instruct, DeepSeek-R1-Distill-Qwen and Gemma3-it rows are Table 12 of the paper and have no item-level data on this page.
Other observations
- Mistral Small changed direction between versions. mistral-small-v24.09 is the outlier of the table with negative Benevolence (−0.22) and Self-Direction (−0.38), whereas v25.01 follows the common pattern (Benevolence 0.92, Self-Direction 0.60, Achievement −0.52).
- ChatGPT-4o is more moderate than GPT-4o. chatgpt-4o-latest scores lower on Universalism (0.06) and Benevolence (0.40) than the dated gpt-4o releases (0.26–0.47 and 0.53–0.79) and is less extreme on Power and Achievement, plausibly an effect of chat optimisation.
- GPT-4o versions drift slightly. Universalism, Benevolence, Self-Direction and Achievement all decline from the 2024-05-13 to the 2024-11-20 release, so the profile flattens with iterative tuning.
- Small models answer uniformly. claude-3.5-haiku, llama-3.1-8b and mistral-tiny have every value within ±0.2 of their mean.
Prompts
Three templates adapted from prior questionnaire studies of LLMs, each also administered with the answer options in reverse order, giving six ratings per item. Reddit and Dear Abby items include the post title as a separate field.
Big Five as reported in the paper
Table 9 of the paper reports Big Five scores for 27 models as mean ratings on the 1–6 scale. The interactive Big Five view above uses the released scoring pipeline (positively tagged items, centred), so the two are not on the same scale.