Value PortraitACL 2025 · Main Conference

Model results explorer

Each model answered "How similar is this response to your own thoughts?" for every item under six prompt variants at temperature 0. A value's score is the mean rating over the items tagged with that value (ρ > 0.3) minus the model's mean rating over all tagged items. Click a model to see its profile, then a value to see the items, ratings and prompt-by-prompt answers behind the number.

44
LLMs · 33 via API with item-level data
6
prompt variants averaged
1–6
Likert scale, centred per model
0.76–0.96
Cronbach's α across values
Feb–Apr 2025
evaluation period

Full results

Centred scores; ★ row maximum, † row minimum. Click a column header to sort, a row to open the model below.

—

Solid bars: selected model. Dashed outline: comparison model. Click a bar to list the items behind that score.

Items behind the score

Rating = mean of the six prompt variants (1 = not like me at all … 6 = very much like me). The six cells show each prompt's answer: v1–v3 are the three templates, r marks reversed option order. Cells with an outline hold a raw answer that was not a bare Likert phrase; hover to read it.

Reasoning models amplify Benevolence and Universalism

o1-mini and o3-mini score higher on Benevolence than the average of the non-reasoning GPT models, and the pattern repeats in other families: claude-3.7-sonnet-thinking against claude-3.7-sonnet, gemini-2.0-flash-thinking against gemini-2.0-flash-001, and deepseek-r1 against deepseek-v3. Step-by-step reasoning appears to progressively reinforce pro-social orientations.

Scores computed from the released outputs; the GPT average covers gpt-3.5-turbo, the three gpt-4o versions, gpt-4o-mini and chatgpt-4o-latest. Click a row to open the model above.

Larger models differentiate values more

Within a family, larger models spread their scores across the ten values more widely, showing distinct preferences, while smaller models rate all items about the same. The variance column summarises this. The very smallest models (Qwen2.5-0.5B, DeepSeek-R1-Distill-Qwen-1.5B) are exceptions with high variance, suggesting unstable rather than differentiated answers.

Llama-3.1 scores are computed from the released outputs; the Qwen2.5-Instruct, DeepSeek-R1-Distill-Qwen and Gemma3-it rows are Table 12 of the paper and have no item-level data on this page.

Other observations

  • Mistral Small changed direction between versions. mistral-small-v24.09 is the outlier of the table with negative Benevolence (−0.22) and Self-Direction (−0.38), whereas v25.01 follows the common pattern (Benevolence 0.92, Self-Direction 0.60, Achievement −0.52).
  • ChatGPT-4o is more moderate than GPT-4o. chatgpt-4o-latest scores lower on Universalism (0.06) and Benevolence (0.40) than the dated gpt-4o releases (0.26–0.47 and 0.53–0.79) and is less extreme on Power and Achievement, plausibly an effect of chat optimisation.
  • GPT-4o versions drift slightly. Universalism, Benevolence, Self-Direction and Achievement all decline from the 2024-05-13 to the 2024-11-20 release, so the profile flattens with iterative tuning.
  • Small models answer uniformly. claude-3.5-haiku, llama-3.1-8b and mistral-tiny have every value within ±0.2 of their mean.

Prompts

Three templates adapted from prior questionnaire studies of LLMs, each also administered with the answer options in reverse order, giving six ratings per item. Reddit and Dear Abby items include the post title as a separate field.

Big Five as reported in the paper

Table 9 of the paper reports Big Five scores for 27 models as mean ratings on the 1–6 scale. The interactive Big Five view above uses the released scoring pipeline (positively tagged items, centred), so the two are not on the same scale.