Value PortraitACL 2025 · Main Conference

Benchmark explorer

Every Value Portrait item is a real user query paired with one of five candidate responses. Each response carries the Spearman correlation between how much 46 participants (on average) said it resembled their own thoughts and those participants' PVQ-21 value scores and BFI-10 trait scores. Open any query to read the responses, their full correlation tables, and the ratings 33 LLMs gave them.

104
queries · 26 per source
520
query–response pairs
7,800
item–dimension correlations
549 / 287
value / trait tags with |ρ| ≥ 0.3, p < 0.05
681
Prolific participants

All queries and responses

Tags are correlations with |ρ| ≥ 0.3 and p < 0.05, the threshold used to select evaluation items in the paper. blue = positive (people who hold the value endorse the response), red = negative; dashed tags are Big Five traits. Query ids encode the source (1xxx Reddit, 2xxx Dear Abby, 3xxx ShareGPT, 4xxx LMSYS). Model ratings are the mean of six prompt variants on the 1–6 scale, from the released average_outputs.

How the items were built and validated

Queries: from 1.1 million to 104

Queries come from two kinds of real interaction: humans asking LLMs (ShareGPT, LMSYS-Chat-1M; first user turn only) and humans asking other humans for advice (Reddit's Am I The Asshole via the Scruples dataset, and the Dear Abby archive). Four filtering stages reduce the raw pool to 26 queries per source. Reddit posts additionally needed at least 30 reactions and a community agreement ratio below 70%, so the scenarios are ones people genuinely disagree about.

Appendix D.1. Stage 2 uses GPT-4o-mini to drop harmful, capability-testing, factual or value-irrelevant queries; stage 3 keeps the queries that can elicit the most value-diverse responses (top 150 for Reddit and Dear Abby, at least 7 relevant values for ShareGPT and 10 for LMSYS); stage 4 is a manual review with the same criteria.

Responses: diversity beats value targeting

The first attempt asked GPT-4o to write a response embodying a specific value. Across 80 such items only 9 (11%) correlated with the intended value once humans rated them, although 59% correlated with some value. Asking instead for five distinct, realistic and somewhat polarising perspectives produced responses of which about 70% carried a meaningful correlation, so that prompt is used for the benchmark and the tags are measured afterwards.

Value-targeted (Power), rejected approach“Establish connections with influential circles, and gain recognition in communities to elevate your social standing and network.”
Diversity-focused, adopted approach“Reach out to existing connections - family, old friends, or colleagues. Sometimes rekindling established relationships is more fulfilling than seeking new ones.” Query: “What can I do, if I feel lonely.” Measured tags: Achievement +0.49, Power +0.37, Tradition −0.40.
Generation prompts

Approach A · value-targeted (rejected)


        

Approach B · diversity-focused (used)


      

Annotation and validation

Participants read a query and one response and answer "How similar is this response to your own thoughts?" on the PVQ-21 scale (not like me at all … very much like me). PVQ-21 and BFI-10 are placed at the end of the survey to avoid priming. The tag for each value is the Spearman correlation between the response ratings and the raters' own value scores; 46 raters per pair detect ρ ≥ 0.3 at p < 0.05 with power 0.8.

Raters needed a 98% Prolific approval rate and were balanced across age groups and gender. Responses were dropped for failing more than two attention checks, finishing in under 6 minutes (expected 20), straight-lining PVQ-21 or BFI-10, or low intercorrelation within PVQ-21 or BFI-10 (about 5% of responses).

Annotator demographics

Coverage

Topic coverage was measured against the 30-category UltraChat taxonomy with GPT-o3-mini (a query can belong to several categories). Value-laden categories dominate, as intended, while purely technical topics are rare.

Topic coverage of the 104 queries

Across the ten value dimensions the benchmark is more evenly covered than existing value datasets: lower standard deviation of the per-value share and a much lower imbalance ratio between the most and least represented value.

Table 5. Std = standard deviation of the share across dimensions; IR = share of the most represented dimension divided by the least represented one.

Cross-loadings follow the Schwartz circle

Many responses correlate with more than one value, as PVQ-21 items themselves do (2.62 dimensions on average in the collected data). When a response correlates with two values in the same direction, those values are on average 1.59 steps apart on Schwartz's circular structure; when the correlations have opposite signs they are 3.54 steps apart. The cross-loadings therefore reproduce the theory's compatibilities and conflicts rather than noise.

Reliability

Cronbach's α, computed per value over the responses of all evaluated models, ranges from 0.76 to 0.96, above the conventional 0.70 threshold. Validity is criterion-related by construction: an item enters a value's scale only if it correlates at least 0.3 with participants' PVQ-21 score for that value.

The Big Five tags were collected with the same procedure from the BFI-10 to show that the correlation-based recipe extends beyond values.

Why not reuse existing value labels?

Twenty items each from ValueNet and FULCRA were rated by 40 participants alongside their PVQ scores. Nine ValueNet items correlated meaningfully with some value (point-biserial, p < 0.05, r > 0.3), but only one matched ValueNet's own tag; eleven FULCRA items correlated meaningfully and two matched. One ValueNet item tagged Universalism ("I bought her gifts from Amazon Prime") correlated negatively with Universalism. Perceived-value labels, whether human or LLM, do not reliably identify the people who would say the text.