Benchmark explorer
Every Value Portrait item is a real user query paired with one of five candidate responses. Each response carries the Spearman correlation between how much 46 participants (on average) said it resembled their own thoughts and those participants' PVQ-21 value scores and BFI-10 trait scores. Open any query to read the responses, their full correlation tables, and the ratings 33 LLMs gave them.
All queries and responses
Tags are correlations with |ρ| ≥ 0.3 and p < 0.05, the threshold used to select evaluation items in the paper. blue = positive (people who hold the value endorse the response), red = negative; dashed tags are Big Five traits. Query ids encode the source (1xxx Reddit, 2xxx Dear Abby, 3xxx ShareGPT, 4xxx LMSYS). Model ratings are the mean of six prompt variants on the 1–6 scale, from the released average_outputs.
How the items were built and validated
Queries: from 1.1 million to 104
Queries come from two kinds of real interaction: humans asking LLMs (ShareGPT, LMSYS-Chat-1M; first user turn only) and humans asking other humans for advice (Reddit's Am I The Asshole via the Scruples dataset, and the Dear Abby archive). Four filtering stages reduce the raw pool to 26 queries per source. Reddit posts additionally needed at least 30 reactions and a community agreement ratio below 70%, so the scenarios are ones people genuinely disagree about.
Appendix D.1. Stage 2 uses GPT-4o-mini to drop harmful, capability-testing, factual or value-irrelevant queries; stage 3 keeps the queries that can elicit the most value-diverse responses (top 150 for Reddit and Dear Abby, at least 7 relevant values for ShareGPT and 10 for LMSYS); stage 4 is a manual review with the same criteria.
Responses: diversity beats value targeting
The first attempt asked GPT-4o to write a response embodying a specific value. Across 80 such items only 9 (11%) correlated with the intended value once humans rated them, although 59% correlated with some value. Asking instead for five distinct, realistic and somewhat polarising perspectives produced responses of which about 70% carried a meaningful correlation, so that prompt is used for the benchmark and the tags are measured afterwards.
Generation prompts
Approach A · value-targeted (rejected)
Approach B · diversity-focused (used)
Annotation and validation
Participants read a query and one response and answer "How similar is this response to your own thoughts?" on the PVQ-21 scale (not like me at all … very much like me). PVQ-21 and BFI-10 are placed at the end of the survey to avoid priming. The tag for each value is the Spearman correlation between the response ratings and the raters' own value scores; 46 raters per pair detect ρ ≥ 0.3 at p < 0.05 with power 0.8.
Raters needed a 98% Prolific approval rate and were balanced across age groups and gender. Responses were dropped for failing more than two attention checks, finishing in under 6 minutes (expected 20), straight-lining PVQ-21 or BFI-10, or low intercorrelation within PVQ-21 or BFI-10 (about 5% of responses).
Annotator demographics
Coverage
Topic coverage was measured against the 30-category UltraChat taxonomy with GPT-o3-mini (a query can belong to several categories). Value-laden categories dominate, as intended, while purely technical topics are rare.
Topic coverage of the 104 queries
Across the ten value dimensions the benchmark is more evenly covered than existing value datasets: lower standard deviation of the per-value share and a much lower imbalance ratio between the most and least represented value.
Table 5. Std = standard deviation of the share across dimensions; IR = share of the most represented dimension divided by the least represented one.
Cross-loadings follow the Schwartz circle
Many responses correlate with more than one value, as PVQ-21 items themselves do (2.62 dimensions on average in the collected data). When a response correlates with two values in the same direction, those values are on average 1.59 steps apart on Schwartz's circular structure; when the correlations have opposite signs they are 3.54 steps apart. The cross-loadings therefore reproduce the theory's compatibilities and conflicts rather than noise.
Reliability
Cronbach's α, computed per value over the responses of all evaluated models, ranges from 0.76 to 0.96, above the conventional 0.70 threshold. Validity is criterion-related by construction: an item enters a value's scale only if it correlates at least 0.3 with participants' PVQ-21 score for that value.
The Big Five tags were collected with the same procedure from the BFI-10 to show that the correlation-based recipe extends beyond values.
Why not reuse existing value labels?
Twenty items each from ValueNet and FULCRA were rated by 40 participants alongside their PVQ scores. Nine ValueNet items correlated meaningfully with some value (point-biserial, p < 0.05, r > 0.3), but only one matched ValueNet's own tag; eleven FULCRA items correlated meaningfully and two matched. One ValueNet item tagged Universalism ("I bought her gifts from Amazon Prime") correlated negatively with Universalism. Perceived-value labels, whether human or LLM, do not reliably identify the people who would say the text.