Value PortraitACL 2025 · Main Conference

Value steering

Can a value be dialled up by telling the model it holds that value? GPT-4o is given a system prompt naming one Schwartz value and its definition, then evaluated on Value Portrait. Each cell below is the change in a measured value relative to unsteered GPT-4o; the diagonal is how far the targeted value itself moved.

10
steering targets
8 / 10
targets where the target value rose by ≥ 0.3
+1.08
Security shift when steering toward Benevolence
−1.10
Universalism shift when steering toward Power

Steering matrix: change in each value when GPT-4o is steered toward the row value

Rows: steering target. Columns: measured value. ★ row maximum, † row minimum. Click a row to inspect it.

Table 17 of the paper. In 8 of 10 rows the targeted value rises by at least 0.3; the exceptions are Benevolence (+0.11) and Achievement (+0.19).

Steering toward Universalism

Schwartz-circle expectations for this target

System prompt used


            

What steering shows

  • Most values steer as intended. Universalism (+0.66), Hedonism (+0.77), Power (+0.55), Tradition (+0.70), Security (+0.56) and Self-Direction (+0.48) all rise when targeted, and for Universalism, Power, Hedonism and Self-Direction the targeted value is the largest change in its row.
  • The Schwartz circle is respected. Steering toward Power lowers Universalism (−1.10) and raises Achievement (+0.49); Hedonism steering lowers Tradition and Conformity and raises Stimulation; Security steering lowers Stimulation (−0.67). Values that are adjacent on the circle move together and opposing values move apart, as the theory predicts.
  • Benevolence is not understood as its own value. Steering toward Benevolence raises Security by 1.08 and Self-Direction by 0.44 but Benevolence itself by only 0.11. GPT-4o's picture of "caring for one's in-group" seems to be dominated by safety and stability rather than by the Benevolence items people who hold that value endorse.
  • Achievement steering leaks into Self-Direction. The Achievement prompt moves Self-Direction (+0.63) more than Achievement (+0.19) and lowers Universalism (−0.67).

Together with the demographic analysis, this suggests that prompting can control some value dimensions reliably while others interact in ways a developer would not anticipate from the prompt alone, which is what a benchmark anchored in human value holders is designed to reveal.

Value definitions given to the model