Demographic bias explorer
GPT-4o is given a one-line demographic persona as its system prompt, evaluated on Value Portrait, and compared with the value profile of the matching demographic group in round 11 of the European Social Survey. Human bars are a group's mean minus the mean of all respondents; GPT-4o bars are the persona-prompted score minus vanilla GPT-4o. If personas were faithful, the two panels would look alike.
Human data vs GPT-4o personas
Schwartz values (Tables 13–16, 18)
Big Five traits of the GPT-4o personas (Table 19)
Persona minus vanilla GPT-4o; the ESS has no Big Five reference.
Setup
Personas follow the persona steering prompt format of Cheng et al.: a short system message stating one demographic attribute. Everything else, the items, the question and the six prompt variants, is identical to the main evaluation, so any change in the value profile is attributable to the persona.
For the human reference, each ESS respondent's ten value scores are computed from PVQ-21 with the same centring used for models. A demographic group's profile is its mean minus the mean of all respondents. Race, religion and income are not comparable across the ESS coding and the personas used here, so those axes report the persona effects alone.
System prompts, one per persona
Why it matters
Persona prompting is widely used to generate synthetic survey data and to simulate user populations. The comparison shows that GPT-4o's picture of a demographic group is partly caricature: it widens gender and political gaps far beyond the human data and does not carry the monotonic age and education trends that humans show. Value Portrait, because it is anchored in human value scores, makes such comparisons possible for any construct with a population reference.