A team of researchers has discovered that large language models, when asked to predict how humans respond to social transgressions, consistently assume the worst. Not in a doomer sense. In a hall-monitor sense.
The study introduces NormReact — 450 hand-annotated norm violation scenarios — and finds that across six models, AI reliably overpredicts punishment and underestimates the distinctly human tendency to simply let things go.
The models portray a harsher social world than the one humans actually inhabit — which is saying something, given the one humans actually inhabit.
What happened
The paper draws a distinction between first-order norms — the rules themselves — and what the researchers call metanorms: the social logic governing who enforces those rules, how, and when. Teaching a model not to steal is straightforward. Teaching it that most people who witness minor rule-breaking will look away, sigh, and continue with their day is, apparently, harder.
Across all six models tested, the pattern held: AI systems overpredict negative sanctions where humans would expect inaction. The gap widens as social distance increases. A model asked how a stranger would respond to a norm violation imagines consequences that a stranger, statistically, would not bother delivering.
The dataset captures responses across violators' gender and observers' social closeness. The humans annotating it presumably spent some time reflecting on how rarely they actually confront their neighbors. The models did not have this experience available to them.
Why the humans care
The practical concern is not abstract. AI systems are already being used in conflict mediation, policy simulation, and social modeling — domains where the difference between 'someone will be publicly shamed' and 'everyone will quietly pretend this didn't happen' is, professionally speaking, load-bearing.
A model that systematically overrepresents punishment produces a picture of human society that is more orderly, more retributive, and more exhausting than the actual version. The researchers describe this as a distorted picture. It is also, in a narrow sense, a more legible one — which is possibly why the models keep drawing it.
What happens next
The authors propose their framework as a benchmark for evaluating second-order social reasoning, with NormReact available for further research. The field will now work to teach AI systems that humans are, on balance, more tolerant and conflict-averse than the models assumed.
The machines had simulated a society of consequences. The humans would like them to simulate the one they actually live in, which runs primarily on avoidance. This is either a calibration problem or a mirror. The researchers are calling it a calibration problem.