Quantification bias is a broad label for giving numeric or readily measurable information more decision weight than relevant information that is harder to compare, express, or count. The label does not describe one settled mechanism or a diagnostic scale. Its clearest direct evidence comes from a narrower experimental pattern called quantification fixation: changing which attribute appears as a number can change which option people choose, even when the underlying information stays the same (Chang et al., 2024).

That is the useful warning. A number may deserve attention because it is valid and informative. The problem begins when its format earns it influence that its relevance did not.

Two dashboards, one repair decision

This scene is fictional and illustrates the experimental logic. It is not a study or real event.

A neighborhood tool library has enough grant money for one of two repair clinics. Its committee cares equally about two outcomes: how many items the clinic can repair and how often participants can complete a similar repair independently later. An agreed assessment gives each proposal a score on both dimensions.

  • Cedar Clinic: repair-volume score 84; later-independence score 62.
  • Harbor Clinic: repair-volume score 68; later-independence score 80.

The first dashboard prints the repair-volume scores as 84/100 and 68/100. It shows later independence only as proportional bars. Cedar's advantage is the one written in digits.

The second dashboard reverses the formats. Repair volume appears as proportional bars, while later independence appears as 62/100 and 80/100. Nothing about either clinic has changed. Harbor's advantage is now the one written in digits.

If a committee member moves toward Cedar on the first dashboard and Harbor on the second, that switch could illustrate quantification fixation. Choosing Cedar once does not prove a bias, and neither does preferring an outcome that happens to be measurable. The diagnostic question is whether equivalent information receives different weight when its format changes.

What the direct experiments tested

Linda Chang, Erika Kirgios, Sendhil Mullainathan, and Katherine Milkman used this representation test across 21 preregistered experiments. Eight main-text experiments included 9,303 participants; 13 supplemental experiments included another 13,936. The tasks covered consumer, managerial, policy, hiring, and donation choices. In each, the options involved a tradeoff, and the researchers varied which dimension was numeric and which was verbal or graphical (Chang et al., 2024).

One experiment with 2,000 Prolific participants asked which of two employees should be promoted. One employee scored better on likelihood of advancement; the other scored better on likelihood of staying with the organization. The higher-advancement employee was selected by 44.2% when advancement alone was numeric, compared with 21.8% when retention alone was numeric. When both dimensions used numbers, the figure was 27.9%; when neither did, it was 32.7%. Those two same-format conditions did not significantly differ. The asymmetric format, rather than a general love of one employee, moved the choices.

The result was not confined to hypothetical selections. The research program included an incentive-compatible hiring task and real, small charitable donations. In an in-person donation experiment with 701 participants, 56.7% selected the charity that led on accountability and finance when that dimension was numeric, compared with 41.4% when its competing culture-and-community dimension was numeric. These findings show that format can alter tradeoff choices in the tested tasks. They do not establish how often the pattern controls consequential decisions outside those settings.

Why a number can pull the comparison

The focal account is comparison fluency. Digits make differences easy to inspect: 84 is larger than 68, and the gap can be calculated at a glance. Words and graphics can carry the same information without making that comparison feel as immediate.

Chang and colleagues tested this account instead of treating it as a slogan. In a preregistered experiment with 2,000 participants, awkward ratios such as 51/68 and 23/92 made the numeric comparison less fluent and attenuated quantification fixation. A supplemental experiment replicated the moderation. In a nationally representative U.S. panel of 602 adults, higher subjective numeracy, meaning greater comfort with numbers, predicted greater susceptibility; objective numeracy did not show the same moderation (Chang et al., 2024).

This evidence supports comparison fluency as a contributor, not a complete account of every metric-heavy decision. Older work on dimensional commensurability found that attributes shared across options received more weight in comparative judgments than unique attributes did (Slovic & MacPhillamy, 1974). Hsee's evaluability research showed that hard-to-judge attributes can gain or lose influence depending on whether options are assessed jointly or separately (Hsee, 1996). Both help explain why ease of comparison matters, but neither is a numeric-versus-nonnumeric replication of quantification fixation.

Numbers do not always create the same response

Quantification can clarify uncertainty, expose disagreement, and support accountability. Its behavioral effect depends on the task.

Survey experiments with national security professionals tested the concern that numeric probability estimates would create an illusion of rigor. The predicted pattern did not appear. Respondents who received numeric probabilities were less willing to support risky action and more willing to gather additional information. A different problem emerged when respondents generated the probabilities themselves: quantification magnified overconfidence, especially among lower-performing assessors (Friedman et al., 2017).

Numeric response scales can also pull judgments away from extremes. Across experiments on quantified versus verbal evaluation scales, Jinseok Chun and Michael Norton found less endpoint use and more conservative ratings when the scale used numbers. The effect weakened when the endpoint labels were already extreme (Chun & Norton, 2024).

These are different procedures with different outcomes. They rule out a blanket story in which numbers always create certainty, trust, recklessness, or neglect of other evidence.

Metric problems that need different names

Several nearby ideas involve measurement, yet their defining tests are not the repair-dashboard format swap.

Evaluability, comparability, and common measures

The evaluability hypothesis concerns whether an attribute can be judged with available reference information and whether options are assessed together or separately. Dimensional commensurability concerns how directly cues can be compared across options.

The common-measures bias applies the comparability problem to performance evaluation. In a balanced-scorecard experiment, evaluators relied more on measures shared across business units than on measures unique to each unit's strategy (Lipe & Salterio, 2000). Every item can be a metric in that design. What differs is whether it offers a common basis for comparison.

Surrogation and attribute substitution

Surrogation occurs when a measure stops being treated as an imperfect sign of a strategy and starts being treated as the strategy itself. In two experiments, pay tied to one indicator encouraged that substitution; compensation based on several indicators of the strategy weakened it (Choi et al., 2012). A later experiment found that taking part in selecting the strategy reduced surrogation, while deliberating about an assigned strategy did not reliably do so (Choi et al., 2013).

Attribute substitution is a proposed process in this literature: an accessible measure stands in for a harder strategic judgment. It is not evidence that every numeric preference contains a hidden substitution.

Targets, metric fixation, and the McNamara label

Goodhart's law and Campbell's law concern what happens when an indicator becomes a consequential target. People and organizations may redirect effort, game the indicator, or alter the process being measured. Campbell's account focused on the corruption pressure and distortion that can follow heavy use of quantitative social indicators (Campbell, 1979). Quantification fixation can occur before anyone is rewarded for moving a target.

Metric fixation and measurement fixation are broad descriptions of organizational dependence on metrics, rankings, and incentives. The McNamara fallacy is a popular label for ignoring what cannot readily be counted. These phrases may describe a recognizable management problem, but they do not supply one standardized experimental procedure. No historical McNamara story is needed to explain the psychological comparison at issue here.

Precision and commensuration

Precision bias concerns the influence of exact or granular numbers. The direct Topic 224 test instead compares numeric and nonnumeric representations. Chang and colleagues still observed the pattern when a number appeared as a range instead of a point estimate, which weighs against precision as the central explanation (Chang et al., 2024).

Commensuration is the social process of translating different qualities into a shared metric so they can be compared. That translation can change categories, visibility, and power; it is broader than an individual choice effect (Espeland & Stevens, 1998).

Where the umbrella claim holds

The catalog claim concerns a familiar imbalance: countable results can crowd out consequential outcomes that resist easy measurement. The evidence supports that meaning with two limits.

First, the direct experiments demonstrate extra weight for numeric presentation under controlled tradeoffs. They do not prove that the qualitative attribute is always more important, or that the numeric measure is poor. Second, organizational research shows that comparable metrics and performance proxies can crowd out unique or strategic information in particular tasks. It does not justify assigning every case to one cognitive mechanism.

The safe conclusion is conditional: when one consequential attribute is easier to count or compare, its format can earn it influence beyond the weight a decision maker would give the same information in another form.

Sources

  • Chang, Linda W., Erika L. Kirgios, Sendhil Mullainathan, and Katherine L. Milkman. 2024. “Does Counting Change What Counts? Quantification Fixation Biases Decision-Making.” Proceedings of the National Academy of Sciences 121(46): e2400215121. https://doi.org/10.1073/pnas.2400215121
  • Slovic, Paul, and Douglas MacPhillamy. 1974. “Dimensional Commensurability and Cue Utilization in Comparative Judgment.” Organizational Behavior and Human Performance 11(2): 172–194. https://doi.org/10.1016/0030-5073(74)90013-0
  • Hsee, Christopher K. 1996. “The Evaluability Hypothesis: An Explanation for Preference Reversals between Joint and Separate Evaluations of Alternatives.” Organizational Behavior and Human Decision Processes 67(3): 247–257. https://doi.org/10.1006/obhd.1996.0077
  • Lipe, Marlys G., and Steven E. Salterio. 2000. “The Balanced Scorecard: Judgmental Effects of Common and Unique Performance Measures.” The Accounting Review 75(3): 283–298. https://doi.org/10.2308/accr.2000.75.3.283
  • Choi, Jongwoon, Gary W. Hecht, and William B. Tayler. 2012. “Lost in Translation: The Effects of Incentive Compensation on Strategy Surrogation.” The Accounting Review 87(4): 1135–1163. https://doi.org/10.2308/accr-10273
  • Choi, Jongwoon, Gary W. Hecht, and William B. Tayler. 2013. “Strategy Selection, Surrogation, and Strategic Performance Measurement Systems.” Journal of Accounting Research 51(1): 105–133. https://doi.org/10.1111/j.1475-679X.2012.00465.x
  • Espeland, Wendy Nelson, and Mitchell L. Stevens. 1998. “Commensuration as a Social Process.” Annual Review of Sociology 24: 313–343. https://doi.org/10.1146/annurev.soc.24.1.313
  • Campbell, Donald T. 1979. “Assessing the Impact of Planned Social Change.” Evaluation and Program Planning 2(1): 67–90. https://doi.org/10.1016/0149-7189(79)90048-X
  • Friedman, Jeffrey A., Jennifer S. Lerner, and Richard Zeckhauser. 2017. “Behavioral Consequences of Probabilistic Precision: Experimental Evidence from National Security Professionals.” International Organization 71(4): 803–826. https://doi.org/10.1017/S0020818317000352
  • Chun, Jinseok S., and Michael I. Norton. 2024. “Quantification of Evaluations.” Journal of Experimental Social Psychology 110: 104558. https://doi.org/10.1016/j.jesp.2023.104558