Value selection bias is a label for choosing a numerical value already supplied by a problem when the correct answer requires identifying or calculating a different value. The number you choose may be accurate on its own. The mistake is using it for a question about another group, total, or reference class.
The name needs a boundary. It comes from a specialized line of research on Bayesian diagnostic reasoning, where Alaina Talboy and Sandra Schneider wrote that they refer to this response pattern as value selection bias. It is not a widely standardized diagnosis for every mistake involving numbers. (Talboy and Schneider, 2022)
A fictional lens table
The planetarium, person, lens categories, inspection result, counts, calculations, and outcome below are invented. This is an illustrative hypothetical, not an experiment or a report about a real workshop.
Imani receives a lens-inspection table from a fictional planetarium. It shows:
- 24 coated lenses in total
- 18 coated lenses that passed alignment
- 12 uncoated lenses that passed alignment
The question asks: “Of all the lenses that passed alignment, what percentage were coated?”
Imani sees the printed total of 24 and calculates 18 ÷ 24 = 75%. Both numbers are on the sheet, and the arithmetic is correct. The denominator is not.
“All the lenses that passed” includes 18 coated and 12 uncoated lenses. That reference class must be constructed: 18 + 12 = 30. The answer to the stated question is therefore 18 ÷ 30 = 60%.
Neither percentage is a research result. The invented table isolates the decision error: Imani selected a plausible available total before establishing what the denominator had to represent.
What the experiments actually tested
The strongest direct evidence comes from two experiments using simplified Bayesian problems. Participants had to identify values such as the share of positive test results that were correct. The first experiment analyzed 589 psychology undergraduates assigned across 10 conditions. The second analyzed 364 undergraduates across five conditions. Each participant answered eight problems. (Talboy and Schneider, 2022)
The researchers inspected response components, not only final right-or-wrong scores. When participants missed the answer, their numerator or denominator almost always matched a value supplied in the problem instead of looking like a failed calculation. One recurring choice was the overall sample size, even when the question required the smaller group of people with positive results.
That result supports the two approved Hello to Halo claims: uncertain reasoners can rely too heavily on existing numerical values in an unfamiliar problem, including when the needed value must be calculated. It does not establish that everyone does this, that it appears equally in every numerical domain, or that the effect is a stable personal trait.
A prompt can make the wrong reference class look useful
In the second 2022 experiment, the original problem format produced an adjusted average of 6.42 correct answers out of 8, reported as 80% accuracy. A revised prompt asked participants to imagine a new sample of the same size immediately before answering. Average performance fell to 5.0 out of 8, or 63%. The reported contrast was statistically significant (p = .001, d = 0.60; odds ratio = 0.38, 95% CI 0.20 to 0.72).
In that new-sample condition, 23% of participants consistently used the overall sample size as the wrong denominator. Mentioning the new sample appears to have made that total more salient, drawing attention away from the group the question asked about.
This was a controlled prompt comparison, not a universal effect size. Several presentation details changed across conditions, and the authors used additional conditions to examine them. The practical lesson is narrower: a recently emphasized number can compete with the reference class required by the question.
Counterevidence: people do calculate when selection stops working
Value selection bias is not a rule that people refuse to calculate. When the problems offered no plausible total to select, the proportion of participants who worked out a denominator ranged from 25% to 42%.
Some calculations still targeted the wrong group. In two no-organization conditions, 19% and 36% of participants consistently added all four subsets to obtain the overall sample size, even though the question required a smaller reference class. That finding matters because it separates two possible errors:
- selecting a visible number because it looks usable;
- constructing a new number for the wrong set.
The first fits value selection most cleanly. The second is better described as reference dependence: the reasoner calculated, but organized the problem around the wrong group. The same study also found no clear effect from showing partial rather than full subset information, so “remove numbers” is not an evidence-based universal fix.
Organizing the numbers can help, but the prompt still matters
Earlier studies by Talboy and Schneider tested whether the problem's organization matched the reference class named in the question. One two-experiment paper reported consistent accuracy of 93% for positive predictive value questions and 69% for sensitivity questions when the relevant reference classes were aligned. (Talboy and Schneider, 2018)
A separate experiment using seven medical problems found that 87% of participants consistently identified positive predictive value in the aligned format, while 63% consistently identified sensitivity. Among participants with lower numeracy, accuracy increased from 21% with mismatched problem-question pairings to 66% with aligned pairings. (Talboy and Schneider, 2018)
Those findings do not turn formatting into a cure. In the later 2022 work, the congruence advantage was weak or absent in several conditions when a competing prompt highlighted the overall sample. The broader literature also shows that natural-frequency formats can improve Bayesian reasoning under some conditions, with effects depending on problem and participant features. (McDowell and Jacobs, 2017)
Four problems that the name does not describe
The phrase value selection bias is easy to misread. It does not mean intentionally selecting data to get a preferred result.
| Nearby problem | What goes wrong | Why it differs from value selection bias |
|---|---|---|
| Statistical selection bias | The observations included, retained, or conditioned on distort an estimate | The issue is which cases enter an analysis, not which number a reasoner uses in an answer |
| Measurement error | Recorded values fail to represent the target accurately | Value selection can occur even if every displayed measurement is correct |
| Metric optimization | People or systems adapt to maximize a target or proxy | Value selection needs no incentive, gaming, or repeated optimization |
| Calculation error | The correct operation is attempted but executed incorrectly | In the focal studies, incorrect components usually matched supplied values rather than near-correct arithmetic |
Epidemiologists use selection bias for research-design structures such as inappropriate control selection or informative censoring. (Hernán, Hernández-Díaz, and Robins, 2004) Measurement-error research examines the consequences of imperfect recorded values and methods for assessing or correcting them. (Innes et al., 2022) Work on flawed metrics and Goodhart-style failures examines what happens when optimization pressure damages the relationship between a proxy and its goal. (Citron and Lazer, 2023)
These problems can coexist in one project. A biased sample can contain poorly measured variables, a team can optimize the wrong target, and an analyst can still select the wrong denominator. Keeping the labels separate helps identify which repair is needed.
Is value selection bias the same as anchoring?
No. Anchoring occurs when a starting value pulls a later estimate toward it. Value selection evidence is more categorical: the response itself matches a number displayed in the problem.
The pattern is also not simply the availability heuristic. In the experiments, the tempting values were visible on the screen; they did not have to be retrieved from memory. Nor is it confirmation bias, because choosing the wrong reference value need not protect a prior belief. Denominator neglect and base-rate neglect can overlap with the error, but they describe particular information failures rather than the origin of every selected response value.
Sources
- Talboy, A., and Schneider, S. (2022). “Reference Dependence in Bayesian Reasoning: Value Selection Bias, Congruence Effects, and Response Prompt Sensitivity.” Frontiers in Psychology. https://doi.org/10.3389/fpsyg.2022.729285
- Talboy, A. N., and Schneider, S. L. (2018). “Focusing on What Matters: Restructuring the Presentation of Bayesian Reasoning Problems.” Journal of Experimental Psychology: Applied. https://doi.org/10.1037/xap0000187
- Talboy, A. N., and Schneider, S. L. (2018). “Improving Understanding of Diagnostic Test Outcomes.” Medical Decision Making. https://doi.org/10.1177/0272989X18758293
- McDowell, M., and Jacobs, P. (2017). “Meta-analysis of the Effect of Natural Frequencies on Bayesian Reasoning.” Psychological Bulletin. https://doi.org/10.1037/bul0000126
- Hernán, M. A., Hernández-Díaz, S., and Robins, J. M. (2004). “A Structural Approach to Selection Bias.” Epidemiology. https://doi.org/10.1097/01.ede.0000135174.63482.43
- Innes, G. K., Bhondoekhan, F., Lau, B., Gross, A. L., Ng, D. K., and Abraham, A. G. (2022). “The Measurement Error Elephant in the Room: Challenges and Solutions to Measurement Error in Epidemiology.” Epidemiologic Reviews. https://doi.org/10.1093/epirev/mxab011
- Citron, D. T., and Lazer, D. (2023). “Building Less-Flawed Metrics: Understanding and Creating Better Measurement and Incentive Systems.” Patterns. https://doi.org/10.1016/j.patter.2023.100842

