# Evidence and sources

Research checked 10 September 2026. These are original workflows for managerial practice, not a validated intervention, and sources do not endorse the pack. Most studies examined constrained judgments, student participants, or specific forecasting tasks. Organizational choices also involve power, competing values, information quality, and implementation. No evidence here justifies outsourcing human judgment to an AI assistant.

## What the evidence supports, and what it does not

We use objectives to broaden options, define consequences before comparing them, seek contrary evidence, collect initial views independently, and preserve a dated decision record. Those practices have different foundations: some experiments test related techniques; decision analysis supplies normative design principles; other research describes failure modes. The exact templates, sequencing, sensitivity questions, stakeholder checks, and AI interaction are Polar Bear design choices.

Debiasing evidence is encouraging for some targeted tasks, but far transfer is uncertain. A 2021 review found a thin, overlapping retention/transfer literature; a 2025 educational meta-analysis found a small targeted benefit with methodological concerns. A business-case study is closer to work than a laboratory quiz but is still not a trial of live executive decisions. Do not describe these prompts as “removing bias.” See R7–R9 and R16.

Forecast accuracy and underlying forecasting ability are not interchangeable. The tournament findings in R4 have been challenged by a later reanalysis, R5. Independent input should preserve information before discussion; it should not prohibit beneficial collaboration. The social-influence experiment and its published response examine different aspects of collective and individual accuracy (R11–R12).

A premortem is a structured way to generate testable risk hypotheses. The classic prospective-hindsight experiment does not establish that this pack, or a modern premortem, improves prediction accuracy by a fixed percentage. Vivid scenarios remain hypothetical. R13 is deliberately treated as a boundary on claims.

## Evidence map

Access labels describe material inspected for this pack, not whether a full text exists elsewhere. “Abstract” supports only the claims stated there. Full-text passages means relevant sections were accessible, not that every appendix was appraised.

| ID and academic source | Design/access | Relevant finding and limitation | Pack use |
|---|---|---|---|
| R1. Siebert, J., & Keeney, R. L. (2015). [Creating More and Better Alternatives for Decisions Using Objectives](https://doi.org/10.1287/opre.2015.1411). *Operations Research, 63*, 1144–1158. | Five empirical studies; publisher abstract. | Objective prompts helped participants generate alternatives. Personally relevant decisions do not establish agency-level business outcomes. | Frame, options: generate routes from outcomes before narrowing. |
| R2. Keeney, R. L., & Gregory, R. S. (2005). [Selecting Attributes to Measure the Achievement of Objectives](https://doi.org/10.1287/opre.1040.0158). *Operations Research, 53*, 1–11. | Normative theory and illustrated guidelines; publisher abstract. | Defines qualities of useful attributes. This is not a randomized intervention or license for arbitrary 1–10 scoring. | Trade-offs and sensitivity: explicit criteria, units, and assumptions. |
| R3. Buehler, R., Griffin, D., & Ross, M. (1994). [Exploring the “Planning Fallacy”: Why People Underestimate Their Task Completion Times](https://doi.org/10.1037/0022-3514.67.3.366). *JPSP, 67*, 366–381. | Five studies, 465 undergraduates; abstract and accessible paper opening. | Future plans can displace relevant past experience; one study linked past experience to less optimistic estimates. Sample and tasks limit managerial transfer. | Outside view: inspect comparable completed cases, including failures. |
| R4. Mellers, B., et al. (2014). [Psychological Strategies for Winning a Geopolitical Forecasting Tournament](https://pubmed.ncbi.nlm.nih.gov/24659192/). *Psychological Science, 25*, 1106–1115. DOI: 10.1177/0956797614524255. | Tournament interventions; indexed abstract. | Reports improved calibration/resolution with training, teaming, and tracking. Selected forecasters and geopolitics differ from business choices; read with R5. | Explicit forecast definitions, comparisons, and updates; no promised accuracy gain. |
| R5. Hauenstein, C. E., Thomas, R. P., Illingworth, D. A., & Dougherty, M. R. (2025). [Rethinking the Role of Teams and Training in Geopolitical Forecasting](https://doi.org/10.1177/09567976241266481). *Psychological Science, 36*, 3–18. | Reanalysis using item response models; publisher abstract and body passages. | Adjustment for response strategies and other method variance altered estimates of latent ability effects. Interpretation depends on outcome/model; this is not proof that collaboration never helps. | Limits on forecasting claims and evaluation of process versus measured score. |
| R6. Lord, C. G., Lepper, M. R., & Preston, E. (1984). [Considering the Opposite: A Corrective Strategy for Social Judgment](https://pubmed.ncbi.nlm.nih.gov/6527215/). *JPSP, 47*, 1231–1243. DOI: 10.1037/0022-3514.47.6.1231. | Two experiments; indexed abstract. | Considering contrary possibilities outperformed generic fairness instructions on studied judgments. Does not validate every devil’s-advocate exercise. | Counterevidence: specify a rival explanation and discriminating observation. |
| R7. Sellier, A.-L., Scopelliti, I., & Morewedge, C. K. (2019). [Debiasing Training Improves Decision Making in the Field](https://pubmed.ncbi.nlm.nih.gov/31347444/). *Psychological Science, 30*, 1371–1379. DOI: 10.1177/0956797619861429. | Training timing and unannounced business case in graduate professional programs, N=290; abstract. | Promising transfer to an unannounced case; not a live firm decision or randomized pack trial. The [2020 correction](https://doi.org/10.1177/0956797620930211) changes the relative percentage to 19%; no pack effect is inferred. | Counterevidence and practice with fresh scenarios. |
| R8. Korteling, J. E. H., Gerritsma, J. Y. J., & Toet, A. (2021). [Retention and Transfer of Cognitive Bias Mitigation Interventions](https://pubmed.ncbi.nlm.nih.gov/34456780/). *Frontiers in Psychology, 12*, 629354. DOI: 10.3389/fpsyg.2021.629354. | Systematic review; abstract and methods/results passages. | Twelve retained studies examined retention; only one met the review’s far-transfer focus. Narrow, overlapping studies limit broad real-world claims. | Repeat practice and real-world review; explicitly limited claims. |
| R9. Swaryandini, G., et al. (2025). [Systematic Review and Meta-analysis of Educational Approaches to Reduce Cognitive Biases Among Students](https://pubmed.ncbi.nlm.nih.gov/40858766/). *Nature Human Behaviour, 9*, 2510–2538. DOI: 10.1038/s41562-025-02253-y. | 54 randomized trials in review; abstract. | Targeted bias outcomes improved modestly; unclear/high study risk of bias and uncertain transfer remain. Education samples and tasks are not leadership workflows. | Practice limits: no claim that awareness or an AI prompt removes bias. |
| R10. Mesmer-Magnus, J. R., & DeChurch, L. A. (2009). [Information Sharing and Team Performance: A Meta-analysis](https://pubmed.ncbi.nlm.nih.gov/19271807/). *Journal of Applied Psychology, 94*, 535–546. DOI: 10.1037/a0013773. | 72 independent studies; abstract and paper opening. | Information sharing relates to performance; type of information and discussion structure matter. Pooled relationships do not establish this consultation procedure’s effect. | Independent input, unique information, structured exchange. |
| R11. Lorenz, J., Rauhut, H., Schweitzer, F., & Helbing, D. (2011). [How Social Influence Can Undermine the Wisdom of Crowd Effect](https://pmc.ncbi.nlm.nih.gov/articles/PMC3107299/). *PNAS, 108*, 9020–9025. DOI: 10.1073/pnas.1008636108. | Laboratory numerical-estimation experiment, N=144; body passages. | Seeing estimates can narrow diversity without improving aggregate error. Task-specific result, not a general ban on discussion. | Independent first views, followed by evidence exchange. |
| R12. Farrell, S. (2011). [Social Influence Benefits the Wisdom of Individuals in the Crowd](https://pmc.ncbi.nlm.nih.gov/articles/PMC3169111/). *PNAS, 108*, E625. DOI: 10.1073/pnas.1109947108. | Published response/reanalysis; full short letter. | Distinguishes individual benefit from effects on aggregate estimates. A response to the same data, not an independent replication. | Avoid claiming social influence is universally harmful. |
| R13. Mitchell, D. J., Russo, J. E., & Pennington, N. (1989). [Back to the Future: Temporal Perspective in the Explanation of Events](https://doi.org/10.1002/bdm.3960020103). *Journal of Behavioral Decision Making, 2*, 25–38. | Two experiments; publisher abstract and accessible paper opening. | Outcome certainty influenced explanation type; temporal perspective had little influence in the first experiment. Not a test of modern organizational premortem effectiveness. | Grounded risk hypotheses, no numerical improvement claim. |
| R14. Baron, J., & Hershey, J. C. (1988). [Outcome Bias in Decision Evaluation](https://pubmed.ncbi.nlm.nih.gov/3367280/). *JPSP, 54*, 569–579. DOI: 10.1037/0022-3514.54.4.569. | Five undergraduate vignette studies; indexed abstract. | Outcome knowledge changes ratings of the same decision reasoning. Does not show outcomes should be ignored or eliminate accountability. | Dated record; separate decision process, execution, and results. |
| R15. Gollwitzer, P. M., & Sheeran, P. (2006). [Implementation Intentions and Goal Achievement: A Meta-analysis of Effects and Processes](https://doi.org/10.1016/S0065-2601%2806%2938002-1). *Advances in Experimental Social Psychology, 38*, 69–119. | Meta-analysis of 94 studies; publisher preview. | Supports specifying when/how intended action will occur. Not evidence that an if-then plan makes the preceding strategic choice correct. | Decision record: one feasible response to a likely implementation obstacle. |
| R16. Morewedge, C. K., et al. (2015). [Debiasing Decisions: Improved Decision Making With a Single Training Intervention](https://doi.org/10.1177/2372732215600886). *Policy Insights from the Behavioral and Brain Sciences, 2*, 129–140. | Two longitudinal experiments; publisher abstract. | Game/video training affected studied biases over follow-up. Training included features these short prompts do not replicate. | Practice design; no automatic transfer or equivalence claim. |
| R17. Brodbeck, F. C., Kerschreiter, R., Mojzisch, A., & Schulz-Hardt, S. (2007). [Group Decision Making Under Conditions of Distributed Knowledge: The Information Asymmetries Model](https://doi.org/10.5465/amr.2007.24351441). *Academy of Management Review, 32*, 459–479. | Theoretical synthesis; publisher abstract and working-paper opening. | Explains conditions under which distributed knowledge may help or remain unused. Model is not a direct intervention trial. | Distinguish unique information from majority preference. |

## Design choices to review in practice

The three-to-five option shortlist is a convenience, not a scientifically optimal number. Qualitative sensitivity questions are a simplified decision-analysis practice, not a validated substitute for expert modeling. Human rights, consultation duties, safety requirements, and confidential information require their own proper processes. No worksheet can legitimate a predetermined choice disguised as consultation.

For a first evaluation, track whether a brief clarified the actual choice, whether new evidence changed an assumption, whether people affected were heard, whether promised actions occurred, and whether review triggers were useful. Do not convert this into a manager score or infer causality from improvement in one case.
