Critical Assumptions & Bias Mitigation

Critical Assumptions & Bias Mitigation

Strategy fails less from bad models than from unseen assumptions and unmanaged bias. Scenarios give you structures; assumptions are the joints. If you don’t surface, test, and track them—especially the ones leadership is subconsciously anchoring on—your ranges harden into wishful thinking and your triggers won’t fire until it’s expensive. This chapter turns assumptions into first‑class objects: plainly worded, traceable to evidence, tied to decisions, and designed to be disconfirmed quickly and cheaply.

Bias mitigation is not a workshop gimmick here. It is operational: a Challenger with real standing, mandatory disconfirming searches, premortems before Gate decisions, and signposts that force action. The central artifact is TPL‑10 Assumptions & Coherence Log—the running ledger of what must be true, why we believe it, how we could be wrong, and what we will do to find out fast. Treat it like a risk register for thinking.

8.1 Assumption Elicitation & Shortlist—Step‑by‑Step

Assumptions are the hinges between narrative and numbers. They state what must be true—within a defined horizon and scenario—for an option, hedge, or gate to be attractive. Treated carelessly, they bloat into a catalog. Treated rigorously, they compress uncertainty into a short list you can test and wire to signposts, ranges, and learning gates.

A. Write assumptions so they can be tested

Use this assumption grammar and apply it consistently:

We believe [subject] will [direction/level] by [amount/range] within [horizon] in [Scenario S] because [mechanism]; falsify if [observable threshold] from [source] by [date].

Mandatory fields: scenario conditioning (S‑##), horizon (monitoring vs. evaluation), unit/measure, mechanism, falsification threshold, data source/cadence, owner. Each card gets an ID [A‑###] and links to drivers [DRV‑###], model blocks [MB‑##], and (when relevant) a trigger TRG‑###.

Good assumptions are decision‑moving: if the statement flipped, you would change an option, hedge, trigger, or gate. If not, it belongs in background context, not on the shortlist.

B. The 90‑minute elicitation & shortlist (mechanics that work)

1) Pre‑work (5 min).
Re‑state the decision sentence and which decision this work could change. FP&A shows a reference class (base rates for adoption ramps, pass‑through lags, supply ramps, enforcement cycles).

2) Silent writing (10 min).
Each participant writes 6–8 candidate assumptions across demand, price/mix, cost/lead time, capacity/quality, regulation/compliance, partner/supplier reliability. No discussion.

3) Cluster & classify (15 min).
On the wall, cluster near‑duplicates; split into two piles: structural (physics/policy; slow to change) and situational (regime‑specific; faster to update). Park “nice to know” items.

4) Make testable (20 min).
Rewrite top candidates in the grammar above. Add scenario tag, horizon, measurement, and a falsification threshold. Example:
“We believe unit demand in S‑02 will recover to 95–105 index within 2 quarters because backlog + channel fill; falsify if 3‑month moving average stays <90 from POS panel X by 30‑Jun.”

5) Score & triage (10 min).
Score each on Impact × Uncertainty × Observability (1–3 each). Impact = effect on VaS/headroom or breach probability; Uncertainty = current spread; Observability = how fast/cheap you can test. Multiply; sort.

6) Form the shortlist (10 min).
Select 6–8 highest composite scores. For each, state the decision impact (“If [A‑014] fails, we expire O‑04; if passes, we raise tranche 2”). Everything else goes to the backlog with a next review date.

7) Wire to actions (10 min).
For each shortlisted [A‑###]:

  • Ranges: set the parameter band in TPL‑12 and tag [MB‑##].
  • Signpost candidate: if the falsification threshold is externally observable and timely, draft an indicator spec (13.1) and a trigger rule (12.4).
  • Learning Gate: if evidence requires a test/pilot, create an LG with pass/fail thresholds, budget, and owner.

8) Log & publish (10 min).
Create or update TPL‑10 cards; publish the Shortlist index at the top of the file with owners, thresholds, and due dates. PMO stamps versions.

C. TPL‑10 Assumption Card (merged with shortlist) — fields to keep

  • ID & Title: [A‑###] short name.
  • Scenario & Horizon: S‑##; Monitoring vs. Evaluation.
  • Statement (grammar‑compliant).
  • Mechanism: behavior/constraint causing it.
  • Falsification test: observable, threshold, window, data source/cadence.
  • Decision impact: option/hedge/gate/trigger that changes if false/true.
  • Links: [DRV‑###], [MB‑##], TRG‑### (if candidate), LG (if required).
  • Owner & Challenger: names; due date.
  • Status: Open / Accepted / Rejected / Superseded; last update.
  • Evidence notes: [CIT‑###] snapshots; confidence tag.

Keep cards to one page; retire or supersede—don’t edit history.

D. Prioritization rules that prevent bloat

  • Prefer high‑impact, observable assumptions over esoteric but interesting ones.
  • If an assumption is structural and hard to test quickly (e.g., long‑run elasticity), treat it as a band in TPL‑12, not a trigger.
  • If an assumption is situational and observable (e.g., enforcement intensity, supplier yield), consider promoting it to a signpost with a numeric trigger.
  • Kill duplications: if two cards share mechanism and test, merge them and expand the threshold notes instead. 

E. Quick checklist (use before Gate 3)

  • 6–8 shortlisted assumptions, each with scenario tag, horizon, mechanism, falsification test, and decision impact.
  • At least two have an attached Learning Gate with pass/fail evidence and budget.
  • At least two have a drafted indicator spec suitable for promotion to a trigger.
  • Ranges updated for parameters tied to the shortlist; [MB‑##] footnotes present.
  • All cards live in TPL‑10 with owners, due dates, and evidence links [CIT‑###].

F. Anti‑patterns—and the counter‑move

  • Vague beliefs. (“Customers will like…”). Counter: force the grammar; add unit, horizon, threshold.
  • Unowned cards. Counter: no owner, no card. Assign both Owner and Challenger.
  • Endless lists. Counter: cap at 6–8; demote the rest with a date to revisit.
  • Falsification by opinion. Counter: specify observables from a named source with cadence.
  • Trigger drift. Turning assumptions into color‑coded dashboards. Counter: promote only via 13.1/12.4 with numeric rules, windows, and unwinds.

Run this sequence and your assumptions stop being ambient beliefs. They become testable bets—each with a threshold, an owner, and a clear consequence for the portfolio—so learning moves capital instead of filling notebooks.

8.2 Debiasing Toolkit + Signal Hygiene

Bias isn’t a character flaw; it’s a system property. The cure is not more debate—it’s mechanics that force outside views, surface disconfirming facts, and keep signals clean enough to drive triggers. This section combines in‑room debiasers with signal hygiene controls in the data pipeline so what you think and what you measure both withstand stress.

A. Operating stance (read aloud before analysis)

  • Decision‑backward. Every analysis and signal must state the decision impact (which option, trigger, gate, or limit could change).
  • One Challenger. Name an independent Challenger with guaranteed airtime and escalation rights.
  • Outside view first. Begin with a reference class/base rate before any bespoke forecast.
  • Traceability or park it. Claims carry [A‑###]/[DRV‑###]; numbers carry [MB‑##]/[CIT‑###] with retrieval dates.

B. In‑room debiasing—short, mechanical, and logged

Run these in every Gate workshop; they take minutes and change outcomes.

  • Silent pre‑commit (3 minutes). Everyone writes their choice and the one assumption that would flip them. Prevents anchoring and groupthink; log to TPL‑10.
  • Premortem micro‑drill (7 minutes). “It’s 12 months later; the decision failed. List three concrete causes.” Cluster causes into assumptions to test and mitigations; create/upgrade cards in TPL‑10 with pass/fail thresholds.
  • Red Team first (5 minutes). Challenger presents the best contrary case before the sponsor speaks; owners respond in one minute each.
  • Devil’s advocate / two‑minute kill. Assign one person to argue to expire or kill the favored option; if no one can, you’re in sunk‑cost drift—tighten kill criteria.
  • Blind scoring. Score options/signals on impact, uncertainty, actionability without source names visible; reveal sources only after scores are logged.
  • Threshold pre‑commit. Write the trigger rule before looking at last month’s line: indicator, threshold, window, hysteresis, and unwind (see 12.4/13.1). Adjust only with a cost‑of‑error note.
  • Last‑to‑speak rule. Sponsor and Decision Owner ask questions early but deliver views last, after Red Team and FP&A.

Log which debiasers you used in the Decision Record (TPL‑25). 

C. Signal hygiene—controls in the pipeline (not in slides)

Most “bias” arrives via dirty signals. Clean them at source and encode rules where they run—TPL‑15I (indicator spec), TPL‑15R (trigger rule), and the dashboard (TPL‑16).

1) Source tiers & quotas (selection bias control).

  • Tier 1: primary, auditable (statistical releases, filings, price feeds, internal systems).
  • Tier 2: vetted secondary (major research, audited vendor data).
  • Tier 3: commentary/estimates (press, blogs, social).
    Rule: An indicator must rely on 2× Tier 1 or 1× Tier 1 + 1× Tier 2; Tier 3 is optional and never alone. Quotas live in TPL‑15I.

2) Recency windows (stale data control).
Define time‑to‑live per source: e.g., vendor quotes ≤30 days, regulatory dockets ≤90 days, macro series next scheduled release ±5 days. The window sits in TPL‑15I and enforces an “expired” state on tiles.

3) De‑dup & gaming defense (correlation + adversarial control).

  • Collapse near‑duplicate feeds (same construct) to one; document rationale.
  • Prefer derived indicators (e.g., freight bookings + dwell times) over single‑counterparty claims.
  • Use m‑of‑n confirmation (e.g., 2 of 3 breach) for high‑impact triggers.

4) Smoothing, hysteresis & windows (noise control).
Set rolling averages/medians and explicit hysteresis (“lift when ≥X for 30 days; unwind when ≤X−δ for 45 days”) in TPL‑15I/15R to avoid whipsaw.

5) Backtests to cost‑of‑error (false alarm vs. miss).
Backtest thresholds and document hit rate, false‑alarm rate, median lead time, time‑under‑water and a short cost‑of‑error rationale in TPL‑15I. Tighten rules where misses cost more than false alarms.

6) Evidence snapshots & lineage.
Store raw source files with timestamps and checksums; every tile links to a [CIT‑###] snapshot. Transform scripts are versioned; no manual edits in the viz layer.

7) Negative sampling & “non‑events.”
For every promoted indicator, store two non‑breach episodes and why they didn’t fire; this disciplines narrative hindsight.

D. Debiasing & Signal Hygiene Runbook (TPL‑17)

Use this 8‑step loop before any Gate:

  1. Decision impact statement. One line: which choice, hedge, or trigger could change.
  2. Reference class. Produce a base‑rate chart for the relevant pattern (cycle length, ramp speed, adoption, price mean‑reversion).
  3. Disconfirming search. Challenger names 3 queries and 2 sources aimed to overturn the favored view; results logged to TPL‑10 with confidence.
  4. Signal intake. For each candidate indicator, fill TPL‑18 Signal Intake Card: definition, tier mix, recency, smoothing/hysteresis proposal, gaming risk, and draft threshold.
  5. Blind scoring. Team scores options/signals; FP&A overlays ranges impact; Risk estimates cost‑of‑error.
  6. Write the rule. Promote to TPL‑15R with numeric threshold, window, hysteresis, unwind, owner, SLA, and budget reference.
  7. Dry‑run. Replay last year; attach evidence and the hit/false‑alarm summary to TPL‑15I.
  8. Log debiasers used. Note in TPL‑25 which tools ran; capture any threshold changes with rationale.

E. Quick checklist (use before Gate 3/4)

  • A reference class slide appears before bespoke numbers.
  • At least one assumption was added or tightened due to the premortem/Red Team.
  • Every promoted indicator meets tier quotas, has a recency window, and shows hysteresis.
  • Thresholds carry a backtest + cost‑of‑error note and a dry‑run result.
  • Trigger rules include unwind and an SLA; owners are named.
  • Decision Record lists which debiasers ran. 

F. Anti‑patterns & counter‑moves

  • Probability theater. Weighting across scenarios to look precise. Counter: work within scenarios; use coverage/regret across worlds.
  • Expert monoculture. One vendor/expert shapes the view. Counter: enforce tier quotas and disconfirming searches.
  • Colored triggers. Red/amber/green without numbers. Counter: rewrite into TPL‑15R with thresholds and windows.
  • Hindsight creep. Thresholds tuned to last quarter. Counter: lock pre‑commit values; adjust only with cost‑of‑error notes.
  • Signal whipsaw. No hysteresis → action flip‑flops. Counter: add δ‑bands and cooling windows.

Run this toolkit and the program becomes self‑correcting: diverse evidence in, dirty signals cleaned, thresholds written like policies, and decisions debias‑by‑design—fast enough to matter, rigorous enough to defend.

8.3 Assumption Log & Traceability—Template

Assumptions are the hidden load‑bearing beams of your scenarios. If you don’t name them, date them, and link them to evidence and actions, they harden into dogma and your triggers won’t fire until it’s costly. The Assumptions & Coherence Log (TPL‑10) is the single source of truth for what must be true, why we believe it, how we could be wrong, and what happens if the belief fails. Treat it as an operational ledger—not a parking lot.

Purpose of TPL‑10

  • Make every belief falsifiable, owned, and dated.
  • Tie beliefs to decisions (alternatives, options, hedges) and to mechanics (model blocks, signposts, gates).
  • Provide an audit trail from statement → evidence → challenge → action.
  • Enable fast updates to ranges and triggers when reality moves.

TPL‑10 entry — fields to capture (one page per assumption)

  • ID and title. Stable code (e.g., A‑017), short name.
  • Plain‑language statement (falsifiable and dated). “If X, then Y by <date>.”
  • Category. Causal / Behavioral / Constraint / Timing.
  • Decision link. Which decision or alternative this affects; value‑at‑stake if wrong (order‑of‑magnitude).
  • Observable & threshold. Metric, source, refresh cadence; the number/event that would disconfirm.
  • Time window. Monitoring horizon and any natural expiry.
  • Evidence pack. Links to driver entry (TPL‑05), source docs/datasets, expert notes.
  • Disconfirming evidence. At least one contrary source or query to run next.
  • Confidence label. High/Medium/Low with a one‑line rationale.
  • Model hooks. Variables and model block references (TPL‑11/12) this drives.
  • Scenario mapping. Which scenarios assume it; flip conditions that reassign membership.
  • Action mapping. Options/hedges/triggers affected (TPL‑14/15) and the if/then rule.
  • Challenge plan. Premortem/Red‑Team/experiment; owner and due date.
  • Status & lifecycle dates. Proposed / Accepted / Under Test / Rejected / Superseded; created/next review/closed.
  • Owner & Challenger. Named individuals with response SLAs.
  • Change log. Date, change, rationale, initials.

Keep each entry to one page. If you need more, the statement isn’t sharp or the evidence isn’t curated.

Lifecycle and governance

  • Proposed → Accepted. Created during framing or architecture; accepted after a quick evidence check and Challenger review.
  • Accepted → Under Test. Elevated when decision‑critical; attach a Learning Gate with pass/fail threshold and date.
  • Under Test → Rejected/Superseded. Rejected if disconfirmed; superseded when a better, more precise statement replaces it.
  • Dormant cleanup. Entries with expired windows are closed or rewritten; no “evergreen” assumptions.

Status changes are gate events: the PMO notes them in the sprint log; FP&A updates ranges; Strategy updates options and triggers.

Traceability conventions (make them mechanical)

  • Anchor tags in decks. Any claim or number in executive materials carries a small tag: [A‑017] for assumption, [DRV‑003] for driver, [MB‑12] for model block. The final page lists tags and owners.
  • Footnote lineage. Every number footnotes to a model block and date; every assertion footnotes to a TPL‑10 entry.
  • Deck hygiene. If a slide has an untagged assertion or number, it doesn’t ship.
  • One model. All quantitative references pass through the shared model; no parallel math.

Step‑by‑step: standing up and running TPL‑10

  1. Harvest and normalize. After framing (3.1), run a 45‑minute harvest across causal/behavioral/constraint/timing. Rewrite each in falsifiable form with a threshold and date.
  2. Rank for decision criticality. Score Impact if wrong (H/M/L) and Reversibility (Hard/Moderate/Easy). Prioritize High + Hard for challenge and Learning Gates.
  3. Wire to actions. For prioritized entries, specify the if/then rule: “If A‑017 fails, let Option O‑05 expire and lift Hedge H‑02 to cap exposure at $X.”
  4. Assign Owner and Challenger. Owner maintains evidence and proposes updates; Challenger files disconfirming items and can escalate.
  5. Install SLAs. Written responses to Challenger notes in 48 hours; trigger breaches acted on within 10 business days (or stricter per register).
  6. Publish and police. PMO publishes a weekly extract: new/changed assumptions, upcoming tests, and items due this week. Unowned or overdue entries escalate to the Sponsor.
  7. Refresh and close. At the quarterly refresh, close or supersede entries; recalibrate confidence and thresholds based on range‑to‑reality.

Quality bar for entries

  • Short and dated. If it can’t fit in two sentences with a date, it isn’t testable.
  • Measurable. Threshold must be observable from a source you actually have or can stand up within two weeks.
  • Decision‑linked. The entry names a decision, option, hedge, or trigger that would change.
  • Challenged. A contrary source or query is logged; silence is not acceptable.
  • Owned. Both Owner and Challenger are named with the next review date.

Anti‑patterns—and the fix

  • Essay entries. Replace narratives with a falsifiable statement and an observable threshold.
  • “It depends.” If the threshold can’t be written, demote to context or reframe the driver.
  • Orphans. Any assumption without a decision or model hook is removed.
  • Zombie entries. Expired windows that linger; PMO runs a monthly purge.
  • Shadow math. Numbers that don’t footnote to the shared model; gate is paused until fixed.
  • One‑sided evidence. Entries without disconfirming sources are sent back.

Acceptance checklist (five minutes before Gate 3)

  • Top assumptions are falsifiable, dated, and labeled High/Med/Low with reversibility noted.
  • Each has an observable and a source; at least one disconfirming item is logged.
  • If/then links to options, hedges, or triggers are written and budgeted.
  • Model blocks and scenario membership are referenced; FP&A can show the delta if the assumption flips.
  • Owner and Challenger have responded to open notes; overdue items escalated.
  • Deck exhibits carry [A‑###] tags; every number footnotes to a model block and date.

Run TPL‑10 this way and you will replace confident storytelling with testable beliefs and fast updates. When a signpost moves or a test fails, the log drives the change—ranges update, options are exercised or killed, hedges adjust—because the traceability has already done the political work.

Scenario Planning Playbook

Request the Scenario Planning Playbook

How to get started

1

arrow-down-blue

Tell us about your project

2

arrow-down-blue

Interview candidates

(We’ll provide bios within 48 hours on average)

3

Select your consultant and start work

Find a Consultant

or email us at: [email protected]