Technology changes what is possible. Culture and change management determine what actually happens. In every AI program that sustains impact, the decisive moments are not in model training or platform builds—they’re in how leaders tell the story, how managers coach the first week after going‑live, how frontline teams learn a new workflow, and how quickly feedback turns into improvements. This chapter focuses on the human system that converts AI design into durable behavior change: narrative, trust, incentives, adoption instrumentation, training, job redesign, and the rituals that keep learning continuous.
Two principles guide the assessment. First, value is realized only when users adopt a different way of working; therefore adoption is part of “definition of done,” not an afterthought. Second, change must be run with the same operational rigor as engineering: owners, SLAs, stage‑gates, telemetry, and drills. When those principles hold, you see faster time to competence, fewer rollbacks driven by social risk, and steady gains in unit economics as teams use AI safely and well. When they don’t, you see “launch and hope,” piloting forever, and a trust deficit that is costly to reverse.
10.1. Diagnostic Criteria
This domain evaluates Culture and Change‑Management Readiness across fifteen sub‑dimensions. Each criterion states “what must be true,” anchored in evidence you can produce now, with measurable thresholds and gating rules where appropriate.
Executive narrative and psychological contract
Leaders have a clear, consistent AI narrative that connects to strategy, names non‑goals, and describes job impact with specificity (what will change, what won’t, and how people will be supported). The narrative includes a plain‑English responsible‑use posture. Leaders show up in the first releases (not just the kickoff) and celebrate adoption wins and safe‑failure learnings. Evidence includes a signed two‑page narrative, artifacts from leader engagements, and examples where scope was shaped by the narrative (e.g., pausing an external GenAI feature until grounding and safety cleared).
Accountability for adoption and value
Adoption and behavior change are owned by the business, not outsourced to “communications.” Each value case has an adoption owner, explicit targets, and acceptance tests in the stage‑gates. Operating reviews track usage, time‑to‑competence, and value realization alongside reliability and cost. Evidence includes the portfolio register with adoption owners, dashboards with targets vs. actuals, and decision logs where funding hinged on adoption evidence.
Change governance and ownership
A named change lead sits in the delivery triad for each priority use case. A cross‑functional change forum—linked to the weekly portfolio forum—manages the change calendar, message cadence, enablement, and feedback. Separation of duties is preserved: product ships features; change ships behavior. Evidence includes the charter, cadence, change calendar, and recent escalations resolved within SLA.
Change network and manager enablement
A distributed network of champions covers ≥80% of the impacted org by headcount, with time allocated, toolkits, and a clear job to do (demonstrate, collect feedback, escalate issues). Frontline managers are prepared to coach specific tasks—not just “be supportive”—and have a manager kit (talk tracks, metrics, troubleshooting). Evidence includes the champion roster with coverage analytics, manager kit, and participation/completion rates.
Adoption instrumentation and measurement
Instrumentation for adoption exists before build completes. For each value case: target population and cohorts; definitions for active use and time‑to‑competence; task‑quality measures; and attribution logic that links adoption to operating KPIs. Dashboards segment usage by role, site, shift, and language; feedback channels are linked in‑app. Evidence includes the measurement plan (Chapter 5), live dashboards, and examples where telemetry triggered design or training changes.
Training and enablement
Role‑based training is hands‑on and scenario‑led (not slideware). It includes secure data handling, responsible‑AI basics, the new workflow, and “what good looks like” in the first week. In‑product guidance and job aids are available, localized, and maintained. People have time to learn (protected hours) and managers track completions and proficiency gains. Evidence includes curricula, completion dashboards, sandbox exercises, and adoption curves by cohort.
Job, process, and control redesign
Workflows, SOPs, and controls are updated before scale. Human‑in‑the‑loop boundaries are explicit; escalation paths and error taxonomies exist; quality checks are integrated. For GenAI, grounding, citation, and safety thresholds are reflected in SOPs and reviewer tools. Evidence includes revised SOPs, change logs, tool UX with review flows, and incident playbooks used in the last 90 days.
Stakeholder, labor, and legal engagement
Where works councils, unions, or sector regulators require consultation, plans exist and are executed early. Employee data use is transparent; monitoring uses are governed and communicated; cross‑border and accessibility needs are addressed. Evidence includes consultation timelines, summaries of feedback and how it shaped design, DPIA notes, and accessibility reviews.
Communications and storytelling
A simple, consistent information architecture exists: who says what, to whom, how often, and where the source of truth lives. Messages are two‑way (AMAs, office hours, in‑app Q&A) with response SLAs. Release notes are written for users, not engineers. Evidence includes the comms plan, open/read rates, Q&A logs with response times, and examples where myths were corrected publicly.
Trust, ethics, and psychological safety
People know how to raise concerns without retaliation, and see action taken. Safety incidents (including hallucination‑driven missteps) are acknowledged, fixed, and shared as learning. Responsible‑AI guidelines are visible in the tools and training—not just on an intranet page. Evidence includes pulse‑survey results, incident communications, and links to guidelines embedded in product surfaces.
Change portfolio hygiene and fatigue management
Concurrency limits exist. The change calendar shows who is impacted by what, when, with “stop doing” items that create capacity. Peak periods (season, quarter close) are protected. Evidence includes a live change calendar, WIP limits, and examples of deferrals to protect frontline bandwidth.
Incentives and performance alignment
Performance goals and recognition reinforce adoption, SLO adherence, cost‑aware operation, and safe behavior—not activity volume. Sales comp and service metrics are updated if AI changes how value is delivered. Evidence includes OKRs/scorecards and recognition artifacts tied to these outcomes.
Experimentation and learning culture
Teams work hypothesis‑first; quick A/Bs and staged rollouts are normal; post‑incident and post‑launch reviews produce changes in process or product within two sprints. Learning is shared across teams via templates and show‑and‑tell rituals. Evidence includes experiment logs, decision memos, and repository analytics showing template reuse.
Inclusion, accessibility, and localization
Materials and tools meet accessibility guidelines; languages and shifts are supported; frontline devices and connectivity realities are considered; user groups at risk of exclusion have tailored support. Evidence includes accessibility checks, localization coverage, device compatibility tests, and adoption by shift/site/language.
Customer and frontline co‑design
Users and customers are co‑authors of solutions. Pilots have named sites and owners; feedback is structured; “nothing about me without me” is visible in decisions. Evidence includes pilot charters, participant rosters, and design changes driven by frontline/customer input.
Measurable thresholds that indicate “ready”
Tune these to context; set the target before scoring and hold to it.
- Ownership and governance
- 100% of priority value cases have a named adoption owner, change lead, and champion coverage ≥80% of impacted headcount.
- A live change calendar exists; concurrency limits are defined and enforced for impacted cohorts.
- Adoption and time‑to‑competence
- For each value case, adoption targets are defined by cohort (e.g., ≥70% weekly active among target users within 90 days of launch; median time‑to‑competence ≤14 days).
- Task‑quality and human‑in‑the‑loop adherence meet thresholds for two consecutive review cycles before scale.
- Training and enablement
- ≥85% of impacted users complete the role‑based training pre‑go‑live; ≥95% within 30 days post‑go‑live.
- 100% of manager cohorts complete the manager kit; office hours/AMA cadence is on calendar with ≥60% attendance or views.
- Instrumentation and value linkage
- 100% of priority value cases have in‑product telemetry for adoption, task quality, and feedback; dashboards segment by role/site/language.
- Value attribution logic exists and is used in operating reviews (leading indicators and lagging benefits).
- Job and control redesign
- 100% of affected SOPs updated and communicated before scale; escalation and error taxonomies documented and tested.
- For GenAI, grounding/citation and content‑safety thresholds are embedded in workflows and reviewer tools.
- Trust and inclusion
- Pulse‑survey response rate ≥60% in impacted cohorts; psychological safety index at or above baseline by +5 points within 90 days of launch.
- Accessibility and localization coverage for ≥95% of impacted users; zero critical accessibility defects outstanding at go‑live.
Domain‑specific gating rules (apply caps if any fail)
If any of the following fail for in‑scope use cases, cap Culture & Change‑Management at 2.0 (per Chapter 3) and flag on the heat map.
- No named adoption owner or change lead for a priority value case.
- No adoption telemetry instrumented (active use, cohort coverage, time‑to‑competence) before go‑live.
- No updated SOPs/human‑in‑the‑loop boundaries before scale, or no escalation path documented.
- No role‑based training plan and manager enablement kit, or training time not protected.
- For GenAI: no grounding/citation and content‑safety thresholds embedded in workflow before external exposure.
- Works‑council/union/regulatory consultation required but not planned or executed.
Minimum evidence required to score confidently
- Two‑page executive AI narrative with non‑goals and job‑impact commitments; recordings or notes from leader engagements.
- Change forum charter; change calendar with concurrency/WIP; roster of champions with coverage analytics.
- Adoption measurement plans and live dashboards (usage, time‑to‑competence, task quality, feedback), segmented by cohort.
- Training curricula, completion dashboards, sandbox exercises, and in‑product guidance/jobaids; manager kits and attendance metrics.
- Updated SOPs, human‑in‑the‑loop boundaries, escalation paths, error taxonomy; incident playbooks and recent PIRs.
- Consultation plans and artifacts (works councils/unions/regulators) where applicable; DPIA decisions and accessibility checks.
- Communications plan and cadence; release notes; AMA/Q&A logs with response SLAs; myth‑busting examples.
- Pulse‑survey snapshots for impacted cohorts; inclusion/localization coverage; adoption vs. value linkage in operating reviews.
Anti‑patterns to watch for
- “Launch and hope”: no adoption targets, no telemetry, no manager kit.
- Change theater: town halls and posters without protected learning time or job redesign.
- Pilot purgatory: perpetual “beta” with no stage‑gates, no decision to scale or stop.
- Competing initiatives: no change calendar; flooding the same cohort with multiple changes in the same week.
- AI trust erosion: shipping ungrounded external content; monitoring uses communicated poorly or not at all.
- Training as compliance: slide completions tracked, skills not acquired; no sandbox, no scenario‑based practice.
- English‑only: no localization or accessibility for large parts of the workforce; night shift ignored.
- Side‑door workarounds: managers encourage “the old way” to hit volume targets; incentives misaligned.
Signals of excellence (what “great” looks like)
- Adoption is visible and owned: operating reviews show usage, time‑to‑competence, task quality, and value in the same pack as reliability and cost.
- First‑week rituals: leaders visit pilot sites; managers run daily huddles around the new workflow; champions collect feedback and close loops fast.
- Learning loops: changes to prompts, retrieval corpora, or UI copy roll out weekly based on telemetry and frontline input; release notes are written for users.
- Trust by design: responsible‑use guardrails are explained in the product; incidents are acknowledged and used to improve SOPs and tests.
- Inclusion is muscle memory: materials are localized; accessibility checks are standard; time zones and devices are considered.
- Fatigue is managed: a live change calendar and concurrency limits exist; “stop doing” items create space for the new way of working.
- Recognition reinforces outcomes: teams are celebrated for adoption, safe operation, and reuse—not just volume shipped.
When these criteria are satisfied with traceable evidence—and the gating rules all read “Yes”—you can scale AI with confidence that people, not just platforms, are ready. In the next sections, you will find the step‑by‑step assessment, the culture levers and playbooks to move the metrics that matter, and the checklists and templates that make change management a high‑velocity, repeatable capability.
10.2. Culture Assessment – Step-by-Step Guide
This guide gives you a practical, time‑boxed way to baseline cultural readiness, prove adoption mechanics before launch, and publish a 90‑day plan that moves behaviors—not just slides. Treat it like an engineering assessment: observable evidence, pass/fail acceptance tests, and dated fixes. You can complete the core run in 10 business days for a business unit; extend to 15 for an enterprise sweep.
Step 1 — Set scope, outcomes, and the clock
Write down exactly which value cases are in scope, the populations they touch, and the business outcomes you must enable in the next 90 days (adoption rates, time‑to‑competence, task quality, and value realization). Map cohorts (roles, sites, shifts, languages) and note any protected periods (seasonal peaks, quarter close). Time‑box the assessment.
Acceptance tests: a one‑page charter lists value cases, cohorts, outcomes, and the end date; the sponsor signs; non‑goals are explicit.
Step 2 — Assemble the culture evidence pack
Collect live artifacts that show how change is led, measured, and learned from today: executive narrative (with non‑goals and job‑impact specifics), change forum charter and cadence, change calendar with concurrency/WIP limits, champion roster with coverage analytics, manager kit, training curricula and completion dashboards, sandbox exercises, adoption dashboards or event schemas, SOPs and HITL boundaries, incident playbooks, consultation records (works councils/unions/regulators), DPIA notes (if relevant), accessibility/localization checks, communications plan and release notes, AMA/Q&A logs, pulse‑survey snapshots.
Acceptance tests: each artifact has an owner and last‑updated date; gaps are logged with owners and due dates.
Step 3 — Stakeholder map and impact analysis
For each value case, map who is affected and how their day changes. Use “jobs‑to‑be‑done” language: what steps get removed, added, or moved? Identify skeptics, influencers, and impacted metrics (volume, quality, compliance).
Acceptance tests: a one‑page stakeholder/impact map exists per value case; it names cohorts, pain points, and the two or three behaviors that must change.
Step 4 — Adoption measurement plan (design before build)
Verify that adoption is measurable: definitions for target population, active use, time‑to‑competence, task‑quality signals, HITL adherence, and feedback channels. Confirm your attribution logic (how adoption ties to operating KPIs) and the event schema you’ll instrument.
Acceptance tests: a signed measurement plan exists per value case with targets by cohort (e.g., ≥70% weekly active within 90 days; median time‑to‑competence ≤14 days); dashboards or placeholders are linked.
Step 5 — Telemetry and privacy dry‑run
Run a telemetry fire drill in a lower environment: emit adoption events, task‑quality signals, and HITL flags; verify they land in dashboards. Confirm consent language and DPIA decisions (if required), opt‑outs, and redaction for prompts/outputs in GenAI flows.
Acceptance tests: a screenshot or log shows events received; privacy approvals are recorded; feedback links appear in‑app.
Step 6 — Change governance and cadence
Review the change forum: membership, quorum, decision rights, escalation ladder, and SLAs. Inspect the change calendar: concurrency limits by cohort, protected periods, and “stop‑doing” items that free capacity.
Acceptance tests: the forum is on calendars; the last two escalations were resolved within SLA; concurrency limits are enforced (evidence: a deferral to protect the frontline).
Step 7 — Champion network coverage and activation
Confirm a distributed champion network covers ≥80% of impacted headcount with time allocation, a playbook (demo, observe, escalate), and a communication loop into product teams.
Acceptance tests: coverage analytics show ≥80%; champions have a monthly cadence and log at least one feedback item each in the last 30 days.
Step 8 — Manager enablement
Open the manager kit: talk tracks, “first‑week coaching” checklist, troubleshooting guide, metrics to watch, and escalation contacts. Check completion and participation.
Acceptance tests: 100% of impacted manager cohorts have completed the kit; office hours exist with ≥60% attendance or views.
Step 9 — Training and enablement readiness
Evaluate role‑based, scenario‑led training: sandbox exercises, in‑product guidance, job aids, localization, and protected learning time. Confirm that training addresses secure data handling, responsible AI, and the new workflow.
Acceptance tests: ≥85% of impacted users have pre‑go‑live training scheduled; materials are localized; job aids are published in the product.
Step 10 — Job, process, and control redesign
Review updated SOPs and controls for the new way of working. HITL boundaries, error taxonomy, sampling, and escalation paths must be explicit. For GenAI, check grounding/citation standards in reviewer tools and content‑safety thresholds wired into workflow.
Acceptance tests: SOPs are versioned and published; a tabletop exercise walks a real case through the new escalation path; at least one reviewer completes a mock review using the tooling.
Step 11 — Communications and story cadence
Check the comms plan: message map (what, who, when, where), source of truth, release notes written in user language, AMAs with response SLAs, and myth‑busting mechanisms.
Acceptance tests: the last release note is user‑centric; Q&A items show response within SLA; a myth was corrected publicly in the last quarter.
Step 12 — Trust, ethics, and psychological safety
Confirm how responsible‑use commitments show up in the product and training (e.g., what AI can/can’t do, data handling, when to escalate). Review incident playbooks and communications patterns (acknowledge, fix, learn).
Acceptance tests: guidance appears in‑product; a recent incident or drill resulted in a visible change to SOPs or tests; pulse scores show stable or improving psychological safety in impacted cohorts.
Step 13 — Labor, regulator, and legal engagement
If relevant, verify consultation timelines, materials, and how feedback shaped design (works councils, unions, sector regulators).
Acceptance tests: meeting notes exist; changes tied to feedback are cited; any required approvals are on file.
Step 14 — Inclusion, accessibility, and localization
Test the experience on real devices, languages, and shifts. Review accessibility checks against policy, and confirm alternate channels for low‑connectivity settings.
Acceptance tests: ≥95% of impacted users have language coverage; zero critical accessibility defects remain at go‑live; a recorded test shows the workflow works on frontline devices.
Step 15 — Pilot design and first‑week rituals
Define pilot sites, sample sizes, and staged rollout. Specify first‑week rituals: leader site visits, daily huddles, champion shadowing, and feedback loops that drive changes within a week.
Acceptance tests: a pilot charter exists with success criteria and exit decisions (scale/iterate/stop); ritual owners and dates are on calendars.
Step 16 — Readiness rehearsal
Run a “day‑in‑the‑life” rehearsal with a real cohort: training, go‑live, coaching, escalation, and comms. Include service‑desk scripts and handoffs.
Acceptance tests: a dry‑run record exists; at least one fix is fed back into training, SOPs, or product copy.
Step 17 — Post‑launch operating rhythm
Confirm the cadence for adoption reviews (weekly for first month), dashboard owners, backlog integration of feedback, and how release notes and training get updated.
Acceptance tests: calendar invites exist; the last change driven by adoption telemetry is visible in release notes.
Step 18 — Synthesize, score, and cap
Using Chapter 3’s rubric and the gating rules in 10.1, score each sub‑dimension with confidence tags; cite evidence. Apply caps if any gating item fails (no adoption owner, no telemetry, no SOP updates/HITL boundaries, no training/manager kit, missing GenAI safety/grounding, missed consultations).
Acceptance tests: a scored workbook with links exists; caps are explicit; top five cultural constraints are quantified for value at risk.
Step 19 — Publish the 90‑day culture and change plan
Turn constraints into funded, dated actions with pass/fail tests. Typical moves: instrument adoption telemetry; expand champion coverage; run manager bootcamps; localize materials; update SOPs and embed HITL tooling; add grounding/citation and safety thresholds to workflows; schedule leader rituals; enforce WIP limits on cohorts; launch feedback‑to‑backlog integration.
Acceptance tests: each action has an owner, budget, and acceptance test; stage‑gates reference these conditions.
Step 20 — Stand up the weekly adoption scorecard
Track a small, durable set of metrics: adoption by cohort; time‑to‑competence; task‑quality and HITL adherence; training and manager‑kit completion; champion activity; comms engagement and response SLAs; ticket volume and time‑to‑resolution; pulse‑survey safety; value linkage (leading and lagging).
Acceptance tests: the scorecard appears in the operating review with named owners; red items generate actions in the decision log.
Probes you can run this week (lightweight, high signal)
- Event fire drill: trigger adoption events in lower environments; ensure all tiles populate in the dashboard.
- Champion coverage test: pick three sites/shifts and show live champion contacts and time allocations.
- Manager Q&A reality check: sample the last 10 AMA items; compute response time; verify clarity and consistency with the narrative.
- SOP recall test: ask five frontline users to explain the escalation path; verify the same answer and the same place to find it.
- GenAI safety & grounding check: run 50 representative prompts; confirm grounded answers, citations, and safety pass rates meet thresholds; show where overrides are logged.
- Accessibility lap: complete the workflow in each target language and on a frontline device; record any friction.
- Concurrency limit test: show one example where a change was deferred to protect a cohort and how that decision was communicated.
Fast‑track plan (5 business days)
- Day 1: Scope, evidence pack, stakeholder/impact maps.
- Day 2: Measurement plan and telemetry dry‑run; change forum and calendar check.
- Day 3: Champion coverage and manager enablement; training and SOP/HITL review.
- Day 4: Comms, trust/ethics, accessibility/localization, consultation status.
- Day 5: Scores with caps; 90‑day plan; adoption scorecard; executive readout.
Deep‑dive (15 business days)
Add cohort‑level experiments (A/B on training or UI copy), qualitative fieldwork at pilot sites, red‑team scenarios for social risk (misuse, over‑reliance), compensation alignment review (if AI affects incentives), and a board‑ready “people risk” statement with mitigations.
10.3. Culture Indicators Checklist
This checklist turns “soft stuff” into hard signals you can review every week. Use it to decide if people are actually adopting the new way of working, if trust is rising or eroding, and if you’re ready to scale without whiplash. Every line is binary—either you can point to live evidence or you have a dated action to close the gap.
How to use this checklist
- Run it at two levels: per value case (pre‑flight and first 90 days) and program‑wide (weekly).
- Set targets up front (by cohort) and keep a single source of truth for each metric.
- When any gating item fails, stop the external scale for that use case and fund the fix first.
Sources of truth you should link
- Product telemetry (adoption, time‑to‑competence, task quality, HITL flags, feedback).
- LMS and HRIS (training completion, time allocation).
- Service desk and incident system (tickets, time‑to‑resolution, post‑incident actions).
- Change calendar and portfolio forum notes (concurrency, deferrals).
- SOP repository and reviewer tools (versioned workflows, escalation logs).
- Communications analytics (opens, reads, Q&A SLAs), pulse surveys, and champion logs.
Pre‑launch readiness canaries (must be green before go‑live)
- Adoption owner and change lead named; champion coverage ≥80% of impacted headcount with time allocated.
- Measurement plan signed: target population, active‑use definition, time‑to‑competence, task‑quality signals, feedback channels.
- Telemetry dry‑run completed; events appear in dashboards; in‑app feedback link works.
- Role‑based training scheduled for ≥85% of impacted users; manager kit assigned to 100% of managers.
- Updated SOPs published with explicit human‑in‑the‑loop (HITL) boundaries, escalation path, and error taxonomy.
- Concurrency limits applied; at least one deferral on the change calendar to protect the cohort.
- For GenAI: grounding/citation and content‑safety thresholds embedded in reviewer tools; prompts/outputs logging confirmed.
First‑week pulse (lead indicators of stickiness)
- Day‑7 activation: ≥50% of target users active at least once; ≥30% complete a full task end‑to‑end.
- Median time‑to‑competence trending to target (default ≤14 days): three consecutive shifts with on‑target quality and cycle time without assisted help.
- Ticket volume × time‑to‑resolution within plan; top five blockers closed or mitigated within 72 hours.
- Manager participation: ≥90% of impacted managers run first‑week huddles; office hours/AMA attendance or views ≥60%.
- Champion activity: ≥1 observation and ≥1 escalated issue per champion logged; turnaround on escalations ≤5 business days.
- Trust signals: zero critical incidents undisclosed; at least one visible “we learned and changed X” note.
30/60/90‑day stickiness and scale indicators
- Weekly active ratio (WAU ÷ target population) hits the cohort target (e.g., ≥70% by day 90); 4‑week retention ≥60%.
- Depth of use: ≥60% of active users complete ≥N key tasks per week; share‑of‑work assisted by AI ≥X% in the targeted flow.
- Task quality: on‑target accuracy/quality for two consecutive review cycles; rework and escalation rates stable or down.
- HITL adherence ≥95% where required; median time from HITL flag to human decision within SLA.
- Value linkage: leading indicators (cycle time, handle time, first‑contact resolution, error rate) move ≥75% toward the modeled benefit; lagging benefits begin to accrue.
- Change fatigue: concurrency index ≤2 overlapping changes per cohort; overtime and PTO deferrals at or below baseline.
Metric definitions and how to compute them
- Target population (TP): count of people expected to use the tool in the defined period (exclude those on leave or out of role).
- Activation rate: unique users who performed at least one target action ÷ TP in the period.
- Weekly active ratio (WAU/TP): unique users completing ≥1 target action in 7 days ÷ TP.
- Depth of use: median count of target actions per active user in 7 days; track 25th percentile for tail risk.
- Time‑to‑competence (TTC): median calendar time from first login to “three consecutive shifts completing the target task within quality/time thresholds without assisted help.”
- Task‑quality score: % of tasks meeting defined quality criteria (by cohort).
- HITL adherence: required human reviews completed ÷ required human reviews triggered.
- Escalation health: escalations acknowledged within SLA ÷ total escalations; % resolved within SLA.
- Feedback closure: feedback items with a published disposition (shipped, rejected with rationale, in backlog) ÷ total feedback items.
- Change concurrency index: unique change events affecting a cohort in a rolling 14‑day window.
- Manager enablement: managers who completed the kit and ran first‑week huddles ÷ impacted managers.
- Champion coverage: impacted headcount with an assigned champion ÷ impacted headcount.
GenAI‑specific indicators (add to the core set)
- Grounded‑answer rate: % of responses that cite approved sources and pass grounding checks for the task class.
- Citation accuracy: % of citations that truly support the answer in reviewer sampling.
- Safety block rate: % of outputs intercepted by safety filters; track false‑positive rate separately.
- Override rate: % of safety blocks overridden by authorized reviewers; reasons and time‑to‑decision logged.
- Ungrounded escalation rate: % of user‑reported hallucinations; mean time to triage and correct corpus/prompt.
- Retrieval reliance: % of answers produced with retrieval (vs. pure LLM); track “retrieval missing” flags that trigger model‑only answers.
- Human‑confidence delta: change in user‑reported confidence (pre/post) that the output is safe and accurate for external use.
Qualitative signals that matter (listen for these)
- Managers coach “what good looks like this week” rather than telling people to “wait for the next version.”
- Champions volunteer real stories—both wins and safe failures—within seven days of launch.
- Frontline users can explain, in their own words, when to trust the system, when to escalate, and where to find the SOP.
- Leaders celebrate adoption, safe operation, and reuse—not just feature volume.
- Slack/Teams shifts from “How do I get access?” to “We changed X in the SOP based on Y signal.”
Red flags that require immediate action
- No adoption owner/change lead; or champion coverage <80% of impacted headcount.
- Telemetry missing or stale; no in‑app feedback channel.
- Training or manager kit completion below 70% pre‑go‑live; no protected learning time.
- SOPs not updated; HITL boundaries unclear; escalation path unknown to users.
- WAU/TP flat <40% for two weeks; TTC not improving by week 2.
- Ticket backlog aging >7 days; repeated issues with no playbook update.
- For GenAI: grounded‑answer rate below threshold; safety overrides without logging; spikes in hallucination reports.
Acceptance tests for “ready to ship” (before first pilot)
- Adoption owner and change lead named; champions cover ≥80% of impacted headcount with time allocation.
- Signed measurement plan with targets and event schema; telemetry dry‑run completed.
- Role‑based training scheduled for ≥85% of users; 100% manager kit assigned; office hours on calendar.
- SOPs and HITL boundaries published; reviewer tools configured; escalation path tested in a tabletop.
- Concurrency limits applied; at least one deferral recorded to protect the cohort.
- For GenAI: grounding/citation and safety thresholds wired into workflow; prompts/outputs logging verified.
Acceptance tests for “ready to scale” (post‑pilot)
- WAU/TP ≥70% for the target cohorts across two consecutive weeks; 4‑week retention ≥60%.
- Median TTC ≤14 days; task‑quality on target across two review cycles; HITL adherence ≥95%.
- Safety indicators stable: grounded‑answer and citation accuracy ≥ thresholds; override rate low with reasons logged.
- Value linkage visible: leading indicators ≥75% toward modeled benefit; first lagging benefits observed.
- Change fatigue under control: concurrency index ≤2; overtime and PTO deferrals at or below baseline.
- Communications and Q&A SLAs met; feedback closure ≥80% within two sprints.
30‑minute field triage (when time is tight)
- Open the adoption dashboard: read WAU/TP, depth of use, TTC trend; screenshot last 14 days.
- Pull training and manager kit completion; list top three gaps and owners.
- Check SOP version date and a live escalation log; ask three users to recite the path.
- Review champion coverage and last 10 escalations; compute response time.
- For GenAI: sample 50 prompts; record grounded‑answer rate, citation accuracy, and safety blocks; log variances vs. thresholds.
Change‑portfolio hygiene (run weekly)
- Concurrency index by cohort; deferrals in the last 7 days; “stop‑doing” items that created capacity.
- Ticket age distribution; top 5 pain points and whether the fix shipped or is dated.
- Communications: open/read, AMA backlog age, myth‑busting examples.
- Pulse: psychological safety index vs. baseline; inclusion and localization coverage by site/shift/language.
Anti‑patterns to stamp out
- Launch‑and‑leave: no first‑week rituals, no daily huddles, no visible fixes.
- Slide‑only training: no sandbox, no scenario practice, no job aids.
- Shadow workflows: managers quietly telling teams to use the old process.
- “RAG on shared drives”: uncurated corpora create trust incidents; no grounding or citation in reviewer tools.
- Metric whiplash: adding metrics without removing any; no single source of truth.
90‑day improvement levers (assign owners and dates)
- Instrument what’s missing; publish adoption and TTC targets by cohort; make them part of “definition of done.”
- Expand and activate the champion network; add office hours; launch a manager bootcamp.
- Localize materials and in‑product guidance; add frontline device testing; fix accessibility gaps.
- Update SOPs based on real incidents; tighten HITL; publish error taxonomy; integrate reviewer tools.
- For GenAI: improve retrieval (curate corpus, adjust chunking/filters), tune prompts, and raise grounded‑answer/citation accuracy; close red‑team findings.
- Enforce concurrency limits and celebrate deferrals that protect cohorts.
Use this checklist as a standing agenda for your weekly operating review. When it’s all green—and stays green for two cycles—you can scale with confidence that people, not just platforms, are ready.
10.4. Change-Management Playbook Template
Use this template to turn a launch into sustained behavior change. It is structured as a set of fill‑in‑the‑blank sections you can complete in under two weeks for any value case. Each section specifies the minimum content, the evidence to attach, and an acceptance test so you know when it is “ready to run.” Link to live systems (dashboards, repos, calendars) instead of pasting screenshots. Treat adoption and change with the same rigor you apply to shipping code.
How to use this template
- Complete Sections 1–10 before pilot go/no‑go; finish the remainder within 30 days of pilot start.
- For external or sensitive GenAI use cases, complete the GenAI addendum before any external exposure.
- If any gating item fails (see Section 22), do not scale. Record the cap in your readiness heat map and fund the fix first.
1) Header and Version Block
Establish ownership, scope, and the clock.
- Playbook name and value case(s) in scope; business owner; change lead; product lead.
- Effective dates; version; next review date and forum.
- Links: portfolio entry, decision log, adoption dashboard, change calendar.
Acceptance test: Owner and change lead are named; links resolve for all reviewers.
2) Purpose, Outcomes, and Non‑Goals
State why this change exists and what success looks like—no buzzwords.
- Business problem and user jobs‑to‑be‑done.
- Outcome targets (by cohort): adoption (WAU/TP), time‑to‑competence, task quality, value KPIs, safety adherence.
- What this playbook will not do (non‑goals and deferred scope).
Acceptance test: Targets are numeric and dated; non‑goals are explicit and referenced in scope decisions.
3) Cohorts and Impact Map
Name who is affected and how their day changes.
- Cohorts (roles, sites, shifts, languages) and size.
- Impact by cohort: steps removed/added/changed; metrics affected; risks and mitigations.
- Named champions per cohort; manager roster for coaching.
Acceptance test: Coverage map shows ≥80% of impacted headcount with an assigned champion and manager.
4) Stage Gates and Adoption Criteria
Define pass/fail tests for movement across discovery → pilot → production → scale.
- Pre‑pilot: telemetry dry‑run; training scheduled; SOP/HITL updated; consultation plan (if required).
- Exit pilot: WAU/TP ≥ target, median time‑to‑competence ≤ target, task quality on target for two cycles, safety adherence ≥95%, leading indicators ≥75% toward model.
- Scale: retention ≥ target, value realized vs. plan for N periods, fatigue under control (concurrency ≤2 per cohort).
Acceptance test: Funding and release decisions reference these gates verbatim.
5) Measurement and Telemetry Plan
Instrument adoption before you build the last feature.
- Definitions: target population, active use, depth of use, time‑to‑competence, task quality, HITL adherence, feedback.
- Event schema (names, payloads, IDs); sampling and retention; privacy/DPIA notes.
- Dashboards segmented by cohort; alert rules and owners.
Acceptance test: Events appear in dashboards from a dry‑run; owners and alert thresholds are published.
6) Training and Enablement Plan
Make it hands‑on and job‑specific.
- Role‑based curricula and sandbox drills (secure data handling, responsible AI, new workflow scenarios).
- Manager kit: talk tracks, first‑week coaching checklist, metrics to watch, escalation contacts.
- Job aids and in‑product guidance; localization and accessibility coverage; protected learning time.
Acceptance test: ≥85% of impacted users are scheduled pre‑go‑live; 100% of managers have the kit; materials are localized.
7) SOPs, Controls, and HITL Boundaries
Ship the new way of working, not just a feature.
- Updated SOPs, error taxonomy, sampling rules, escalation path, and service‑desk scripts.
- HITL boundaries (what must a human decide), reviewer tooling, and audit logs.
- Control updates (privacy, security, content safety) reflected in workflow.
Acceptance test: A tabletop exercise walks a real case through the new SOP/HITL path; logs prove execution.
8) Communications and Story Cadence
Design a simple, repeatable message architecture.
- Message map: what the change is, what it isn’t, why now, what success looks like, how to get help.
- Channels and cadence (leader notes, site huddles, AMAs, in‑app messages); source of truth.
- Release notes written for users; Q&A response SLAs; myth‑busting mechanism.
Acceptance test: A complete set of messages is scheduled on calendars; the last AMA shows responses within SLA.
9) Champion Network and Field Ops
Turn champions into a distributed change team.
- Roster, coverage analytics, and time allocation agreements.
- Playbook: demo, observe, collect feedback, escalate; monthly cadence.
- Feedback workflow into product backlog with dispositions (shipped/rejected/queued).
Acceptance test: Each champion logged at least one observation and one escalation in the last 30 days.
10) Pilot Design and First‑Week Rituals
Prototype the future and rehearse the launch.
- Pilot sites, sample size, and selection criteria; success metrics and exit decisions.
- First‑week rituals: leader visits, daily huddles, champion shadowing, “fix it fast” loop, release‑note cadence.
- Support readiness: service‑desk scripts, surge staffing, routing rules.
Acceptance test: Pilot charter signed; first‑week rituals are on calendars with named owners.
11) Inclusion, Accessibility, and Localization
Make the experience usable for everyone affected.
- Languages, reading level, and device constraints; assistive tech support.
- Offline/low‑bandwidth alternatives; shift coverage.
- Testing plan and remediation backlog.
Acceptance test: Zero critical accessibility defects before go‑live; language coverage ≥95% of impacted users.
12) Labor, Legal, and Regulator Engagement
Engage early where required.
- Works‑council/union plan and materials; regulator touchpoints (if applicable).
- Employee data use transparency; monitoring policy communication.
- Decisions and changes driven by feedback.
Acceptance test: Consultation artifacts exist; required approvals on file before pilot.
13) Trust and Responsible‑Use Guardrails
Show, don’t just say, how you keep people safe.
- Responsible‑use commitments embedded in product copy and training.
- Incident communication pattern (acknowledge, fix, learn); playbook for social risk (misuse, over‑reliance).
- Feedback and appeal channels for users.
Acceptance test: Guidance appears in‑product; a recent incident or drill changed SOPs/tests.
14) Change Calendar and Concurrency Limits
Prevent fatigue by design.
- Integrated change calendar across initiatives; cohort‑level WIP limits.
- Protected periods and deferral process; “stop‑doing” items that create space.
Acceptance test: At least one deferral recorded to protect an impacted cohort.
15) Incentives and Performance Alignment
Reward the behaviors you want.
- Updated goals (adoption, SLOs, safety, unit cost, reuse) for managers and teams.
- Sales/service metric adjustments if AI changes value delivery.
- Recognition mechanisms for safe operation and reuse.
Acceptance test: Recent recognitions and reviews cite adoption or safe operation, not just volume shipped.
16) Post‑Launch Operating Rhythm
Keep learning continuously.
- Weekly adoption review for the first month; then biweekly.
- Decision rules for when telemetry triggers design/training/SOP changes.
- Release‑note and training‑material refresh cadence.
Acceptance test: The last change driven by adoption telemetry is visible in release notes and training updates.
17) Support, Ticketing, and Escalation
Handle friction fast and visibly.
- Support tiers, routing, SLAs, and escalation ladder.
- Top‑issue playbooks; macros and knowledge base; ownership of article currency.
- Feedback loop from tickets to backlog and SOPs.
Acceptance test: Ticket age distribution within target; top 5 issues have owners and dated fixes.
18) Budget, Capacity, and Roles
Fund behavior change—not just build.
- Budget split: enablement, comms, support surge, localization, champion time.
- Capacity by week for the first 30 days; backfills for frontline time in training.
- RASCI for changing decisions.
Acceptance test: Capacity is booked; no training is scheduled without backfill or protected time.
19) Risk Register and Mitigations
Name the people’s risks and price them in.
- Top risks (adoption stall, trust incident, champion attrition, reviewer bottleneck) with owners and mitigations.
- Early warning indicators and triggers; rollback or pause rules.
Acceptance test: Risks map to actions in the 90‑day plan; triggers and responses are documented.
20) Evidence Index
Paste links only; keep everything auditable.
- Dashboards (adoption, TTC, task quality, HITL); training and manager‑kit completion; SOPs and reviewer tool configs.
- Change calendar; comms artifacts; AMA/Q&A logs; champion logs.
- Consultation artifacts; accessibility checks; ticketing views; decision log.
Acceptance test: A new auditor can reconstruct the last 14 days of change activity in <60 minutes.
21) First‑Week Cutover Runbook (T‑14 to T+14)
Turn the playbook into a timeline teams can run.
- T‑14 to T‑7: finalize SOPs and training; telemetry dry‑run; schedule huddles and leader visits; publish release notes v1.
- T‑3 to T‑1: manager briefings; champion rehearsal; service‑desk scripts live; freeze non‑essential changes for the cohort.
- T‑0: go‑live; huddle script; champion shadowing; in‑app feedback enabled; comms sent; surge support on.
- T+1 to T+7: daily adoption review; fix‑fast loop; updated release notes; highlight wins and lessons.
- T+8 to T+14: weekly adoption review; training refresh; SOP/HITL tweaks; decide scale/iterate/stop.
Acceptance test: Calendar holds each event with owner and acceptance criteria; a dry‑run completed.
22) Gating Checklist (must be “Yes” to scale)
- Adoption owner and change lead named with time allocated.
- Measurement plan signed; telemetry live; in‑app feedback works.
- Role‑based training ≥85% complete pre‑go‑live; manager kit 100% assigned.
- SOPs and HITL boundaries updated and rehearsed; escalation path tested.
- Champion network covers ≥80% of impacted headcount; activity logged.
- Concurrency limits applied; at least one deferral recorded to protect cohorts.
- Consultation completed where required; approvals on file.
- Accessibility/localization coverage ≥95% of impacted users.
- For GenAI: grounding/citation and safety thresholds embedded in workflow; prompts/outputs logged and protected.
Acceptance test: If any item is “No,” the playbook status is “Approve with conditions” and a dated fix is funded.
23) GenAI Addendum (apply when relevant)
Layer these into the core, not as a separate process.
- Corpus governance and retrieval thresholds for representative tasks; grounded‑answer and citation‑accuracy targets; safety pass‑rate targets; evaluator roles and SLAs.
- Reviewer tooling for citations and override logging; human‑in‑the‑loop sampling; red‑team cadence and closeout SLAs.
- User education: “what AI can/can’t do,” when to escalate, and how to interpret citations.
Acceptance test: Retrieval and safety metrics meet thresholds for two cycles before external exposure.
24) Templates You Can Copy
Keep these lightweight and repeatable.
Message Map (fill in 5 boxes)
- What’s changing (user‑centric).
- Why now (link to KPI/strategy).
- What stays the same (reassurance).
- What success looks like (targets and date).
- Where to get help (channel, SLA).
Manager First‑Week Huddle Script (10 minutes)
- Today’s goal and “what good looks like.”
- Yesterday’s wins + blockers; who owns fixes and by when.
- One SOP reminder; where to find the job aid.
- Safety moment (HITL or grounding example); how to escalate.
Champion Log (per shift)
- Demo given to [cohort]; # attendees.
- Observed friction (task, step, device); evidence link.
- Escalation filed (ID, owner, due date).
- One story (win or safe failure) shared with the team.
Escalation Card (users and managers)
- Trigger condition; where to log; expected SLA.
- What happens next; who is on point; how we’ll close the loop.
25) 90‑Day Improvement Plan
Convert constraints into dated, funded moves.
- Instrument missing telemetry; expand champion coverage; manager bootcamps; localize materials; fix accessibility gaps.
- Update SOPs from incidents; tighten HITL; integrate reviewer tools; publish error taxonomy.
- For GenAI: curate corpora; adjust chunking/filters; tune prompts; raise grounded‑answer and citation accuracy; close red‑team findings.
Acceptance test: Each item has an owner, budget, acceptance test, and a date on the calendar.
26) One‑Page Executive Summary (for approval)
Keep this crisp and decision‑ready.
- Why now and what outcomes (targets + dates).
- Cohorts and owners; champion coverage.
- Gating status (green/amber/red with dates).
- First‑week rituals; telemetry readiness.
- Top risks and mitigations; budget and capacity.
- Scale decision date and criteria.
Acceptance test: Executives can approve in one sitting; conditions become part of the decision log.
Complete this playbook with real owners, dates, and links to live evidence. When every acceptance test passes—and every gating question is a confident “Yes”—you have a change system that makes adoption predictable: fast starts, safe operation, continuous learning, and visible value.