Even the best system selection and global design will fail if delivery quality drifts during build and test. In most finance ERP programs, the system integrator (SI) performs a large share of configuration, integration, data work, and testing support. That can be a strength when the SI brings proven methods and disciplined execution. It becomes a risk when staffing is thin, scope expands quietly, or quality gates become optional in the rush to meet a go-live date.
Oversight is not about “policing the vendor.” It is about protecting outcomes with evidence: what is being built, how it is being tested, and whether the solution will be operable under close pressure. That requires objective assurance, structured design and build reviews, rigorous change control, business-owned testing, and fact-based performance management. When those mechanisms are in place, issues surface early and can be corrected cheaply. When they are absent, issues surface late and the program pays through rework, delayed go-lives, and unstable operations after cutover.
This chapter provides practical methods to oversee SI delivery without creating bureaucracy. The goal is simple: the right work gets done, in the right sequence, to an explicit quality standard, with proof that the business can run the process on Day 1 and close without heroics.
10.1 Independent Delivery Assurance / Quality Assurance
Independent delivery assurance (IDA): an objective, periodic assessment of program health and deliverable quality performed by a team that is not responsible for the day-to-day delivery. The word “independent” is not ceremonial. If the same people who build the solution judge whether it is ready, optimism and pressure will win. IDA exists to bring disciplined skepticism and evidence-based conclusions to leadership.
In practice, IDA should answer four questions repeatedly. Are we building the right thing (aligned to outcomes and standard design)? Are we building it well (quality, controls, and operability)? Are we building it fast enough (realistic plan, critical path integrity)? And are we setting the business up to run it (readiness, adoption, and support model)? Those questions map to the two reasons programs fail: poor design choices made early, and weak execution discipline that allows defects to accumulate late.
Scope and cadence should match risk. In early phases, IDA focuses on governance, decision rights, scope control, and design authority effectiveness. During build, the focus shifts to configuration quality, integration operability, data readiness, and defect trends. During testing and cutover, the focus shifts to readiness evidence, test coverage, reconciliation proof, and cutover rehearsal results. A common cadence is monthly deep reviews plus targeted spot checks on high-risk areas such as data conversion and close rehearsal, tightening as the program approaches go-live.
The most valuable IDA output is not a long report. It is a short set of findings that are specific, prioritized, and actionable. Each finding should include the evidence observed, the impact if unaddressed, and the recommended action with an owner and deadline. Avoid generic language like “strengthen change management.” Instead, state what is missing: role mapping incomplete for key roles, training scenarios not aligned to the future process, or local change leads not mobilized in high-volume locations.
IDA should be anchored in artifacts, not opinions. Expect the assurance team to review and sample deliverables: decision logs, design standards, interface control designs, conversion results, reconciliation scripts, test evidence, and readiness assessments. Interviews matter too, because misalignment often shows up as conflicting narratives about what “done” means.
IDA is most effective when it treats delivery as a set of controllership-critical systems rather than as a single “project.” Focus on the few areas that routinely break finance ERP programs: conversions, integrations, security and SoD, close operability, reporting certification, and business readiness. For each area, define the minimum evidence that must exist and sample it consistently over time. For example, conversion is not “on track” because mapping is written; it is on track when mock conversions reconcile with documented scripts and exceptions are trending down. Integration is not “done” because a file is transferred; it is done when control totals, alerts, and replay procedures work under failure conditions.
- Common IDA red flags: repeated defects in the same design area, mock conversions that do not reconcile, interfaces that “work” only with manual fixes, rising late change volume, and UAT that runs with unrealistic data or low business participation.
Make IDA conclusions visible in governance. The assurance team should present directly to the steering committee (or an empowered subcommittee), not only to the program team. Leaders should respond with explicit actions: accept risk with rationale, assign remediation, or change scope/timing. Use a simple rating model only if criteria are clear and evidence-based (for example, “data readiness” is green only when mock conversions reconcile above a defined threshold).
Finally, protect independence. If the assurance team starts writing test scripts, managing defects, or filling staffing gaps for the SI, it has crossed from assurance into delivery. If the program needs additional delivery capacity, resource it openly and keep IDA separate.
- IDA setup checklist: independent reporting line to steering, defined scope by phase, evidence-based criteria, monthly cadence with pre-go-live intensification, actionable findings with owners, and clear separation between assurance and delivery.
10.2 Design and Build Review Governance
Design and build reviews are the program’s quality gates. They exist to catch misalignment and defects when they are still cheap to fix, before configuration hardens and downstream testing becomes a rework factory. Reviews must be structured, role-based, and tied to acceptance criteria. Otherwise, they become optional meetings where teams present slides and continue building regardless of feedback.
Design review: a structured checkpoint where process patterns, data definitions, controls, and reporting implications are validated against global standards before build begins. The biggest design review mistake is focusing on screen walkthroughs instead of decisions. A good review examines the intent: what is the standard pattern, what variants are proposed, what controls are embedded, and how the pattern will reconcile and report at period end.
Organize design reviews around a small number of review lenses. Process lens: end-to-end steps, roles, approvals, and exception handling. Controls lens: segregation of duties, workflow approvals, audit trail, and evidence requirements. Reporting lens: required outputs, dimensional coding, hierarchies, and certification. Operability lens: monitoring, work queues, runbooks, and support ownership. Each lens should have reviewers with authority to approve or reject.
Build reviews are different. Build review: validation that configured solution components meet design intent, adhere to configuration standards, and are supportable. Build reviews should sample configured objects, interfaces, reports, workflows, and security roles, including the “nonfunctional” outcomes that determine operability: monitoring, error handling, reconciliation reports, and security logging.
Make reviews evidence-based. Require the SI to demonstrate behaviors, not describe them. For an integration, require control totals, rejection handling, and replay procedures. For a workflow, require approvals, delegation, escalation, and audit trail. For a report, require that totals tie back to the ledger and that drill-down is possible.
To keep reviews efficient, require a standardized review pack. At minimum, the pack should show the design decision being implemented, the configured objects that realize it, and the downstream impacts on reports, controls, and integrations. Design-to-build traceability: a lightweight mapping that links each approved design pattern (and any approved variant) to the specific configurations, interfaces, and reports that implement it. Traceability prevents a common failure mode: features being built that no one remembers approving, while true requirements are missed because they were never tied to concrete build objects.
Timing matters. Run design reviews early enough that changes do not cause cascading rework. For complex domains (intercompany, conversion, reporting), use iterative reviews rather than one big sign-off. Establish basic standards (naming, documentation, extension patterns) so the build remains coherent and supportable.
Design and build reviews must include data and security, not just process. Data reviews validate master data workflows, ownership, and required attributes. Security reviews validate role design, SoD rules, and privileged access handling. These areas are common late blockers because they are often deferred until the system “works.”
Close readiness should be treated as a design and build review topic, not only a testing topic. Require each process design to include its month-end implications: cutoffs, reconciliations, expected accruals and adjustments, and the reports used to prove completeness and accuracy.
Use structured sign-off, but avoid “paper sign-off” without understanding. Limit sign-off to empowered owners, and make it conditional on closing open items by specified dates.
- Review governance checklist: standard review lenses, named approvers with authority, evidence-based demonstrations, required gates before build, iterative reviews for complex areas, configuration standards and peer review, and explicit month-end implications captured in design.
10.3 Scope Control and Change Management for Build
Scope control during build is where many programs lose their trajectory. Once configuration starts, every new idea feels feasible, and stakeholders can see the system enough to request “small enhancements.” Meanwhile, the SI’s default posture may be to accommodate requests. Without disciplined change control, the program accumulates variants and customizations that explode testing and destabilize go-live.
Build scope control: the disciplined management of solution changes after design sign-off, using explicit value tradeoffs and a defined approval process. The objective is not to stop all changes. The objective is to ensure changes are deliberate, justified, and affordable within the program’s capacity and risk tolerance.
Start by establishing a baseline. Baseline scope includes approved global process designs, the approved variant register, the defined integrations and reporting inventory, and the agreed conversion objects. Anything outside that baseline is a change, even if it feels “small.”
Implement a structured intake that triages requests quickly. A practical triage model classifies each request as one of four types. Defect: the system does not meet an agreed requirement or design. Compliance need: a regulatory or audit requirement that must be met for go-live. Gap closure: a missing capability required to run the business on Day 1. Enhancement: an improvement beyond Day 1 needs.
For each change request, require impact assessment in three dimensions. Delivery impact: design, build, and testing effort, including rework. Control impact: effects on SoD, audit trail, approvals, and evidence. Lifecycle impact: upgradeability, support burden, and integration complexity.
Protect change control from being undermined by language. Teams will sometimes label new scope as a “defect” to avoid approvals, while others will label true defects as “change” to shift cost. Set a simple rule: a defect is failure to meet an approved design or requirement; a change is anything that alters that approved baseline. Then enforce it consistently in triage, with the design authority as the referee when classification is disputed.
Also establish a change budget: a deliberate allocation of time and capacity for unavoidable change within the release (for example, newly discovered statutory needs or genuinely missing Day 1 capability). A change budget is not permission to expand scope; it is a control mechanism that forces prioritization. When the budget is consumed, leaders must either de-scope something else, add capacity, or move the date. This prevents the slow, quiet accumulation of “just one more thing” that eventually collapses testing.
Demand tradeoffs. If a change adds scope, something else must move. If leaders refuse to trade off, they are implicitly choosing schedule slip or quality degradation. Make that explicit and require steering to accept the consequence knowingly.
Use time-based controls as you approach testing. Scope freeze: after this date, new functionality enters the release only with executive approval. Build freeze: after this date, changes are limited to defect fixes and critical compliance needs so testing can stabilize. Testing against a moving target produces false results and repeated cycles.
Control variants and customizations aggressively. Each variant multiplies training, testing, and support. Require a clear rationale, define minimum scope, and capture reporting and close impacts. Customization decision: an explicit choice to deviate from standard capabilities in a way that increases long-term maintenance and upgrade burden; require additional approval and a clear support plan.
Make change control operational with a weekly forum that includes finance process owners, IT architecture leads, and the SI. Approve or reject quickly, publish decisions, and trend change volume over time; rising change volume late in build is a warning sign that should trigger intervention.
- Build change control checklist: clear baseline scope, triage categories, impact assessment including lifecycle, required tradeoffs, scope and build freezes, and explicit governance for variants and customizations.
10.4 Testing Governance and Business Participation Model
Testing is where the program proves the business can run the solution, not just that the system can execute transactions. The most common testing failure is “IT-only testing”: scripts run by technical teams that validate configuration but do not validate end-to-end operations, controls, reporting, and close behavior. The result is a system that passes tests and fails in production.
Testing governance: the structures, roles, standards, and exit criteria that ensure testing is comprehensive, evidence-based, and owned by the business where it matters. Testing governance is also where you protect the schedule, because uncontrolled defects and repeated cycles are a primary driver of late delays.
Start with ownership. Business teams must own user acceptance testing because they are the ones who will run the process. IT and the SI support environments, defect resolution, and technical testing, but they cannot certify that finance can close. Define a testing RACI: finance process owners own scenario coverage and sign-off; IT owns environment stability and integration testing; the SI owns defect fixes and technical test execution; internal audit may own or observe controls testing depending on compliance needs.
Define test phases and their purpose, and do not blur them. System testing validates configuration components. Integration testing validates end-to-end flows across systems with control totals and monitoring. UAT validates business execution, controls, and reporting. Performance and security testing validate nonfunctional requirements. Cutover rehearsals validate the transition plan. Close rehearsals validate month-end reality.
Make coverage explicit with a scenario coverage matrix: a simple mapping from critical business outcomes and risks to the test scenarios that prove them. For example, if close acceleration is a target, the matrix should include scenarios that test period close sequencing, intercompany matching, and reconciliation completion within the calendar. If controls uplift is a target, the matrix should include scenarios that prove approval thresholds, SoD restrictions, and audit trail evidence for high-risk transactions. The matrix prevents a common gap: hundreds of scripts executed, yet no one can prove the program tested what actually matters.
Define exit criteria in business language and enforce them consistently across cycles. Exit criteria should include both completion and quality, and they should be supported by retained evidence. A practical set of exit criteria categories is below.
- Coverage: all critical end-to-end scenarios executed, including at least one exception path per scenario.
- Defects: severity thresholds met, with aged defects and reopen rates within agreed limits.
- Controls: approval workflows, SoD constraints, and evidence retention proven for key controls.
- Reporting: trial balance and certified reports tie to the ledger with drill-down traceability.
- Operability: monitoring, runbooks, and support triage demonstrated for close-critical interfaces and workflows.
Testing needs realistic data and realistic timing. Create a test data strategy that includes representative master data, transaction volumes, and period-end timing. Where possible, use masked production-like data to expose coding and hierarchy issues. For integrations, test with realistic file timing and failures.
Defect management is a governance system by itself. Define severity levels, triage cadence, and decision rights for workarounds. Require root-cause tagging so trends are visible (data, configuration, integration, security, training, process). Track defect aging and reopen rates, and enforce exit criteria; unresolved high-severity defects at the end of UAT are a readiness failure.
Business participation must be designed, not assumed. Define required testers by role and location, allocate time, and plan backfills. Provide role-based training before UAT so testers can execute scripts without learning basic navigation. Make participation visible; if participation is low, testing results are not credible and leadership must intervene.
Include controls and reporting validation explicitly. Controls testing should validate SoD rules, approval workflows, audit trail, and evidence retention. Reporting validation should prove that trial balances, key subledger reports, and certified management reports tie to the ledger and support drill-down. Close rehearsal is the ultimate test: run at least one rehearsal that executes the close calendar end to end, with daily triage and evidence capture.
- Testing governance checklist: business-owned UAT sign-off, clear phase purposes, realistic data and timing, defect triage with root-cause tagging, evidence standards, explicit participation model with backfills, controls and reporting validation, and at least one end-to-end close rehearsal as a readiness gate.
10.5 Vendor Performance Management
Vendor performance management is how you ensure the SI and other vendors deliver what they promised, with the quality and transparency required for a high-stakes finance program. It is not about adversarial relationships. It is about clear expectations, measurable performance, and timely intervention when performance slips.
Performance management: a fact-based governance approach that tracks vendor delivery against milestones, quality measures, staffing commitments, and contractual obligations, and uses escalation and commercial levers when needed. Performance management should be integrated with the PMO’s plan and quality dashboards so leaders see one version of truth.
Start with staffing quality and stability. Many programs sign a contract based on a strong proposed team and then receive a different team in delivery. Require named key roles, onboarding plans, and continuity expectations. Track actual staffing against the plan weekly, including the mix of senior and junior resources and turnover in critical roles.
Measure delivery performance with evidence-based metrics: on-time deliverable completion with accepted quality, defect density by component, rework rates, conversion success rates, interface stability, and the percentage of deliverables accepted without major rework. Avoid metrics that can be gamed, such as “tasks completed.”
Run structured governance cadences with the SI. Weekly delivery governance should focus on plan, dependencies, and blockers. Monthly executive governance should focus on trend metrics, staffing stability, and scope/quality tradeoffs. Manage commercial controls with transparency: validate timesheets and billing against deliverables and agreed roles, and track burn rate against plan.
Use a simple vendor scorecard and review it openly. Include staffing stability, deliverable acceptance rate, defect escape rate, responsiveness to triage, and transparency of reporting. The point is not to punish; it is to create shared facts that trigger early corrections before problems become contractual disputes at executive and delivery levels.
Escalation should be structured and timely. When performance issues arise, require a corrective action plan with specific actions, owners, and dates. If the issue is staffing, require replacement with defined qualifications. If the issue is quality, require stabilization sprints and, when needed, temporary build freezes. If the issue is planning realism, require a re-baseline based on evidence, not optimism.
Incentives and penalties can help, but only if “done” is defined objectively. If you choose to use incentives, tie them to readiness criteria (for example, successful rehearsal close and defect thresholds) rather than calendar dates. Dates can be met by cutting quality; readiness criteria are harder to fake.
Finally, manage the relationship without losing independence. Collaboration is essential, but the client must retain control of decisions, standards, and readiness. The healthiest relationships are transparent and fact-based: risks are surfaced early, tradeoffs are explicit, and there is a shared commitment to operability after go-live.
- Vendor performance checklist: named key roles with continuity tracking, evidence-based delivery metrics, integrated governance cadence, burn-rate and billing transparency, corrective action plans for performance issues, and incentives aligned to readiness evidence rather than dates.