Back to Thinking

AI implementation · Measurement

AI Impact You Can Defend

How to find measurable AI opportunities, build the evidence into the workflow, and report business value without inflating ROI.

Ross McLaughlin, Rivington Solutions· 28 min read

AI impact is not measured at the tool. It is measured in what happens differently to a unit of work. This guide shows how to find the right workflow, build a credible measurement plan, and report value without confusing activity, capacity, and ROI.

A common AI impact model is arithmetic disguised as evidence.

Estimate the minutes an employee saves. Multiply by the number of employees. Multiply again by loaded salary. Put the result in a slide labeled ROI.

That calculation may describe theoretical capacity. It does not show that the work changed, that quality held, that the time was redeployed, or that the company captured any economic value.

The problem usually begins earlier. Measurement is added after the tool has been selected and the pilot has started. There is no reliable baseline. The intervention is not tied to a specific workflow constraint. Usage data is available, but operating data is not. When leadership asks what changed, the team can report licenses, prompts, outputs, and employee sentiment—but not the result.

Those metrics can help manage a rollout. They cannot, by themselves, prove impact.

AI impact is measured in what happens differently to a unit of work:

  • Did the claim reach the right adjuster faster?
  • Did the qualified account reach the right rep while it was still active?
  • Did the underwriter spend less time assembling information without increasing referral errors?
  • Did the support agent resolve more cases without lowering resolution quality or customer trust?
  • Did an expert review more submissions while preserving the judgment the review exists to provide?

When the answer can be observed at the workflow level, value can be measured. When it cannot, the impact story rests on assumptions.

Design backward from the business outcome. Measure forward from the workflow change.

This guide lays out the method I use in Rivington's work: Find a consequential workflow constraint, Prove the change by designing the evidence before the build, and Claim only the value the organization can defend.

Why no universal AI productivity number will answer the question

AI does not create a uniform effect across every task, person, or operating environment.

In the final published study of 5,172 customer-support agents, access to a generative AI assistant increased issues resolved per hour by 15% on average, with larger gains among less-experienced and lower-skilled workers. In a preregistered experiment with 758 BCG consultants, GPT-4 users worked faster and produced higher-quality results on the tested tasks inside its capability frontier; on one deliberately outside-frontier task, they were 19 percentage points less likely to produce a correct solution. In a small randomized 2025 trial, 16 experienced open-source developers completed 246 tasks in familiar repositories; allowing early-2025 AI tools increased completion time by 19% even though the developers believed they were faster.

The point is not that AI works or does not work. The point is that its effect is conditional. External studies can help form a hypothesis. They cannot tell you what happened inside your workflow.

That requires evidence designed around the work itself.

Four distinctions that make AI measurement more credible

Common reportWhat it can showWhat it cannot establish by itself
Licenses, active users, prompts, or outputsAvailability and activityA better operating or business result
Model accuracy or task evaluation scoreCapability under defined test conditionsPerformance of the full human-and-system workflow
Estimated hours savedPotentially released capacityCost savings, incremental output, or captured value
A precise ROI estimateThe result of a model and its assumptionsThat the change was caused by the intervention

These distinctions lead to a more useful measurement system.

Activity is not impact. Adoption matters because an intervention cannot create value if it is not used. It is an exposure measure, not the outcome.

Model performance is not workflow performance. A strong model can still create a weak result if inputs are incomplete, handoffs fail, people do not trust it, or exceptions have nowhere to go.

Released capacity is not captured value. A measured reduction in touch time, with comparable output and guardrails held, creates an option. The organization must use that option to process more volume, improve service, avoid hiring or contractor spend, increase contribution margin, or reduce loss before it becomes economic value.

Precision is not confidence. A precise number built on self-reported time savings and an untested counterfactual is still a weak claim.

The method: Find, Prove, Claim

PhaseGoverning questionPractical output
[Find](#1-find-a-workflow-worth-measuring)Which workflow constraint is consequential, observable, and worth changing?Opportunity Map
[Prove](#2-prove-the-change-by-designing-the-evidence-before-the-build)What evidence will tell us whether the intervention worked?Measurement Contract
[Claim](#3-claim-only-the-value-the-company-can-defend)What changed, how confident are we, and what value did the company capture?Impact Ledger

The order matters. If the company cannot describe how it would prove value, that uncertainty belongs in the opportunity decision—not in the final reporting exercise.

1. Find a workflow worth measuring

Start with one consequential workflow, not a list of AI use cases

An AI opportunity is not “give the sales team a copilot” or “use an agent for support.” Those are possible interventions. They say nothing about where performance is being lost or what result should change.

A better starting point is a recurring workflow with:

  • A visible business consequence
  • A clear beginning and end
  • A repeated unit of work
  • Enough volume or complexity to matter
  • Multiple people, systems, decisions, or handoffs
  • Observable delay, manual effort, rework, inconsistency, or risk
  • An owner who can change how the work operates

Look for places where manual review is the default, information is scattered, handoffs lose context, the same decision is made differently by different people, exceptions escalate informally, or senior employees repeatedly intervene.

One question is especially productive:

Where is expensive judgment being used as middleware?

If experienced people spend much of their time retrieving information, rebuilding context, checking routine inputs, moving data between systems, coordinating handoffs, or sorting work before they can exercise judgment, there may be a strong workflow opportunity.

The objective is not maximum automation. It is maximum leverage of human judgment.

Work backward from the outcome and decision

Rivington's workflow architecture uses nine questions in a structured pass:

  1. Outcome: What economic or operating result matters?
  2. Decision: What consequential decision creates that outcome?
  3. Information: What context does the decision-maker require?
  4. Judgment: Where is real expertise necessary?
  5. AI: Where can probabilistic reasoning improve the work?
  6. Automation: What is deterministic and repeatable?
  7. Exceptions: Where should a human take over?
  8. Routing: Who owns the next action, and when?
  9. Measurement: How will we know whether the redesigned workflow improved?

Measurement appears ninth because it stress-tests the proposed design. Once the team answers it, revisit the earlier questions and change the workflow if it cannot produce the evidence needed to assess performance.

Define the unit of work

Company-level questions such as “Is AI making us more productive?” are usually too broad to answer well.

Choose a unit that can pass through both the current and future workflow:

  • One claim
  • One lead or account
  • One support case
  • One submission
  • One proposal
  • One approval request
  • One onboarding
  • One review decision

The unit makes the baseline, comparison, and economics coherent. It lets the team ask what happened to comparable work—not whether people generally felt faster.

Observe the real workflow

Documented process is a useful input. It is rarely the full workflow.

To establish the current state:

  • Follow a sample of real work items from entry to outcome.
  • Interview the people who perform, review, route, and receive the work.
  • Inspect queue, system, calendar, and communication timestamps where available.
  • Record workarounds, copied data, informal approvals, and decisions made in Slack or email.
  • Separate touch time from wait time.
  • Identify routine paths, exceptions, and repeated rework.
  • Note where decision criteria live: policy, system logic, individual memory, or nowhere explicit.

The goal is not a perfect process map. It is a defensible account of how work actually moves and where performance is lost.

Quantify friction before estimating upside

For each unit of work, capture the variables that determine current performance:

DimensionUseful measures
VolumeEligible units per day, week, or month; mix by type and complexity
EffortHuman touch time; number and seniority of people involved
DelayEnd-to-end cycle time; queue time; handoff latency; p50 and p90
QualityError, rework, correction, reopen, approval, or acceptance rate
ExceptionsException rate; override rate; escalation rate; resolution latency
OutcomeThroughput, conversion, service level, customer result, loss, or risk event
CostLabor, model, software, integration, human review, retries, and support

A simple capacity calculation can expose the scale of a constraint:

Modeled capacity at stake per period = eligible units in the period × modeled avoidable touch-time hours per unit

That is a useful operating estimate. It is not yet ROI.

Delay, quality, revenue, customer experience, and risk may be more important than labor time. Keep them visible rather than forcing every opportunity into a salary-savings model.

Build an Opportunity Map

Use one row for each candidate workflow.

Opportunity Map
FieldQuestion to answer
WorkflowWhat recurring work are we examining?
OwnerWho is accountable for its operating result?
Unit of workWhat enters, moves through, and exits the workflow?
Consequential decisionWhat decision most directly affects the outcome?
Current constraintWhere are time, capacity, quality, revenue, or control being lost?
Baseline evidenceWhat data or sample can establish current performance?
Proposed changeWhat should the system do differently? What should the human do differently?
Primary operating measureWhat should move first if the change works?
Quality or risk guardrailWhat must not get worse?
Value pathwayHow could the operating change become business value?
ReadinessAre data, ownership, access, and implementation support available?

Then use a qualitative 1-to-5 triage score for each dimension: 1 is weak, 3 is workable, and 5 is strong.

  1. Consequence: Would improvement affect a meaningful operating or economic result?
  2. Recurrence: Is there enough repeated work to learn and compound value?
  3. Friction: Is the current constraint material and observable?
  4. Evidence: Can the company establish a baseline and a reasonable comparison?
  5. Readiness: Are the owner, data, systems, and users accessible?

Do not let a high total hide a fatal weakness. If consequence or evidence scores below 3, or there is no accountable owner, the workflow is unlikely to be a strong first pilot.

The best first opportunity is not always the one with the largest theoretical upside. A somewhat smaller workflow with clear evidence and an accountable owner can produce more learning—and more defensible value—than a broad, impressive idea no one can measure.

2. Prove the change by designing the evidence before the build

Write the impact hypothesis

Before selecting a model, tool, or interface, state what should change and why.

Use this structure:

If we [change the workflow] for [eligible unit or population], then [primary operating measure] should move from [baseline] to [target or direction], without worsening [quality or risk guardrail], producing [business consequence] because [causal mechanism].

Example:

If the system assembles account signals and produces a decision-ready qualification brief for eligible enterprise leads, research touch time should fall from 20 minutes to 5 minutes while routing accuracy remains at or above baseline. At 200 leads per week, that would release roughly 50 hours of selling capacity for higher-value qualification and outreach.

Those numbers are illustrative. Their value is the logic they make visible: unit, population, workflow change, operating measure, guardrail, volume, and intended consequence.

Create a Measurement Contract

A Measurement Contract is the agreement the sponsor, workflow owner, builder, and analyst make before launch about what evidence will govern the next decision.

Measurement Contract
FieldDefinition
DecisionWhat will the evidence determine: stop, redesign, expand, or scale?
Unit and eligible populationWhich work items count, and which do not?
InterventionWhat exactly changes in the workflow?
Primary measureWhat single operating result best represents improvement?
BaselineWhat is current performance, over what period, and from which source?
ComparisonHow will we estimate what would have happened without the change?
Assignment and exposureWhich work was eligible or assigned, and what counts as actual use? Report both rather than analyzing only adopters.
Analysis and observation planWhat sample or effect is decision-relevant? How will exclusions, missing data, uncertainty, and outcome-maturation time be handled?
Guardrails and toleranceWhich quality, customer, safety, compliance, or risk measures must hold, and within what preset limit?
InstrumentationWhich events, timestamps, reasons, and outcomes must be recorded?
Data governanceWhat minimum data is required, who may access it, how long is it retained, and which privacy, security, legal, or employee-notice reviews apply?
Value conversionHow could an observed operating change become economic value?
Full costBuild, model, software, integration, review, support, and change cost
Financial validationWhich finance owner, reporting period, valuation assumptions, and one-time versus recurring costs govern the claim?
ThresholdsWhat results cause the team to scale, redesign, pause, or stop?
Ownership and dateWho reviews the evidence, and when is the decision made?

The contract reduces the risk of changing the definition of success after seeing the result and makes any later changes visible. It also keeps a technically successful pilot from being mistaken for a valuable operating change.

Use a metric stack, not one headline number

A credible pilot needs measures at several levels.

LevelGoverning questionExamples
ExposureDid eligible work actually enter the new workflow?Workflow coverage, eligible units exposed, completion through new path
WorkflowDid work move differently?Touch time, cycle time, throughput, queue depth, straight-through rate
Human-systemHow did judgment and exceptions behave?Acceptance, material edit, override, escalation, exception rate, override latency
Quality and riskWas the result still acceptable?Accuracy, rework, reopen rate, policy adherence, incidents, customer outcome
BusinessDid the operating change affect the company?Capacity used, cost per completed unit, conversion, contribution margin, observed loss change, modeled expected-loss reduction

Pair speed with quality. Pair automation with exceptions. Pair adoption with outcomes.

No single metric is sufficient. A rising straight-through rate can look positive while rework quietly increases. A falling override rate can indicate improved recommendations—or automation bias. A high acceptance rate can mean the system is useful, or that review is superficial.

Metrics become meaningful in relation to one another and to the decision the workflow exists to make.

Establish the baseline before the workflow changes

A baseline reconstructed after launch is vulnerable to selective memory, inconsistent definitions, and unavailable data.

Before the pilot:

  • Use the same unit, eligibility rules, and outcome definition planned for the new workflow.
  • Cover enough time or volume to represent normal variation.
  • Account for seasonality, campaign changes, staffing, backlog, and mix shifts.
  • Segment by complexity, channel, customer type, employee experience, or other factors likely to affect the result.
  • Report distributions where averages hide the operating reality; median and p90 cycle time are often more useful than the mean alone.
  • Record data limitations explicitly.

When system data is missing, use a structured sample. A carefully reviewed set of real work items is a stronger baseline than an unsupported estimate. The first phase of the work may simply be making the workflow observable.

Choose the strongest comparison the operating context can support

The central question is counterfactual: what would likely have happened to comparable work without the intervention?

Use the strongest feasible design:

  1. Randomized comparison: Assign eligible work items, users, or teams to current and redesigned conditions when it is practical and appropriate.
  2. Staggered rollout: Introduce the workflow to comparable teams, regions, or queues at different times.
  3. Matched comparison: Compare exposed work with similar unexposed work using relevant characteristics.
  4. Interrupted time series: Use a stable pre-change trend and enough post-change observations to identify a break.
  5. Simple pre/post comparison: Use when nothing stronger is feasible, but document concurrent changes and limit the claim.
  6. Self-report or modeled estimate: Use as supporting evidence, not as proof of attributed impact.

The best design is not the most academically impressive one. It is the strongest comparison the company can execute without damaging service, creating unacceptable risk, or making the pilot impossible to operate.

Avoid comparing enthusiastic power users with nonusers and treating the difference as causal. People choose whether and how to adopt AI, so the groups may differ before the intervention. Where work and practices spill across colleagues, compare at the team, queue, or location level rather than pretending each employee is an isolated unit.

Design strength still depends on execution, sample size, contamination between groups, learning effects, attrition, outcome maturity, and the assumptions each method requires. The ordering above is a guide, not an automatic grade.

Instrument the whole workflow, not only the model

A model call is not an outcome. Record enough of the workflow to connect the intervention to the final result.

For each unit, capture where appropriate:

  • A persistent work-item identifier
  • Eligibility, assignment, comparison group, and actual exposure
  • Entry, queue, handoff, review, decision, and completion timestamps
  • Completion status, including failed, abandoned, rerouted, and fallback cases
  • Whether AI was available and actually used
  • Model or workflow version
  • Recommendation, classification, or draft produced
  • A validated or calibrated confidence or uncertainty signal, if used; an LLM's self-reported confidence is not a probability
  • Human action, material edits, and final decision
  • Override or rejection reason
  • Exception type, owner, and resolution time
  • Downstream quality and business outcome
  • Model, software, human-review, and exception-handling cost

This is the connective tissue between technical capability and operating value. Without it, the company can inspect the model or the outcome, but not the mechanism linking them.

Treat exceptions as measurement infrastructure

Exceptions are where many AI-enabled workflows reveal their real cost.

If low-confidence or unusual cases disappear into Slack, email, or an informal manager review, the system loses both control and learning. Every material exception should have a defined route, owner, service expectation, resolution, and reason code.

Two useful early-warning measures are:

  • Exception rate: the share of work requiring human intervention outside the routine path
  • Override latency: the time from exception trigger to accountable resolution

An increasing exception rate can signal model stress, missing inputs, poor thresholds, or an expanding scope. Increasing override latency can signal that the human review layer is becoming the new bottleneck.

Do not hide this labor. Human review, retries, corrections, and exception handling belong in both the workflow result and the full cost of completed work.

Decide the scale, redesign, and stop thresholds in advance

A pilot should not end with a presentation. It should end with a decision.

Example thresholds:

  • Scale: Primary measure improves beyond the agreed threshold, guardrails hold, workflow coverage is sufficient, and the economics remain credible after full run cost.
  • Redesign: Users engage, but exceptions, edits, or a specific segment reveal a correctable workflow or system problem.
  • Pause: Quality, compliance, customer, or safety guardrails breach their tolerance.
  • Stop: The intervention creates activity but no meaningful operating improvement, or the cost to sustain it exceeds the value pathway.

A negative result is not a failed measurement program. It is information that can prevent the company from scaling the wrong thing.

3. Claim only the value the company can defend

Separate modeled, observed, attributed, and captured value

AI impact reports often collapse four different claims into one number.

Value layerWhat it meansAppropriate language
Modeled potentialUpside if assumptions about volume, adoption, performance, and conversion hold“The modeled annual opportunity is…”
Observed operational changeA measured difference in the workflow for exposed work“Touch time fell by… while quality remained…”
Attributed impactThe change estimated to have been caused by the intervention relative to a credible counterfactual“Compared with similar unexposed work, the intervention changed…”
Captured business valueThe company demonstrably realized throughput, avoided spend, margin, reduced loss, or another bankable result after the workflow changed; attribution strength determines how confidently that value can be credited to the intervention“The company processed… avoided… generated… or reduced…”

All four can be useful. They should not be presented as though they have the same evidentiary weight. When attribution is weak, report an observed business result after the change—not value caused by the intervention—and show attribution strength separately.

Be precise about capacity, use, and financial value

Use these calculations as a starting point:

Gross capacity released in the period = comparison-adjusted change in mean primary-worker touch-time hours per eligible assigned unit × all eligible units assigned to the redesigned workflow

Net capacity released in the period = gross capacity released - incremental review, exception, correction, and maintenance human hours in the same period

Use mean touch time for aggregate capacity calculations. Use median and p90 to describe the operating distribution and expose long-tail delay.

Released capacity becomes captured value only when the organization can show what happened to it.

Examples include:

  • More cases completed with the same team
  • Faster response that improves conversion or retention
  • Overtime, contractor, or outsourced work reduced
  • A contemporaneously approved hire canceled or deferred, with finance confirming the cash effect and service and quality holding
  • Scarce experts shifted to higher-value or higher-risk work
  • Backlog reduced and a service-level commitment restored

If none of those has happened yet, report capacity honestly as capacity. Do not multiply hours by salary and call the result savings.

Convert operating results using the right economic pathway

Operating changePossible value pathwayEvidence required
Less touch timeAdded throughput, avoided overtime or contractors, credible hiring avoidanceCompleted volume, staffing/spend decision, service and quality held
Shorter cycle timeConversion, retention, earlier revenue, lower working capitalA demonstrated relationship to the downstream business outcome
Lower rework or errorLabor avoided, fewer credits/refunds, lower remediation costError definition, baseline frequency, unit cost of correction
Better routing or prioritizationHigher conversion, service level, or scarce-capacity allocationRouting accuracy plus downstream outcome by comparable cohort
Lower incident or exception rateModeled expected-loss reduction, less review burden, lower disruptionValid event frequency and loss assumptions reviewed with finance or risk; use stronger language only with a defensible comparison and validated loss model

Revenue claims should use incremental contribution where possible, not gross revenue. Risk claims should disclose the event-probability and loss assumptions. Cost claims should include the full cost of the AI-enabled workflow—not only model tokens or licenses.

Additional throughput is financial value only when demand exists and the incremental contribution is demonstrated. Otherwise, report it as captured operating value.

At minimum, include:

  • Model and software cost
  • Retrieval, orchestration, and infrastructure
  • Implementation and integration
  • Human review and exception handling
  • Evaluation and monitoring
  • Training and change support
  • Ongoing maintenance and workflow ownership

Then calculate:

Net realized value = gross realized financial benefit - full incremental build, run, review, and change cost for the same period

Track incremental review and exception labor in both the operating and financial views, but subtract it only once in the economic model. Keep one-time build cost separate from recurring run cost while using a common reporting horizon.

Only calculate ROI when the captured benefit can also be attributed with appropriate confidence:

ROI for a defined period = attributed net realized value ÷ full incremental investment cost for the period

Have finance approve the conversion method, and report the numerator, denominator, period, assumptions, and attribution grade alongside the percentage.

Grade attribution separately from value status

The attribution grade below is an internal decision aid, not an industry standard. Assign it to each material claim—not to the initiative as a whole—and consider design, execution, data quality, sample adequacy, and uncertainty. A pilot can have strong evidence for a touch-time effect and weak evidence for a financial conversion.

GradeEvidence basisClaim posture
ARandomized or similarly strong counterfactual with adequate sample and reliable outcome dataStrong causal support if execution, contamination, attrition, and outcome-data checks pass
BStaggered rollout or well-matched comparison with reliable outcome dataCredible attribution with stated limitations
CAdjusted pre/post evidence with known concurrent changesDirectional association; causal attribution limited
DModeled estimate, survey, or self-report without a credible comparisonNo causal attribution; planning value only

A lower grade does not make the information useless. It limits how confidently and precisely the organization should speak about it.

“Modeled annual potential, attribution grade D” is more useful than a precise ROI percentage when the inputs are still assumptions.

Use an Impact Ledger, not a collection of success stories

An Impact Ledger gives every initiative the same reporting structure.

Impact Ledger
FieldWhat to report
Initiative and workflow ownerName the change and the accountable operator
Unit, eligible population, and periodDefine exactly what the result covers
Intervention and exposureState what changed and how much work used it
Baseline and comparisonShow the reference point and method
Primary resultReport absolute and percentage change
GuardrailsShow quality, risk, customer, and exception results
Operational impactCapacity, speed, throughput, quality, or risk movement
Captured valueState what reached output, spend, margin, or loss—with a range if appropriate
Full cost and net valueInclude build and ongoing operating cost
Value statusModeled potential, observed change, attributed impact, or captured value
Attribution grade by claimA, B, C, or D for each material operating or financial claim
Assumptions and limitationsName what could change the conclusion
Next decisionScale, redesign, pause, stop, or continue measuring

The ledger should contain weak, negative, and stopped initiatives as well as successes. Otherwise it becomes marketing rather than management.

Report differently for operators and executives

AudienceWhat they need
Workflow ownerWhich parts of the flow improved, where exceptions moved, and what to change next
Builder or technical ownerReliability, latency, cost, edits, failures, version effects, and instrumentation gaps
Finance or riskAssumptions, full cost, attribution, downside, and audit trail
Executive sponsorCaptured value, confidence, guardrails, and the decision required

A practical cadence is:

  • Weekly operating review: Workflow coverage, exceptions, latency, quality, incidents, and immediate corrections
  • Monthly impact review: Baseline comparison, operational impact, cost, attribution strength, and scale/redesign decision
  • After a material model, prompt, retrieval, policy, or workflow change: Rerun relevant evaluations and revalidate operating thresholds
  • Quarterly portfolio review: Captured value across initiatives, cumulative cost, stopped work, concentration risk, and where to invest next

The report is successful when it changes a decision—not when the dashboard gains another chart.

What I would not call ROI

On their own, I would not call any of the following ROI:

  • Employees trained
  • Licenses provisioned
  • Weekly active users
  • Prompts, calls, or outputs generated
  • A model evaluation score
  • Gross employee estimates of hours saved
  • Salary multiplied by theoretical time savings
  • Headcount avoidance no one changed a hiring plan to capture
  • Benefits reported without human-review, exception, implementation, and run costs

These can be inputs to a measurement system. They are not the final business result.

What this looks like in practice

In a prior operating role, before Rivington, I redesigned a high-volume expert-review workflow that supported more than 5,000 monthly inbound submissions.

The workflow was rebuilt around structured intake, explicit qualification logic, consistent routing, concise expert summaries, and human review at the points where judgment mattered. The reported result was an approximately 80% reduction in expert assessment time and more than 50 hours of manual work removed each week, with the expert-review step retained for material cases.

The important measure was not the volume of summaries produced. It was the reported change in expert assessment time across live workflow volume, with the judgment requirement intact.

The claim should stop there unless the organization can also show what the released capacity produced. If those hours supported more completed assessments, reduced a backlog, avoided external spend, or prevented a planned hire, that outcome can be measured separately as captured value. If not, the result remains a meaningful operating improvement—not automatic cash savings.

The public facts for this example do not, by themselves, establish a controlled counterfactual, a measured quality effect, or captured financial value. Retaining an expert-review step describes the workflow design; it does not prove quality held. In an Impact Ledger, the result should include the exact baseline period and comparison method supported by the underlying record, report any measured guardrail separately, assign the corresponding attribution grade, and avoid an ROI claim unless redeployment and value capture are documented.

That distinction makes the impact story more credible, not less impressive.

A 30-day way to begin

Thirty days is enough to design the measurement system, not necessarily to demonstrate impact. The observation period depends on workflow volume, normal variation, and how long downstream outcomes take to mature.

Week 1: Choose the work

  • Select one recurring workflow with a visible consequence.
  • Name the owner and the unit of work.
  • Define the business outcome and consequential decision.
  • Complete the Opportunity Map.

Week 2: Baseline reality

  • Follow real work from entry to outcome.
  • Quantify volume, effort, delay, quality, exceptions, and current cost.
  • Identify routine work, genuine judgment, and informal workarounds.
  • Record gaps that make the workflow difficult to observe.

Week 3: Design the evidence

  • Write the impact hypothesis.
  • Complete the Measurement Contract.
  • Choose the comparison method.
  • Define primary, exposure, guardrail, exception, and business measures.
  • Add required events and reason codes to the build plan.

Week 4: Set up the decision

  • Agree on scale, redesign, pause, and stop thresholds.
  • Assign operating, technical, data, finance, and risk responsibilities.
  • Create the first Impact Ledger entry using the baseline and modeled potential.
  • Schedule the operating and impact reviews before the pilot begins.

At the end of 30 days, the company should have more than an AI idea. It should have a defined workflow opportunity, a baseline, an evidence plan, and a decision process.

The five-question impact-claim audit

Before an AI result reaches an executive or board report, ask:

  1. What unit of work changed?
  2. Compared with what?
  3. Did quality, customer, and risk guardrails hold?
  4. How did the operating change create economic value?
  5. What value did the company actually capture, net of cost?

If the team cannot answer all five, the result may still be promising. It is not yet a defensible ROI claim.

Start with one workflow that matters

Many companies do not need another enterprise-wide AI scorecard before they can learn. They need one consequential workflow, observed honestly and designed so improvement can be measured.

Rivington's free Workflow Friction Check can help identify where ownership, inputs, policy, exceptions, and measurement are breaking down.

For a deeper answer, the AI Workflow Diagnostic examines one recurring workflow over three weeks. The work includes a current-state map, quantified friction analysis, future-state design, AI and automation recommendations, a baseline and measurement plan, and a practical 90-day implementation roadmap.

Bring one workflow that is slow, manual, inconsistent, or failing to produce a measurable result. We will determine whether the problem is specific enough, consequential enough, and accessible enough to fix.


This guide provides general operating guidance, not legal, regulatory, accounting, employment, or compliance advice. Workflows involving employment, insurance, credit, healthcare, safety, or other consequential decisions require review by the organization's appropriate legal, compliance, security, privacy, risk, finance, and domain owners.


Related reading