AI impact is not measured at the tool. It is measured in what happens differently to a unit of work. This guide shows how to find the right workflow, build a credible measurement plan, and report value without confusing activity, capacity, and ROI.
A common AI impact model is arithmetic disguised as evidence.
Estimate the minutes an employee saves. Multiply by the number of employees. Multiply again by loaded salary. Put the result in a slide labeled ROI.
That calculation may describe theoretical capacity. It does not show that the work changed, that quality held, that the time was redeployed, or that the company captured any economic value.
The problem usually begins earlier. Measurement is added after the tool has been selected and the pilot has started. There is no reliable baseline. The intervention is not tied to a specific workflow constraint. Usage data is available, but operating data is not. When leadership asks what changed, the team can report licenses, prompts, outputs, and employee sentiment—but not the result.
Those metrics can help manage a rollout. They cannot, by themselves, prove impact.
AI impact is measured in what happens differently to a unit of work:
- Did the claim reach the right adjuster faster?
- Did the qualified account reach the right rep while it was still active?
- Did the underwriter spend less time assembling information without increasing referral errors?
- Did the support agent resolve more cases without lowering resolution quality or customer trust?
- Did an expert review more submissions while preserving the judgment the review exists to provide?
When the answer can be observed at the workflow level, value can be measured. When it cannot, the impact story rests on assumptions.
Design backward from the business outcome. Measure forward from the workflow change.
This guide lays out the method I use in Rivington's work: Find a consequential workflow constraint, Prove the change by designing the evidence before the build, and Claim only the value the organization can defend.
Why no universal AI productivity number will answer the question
AI does not create a uniform effect across every task, person, or operating environment.
In the final published study of 5,172 customer-support agents, access to a generative AI assistant increased issues resolved per hour by 15% on average, with larger gains among less-experienced and lower-skilled workers. In a preregistered experiment with 758 BCG consultants, GPT-4 users worked faster and produced higher-quality results on the tested tasks inside its capability frontier; on one deliberately outside-frontier task, they were 19 percentage points less likely to produce a correct solution. In a small randomized 2025 trial, 16 experienced open-source developers completed 246 tasks in familiar repositories; allowing early-2025 AI tools increased completion time by 19% even though the developers believed they were faster.
The point is not that AI works or does not work. The point is that its effect is conditional. External studies can help form a hypothesis. They cannot tell you what happened inside your workflow.
That requires evidence designed around the work itself.
Four distinctions that make AI measurement more credible
| Common report | What it can show | What it cannot establish by itself |
|---|---|---|
| Licenses, active users, prompts, or outputs | Availability and activity | A better operating or business result |
| Model accuracy or task evaluation score | Capability under defined test conditions | Performance of the full human-and-system workflow |
| Estimated hours saved | Potentially released capacity | Cost savings, incremental output, or captured value |
| A precise ROI estimate | The result of a model and its assumptions | That the change was caused by the intervention |
These distinctions lead to a more useful measurement system.
Activity is not impact. Adoption matters because an intervention cannot create value if it is not used. It is an exposure measure, not the outcome.
Model performance is not workflow performance. A strong model can still create a weak result if inputs are incomplete, handoffs fail, people do not trust it, or exceptions have nowhere to go.
Released capacity is not captured value. A measured reduction in touch time, with comparable output and guardrails held, creates an option. The organization must use that option to process more volume, improve service, avoid hiring or contractor spend, increase contribution margin, or reduce loss before it becomes economic value.
Precision is not confidence. A precise number built on self-reported time savings and an untested counterfactual is still a weak claim.
The method: Find, Prove, Claim
| Phase | Governing question | Practical output |
|---|---|---|
| [Find](#1-find-a-workflow-worth-measuring) | Which workflow constraint is consequential, observable, and worth changing? | Opportunity Map |
| [Prove](#2-prove-the-change-by-designing-the-evidence-before-the-build) | What evidence will tell us whether the intervention worked? | Measurement Contract |
| [Claim](#3-claim-only-the-value-the-company-can-defend) | What changed, how confident are we, and what value did the company capture? | Impact Ledger |
The order matters. If the company cannot describe how it would prove value, that uncertainty belongs in the opportunity decision—not in the final reporting exercise.
1. Find a workflow worth measuring
Start with one consequential workflow, not a list of AI use cases
An AI opportunity is not “give the sales team a copilot” or “use an agent for support.” Those are possible interventions. They say nothing about where performance is being lost or what result should change.
A better starting point is a recurring workflow with:
- A visible business consequence
- A clear beginning and end
- A repeated unit of work
- Enough volume or complexity to matter
- Multiple people, systems, decisions, or handoffs
- Observable delay, manual effort, rework, inconsistency, or risk
- An owner who can change how the work operates
Look for places where manual review is the default, information is scattered, handoffs lose context, the same decision is made differently by different people, exceptions escalate informally, or senior employees repeatedly intervene.
One question is especially productive:
Where is expensive judgment being used as middleware?
If experienced people spend much of their time retrieving information, rebuilding context, checking routine inputs, moving data between systems, coordinating handoffs, or sorting work before they can exercise judgment, there may be a strong workflow opportunity.
The objective is not maximum automation. It is maximum leverage of human judgment.
Work backward from the outcome and decision
Rivington's workflow architecture uses nine questions in a structured pass:
- Outcome: What economic or operating result matters?
- Decision: What consequential decision creates that outcome?
- Information: What context does the decision-maker require?
- Judgment: Where is real expertise necessary?
- AI: Where can probabilistic reasoning improve the work?
- Automation: What is deterministic and repeatable?
- Exceptions: Where should a human take over?
- Routing: Who owns the next action, and when?
- Measurement: How will we know whether the redesigned workflow improved?
Measurement appears ninth because it stress-tests the proposed design. Once the team answers it, revisit the earlier questions and change the workflow if it cannot produce the evidence needed to assess performance.
Define the unit of work
Company-level questions such as “Is AI making us more productive?” are usually too broad to answer well.
Choose a unit that can pass through both the current and future workflow:
- One claim
- One lead or account
- One support case
- One submission
- One proposal
- One approval request
- One onboarding
- One review decision
The unit makes the baseline, comparison, and economics coherent. It lets the team ask what happened to comparable work—not whether people generally felt faster.
Observe the real workflow
Documented process is a useful input. It is rarely the full workflow.
To establish the current state:
- Follow a sample of real work items from entry to outcome.
- Interview the people who perform, review, route, and receive the work.
- Inspect queue, system, calendar, and communication timestamps where available.
- Record workarounds, copied data, informal approvals, and decisions made in Slack or email.
- Separate touch time from wait time.
- Identify routine paths, exceptions, and repeated rework.
- Note where decision criteria live: policy, system logic, individual memory, or nowhere explicit.
The goal is not a perfect process map. It is a defensible account of how work actually moves and where performance is lost.
Quantify friction before estimating upside
For each unit of work, capture the variables that determine current performance:
| Dimension | Useful measures |
|---|---|
| Volume | Eligible units per day, week, or month; mix by type and complexity |
| Effort | Human touch time; number and seniority of people involved |
| Delay | End-to-end cycle time; queue time; handoff latency; p50 and p90 |
| Quality | Error, rework, correction, reopen, approval, or acceptance rate |
| Exceptions | Exception rate; override rate; escalation rate; resolution latency |
| Outcome | Throughput, conversion, service level, customer result, loss, or risk event |
| Cost | Labor, model, software, integration, human review, retries, and support |
A simple capacity calculation can expose the scale of a constraint:
Modeled capacity at stake per period = eligible units in the period × modeled avoidable touch-time hours per unit
That is a useful operating estimate. It is not yet ROI.
Delay, quality, revenue, customer experience, and risk may be more important than labor time. Keep them visible rather than forcing every opportunity into a salary-savings model.
Build an Opportunity Map
Use one row for each candidate workflow.
| Field | Question to answer |
|---|---|
| Workflow | What recurring work are we examining? |
| Owner | Who is accountable for its operating result? |
| Unit of work | What enters, moves through, and exits the workflow? |
| Consequential decision | What decision most directly affects the outcome? |
| Current constraint | Where are time, capacity, quality, revenue, or control being lost? |
| Baseline evidence | What data or sample can establish current performance? |
| Proposed change | What should the system do differently? What should the human do differently? |
| Primary operating measure | What should move first if the change works? |
| Quality or risk guardrail | What must not get worse? |
| Value pathway | How could the operating change become business value? |
| Readiness | Are data, ownership, access, and implementation support available? |
Then use a qualitative 1-to-5 triage score for each dimension: 1 is weak, 3 is workable, and 5 is strong.
- Consequence: Would improvement affect a meaningful operating or economic result?
- Recurrence: Is there enough repeated work to learn and compound value?
- Friction: Is the current constraint material and observable?
- Evidence: Can the company establish a baseline and a reasonable comparison?
- Readiness: Are the owner, data, systems, and users accessible?
Do not let a high total hide a fatal weakness. If consequence or evidence scores below 3, or there is no accountable owner, the workflow is unlikely to be a strong first pilot.
The best first opportunity is not always the one with the largest theoretical upside. A somewhat smaller workflow with clear evidence and an accountable owner can produce more learning—and more defensible value—than a broad, impressive idea no one can measure.
2. Prove the change by designing the evidence before the build
Write the impact hypothesis
Before selecting a model, tool, or interface, state what should change and why.
Use this structure:
If we [change the workflow] for [eligible unit or population], then [primary operating measure] should move from [baseline] to [target or direction], without worsening [quality or risk guardrail], producing [business consequence] because [causal mechanism].
Example:
If the system assembles account signals and produces a decision-ready qualification brief for eligible enterprise leads, research touch time should fall from 20 minutes to 5 minutes while routing accuracy remains at or above baseline. At 200 leads per week, that would release roughly 50 hours of selling capacity for higher-value qualification and outreach.
Those numbers are illustrative. Their value is the logic they make visible: unit, population, workflow change, operating measure, guardrail, volume, and intended consequence.
Create a Measurement Contract
A Measurement Contract is the agreement the sponsor, workflow owner, builder, and analyst make before launch about what evidence will govern the next decision.
| Field | Definition |
|---|---|
| Decision | What will the evidence determine: stop, redesign, expand, or scale? |
| Unit and eligible population | Which work items count, and which do not? |
| Intervention | What exactly changes in the workflow? |
| Primary measure | What single operating result best represents improvement? |
| Baseline | What is current performance, over what period, and from which source? |
| Comparison | How will we estimate what would have happened without the change? |
| Assignment and exposure | Which work was eligible or assigned, and what counts as actual use? Report both rather than analyzing only adopters. |
| Analysis and observation plan | What sample or effect is decision-relevant? How will exclusions, missing data, uncertainty, and outcome-maturation time be handled? |
| Guardrails and tolerance | Which quality, customer, safety, compliance, or risk measures must hold, and within what preset limit? |
| Instrumentation | Which events, timestamps, reasons, and outcomes must be recorded? |
| Data governance | What minimum data is required, who may access it, how long is it retained, and which privacy, security, legal, or employee-notice reviews apply? |
| Value conversion | How could an observed operating change become economic value? |
| Full cost | Build, model, software, integration, review, support, and change cost |
| Financial validation | Which finance owner, reporting period, valuation assumptions, and one-time versus recurring costs govern the claim? |
| Thresholds | What results cause the team to scale, redesign, pause, or stop? |
| Ownership and date | Who reviews the evidence, and when is the decision made? |
The contract reduces the risk of changing the definition of success after seeing the result and makes any later changes visible. It also keeps a technically successful pilot from being mistaken for a valuable operating change.
Use a metric stack, not one headline number
A credible pilot needs measures at several levels.
| Level | Governing question | Examples |
|---|---|---|
| Exposure | Did eligible work actually enter the new workflow? | Workflow coverage, eligible units exposed, completion through new path |
| Workflow | Did work move differently? | Touch time, cycle time, throughput, queue depth, straight-through rate |
| Human-system | How did judgment and exceptions behave? | Acceptance, material edit, override, escalation, exception rate, override latency |
| Quality and risk | Was the result still acceptable? | Accuracy, rework, reopen rate, policy adherence, incidents, customer outcome |
| Business | Did the operating change affect the company? | Capacity used, cost per completed unit, conversion, contribution margin, observed loss change, modeled expected-loss reduction |
Pair speed with quality. Pair automation with exceptions. Pair adoption with outcomes.
No single metric is sufficient. A rising straight-through rate can look positive while rework quietly increases. A falling override rate can indicate improved recommendations—or automation bias. A high acceptance rate can mean the system is useful, or that review is superficial.
Metrics become meaningful in relation to one another and to the decision the workflow exists to make.
Establish the baseline before the workflow changes
A baseline reconstructed after launch is vulnerable to selective memory, inconsistent definitions, and unavailable data.
Before the pilot:
- Use the same unit, eligibility rules, and outcome definition planned for the new workflow.
- Cover enough time or volume to represent normal variation.
- Account for seasonality, campaign changes, staffing, backlog, and mix shifts.
- Segment by complexity, channel, customer type, employee experience, or other factors likely to affect the result.
- Report distributions where averages hide the operating reality; median and p90 cycle time are often more useful than the mean alone.
- Record data limitations explicitly.
When system data is missing, use a structured sample. A carefully reviewed set of real work items is a stronger baseline than an unsupported estimate. The first phase of the work may simply be making the workflow observable.
Choose the strongest comparison the operating context can support
The central question is counterfactual: what would likely have happened to comparable work without the intervention?
Use the strongest feasible design:
- Randomized comparison: Assign eligible work items, users, or teams to current and redesigned conditions when it is practical and appropriate.
- Staggered rollout: Introduce the workflow to comparable teams, regions, or queues at different times.
- Matched comparison: Compare exposed work with similar unexposed work using relevant characteristics.
- Interrupted time series: Use a stable pre-change trend and enough post-change observations to identify a break.
- Simple pre/post comparison: Use when nothing stronger is feasible, but document concurrent changes and limit the claim.
- Self-report or modeled estimate: Use as supporting evidence, not as proof of attributed impact.
The best design is not the most academically impressive one. It is the strongest comparison the company can execute without damaging service, creating unacceptable risk, or making the pilot impossible to operate.
Avoid comparing enthusiastic power users with nonusers and treating the difference as causal. People choose whether and how to adopt AI, so the groups may differ before the intervention. Where work and practices spill across colleagues, compare at the team, queue, or location level rather than pretending each employee is an isolated unit.
Design strength still depends on execution, sample size, contamination between groups, learning effects, attrition, outcome maturity, and the assumptions each method requires. The ordering above is a guide, not an automatic grade.
Instrument the whole workflow, not only the model
A model call is not an outcome. Record enough of the workflow to connect the intervention to the final result.
For each unit, capture where appropriate:
- A persistent work-item identifier
- Eligibility, assignment, comparison group, and actual exposure
- Entry, queue, handoff, review, decision, and completion timestamps
- Completion status, including failed, abandoned, rerouted, and fallback cases
- Whether AI was available and actually used
- Model or workflow version
- Recommendation, classification, or draft produced
- A validated or calibrated confidence or uncertainty signal, if used; an LLM's self-reported confidence is not a probability
- Human action, material edits, and final decision
- Override or rejection reason
- Exception type, owner, and resolution time
- Downstream quality and business outcome
- Model, software, human-review, and exception-handling cost
This is the connective tissue between technical capability and operating value. Without it, the company can inspect the model or the outcome, but not the mechanism linking them.
Treat exceptions as measurement infrastructure
Exceptions are where many AI-enabled workflows reveal their real cost.
If low-confidence or unusual cases disappear into Slack, email, or an informal manager review, the system loses both control and learning. Every material exception should have a defined route, owner, service expectation, resolution, and reason code.
Two useful early-warning measures are:
- Exception rate: the share of work requiring human intervention outside the routine path
- Override latency: the time from exception trigger to accountable resolution
An increasing exception rate can signal model stress, missing inputs, poor thresholds, or an expanding scope. Increasing override latency can signal that the human review layer is becoming the new bottleneck.
Do not hide this labor. Human review, retries, corrections, and exception handling belong in both the workflow result and the full cost of completed work.
Decide the scale, redesign, and stop thresholds in advance
A pilot should not end with a presentation. It should end with a decision.
Example thresholds:
- Scale: Primary measure improves beyond the agreed threshold, guardrails hold, workflow coverage is sufficient, and the economics remain credible after full run cost.
- Redesign: Users engage, but exceptions, edits, or a specific segment reveal a correctable workflow or system problem.
- Pause: Quality, compliance, customer, or safety guardrails breach their tolerance.
- Stop: The intervention creates activity but no meaningful operating improvement, or the cost to sustain it exceeds the value pathway.
A negative result is not a failed measurement program. It is information that can prevent the company from scaling the wrong thing.
3. Claim only the value the company can defend
Separate modeled, observed, attributed, and captured value
AI impact reports often collapse four different claims into one number.
| Value layer | What it means | Appropriate language |
|---|---|---|
| Modeled potential | Upside if assumptions about volume, adoption, performance, and conversion hold | “The modeled annual opportunity is…” |
| Observed operational change | A measured difference in the workflow for exposed work | “Touch time fell by… while quality remained…” |
| Attributed impact | The change estimated to have been caused by the intervention relative to a credible counterfactual | “Compared with similar unexposed work, the intervention changed…” |
| Captured business value | The company demonstrably realized throughput, avoided spend, margin, reduced loss, or another bankable result after the workflow changed; attribution strength determines how confidently that value can be credited to the intervention | “The company processed… avoided… generated… or reduced…” |
All four can be useful. They should not be presented as though they have the same evidentiary weight. When attribution is weak, report an observed business result after the change—not value caused by the intervention—and show attribution strength separately.
Be precise about capacity, use, and financial value
Use these calculations as a starting point:
Gross capacity released in the period = comparison-adjusted change in mean primary-worker touch-time hours per eligible assigned unit × all eligible units assigned to the redesigned workflow
Net capacity released in the period = gross capacity released - incremental review, exception, correction, and maintenance human hours in the same period
Use mean touch time for aggregate capacity calculations. Use median and p90 to describe the operating distribution and expose long-tail delay.
Released capacity becomes captured value only when the organization can show what happened to it.
Examples include:
- More cases completed with the same team
- Faster response that improves conversion or retention
- Overtime, contractor, or outsourced work reduced
- A contemporaneously approved hire canceled or deferred, with finance confirming the cash effect and service and quality holding
- Scarce experts shifted to higher-value or higher-risk work
- Backlog reduced and a service-level commitment restored
If none of those has happened yet, report capacity honestly as capacity. Do not multiply hours by salary and call the result savings.
Convert operating results using the right economic pathway
| Operating change | Possible value pathway | Evidence required |
|---|---|---|
| Less touch time | Added throughput, avoided overtime or contractors, credible hiring avoidance | Completed volume, staffing/spend decision, service and quality held |
| Shorter cycle time | Conversion, retention, earlier revenue, lower working capital | A demonstrated relationship to the downstream business outcome |
| Lower rework or error | Labor avoided, fewer credits/refunds, lower remediation cost | Error definition, baseline frequency, unit cost of correction |
| Better routing or prioritization | Higher conversion, service level, or scarce-capacity allocation | Routing accuracy plus downstream outcome by comparable cohort |
| Lower incident or exception rate | Modeled expected-loss reduction, less review burden, lower disruption | Valid event frequency and loss assumptions reviewed with finance or risk; use stronger language only with a defensible comparison and validated loss model |
Revenue claims should use incremental contribution where possible, not gross revenue. Risk claims should disclose the event-probability and loss assumptions. Cost claims should include the full cost of the AI-enabled workflow—not only model tokens or licenses.
Additional throughput is financial value only when demand exists and the incremental contribution is demonstrated. Otherwise, report it as captured operating value.
At minimum, include:
- Model and software cost
- Retrieval, orchestration, and infrastructure
- Implementation and integration
- Human review and exception handling
- Evaluation and monitoring
- Training and change support
- Ongoing maintenance and workflow ownership
Then calculate:
Net realized value = gross realized financial benefit - full incremental build, run, review, and change cost for the same period
Track incremental review and exception labor in both the operating and financial views, but subtract it only once in the economic model. Keep one-time build cost separate from recurring run cost while using a common reporting horizon.
Only calculate ROI when the captured benefit can also be attributed with appropriate confidence:
ROI for a defined period = attributed net realized value ÷ full incremental investment cost for the period
Have finance approve the conversion method, and report the numerator, denominator, period, assumptions, and attribution grade alongside the percentage.
Grade attribution separately from value status
The attribution grade below is an internal decision aid, not an industry standard. Assign it to each material claim—not to the initiative as a whole—and consider design, execution, data quality, sample adequacy, and uncertainty. A pilot can have strong evidence for a touch-time effect and weak evidence for a financial conversion.
| Grade | Evidence basis | Claim posture |
|---|---|---|
| A | Randomized or similarly strong counterfactual with adequate sample and reliable outcome data | Strong causal support if execution, contamination, attrition, and outcome-data checks pass |
| B | Staggered rollout or well-matched comparison with reliable outcome data | Credible attribution with stated limitations |
| C | Adjusted pre/post evidence with known concurrent changes | Directional association; causal attribution limited |
| D | Modeled estimate, survey, or self-report without a credible comparison | No causal attribution; planning value only |
A lower grade does not make the information useless. It limits how confidently and precisely the organization should speak about it.
“Modeled annual potential, attribution grade D” is more useful than a precise ROI percentage when the inputs are still assumptions.
Use an Impact Ledger, not a collection of success stories
An Impact Ledger gives every initiative the same reporting structure.
| Field | What to report |
|---|---|
| Initiative and workflow owner | Name the change and the accountable operator |
| Unit, eligible population, and period | Define exactly what the result covers |
| Intervention and exposure | State what changed and how much work used it |
| Baseline and comparison | Show the reference point and method |
| Primary result | Report absolute and percentage change |
| Guardrails | Show quality, risk, customer, and exception results |
| Operational impact | Capacity, speed, throughput, quality, or risk movement |
| Captured value | State what reached output, spend, margin, or loss—with a range if appropriate |
| Full cost and net value | Include build and ongoing operating cost |
| Value status | Modeled potential, observed change, attributed impact, or captured value |
| Attribution grade by claim | A, B, C, or D for each material operating or financial claim |
| Assumptions and limitations | Name what could change the conclusion |
| Next decision | Scale, redesign, pause, stop, or continue measuring |
The ledger should contain weak, negative, and stopped initiatives as well as successes. Otherwise it becomes marketing rather than management.
Report differently for operators and executives
| Audience | What they need |
|---|---|
| Workflow owner | Which parts of the flow improved, where exceptions moved, and what to change next |
| Builder or technical owner | Reliability, latency, cost, edits, failures, version effects, and instrumentation gaps |
| Finance or risk | Assumptions, full cost, attribution, downside, and audit trail |
| Executive sponsor | Captured value, confidence, guardrails, and the decision required |
A practical cadence is:
- Weekly operating review: Workflow coverage, exceptions, latency, quality, incidents, and immediate corrections
- Monthly impact review: Baseline comparison, operational impact, cost, attribution strength, and scale/redesign decision
- After a material model, prompt, retrieval, policy, or workflow change: Rerun relevant evaluations and revalidate operating thresholds
- Quarterly portfolio review: Captured value across initiatives, cumulative cost, stopped work, concentration risk, and where to invest next
The report is successful when it changes a decision—not when the dashboard gains another chart.
What I would not call ROI
On their own, I would not call any of the following ROI:
- Employees trained
- Licenses provisioned
- Weekly active users
- Prompts, calls, or outputs generated
- A model evaluation score
- Gross employee estimates of hours saved
- Salary multiplied by theoretical time savings
- Headcount avoidance no one changed a hiring plan to capture
- Benefits reported without human-review, exception, implementation, and run costs
These can be inputs to a measurement system. They are not the final business result.
What this looks like in practice
In a prior operating role, before Rivington, I redesigned a high-volume expert-review workflow that supported more than 5,000 monthly inbound submissions.
The workflow was rebuilt around structured intake, explicit qualification logic, consistent routing, concise expert summaries, and human review at the points where judgment mattered. The reported result was an approximately 80% reduction in expert assessment time and more than 50 hours of manual work removed each week, with the expert-review step retained for material cases.
The important measure was not the volume of summaries produced. It was the reported change in expert assessment time across live workflow volume, with the judgment requirement intact.
The claim should stop there unless the organization can also show what the released capacity produced. If those hours supported more completed assessments, reduced a backlog, avoided external spend, or prevented a planned hire, that outcome can be measured separately as captured value. If not, the result remains a meaningful operating improvement—not automatic cash savings.
The public facts for this example do not, by themselves, establish a controlled counterfactual, a measured quality effect, or captured financial value. Retaining an expert-review step describes the workflow design; it does not prove quality held. In an Impact Ledger, the result should include the exact baseline period and comparison method supported by the underlying record, report any measured guardrail separately, assign the corresponding attribution grade, and avoid an ROI claim unless redeployment and value capture are documented.
That distinction makes the impact story more credible, not less impressive.
A 30-day way to begin
Thirty days is enough to design the measurement system, not necessarily to demonstrate impact. The observation period depends on workflow volume, normal variation, and how long downstream outcomes take to mature.
Week 1: Choose the work
- Select one recurring workflow with a visible consequence.
- Name the owner and the unit of work.
- Define the business outcome and consequential decision.
- Complete the Opportunity Map.
Week 2: Baseline reality
- Follow real work from entry to outcome.
- Quantify volume, effort, delay, quality, exceptions, and current cost.
- Identify routine work, genuine judgment, and informal workarounds.
- Record gaps that make the workflow difficult to observe.
Week 3: Design the evidence
- Write the impact hypothesis.
- Complete the Measurement Contract.
- Choose the comparison method.
- Define primary, exposure, guardrail, exception, and business measures.
- Add required events and reason codes to the build plan.
Week 4: Set up the decision
- Agree on scale, redesign, pause, and stop thresholds.
- Assign operating, technical, data, finance, and risk responsibilities.
- Create the first Impact Ledger entry using the baseline and modeled potential.
- Schedule the operating and impact reviews before the pilot begins.
At the end of 30 days, the company should have more than an AI idea. It should have a defined workflow opportunity, a baseline, an evidence plan, and a decision process.
The five-question impact-claim audit
Before an AI result reaches an executive or board report, ask:
- What unit of work changed?
- Compared with what?
- Did quality, customer, and risk guardrails hold?
- How did the operating change create economic value?
- What value did the company actually capture, net of cost?
If the team cannot answer all five, the result may still be promising. It is not yet a defensible ROI claim.
Start with one workflow that matters
Many companies do not need another enterprise-wide AI scorecard before they can learn. They need one consequential workflow, observed honestly and designed so improvement can be measured.
Rivington's free Workflow Friction Check can help identify where ownership, inputs, policy, exceptions, and measurement are breaking down.
For a deeper answer, the AI Workflow Diagnostic examines one recurring workflow over three weeks. The work includes a current-state map, quantified friction analysis, future-state design, AI and automation recommendations, a baseline and measurement plan, and a practical 90-day implementation roadmap.
Bring one workflow that is slow, manual, inconsistent, or failing to produce a measurable result. We will determine whether the problem is specific enough, consequential enough, and accessible enough to fix.
This guide provides general operating guidance, not legal, regulatory, accounting, employment, or compliance advice. Workflows involving employment, insurance, credit, healthcare, safety, or other consequential decisions require review by the organization's appropriate legal, compliance, security, privacy, risk, finance, and domain owners.
Sources and related reading
- Erik Brynjolfsson, Danielle Li, and Lindsey R. Raymond, “Generative AI at Work”, Quarterly Journal of Economics 140(2), 2025.
- Fabrizio Dell'Acqua et al., “Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of Artificial Intelligence on Knowledge Worker Productivity and Quality”, Organization Science, published online March 2026.
- Joel Becker et al., “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity”, METR, 2025. METR later described limitations in its follow-on experimental design as tools and usage patterns changed.
- National Institute of Standards and Technology, AI Risk Management Framework Core, especially the cross-cutting Govern function and the Map, Measure, and Manage functions.
- Rivington Solutions, Workflow Teardowns.
Related reading