Why your AI workflow still needs so much supervision
AI can finish a draft in seconds while the work around it still takes hours. A practical guide to finding that effort, deciding what should change, and measuring whether your team has gained capacity.
By Rivington Solutions
An onboarding manager opens an AI-generated welcome package. The summary is clear. The checklist looks complete. The customer email is ready.
Before using it, she checks the signed agreement, reconciles a conflicting start date, corrects an access requirement, confirms who owns implementation, and updates the account record. Later, she follows up to find out whether anyone completed the setup.
The AI produced something useful. She still carried the workflow.
That distinction matters when deciding whether to expand an AI pilot. Faster output creates an opportunity. The work required to turn that output into a completed result determines how much of the opportunity the business captures.
Find the work surrounding the output
When reviewing an AI workflow, follow one case from arrival to completion. Watch what people do before the agent starts, between its actions, and after it reports success.
Four types of effort deserve attention:
- •Preparation: finding records, resolving conflicting information, and explaining the task.
- •Review and correction: checking sources, spotting omissions, and repairing the output.
- •Coordination: assigning owners, securing approvals, and moving information between systems.
- •Completion checks: confirming that actions succeeded and the intended outcome occurred.
Some of this work protects quality. Some compensates for missing context, unclear responsibility, or incomplete system connections. The useful question is which kind you are observing.
Microsoft’s May 2026 Work Trend Index found that 50% of AI users surveyed identified quality control of AI output as a human skill becoming more important. Review deserves an explicit place in the workflow and its economics. Source: Microsoft Work Trend Index
Measure the whole job
Suppose an onboarding team compares its existing process with an AI-assisted version:
| Human effort per onboarding | Existing process | AI-assisted process |
|---|---|---|
| Prepare information and create the draft | 30 minutes | 8 minutes |
| Review and correct the work | 10 minutes | 16 minutes |
| Coordinate handoffs and update records | 10 minutes | 8 minutes |
| Follow up and confirm completion | 10 minutes | 10 minutes |
| Total human time | 60 minutes | 42 minutes |
Illustrative assumptions, not benchmarks or a forecast.
Preparation and drafting became much faster. Review increased. Follow-up stayed the same. The net reduction is 18 minutes per completed onboarding.
At 100 onboardings a month, that would free 30 hours of capacity, assuming comparable case complexity and acceptable quality. Its financial value depends on how the team uses those hours, alongside the costs of software, implementation, and ongoing maintenance.
Also measure elapsed time and errors. A workflow can consume fewer working minutes while customers wait longer for an approval. It can move faster while creating problems that only become visible later.
Give the agent a workflow it can carry
In the example above, a person connects almost every step:
Current Workflow
A redesigned workflow gives the system a more explicit job:
Proposed Workflow
Missing or conflicting information routes to a named owner. The case waits for resolution.
Actions requiring approval route to an authorized reviewer. Customer commitments and access changes follow this route.
Failed verification routes to the responsible owner for correction.
Missing information goes to a named owner. A conflict in the contract goes to someone authorized to resolve it. Customer commitments and access changes follow defined approval rules. Failed checks keep the case open.
This is where harness development becomes useful.
An agent harness is the surrounding software and configuration that manages how an AI agent receives context, uses tools, retains progress, and checks its work. For a business workflow, those mechanics need to reflect actual responsibilities and operating rules.
Anthropic’s engineering work on long-running agents illustrates the importance of recording progress between sessions and testing whether completed work actually functions. Applying those principles to onboarding means preserving case status and checking real system outcomes. Source: Anthropic
The agent should be able to pick up an open case without losing a prior approval, repeating a completed action, or treating an attempted update as a successful one. Permissions and approval requirements need enforcement in the tools and workflow; written instructions alone are insufficient.
Put judgment where it changes the outcome
Reducing supervision starts with understanding why each intervention exists.
If someone repeatedly corrects the same account field, investigate the source and validation rule. If every case needs a manager to choose an owner, make the routing criteria explicit. If an unusual contractual commitment changes what the company must deliver, preserve accountable human review.
An approval request should arrive with the relevant evidence, the issue requiring judgment, and the consequence of the available choices. Otherwise, the reviewer has to reconstruct the case before deciding.
Completion also needs a definition. In this onboarding example, the team might require access checks to pass, required configuration to be verified, and the receiving owner to accept the customer handoff. The workflow remains open until those conditions are evidenced.
Inspect one workflow this week
Choose a small sample of recent cases, including routine work and exceptions. Ask five questions:
- 01How many human minutes did each completed case require, including preparation, review, corrections, and follow-up?
- 02Which interventions required judgment, and which repaired missing information or broken handoffs?
- 03What could the agent finish within its authority, and where did it need approval?
- 04When work paused or failed, could the next person or session resume from a reliable record?
- 05What evidence showed that the intended outcome had occurred?
Bring one workflow you still have to carry
A small sample can reveal failure patterns. Use a broader, representative baseline before projecting savings. Test proposed changes against ordinary cases, missing information, conflicting records, and failed actions before expanding the agent’s responsibility.
The most useful measure is total effort per completed case at an acceptable standard. It connects AI capability to the work your team actually needs to deliver.
Rivington helps teams identify where human effort accumulates around AI, clarify decision ownership, and design a practical path to implementation. The AI Workflow Diagnostic examines one consequential workflow over three weeks and produces a redesign and a 90-day implementation roadmap.
Is this happening inside a live workflow?
Rivington diagnoses and redesigns one important workflow in three weeks.