Operating strength / A practical guide

Test one AI workflow before you automate it

Illustrative situation: an AI tool produces a convincing summary of customer complaints. The manager spots that two different problems have been merged and one important complaint is missing. The summary reads well. Whether it helps the work is a separate question.

For an owner or operations lead considering AI for recurring work that a person currently checks.

01Name the task and the boundary

“Use AI in operations” cannot be tested. “Draft a weekly grouping of approved, de-identified complaint examples for a manager to inspect” can. State the input, proposed output and decision that follows. Name actions outside the trial, such as contacting customers, altering account records or deciding compensation.

Begin with synthetic examples when permitted real data and an approved tool environment are not available. Synthetic inputs can test mechanics, but they cannot prove performance on the business’s actual work. Removing names alone does not establish that a record is safe to upload. Check the tool, access, retention and data-use arrangements for the intended information. Whether customer or employee information may be used this way is a legal and contractual question for your own attorney.

02Make the baseline and errors visible

Have a competent person complete the task on a small, representative set before comparing the tool. Preserve their reasoning and the original inputs. Include incomplete, ambiguous and unusual examples. Define unacceptable errors before seeing the output; otherwise a fluent answer can encourage you to lower the standard after the fact.

NIST’s AI Risk Management Framework calls for documented task scope, human oversight and testing before deployment and during operation.1 It is a voluntary framework, not a certification of this worksheet or any product. The adaptation here is a practical record of what the trial is allowed to do and how a person will check it.

The workflow trial card

A trial needs a return path

  1. Bound

    Approved inputs, task and excluded actions

  2. Compare

    Human baseline and representative cases

  3. Inspect

    Omissions, errors and total review effort

  4. Decide

    Your call: continue narrowly, revise or stop

A failed input, quality or oversight check returns the work to the existing process.

03Count review and correction as work

Observe preparation, generation, checking, correction and exception handling. A short generation time does not show that the complete workflow is faster. If the reviewer must reconstruct every claim from the input, the draft may add work. Compare quality and effort at the level of an accepted output, not the moment text appears.

Ask whether the reviewer can reliably spot the important errors. “A human will check” is weak protection if that person lacks time, context or authority. Give the reviewer a way to reject the output and return to the previous process. Keep unsupported statements visibly unresolved rather than allowing the tool to fill gaps.

04An illustrative trial

For the complaint-summary example, the manager prepares invented records that include scheduling, billing and site-care issues. The requested output groups them, references the input record identifiers and lists uncertain cases. The manager checks omissions, incorrect groupings and claims that were never in the records.

If the output merges “late arrival” with “late invoice,” the manager records why that matters: the two issues have different process owners. A revised prompt might help, but repeated prompt editing on the same examples can overfit the trial. Keep a separate set of unseen examples for the next check. None of this establishes production readiness on customer data.

05Decide whether to continue, change or stop

Continue only within the tested scope and only if the accepted work meets the standard, the review burden is worthwhile and the responsible people can operate the controls. Record the model or product version and relevant settings because later changes can alter behavior. Define the event that triggers a new evaluation.

The strongest case against a trial is that the task cannot be checked safely or cheaply enough. High-consequence work may need specialist assessment or may be unsuitable for the proposed tool. Manual work or ordinary automation can be the better answer. Use the worksheet to make that decision before a pilot becomes an unexamined dependency.

Use this now

One AI workflow trial card

Use invented or properly approved inputs. Do not paste confidential records into this worksheet.

No submission or account is needed. This page does not send or store your answers. Download or print them before leaving; they are not saved by this tool. Your browser or extensions may retain form data.

What decision follows the output?
Synthetic or approved data; access and retention checked by whom?
Include difficult cases and a separate later check set.
State the standard before seeing results.
Preparation, checking, correction and exception handling.
Who can stop it and how does work continue?
Model/product, date and relevant settings.

Sources and limits

Source pages checked 10 October 2026 UTC. The worksheets are editorial aids; they have not been validated as predictive assessments.

  1. National Institute of Standards and Technology · AI Risk Management Framework: Core. AI RMF 1.0, 2023; checked 10 October 2026. Primary risk-management framework. Supports defining the task and context, measuring performance and risks, and assigning human oversight. This worksheet is not a compliance assessment or a certification of an AI system. NIST states that AI RMF 1.0 is being revised; this guide reflects the version checked on 10 October 2026.

Our interest

Heritage Intelligence is taking first conversations about paid operations, data and AI work. You can use this guide and its worksheet independently, without engaging Heritage.

Service work, The Read included, is a separate relationship: if we become interested in buying a business we are working with, we stop that work and tell the owner plainly before any purchase discussion begins.

This article is educational. It is not individual legal, tax or investment advice, an offer or a promise of results.

Your accountant, attorney, and family make every real decision with you.

A useful next readA useful metric ends with someone doing something →Choose measures that can change the decision to continue, revise or stop.