AI Transformation Case Study: Creating Metrics for AI-Supported Workflows

🧭 Dojo Compass

Module: Finance, Risk Management and Long-Term Resilience

Focus Area: Technology, AI and Future Readiness

Key Issue

BrightPath Services was a growing SME that had enthusiastically adopted artificial intelligence across multiple parts of its business.

Employees used AI to draft documents, summarize information, analyze data, prepare customer communications, conduct research, support software development, and assist with internal decision-making. Managers regularly heard positive reports about productivity gains.

“This used to take me three hours.”

“AI prepared the first draft in five minutes.”

“We can now analyze much more information than before.”

The company viewed these developments as evidence that its AI transformation was succeeding.

However, senior management eventually recognized an important problem.

The company had no consistent way to determine how much value its AI-supported workflows were actually creating.

Individual employees were reporting time savings, but those reports did not capture the entire workflow. A task might take less time to complete initially but require additional review and correction. An AI-generated output might be produced quickly but have limited practical usefulness. One workflow might be highly replicable across the company, while another depended heavily on a particular employee’s expertise.

BrightPath was measuring activity and enthusiasm.

It was not consistently measuring value.

The company’s leadership therefore faced a new question:

How could an SME create a practical framework for evaluating whether AI-supported workflows were actually improving the way work was performed?

Facts

BrightPath employed approximately 250 people across operations, sales, finance, customer service, technology, and management.

AI adoption had emerged from the bottom up.

Some employees had integrated AI deeply into their daily workflows. Others used it occasionally. Different teams had adopted different tools and developed their own approaches to prompting, reviewing, and incorporating AI outputs into their work.

The company had collected several examples of apparent success.

The finance team reported that AI-assisted reporting had reduced the time required to prepare certain management reports. The sales team used AI to accelerate the preparation of customer materials. The technology team used AI-supported coding tools. Customer service employees used AI to help draft responses and summarize customer interactions.

Yet when management attempted to compare these use cases, it encountered a problem.

There was no common measurement framework.

A workflow that saved 30 minutes might appear successful, but management did not know whether the output required extensive human revision. Another workflow might save little time but significantly improve the quality or consistency of the output. A third might work extremely well for one employee but prove difficult to replicate across the broader organization.

The company also discovered that measuring task time alone could be misleading.

For example, AI reduced the initial drafting time for a customer report from two hours to 20 minutes. However, the employee then spent 45 minutes reviewing, correcting, and adapting the output. The workflow was still faster, but the actual gain was substantially smaller than the initial headline suggested.

In another case, an AI system produced an analysis in minutes, but the analysis was sufficiently useful that managers could evaluate opportunities they previously would not have had the time or resources to consider.

The value was not simply time saved.

It was additional capability created.

BrightPath concluded that it needed to evaluate AI-supported workflows as complete systems rather than measuring only the speed of individual AI tasks.

Solution

The company created a simple AI-Supported Workflow Metrics Framework.

Rather than attempting to develop a complex enterprise-wide measurement system, BrightPath identified several factors that could be applied consistently across different functions.

The first metric was total workflow time.

Employees measured the approximate time required to complete a task before and after AI support. Importantly, the measurement included preparation, prompting, review, correction, and finalization.

The relevant question was not:

How quickly did the AI produce an answer?

It was:

How much total human and technological time is required to produce a usable result?

The second metric was human intervention.

BrightPath assessed how much human involvement remained necessary at different stages of the workflow. This included the need for human preparation, prompting, supervision, correction, judgment, and approval.

A workflow requiring minimal human intervention after initial setup represented a different form of value from one that required constant supervision.

Third, the company measured workflow dependencies.

Management examined how many people, systems, information sources, and approval stages were required to complete a process.

AI could create value not only by accelerating an individual task but also by reducing organizational complexity.

For example, if employees previously needed to contact three colleagues to gather information but could obtain and synthesize that information through an AI-supported knowledge system, the value included a reduction in dependencies and interruptions.

Fourth, BrightPath measured cost.

This included direct software costs but also implementation, training, human review, maintenance, and other relevant expenses.

A workflow that produced modest efficiency gains at very low cost might create greater overall value than a more sophisticated system requiring substantial investment and ongoing support.

Fifth, the company introduced a measure of output utility.

Employees and managers assessed whether the AI-supported output was actually useful for its intended purpose.

This was particularly important because faster production of low-value output created little real benefit.

The company considered questions such as:

  • Is the output accurate enough for its intended use?
  • Does it improve the quality of the underlying work?
  • Does it enable a better decision?
  • Does it increase the number of useful tasks that can be completed?
  • Would employees choose to use the workflow again?

Sixth, BrightPath evaluated replicability.

Some AI-supported workflows depended on highly skilled employees who had developed sophisticated personal prompting techniques. These workflows could be valuable but difficult to scale.

Others could be documented, standardized, and transferred across teams.

The company therefore asked:

Can another employee reproduce this result with reasonable consistency?

Finally, BrightPath introduced two broader strategic measures: scalability and value concentration.

Scalability considered whether the workflow could handle a greater volume of work without requiring proportional increases in human effort.

Value concentration examined whether AI was helping employees shift their time toward higher-value activities.

This last measure became particularly important to management.

The company did not want AI merely to help employees complete a larger volume of low-value work.

It wanted to know whether AI was allowing people to spend more time on activities involving customer relationships, judgment, creativity, problem-solving, and strategic decision-making.

Creating an AI Workflow Scorecard

BrightPath translated these principles into a simple scorecard that could be used to compare different AI-supported workflows.

The company assessed each workflow across the following dimensions:

  • Task and total workflow time: How much time does the complete process require?
  • Human intervention: How much preparation, supervision, correction, and approval is required?
  • Dependencies: How many people, systems, or process steps are necessary?
  • Cost: What is the total economic cost of operating the workflow?
  • Output utility: How useful, accurate, and actionable is the result?
  • Quality and reliability: How consistently does the workflow produce an acceptable result?
  • Replicability: Can other employees reproduce the process successfully?
  • Scalability: Can the workflow handle greater volume without proportional increases in resources?
  • Risk: Does the workflow create material technical, operational, legal, confidentiality, or other risks?
  • Value concentration: Does the workflow allow human time and attention to move toward higher-value activities?

The company did not attempt to create a false level of mathematical precision.

Some metrics could be measured numerically. Others required managerial judgment.

The objective was consistency rather than perfection.

Over time, BrightPath began creating a portfolio view of its AI-supported workflows.

Some were identified as high-value and scalable and became candidates for broader implementation.

Others produced useful results but required significant human intervention and remained specialized tools.

A number of experiments were discontinued because their apparent benefits did not justify their cost or complexity.

Key Takeaways

BrightPath’s experience illustrates several important principles for SMEs seeking to measure the value of AI.

First, AI activity is not the same as AI value. A company may have many employees using AI without knowing whether those uses are creating meaningful operational or strategic benefits.

Second, measure the entire workflow rather than the AI task in isolation. Fast output generation has limited value if the result requires extensive human correction, creates new dependencies, or introduces additional complexity elsewhere.

Third, time savings are only one dimension of value. AI can also create value by improving output quality, reducing dependencies, increasing capacity, improving consistency, enabling new forms of analysis, or allowing employees to focus on higher-value work.

Fourth, human intervention should be measured explicitly. One of the most important questions is not simply whether AI is involved, but how much human effort remains necessary to convert AI output into useful work.

Fifth, replicability determines whether a successful experiment can become an organizational capability. An AI workflow that works brilliantly for one individual is not necessarily a scalable company asset.

Sixth, SMEs do not need perfect metrics to make better decisions. A practical and consistently applied framework can be more useful than an elaborate measurement system that employees lack the time or resources to maintain.

For BrightPath, the result was a shift in how management thought about AI transformation.

The company stopped asking simply:

“Where are we using AI?”

It began asking:

“Which AI-supported workflows are creating the greatest value, why are they creating that value, and how can we improve or replicate them?”

This created a more disciplined approach to AI adoption.

Successful experiments could be identified and scaled. Weak experiments could be redesigned or discontinued. Management could compare workflows across different functions and begin allocating resources toward the areas of greatest demonstrated value.

The broader lesson was straightforward:

AI transformation becomes more powerful when companies stop measuring whether AI is being used and begin measuring how AI changes the complete workflow through which value is actually created.

Case Study Note

The case studies published by Business Warrior’s Dojo are intended primarily as tools for learning, discussion, and analysis.

They may be based on real business situations, publicly available case studies, professional experiences, or entirely hypothetical scenarios. In some cases, names and identifying details have been changed to preserve confidentiality. In others, facts, circumstances, timelines, or outcomes may have been substantially modified, combined, or simplified to better illustrate particular business issues or support discussion. Some case studies are entirely fictional and have been developed solely for educational purposes.


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *