
Principal Basil C. Puglisi reads this week’s AI return research as a budgeting problem hiding inside a technology story. On October 8, Bain & Company reported that across 951 companies and 63 enterprise processes, organizations are converting roughly 38% of AI’s potential into deployed, measurable value. Bain calls that share the realization rate. A day later, a Harvard study of software firms drew attention for finding that AI coding agents produce more code without a matching rise in finished software, because human code review becomes the bottleneck.
Put the two side by side and the lesson for a business owner is plain. The value promised in the business case and the value that shows up in the ledger are different numbers, and the gap usually sits at the step where a person has to check the work. Puglisi’s consulting position is that the gap is not a sign the review step is wrong. It’s a sign the business case never priced the review step in the first place.
That’s the discipline behind Factics, the method Puglisi has used since 2012 to tie every claim to a fact, a tactic, and a KPI. A projected saving with no measured KPI behind it is an assertion. Bain’s realization rate is, in effect, a Factics KPI for the whole AI portfolio: what was promised, what was measured, and the ratio between them.
What Bain found about realized AI value
Bain’s numbers come from its Automation and AI Pathfinder Survey 2026. Traditional automation, after two decades of process mapping and change management, has realized 52% of its potential value, while generative AI is realizing 38% across functions. Nearly 40% of companies that measured AI cost savings landed below 10% despite targeting savings of 11% to 20%. Most are still increasing AI spending, and Bain notes that many fund new initiatives with savings from earlier automation efforts that fell short.
The finding Puglisi would put in front of a CFO first is about oversight. Bain reports that only about 7% of companies run fully autonomous agents in production, yet many investment cases are built on the economics of full autonomy. Its advice is direct: if a business case assumes full automation but production still needs human approval, the expectations should reflect that. Otherwise, as Bain puts it, a company may end up approving one model while operating another.
What the Harvard study found about review
Harvard economists Fiona Chen and James Stratton studied AI coding assistants and agents using engineering analytics data from Jellyfish. In their paper, Artificial Intelligence in the Firm: Bottlenecks in Software Production, they find little evidence that firms increase software output or reduce employment after adopting the tools. For AI agents, they trace that incomplete pass-through to code review. The average time to review a pull request rises 49%, the share of pull requests with changes requested nearly doubles, and reviews carry more comments. Firms increasingly use AI to help with review, but human reviewers still play the central role.
The authors’ model is the useful part for anyone outside software. AI raises the volume of draft work, and it also changes how much checking each unit of work needs. Both forces land on the review stage. Speed up the drafting and leave the review capacity alone, and the queue just moves.
Price the checkpoint, then measure it
Puglisi treats the review step as a design decision rather than friction. Under Checkpoint-Based Governance, a named human accepts, modifies, or rejects AI output at a binding decision point, and the record shows who decided and why. That is AI Governance. Responsible AI, in his vocabulary, is the machine-side checking that runs without human review, such as an automated test suite or an AI reviewer flagging a pull request. The Harvard data shows firms leaning on both, with people still carrying most of the review load. A business case that counts the agent’s output and ignores the checkpoint’s cost will overstate the return every time.
His recommended fix is a short Factics sheet for each AI use case, written before funding. The fact is the baseline: current cycle time, cost per transaction, and error rate. The tactic is the deployment, including where the checkpoint sits and who owns it. The KPIs are the realized numbers, tracked on the same baseline: end-to-end cycle time rather than drafting time, the share of AI output accepted without modification, review hours per unit of output, and the realization rate against the original projection. A marketing team using AI to draft campaign copy, for example, would track time from brief to approved asset, not words produced, and would log how often the reviewer had to modify or reject the draft.
Two rules follow. Fund the next wave at the realization rate the last wave actually delivered, as Bain suggests, not at the vendor’s projection. And when the review queue grows, treat it as the signal to invest in review capacity or tighten the agent’s scope, not as permission to drop the checkpoint.
What to watch next
Puglisi would watch two things through the end of the year. The first is whether more studies measure output at the finished-work level, features shipped or cases closed, instead of activity counts like lines of code or drafts generated. The second is whether vendors start reporting realization rates and review cost alongside hours saved. Until that reporting is normal, his guidance for clients is to build it themselves: count the checkpoint as part of the cost, measure the outcome the business actually sells, and let the realized number set the next budget.
Sources
- Chen, F., & Stratton, J. (2026). Artificial intelligence in the firm: Bottlenecks in software production [Working paper]. Harvard University. https://fion.ac/jellyfish.pdf
- Doddapaneni, P., & Heric, M. (2026, October 8). Realization rate: The metric reshaping enterprise AI. Bain & Company. https://www.bain.com/insights/realization-rate-the-metric-reshaping-enterprise-ai/
Questions readers ask
What is Bain’s AI realization rate?
Bain & Company defines the realization rate as the share of AI’s expected value that turns into measurable business results. In its October 2026 analysis of 951 companies and 63 enterprise processes, generative AI was realizing about 38% of its potential value, compared with 52% for traditional automation.
Why do many AI business cases miss their targets?
Bain reports that many investment cases assume the economics of fully autonomous agents, while only about 7% of companies run fully autonomous agents in production. Most AI still needs human approval, so a case that leaves out the cost of oversight overstates the return.
What did the Harvard study find about AI coding agents?
Fiona Chen and James Stratton found little evidence that firms increase software output or reduce employment after adopting AI coding tools. For AI agents, code review became the bottleneck: review time per pull request rose 49% and the share of pull requests needing changes nearly doubled.
What is Factics?
Factics is Basil C. Puglisi’s method, in use since 2012, for tying every claim to a fact, a tactic, and a KPI. Applied to AI ROI, it means setting a baseline before deployment and measuring realized results against it rather than projected savings.
How should a business measure AI ROI when people still review AI output?
Measure end-to-end outcomes such as cycle time and cost per finished unit, not drafting speed. Count review hours as a cost, track how often reviewers accept, modify, or reject AI output, and fund the next project at the realization rate the last one actually delivered.
#AIg
Leave a Reply