Skip to content

AI

How Will Agentic AI Improve My Business?

Agentic AI improves a business by finishing bounded, repetitive work end to end. What the 2026 evidence shows about where it pays and where it fails.

By Espen Hareide17 min read
An open-plan office where three people work through routine casework.

Agentic AI improves a business in one specific way. It takes a bounded, repetitive, multi-step job, such as matching an invoice, resolving a tier-1 support case, or chasing a document through an approval chain, and finishes it end to end, including the write-back into your system of record. That is the whole mechanism. Everything else sold as agentic AI is either a chatbot with a new label or a project that will be canceled.

The distinction matters more in 2026 than it did a year ago, because the evidence has caught up with the pitch. Adoption is climbing. Reported profit impact is not. Understanding why that gap exists is the difference between a deployment that pays for itself in a quarter and one that joins the 40% of agentic projects Gartner expects to be canceled by the end of 2027.

Key Takeaways

  • Agentic AI pays where the job is bounded and the output is checkable. Document-heavy back-office work, tier-1 service resolution, and software delivery are where returns show up consistently. Open-ended "transform the company" mandates are where they do not.
  • Agent adoption is rising faster than measured profit. In McKinsey's 2026 survey of 1,719 leaders, 40% of organizations above $1B in revenue are scaling AI agents, up from 27% a year earlier, while the share reporting any EBIT impact stayed flat at 37% and true high performers stayed flat at 6%.
  • Mid-market firms are ahead on general AI adoption, not yet on agents. US Census Bureau data puts overall AI use at 37% among firms with 250+ employees and 32% at 100 to 249 employees, against a 19.8% national rate. Agent deployment specifically remains in single digits across nearly every business function.
  • The failure modes are predictable: buying agent-washed software, layering agents onto a workflow nobody redesigned, unbounded inference cost, and no identity or audit controls.
  • The EU AI Act's high-risk deadline moved, but not the whole Act. Annex III obligations shifted to 2 December 2027, while Article 50 transparency duties still applied from 2 August 2026.

What agentic AI actually is, and what it is not

An AI agent is software that is given a goal rather than a script, decides for itself which tools and systems to call in which order, and keeps going until the goal is met or it hits a boundary you set. Three properties distinguish it from what came before:

  • Goal-directed sequencing. It plans the steps rather than replaying a recorded sequence.
  • Tool access with write permission. It does not just read your CRM and summarize it. It creates the credit note, updates the ticket, posts the journal entry.
  • Exception handling. When the input is malformed or the case is unusual, it adapts or escalates rather than failing the batch.

Set against the two things it is most often confused with, the line is clear enough. A chatbot answers; an agent acts. Robotic process automation replays a fixed click path and breaks the moment a field moves; an agent re-plans. That third property, what happens on the exceptions, is where most of the operational value and most of the operational risk sit.

The common misconception is that agentic AI is "AI that replaces a department." It is not. It is a way to remove the sequencing, handoffs and waiting from a process that a department currently performs. That is a smaller claim, and a far more bankable one. For a plain-language primer on the models underneath, see our beginner's guide to NLP and LLMs, and for the shorter framing of what an agent is, AI agents: your digital assistants of the future.

Agent washing is a real procurement problem

Gartner has been blunt about the supply side. In a June 2025 assessment, the firm estimated that only around 130 of the thousands of vendors marketing agentic AI actually qualify, with the rest rebranding assistants, chatbots and RPA without substantive agentic capability. Gartner named the practice "agent washing." Senior Director Analyst Anushree Verma put the underlying constraint plainly: "Most agentic AI propositions lack significant value or return on investment, as current models don't have the maturity and agency to autonomously achieve complex business goals or follow nuanced instructions over time."

For a mid-market buyer, that single statistic reframes the evaluation. The base rate says the product in front of you is probably not what it claims. Ask for the agent's plan trace, its tool-call log, and its behavior on a malformed input. Those three artifacts separate the real thing from the relabel.

The honest picture of returns in 2026

Two facts sit uncomfortably next to each other, and any business case has to reckon with both.

Fact one: AI adoption is genuinely broad, and mid-market firms are ahead of the national average. The US Census Bureau's Business Trends and Outlook Survey samples actual businesses rather than conference attendees. It found AI use running at 19.8% nationally in the collection period ending 3 May 2026. The size gradient is sharp: 37% of firms with at least 250 employees and 32% of firms with 100 to 249 employees, against under 20% for firms with fewer than 20 people. Between December 2025 and May 2026, use rose among firms with at least 20 employees and did not move meaningfully among smaller ones.

Bar chart of AI use by US firm size in May 2026: 37 percent of firms with 250 or more employees, 32 percent of firms with 100 to 249 employees, and 19.8 percent across all firms nationally.
Mid-market and larger firms adopt at roughly twice the national rate. Source: US Census Bureau Business Trends and Outlook Survey, 26 May 2026.

Note what that measures: AI use of any kind, not agents specifically. The two numbers are far apart, which is the point of the next fact.

Fact two: the profit line has not moved with adoption. McKinsey's The state of AI in 2026: On the road to ROI, published 25 August 2026 from a survey of 1,719 leaders, found that 40% of respondents at organizations above $1 billion in revenue were scaling AI agents, up from 27% the year before. Over the same period the share attributing any EBIT impact to AI held flat at 37%, and the share qualifying as AI high performers, meaning at least 5% of EBIT attributed to AI with impact described as significant, stayed flat at 6%. Meanwhile 80% of people using AI in their own role said it improved their individual productivity.

Stanford HAI's 2026 AI Index Report describes the same shape from a different angle. 88% of surveyed organizations use AI in at least one capacity and 70% have generative AI in at least one business function, yet agent deployment remains in single digits across nearly all business functions.

Lollipop chart showing that 88 percent of organizations use AI somewhere and 80 percent of users report personal productivity gains, while only 40 percent of billion-dollar organizations are scaling agents, 37 percent report any EBIT impact, and 6 percent qualify as AI high performers.
Individual productivity gains are widely reported. Organizational profit impact is not. Sources: Stanford HAI 2026 AI Index; McKinsey, The state of AI in 2026.

The most-quoted version of this gap comes from MIT's NANDA initiative, whose 2025 report The GenAI Divide: State of AI in Business concluded that 95% of enterprise generative AI pilots delivered no measurable P&L impact. That figure deserves its caveat. It rests on 52 executive interviews, surveys of 153 leaders and analysis of roughly 300 public deployments, and it measures pilots rather than mature deployments. Treat it as a directionally useful warning about pilot design, not as a precise failure rate.

What all three datasets agree on is the operative point. Individual productivity gains are real and widely felt, and they do not automatically become company profit. Something has to convert saved minutes into either removed cost or added revenue. That conversion is a management act, not a software feature.

Where agentic AI is measurably improving businesses

Four patterns recur in the 2026 evidence. They share a structure: a bounded job, a checkable output, and a system of record to write to.

1. Document-heavy back-office work

Accounts payable matching, claims intake, procurement review, contract abstraction, compliance checks. This is the least glamorous category and the most consistently profitable one. MIT's NANDA analysis found the highest returns in back-office automation, specifically document processing, procurement and risk review, and found that deployments built with specialized external vendors succeeded at roughly twice the rate of internal builds, because domain fluency and workflow integration mattered more than interface polish. That ratio comes from the same contested v0.1 deck as the 95% figure, so treat it as a strong directional signal rather than a measured constant.

The reason this category works is structural. The input arrives in a queue, the output is verifiable against a source document, and the write-back target is a single system. An agent that is wrong gets caught by a reconciliation step you already run.

2. Tier-1 service resolution, with write access

The value in customer service comes from resolution, not deflection. Deflection counts conversations a human never touched. Resolution counts problems actually solved. An agent that can only retrieve answers moves the first number. An agent that can issue the refund, reschedule the delivery or reset the entitlement moves the second.

Klarna is the instructive case in both directions. In February 2024 the company announced that its OpenAI-powered assistant had handled 2.3 million conversations in its first month, two-thirds of its customer service chats, doing the equivalent work of 700 full-time agents. Average resolution time fell from 11 minutes to under 2, repeat inquiries dropped 25%, and Klarna estimated a $40 million profit improvement for 2024. Worth reading precisely: Klarna framed this as work-equivalency, a throughput comparison, not as 700 people dismissed.

By May 2025, CEO Sebastian Siemiatkowski told Bloomberg that cost had been "a too predominant evaluation factor," producing "lower quality," and Klarna began rehiring human agents into a hybrid model, with AI on routine high-volume queries and people on escalations and high-value accounts.

The lesson is not that the automation failed. Two-thirds containment is a real result. The lesson is that the optimization target was wrong. The company optimized for cost removed rather than for problems resolved at acceptable quality and cost.

3. Software delivery

Stanford's 2026 AI Index Report puts measured productivity gains at around 26% in software development, against 14% to 15% in customer support, a notably wide spread between two task types. McKinsey's 2026 survey adds a second-order effect worth noting in budget planning: nearly a third of respondents said their organization decided against buying at least one software product or feature because they could build the functionality in-house with agentic coding tools.

That is a genuine change in the build-versus-buy calculus for the specific class of internal tooling that used to be too small to justify a project and too annoying to do by hand.

4. Coordination-heavy workflows

Some processes are slow not because any step is hard, but because of sequencing, handoffs and waiting. Order exceptions that bounce between sales, credit and logistics. Onboarding that stalls between HR, IT and facilities. McKinsey's Seizing the agentic AI advantage points at exactly this profile when it describes which processes justify redesign: high coordination overhead, rigid sequences that delay responsiveness, and frequent human intervention for decisions that could be data-driven. Where the bottleneck is coordination rather than cognition, an agent that can hold the whole thread, chase the missing input, update each system and escalate on a clock compresses cycle time without anyone doing less work.

Why more than 40% of agentic projects get canceled

Gartner, from a poll of more than 3,400 organizations actively investing in the technology, forecast that over 40% of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls.

40%+
of agentic AI projects are forecast to be canceled by the end of 2027, on cost, unclear value and weak risk controls. Not on model capability. Source: Gartner, 25 June 2025.
Donut chart showing Gartner's forecast that over 40 percent of agentic AI projects will be canceled by the end of 2027, with the remainder continuing.
Gartner attributes cancellations to escalating cost, unclear business value and inadequate risk controls. Source: Gartner, 25 June 2025.

Underneath those three headings, four concrete failure modes account for most of it.

Buying agent-washed software. You paid for autonomy and received a retrieval chatbot. The symptom: the vendor demonstrates conversation quality and avoids demonstrating a write-back to your system of record.

Layering agents onto an un-redesigned workflow. This is the one that quietly kills the business case. If an agent drafts the response and a human still reviews every draft, re-keys the result and sends it manually, you have added a step and removed nothing. In Seizing the agentic AI advantage, McKinsey argues the same point from the other direction: the organizations getting impact rebuilt the workflow around the agent instead of inserting the agent into the existing sequence.

Unbounded inference economics. Agentic workflows are iterative by construction. Plan, call a tool, evaluate, re-plan. A single completed task therefore consumes many multiples of the tokens a one-shot chatbot answer does. Per-token prices have fallen steeply, and Stanford's 2025 AI Index measured a more than 280-fold drop in the cost of querying a GPT-3.5-level model between November 2022 and October 2024, from about $20 to $0.07 per million tokens. Per-task cost has not followed automatically. A pilot priced on chatbot-shaped usage can be badly wrong at production volume, so price the unit of completed work rather than the unit of text.

No control plane. Permissions, identity, logging and a kill switch, discussed below. Its absence is what turns a successful pilot into a project security refuses to let scale.

What separates the businesses that get value

Five things, drawn from where the data consistently points.

  1. Redesign the workflow, do not decorate it. Pick the process, map where time actually goes, and rebuild the sequence assuming the agent does the middle. If your post-agent process diagram looks like the old one with a robot icon added, stop.

  2. Choose jobs with a checkable output. If you cannot state in one sentence how you would know the agent got it right, you cannot operate it, measure it, or defend it. Getting that sentence written down is most of the work, which is the same problem we hit in AI-driven project specifications. Invoice matched to PO and receipt. Ticket resolved without reopen within 14 days. Contract clause extracted and confirmed against source.

  3. Buy vertical, build thin. The evidence, chiefly MIT NANDA's vendor-versus-internal comparison with the caveats noted above, favors specialized vendors for the domain logic and internal build for the glue. The reverse, a generic platform plus heavy internal domain build, is the expensive path.

  4. Instrument before you scale. Baseline the current cost and cycle time of the process before the agent touches it. Teams that skip this can never prove impact afterwards, which is a large part of why "unclear business value" kills projects.

  5. Keep a human at the consequential step. Not at every step, which destroys the economics, but at the ones with irreversible external effect: money leaving, contracts binding, customers being told something final. Klarna's reversal is what the alternative costs.

The controls to put in place before you scale

An agent with write access to your systems is a new class of privileged identity. Three areas need to be settled before volume goes up, not after.

Treat each agent as a non-human identity. Scoped permissions, short-lived credentials, rotation, and a revocation path that works in minutes. Machine identities already outnumber human ones heavily in most estates. Palo Alto Networks' 2026 Identity Security Landscape, a survey of more than 2,900 cybersecurity decision-makers, puts the ratio at 109 machine identities per human, AI agents included, and reports that nine in ten organizations suffered a successful identity-related breach in the previous twelve months. An agent that inherits a long-lived service account with broad rights is a common way a useful pilot becomes an unacceptable production risk.

Design for the known attack classes. The OWASP GenAI Security Project's Q1 2026 exploit round-up cataloged eight significant incidents between January and April 2026, spanning identity and privilege abuse, unsafe autonomy, supply-chain compromise, prompt injection, remote code execution and internal data exposure. The recurring themes were excessive agency, misconfigured permissions and weak validation, plus human over-trust in agent output. Assume the agent's inputs are hostile, cap what it can do without approval, and log every tool call with enough context to reconstruct a decision.

Read the regulatory timeline correctly. The EU AI Act's high-risk deadline moved, and it is easy to draw the wrong conclusion from that. Regulation (EU) 2026/1744, the Digital Omnibus on AI, was published in the Official Journal on 24 July 2026 and entered into force on 27 July 2026. It deferred obligations for standalone Annex III high-risk systems from 2 August 2026 to 2 December 2027, and for AI embedded in products under Annex I to 2 August 2028. Three things did not move: Article 50 transparency and AI-content-labeling duties still applied from 2 August 2026, the prohibited-practices regime has been in force since 2 February 2025, and GPAI provider obligations since 2 August 2025.

The practical reading for a mid-market operator is that you have more time to complete conformity work on high-risk use cases, and no additional time at all on disclosure. Norwegian companies have a further wrinkle. The AI Act is EEA-relevant but is not automatically Norwegian law: it has to be incorporated into the EEA Agreement by a Joint Committee decision and then given effect domestically. EFTA's own EEA-Lex record for the regulation still listed it as pending incorporation as of September 2026, so the transparency duties binding an EU competitor from August 2026 do not yet bind a Norwegian one on the same date.

Either way, the controls the high-risk regime will eventually require, meaning human oversight at consequential steps, activity logs, permission review and a documented kill switch, are the same controls that make an agent operable and auditable. Building them now is not compliance theater. It is the thing that lets you scale.

A 90-day sequence for a mid-market team

  1. Days 1–15

    Pick one process and baseline it. Choose a high-volume, rules-heavy workflow with a verifiable output and a single system of record. Measure current volume, cost per unit, cycle time, and current error and rework rate. No agent yet.

  2. Days 16–45

    Redesign, then build or buy. Draw the target workflow assuming the agent performs the middle and a human approves the consequential step. Only then evaluate vendors, and require a plan trace, a tool-call log, and a demonstration on deliberately malformed input. Set the permission scope and the escalation rule before the first production call.

  3. Days 46–75

    Run in shadow, then in a narrow lane. The agent proposes, a human decides, and you record the agreement rate. When agreement is stable, let it act autonomously on the lowest-risk slice only, with every action logged and reversible.

  4. Days 76–90

    Prove the unit economics. Compare cost per completed unit of work against the baseline, with inference and human-oversight time both included. Decide expansion on that number. If it is not better than the baseline, the answer is a different process, not a bigger model.

Timeline of a 90-day agentic AI deployment sequence: baseline the process on days 1 to 15, redesign then build or buy on days 16 to 45, run in shadow then a narrow lane on days 46 to 75, and prove the unit economics on days 76 to 90.
The agent does not touch the live process until day 46. Sequence synthesized from McKinsey, MIT NANDA and OWASP guidance cited above.

How to measure it so the number survives a CFO conversation

MetricWhat it tells youCommon trap
Resolution rateShare of cases the agent finished, end to endReporting deflection (untouched by a human) as if it were resolution
Cost per completed unitInference, tooling and human oversight, per finished itemCounting tokens only, omitting the review time you created
Exception rateShare escalated to a humanA falling exception rate can mean better handling, or silent wrong answers
Cycle timeElapsed time from intake to closedImproving step time while total handoff time is unchanged
Rework rateShare of agent outputs later correctedMeasured too early, before downstream effects surface
Quality at constant costCSAT, accuracy or reopen rate, tracked against costOptimizing cost alone, which is the Klarna failure mode

The metric that decides the business case is cost per completed unit of work at constant or better quality. Every other number is diagnostic.

The bottom line

Agentic AI will improve your business to the exact extent that you point it at a bounded job with a checkable output, then redesign the surrounding process so the saved time actually leaves the cost base. That is a narrower promise than the market makes, and it is the one the 2026 evidence supports.

The organizations converting AI use into measured profit are still a minority, 6% by McKinsey's definition and unchanged year over year. What separates them is not model access, which everyone has, or ambition, which everyone claims. It is scope discipline, workflow redesign, instrumentation before deployment, and a control plane built early enough that scaling is a decision rather than an argument with security.

Start with one process. Baseline it first. Measure cost per completed unit of work. Expand only on that number.

Related reading

Sources

  • US Census Bureau, Large Firms With at Least 20 Employees Biggest AI Users, Business Trends and Outlook Survey, 26 May 2026. census.gov
  • McKinsey & Company, The state of AI in 2026: On the road to ROI, 25 August 2026. mckinsey.com, also reported by The Register
  • McKinsey & Company, Seizing the agentic AI advantage, June 2025. mckinsey.com
  • Stanford HAI, 2026 AI Index Report, Economy chapter, April 2026. hai.stanford.edu
  • Stanford HAI, 2025 AI Index Report, April 2025, on inference-cost decline. hai.stanford.edu
  • Gartner, Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027, 25 June 2025. gartner.com
  • MIT NANDA, The GenAI Divide: State of AI in Business 2025, July 2025 (v0.1 deck, third-party mirror). report PDF
  • Klarna, Klarna AI assistant handles two-thirds of customer service chats in its first month, 27 February 2024. prnewswire.com
  • OpenAI, Klarna's AI assistant does the work of 700 full-time agents, February 2024. openai.com
  • Forbes, Klarna Reverses On AI, Says Customers Like Talking To People, 18 May 2025, reporting Siemiatkowski's Bloomberg interview. forbes.com
  • Palo Alto Networks, 2026 Identity Security Landscape, survey of 2,900+ cybersecurity decision-makers. paloaltonetworks.com
  • OWASP GenAI Security Project, Exploit Round-up Report Q1 2026, 14 April 2026. genai.owasp.org
  • Regulation (EU) 2026/1744 (Digital Omnibus on AI), Official Journal, 24 July 2026. eur-lex.europa.eu
  • Gibson Dunn, EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines and Other Key Changes, July 2026. gibsondunn.com
  • EFTA, EEA-Lex factsheet for Regulation (EU) 2024/1689 (AI Act), incorporation status. efta.int
  • EU Artificial Intelligence Act, Article 14: Human Oversight. artificialintelligenceact.eu

About this article. Written by Espen Hareide, co-founder and partner at Apps. This is a research synthesis rather than a report of first-hand deployment experience: every figure above is drawn from the named published sources, with the retrieval date of 16 September 2026. Where a source is contested, such as the MIT NANDA 95% figure, the methodology and its limits are stated in the text.

FAQ

Is agentic AI worth it for a company with 100 to 250 employees?
Often yes, and Census data suggests firms this size are already ahead of the national average on AI adoption generally. The constraint is not company size but process shape. You need a workflow with enough volume to justify the setup, a verifiable output, and a system of record to write to. One well-chosen process beats a broad rollout.
Should we build agents ourselves or buy them?
The evidence favors buying specialized, domain-specific tooling and building only the integration layer. MIT's analysis found vendor-partnered deployments succeeded at roughly twice the rate of internal builds, on the same contested dataset as its headline 95% figure. The exception is narrow internal tooling, where agentic coding tools have made building cheap enough that nearly a third of McKinsey's respondents declined to buy software they previously would have.
How do we tell real agentic AI from a rebranded chatbot?
Ask for three things: the plan trace showing how the system decided its sequence, the tool-call log showing what it wrote to which system, and a live demonstration on a malformed input. Retrieval chatbots fail the second and third.
Will agentic AI reduce our headcount?
Stanford's 2026 AI Index reports that one-third of organizations expect AI to reduce their workforce in the coming year, and it documents real displacement in specific segments, with employment for software developers aged 22 to 25 down nearly 20% from 2024. Headcount reduction is still a poor primary target for a deployment. Klarna optimized for it, got lower quality, and rehired. Capacity redeployment is the more durable outcome.
What is the single most common reason these projects fail?
Inserting an agent into a workflow nobody redesigned. The agent works, the demo is impressive, and the process still takes the same number of human touches, so no cost leaves the business and the project loses its sponsor.