State of AI in Customer Support / Edition 01

The State of AI in Customer Support: A Mid-2026 Field Report

What vendors actually bill for, what the evidence can and can't prove, and what decides whether a deployment works.

Published
Author
Maks Zhashkevych
Reading time
25 minutes
12Vendors priced
75Sources catalogued
41Primary documents
Evidence cutoffJuly 31, 2026
Report contents

Zendesk's "automated resolution" includes 72 hours of customer silence, and bills at an estimated $1.50–2.00 (third-party estimates; no official list price). HubSpot's is "no human handoff within 72 hours," at $0.50. Crescendo charges $2.99 and puts a 24/7 human team behind every one. Three vendors, three prices, three different things sold under the same word.

Most "state of AI in CX" reports are written by support platforms surveying their own customers. This one is written by someone who had to make the thing work. I spent two months of 2026 building and operating an agentic customer-support system for two e-commerce pilot deployments — the integrations, the escalation logic, the approval flows, the parts that break. Around that work I compiled a research corpus: teardowns of 14 companies in and around the category, vendor pricing pages checked and re-checked between June 22 and July 31, a frozen funding tracker, and deployment case documentation.

Where I stand: Cadenflow is a commercial product-development and AI-implementation studio. I sell implementation work; I do not sell any of the support platforms compared here, and no client data appears in this report. Two pilots is two pilots — the operating experience behind this report is real but bounded, and I say so where it matters.

The promise this report makes: facts, vendor claims, third-party estimates, my own calculations, interpretations, and forecasts are explicitly distinguished. Every material number traces to a source or a published calculation — the claim ledger, frozen datasets, and source register are published alongside this report. Three widely-quoted statistics get corrected in Section 03, against their original sources.


The short version

I came at this market from five directions — what things cost, where the capital went, whether the performance claims survive their own sources, what happened after deployment, and what the law now requires. Eight findings came back. They converge on one:

The model is the cheap part. A text support resolution runs on the order of half a cent to a few cents at mid-2026 list prices. It sells for $0.05 to $2.99. The spread isn't inference — it's accountability: who defines the outcome, who maintains it when your refund policy changes, and who answers when a customer acts on what the bot told them. Every finding below is a version of that sentence.

  1. 1 · Twelve vendors, one word, no shared meter. Published prices for AI-handled support run ~$0.05 to $2.99, across at least seven commercial models — metered per token, per reply, per action, per conversation, per resolved conversation, per resolution, per verified resolution, or per seat. No honest cross-vendor multiple can be computed from that range, and this report doesn't compute one. → §2
  2. 2 · The capital went to the top three. In my frozen tracker of 23 disclosed equity rounds, 2026's three largest — Sierra $950M, Parloa $350M, Decagon $250M — account for ~87% of tracked 2026 dollars, all Series D/E at $3B–$15.8B valuations. The largest first-institutional round in the tracked 2026 set is $5M. That the platform race is effectively over is my strategic read, not a measured fact. → §1
  3. 3 · Consolidation stopped being a thesis in one quarter. Salesforce signed ~$3.6B for Fin. Zendesk closed Forethought, its sixth AI acquisition since 2024. NICE paid a reported $955M for Cognigy. Shopify shipped a free AI sales associate into every merchant's inbox; Klaviyo turned its Customer Agent on across a ~196,000-customer base. On July 22, OpenAI entered the application layer itself. → §1
  4. 4 · The headline percentages measure five different things. Containment, deflection, automated resolution, verified resolution, and human-free resolution are not the same number, and the marketed ones are gameable by the party reporting them. Against vendor claims of 67–85%: only ~10% of 2,400+ surveyed support professionals describe their own deployment as mature (Intercom, Jan 2026), and 14% of 700+ US support leaders say AI significantly improved resolution times (Hiver, Mar 2026). → §3
  5. 5 · Unaudited private revenue is a claim, not a number. Decagon has reported $30–35M annualized; Forbes, investigating independently, estimated ~$12M for 2025 — at a company valued at $4.5B. Decagon and Fin both publicly claim to win bake-offs against each other. The 11x.ai episode — fake customer logos, full-year ARR booked on three-month break clauses — showed how far the incentives stretch. Billions in valuation are priced on multiples of figures nobody has audited. → §3
  6. 6 · The rollbacks and the results are both real. Klarna's assistant now does work its CEO equates to 853 full-time agents (Q3 2025 earnings call) — and Klarna has been recruiting human agents since May 2025 as a quality tier. Both numbers grew. Commonwealth Bank reversed AI-driven call-center cuts within weeks. Gartner predicts both >40% of agentic AI projects canceled by end-2027 and 80% autonomous resolution of common issues by 2029. → §4
  7. 7 · Deployment quality, not model choice, is where outcomes diverged. Frontier models score near the ceiling of the category's own (vendor-authored) benchmark; a demo agent is a weekend project. The recurring failure factors in the documented cases are knowledge-base rot, escalation design, scope selection, and absent QA. All fourteen companies in my teardown set — from $15.8B Sierra to bootstrapped Crisp — attach mandatory onboarding, implementation, or account management to their core price point. The deployment is the product. → §5
  8. 8 · August 2, 2026 gives AI support its first hard compliance date in the EU. AI Act Article 50 requires people to be told they're interacting with an AI system unless that's obvious to a reasonably well-informed person. The Digital Omnibus delayed the high-risk obligations — explicitly not this one. Transparency violations can draw fines up to €15M or 3% of worldwide turnover (higher of the two; lower for SMEs). Operational analysis, not legal advice. → §6

1 · The consolidation year

AI customer support is the one enterprise-AI category where the revenue, the disappointment, and the consolidation are all real at the same time.

The revenue is real — with labels. Sierra reports $100M ARR (November 2025) and $150M (February 2026) — company-reported figures — and raised $950M in May at a post-money valuation above $15B, reported as $15.8B. Salesforce reported $1.2B of Agentforce ARR, up 205% year over year, in its Q1 FY27 earnings. Zendesk told press it reached ~$200M of AI ARR in 2025, targeting up to $500M in 2026 (private company; unaudited). Fin crossed a company-reported $100M ARR — enough that Intercom renamed the entire company after the product on May 14, one month before Salesforce agreed to acquire it for ~$3.6B.

The funding data shows concentration. In my frozen tracker — 23 disclosed equity rounds in AI customer support and service from January 2024 through July 2026, published alongside this report — the three largest 2026 rounds account for approximately 87% of tracked 2026 dollars ($1.55B of $1.79B), and all three are Series D/E rounds at $3B–$15.8B valuations. The largest first-institutional financing in the tracked 2026 set is $5M (Cue, a regional CX-automation startup). The tracker under-captures small seed rounds, so deal-count trends can't be read from it — but the dollar-concentration conclusion is robust to that bias.

Cadenflow / Fig. 01

2026 capital concentration

Funding concentration data for tracked 2026 AI customer support equity rounds.
MeasureValue
Sierra round$950 million
Parloa round$350 million
Decagon round$250 million
Top three total$1.55 billion, approximately 87%
Tracked 2026 total$1.79 billion
Top three stagesSeries D / E
Top three valuations$3 billion–$15.8 billion
Largest first-institutional round$5 million
Cadenflow calculation from a frozen tracker of 23 disclosed equity rounds, January 2024 through July 2026. The tracker under-captures small seed rounds, so dollar concentration is robust to that bias; deal-count trends are not inferred.

The data does not establish that early-stage funding has stopped or that no new entrant can succeed. My read: the horizontal platform position is now contested from three directions at once — funded leaders above, free bundles below, and model labs moving up-stack. That is a hard place to start a company, and a fine place to sell implementation.

The incumbents bought instead of building. Zendesk has acquired six AI companies since 2024, announcing its largest deal in two decades (Forethought) in March 2026, since closed. NICE paid a reported $955M for Cognigy (~25x revenue, per press coverage). Salesforce closed Qualified in April and signed Fin in June — signed, not closed; regulatory clearance is expected to run into early 2027.

The platforms bundled the baseline away. Within thirteen months, "AI answers your customers using your data" became a free or near-free checkbox inside software merchants already pay for: Shopify's June 2026 edition put a free AI sales associate inside Shopify Inbox. Klaviyo turned on its Customer Agent across a ~196,000-customer base. HubSpot cut its Breeze agent to $0.50 per resolved conversation. Gorgias, Zendesk, and Salesforce all sell native agents inside their own suites.

And the model labs entered the application layer. On July 22, OpenAI launched Presence — a first-party enterprise product for AI voice and chat support agents, with design partners including BBVA, SoftBank, and IAG, deployed by OpenAI's own forward-deployed engineers at what The Register called "boots-on-the-ground prices." Anthropic took the opposite route — powering the application layer (default build-agent model in ServiceNow; the model behind Fin) rather than competing with it.

On valuations, with bases stated: Sierra's round prices at roughly 79x forward revenue (press-reported valuation over a ~$200M forward estimate); Decagon at ~130x its own reported annualized revenue — or ~375x Forbes's independent estimate; Wonderful at ~110x company-reported revenue. For contrast: Freshworks, a public helpdesk vendor, trades near ~3.7x trailing sales (my calculation from market data, July 31, 2026 — Zendesk has been private since November 2022, so no current public multiple exists for it). And Influx, a 13-year-old, ~1,000-person support BPO, carries a third-party enterprise-value estimate of $12–18M — under 1x estimated revenue. The spread between what the market pays for "AI agent" revenue and "support labor" revenue, on these estimates, is about two orders of magnitude.


2 · What a resolution costs — and what vendors actually bill for

Support became the proving ground for outcome-based pricing because it's the one AI domain with a countable outcome. The result, in mid-2026, is not one price for one unit — it's at least seven different commercial models wearing one price-per-something costume.

The pricing models, June–July 2026

Billing modelVendor (product)PriceWhat the unit actually is
Token-metered conversationCrisp / Hugo~$0.05–0.10 per AI-handled conversationCredits burn per token; conversation = first AI message to resolve/route/escalate
Per action, conversation, or seatSalesforce Agentforce~$0.10/action; ~$2/conversation; $125/user/moThree meters sold side by side, in one product
Per replyZipchat$0.20 per replyEvery AI reply billed, resolved or not
Per resolved conversationHubSpot Breeze$0.50AI action with no human handoff within 72h; cut from $1.00 on Apr 14, 2026
Per resolved ticket (outcome)Yuma~$0.60–0.70 (third-party estimate)No public list price; $850–1,200/mo packages; ROI guarantee
Per resolution, bundledRedo$0.85One module inside a returns platform whose core is free
Per automated resolutionGorgias$0.90–1.00 — plus the normal per-ticket feeThe same ticket is metered twice
Per resolutionFin (Intercom)$0.99The number that anchored the market; 50-resolution monthly minimum
Per conversation or per resolutionDecagon~$0.99 / ~$1.50 + platform fee (estimates)Sales-negotiated; vendor concedes definitions have "gray areas"
Per automated resolutionZendesk~$1.50–2.00 (third-party estimates)No official list price; "resolution" includes 72h of customer silence
Outcome contractSierra~$1.00–2.50 per resolution (reported)$150–200K/yr platform minimums; implementation billed separately
Verified resolution, humans includedCrescendo$2.9924/7 human backup in the price; the vendor bought a BPO to deliver it
Cadenflow / Fig. 02

Pricing models, June–July 2026

Published AI customer support prices and billing definitions, June to July 2026.
VendorUnit and price
Crisp / Hugo~$0.05–0.10 per AI-handled conversationToken-metered conversation. Credits burn per token. A conversation ends at resolution, routing, or escalation.
Salesforce Agentforce~$0.10/action; ~$2/conversation; $125/user/monthPer action, conversation, or seat. Three meters are sold side by side. The seat price is not plotted on the unit-price axis.
Zipchat$0.20 per replyPer reply. Every AI reply is billed, whether the conversation resolves or not.
HubSpot Breeze$0.50Per resolved conversation. An AI action with no human handoff within 72 hours.
Yuma~$0.60–0.70Per resolved ticket. Third-party estimate. No public list price; packages run $850–1,200/month.
Redo$0.85Per resolution, bundled. One module inside a returns platform whose core product is free.
Gorgias$0.90–1.00 plus the normal per-ticket feePer automated resolution. The same ticket is metered twice.
Fin (Intercom)$0.99Per resolution. Includes a 50-resolution monthly minimum.
Decagon~$0.99/conversation; ~$1.50/resolution plus platform feePer conversation or resolution. Sales-negotiated third-party estimates; the two points are separate meters.
Zendesk~$1.50–2.00Per automated resolution. Third-party estimate. No official list price; 72 hours of customer silence can count.
Sierra~$1.00–2.50 per resolutionOutcome contract. Reported range, with $150,000–200,000 annual platform minimums and separate implementation.
Crescendo$2.99Verified resolution, humans included. The price includes 24/7 human backup.
These are different units, fees, and definitions. No single cross-vendor multiple can be computed honestly from this range. Prices were checked between June 22 and July 31, 2026; estimates are labeled in the data.

Full snapshot — twelve vendors with fees, minimums, definitions, check dates, and confidence levels — in pricing-snapshot.csv. These are different units. No single "price of a resolution" multiple can honestly be computed across this table, and this report doesn't.

Two reference points bracket that table, both with the math shown. Inference cost: a text support resolution runs on the order of half a cent to a few cents at mid-2026 list prices — e.g., a 4-turn, tool-using resolution at ~6,000 input / ~1,200 output tokens costs ≈$0.005 on GPT-5.4-mini ($0.375/$2.25 per million tokens) or ≈$0.012 on Claude Haiku 4.5 ($1/$5), before retries, evaluation overhead, and orchestration (scenarios in calculation-notes.md; real deployments run higher). Human cost: commonly cited industry estimates put a human-handled ticket at roughly $5–12 fully loaded in Western teams (third-party estimates; varies by channel and geography), and an offshore support VA at $900–1,300/month all-in before management time.

What the data establishes: list prices for AI-handled work sit far above raw inference and far below human handling, on meters that differ vendor by vendor. My read: the price tracks the accountability wrapped around the model. At $0.05 you buy metered tokens with a UI. At $0.99 you buy guardrails, integrations, and a definition someone will argue about. At $2.99 you buy an outcome with a human safety net and someone to call. Consistent with that read: Redo bundles AI resolutions at $0.85 inside a free returns product, and PolyAI — after eight years and $200M+ invested in proprietary voice models — opened its platform to self-serve builders in May 2026, free for the first two months. The dialog layer, on its own, is pricing like a commodity.

The pricing-model war has named camps

Pro-outcome: Sierra (outcome contracts down to saved cancellations and upsells), Yuma ("you only pay for tickets the AI fully resolves," with a 100% ROI guarantee), Zendesk (per resolution confirmed by an independent evaluation model — verification itself becoming product).

Anti-outcome, in public: Parloa published a manifesto in Forbes in January — sponsored content, worth reading anyway — calling outcome-based pricing "the most expensive myth in enterprise AI": attribution complexity creates "endless reconciliation meetings," and vendors capture the customer's efficiency gains. Siena published "Why Outcome Based Pricing in AI Hurts Customer Service," arguing vendors game what counts as a resolution.

The buyers, voting with contracts: Decagon offers both meters and says most of its customers choose per-conversation — for predictability — while conceding that resolution definitions have "gray areas [that] lead to billing disagreements."

The dispute is structural, because "resolution" is defined by the party that bills for it. Zendesk's definition includes 72 hours of customer silence — indistinguishable from abandonment. HubSpot's includes "no human handoff within 72 hours." And when the automation works, the bills surprise: one Gorgias merchant reported a $14K overage on top of a $13.5K/yr contract. The community's name for it is the success tax — the better it works, the more you pay, at a rate fixed when the work was still expensive.

What buyers should take from this

The $0.50–0.99 anchors are marketing prices for the easy half of the volume. The questions that matter in 2026: who defines a resolution, what happens to the price as inference keeps deflating, and what share of your bill is the wrapper rather than the work. The most honest pricing designs this year either meter against actual model cost with a hard stop, or price the outcome with the humans included. Everything in between is negotiating leverage.


3 · The claims vs. the evidence

Five metrics, one percentage

Every sales deck leads with a percentage. These are the five things that percentage might actually measure — and how each one flatters:

MetricWhat it measuresHow it flatters
ContainmentConversation stayed in the AI channelCounts customers who gave up and left
DeflectionNo human ever touched itCounts abandonment as success
Automated resolutionAI marked it resolvedVendor defines "resolved"; silence can count (Zendesk: 72h)
Verified resolutionAn evaluation layer confirmed the fixBest current standard; verifier often belongs to the vendor
Human-free resolutionProblem actually solved, zero human involvementThe rarest and least-marketed number

Marketed figures in the 67–85% range typically live in the first three rows. They are not comparable to each other, and a gap between two vendors' headline numbers is as likely to be a definition difference as a product difference.

The honest denominators

Three surveys, each with its sample and field date attached, anchor where the market actually is. Intercom's survey of 2,400+ support professionals (published January 2026): 87% of leaders plan AI investment in 2026 — and only ~10% describe their current deployment as mature. Hiver's survey of 700+ US support leaders (March 2026): 14% say AI significantly improved resolution times; roughly half say it hasn't meaningfully lowered cost per ticket. And Gartner's survey of 5,728 customers (fielded December 2023): only 14% of customer service issues were fully resolved in self-service — a pre-agent self-service baseline, not a measure of today's AI agents, but a standing reminder that most "self-served" journeys historically didn't end in resolution.

Vendor case files show the same shape from inside. Hugo markets "up to 60%" automated resolution while its own named case studies land at 40–60%. Salesforce's acquisition PR cites Fin resolving an average of 76% of support volume end-to-end (vendor-reported), while a competitor-authored teardown — incentives noted — estimates 45–53% in production deployments. The words "up to" are doing structural work across this category. Directionally, practitioners agree that structured intents (order status, resets) automate far better than complaint-type, emotional, or policy-exception conversations; I publish no precise per-intent percentages because I could not trace the circulating ones to a primary source.

The credibility problem is documented

In March 2025, TechCrunch reported that 11x.ai — backed by a16z and Benchmark — had displayed logos of companies that weren't customers (ZoomInfo's lawyers threatened action) and booked full-year ARR on contracts with three-month break clauses; ex-employees described 70–80% churn. The company survived. The episode set a marker for what unaudited category metrics can hide.

The pattern persists at larger scale, more politely. Decagon has reported $30–35M in annualized revenue; Forbes, investigating independently, estimated ~$12M for 2025 — a 3x gap at a company valued at $4.5B. Decagon says it wins "almost all" competitive bake-offs; Fin's CEO claims the opposite ("every performance bake-off we have with them, we win"). Both statements are marketing; neither is audited; billions in valuation are priced on multiples of exactly these numbers.

The rule this earns: treat every vendor-reported performance rate as a ceiling, and every unaudited private-company revenue figure as a claim until triangulated. The eval you run on your own tickets is the one benchmark in this market whose definition you control.

Three numbers to retire — against their original sources

"95% of AI pilots fail" (MIT/NANDA). The GenAI Divide report (August 2025) actually says ~95% of integrated, task-specific enterprise GenAI pilots showed zero measurable P&L return within its observation window. Its method: 300+ public initiatives analyzed, interviews across 52 organizations, 153 surveyed leaders. Two findings usually mangled in citation: ~40% of organizations report deployment of general-purpose LLM tools (ChatGPT/Copilot class) while only ~5% of task-specific tools reached production — that's a generic-vs-embedded funnel, not a build-vs-buy score; and in the interview sample, external partnerships reached deployment roughly twice as often as internal builds (~67% vs ~33%, self-reported). The authors themselves flag selection bias, a six-month window, and inconsistent success metrics. Quote it precisely or not at all.

"Responding in 5 minutes makes you 100x more likely to win the lead." The origin is a 2007 phone-call study (Oldroyd, MIT/InsideSales) measuring contact and qualification odds — routinely misattributed to Harvard. The nearest modern datapoint, Optifai's 2026 commercial benchmark of 939 B2B companies, is observational, measures a different outcome (close rate), and finds 2.6x. Directionally supportive; nothing like 100x; not a formal replication. Meanwhile the share of inbound leads never contacted at all rose from 23% (a 2011 HBR-published audit) to 63.5% (a 2024 vendor audit of 1,000 demo requests — different sample and method, same direction). Two decades of speed tooling and the base rate got worse: the bottleneck looks operational, not technological.

"The AI agents market will grow from $7B to $199B." Research vendors publishing these curves disagree with each other by 2–3x on the same category, with methodologies undisclosed. No analyst firm publishes an "AI for customer support" market size worth citing. The frame investors actually use — pricing these companies against the $300–400B/yr global contact-center labor pool rather than a software TAM — is an interpretation, but at least an honest one: it says the bet is on labor budgets, not software budgets. The hockey-stick charts say less.


4 · The rollback file

2025–26 produced the first meaningful public record of what happens after deployment. What follows is a selected case file, not a dataset — cases surfaced through press coverage, with all the selection bias that implies. It still cuts both ways instructively.

Klarna is the whole story in one company — with each number labeled. February 2024, company announcement: the assistant handled 2.3M conversations in its first month, "equivalent to the work of 700 full-time agents," with a ~$40M projected profit improvement. May 2025, CEO Sebastian Siemiatkowski, publicly: cost was "a too predominant evaluation factor… what you end up having is lower quality," followed by re-hiring of human agents and the line "there will always be a human if you want." Q3 2025 earnings call: the agent "can now do the work of more than 853 full-time agents." Total company headcount fell from ~7,400 to ~3,000 over several years — driven mainly by slowed hiring and attrition, across the whole company, and not attributed by Klarna solely to the support AI. Nobody turned the AI off; nobody kept it exclusive; the automation number and the human-quality investment grew at the same time.

Cadenflow / Fig. 03

Klarna: automation and human quality

Klarna customer support AI timeline from February 2024 through the third quarter of 2025.
DateReported development
Launch claimThe assistant handled 2.3 million conversations in its first month, described as the work of 700 full-time agents.~$40 million projected profit improvement, according to Klarna's announcement.
Quality correctionThe CEO said cost had become too predominant and that the result was lower quality.Klarna began recruiting human agents and promised a human option for customers who wanted one.
Hybrid modelKlarna said the agent could now do the work of more than 853 full-time agents.Automation equivalence rose while human-agent recruiting continued as a quality tier.

Both expanded: automation equivalence and human-quality investment.

Automation figures come from Klarna announcements and its Q3 2025 earnings call; the quality correction comes from the CEO's May 2025 remarks. Total company headcount fell from roughly 7,400 to 3,000 over several years, mainly through slowed hiring and attrition, and was not attributed solely to support AI.

The adjacent cases rhyme — across different functions. Commonwealth Bank of Australia cut 45 call-center roles for a voicebot, watched call volumes rise, and reversed within three weeks, with an apology (customer service). Ford rehired ~350 veteran engineers after automated systems underperformed (engineering, not CX — evidence of an automation-overreach pattern, not of support outcomes). IBM automated ~94% of routine HR requests and announced plans to triple entry-level hiring (HR, same caveat). Salesforce cut support headcount from ~9,000 to ~5,000 as Agentforce took roughly half of conversation volume — redeploying hundreds of people to sales (per Fortune's reporting of company statements).

The forecasts bracket the field. Gartner predicts both: over 40% of agentic AI projects canceled by end-2027 (June 2025 — cost, unclear value, inadequate risk controls) and 80% of common customer service issues autonomously resolved by 2029 (March 2025). Forrester's 2026 prediction: one-third of brands will erode customer trust through self-service AI. All three are predictions, not measurements — and they can all come true together.

Customer sentiment, from one properly-sized survey: Gartner's December 2023 survey of 5,728 customers found 64% would prefer companies didn't use AI in customer service, and 53% would consider switching if they learned a company was about to. That predates the current agent generation — treat it as the trust baseline vendors are deploying into, not a verdict on today's tools. A visible "proudly human" marketing wave (Aerie, Polaroid, LEGO, Dove) is already monetizing that sentiment, and Klarna's CEO named the endgame: if AI does customer service, AI is the cheap customer service — human attention becomes the premium tier.

The synthesis, labeled as such: the reviewed cases and the major vendors' own product designs increasingly converge on a hybrid operating model — AI on routine, high-structure volume; instant human escape; humans repositioned as the premium layer. Vendors pitching "replace your support team" are selling against their own category's case record.


5 · What separates deployments that work

Model capability has stopped being the interesting variable. Frontier models score near the ceiling of τ²-bench — the policy-following, tool-using support benchmark authored by Sierra itself (vendor-authored; benchmark performance is not production reliability, but the direction is unambiguous, up from ~50–70% scores a year earlier). A competent demo agent is now a weekend project on open frameworks. Reliable production systems are not — and in every failure case I reviewed, the gap was operational, not model-quality.

The recurring failure factors

Across the documented cases (Air Canada, Cursor, CBA, the drive-thru experiments, and the deployment post-mortems in my corpus), four factors recur. This is a pattern in a selected case file — strong enough to design against, not a proven causal law:

  1. Knowledge-base rot. Support-content audits routinely surface outdated articles; an agent grounded on stale content distributes the errors at scale, confidently.
  2. Escalation design. Customers tolerate automation when the exit to a human is instant and obvious; automation loops with no exit are the most reliable rage-generator in the case file.
  3. Scope selection. Programs that deploy on complaint-heavy, emotional, or policy-exception intents fail where the same product would have thrived on order status and resets.
  4. Absent QA. Containment metrics hold while answer quality drifts. Tuning is a weekly discipline, not a launch task.

The liability record backs the design rules. In Moffatt v. Air Canada (BC Civil Resolution Tribunal, February 2024), the tribunal found negligent misrepresentation after Air Canada's chatbot gave a passenger wrong bereavement-fare information, rejecting the airline's suggestion that the chatbot was "a separate legal entity that is responsible for its own actions" — "a remarkable submission," in the tribunal's words. The award was modest (CA$812 total) and a small-claims tribunal decision binds nobody else — but the lesson generalizes: companies can incur liability when customers reasonably rely on inaccurate chatbot information. Add Cursor — an AI company whose own support bot invented a nonexistent login policy in April 2025, prompting public cancellations — and the Chevrolet dealership bot talked into a $1 Tahoe (December 2023), and you have the incident file that justifies grounding answers in live systems, validating outputs against source-of-truth policy, and gating anything that touches money behind a human.

The interface consensus: graduated autonomy

Gorgias, Fin, and Yuma — and Triple Whale outside support — have all shipped the same trust pattern, per their own product documentation: a review queue as the home screen; the AI draft pre-loaded in the composer so approving and editing are one gesture; a per-intent autonomy dial graduating from draft-onlyapprove-to-sendautonomous as measured accuracy earns it; a feedback affordance on every draft. Four vendors shipping one pattern isn't an industry standard — but it's a strong signal of where agentic trust UX has settled, and it's the pattern I'd build to.

The service layer ate the category

Follow the job titles and the deal structures. OpenAI ships Presence with its own forward-deployed engineers at consulting prices. Sierra's implementation services are reported to equal or exceed its year-one license in some contracts. Crescendo — past a company-reported $100M ARR — sells $2.99 verified resolutions with the humans included and acquired a BPO (PartnerHero) to deliver them. Wonderful, the fastest company in the tracked set to a $2B valuation, runs 28 offices for ~80 enterprise customers; its CEO says the quiet part: "The product itself isn't worth much if you don't connect it to dozens of systems in the organization." "Forward-deployed engineer" is one of the fastest-growing AI job titles by industry accounts. At Yuma — the YC-backed e-commerce leader — a July 2026 snapshot of LinkedIn-visible headcount (~30 profiles; method in the source register) shows roughly two-thirds in sales, account management, and implementation, with about 6–7 visible engineers. Indicative, not audited — and consistent with everything else in this section.

And across the fourteen companies in my teardown set — from $15.8B Sierra to bootstrapped Crisp — every one pairs its product with mandatory onboarding, implementation, or account management at its core price point, per their own public materials ("plug-and-play" here meaning: buy, connect, run at the advertised price without vendor-side humans). That is the market quietly admitting what the margin data suggests: AI application companies project ~52% gross margins for 2026 (actuals: 45% in 2025, 41% in 2024, per ICONIQ's survey of ~300 executives) against the classic-SaaS 80–90% benchmark. The data doesn't decompose where the margin goes. My read: a large share of it goes into the delivery and accountability layer — because that's what makes deployments work, and it's what buyers at 79x-revenue valuations are actually paying for.


6 · August 2: the EU transparency rules arrive

Operational analysis, not legal advice. Have counsel review any compliance decision.

The rule. From August 2, 2026, EU AI Act Article 50(1) requires providers of AI systems intended to interact directly with natural persons to design them so people are informed they're interacting with AI — unless that's "obvious to a natural person who is reasonably well-informed, observant and circumspect" given the context. The European Commission adopted final guidelines on these transparency obligations on July 20, 2026. For a customer-facing support bot serving EU users, the safe reading is simple: disclose at first contact, clearly, in every channel, and don't plan to litigate "obvious."

What the Digital Omnibus did and didn't change. Regulation (EU) 2026/1744 (Official Journal July 24, in force July 27) moved the high-risk obligations: Annex III systems to December 2, 2027, Annex I embedded systems to August 2, 2028. It did not move Article 50 — ordinary support chatbots were never high-risk, and their transparency duty proceeds on schedule. Separately, machine-readable marking of synthetic content applies immediately to new generative systems, while systems already on the market before August 2, 2026 have until December 2, 2026 to comply. Four different clocks; don't compress them into one deadline.

Penalties, precisely. Article 99(4)(g): non-compliance with Article 50 transparency obligations can draw administrative fines up to €15M or 3% of total worldwide annual turnover, whichever is higher — except for SMEs and startups, where it's whichever is lower (Art. 99(6)). Those are statutory maximums; actual enforcement depends on the violation and the authority. A missing bot label does not automatically cost €15M — and it's still not a risk worth taking for the cost of a disclosure line.

One correction to a common claim, including my own July 30 draft: Article 50(4)'s human-review exemption applies to AI-generated text published to inform the public on matters of public interest where a human exercised editorial control. It is not a disclosure exemption for customer-support replies. A human-approval queue remains worth running — for quality, for evidence, for governance, and for anything touching money — but it does not switch off the support bot's disclosure duty.

The US picture, scoped correctly. The recent state laws generating headlines — California's SB 243 (signed October 2025; private right of action, $1,000 minimum statutory damages) and New York's GBS §1702 (disclosure at the start and every three hours) — regulate companion chatbots, and both expressly exclude bots used only for customer service. California's older B.O.T. Act prohibits bots that mislead about their artificial identity with intent, in commercial-incentive or electoral contexts — disclosure is a safe harbor. Adjacent law, worth watching, not currently a general customer-service-bot mandate. The trend line — from transparency duties toward risk-management duties, with Gartner predicting a legislated "right to talk to a human" by 2028 — is the thing to plan for.

What falls out of the 2026 record: disclose the bot at first contact in EU-facing channels; keep an instant, obvious path to a human; ground answers in live source-of-truth systems, because the Air Canada tribunal shows reliance on wrong bot answers can cost you; log human review where you rely on it — as evidence and governance, not as a disclosure exemption; and note that EU-hosted deployment has been moving from trust signal toward procurement requirement in parts of the European market (sovereign-AI procurement is now a real category — e.g., SAP's EU AI Cloud), an observation about direction, not a statutory rule.


7 · Where this goes next

Four things worth watching through mid-2027 — each labeled for what it is.

Voice (observed trend + thesis). Voice overtook text as Sierra's primary channel in October 2025 (company-reported). Phone is the most expensive human channel (~$5–16/call, third-party estimates), and voice infrastructure costs have collapsed to roughly $0.05–0.10/minute (third-party estimates). The platforms are chasing enterprise contact centers. My thesis: a credible phone-answering agent for small business is the biggest under-served surface, and the tooling already exists.

Machine customers (early signal). Gartner warns service organizations to prepare for inbound AI — assistants contacting businesses on their owners' behalf. Zendesk already exposes MCP as a support channel; Shopify ships merchant-verified knowledge specifically so shopping agents don't hallucinate store policies; MCP passed 10,000 public servers and moved under the Linux Foundation in December 2025. These are early signals of a support-API-for-agents pattern — growing, not yet standard.

Verification as product (observed pattern + forecast). Zendesk bills only resolutions confirmed by an evaluation model; a fresh vendor category (simulation testing, LLM-as-judge QA, policy-compliance regression) is forming under the agents. My forecast: as autonomy rises, "prove the AI did the job" compounds in value — and verification needs an accountable party, which makes it the hardest layer to bundle away.

Prices (historical series + forecast). Inference cost for constant capability has fallen roughly 10x per year over the measured window (a16z/Epoch series) — a historical series, not a law of nature. My forecast: the $0.99-tier prices sag, platforms bundle more of the baseline for free, and the market's value keeps concentrating where this report kept landing: integration depth, honest measurement, compliance, and the humans who make deployments work.


8 · If you're evaluating AI support in 2026: ten questions

  1. What exactly counts as a "resolution" in the contract — and does customer silence count toward it?
  2. What's the price after the AI works — is there a cap, or a success tax?
  3. Can we run a bake-off on our own tickets before signing? (The one benchmark whose definition you control.)
  4. Which of our intents are high-structure (order status, resets) vs. complaint-shaped — and are we deploying on the right ones first?
  5. Who maintains the knowledge base, on what cadence, and what happens when policy changes?
  6. How does a customer reach a human, and how many steps does it take?
  7. Does the system draft-then-approve on anything touching money — refunds, discounts, account changes?
  8. Is our EU bot disclosure visible at first interaction, in every channel? (Article 50 applies from August 2, 2026.)
  9. What's logged — and would the log show what the bot told a customer if a dispute ever reached a tribunal?
  10. If the vendor disappeared tomorrow, what do we own — the prompts, the workflows, the conversation data?

9 · Method & limitations

Period and cutoff. Research corpus built June–July 2026 on top of implementation work running April–July; final verification pass completed July 31, 2026, which is this report's evidence cutoff.

The teardown set (14 companies): Sierra, Decagon, Parloa, PolyAI, Wonderful, Siena, Yuma, Redo, Chatwoot, Crisp, Hugo (by Crisp), Influx, Smith.ai, and Hyros (an adjacent measurement case). Selection: vendors I studied while building in the category — enterprise leaders, e-commerce specialists, open-source, BPO, and bootstrapped counter-examples. A teardown = funding, revenue (labeled by source quality), pricing, GTM, product mechanics, and public case studies, from company materials, press, and third-party research.

Datasets. Funding claims come from a frozen tracker (23 disclosed equity rounds, Jan 2024–Jul 2026, with sources and exclusion rules — debt, AI-SDR, and eval-tooling rounds excluded). It under-captures small seed rounds; dollar-concentration conclusions are robust to that, deal-count trends are not, and none are published. Pricing claims come from a 12-vendor snapshot with units, fees, definitions, and check dates (June 22–July 31). Calculations (inference scenarios, the Freshworks multiple, funding aggregates) are published with formulas.

Source standard. 75 catalogued sources, of which 41 are primary documents (legislation, tribunal decisions, earnings materials, company announcements, original research, official pricing pages — vendor announcements count as primary evidence that the vendor made the claim, not that the claim is true). Vendor-reported metrics are labeled and should be read as ceilings. Estimates are labeled with their estimator.

What I could not verify, and therefore removed: a widely-circulated "41.2% median deflection" benchmark attributed to Zendesk (no primary source found); per-intent deflection percentages (stat-aggregator provenance); a "70–85% mature deployment" benchmark (unsourced); a consumer-sentiment tracking series (methodology unconfirmable). Where a removed number left a hole, the replacement is either a sourced alternative or an honest "directional" statement.

Known limitations. Two pilot deployments is operating experience, not a statistical base. The rollback file is press-selected cases, not a systematic sample. Several revenue figures are company-reported and unaudited — labeled wherever used. The regulation section was checked against official texts and the Commission's July 20 guidelines but has not been reviewed by counsel; treat it as operational analysis, not legal advice. Pricing changes fast in this market — re-verify before relying on any number here after Q3 2026.

Changes from the July 30 draft: pricing reframed from a single "price of a resolution" ladder to a pricing-models comparison; funding claims rebuilt on the frozen tracker; the "30-point gap" chart replaced with a metric taxonomy; MIT/NANDA, Gartner 14%, Klarna headcount, Article 50(4), US state-law scope, and the Zendesk public-multiple error corrected; source counts regenerated from the register.


About the author. Maks Zhashkevych is the founder of Cadenflow, a product-development and AI-automation studio. He has ten years of software engineering behind him and spent two months of 2026 building and operating an agentic customer-support system for two e-commerce pilot deployments — which is where most of this report's skepticism was earned. Cadenflow builds production AI systems — integration, evaluation, guardrails, deployment — from idea to production in weeks: cadenflow.com.

This report is planned as a recurring series; Edition 02 is planned for the end of August 2026. Get it by email: maks@cadenflow.com.

From report to production

Need the system behind the claim?

Cadenflow designs, builds, and operates production AI workflows with the integration, evaluation, and human controls included.

See how we build