AISSIST is awarded Best Agentic AI for Business from CIOReview.
AissistAissist
Back to Insights
Customer Support·Buyer Guide·Agentic AI·Vendor Evaluation

AI Washing in Customer Service: 5 Tests for Buyers

Gartner counts roughly 130 genuinely agentic vendors among the thousands claiming it. Five tests — resolution denominator, named action, model-failure behaviour, price of a failure, per-decision audit trail — each with a verbatim question and a passing and failing answer, applied to seven AI support agents from their own live pages. Aissist.io fails the first one.

Alex Gomez · Sep 21, 2026 · 15 min read · Updated Sep 21, 2026

AI Washing in Customer Service: How to Tell Real Agentic AI From a Rebranded Chatbot

Gartner counts roughly 130 vendors with genuine agentic capability among the thousands claiming it. Five questions, asked in twenty minutes on a vendor call, tell you which kind you are talking to.

TL;DR: AI washing in customer service is caught by disclosure, not by demo — Gartner puts genuinely agentic vendors at about 130 of thousands, and the five tests below expose the rest before you sign anything.

Methodology & sources

  • Keyword data: Semrush, US database, pulled 21 September 2026 — "ai washing" 1,300 searches/month at keyword difficulty 34; "agent washing" 70/month at KD 36 with a $6.91 CPC.
  • Vendor figures read from each vendor's own live pricing page, docs or help centre on 21 September 2026 — never from a competitor's comparison article. Re-verify prices by December 2026.
  • Performance figures read from the publishing vendor's own page, labelled vendor-claimed throughout. Re-verify by March 2027.
  • Enforcement counts from SEC and FTC primary releases and law-firm analysis. Re-verify by March 2027.
  • Disclosure: this is Aissist.io's blog. Aissist.io sells an agentic AI platform for customer service and sales, and is scored on the same five tests as every other vendor here — including the one it fails.

Five-test AI washing scorecard comparing Intercom Fin, Zendesk, Salesforce Agentforce, Ada, Sierra, Decagon and Aissist.io on published resolution rate, named actions, model dependency, failure pricing and audit trail

This article scores what vendors publish, not what they can demo. For the metric underneath test T1, see deflection rate versus resolution rate. For the architecture distinction the whole argument rests on, see traditional chatbots versus agentic AI.

What is AI washing in customer service?

AI washing in customer service is the practice of marketing a support system as autonomous AI when the customer-facing work is still done by scripted flows, retrieval-only answers, a thin wrapper over someone else's model, or hidden human labour — and, most often, of reporting a success metric that counts conversations the AI never actually resolved. Cornell Law School's Legal Information Institute defines the parent term plainly: "AI washing occurs when companies claim to be using artificial intelligence (AI) technology to enhance their services but, in fact, are not."

In customer service the claim is usually softer than an outright fabrication, which is exactly what makes it hard to catch. The AI is real. The number describing it is the thing that has been laundered.

Gartner gave the software version of this its own name. "Agent washing," the firm wrote in its June 2025 forecast, is the "rebranding of existing products, such as AI assistants, robotic process automation (RPA) and chatbots, without substantial agentic capabilities." Gartner's own count of vendors that genuinely deliver — an analyst estimate, not a census — is about 130, out of thousands making the claim.

"Most agentic AI propositions lack significant value or return on investment (ROI), as current models don't have the maturity and agency to autonomously achieve complex business goals." — Anushree Verma, Senior Director Analyst, Gartner

Regulators arrived before the buyers did. The SEC charged two investment advisers over false AI claims in March 2024, settling for $225,000 from Delphia and $175,000 from Global Predictions — $400,000 in total — with then-Chair Gary Gensler stating that "investment advisers should not mislead the public by saying they are using an AI model when they are not." The FTC brought "more than a dozen cases associated with AI washing" in the year to August 2026, according to Holland & Knight's review of Operation AI Comply, published 18 August 2026 — and, importantly for anyone buying software rather than consumer goods, the firm notes that "the same substantiation requirements and deception standards apply, regardless of whether the audience is a consumer or sophisticated business purchaser."

That last line is the one to keep. A support platform's resolution-rate claim is a substantiable marketing claim, not a vibe.

Why has AI washing got harder to spot in 2026?

Because buying accelerated faster than evaluation did. Verint's State of Contact Center AI 2026, reported by CX Today on 3 September 2026, found that 87% of contact-centre leaders plan to increase spending on automated agent assistance — while the average contact centre already runs technology from three or more providers and 20% run five or more. A market adding vendors that quickly has stopped comparing them.

Then there is the naming. Salesforce spent Dreamforce week shipping labels: AIforce, Claudeforce, Slackforce, Agentforce Coworker and Koa — a CRM reasoning model the company says was trained on 27 years of deployment scenarios — all announced inside four days, per CXM's 18 September 2026 roundup. Some of those are real products with real mechanisms behind them. The problem is that a four-day naming sprint teaches the whole market that a new capability and a new suffix look identical from the outside.

Buyers can feel it. The tension in every agentic AI evaluation right now is between a board that wants the spend announced and an operator who has to run the thing on Monday.

"But we hear that and people say, 'wait a minute, I want to make sure before I put this in production that it's actually going to do a positive interaction with my customer if it's customer facing, and that it's not burning through tokens.'" — Rebecca Wettemann, CEO and Principal Analyst, Valoir

Gartner's matching prediction is that over 40% of agentic AI projects will be cancelled by the end of 2027, on "escalating costs, unclear business value or inadequate risk controls." Note what is not on that list: the model was not good enough. Projects die because nobody agreed in advance what success would be measured against.

Which is the good news, in a way. If AI washing were a technology problem, a buyer could not detect it from the outside. It is a disclosure problem, and disclosure can be tested on a call.

What five questions expose AI washing on a vendor call?

Five questions expose it: what the vendor publishes as a resolution rate and against what denominator, what action it takes inside your system, what happens when its model degrades, what it charges for a conversation it did not resolve, and whether it can show the audit trail for one decision. Each has a passing answer and a failing one, and the failing answers are not evasive — they are confident, fluent and specific about the wrong thing.

Ask them verbatim. Vendors answer rehearsed questions with rehearsed answers.

T1 — The resolution rate and its denominator

Ask, verbatim: "What is your average resolution rate, across how many customers, and what is excluded from the denominator?"

Passing answer: a single number, the population it is measured across, and an explicit list of what does not count. Intercom's Fin publishes 76% average resolution across 8,000+ customers (vendor-claimed, fin.ai, page dated 17 June 2026, verified 21 September 2026), and states its exclusions: "Fin only counts genuine, positive resolutions. Conversations where customers express dissatisfaction, where the agent fails to address the core issue, or where the customer returns with the same problem are not counted."

Failing answer: a deflection rate, a containment rate, an engagement rate, or a resolution rate with no population attached.

The gap between those two metrics is not a rounding error. Ada reports containment averaging around 72% and automated resolution averaging around 52% measured on the same conversations across 550+ deployments (vendor-claimed, Ada, verified 21 September 2026) — a twenty-point spread that Ada, to its credit, publishes itself. Zendesk is more literal still: its help centre states an automated resolution is counted "after 72 hours of inactivity if AI evaluation has confirmed that the AI agent's response was relevant," and a customer who gives no feedback at all does not prevent the count (Zendesk help centre, verified 21 September 2026).

Silence is not a resolution. It is a customer who gave up somewhere you cannot see. Our own breakdown of deflection rate versus resolution rate walks the arithmetic.

T2 — The action inside your system

Ask, verbatim: "Name the action your agent takes inside my system on this ticket type — the field it writes, the record it changes, the API it calls."

Passing answer: the object, the system and the trigger, named without hedging. "It writes the RMA number to the order record in Shopify and sets the ticket's refund_status field."

Failing answer: "It surfaces the right answer from your knowledge base." That is retrieval. Retrieval is useful, and it is not an action — it is the whole of the chatbot versus AI agent distinction in one sentence, and the reason we argued back in 2025 that the chatbot was already dead.

Salesforce, of all vendors, makes this testable because it prices the unit: one Agentforce action consumes 20 Flex Credits, which is $0.10, from packs at $500 per 100,000 credits, and the documentation says "Flex Credits ensure you only pay for the actions Agentforce performs, such as updating customer records, automating workflows, or resolving cases" (Salesforce Help, verified 21 September 2026).

A vendor that bills per action has been forced to define one. That is not a coincidence; it is a disclosure.

T3 — Behaviour when the model degrades

Ask, verbatim: "Which models does this run on, and what happens to my resolution rate the week a provider ships a regression?"

Passing answer: names the models, the routing rule, the fallback, and the regression suite run before a new model is promoted into production traffic.

Failing answer: "We're model-agnostic" with no routing rule behind it, or "we always use the best model" with no model named. Model-agnostic and model-indifferent are not the same posture.

Cisco's own deployment shows what the passing answer sounds like from the buyer's side of the table, and at a scale worth quoting: 145,000 support cases resolved entirely by AI with zero human intervention in FY2026 (company-reported, stated by Cisco's CEO on its earnings call).

"Circuit runs on our secure AI factory infrastructure, which improves GPU utilization and automatically routes each task to the appropriate large language model, allowing us to manage token consumption." — Chuck Robbins, Chair and CEO, Cisco (reported 13 August 2026)

If the largest buyers are specifying routing at the task level, a vendor that cannot describe its own is behind its customers.

T4 — The price of a failure

Ask, verbatim: "What do you charge me for a conversation your AI did not resolve — and what platform tier do I have to be on first?"

Passing answer: nothing, or a stated lower rate, plus the host-plan requirement said out loud. Intercom's Fin is $0.99 per Fin outcome, and Intercom's plans — Essential at $29, Advanced at $85 and Expert at $132 per seat/month — include Fin at no additional seat cost, with a standalone option at the same $0.99 and no seats required (Intercom pricing, verified 21 September 2026). Zendesk includes AI agents in Suite plans — Team at $55 and Professional at $115 per agent/month billed yearly — and bills automated resolutions beyond the plan allowance (Zendesk pricing, verified 21 September 2026).

Failing answer: per conversation, per message, per interaction with no outcome cap, or "quote only." Salesforce's alternative to per-action pricing is $2 per conversation, which is charged whether or not the conversation goes anywhere. Sierra says only that customers "pay for the value Sierra delivers with outcome-based pricing" and publishes no rate (sierra.ai, verified 21 September 2026). Decagon explains resolution-based pricing at length — and concedes the hard part, that "defining what a resolution is can be tricky, as not all cases end wrapped up in a bow" — without publishing its own number (Decagon glossary, verified 21 September 2026).

"Quote only" is an honest answer. It is just not a comparable one, and it means the resolution definition arrives in a contract rather than on a web page.

T5 — The audit trail for one decision

Ask, verbatim: "Open one conversation from last week and show me why the AI said what it said — the source it retrieved, the policy check it ran, the rule that escalated or didn't."

Passing answer: someone shares a screen and walks the chain for a specific, real conversation, including a case where the AI was wrong.

Failing answer: a dashboard. Aggregate accuracy is a report card; an audit trail is evidence. Only one of them tells you why a customer got the answer they got.

This is the test the whole category fails most often. Of the seven vendors compared below, not one publishes a per-decision audit-trail specification on its public site as of 21 September 2026 — Aissist.io included. Our evaluable AI page puts the share of AI support vendors offering no evaluation or only superficial metrics at roughly 90%, which is a strong claim about an industry that Aissist.io is a member of.

How do the named vendors score on the five tests?

On public disclosure as of 21 September 2026, Intercom's Fin and Ada score best on T1, Salesforce and Intercom score best on T4, and every vendor in the set — Aissist.io included — fails T5. The columns below score what each vendor publishes, not what its product can do in a demo. A vendor may well pass a test privately and score Partial here; that is the point of asking on a call.

Vendor (verified 21 Sep 2026)T1 Rate + denominatorT2 Names the actionT3 Model-failure behaviourT4 Charge if unresolvedT5 Per-decision audit trail
Intercom FinPass — 76% across 8,000+ customers, exclusions statedPartialNot publishedPass — $0.99 per outcome onlyNot published
Zendesk AI agentsFail — resolution counted after 72h of inactivityPartialNot publishedPartial — AR-based, above plan allowanceNot published
Salesforce AgentforceNot publishedPass — action is the priced unitPartial — Koa model namedPartial — $0.10/action or $2/conversationNot published
AdaPass — 52% AR vs 72% containment, 550+ deploymentsPartialNot publishedNot publishedNot published
SierraNot publishedPartialNot publishedPartial — "outcome-based," no rateNot published
DecagonNot publishedPartialNot publishedPartial — resolution-based, no rateNot published
Aissist.ioPartial — 83% with no denominatorPartial — 5 named Zendesk actionsPartial — routing + governor publishedPass — unresolved costs nothingNot published

Pricing, normalised, since T4 is the test buyers get wrong most:

VendorBilling unit (Sep 2026)Published rate (Sep 2026)Charged on failure?Required host tier (Sep 2026)
Intercom FinPer Fin outcome$0.99 USDNoAny Intercom plan from $29/seat/mo, or standalone
Zendesk AI agentsPer automated resolutionNot published on pricing pageOnly if AR is counted (incl. 72h silence)Suite Team $55 or Professional $115/agent/mo, billed yearly
Salesforce AgentforcePer action, or per conversation$0.10/action (20 Flex Credits); $2/conversationYes, on the conversation modelNot stated in pricing doc
AdaPer automated resolutionNot publishedNot publishedNot published
SierraOutcome-basedNot publishedNot publishedNot published
DecagonResolution-basedNot publishedNot publishedNot published
Aissist.ioPer interaction, capped per resolution$0.09/interaction capped at $0.60/resolution — lower of the twoNoNone; free tier to 1,000 tickets/mo

Two fairness notes. "Not published" means we could not find it on the vendor's own site on 21 September 2026 — it is a statement about disclosure, not about capability. And the head-to-head resolution rates Fin publishes for competitors (Decagon at 49%, Forethought at 50%) are Fin-claimed comparative testing, which is exactly the kind of number T1 exists to interrogate, whoever publishes it.

Three decision rules fall out of the table. If you need a resolution number you can benchmark against your own, shortlist the vendors that publish one with a population attached — today that is Intercom Fin and Ada — because a rate without a denominator cannot be compared to anything. If your automation depends on writing to a system outside the helpdesk, weight T2 hardest and make the vendor name the API call, because an integrations page proves a connection exists, not that the agent can write through it. And if your budget cannot absorb paying for failure, rule out per-conversation pricing before the demo rather than after, because that is the one contract term no amount of tuning improves.

Ada's containment rate of 72 percent against its automated resolution rate of 52 percent on the same conversations, with Zendesk's 72-hour silence rule shown as the mechanism that inflates AI washing metrics

Where does Aissist.io fail its own test?

Aissist.io fails T1, the test this article leads with. Its homepage publishes an 83% average resolution rate and a 4.8/5 CSAT on AI-resolved conversations (vendor-claimed, aissist.io, verified 21 September 2026) with no customer count and no conversation count attached. Intercom publishes 8,000+ customers behind its 76%. On the denominator question, Fin answers better than we do.

It gets more awkward under the hood, and the awkwardness is instructive. Aissist.io's own AI Customer Service Benchmark 2026, updated June 2026 from 40+ sources across six industries, puts the cross-program verified median resolution rate at roughly 41%, with a top quartile near 59% (independently reported, compiled by Aissist.io from third-party programs). Those are different populations — an industry-wide median drawn from third-party programs versus Aissist.io's own deployments — and both can be true at once. But a reader who sees 83% in the hero and 41% in the benchmark deserves the sentence explaining which is which, and until this paragraph, we had not written it.

Where Aissist.io does pass is T4, and unambiguously: billing is $0.09 per interaction capped at $0.60 per resolution, whichever is lower, and a conversation handed to a human costs nothing (Aissist.io pricing, verified 21 September 2026). T3 is a Partial — the reliable AI page documents a multi-layer approach with a separate governor monitoring outputs against policy, but no published regression gate for promoting a new model. T2 is a Partial: the Zendesk integration page names five concrete actions inside Zendesk, and then describes external systems generically as "CRM, ecommerce, shipping, ERP, or internal tools," which is the hedge T2 is designed to catch. T5 is a fail for everyone, us included.

A page that grades its own author top marks on every line is a sales sheet. Score us on the five; the 83% needs a denominator, and we will publish one.

Verdict: AI washing is a disclosure failure, not a technology failure

AI washing is caught by disclosure, not by demonstration: what separates the roughly 130 genuinely agentic vendors Gartner counted from the thousands claiming the label is not capability alone but whether the vendor has published a number specific enough to be wrong. The instinct when a category fills with identical claims is to demand a better demo. That is the wrong instinct, because the demo is the one artefact every vendor has already perfected.

So make the five tests the first twenty minutes of the call, before the screen share. A vendor that names its denominator, names the field it writes, names its models, prices its own failure and opens a single decision to inspection has handed you five falsifiable claims. A vendor that answers all five fluently without naming anything has handed you a suffix. Both will call it agentic AI. Only one of them can be checked — and checking is cheaper before the contract than after the quarter.

Run the five tests on us. Aissist.io publishes its per-resolution price, its resolution definition and its error rate — and fails T1 on the denominator, which we are fixing. Book a consultation →

Frequently asked questions

What is AI washing in customer service?

AI washing in customer service is marketing a support system as autonomous AI when the work is still done by scripted flows, retrieval-only answers, a thin wrapper over another company's model, or hidden human labour. Most often it takes the form of a success metric that counts conversations the AI did not resolve.

Is AI washing the same thing as agent washing?

No — agent washing is the narrower software version. Gartner uses "agent washing" for the rebranding of existing assistants, RPA and chatbots without substantial agentic capability, and estimates only about 130 of the thousands of vendors claiming agentic AI genuinely deliver it.

How can I tell a real AI agent from a rebranded chatbot?

Ask the vendor to name the action its agent takes inside your own system — the field it writes, the record it changes, the API it calls. A chatbot retrieves and replies; an agent changes state in a system of record and can tell you which one.

Why is deflection rate a red flag in an AI vendor's pitch?

Deflection counts conversations that never reached a human, including customers who gave up. Ada publishes containment at around 72% and true automated resolution at around 52% on the same conversations across 550+ deployments — a twenty-point gap that deflection hides entirely.

Does a 72-hour no-reply window really count as a resolution?

At Zendesk, yes. Its help centre states an automated resolution is counted after 72 hours of inactivity once AI evaluation confirms the response was relevant, and a customer giving no feedback at all does not stop the count. Ask every vendor what their equivalent window is.

Is AI washing illegal?

It can be. The SEC settled charges against two investment advisers for $400,000 in March 2024 over false AI claims, and the FTC brought more than a dozen AI-washing cases in the year to August 2026. Holland & Knight notes the same deception standards apply to business buyers as to consumers.

What should I ask about pricing to avoid paying for AI washing?

Ask what you are charged for a conversation the AI did not resolve, and what platform tier you must buy first. A $0.99-per-resolution rate on top of a $115-per-seat plan is not a $0.99 product, and a per-conversation rate charges you identically for success and failure.

Read Next

AG

Alex Gomez

Sr. Analyst

Alex is senior analyst at Aissist.io, covering AI vendor pricing, benchmarks and market structure. He has 5 years of experience in product management and marketing within the AI industry.