Back to blog
AI Research

How much support can AI actually resolve? Setting honest expectations

What share of support can AI honestly resolve? Realistic ranges by question type, the four factors that set your rate, and the levers that raise it without gaming the metric.

ReplyPool TeamJanuary 14, 20265 min read

Key takeaways

  • Count "resolved" strictly: a real answer, zero human touches, and no repeat contact within seven days on any channel — the only version of the metric you can budget on.
  • Routine, documented questions — order status, returns mechanics, product and billing basics — are 55–70% of e-commerce inbound and the honest automation ceiling.
  • Your rate is set by ticket mix, knowledge base coverage, order-data access and counting rules — far more than by which AI model sits underneath.
  • Expect 30–50% resolution in month one and 55–70% at maturity under strict counting; a 90% pitch is measuring something else.
  • Raise the rate by closing knowledge gaps weekly and connecting order data — never by loosening the definition of resolved.

Ask five vendors what share of your support an AI agent will resolve and you'll hear five confident numbers, most of them starting with a 7 or an 8. Ask five support leads who actually run one and you'll hear something more useful: "it depends — and here's on what." This article is the second answer, written down: which resolution rates are realistic, which question types automate well, which never should, and which levers actually move your number.

First, define "resolved" — strictly

A resolution rate is only as honest as its definition. Before you compare any numbers — a vendor's, a competitor's, your own last quarter — pin down what counts as resolved:

  • The customer got a real answer, not a link dump or a "please contact support" loop.
  • No human touched the conversation, before or after the AI reply.
  • The customer didn't come back about the same issue within a sensible window — seven days is a common choice — on any channel.

Under that strict definition, most spectacular numbers deflate. A bot that "resolves 85% of chats" often turns out to resolve 85% of the chats it was allowed to finish, after routing rules quietly filtered the hard ones to humans, with abandoned conversations counted as wins. Strict counting is worth the discomfort: it's the only version of the metric you can build staffing and budget decisions on.

What automates well: the routine majority

E-commerce support has a favorable shape for automation, because a large share of inbound volume is a small set of questions repeated thousands of times:

  • Order status and delivery ("where is my order?"). The single biggest category in most stores — commonly 25–40% of all contacts. Fully automatable if the AI can read live order and tracking data.
  • Returns and exchange mechanics. "How do I return this?", "What's the window?", "Who pays return shipping?" — policy questions with documented answers.
  • Product questions your pages already answer. Sizing, materials, compatibility, care instructions — anything a good product page or FAQ covers.
  • Account and billing mechanics. Password resets, address changes, invoice copies, updating a payment method.
  • Pre-sales basics. Shipping costs and timelines, payment options, stock and restock questions.

Add those up and you land at the familiar 55–70% of inbound volume that's routine and documented. That's the honest ceiling for most stores — not because AI can't read harder questions, but because the harder ones shouldn't be closed without a human.

What AI should not resolve — by design

Some categories can be answered by an AI but shouldn't be finished by one:

  • Judgment calls. Refund exceptions, goodwill gestures, "the courier says delivered but nothing arrived." The right answer depends on a decision only your team can own.
  • Angry escalations. Automation applied to an already frustrated customer multiplies the frustration. The AI's job here is to detect the tone and hand off fast, with context.
  • Complex or novel troubleshooting. If the answer doesn't exist in your knowledge base, a well-behaved AI says so and routes the conversation. A confident improvisation is worse than a queue.
  • Signal. Bug reports, defect patterns, feature requests, churn warnings. AI can tag and acknowledge them, but a person should read every one — this is product feedback arriving free of charge.

Teams that push these categories into automation buy a higher headline rate and pay for it in reopens, chargebacks and quiet churn. The goal is not the maximum resolution rate; it's the maximum rate at or above your quality bar.

The four factors that set your actual number

Two stores can run the same AI and land thirty points apart. The difference is almost always one of these:

  1. Ticket mix. A store where delivery questions dominate automates far more than a store selling complex, configurable products. Pull 200 recent conversations and sort them by type — the routine share you find is a better predictor than any industry benchmark.
  2. Knowledge base coverage. The AI resolves what your documentation grounds. Thin docs mean the honest rate starts low; every documented gap you close moves it up. This is the highest-leverage work of the first three months.
  3. Data access. Without a live connection to orders and tracking, "where is my order?" — your biggest category — is unautomatable. With it, that category alone can carry a third of the total.
  4. Counting rules. Reopen windows, abandonment handling, which channels count. Two teams with identical operations can report 55% and 80% from the same month of conversations.

Realistic ranges, by stage

With strict counting, here's what healthy looks like for an e-commerce operation:

  • First month: 30–50%. The knowledge base has gaps, integrations are new, and you should be running the AI on narrow categories anyway.
  • Months two to four: 45–60%, driven almost entirely by closing the documented gaps the AI reports.
  • Mature, six months and beyond: 55–70% for typical stores. Above 70% happens, but usually in businesses whose mix is unusually routine — high order volume, simple products, generous and well-documented policies.

If a pitch quotes 90%+ without asking about your mix, your documentation and your definition of "resolved," it's quoting a different metric than the one you'll live with.

Raising the rate without gaming it

The honest levers, in order of return:

  1. Close knowledge gaps weekly. Review the questions the AI couldn't answer, write the missing article, repeat. Nothing else compounds like this.
  2. Connect order data. It converts your largest category from "handoff" to "resolved in seconds."
  3. Fix root causes. If 15% of contacts are "confusing size chart," the best automation is a better size chart. A falling contact rate is a win even when it makes the resolution rate look flat.
  4. Expand category by category. Turn the AI loose on one new question type at a time, audit a sample of transcripts, then move on.

And one lever to refuse: loosening the definition. The teams that regret automation are rarely the ones with an honest 55%; they're the ones with a reported 85% that customers experience as 50.

Where ReplyPool fits

ReplyPool's economics are built for honest numbers. Every plan is a fixed monthly pool of AI answers — 1,000, 2,500 or 7,500 — so there is no meter that earns more when a "resolution" gets counted generously. The AI agent answers from your knowledge base, hands off what it can't ground, and the dashboard shows the questions it couldn't answer so you can close the gaps that actually raise your rate. When the pool runs out, conversations route to your team inbox and the bill stays where the pricing page said it would.

Want to know your real number? Sort 200 of your recent conversations into routine and judgment, and you'll have a better forecast than any benchmark — then run one calibration month to confirm it.

Share this article

Frequently asked questions

With strict counting — a real answer, no human involvement, no repeat contact within seven days — expect 30–50% in the first month and 55–70% once the knowledge base matures, for a typical e-commerce mix. Rates above 70% are possible but usually reflect an unusually routine ticket mix: high order volume, simple products and well-documented policies.

The routine, documented majority: order status and delivery questions (25–40% of contacts in most stores, automatable when the AI reads live order data), returns and exchange mechanics, product questions your pages already answer, account and billing mechanics, and pre-sales basics like shipping costs and stock. Together these are typically 55–70% of inbound volume.

Judgment calls such as refund exceptions and goodwill gestures, conversations with angry customers, troubleshooting that is not covered by your knowledge base, and signal like bug reports and feature requests. AI can acknowledge and tag these, but a person should own the outcome — automating them buys a higher headline rate and pays for it in reopens and churn.

Loose counting. Common inflators: measuring only the conversations the bot was allowed to finish after hard ones were routed away, counting abandoned chats as resolved, using no reopen window, and ignoring customers who come back through another channel. Two teams with identical operations can report 55% and 80% from the same month.

Four levers, in order of return: close knowledge base gaps weekly using the list of questions the AI could not answer; connect live order and tracking data; fix root causes so the demand itself shrinks; and expand automation one category at a time with transcript audits. The one lever to refuse is loosening the definition of resolved.

No. The goal is the maximum rate at or above your quality bar, not the maximum rate. Pushing judgment calls, angry escalations and undocumented questions into automation raises the number while degrading the experience — the damage shows up later as repeat contacts, chargebacks and quiet churn.

Ready to put AI support to work?

14 days free. Full platform. We move your data for you.