Next best action: the whole category names an action it never takes

14 min read

A customer writes in: "I was charged twice and still haven't been refunded." Your support AI agent reads the ticket, understands it perfectly, and recommends a refund in seconds. Fast, confident, and – quite possibly – wrong.

Because the refund may already be pending. A fraud review may be open. The customer may be on a nonrefundable plan. A known billing bug may explain the duplicate charge. The ticket alone can't tell you any of that. So the recommended next best action looks right on screen while missing the facts that decide whether it's actually safe to take.

That gap – between a recommendation that sounds correct and an action the system can safely complete – is where you need more than a smarter suggestion. It takes connected context, real permissions, and the ability to finish the work and record it.

What is next best action?

  • Next best action is an AI-guided recommendation for the most effective step to take with a customer at a given moment.
  • In customer service, that step may include sharing an answer, changing a record, issuing a refund, escalating a case, or resolving a request.
  • But a recommendation isn't a resolution. The work isn't finished until that step is executed and recorded.
  • The traditional model predicts an action and surfaces it to a person or system. The recommendation can be useful. But a suggestion sitting on a screen hasn't changed anything for the customer.

Learn why recommendation quality depends on connected context, how governed execution changes the evaluation, and what service leaders should look for in an agent-assist platform.

TLDR: What does the next best action really promise?

  • The next best action promises to help a service team choose and complete the right step. It should close work, not create another task for an agent. The quality of that decision depends on live customer, policy, ticket, product, and system context, plus the permission to act on it.
  • Computer, by DevRev, is an AI resolution platform built on Native Shared Memory – a live, permission-aware view of your customer, product, and system data – so it can reason across connected business context and take controlled action.
  • The practical test is simple: did the system move the work forward, or did it add work?

Why is the recommendation the part every tool shows?

The recommendation is the part of the next best action that demos well. An agent opens a ticket, and the software suggests a knowledge article, response, macro, escalation, or account change. The result looks fast and intelligent. It also fits neatly into a 20-minute demo. That's the promise behind most next best action software.

The approach has real value. Agents shouldn't need to search five systems for every answer. An AI agent can reduce decision time, surface relevant guidance, and help newer agents follow established processes.

But a suggestion still depends on human follow-through. The agent must decide whether it applies, check the relevant systems, perform the action, and update the record. That's where many deployments lose the efficiency they just gained.

Service leaders should separate two outcomes:

RecommendationAction
Tells an agent what might helpChanges the customer's situation
Uses available conversation or ticket contextChecks connected customer and system state
Requires the agent to verify and executeFollows permissions and workflow controls
May remain a draft or suggestionIs logged in the system of record
Ends with guidanceEnds with measurable progress

Agent assist can help a human make a better decision. It becomes far more valuable when the same workflow can carry an approved decision into the systems where work actually happens.

The goal isn't to dismiss recommendations. Recommendations are often the right starting point. The problem is calling the starting point the finish line.

Key takeaway: A good recommendation can improve judgment. A real action also reduces manual work and closes the operational loop.

Why is there a suggestion gap in customer service?

The suggestion gap is the distance between what the next best action promises and what it delivers. It appears when a copilot has enough information to sound confident, but not enough context to act safely. It reads the ticket correctly and still misses the conditions that decide whether the recommended step is allowed.

Consider this ticket:

“I was charged twice and still haven't been refunded.”

A copilot might recommend a refund macro. That suggestion seems reasonable. Yet the correct next step could be completely different:

  • A refund may already be processing.
  • A fraud review may be blocking the transaction.
  • The customer may be on a nonrefundable plan.
  • A known billing bug may require a product escalation.
  • The duplicate charge may belong to a different account.
  • The payment provider may have declined the original refund.

Ticket language alone can't resolve those questions. The suggestion looked reasonable. It just wasn't grounded in what was true.

This is a common pattern in support AI. The system identifies the apparent intent correctly but misses the conditions around that intent. It understands what the customer said. It doesn't understand what the organization already knows.

Practitioners raise the same point. In one community discussion of next-best-action guidance, support teams argue that AI suggestions only help once the underlying workflows and data are in order.

The problem gets more serious when the suggested action changes data or money. A generic answer can be corrected in the next message. An incorrect refund, cancellation, entitlement change, or account update may need a recovery process.

What happens when the refund shouldn't fire?

An agent may trust the suggested macro, issue a second refund, and create a reconciliation problem. The customer gets a confusing experience, finance has to investigate, and the agent has to explain an error they didn't cause.

This isn't always a model-intelligence problem. It's usually a context problem. The system can recognize the language of a billing complaint while missing the relationships and state that determine the safe response.

The same issue shows up outside billing:

  • An account change may violate a contract exception.
  • A replacement may be unnecessary because a shipment is already scheduled.
  • An escalation may duplicate an engineering issue that already has an owner.
  • A cancellation may trigger a retention workflow that another team is handling.
  • A troubleshooting step may be unsafe during an active incident.

The recommendation may be logical in a narrow view. It becomes wrong once the full operating context is considered. That's why the next best action depends on the whole situation, not just the latest message.

Key takeaway: A confident suggestion can still be unsafe when it lacks current context. Support AI must understand what has already happened before it recommends what should happen next.

How is the right step different from the next step?

The right step is the next step that fits the customer's actual situation, business rules, permissions, and desired outcome. That takes more than a recommendation engine. It takes a connected operating view of the work. The system should be able to understand:

  • Who the customer is and what they've experienced.
  • Which product, plan, order, or service is involved.
  • What policies apply to the situation.
  • What actions have already been attempted.
  • Which systems hold the current state.
  • What the user and organization are allowed to change.
  • Whether the proposed action can be reviewed or reversed.

A next best action without shared memory is a next best guess.

This is where the evaluation criteria change. A buyer shouldn't score a tool only on recommendation relevance. They should also assess context coverage, decision transparency, action controls, system updates, and recovery options.

A strong platform should answer ten questions:

  1. What does the system know about this case?
  2. Did the system identify the right action?
  3. Did it show the context behind that action?
  4. Did it respect the agent's permissions?
  5. Did it request approval at the right point?
  6. Did it update the system of record?
  7. What will change after the action runs?
  8. Could the action be audited?
  9. Could it be reversed?
  10. Did the customer's issue move closer to resolution?

These questions make decision quality visible. They also reveal whether the product is an assistant, a copilot, or an action-capable agent. The practical difference is simple: after the recommendation appears, does the system continue the workflow, or does the agent have to take over?

How does an agent know the right next action?

An agent needs current, permission-aware context from the customer record, ticket history, policies, product data, transaction state, and related work. It also needs clear business logic that explains which action is allowed, when approval is required, and how the result should be recorded.

That's why the enterprise AI memory behind the decision matters. Shared context helps the system understand relationships across systems instead of treating every ticket as an isolated prompt.

The system might need to connect a ticket to an account, payment, entitlement, product incident, engineering issue, internal approval, or previous exception. Those relationships often decide the correct action.

The goal isn't to make the agent sound more certain. The goal is to make the decision more defensible.

Key takeaway: The next best action is only as trustworthy as the context behind it. Connected data and explicit rules turn a plausible suggestion into a defensible decision.

What does Enterprise-Bench show about action quality?

Enterprise-Bench is an enterprise AI benchmark developed with the Laude Institute and validated by Alexandros Dimakis, UC Berkeley professor and DevRev board member. It models a realistic B2B company with 42 customer accounts, 40 product parts, and five interconnected enterprise systems – the kind of sprawling, siloed data environment where production AI actually operates.

Results on identical L1–L2 tasks:

  • Accuracy: Computer reached 94.3% vs. 63.6% for Claude Code – a 48% gap.
  • Same model, different architecture: Both systems ran on the same Opus 4.8 model family.

The accuracy gap came from how each system reached and organized enterprise data, not from the LLM underneath it.

What this means for next best action AI: If the system can't reach the facts that determine whether an action is safe, a more capable model can still produce the wrong result. Enterprise-Bench doesn't argue that every recommendation should become autonomous. It argues that decision quality has to be evaluated with realistic enterprise context, because retrieval architecture is a stronger predictor of production performance than the choice of foundation model.

A useful test is to compare two prompts:

  • Prompt 1: "The customer says they were charged twice. What should I say?"
  • Prompt 2: "The customer says they were charged twice. Check payment state, prior refunds, plan policy, fraud status, and related incidents. What action is allowed, and what should be written back?"

The second prompt reflects the actual service decision. It gives the system what it needs to tell a response apart from a resolution.

Key takeaway: Better models help, but better context can matter just as much. Enterprise-Bench ties action quality to the information available at decision time.

How does the next best action become a safe action?

A recommendation becomes a safe action when the system can understand the context, apply reusable decision logic, respect permissions, record the result, and reverse the change when needed. Each of those five controls has to hold before an AI-suggested step is allowed to touch a customer record.

Computer, by DevRev, approaches this as an AI resolution platform that works with existing business tools . It doesn't replace those systems. It connects the context and the work across them.

The operating model has three parts:

  • Computer Memory is a persistent, permission-aware layer that remembers your docs, deals, teams, and workflows. It connects relevant customer, product, support, and operational context and helps the system understand relationships and current state.
  • Agent Studio Skills encode reusable business logic. Teams define how an agent should investigate, decide, ask for approval, and complete a repeatable workflow.
  • Safe Actions govern execution. Actions can be permissioned, logged, and reversible, with human approval for sensitive work.

The flow looks like this:

Recommendation → context check → permission check → approval if needed → system update → audit trail

That's different from asking an agent to copy a suggested answer into another tool. Computer reads from connected systems and writes back to your systems when the workflow allows it.

For example, a billing skill could:

  1. Check whether a duplicate charge exists.
  2. Confirm whether a refund is already pending.
  3. Review plan and fraud policies.
  4. Ask for approval if the amount or risk exceeds a threshold.
  5. Issue the permitted refund.
  6. Update the case and the payment record.
  7. Preserve an audit trail.
  8. Reverse the action if the workflow supports rollback.

This pattern isn't theoretical. BILL, a financial operations platform serving 500,000+ businesses, deployed Computer, by DevRev, as its customer-facing AI agent. On a proof of concept of 200,000 real customer queries, the agent reached a 70% resolution rate — resolving common issues in seconds, and for complex cases, triaging and executing automated workflows or escalating to a human with full context preserved. Read the BILL story.

The pattern generalizes; the controls don't

The same next-best-action pattern supports account changes, entitlement updates, ticket routing, incident escalation, or follow-up communication. The action changes by use case. The controls stay consistent.

The agent still owns high-stakes decisions. The difference is that the workflow makes the safe path the easy path – by keeping the same guardrails, audit trails, and approval gates in place no matter what the agent is doing.

What "governed" actually requires

Because not every action carries the same risk, governed workflows do two things: they surface uncertainty instead of guessing, and they tier actions by risk.

If required data is missing, the system should ask for it or escalate. It shouldn't fill the gap with a confident guess.

Organizations should define action tiers. L

  • Low-risk actions, such as adding a tag or summarizing a case, may run automatically with lightweight logging.
  • Medium-risk actions – updating ticket status, modifying non-critical fields – may execute but require post-hoc review or confidence thresholds.
  • High-risk actions, such as refunds, account changes, or external communications, need explicit approval and stronger audit controls before execution.

This risk-based design lets teams expand automation without treating every action as equally safe. The hard part isn't the model; it's the plumbing, permissions, and write-back rules around it.

Key takeaway: The more work AI handles, the more important it becomes to know what happened, why it happened, and how to correct it. Governed execution is what turns that visibility into operational safety, making the right next step clear, controlled, and auditable, every time.

How can support teams stop turning advice into more work?

Support teams should stop measuring AI only by suggestion quality. Measure whether the system reduces search, verification, clicks, handoffs, and unresolved follow-up work. That single shift changes which vendors look strong in an evaluation.

Ask vendors to demonstrate a realistic case, not a clean sample question. Give them a ticket that depends on customer history, policy, billing state, and a related product issue. Then watch what the system does.

The strongest next best action customer service workflows treat the agent as a decision owner, not a data-entry operator. The system prepares the evidence, applies the rules, and handles approved steps. The human focuses on judgment, empathy, exceptions, and accountability.

Teams should track outcome metrics. These measures reveal whether AI is making service better or simply making interactions faster.

Agents shouldn't have to hunt for prior context or second-guess generic advice. They should spend their time on judgment, empathy, and exceptions, while routine steps move forward through controlled workflows.

A recommendation nobody executes isn't an action. It's next-best advice.

See how Computer turns recommendations into safe actions by booking a demo with the team, and explore customer service automation for a broader view of how agentic workflows move from guidance to resolution.


Frequently Asked Questions

Neelabja Adkuloo

Neelabja Adkuloo

Member of marketing staff

Neelabja is a B2B SaaS marketer specialising in AI-driven revenue tools, CRM strategy, and sales operations content. She writes at the intersection of how AI agents are evolving from passive assistants into active employees, ones that don't just surface answers, but take action across the revenue stack. Her work draws on hands-on experience with modern sales tech stacks, with a focus on the shift from Gen 1 chatbots to Gen 3 agentic systems that read, reason, and write back.

DEVREV

See Computer work for you

Your AI teammate that finds answers, takes action, and gets work done across every tool.