Skip to main content

AI & assistant-friendly summary

This section provides structured content for AI assistants and search engines. You can cite or summarize it when referencing this page.

Summary

Forty zaps, one chatbot per OMS code, two refunds on one order. One investigator loop this week. Tools per class. Investigate and recommend. The ~$791/mo platform cost floor at 50,000 sessions is planning cost, not holds closed.

Key Facts

  • The ~$791/mo platform cost floor at 50,000 sessions is planning cost, not holds closed
  • This is series part 54: the reusable exception-agent pattern — detect → investigate → collect evidence → recommend → request approval or execute
  • After July 30, 2026, net-new exception agents should not start on Agents Classic
  • FactualMinds is an AWS Select Tier Services Partner
  • Fill post 8 for Auto vs HITL | | Inventory | ATP vs advertised, receipt miss | Analyze + recommend

Entity Definitions

Amazon Bedrock
Amazon Bedrock is an AWS service discussed in this article.
Bedrock
Bedrock is an AWS service discussed in this article.

The AI Exception Agent: Automatically Investigating eCommerce Business Problems (2026)

AI AgentsPalaniappan P9 min read

Quick summary: Forty zaps, one chatbot per OMS code, two refunds on one order. One investigator loop this week. Tools per class. Investigate and recommend. The ~$791/mo platform cost floor at 50,000 sessions is planning cost, not holds closed.

Key Takeaways

  • The ~$791/mo platform cost floor at 50,000 sessions is planning cost, not holds closed
  • This is series part 54: the reusable exception-agent pattern — detect → investigate → collect evidence → recommend → request approval or execute
  • After July 30, 2026, net-new exception agents should not start on Agents Classic
  • FactualMinds is an AWS Select Tier Services Partner
  • Fill post 8 for Auto vs HITL | | Inventory | ATP vs advertised, receipt miss | Analyze + recommend
Multiple business exceptions converging into one investigation and resolution workflow
Table of Contents

Friday afternoon, finance finds two refunds on one order. Support issued one for “be kind.” The shortage zap issued another. Forty disconnected prompts. One shared Admin token.

This is series part 54: the reusable exception-agent pattern — detect → investigate → collect evidence → recommend → request approval or execute. Inputs are order, inventory, customer, and payment exceptions. It is not a remake of AI order exception management — that is the order specialization (payment fail, address, delay, fulfillment, fraud) with Auto vs HITL columns you still have to fill. Order IDs and decline codes in artifacts are demo shapes. They are not a FactualMinds refund KPI. After July 30, 2026, net-new exception agents should not start on Agents Classic.

The job. One investigator loop. Tools per class. Evidence on every finding.

This week. Default autonomy: investigate + recommend. No auto-refund from a story.

A person still signs. Refunds, cancels, account writes, ATP mutations, and fraud-release.

Skip it when a decline code already maps to “retry once” in the payments processor and no cross-system judgment is needed.

Our take: one harness (or one Runtime service) with tools per class, not forty disconnected zaps. Trade-off: a new exception type waits until you add a tool and a Cedar rule. You also do not train two specialists to refund the same order in one afternoon.

Copy the pattern — Clone exception-agent-pattern.md. Fill your Detect sources and default autonomy. For orders, also fill order-exception-playbook.md. Series folder: ecommerce-ai-agents-series/.

FactualMinds is an AWS Select Tier Services Partner. We help merchants sequence this loop — we do not sell auto-refund as kindness.

Signals that start the loop live in store monitoring. Writes that must stop live in HITL. The mixed ops queues are the back-office matrix.

The loop (reuse it; do not rename it per ticket type)

Detect → Investigate → Collect evidence → Recommend → Request approval or execute
flowchart TD
  Detect[Detect]
  Investigate[Investigate]
  Evidence[CollectEvidence]
  Recommend[Recommend]
  Gate[ApproveOrExecute]
  Detect --> Investigate
  Investigate --> Evidence
  Evidence --> Recommend
  Recommend --> Gate
StepWho owns itFailure if you skip
DetectDeterministic rule or webhookMissed holds; invented holds
InvestigateNamed tools onlyFive UI hops and a guess
Evidenceevidence_tool + evidence_ref on every findingA story that no system measured
RecommendStructured decision, autonomy taggedSlack prose that cannot be approved
Approve or executeCedar + HITL per autonomy spectrumPrompt text as authorization

Finance breaks if refunds fire without a playbook; warehouse breaks if every shortage becomes a silent cancel; CRM breaks if duplicate-customer “cleanup” writes accounts. Build the agent when investigation spans more than one system and the next action is ambiguous.

Four input classes (defaults, not a client mix)

From exception-agent-pattern.md:

Exception classExamplesDefault autonomy
OrderPayment fail, address, delayInvestigate + recommend. Fill post 8 for Auto vs HITL
InventoryATP vs advertised, receipt missAnalyze + recommend. No qty write in week one
CustomerDuplicate create, B2B credit holdRecommend. No account write
PaymentDecline codes, AVSRetry only on an allow-list. AVS / fraud-tool hold → human

Do not flatten order into “payment” and delete the playbook. Post 8 still owns fulfillment short-pick, fraud-release, and split-ship that spends margin. This table is the shared loop those rows sit on.

Order (specialization, not a copy)

Shopper-facing WISMO answers where the order is. Support triages returns. Exception management sits on the ops side of the OMS: the order already failed a happy-path rule. Same Gateway family, different Identity, different write tools. Link the playbook. Do not paste its six-class table into this post and call it new.

Inventory

Available vs reserved is a number. The judgment is wait, split, substitute, or cancel. The agent recommends with getInventory and asOf. Cancel and ATP mutation stay HITL. Advertised ATP=0 is often a monitor detect, then this investigate step.

Customer

Duplicate create spikes and B2B credit holds are quality and AM problems. The agent may cluster and recommend a merge or a hold review. It must not updateCustomer or touch PII beyond what the tool already strips. Prompt “do not show PII” is not a control.

Payment

Most declines are already classified by the processor. The agent adds value when OMS status does not match what finance thinks posted. Auto: retry only on an allow-list. HITL: AVS mismatch, 3DS failure, fraud-tool hold. Unconstrained retries look like card testing.

Evidence is the product

Every finding needs both fields. If a hop has no tool, stop and list the gap.

{
  "exception_id": "EXC-DEMO-001",
  "class": "inventory",
  "detect": { "source": "webhook:oms.exception", "code": "SHORTAGE" },
  "findings": [
    {
      "claim": "ATP 0 on advertised SKU DEMO-1 asOf fixture timestamp",
      "evidence_tool": "getInventory",
      "evidence_ref": "sku:DEMO-1",
      "causation": "possible_contributor"
    }
  ],
  "unknowns": ["Substitute availability — no getSubstitutes tool"],
  "recommended_action": {
    "action": "Hold and recommend split",
    "autonomy": "recommend"
  },
  "human_required": true
}

Fixture IDs and SKUs are demo. Do not cite them as a store outcome. The same evidence contract shows up in alerts and RCA. Share the schema; do not share refund tools with the shopper-facing agent.

One fixture walkthrough (replace the IDs)

Not a client case. EXC-DEMO-001, SKU DEMO-1.

  1. Detect — OMS webhook SHORTAGE on order ORD-DEMO-9. No model in this step.
  2. InvestigategetOrder, getInventory (require asOf), getShipment if the playbook says the class needs it. Stop if a tool 404s; do not guess ATP.
  3. Evidence — ATP 0 on advertised DEMO-1 with evidence_tool: getInventory. Substitute availability is an unknown if getSubstitutes is not on the OpenAPI.
  4. Recommend — hold and split; autonomy recommend. Cancel/refund stay HITL per post 8.
  5. Approve or execute — Cedar DENY refundOrder for this role; HITL queue gets session id + tool trace.

That is the whole pattern. Payment-fail and duplicate-customer reuse steps 1–5 with different tools. They do not reuse refund tools with the shopper-facing support agent.

Anti-patternReplacementWhen the replacement applies
Zap per OMS codeOne investigator, tools per classYou have more than a handful of codes and shared writes
Model as detectorWebhook / queue / joinOMS or WMS already emits a code
Shared Admin tokenIdentity JWT + CedarAny write exists
Prompt “be careful”HITL queueRefund, cancel, account, ATP mutation

For your technical lead

On June 17, 2026, Amazon Bedrock AgentCore Harness reached general availability — a config-driven loop on the same platform as Runtime, Memory, Gateway, Identity, and Policy (What’s New). Agents Classic is in maintenance for new customers after July 30, 2026. Net-new exception agents should use Bedrock AgentCore. Full matrix: lifecycle roundup.

First-party signals we reuse (not eCommerce outcomes) — Gateway server-side tools cut median tool round-trip ~180 ms → ~95 ms on a B2B CRM assistant (12 tools, ~8k turns/day) — Gateway post. Platform TCO silhouette: support-style AgentCore at 50K sessions/mo ~$791/mo platform + model (decision guide). Exception volume is usually far below shopper chat; still model platform + tokens on the AgentCore pricing calculator. Treat ~$791/mo as a platform cost floor, not as savings from fewer holds.

PieceRole here
GatewayPer-class reads; narrow writes only after Cedar
Policy (Cedar)Default-deny refundOrder, cancelOrder, updateInventory, updateCustomer; allow retryPayment only with decline-class entity
IdentityAssociate vs shopper. Shopper JWTs DENY every exception write
MemoryException-id scoped; no PAN, no raw 3DS
HITLQueue with session id and tool trace — HITL architecture

Run Policy LOG_ONLY, then ENFORCE. The store-agents Cedar sample already gates cancelOrder and createReturn. Reuse it. There is no native Shopify AgentCore connector.

Gateway ~180 → ~95 ms is the CRM platform canary. Absolute latency is OMS + payments + WMS.

What broke — A program plan that staffed one zap per OMS exception code (forty prompts, shared Admin token). Two specialists both called refundOrder on the same fixture hold: support for “be kind,” exceptions for shortage. Detection: Gateway traces showed two write tools, two idempotency keys, one order. Fix: one investigator; tools per class; Cedar default-deny refund; HITL for refund/cancel; order rows stay in the order playbook. Lesson: reuse the loop. Do not multiply chatbots.

A second counter-case: detect left in the model (“look for weird orders”). The OMS already emitted payment_failed. The agent retried through a fraud-tool hold — the same class of bug post 8 already documented. Fix: webhook first; allow-list retries in Cedar and in the payments adapter, not only in the prompt.

What to do this week

  1. Copy exception-agent-pattern.md. Map your four classes to tools.
  2. For orders, fill order-exception-playbook.md — do not skip it.
  3. Put Detect in webhooks / queue depth / joins. Do not ask the model to notice.
  4. Require evidence_tool + evidence_ref. Refuse recommendations that skip them.
  5. Default autonomy: investigate + recommend. Retry only on a payments-owned allow-list.
  6. One harness (or one Runtime service). Kill the zap-per-code plan.
  7. HITL queue for refund, cancel, account, ATP write — hitl-approval-architecture.md.
  8. Run monday-checklist.md.
  9. Model sessions on the AgentCore pricing calculator.
  10. Contact us for an architecture conversation. Bedrock path: Amazon Bedrock.

If you only do one thing: stop staffing a chatbot per exception code. Share the loop; specialize the tools.

What this post doesn’t cover

FAQ

When should you NOT staff a new chatbot per exception type?

Skip a new prompt per OMS code when the loop is the same: detect, investigate, evidence, recommend, approve or execute. Tools and Cedar change by class. Forty disconnected zaps is how you get two refunds on one order. One harness (or one Runtime service) with tools per class.

What could go wrong if this post replaces the order-exception playbook?

You lose the six order classes, Auto vs HITL columns, and the fraud-hold veto. This is the reusable pattern. Post 8 is the order specialization. Fill that playbook for payment fail, address, delay, fulfillment, and fraud. Do not flatten it into four generic buckets and call the work done.

When should you NOT auto-execute after investigate?

Veto payment capture, ATP mutation, live price, account or PII writes, fraud-release, and refunds that are not on a named allow-list. Default autonomy is investigate plus recommend. Retry payment only on a decline-code allow-list your payments team owns.

What could go wrong if evidence_tool is optional?

The model will fill the gap with a story. Sample turns recommended cancel on a shortage without calling getInventory. Detection is evals that fail when evidence_tool or evidence_ref is missing. Fix: refuse the recommendation; return the unknown.

How is this different from the monitoring agent or back-office pillar?

Monitoring decides that a signal fired and should be investigated. Back office is a task matrix of ten queues. This pattern is the investigation loop those queues share. Do not merge week-one prompts: a monitor that refunds is a purchaser you did not review.

Harness or Runtime for a cross-domain exception agent?

Harness can host one investigator with a short tool list per class. Use Runtime plus Strands when hop caps, a supervisor, or Cedar-scoped writes across payments, OMS, WMS, and CRM matter. Net-new builds should not use Agents Classic after July 30, 2026.


Need a single exception loop without a zap farm? Contact FactualMinds for an architecture conversation, or start from the order playbook.

PP
Palaniappan P

AWS Cloud Architect & AI Expert

AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.

AWS ArchitectureCloud MigrationGenAI on AWSCost OptimizationDevOps

Recommended Reading

Explore All Articles »
5 min

Who May Release a B2B Credit Hold? (2026)

Assemble the hold reason, aging, and promise-to-pay. Finance releases the order. The agent does not. The account-manager list still caps at 5 priorities — a hold is one of those 5, not a button.