Industry thesis

Keep routine commerce questions off premium models

Retail teams repeatedly answer questions about products, orders, policies and returns. Qua would resolve stable knowledge cheaply, retrieve changing facts within the perimeter and reserve premium reasoning for authorised exceptions.

Industry thesis — modelled, not deployed. Figures are potential savings.
$1,192.6bnUS retail e-commerce sales, 2024US Census Bureau, 2025
16.1%E-commerce share of US retail, 2024US Census Bureau, 2025
$850bnEstimated merchandise returns, 2025National Retail Federation/Happy Returns, October 2025
$316.1bnUS e-commerce sales, Q4 2025US Census Bureau, March 2026
Abstract artwork representing retail & e-commerce workflows
Retail & e-commerce — the cheapest trustworthy source answers first; frontier reasoning is paid for only when it earns its price.
Where the money goes today

Retail & e-commerce: the cost of asking

The US Census Bureau's 2025 fourth-quarter release estimated US retail e-commerce sales at $1,192.6 billion in 2024, or 16.1% of total retail sales. The figures establish channel scale, not the number of questions suitable for AI.

The National Retail Federation and Happy Returns estimated in 2024 that merchandise returns would total $850 billion that year, representing 15.8% of annual retail sales. Returns create policy and evidence questions, but their merchandise value is not an AI savings opportunity.

The National Retail Federation's 2023 National Retail Security Survey estimated shrink at $112.1 billion in 2022. That provides context for controlled access to loss-prevention information, not a claim that Qua would prevent or recover those losses.

Three pressures

What makes this industry different

Repeated Questions At Scale

Policy and product questions recur across customer service and store operations. Premium-only routing would charge repeatedly for answers that may already be approved.

Changing Facts Need Grounding

Stock, order and promotion data can change quickly. Cached answers would need scope and freshness checks before reuse.

Exceptions Require Bounded Authority

Returns and complaints can need judgement without justifying unrestricted automation. A reasoning assistant should not inherit authority to issue refunds or alter accounts.

Five applications

How Qua would run inside retail & e-commerce

Each application is a real workflow, mapped to the tier that could answer it. The waterfall tries the cheapest trustworthy source first and only pays a frontier model when the expected value clears the gate.

Returns policy FAQ lookup

Resolves at Personal Knowledge~100,000 questions/month

Agents would reuse approved answers without generating policy from memory. Personalised eligibility would remain a separate check against authorised order facts.

Potential saving — 90-100% of query spend.
  • Apply market, channel and policy-date filters with an Enterprise Search tier clamp.
  • Check verified Personal Knowledge answers for the relevant policy version.
  • On a miss, search current policy documents inside the perimeter or abstain.

Order and fulfilment evidence lookup

Resolves at Enterprise Search~60,000 questions/month

Support teams would see the evidence behind an order-status answer. Each receipt would show exact query cost, the premium baseline cost, savings and what stayed private.

Potential saving — 85-100% of query spend.
  • Verify agent and customer scope, mask payment data and block external providers.
  • Check freshness-qualified answers, then retrieve authorised order and fulfilment records.
  • Return grounded status with timestamps or flag unavailable current data.

Product attribute normalisation

Resolves at Fast Models~30,000 questions/month

Catalogue teams would review structured drafts against a controlled taxonomy. The workflow would not invent dimensions, certifications or safety claims absent from the source.

Potential saving — 70-90% of query spend.
  • Apply supplier-data permissions, confidentiality masking and a Fast Models tier clamp.
  • Check verified mappings, then retrieve the taxonomy and source product attributes.
  • Use Fast Models for unresolved normalisation and flag unsupported attributes.

Customer service reply drafting

Resolves at Fast Models~8,000 questions/month

Agents would start with a response grounded in the case rather than a generic model answer. Human approval would remain necessary for commitments, compensation and outbound messages.

Potential saving — 70-90% of query spend.
  • Apply case-level access, personal-data masking and approved-provider rules.
  • Check approved replies, then retrieve relevant policy and authorised case facts.
  • Use Fast Models for the remaining draft without sending or executing an action.

Complex returns exception review

Resolves at Pro Models~2,000 questions/month

Specialists would receive a comparison of evidence, applicable policy and unresolved questions. The assistant would neither issue refunds nor make autonomous fraud allegations.

Potential saving — 0-30% of query spend.
  • Apply case permissions, payment-data exclusion and permitted-provider restrictions.
  • Check approved precedents, retrieve policy and transaction evidence, then assess Fast Models capability.
  • Use Pro Models only when expected review value exceeds incremental cost; route recommendations to a human.
The modelled ledger

What the same year of questions could cost

A thesis, not a case history. The assumptions are stated so you can replace them with your own numbers — which is exactly what a pilot does in week one.

xAI Grok 4$3.00 in · $15.00 out / 1M tokensxAI published API pricing
xAI Grok 4 Fast$0.20 in · $0.50 out / 1M tokensxAI published API pricing
OpenAI GPT-5.5$5.00 in · $20.00 out / 1M tokensOpenAI published API pricing
Google Gemini 3.7 Flash$0.20 in · $0.80 out / 1M tokensGoogle AI published pricing

Ledger rates are the vendors’ own published list prices: $0.0150 per premium question and $0.00065 per fast question at 1,500 input / 700 output tokens.

Modelled annual comparison
  • Workload: 200,000 questions/month for 12 months, at 1,500 input and 700 output tokens per question.
  • Published list prices, not estimates — premium baseline xAI Grok 4 at $3.00/$15.00 per 1M tokens = $0.0150/question.
  • Cheap metered tier xAI Grok 4 Fast at $0.20/$0.50 per 1M tokens = $0.00065/question.
  • Terminal-tier mix: 80% verified or in-perimeter, 19% Fast Models, 1% Pro Models. Rows exclude Qua fees, retrieval infrastructure, integration and human review.
WorkloadFrontier-onlyWith Qua
Verified answers and in-perimeter search1,920,000 annual questions resolved with no external model call.$28,800$0
Fast Models (Grok 4 Fast)456,000 annual questions at xAI's published Grok 4 Fast rate ($0.00065/question).$6,840$296
EV-gated Pro Models (Grok 4)24,000 annual questions keep frontier reasoning at the published Grok 4 rate.$360$360

At published vendor prices this thesis models $35,344 of avoided annual inference spend — $36,000 down to $656, a 98.2% reduction — before the excluded costs above.

Controls that matter here

Policy runs before routing, not after

Payment Data Routing Restrictions

For PCI DSS v4.0.1, the proposed design would keep payment credentials out of prompts and exclude sensitive authentication data from retrieval and receipts. Masking and provider blocks would run before routing, with payment-environment scope assessed separately.

Customer Privacy And Review

For GDPR Article 5, case-scoped access and data minimisation would limit what the waterfall could retrieve or disclose. Decisions within Article 22's scope would require a separate legal assessment; exception recommendations would receive meaningful human review.

Authentic Reviews And Claims

For the US FTC's 16 CFR Part 465 rule on consumer reviews and testimonials, catalogue and service workflows would prohibit fabricated reviews and false testimonial generation. Policy controls would constrain routes and inputs, but content governance and human review would still be required.

  • In week one, measure verified-answer reuse and search resolution by workflow, including stale-answer rejection.
  • In week one, compare query spend and agent acceptance with a matched premium-only baseline, including retries and escalations.
  • In week one, test cross-customer access denial, payment-data exclusion and blocked-provider behaviour before any live rollout.
Sources

Every figure on this page, traceable

Market figures come from the publishers below. Qua savings are modelled from the vendors’ published list prices — they are not customer results.

  1. [1]US Census Bureau, Quarterly Retail E-Commerce Sales: 4th Quarter 2024 (2025). Source
    The fourth-quarter 2024 release estimated full-year US retail e-commerce sales at $1,192.6 billion, or 16.1% of total retail sales; these are the release's contemporary estimates.
  2. [2]National Retail Federation / Happy Returns, 2024 Consumer Returns in the Retail Industry (2024). Source
    NRF and Happy Returns projected $850 billion in merchandise returns for 2024, equivalent to 15.8% of annual retail sales; merchandise value is not an AI savings estimate.
  3. [3]US Census Bureau, Quarterly Retail E-Commerce Sales, 4th Quarter 2025 (2026). Source
    US retail e-commerce sales reached $316.1 billion in the fourth quarter of 2025 (seasonally adjusted); this measures channel scale, not AI-suitable question volume.
  4. [4]PCI Security Standards Council, Payment Card Industry Data Security Standard: Requirements and Testing Procedures, Version 4.0.1 (2024). Source
    PCI DSS requires protection of account data, business-need access restrictions and security logging; external AI routing must account for cardholder-data scope and service-provider responsibilities.
  5. [5]European Union, Regulation (EU) 2016/679 — General Data Protection Regulation, Articles 5, 22, 25 and 32 (2016). Source
    GDPR requires privacy safeguards for covered customer data and provides protections for certain solely automated decisions with legal or similarly significant effects; not every retail recommendation triggers Article 22.
  6. [6]Federal Trade Commission / Electronic Code of Federal Regulations, 16 CFR Part 465 — Rule on the Use of Consumer Reviews and Testimonials (2024). Source
    The FTC's rule prohibits specified fake or false reviews, sentiment-conditioned review purchases and certain review suppression practices; it took effect on October 21, 2024.
  7. [7]Federal Trade Commission, FTC Policy Statement Regarding Advertising Substantiation (1984). Source
    Objective advertising claims require a reasonable basis before publication, supporting substantiation and review of generated product and promotional claims.