Repeated Questions At Scale
Policy and product questions recur across customer service and store operations. Premium-only routing would charge repeatedly for answers that may already be approved.
Retail teams repeatedly answer questions about products, orders, policies and returns. Qua would resolve stable knowledge cheaply, retrieve changing facts within the perimeter and reserve premium reasoning for authorised exceptions.

The US Census Bureau's 2025 fourth-quarter release estimated US retail e-commerce sales at $1,192.6 billion in 2024, or 16.1% of total retail sales. The figures establish channel scale, not the number of questions suitable for AI.
The National Retail Federation and Happy Returns estimated in 2024 that merchandise returns would total $850 billion that year, representing 15.8% of annual retail sales. Returns create policy and evidence questions, but their merchandise value is not an AI savings opportunity.
The National Retail Federation's 2023 National Retail Security Survey estimated shrink at $112.1 billion in 2022. That provides context for controlled access to loss-prevention information, not a claim that Qua would prevent or recover those losses.
Policy and product questions recur across customer service and store operations. Premium-only routing would charge repeatedly for answers that may already be approved.
Stock, order and promotion data can change quickly. Cached answers would need scope and freshness checks before reuse.
Returns and complaints can need judgement without justifying unrestricted automation. A reasoning assistant should not inherit authority to issue refunds or alter accounts.
Each application is a real workflow, mapped to the tier that could answer it. The waterfall tries the cheapest trustworthy source first and only pays a frontier model when the expected value clears the gate.
Agents would reuse approved answers without generating policy from memory. Personalised eligibility would remain a separate check against authorised order facts.
Support teams would see the evidence behind an order-status answer. Each receipt would show exact query cost, the premium baseline cost, savings and what stayed private.
Catalogue teams would review structured drafts against a controlled taxonomy. The workflow would not invent dimensions, certifications or safety claims absent from the source.
Agents would start with a response grounded in the case rather than a generic model answer. Human approval would remain necessary for commitments, compensation and outbound messages.
Specialists would receive a comparison of evidence, applicable policy and unresolved questions. The assistant would neither issue refunds nor make autonomous fraud allegations.
A thesis, not a case history. The assumptions are stated so you can replace them with your own numbers — which is exactly what a pilot does in week one.
Ledger rates are the vendors’ own published list prices: $0.0150 per premium question and $0.00065 per fast question at 1,500 input / 700 output tokens.
At published vendor prices this thesis models $35,344 of avoided annual inference spend — $36,000 down to $656, a 98.2% reduction — before the excluded costs above.
For PCI DSS v4.0.1, the proposed design would keep payment credentials out of prompts and exclude sensitive authentication data from retrieval and receipts. Masking and provider blocks would run before routing, with payment-environment scope assessed separately.
For GDPR Article 5, case-scoped access and data minimisation would limit what the waterfall could retrieve or disclose. Decisions within Article 22's scope would require a separate legal assessment; exception recommendations would receive meaningful human review.
For the US FTC's 16 CFR Part 465 rule on consumer reviews and testimonials, catalogue and service workflows would prohibit fabricated reviews and false testimonial generation. Policy controls would constrain routes and inputs, but content governance and human review would still be required.
Market figures come from the publishers below. Qua savings are modelled from the vendors’ published list prices — they are not customer results.