Routine questions accumulate cost
Policy interpretation, onboarding and servicing generate repeated questions. Premium-only access would charge repeatedly for answers the institution already holds.
Banks, asset managers and insurers could resolve routine questions without paying for premium inference. This thesis models how Qua would reserve expensive reasoning for permitted work where its expected value exceeds its incremental cost.

McKinsey's 2023 report, The Economic Potential of Generative AI, estimated annual value of $200–340 billion for banking. That is an industry opportunity estimate, not evidence of realised savings or a forecast for Qua.
The Bank for International Settlements' 2025 Triennial Central Bank Survey put daily foreign-exchange turnover at $9.6 trillion in April 2025. At that scale, operational questions need controlled access to current procedures rather than an expensive model by default.
IBM's 2025 Cost of a Data Breach Report put the global average breach cost at $4.44 million, and found breaches involving ungoverned AI access added about $670,000. Any routing thesis must therefore treat data exposure as a constraint, not merely a cost variable.
Policy interpretation, onboarding and servicing generate repeated questions. Premium-only access would charge repeatedly for answers the institution already holds.
Research, customer records and claims evidence have different access rules. Retrieval and provider selection must respect those differences before routing.
Credit, investment and claims decisions require accountable judgement. A more expensive answer is not automatically a more defensible one.
Each application is a real workflow, mapped to the tier that could answer it. The waterfall tries the cheapest trustworthy source first and only pays a frontier model when the expected value clears the gate.
Staff could reuse approved onboarding explanations instead of regenerating them. Expiry checks would prevent superseded policies from remaining trusted answers.
Portfolio teams could find the relevant restriction and its source clause. The answer would support review, not authorise a trade.
Claims handlers could locate relevant wording without sending claim files to an external model. Coverage determinations would remain with authorised staff.
A cheap model could adapt approved material to a specific servicing request. Staff would review the draft before release.
Complex exceptions could receive deeper comparison of evidence and policy trade-offs. The committee would retain decision authority and review all material claims.
A thesis, not a case history. The assumptions are stated so you can replace them with your own numbers — which is exactly what a pilot does in week one.
Ledger rates are the vendors’ own published list prices: $0.0150 per premium question and $0.00065 per fast question at 1,500 input / 700 output tokens.
At published vendor prices this thesis models $19,253 of avoided annual inference spend — $21,600 down to $2,347, a 89.1% reduction — before the excluded costs above.
To support GDPR data minimisation and security obligations, masking, provider blocks and maximum-tier clamps would run before routing, while Enterprise Search would stay inside the perimeter. Qua Cloud, customer VPC or air-gapped deployment would require an institution-specific assessment.
For DORA ICT risk-management and oversight work, every answer would carry a receipt showing exact inference cost, the configured premium baseline cost, savings and what stayed private. This evidence would support oversight, not establish DORA compliance.
For institutions applying Federal Reserve SR 11-7, a pilot would document intended use, limitations, validation and human review. Pro Models would require both policy permission and an expected-value gate; Open Models would be user-picked only.
Market figures come from the publishers below. Qua savings are modelled from the vendors’ published list prices — they are not customer results.