Industry thesis

Reserve premium reasoning for material financial decisions

Banks, asset managers and insurers could resolve routine questions without paying for premium inference. This thesis models how Qua would reserve expensive reasoning for permitted work where its expected value exceeds its incremental cost.

Industry thesis — modelled, not deployed. Figures are potential savings.
$200–340bnEstimated annual banking AI valueMcKinsey, 2023
$9.6tnDaily foreign-exchange turnover, April 2025Bank for International Settlements, 2025
$4.44mAverage data-breach cost, globalIBM, 2025
$670kAdded cost when AI access is ungovernedIBM, 2025
Abstract artwork representing finance workflows
Finance — the cheapest trustworthy source answers first; frontier reasoning is paid for only when it earns its price.
Where the money goes today

Finance: the cost of asking

McKinsey's 2023 report, The Economic Potential of Generative AI, estimated annual value of $200–340 billion for banking. That is an industry opportunity estimate, not evidence of realised savings or a forecast for Qua.

The Bank for International Settlements' 2025 Triennial Central Bank Survey put daily foreign-exchange turnover at $9.6 trillion in April 2025. At that scale, operational questions need controlled access to current procedures rather than an expensive model by default.

IBM's 2025 Cost of a Data Breach Report put the global average breach cost at $4.44 million, and found breaches involving ungoverned AI access added about $670,000. Any routing thesis must therefore treat data exposure as a constraint, not merely a cost variable.

Three pressures

What makes this industry different

Routine questions accumulate cost

Policy interpretation, onboarding and servicing generate repeated questions. Premium-only access would charge repeatedly for answers the institution already holds.

Entitlements cross business boundaries

Research, customer records and claims evidence have different access rules. Retrieval and provider selection must respect those differences before routing.

Material decisions need scrutiny

Credit, investment and claims decisions require accountable judgement. A more expensive answer is not automatically a more defensible one.

Five applications

How Qua would run inside finance

Each application is a real workflow, mapped to the tier that could answer it. The waterfall tries the cheapest trustworthy source first and only pays a frontier model when the expected value clears the gate.

Approved onboarding policy answers

Resolves at Personal Knowledge~30,000 questions/month

Staff could reuse approved onboarding explanations instead of regenerating them. Expiry checks would prevent superseded policies from remaining trusted answers.

Potential saving — 90–100% of query spend.
  • Apply role entitlements, masking and the maximum-tier policy.
  • Match the question to a verified, current Personal Knowledge answer.
  • Return the cached answer and receipt; escalate misses through permitted tiers.

Investment mandate restriction lookup

Resolves at Enterprise Search~30,000 questions/month

Portfolio teams could find the relevant restriction and its source clause. The answer would support review, not authorise a trade.

Potential saving — 80–100% of query spend.
  • Enforce portfolio entitlements and provider restrictions before routing.
  • Check verified answers, then search current mandates inside the perimeter.
  • Return cited restrictions and a receipt without external model inference.

Insurance coverage clause retrieval

Resolves at Enterprise Search~24,000 questions/month

Claims handlers could locate relevant wording without sending claim files to an external model. Coverage determinations would remain with authorised staff.

Potential saving — 80–100% of query spend.
  • Apply claim-level access controls and mask unnecessary personal data.
  • Check verified answers, then retrieve the applicable policy wording and endorsements.
  • Return source-grounded clauses and a receipt within the perimeter.

Customer servicing correspondence drafts

Resolves at Fast Models~24,000 questions/month

A cheap model could adapt approved material to a specific servicing request. Staff would review the draft before release.

Potential saving — 80–95% of query spend.
  • Mask customer identifiers and apply channel-specific provider rules.
  • Reuse verified templates and retrieve the relevant servicing procedure.
  • Use a permitted Fast Model for the remaining drafting task and issue a receipt.

Credit exception committee preparation

Resolves at Pro Models~12,000 questions/month

Complex exceptions could receive deeper comparison of evidence and policy trade-offs. The committee would retain decision authority and review all material claims.

Potential saving — 0–20% of query spend.
  • Apply borrower-data masking, provider blocks and maximum-tier limits.
  • Check verified answers and retrieved evidence for a sufficient lower-tier response.
  • Use an allowed Pro Model only if the expected-value gate clears; attach a receipt.
The modelled ledger

What the same year of questions could cost

A thesis, not a case history. The assumptions are stated so you can replace them with your own numbers — which is exactly what a pilot does in week one.

xAI Grok 4$3.00 in · $15.00 out / 1M tokensxAI published API pricing
xAI Grok 4 Fast$0.20 in · $0.50 out / 1M tokensxAI published API pricing
OpenAI GPT-5.5$5.00 in · $20.00 out / 1M tokensOpenAI published API pricing
Google Gemini 3.7 Flash$0.20 in · $0.80 out / 1M tokensGoogle AI published pricing

Ledger rates are the vendors’ own published list prices: $0.0150 per premium question and $0.00065 per fast question at 1,500 input / 700 output tokens.

Modelled annual comparison
  • Workload: 120,000 questions/month for 12 months, at 1,500 input and 700 output tokens per question.
  • Published list prices, not estimates — premium baseline xAI Grok 4 at $3.00/$15.00 per 1M tokens = $0.0150/question.
  • Cheap metered tier xAI Grok 4 Fast at $0.20/$0.50 per 1M tokens = $0.00065/question.
  • Terminal-tier mix: 70% verified or in-perimeter, 20% Fast Models, 10% Pro Models. Rows exclude Qua fees, retrieval infrastructure, integration and human review.
WorkloadFrontier-onlyWith Qua
Verified answers and in-perimeter search1,008,000 annual questions resolved with no external model call.$15,120$0
Fast Models (Grok 4 Fast)288,000 annual questions at xAI's published Grok 4 Fast rate ($0.00065/question).$4,320$187
EV-gated Pro Models (Grok 4)144,000 annual questions keep frontier reasoning at the published Grok 4 rate.$2,160$2,160

At published vendor prices this thesis models $19,253 of avoided annual inference spend — $21,600 down to $2,347, a 89.1% reduction — before the excluded costs above.

Controls that matter here

Policy runs before routing, not after

Enforce financial data boundaries

To support GDPR data minimisation and security obligations, masking, provider blocks and maximum-tier clamps would run before routing, while Enterprise Search would stay inside the perimeter. Qua Cloud, customer VPC or air-gapped deployment would require an institution-specific assessment.

Record operational routing evidence

For DORA ICT risk-management and oversight work, every answer would carry a receipt showing exact inference cost, the configured premium baseline cost, savings and what stayed private. This evidence would support oversight, not establish DORA compliance.

Preserve model risk accountability

For institutions applying Federal Reserve SR 11-7, a pilot would document intended use, limitations, validation and human review. Pro Models would require both policy permission and an expected-value gate; Open Models would be user-picked only.

  • Week one would measure the share of questions resolved by verified answers or perimeter search without external inference.
  • Week one would compare receipt-level costs with a matched premium baseline and test answer quality through blinded reviewer scoring.
  • Week one would test entitlement leakage, masking failures and blocked-provider attempts using seeded sensitive records.
Sources

Every figure on this page, traceable

Market figures come from the publishers below. Qua savings are modelled from the vendors’ published list prices — they are not customer results.

  1. [1]McKinsey & Company, The economic potential of generative AI: The next productivity frontier (2023). Source
    Estimated annual generative-AI value for banking of $200–340 billion; this is potential industry value, not realized savings or a company forecast.
  2. [2]Bank for International Settlements, Triennial Central Bank Survey: OTC foreign exchange turnover in April 2025 (2025). Source
    Global foreign-exchange trading averaged $9.6 trillion per day in April 2025, up from $7.5 trillion in 2022.
  3. [3]IBM / Ponemon Institute, Cost of a Data Breach Report 2025: The AI Oversight Gap (2025). Source
    IBM's 2025 study reported a global average breach cost of $4.44 million, down 9%, and an additional $670,000 where AI access was ungoverned.
  4. [4]European Union, Regulation (EU) 2016/679 — General Data Protection Regulation, Articles 5, 25 and 32 (2016). Source
    GDPR requires purpose limitation, data minimization and appropriate security for covered personal-data processing, supporting financial-data access boundaries.
  5. [5]Board of Governors of the Federal Reserve System, SR 11-7: Guidance on Model Risk Management (2011). Source
    For banking organizations within scope, model-risk guidance establishes expectations for validation, governance, documentation and effective challenge.
  6. [6]European Union, Regulation (EU) 2024/1689 — Artificial Intelligence Act (2024). Source
    Applicable high-risk AI systems face documentation, logging and human-oversight duties under the Act's phased application; these are not blanket requirements for every financial AI router.