AI adoption in Financial Services is entering a new phase, and the cost of running it is coming under closer scrutiny.

In tests shown at a recent NextWave breakfast forum, combining deterministic rules with an LLM used 20 times fewer tokens than an LLM working alone. It also caught every error, whereas the LLM alone had an error rate of around 2%.

The forum, Optimising LLM & Agentic AI token usage to maximise business value in Banking, Finance and Insurance, was hosted by NextWave Consulting in collaboration with Alteryx at The Ivy City Garden on 1 October. It brought together senior leaders from across banking, Financial Services and insurance.

The discussion was held under the Chatham House Rule, so nobody or their organisation is named here. What came through clearly was a sector becoming more practical about AI. The enthusiasm remains, but it now comes with harder questions about cost, governance, accuracy and accountability.

Moving beyond experimentation

For many firms, AI started with contained use cases and generative AI models. Staff use it to write, summarise, research and analyse information, and these tools show value quickly.

However, giving employees a generative AI tool is fairly simple, whereas putting AI into a controlled production environment requires decisions about ownership, security, data, governance, support and cost.

In Financial Services, this matters more than in most sectors because AI outputs may shape customer interactions, financial calculations, regulatory processes or important operational decisions. Attendees drew a clear line between using AI as a productivity tool and using it to run part of a business process. They also noted that AI spend is growing faster than the value firms can measure, and the gap is widening.

Where AI costs balloon

As AI usage grows, so does token consumption, and the cost that comes with it. Much of that spend is avoidable.

Poor prompts are part of the problem, but attendees pointed to several other causes:

  • Using the top model for everything: Large frontier models are often the default, even for simple tasks that a smaller model could handle.

  • Repeat submissions: Long chats, unnecessary history and agentic loops keep sending the same context back to the model.

  • Compliance-driven redundancy: Dual-model checks, using one LLM to judge another and verbose reasoning for explainability all add tokens.

  • Weak ownership: AI adopted without governance or budget monitoring makes it hard to see where the money goes.

This means managing AI costs can't be left to procurement teams negotiating token prices. The more important question is what the model is being asked to do. When an LLM works out the same rules on every run, the firm pays for that reasoning every time.

Exact answers need exact tools

Much of financial services is deterministic, with reconciliations, calculations, data transformations and policy checks following defined rules with expected outcomes. LLMs, on the other hand, are probabilistic. They're strong at interpretation, summarisation and generation, but they aren't the best tool for calculations that need to be exact.

Keeping that work in deterministic processes has clear advantages:

  • Known rules: There's no cost in asking a model to work out rules that are already written down.

  • Repeatability: The same inputs give the same outputs, with a clear trail.

  • Structured data: Many inputs and outputs are already structured, so they don't need a model to interpret them.

  • Speed and cost: Deterministic processes run faster and cost far less.

  • Lower error rates: Probabilistic outputs need extra controls to catch mistakes, which adds cost.

  • Simpler change control: Probabilistic models need ongoing evaluation, drift monitoring and re-validation whenever the model changes.

That opens the door to combining technologies. A demonstration at the breakfast showed this in practice using a front-office to back-office reconciliation.

An Alteryx workflow handled the comparison itself, and the LLM was then used to explain and summarise the results.

The AI didn't need to decide whether two values matched, because a deterministic calculation could settle that with certainty. Instead, it focused on explaining the exceptions and helping people see what needed attention.

 

What we measured: 20x fewer tokens and no missed errors

In the reconciliation test, combining deterministic and probabilistic processing used 20 times fewer tokens than an LLM working alone. It also found every break, whereas the LLM alone missed 2 in 120 records, an error rate of around 2%.

We ran the same tasks in two ways:

  • Deterministic + probabilistic: Claude was connected to Alteryx through an MCP server. Alteryx ran the rules and calculations, and Claude explained the results.
  • Probabilistic only: Claude worked through the whole task by itself.


Alongside the reconciliation of 120 records, we tested expense validation on batches of 50 and 100 claims.

Test

Deterministic + probabilistic: total tokens

Probabilistic only: total tokens

Token reduction

Deterministic + probabilistic: missed

Probabilistic only: missed

Reconciliation, 120 records (11 breaks)

4,905

99,383

95% (20x fewer)

0 of 11

2 of 11

Expense validation, 50 claims (24 to flag)

7,161

76,417

91%

0 of 24

2 of 24

Expense validation, 100 claims (46 to flag)

8,852

156,224

94%

0 of 46

3 of 46

 

The combined approach was also much faster. Expense validation took 1 to 2 minutes, compared with 6 to 13 minutes for the LLM alone, a time saving of 83% to 87%.

Accuracy matters as much as cost

Cutting costs can't come at the expense of accuracy, and the group returned several times to this point.

A model that's right 95% of the time sounds impressive in a general-purpose tool. However, in a daily control process, that remaining 5% is a stream of exceptions someone must find and fix. Even the 2% error rate in our test meant missing breaks that a reconciliation exists to catch. Therefore, the real test for Financial Services is whether an AI system is reliable enough for a specific process, and whether the organisation can spot and manage the exceptions.

Where an exact answer is needed, the calculation should stay in a controlled, deterministic workflow. AI can then support the work around it, such as interpreting unstructured information, drafting commentary or prioritising items for human review.

Data preparation is part of AI strategy

AI can't make up for poor processes or insufficient data forever. It's tempting to put an LLM on top of an existing process and expect it to sort out the complexity underneath, but in practice, that often only moves the problem into the AI layer.

If a workflow sends large amounts of raw data to an LLM and asks it to work out what matters, much of that processing should have happened earlier. Preparing data before it reaches a model makes an AI workflow cheaper to run and easier to govern. That might mean:

As one theme of the morning put it, an agent shouldn't have to rediscover the organisation's business rules every time it runs.

Right-sized intelligence: matching the model to the task

Not every task needs the same level of intelligence. The aim is “task-matched AI”, where each task gets the model it really needs. The model market is changing fast and firms are looking beyond a single frontier model to smaller models, open-weight models and models built for particular tasks.

Frontier models are charged per input and output token. They're the best choice for:

  • Hard reasoning

  • Large documents and problems that mix text, images and other formats

  • Consistent instruction following

  • Built-in safety and compliance features from the provider

  • Ease of use

Open-weight models suit simpler tasks and offer:

  • Control over data and where it's processed

  • Running costs only, with no per-token charges

  • The option to train and specialise the model

  • A stable version that won't change without warning

  • No vendor lock-in

The trade-off is that they take more effort to set up.

More choice also means more decisions. Firms need to consider which models are approved, where data can be processed, whether a model can be hosted internally and how much capability a task really needs. The takeaway was that model selection belongs in workflow design from the start.

A second demonstration showed this in practice. Alteryx triaged insurance first notification of loss (FNOL) emails and extracted the key details, using a small open-weight model hosted on a cloud server.

fnol processing

The demonstration gave some useful lessons:

  • Good enough is enough: Responses were slower and not as polished as a frontier model's, but they were sufficient for the task.
  • Costs move up front: Spend shifted from ongoing token charges to a one-off engineering effort.
  • Speed fits the job: Each call took 20 to 40 seconds, which is fine for work that doesn't need an instant response.
  • Memory sets the limit: A model needs about 1 GB of memory per billion parameters, and the model used here had 0.8 billion.
  • Design around the model's size: Shape the task to fit a smaller model, instead of paying for a bigger one.

Agentic AI raises the governance stakes

Agentic AI makes this all the more pressing. An agent can call tools, access data and carry out actions on a user's behalf, which makes it possible to automate more complex processes. It also makes visibility and accountability harder.

Attendees saw a growing need for human oversight and clear limits on what agents are allowed to do. The direction of travel is towards agents doing more work, but inside controlled workflows. In a regulated environment, if an AI-driven process produces an outcome, the organisation has to be able to show how it was reached, what data was used and where a person could have stepped in.

VURA: a practical framework for enterprise AI

This is where the VURA principle discussed at the breakfast comes in. A VURA workflow is:
•    Visible: The organisation can see what the workflow is doing and what data is moving through it.
•    Understandable: People can follow the logic, instead of it being hidden inside a model.
•    Repeatable: The same process gives consistent results, regardless of how someone interacts with it.
•    Auditable: There's a record of the process, inputs and outputs, so the organisation can investigate what happened.
Instead of placing all the business logic inside prompts or agents, firms can keep the important logic in governed workflows and let AI work with those trusted processes.

How NextWave and Alteryx can help

NextWave works with banks, insurers and financial institutions to design AI workflows that cost less to run and stand up to scrutiny. Using Alteryx as a business logic layer, we help firms tackle both problems at source:
•    Preparing data before it reaches the model: Filtering, deduplicating and summarising data so the model receives only what it needs.
•    Reducing repeated work: Batching and caching calls so the same information isn't processed repeatedly.
•    Routing tasks to the right model: Sending simpler tasks to lower-cost or open-weight models and keeping larger models for work that needs them.
•    Exact calculations: Handling reconciliations, calculations and rule-based checks in deterministic workflows, with AI used for explanation and interpretation.
•    Building VURA workflows: Moving data preparation and orchestration out of opaque prompts and into governed Alteryx workflows.

Our approach is to deliver benefits step by step, using the right tool for each problem:

  1. Straight workflow: Where pain points sit in email hand-offs, spreadsheet tracking and manual calculations, an LLM helps move the process onto a process platform quickly, avoiding ongoing token costs.

     

  2. Personal productivity: LLMs support analysis, validation and risk identification at scale, with prepared prompts and supporting material to speed up adoption.

     

  3. Deterministic processes built with LLMs: Calculations, comparisons and data validations move onto a data platform, with an LLM speeding up the move.

     

  4. Agentic operations: Agents automate low-risk, low-value steps such as summarising commentary, validation, risk identification and pulling information together, coordinated on an agentic platform.

The result is lower AI costs and a stronger audit trail, which regulators, risk teams and governance functions increasingly expect.

Smart where it matters, lean where it doesn't

The message from the morning wasn't that financial services should use less AI. It was that firms need to be more deliberate about where and how they use it.

Model prices and capabilities will keep changing, but falling token costs won't fix a poorly designed process. The bigger gains come from letting deterministic tools handle the work that needs precision and saving AI for the work where interpretation adds value. In our tests, that meant 20 times fewer tokens and no missed errors.

That means moving the conversation from AI adoption to AI operating models, and asking:

  • Which processes should be automated, and which should stay under human control?

     

  • Where does deterministic logic belong?

     

  • Which model suits each task?

     

  • How much context does the model really need?

     

  • How can the organisation prove what happened after a process has run?

If your organisation is exploring how to reduce AI costs and build more auditable AI workflows, get in touch with NextWave to discuss your transformation goals.

 

Maya Kokerov
Post by Maya Kokerov
October 8, 2026
Maya is NextWave's Marketing Manager, Head of Alteryx Sales, and Editor of The London Fintech Podcast. She is also a published journalist and media expert, with two first-class degrees from Warwick and LSE. She has a decade of experience in marketing across a variety of industries, including tech and fintech.