Lessons from rebuilding the FNOL process on an open-weight, locally hosted model instead of commercial LLM APIs — and why smaller was the right call.
Ask most technology leaders how they arrived at their AI model of choice, and the answer is almost always about capability. Which provider reasons best, scores highest on the latest benchmark, handles the widest range of tasks. Cost and control tend to enter the conversation later, if at all, and are treated as constraints to manage rather than design principles to build from.
Rebuilding the First Notification of Loss (FNOL) process around an open-weight model with well under a billion parameters, running on ordinary server hardware with no graphics card in sight, forced us to invert that order. What we learned has changed how we now think about agentic AI in regulated environments.
The benefits that rarely make the pitch deck
Most conversations about open-weight, locally hosted LLM models start and end with token cost. It is real. Running inference in-house converts a variable, metered bill into a fixed, plannable capacity cost, and, at scale, that difference compounds. But in financial services, it is rarely the deciding factor.
The stronger case is architectural. When a model runs inside your own estate, customer data never crosses a third-party boundary. There is no processor to assess, no cross-border transfer to document, no contractual assurance to chase about whether correspondence might be retained or used to retrain someone else's model. The control becomes structural rather than something you negotiate for in a supplier contract, which is a materially stronger position when you have to answer to a regulator or an internal audit function.
Reproducibility follows the same logic. A hosted model can change beneath you between one quarter and the next, and the same input can produce a different output without warning. A model you version and pin yourself gives you a decision you can still explain a year later, with the same weights that made it.
There is also a quieter benefit: availability. A process with no external dependency in its critical path has no rate limits to hit, no supplier outage to plan around, and no data leaving the network at all. It simply runs where the data already lives.
None of this argues that open-weight models are cheaper in isolation. At realistic volumes, the unit cost of a commercial API is often immaterial. The argument is that cost, control, reproducibility and availability move together, and for a narrow, repeatable, sensitive task, they tend to point in the same direction.
Where the intelligence actually needs to live
A small model only works if you are ruthless about what you ask it to do. In our rebuild, the tasks of the model (Qwen3.5:0.8B) were narrow by design: read an email written by a member of the public, turn it into structured data and draft a response email.
Every consequential decision sits in ordinary, testable code. Whether a policy is valid, for instance, is resolved by five explicit checks run against the data platform, not inferred by the model. Which route a claim takes, whether it can be automatically accepted: all of it is a rule in a place a developer can point to.

This is also what makes the approach portable across agentic frameworks. Because the model's role is deliberately thin, and the workflow is exposed as a callable tool through an open protocol (LiteLLM + Ollama) rather than wired to one provider, swapping the interpreting model — cloud to local, or small to larger — becomes a configuration change rather than a rebuild. Engineering still decides exactly which workflows are exposed and to whom; where the model happens to run does not change that.
Sizing the model, and the hardware, to the job
The cheapest hardware decision is usually an architecture decision made earlier. Once you have been strict about what the model actually needs to do, the infrastructure question mostly answers itself.
As a rough guide, memory, not processor count, is the binding constraint, at roughly one gigabyte per billion parameters, plus headroom for the length of text involved.
We deliberately sat on the bottom rung. Every step up buys capability we had already decided not to ask the model for. And where inference shares a host with other workloads, as ours did, it needs a deliberate cap, otherwise, a batch of claims will happily starve the very platform that triggered them.

What this looked like in practice: rebuilding first notification of loss
First Notification of Loss (FNOL) is the moment a customer first tells an insurer something has gone wrong: high in volume, repetitive, and for the overwhelming majority of cases, light on judgement. An email arrives; someone reads it, decides whether it is a claim at all, extracts the policy number and circumstances, checks the policy was in force, refers it for fraud scoring, opens a case and replies to the customer.

We rebuilt that sequence end to end, pairing an Alteryx workflow — running the checks, the fraud referral logic and the case creation — with a local model doing nothing but the reading and the drafting.

The measured results were modest by design, and that was the point. A complete claim assessment took three model calls and consumed roughly 200 tokens: about 1600 going in, 400 coming back. Each call took around 20 - 40 seconds, so a claim was fully processed in about 75 seconds once the model was warm, stretching to closer to two minutes straight after a restart.

It is not free of trade-offs. 40 seconds is fine for a process measured in hours and would be unworkable for anything interactive. The model follows formatting instructions unreliably enough that we added an automatic correction step to catch it. And the investment shifts from the invoice to the engineering: the safeguards, the deterministic logic and the test suite are where the real cost sits.


None of this makes commercial, cloud-hosted models the wrong choice generally. They remain the better answer where a task requires genuine reasoning, where the work is novel or exploratory, or where volumes are too low to justify the engineering effort a local build demands.
Final thoughts
The interesting finding here was never that a small model is inexpensive to run. It is that model capability turned out to be the wrong variable to optimise. Once every consequential decision moved into code that can be tested, versioned and explained, the model's job shrank to something a modest one could do — and that, in turn, made it possible to keep every piece of customer data inside the estate it started in.
For regulated businesses experimenting with agentic AI, that sequence is worth holding onto: model size is something to design around, not a capability to purchase your way past.
I would be curious whether others weighing open-weight models are finding the same trade-off between capability and control, and where your organisation is drawing that line. Follow along if you would like more on agentic AI architecture in regulated environments.
If your organisation is exploring how to modernise its analytics estate or get more value from its investment in Alteryx, get in touch with NextWave to discuss your transformation goals.
September 30, 2026