LLMs Are Overkill for Up to 70% of Your AI Agent's Calls. Enter SLMs

LLMs Are Overkill for Up to 70% of Your AI Agent's Calls. Enter SLMs

Zara Elowen

Picture your AI agent on an ordinary Tuesday. A customer email arrives. The agent decides which team should handle it, pulls an invoice number out of the text, calls a billing tool and formats the result as JSON. Every one of those steps probably went to the same frontier model, the one with hundreds of billions of parameters and a per-token price to match.

Now for the uncomfortable part. Researchers at NVIDIA studied three open-source agents and estimated that 40% to 70% of their LLM calls could be handled by a specialized Small Language Model (SLM) instead. That means you may be paying premium rates for work a much smaller, cheaper model can do just as well. At Jalsonic Networks, we see this pattern in almost every agent project we review, and it is one of the fastest ways to cut AI spend without hurting quality.

The Ferrari and the Grocery Run: What Actually Separates LLMs from SLMs

A large language model runs on hundreds of billions of parameters and needs GPU clusters to serve. It can write a sonnet, debug a distributed system and summarize a legal contract in the same afternoon. That range is exactly what you pay for, and exactly what you don't need when your agent is deciding whether a ticket belongs to sales or support. Using a frontier model for that is like taking a Ferrari to the grocery store: it works, but you are burning money and fuel for a trip a hatchback would handle.

A small language model sits below roughly 10 billion parameters. Many are distilled from larger models and quantized to 4-bit precision, which shrinks their memory footprint enough to run on a laptop, a phone or even a Raspberry Pi. There is no cluster to provision and no network round trip to a third-party API. For teams in regulated industries, that also means sensitive data can stay on your own hardware.

The pricing gap is where this gets hard to ignore. Here is how the numbers look per million tokens:

Model Input cost Output cost Where it runs
GPT-6 Astra (OpenAI's top model) $10 $50 OpenAI's cloud
GPT-6 Luna (OpenAI's cheapest) $0.10 $0.50 OpenAI's cloud
Gemma 4 E2B No per-token bill About 2GB of RAM Your own machine

Look at the first two rows. Inside a single vendor's lineup, the cheapest model costs one-hundredth of the flagship on both input and output. You haven't even left OpenAI yet, and you have already found a 100x difference. Once you bring open models into the picture, the per-token bill can disappear entirely, replaced by the fixed cost of hardware you control.

For an agent making thousands or millions of calls a day, that difference shows up quickly in the monthly invoice. A support triage agent that fires off several calls per ticket can turn a manageable pilot budget into a serious line item at production volume. Routing the repetitive calls to a small model is often the single biggest lever you have.

Where Small Models Win, and Where Frontier Models Still Rule

The honest answer is that SLMs are not a replacement for LLMs. They are a better fit for a specific slice of your workload, and the skill lies in knowing which slice.

Small models shine on repetitive, well-defined jobs. Routing a request to the right workflow, extracting fields from text, calling a tool with the right arguments and returning clean JSON are all tasks where the input and output follow a predictable shape. For that kind of work, you can expect serving costs 10 to 30 times lower than a 70B+ model. The task doesn't need a model that has read the entire internet. It needs one that does a narrow job reliably, thousands of times in a row.

Frontier models keep their crown on open-ended work. Multi-step reasoning across unfamiliar problems, creative writing and deep research still benefit from the scale and breadth of a large model. A Carnegie Mellon study found that small models struggle most when every request looks different, because they have less general knowledge to fall back on when the input drifts away from what they know.

Here is a simple way to sort your agent's calls:

Type of call Best fit Why
Intent routing and classification SLM Fixed set of outcomes, high volume
Data extraction from documents or emails SLM Predictable structure, easy to verify
Tool calls and structured JSON output SLM Narrow format, low ambiguity
Open-ended research and analysis Frontier LLM Needs broad knowledge and flexible reasoning
Creative or long-form writing Frontier LLM Quality depends on scale and nuance
Unusual or messy customer requests Frontier LLM Every request looks different

The winning architecture is rarely "SLM only" or "LLM only." It is a hybrid, where a lightweight model handles the predictable bulk of traffic and hands the strange, difficult cases up to a frontier model. Gartner points the same way, expecting companies to use small task-specific models three times more than general-purpose LLMs by 2027. The teams that build that routing layer now will spend less and move faster than the ones that wait.

The Harness Matters More Than the Model

If you take one lesson from the research, make it this one. The most striking result from the CMU study (July 2026) came from a budget-approval agent, a very ordinary business workflow, tested on both a frontier model and a small one:

Metric Gemini 3.1 Pro Gemma 4 26B-A4B (~4B active)
Accuracy 97.3% 98.3%
Cost per query $0.22 $0.017
Response time 70 seconds 33 seconds

The small model was more accurate, about thirteen times cheaper per query and more than twice as fast. That is not a rounding error in favor of the big model. It is the reverse of what most teams assume.

Here is the detail that matters most for anyone planning a rollout. Out of the box, the small model scored only 75%. The jump to 98.3% did not come from a bigger model or a long fine-tuning project. It came from the harness, the scaffolding around the model. The researchers gave it a step-by-step workflow instead of one giant instruction, cut down the number of tools it could choose from, and added a hook that stopped it from looping. In other words, they made the job easier to do correctly.

This is good news, because harness work is engineering, not research. It is something a capable software team can do in weeks. It also explains why so many teams write off small models after a bad first test. They judged the model in a messy setup and blamed the model.

Here is a practical path to try this on your own agent:

Step Action What you get
1 Log your agent's LLM calls and group the repetitive ones A clear picture of where the volume sits
2 Calculate what those calls cost on a frontier model A baseline to measure savings against
3 Test Gemma 4 E4B, Phi-4-mini or Qwen 3.5 4B, all free in Ollama Fast, low-risk comparison on your real data
4 Fix the harness first, then fine-tune on 10k to 100k logged examples later Accuracy gains before you spend on training
5 Keep a frontier model for the messy requests A safety net for unusual inputs

Notice the order in step 4. Fine-tuning gets most of the attention, but it needs a healthy pile of logged examples and it costs time and money. Harness fixes cost neither. Start there, see how far you get, and let your logs tell you whether fine-tuning is worth it.

Key Takeaway

Most agent calls are routine, and routine work doesn't need a frontier model. NVIDIA estimates 40% to 70% of calls in the agents it studied could move to a small model. OpenAI's own pricing shows a 100x gap between its top and cheapest models, and a CMU study showed a small model beating a frontier model on a real business task at a fraction of the cost. Log your calls, test a free small model, fix the harness before you fine-tune, and keep a frontier model on standby for the hard cases.

Ready to Stop Overpaying for Your AI Agent?

Knowing that SLMs can do the job is the easy part. Working out which of your calls to move, how to build a harness that keeps a small model on track and where to draw the line with a frontier model takes hands-on experience. That is what our team at Jalsonic Networks does. We audit agent workflows, identify the calls that are costing you the most, and build hybrid architectures that cut spend while keeping accuracy where it needs to be.

Book a consultation with our team and we will walk through your agent's call patterns, estimate your potential savings and map out a low-risk pilot. Bring your logs if you have them, and if you don't, we will help you start collecting them.

One question to leave you with: which task are you still overpaying a frontier model to handle?

Frequently Asked Questions

What is a Small Language Model (SLM)? An SLM is a language model with roughly 10 billion parameters or fewer. Many are distilled from larger models and compressed to 4-bit precision, so they can run on ordinary hardware such as laptops, phones and small edge devices.

Will an SLM be as accurate as an LLM for my agent? For narrow, repetitive tasks such as routing, extraction and tool calling, it can be. In the CMU budget-approval test, a small model reached 98.3% accuracy against 97.3% for a frontier model. For open-ended reasoning or highly varied requests, frontier models still perform better.

How much can I actually save? It depends on your workload. SLMs can cost 10 to 30 times less to serve than a 70B+ model on suitable tasks, and OpenAI's own top and cheapest models differ by 100x in price. NVIDIA estimates 40% to 70% of calls in the agents it studied could be handed to a small model.

Do I need to fine-tune a small model? Not at first. The CMU study showed that a better harness, meaning a step-by-step workflow, fewer tools and a loop-stopping hook, lifted a small model from 75% to over 98%. Fine-tuning on 10k to 100k logged examples makes sense later, once you have data and a clear gap to close.

Can I run SLMs on my own infrastructure? Yes. Models like Gemma 4 E4B, Phi-4-mini and Qwen 3.5 4B are free to run through Ollama, and Gemma 4 E2B needs only about 2GB of RAM. Running locally can also help with data privacy and compliance.

Should I replace my frontier model entirely? No. A hybrid setup works best: small models handle the predictable bulk of calls, and a frontier model handles the messy, unusual or high-stakes requests.

How does Jalsonic Networks help? We review your agent's call logs, identify the best candidates for SLMs, design the harness and routing layer, and help you test the results before you commit. You can start by booking a consultation with our team.