AI traffic management: Load balancing vs model routing

Lori MacVittie (00:05.283)
Welcome back to Pop Goes the Stack, where just AI is the new, just reboot it, only with more GPUs and fewer guarantees. I am Lori MacVittie, standing by with questions and a very skeptical threat model. Right? Alright, so today, time to put on your headphones. Pretend we're not just arguing about fancy routers in the sky when we're talking about scaling AI.

Load balancing has always been about variables, parameters, protocols, and algorithms and they have always had to adapt to new types of traffic. Well, guess what? AI is new traffic. Ergo, thusley, and therefore, they must adapt. Okay? So today we're gonna talk about that. Like, what does that even mean? Like, what do we need? How are we gonna do it? What are the new variables, new parameters?

To do that, we've got our co-host Joel Moses, who didn't read the article, 'cause there wasn't one.

Joel Moses (01:05.227)
Ha ha ha. Thank you, Lori, of course.

Lori MacVittie
You're welc-

Joel Moses
By the way, thank you for calling me a threat model earlier.

Lori MacVittie
Ha ha ha.

Joel Moses
I'm just gonna focus on the model part. I don't know what the threat part means, but okay.

Lori MacVittie (01:18.819)
Yeah, alright. Well, and we're joined by Scott Calvet today. Welcome.

Scott Calvet (01:24.286)
Thank you,

Lori MacVittie
He's scared. He should be.

Scott Calvet
thank you, it's good to be here.

Lori MacVittie
Say hi.

Scott Calvet
Yeah, I'm a little scared, I'm worried what kind of model I am.

Lori MacVittie/Joel Moses
Ha ha ha.

Scott Calvet
Probably not one at all, so there you go.

Lori MacVittie (01:37.029)
We don't know. Yeah. Yeah. And that's, and interestingly, that kind of ties in to the whole topic, right? When you talk about load balancing models, it's not just about scale, it's also about picking the right model. Joel, you were kind of leading toward this earlier.

Joel Moses (01:56.209)
Yeah, absolutely. I mean load balancing is kind of giving way to a concept that is being called model routing. They are similar in some respects, but the idea of model routing digs a little deeper. Load balancing is like directing trucks to open loading docks based on the contents of the truck, but assuming the truck carries the same cargo--you're just routing for volume.

Model routing is like opening the back of each truck at the gate and inspecting what's in it and then sending it to the loading dock based on whether it's high sensitivity or whether it need it needs special handling. Right? That's kind of what model routing is doing. And to do that it actually has to dig down into the intent of what's being sent. I've heard this referred to as Layer 8.

I resist that as much as I possibly can. Layer 8 to me is the budget layer, but it is digging into things that are contained within a Layer 7 transaction and making decisions based on that. So it is fundamentally different.

Lori MacVittie (02:57.706)
It is. And it is a layer above.

Scott Calvet (02:58.044)
It's like the yield mana-. Absolutely. I was gonna say it's almost it's becoming like the yield management system of AI, in my opinion. Yes. So and you know we brought up the

Joel Moses (03:11.958)
Yeah, explain that a little bit, yield management.

Lori MacVittie (03:12.324)
Yeah, what i-...

Scott Calvet (03:14.64)
Yeah, absolutely. So, you know, if you think about like an airline, they don't send every passenger to first class. They can't do that. It wouldn't help their business model, right? So you need to make sure that you're steering the right customers toward the right seats. You're also maximizing the revenue you're getting on those seats, but you really need to think about all that to begin with.

So AI infrastructure, right, shouldn't be sending every prompt to the most expensive model. Even if users always want to use those. And I think it's that governance that is often not calculated into what kind of, you know, you think about your cost per token, but you're making this big investment in your AI infrastructure. You want to get the most out of it.

And there's a number of ways that that problem has to be dealt with, and this is one aspect of that. Model routing

Joel Moses
Right.

Scott Calvet
is a very important aspect of that.

Lori MacVittie (04:09.238)
I think it's interesting that you said AI infrastructure because in that simple two words kind of implies the need for something other than what we have. Like today's load balancing algorithms are not appropriate for AI traffic in general. Could they work?

Scott Calvet (04:28.954)
No.

Lori MacVittie
Yes. No? Okay. All right, I love it. I love it.

Scott Calvet (04:34.081)
Right.

Lori MacVittie
Why not? Why will they not work?

Scott Calvet
Yeah, so I think, you know, if you look at it what's interesting is AI is running on a lot of existing technologies. We're definitely updating them. We're modernizing them. But the traffic type is different than anything we've dealt before or dealt with before. It's not like your traditional Kubernetes traffic that may be more staple or whatever. This is very bursty.

It's, you could have lulls, you know, periods of low utilization, and then all of a sudden you could have a ton of requests and prompts come into your AI infrastructure at the same time. And those could be different types of requests as well. So to go back to what Joel was saying, it's like you have those trucks coming in, but maybe, you know, nobody's coming in through the gates, and then all of a sudden you have 18 semis sitting outside of there,

Joel Moses (05:23.064)
Mm-hmm.

Scott Calvet (05:24.931)
and you need to inspect these all and get them to the right place where they need to go as fast as possible. And that's gonna be a big part of the yield from your business. But it's

Joel Moses (05:34.501)
Yeah, I think fun

Scott Calvet
different, yeah.

Joel Moses
Yeah, I think fundamentally you're right. It is different. Now there are some elements about traditional load balancing that are still entirely useful. When you're trying to send AI queries or AI prompts to a system, you wanna do it you wanna send it to the most available system. And so things that check idleness or availability, those still have value.

So there, it's not that you're throwing away traditional load balancing--you still wanna make a judgment call at the gate as to what the best basic path is--but one thing you do want to do is inspect the interior. Traditional load balancers tend to think that all requests are equal. That's just par for the course.

And so a two token prompt and a 100,000 token context document is sent through and it'll treat them like two identical trucks at the booth waiting to be routed. And that's where model routing is fundamentally different than load balancing.

Lori MacVittie (06:33.194)
Well, and it's

Scott Calvet (06:33.474)
Yes.

Lori MacVittie
I think, you know, in the past we've looked at requests in terms of well, what size is it? Because really the processing you needed, you know, the memory was determined by, well, is this a 1K or is this, you know, one hundred byte request? And the problem is with AI, "summarize this document" and, right, "write me x" are very small,

Lori MacVittie (06:59.102)
very similar in size in terms of the actual traffic characteristic,

Joel Moses (07:04.55)
Correct.

Lori MacVittie
but the intent is way different. So maybe you want to route those to two different models, one to actually generate and one to just do a summary.

Joel Moses (07:13.192)
Well and of course there are models that are way more expensive than other models. And they differ in the number of parameters that they're trained on. And you know, it may be inappropriate to use a frontier reasoning model to check if a factory log line says okay.

Lori MacVittie (07:28.072)
Ha ha ha.

Joel Moses
That's like hiring a PhD in astrophysics to flip a light switch every time you come in the room. It says more about you than it does the actual task if you're wasting that kind of money.

Joel Moses (07:40.253)
You don't need a super sophisticated reasoning model to do fairly simple activities. Now, I'm not blaming the people who write these prompts. They're doing what they do, which is they're exploring the bounds of new technology. And the people who are offering these models to people are also unaware of the cost that some of these models incur.

And so we have to kind of move up and look at not just the content, but also the intent that's being passed in the prompt. Because you may have a maybe a 100-token query, and maybe that's better serviced and more cheaply serviced by a lightweight model, and maybe another 100-token query might be better, you know, satisfied by a frontier reasoning model. So you have to look at the intent, not just the length.

Lori MacVittie (08:35.282)
Well and cost, right, Scott, is going to be not only part of right how many tokens did I just ingest, right? It's also the complexity of the request could generate a lot of tokens, which is going to cost more. So there's more to it than just how many tokens are in this prompt. Right?

Scott Calvet (08:55.337)
Oh yeah, absolutely. And not every token is equal. And I think that's a tough concept to understand because we want to think about tokens as just a unit of measurement. Right? And I know a lot of people talk about this, but you know, it gets into what is the cost of generating that token. And even cheap tokens can become very expensive when you're paying for tokens that don't accomplish anything, which is kind of what you're talking about. And that can happen as well.

And I would say the expensive token inversely is not always, it's not necessarily the one from the most expensive model. It's also the token you paid to generate that nobody used, right? Or nobody could use. So I think that's another aspect that you have to think about it. And if cost of tokens, you know, or cost per token is the price of production, then you know, it isn't necessarily the economics of the business.

You need to start looking at like throughput, which is measuring how many tokens you can get from your AI infrastructure. And I think that's like stage one as everybody's getting into deploying their systems. But as we mature, we're all going to have to start looking at good-put, which is essentially what kind of outcomes I'm getting into. I know Gartner is starting to talk about like cost per, excuse me, cost per successful outcome as a way to think about measuring the output

Joel Moses (10:19.97)
Yep.

Scott Calvet
you're getting from your AI systems as well. So and I think, you know, what you guys were talking about with intent and going back to Joel's analogy with the truck, it's like, okay, what's the intent of that too? But you also need to think about it intelligently from the other end.

So if you have a warehouse with a bunch of loading bays, if you know what's in those trucks, right, maybe you need to send some trucks to the loading bay that has the forklift because it's got a bunch of pallets in there, and you have another truck with a bunch of boxes, and you need to send that to the loading bay with people who can come in and lift those off manually.

So you have to think about what is the capability of the GPU where you're sending it as well. And what, is it under stress? Have those guys maybe haven't gotten a break in 10 hours and they're hungry, they're not gonna work as fast as the guys who just came back from lunch.

Scott Calvet (11:16.825)
So that's kind of, you know, at F5 I know we have products to do this and not to get into that, but that's how people are starting to think about this. And I look at it as like I like to use the word like "AI service plane". It's that place, if you think about what Jensen talks about, he describes AI factories or this AI infrastructure as a five layer cake. So at the top, the top layer you have your users, your apps, your agents, etc.

That goes down and then you have your models, right, you have your infrastructure. And then he's got chips and power at the bottom. But let's think about those two layers that sit right underneath the users who are creating this traffic. They're the trucks. Right? There needs to be a service plane, something.

If you're gonna maximize the value, getting back to what we were talking about with cost per token, there needs to be a service plane sitting in between that that's governing those interactions and thinking about everything, the intent to what is the capability on the back end.

So to give like another example, like we were talking about the bursty nature of AI traffic, you have a lot of requests coming in, maybe all at once, all of a sudden. A lot of them are like, hey, rewrite this email, help me create a funny joke to send to my coworker. That's nothing. But if somebody's asking for a video or an image to be made, you want that to go to the not only the right model but the appropriate GPU for the process.

And we all know that GPUs have different generations. And that gets back to my point about the loading bays. Yeah, so there's a lot of different aspects I think to think about this. Yeah.

Lori MacVittie (12:55.996)
You are, but I mean you can call it a service plane, you can call it model routing, but what we're talking about is traffic management

Scott Calvet
Mm-hmm.

Lori MacVittie
at various layers. It's always been

Scott Calvet (13:05.602)
Yes.

Lori MacVittie
traffic management, it will always be traffic management. The difference here, I think, is that we have to get into the payload in a new way. This isn't matching words or matching, you know, bits in an URI using regex--thank God, 'cause I hate regex--but

Lori MacVittie (13:24.722)
right, it's understanding more than just the we're not looking for keywords. We're looking for meaning. And that is something that is, it doesn't exist. You can't put that in an algorithm. You can't say, "Oh well, you know, if this word or that word," I mean, eventually the languages are too big to be able to hold that kind of a, you know, an index, if you will.

You run into the, you know, overflowing IP tables problem. Right? Where you're just like,

Joel Moses (13:51.214)
Ha ha ha.

Lori MacVittie
I can't anymore, I give up. So there has to be other ways to do it. Joel, like how are we gonna do this?

Joel Moses (13:57.543)
Yeah.

Lori MacVittie
how are we gonna do this?

Joel Moses
Well, I mean, you know, as we were all talking together, I was sort of, you know, spinning over in my own mind what the impact of agentic AI actually has on this problem set as well. You know, it's very easy to judge, you know, when you've got a chat bot scenario, you've got a user who's sending through a prompt and you can kind of assess the prompt and then you can continue to the next one.

But when you stand up an agent, an agent is essentially going to create not just prompts, but it's going to create a session. And it's going to create a session towards multiple tools and it's going to create multiple tasks and each one of those tasks is a new place where model routing becomes uniquely relevant.

Now I tend to think about the use of agentic AI in software development and I can tell you it's very easy for an entire month's engineering budget to be spent in a single debug loop just by feeding ten line typo fixes to massive frontier models. You know, to using a one hundred million dollar model to fix a misplaced semicolon is a complete misuse of the tool set.

But an agent doesn't select necessarily the model today. And maybe either the agent needs to select the model or something in the middle needs to route the model request appropriately to something that is more attuned to doing things like typo fixes and less generation of sophisticated code libraries.

Lori MacVittie (15:22.729)
Now I want to develop a solution that just has like snarky answers and never uses a model. Like "write your own code," you know,

Joel Moses (15:30.101)
Yeah.

Lori MacVittie
"fix your own words."

Scott Calvet (15:32.944)
Ha ha.

Joel Moses
Well, I mean, look, I look at it this way, without intelligent model routing, for the purposes of software development, it's like giving an autonomous coding agent...it's like handing a teenager an uncapped corporate credit card at an open bar. You know? Just,

Lori MacVittie (15:46.483)
In Wisconsin.

Joel Moses (15:46.551)
you're supplying it with something that it can greatly misuse unless there are controls related to that and routing not just at the prompt level but also at the session and the task level. And I think that that's kind of the next order of business for model routing.

Lori MacVittie (16:05.726)
Yeah. And to Scott, to your point,

Scott Calvet (16:06.087)
Absolutely.

Lori MacVittie
the understanding the variables on the back end, they are different. Like GPUs, like what's the utilization, what's the Q depth, right? Where's the, what about the KV cache? Like all of these different

Scott Calvet (16:20.791)
What's the temperature?

Lori MacVittie
variables.

Scott Calvet
Yeah.

Lori MacVittie (16:23.598)
Yeah, okay. Yeah, temperature, all of these have to be taken into consideration. And that doesn't necessarily exist seamlessly right now. There's no standard that says, here's the list, here's what we're gonna give you, and here's how to get it.

So everybody's still trying to figure out how do we get the data that we need to make intelligent decisions, right, before we even start worrying about what's the meaning in this here. So it's not very advanced at the moment, is it?

Scott Calvet (16:52.717)
No. No, it's getting more advanced, but traditionally it was like typical round robin where it's just like sending it

Lori MacVittie
Agh!

Scott Calvet
down the line. Here, you first,

Joel Moses (17:00.518)
Mm.

Scott Calvet
you second, you third. Right? And it's not taking any of that into account, but it's so important. I mean, you can imagine. I saw

Lori MacVittie
That hurt.

Scott Calvet
I saw a report from Gartner. They were talking about most enterprise AI deployments only get like 25% of their total efficiency rate allotted to the infrastructure. What actually,

Scott Calvet (17:21.591)
the actual output is 75% less than what it needs to be. Now I thought that was an

Lori MacVittie
Wow.

Scott Calvet
extreme number, and I, you know, I kind of bantered with them and argued about that, but even if it's half, that's a lot of waste. And that just comes down to not thinking about things intelligently. And you know, I've been in IT a long time and I think about what I used to pay for a server 20 years ago versus what some of these AI servers cost.

I mean, it's insane. It's multiple, multiple times over. And I can tell you, you know, from having to go to my CFO back in the day and justify the cost of these server purchases. He gave me a hard time then, I can't imagine what he would do today. Especially if you find out I'm only getting like 25 to 50% usable output out of it. That wouldn't fly. You gotta do something about that.

And just throwing more hardware at the problem, which is what I think how a lot of people are thinking about dealing with it,

Lori MacVittie (18:18.945)
Oh, yeah.

Scott Calvet
is not the solution. It's, software plays a major role, which is what we're talking about. Yes.

Joel Moses (18:24.99)
Yeah. Yeah, you know, my takeaway here is, well, first of all, you said 25% and the old Java guy in me was like, "Hey, that sounds pretty good."

Lori MacVittie
Ha ha ha.

Joel Moses
But you know, it's yeah, we have a need to get to better efficiencies across all of these this AI infrastructure, not just because it lowers costs, but also it's just better for the earth if we're able to handle these things much more efficiently. I look at it this way, load balancing is not enough for AI.

You also have to combine it with advanced load balancing techniques that sample load and model routing. Think about load balancing as spreading the weight, while model routing helps you pick the right muscles to move that weight around. And that, the combination of the two, that's what spells success in AI.

Lori MacVittie (19:16.609)
Scott, why don't you, yeah, I hope you

Scott Calvet
100%.

Lori MacVittie
have a takeaway 'cause I'm still triggered by the whole round robin mention. I can't. So go ahead.

Scott Calvet
Ha ha ha.

Lori MacVittie
It hurts.

Scott Calvet (19:28.963)
No, I mean going back to what we were saying with agents, like agents are token multipliers. So it's, I mean, think about what round robin would do with agents generating, you know, millions and

Lori MacVittie
Ha ha, ahhh.

Scott Calvet
millions of tokens, right, in a single instance. But you know, it's something people have to think about and deal with. And like I said, like we were talking about GPU temperature, like are people really thinking about all these metrics that play a role into where you're sending this traffic to get processed. And what's gonna do that at agentic speed and with all this payload coming in? I don't think anybody is.

Joel Moses
Yeah.

Scott Calvet
And round robin's

Lori MacVittie (20:13.368)
Yeah.

Scott Calvet
definitely not gonna do it for you.

Lori MacVittie
Yeah. No.

Scott Calvet (20:16.743)
But it gets into, you know, how are you setting up your GPU infrastructure as well? I think, you know, is do people have enough, you know, "Hey, I got, you know, a hundred rSeries sitting over here." And I mean, I think it depends on the customer and the type of business, but a lot of people are just pulling together different generations of GPUs.

And I think that plays a role as well. And you know, that's something they have to think about.

Joel Moses
Yeah.

Scott Calvet
Again, going into I want to send the truck that has all the pallets over to the right bay with the forklift. I want the Blackwells or the Vera Rubens to be handling this gigantic agentic request that just came in from my engineering team versus rewriting the CEO's email.

As special as he is, that can go over to an rSeries and be dealt with just fine on maybe not a frontier model, right? So you combine those two things, and this is where you start to see that governance make a big dent in your cost per token.

Joel Moses (21:16.406)
And speaking of round robin, it's been to me, it's been to Scott, and now to Lori

Scott Calvet (21:21.958)
Ha ha ha.

Joel Moses
for your conclusion.

Lori MacVittie (21:23.53)
I'm, you're gonna make me do it. I, you know, it goes without saying don't use round robin for anything, ever. Never. Just, just don't. I think one of the takeaways actually came from something you said earlier, Joel, about success.

Lori MacVittie (21:39.616)
Like the first thing is how are you measuring success for your AI here? Is it how fast it responds, how accurately it responds, how cheap it is? You know, it's utilization? Like what are your measures of success? Because without understanding what number you're trying to hit, you're never gonna be able to tune all of the different variables and do the routing that you want. Right? It, you can't.

So, you know, start figuring out what is success when you're doing your AI deployments. And then you can start working backwards and figuring out, well, how are we going to actually get there using the technology that we have?

So that is my takeaway and that is now a wrap for this week's Pop Goes the Stack. Please hit subscribe before your roadmap becomes a prompt and your architecture becomes just a suggestion.

Creators and Guests

Joel Moses
Host
Joel Moses
Distinguished Engineer and VP, Strategic Engineer at F5, Joel has over 30 years of industry experience in cybersecurity and networking fields. He holds several US patents related to encryption technique.
Lori MacVittie
Host
Lori MacVittie
Distinguished Engineer and Chief Evangelist at F5, Lori has more than 25 years of industry experience spanning application development, IT architecture, and network and systems' operation. She co-authored the CADD profile for ANSI NCITS 320-1998 and is a prolific author with books spanning security, cloud, and enterprise architecture.
Scott Calver
Guest
Scott Calver
Director, Product Marketing at F5
Tabitha R.R. Powell
Producer
Tabitha R.R. Powell
Technical Thought Leadership Evangelist producing content that makes complex ideas clear and engaging.
AI traffic management: Load balancing vs model routing
Broadcast by