Full Tech Ahead
On this podcast, I sit down with business leaders, researchers and executives to explore innovative technology solutions and products, whether they’re transforming industries today or still in development. But we go far beyond the tech itself. From real-world use cases and business implementation journeys to cybersecurity challenges and future trends, we uncover what’s shaping the digital landscape.
We also dive into topics that matter to every tech professional: Work/life balance, business communication, education and training. Think of it as your one-stop shop for meaningful technology discussions that inspire and inform.
Full Tech Ahead
Control Your AI Spending
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
In this episode of "Full Tech Ahead," host Amanda Razani interviews Matthew Shaxted, CEO of Parallel Works. The conversation centers on a major obstacle facing enterprises today: skyrocketing token consumption and the ballooning costs of using frontier AI models. Shaxted explains that opening up unrestricted API access to hundreds or thousands of users leads to rapid budget depletion, citing recent industry examples like Uber. Drawing a parallel to the high-performance computing (HPC) and cloud migration trends over the past decade, Shaxted predicts a cyclical shift: while firms currently rely heavily on public cloud endpoints, economic pressures and massive utilization rates will drive them to bring data workloads back on-premise using increasingly powerful open-weight models (like the recently released GLM 5.2). To combat initial adoption chaos, Parallel Works offers a computing control plane called Activate, providing a single pane of glass to enforce visibility, tracking, and strict "token budgets" that automatically deny requests once expenditure thresholds are met.
Key Quotes
- "Unless [token usage] is thought about in the very beginning in terms of how are you going to control and monitor token usage... it really becomes a big problem. We've seen the Uber story recently, where you burn through the entire budget in a few months."
- "Having strong visibility in where the tokens are going... make sure from the very beginning you have visibility into who's doing what, because that's going to start growing very quickly."
- "As soon as it basically is out of budget, it will deny the request until you get more allotted. That's exactly what's happening."
- "As the open weights get better and better—which I think we're starting to see with GLM 5.2 coming out recently—you can start running open weight models for certain classes of things with much more predictable cost."
Takeaways
- Establish Financial Guardrails Early: Unrestricted enterprise AI access creates a cash burn. Implement central governance and a "computing control plane" from day one to enforce hard spending limits (token budgets) at the user or team level, preventing unexpected tech invoicing.
- Prepare for the On-Premise AI Hybrid Shift: Much like cloud computing evolved, enterprise AI will hit a baseline utilization rate (e.g., 8k/80% load) where renting per-token public APIs becomes economically unsustainable. Companies should plan a hybrid stack that shifts routine tasks to dedicated on-premise infrastructure running open-weight models to cut costs up to 6X.
- Incentivize Token Awareness: End-users rarely optimize resource usage unless faced with explicit constraints. Providing visible token limits encourages developers and practitioners to delegate simpler, mundane prompts to lighter, less expensive internal models rather than burning resources on premium frontier labs.
- Architect for Agentic Workloads: With the industry shifting toward massive agentic systems where hundreds of autonomous agents execute tasks 24/7, compute demands are projected to scale up to 1,000X. Managing this scale without breaking corporate cost structures requires unified virtualization and gateway gateways across cloud and hardware assets.
Find Amanda Razani on LinkedIn. https://www.linkedin.com/in/amanda-razani-990a7233/
Follow the FTA LinkedIn Page: https://www.linkedin.com/company/full-tech-ahead/
Visit the FTA website: https://fulltechahead.com/
Check out the Substack Channel: https://fulltechahead.substack.com/
Hello and welcome to Full Tech Ahead. I am your host, Amanda Razzani, and with me today, I'm so excited to have Matthew Shakstead. He is the CEO of Parallel Works. How are you doing today?
SPEAKER_01Hey, Amanda, doing great. Thanks for having me.
SPEAKER_00Thank you for coming on the show. Can you share a little bit about Parallel Works? What services do you provide?
SPEAKER_01Sure. Yeah. We are a software company that makes a product called Activate that I call a computing control plane. So it's a piece of software that organizations deploy, usually in their own cloud or on-prem boundaries. They connect all their various computing resources together into it. And it becomes the single pane of glass that the end users access the computing through. And we do a bunch of more admin and operational capabilities as well for IT organizations and such. So can talk about that.
SPEAKER_00Great. So adding that uh visibility all in one place.
SPEAKER_01Correct. Yes, that is definitely a big aspect of it.
SPEAKER_00Well, that brings us in our topic for the day, which is the incorporation of AI. Of course, it's been swift and furious. And one of the problems business leaders face is how many tokens are being used with this AI use. And we know that the cost of AI use is increasing. The amount of tokens needed for the AI is increasing. And this is presenting a little bit of a challenge to companies. So to start off, from your experience, where is the biggest problem that business leaders face right now?
SPEAKER_01Yeah, great question. Uh, and our company is kind of right in the middle of this, as many of the organizations that we work with start bringing in frontier models from public providers, or they're building on-prem private infrastructure and deploying models there. There's different sets of challenges in both of those. But I'd say the biggest challenge, and I think everybody's probably seen it in the news, is you know, you open up frontier models and API keys to end users within an organization, you know, tens, hundreds, even thousands of people now starting to use these enterprise approved boundaries, you know, for their models, and the costs start skyrocketing. And unless that is kind of thought about in the very beginning, in terms of how are you going to control and monitor token usage and capability, it it really becomes a big problem. Like we've seen the Uber story recently, where you burn through the entire budget in a few months that was set because uh a lot of these things start as somewhat unrestricted. And and you hit it right. The the cost of these frontier models when you're going to public providers is expensive and getting more and more expensive kind of by the day. So that's kind of the world that I find myself in now is one of governance around these AI models and using them effectively and giving the operators of the models and the administrators a way to kind of monitor usage for one, but actually go towards cost control and basically allocation, which I'll kind of talk about as well in the in the worlds that we came from previously. Yeah.
SPEAKER_00Yes, absolutely. And there are so many AI tools being used, it's hard for departments to realize, okay, who's using what tools at what level, what's really needed, what isn't, what's maximizing the best end result, what really isn't efficient and not giving a ROI. So what advice do you have for business leaders to tackle this?
SPEAKER_01Well, I I'd say I'm I'm seeing it kind of as a trend that I've seen take place from kind of on-prem high performance computing resources where you know you bring these systems in to an internal boundary and you know, a data center that you manage, maybe, and then you open it up for pioneer pioneer usage where hey, everybody come in and use this thing, and then eventually you need to start uh allocating out time for units out, cost tokens, you know, kind of similar units, I'd say, uh, because the resources start to get constrained within some set budget. And you know, I think what you know, maybe 10 years ago now, there's a big movement into cloud resources, at least in the the world I'm from, which is high performance computing resources in the cloud, and started moving over with kind of constrained budgets as well. And you have, you know, it was a similar challenge. I'd say, how do we keep control of our ballooning cloud costs? Which starts with tracking of what's actually being used across resources, tagging policies, and then the costs start reaching a point where economics would actually let them bring certain aspects of their operations back on-prem. And you go back to these cycles of hey, we're actually going to purchase a data center and the whole operational team to get this thing running, uh, because we can run it at a base load of 90%. And I think that probably is a similar trend we're going to see in these AI models where organizations can get started very quickly running in these frontier provider models, you know, the public endpoints, if you will. And you can put cost around them, and then you'll reach a point where, hey, we're at hundreds of thousands of people leveraging these things now, using it, you know, 90% baseload or whatever that that metric is that makes sense for a given organization. And hey, let's bring it on-prem. And we're gonna, you've seen the open weight models get better and better month over month, and that trend will continue, and there will be certain sets of things running, you know, on dedicated infrastructure on-prem, which change you know the economic story a little bit. You still want to be able to have control and you know usage tracking on there, but you know, it's kind of a similar transition I thought that we I've seen with like cloud HPC in the last you know 10 years or so. So uh so back to your point though, what what to keep in mind? I'd say from the very beginning, having strong visibility in where the tokens are going. You know, you connect in a cloud API endpoint or you know, bedrock or pick your provider, make sure from the very beginning you have visibility in who's doing what, because that's going to start growing very quickly. If you can put guardrails around these things, uh either from the provider themselves or from like a control plane, what my company does, so you can actually plug in all the different models you want and put the same type of cost control guardrails around them, and then start watching the behavior of it and you know, have in mind that hey, maybe in a year it does make sense to actually bring some dedicated resources on-prem if you're not already doing that. So and kind of manage it that way. So that's my my tip kind of have visibility and guardrails set up from the very beginning. Otherwise, you get into kind of this untethered access, which as we've seen, makes really good news stories lately.
SPEAKER_00So yeah, absolutely. It's interesting how quickly we saw the shift to cloud, you know, migrate everything to cloud, and now how it looks like okay, actually, let's bring some of this back on premises. Uh, there's this maybe there's a balance that's needed, not everything needs to be out going to cloud.
SPEAKER_01Yeah, I think there's a balance for sure. And I'm, you know, my software enables organizations to create hybrid computing environments, so a mix of on-prem and cloud. So I'm I have a have a vested interest in that positioning, but I really do believe that there's a place for both. And as an organization kind of matures on their computing journey, just economically and kind of baseload, maybe security-wise, sovereignty-wise, there's reasons to put infrastructure inside of your own boundary, you know, on-prem and operate it that way with all the costs assumed. And then there's really great places for cloud where, you know, they're typically getting the latest generation technology on the floor faster than anybody, faster than most organizations can get it themselves. You're able to really leverage that uh effectively. And you know, it kind of becomes a seamless transition to kind of move that technology from the cloud back to on-prem when you can, and then, you know, it's a continuous cycle that way. So that's how I I've been seeing it that way. There's places for both for sure.
SPEAKER_00Right. And when you were talking about putting in these expenditure guardrails, is are you saying that there's a way that uh company leaders could, you know, say somebody's using these different AI tools, and once they have basically gone to the threshold of the allotted budget, they just suddenly wouldn't be able to use the tools or move forward without getting some sort of approval?
SPEAKER_01Yep, that's exactly right. You summed it up. So a system like ours, and there's others out there too, I'll say. But uh we we're different in that we bring a lot of other infrastructure types into the same common allocation and kind of cost guardrail framework beyond just LLMs, I'd say. So we're we're touching Kubernetes clusters, you know, virtualized environments, batch schedulers, cloud resources. And we rolled out basically an AI LLM gateway because our customers are asking us, hey, how do we start actually cost controlling those as well? So, what it allows you to do is basically from the admin perspective, plug in the different models you want your teams or your end users to use. So you may just you may have just purchased like an OpenShift cluster and you're serving up your own open weights or you know, private weights that have been fine-tuned, and you want to make those available to your users. You want to bring in bedrock or Azure Foundry resources, you want to bring in your enterprise cloud account, whatever it is. You can plug those all into our platform from an admin perspective, and then start providing access to specific groups of who you want to be able to access to particular models, or users can even serve them up. We deliver all those models through a single OpenAI compatible endpoint. So you plug that into your code assist tool or your chat or wherever you want it, it lists all the models you have access to and the given token budget that you've been allotted. And so users are using their models, and as soon as it basically is out of budget, it will deny the request until you get more allotted. So that's exactly what's happening. Yep.
SPEAKER_00So that brings me to my next question is I think that the end result of this would obviously be some really good company insight and realizing what tools are needed and at what level. But in the beginning, is there maybe a little bit of chaos as people that are used to using certain tools suddenly are stopped in the middle of a project because they can't move further. Um, and how do you advise business leaders deal with that process? You know, that what I think would be a first initial chaos before, you know, it irons out.
SPEAKER_01Yeah, there's even a chaos before that, too, at least that I've been seeing where a lot of end users, like the practitioners, people doing work, uh, maybe start with even like a personal account, you know, and they're they're on their own, you know, co-pilot account or whatever, and they're using that. And then suddenly their organization says, Hey, you need to use this enterprise mandated, you know, set of tools. And hey, you're moving to Cloud or Pick Your Provider, OpenAI, Enterprise, et cetera. So that's like the first level of transition. And then, hey, you're being allocated the equivalent of $10,000 a quarter or $10,000 a month or whatever for your work, use it sparingly. Or, and that, and that's where I think end users do need to start actually thinking about how are they using their models and actually maybe having process these that can use less expensive models for certain things, or if you're bringing this all into a hybrid infrastructure, hey, we're gonna send certain sets of tasks to our on-prem dedicated infrastructure where the token rates, what I'm being charged, are you know maybe six X less than some of the public providers. Move these tasks over there. So you can start to be aware of that. I've seen usually users not unless they have a economic incentive to do so. So, oh, they see what they're actually spending, they know they have this work to be done in the next month. So I got I need to make that last by maybe using a, you know, our private internal model, for example. Yeah, I guess that's my my comments there. There is definitely chaos, but I think it just brings awareness to the actual end users, how they're actually interacting with the models, I would say more so, which is similar to what cloud, I think, was it's like here, you know, you can spin up whatever you want in the cloud, and then suddenly, oh, we have a $50 million budget a year. Uh something needs to change.
SPEAKER_00Right.
unknownYeah.
SPEAKER_00Do you think though that there's gonna be, because I'm hearing a lot about the increased token costs and the amount needed for these tools, that some of them were initially free, everybody started using them heavily. Now suddenly they're costing money. Um, other ones are costing more. Do you think there's gonna be some pushback, especially when companies do put in these regulations and start cutting the fat, if you will? Do you think that maybe we'll see prices go back down as they say, oh shoot, nobody wants to use these tools anymore because how expensive they are, they're cutting back. And we'll see those prices fall back down. Because I think right now it's sort of a money grab. Oh, they need these tools, we're gonna make them very expensive now.
SPEAKER_01Yeah, or or actually, I mean, subscription models before you reach these enterprise plans for the frontier lab providers and such that are being subsidized, essentially, right? So you're getting if you were to run the same processes with like an API key and pay like the per token input output rates, you're getting way, way more value until you transition to these like enterprise API keys. And then suddenly it's like, oh, we're spending, yeah. You've seen it in the news, you know, we just blew our entire budget. I think that will definitely be a drive, you know. I I want to say it will be a strong driver for organizations to seriously think about this transition from moving from just frontier models using the public providers to on-prem infrastructure, where you have a much more predictable cost control, like you saw with you know on-prem HPC and everything else. And as the open weights get better and better, which I think we're starting to see, but you know, GLM 5.2 coming out recently. That's that was like big news where it's really catching up in terms of capability. You can start running open weight models for certain classes of things with much more predictable cost. And then that's going to be a direct driver to the frontier models because, like, why is someone gonna run that if I can run similar things in on-preme infrastructure that I control three month or for three-year cadence, right? Refresh cadence, uh, or you go rent it from a neo cloud provider or a hyperscaler, but you rent the infrastructure and host it yourself. But at least you have very clear cost control. You know, you know what, you know what the max consumption is gonna be. And yeah, you still need to control the token usage because there's contention. But yeah, so I think it's a it's definitely gonna be a big challenge as they want to increase those rates to recap all recoup all the massive investments they did, but you have you know open weight models coming in with you know on-prem infrastructure that could stand up at least some portion of what these things are doing, especially for those like base loads, where hey, we can run this thing, you know, 80% utilization reliably, and this model's good enough for what it needs to do. That will get that will get challenged. And I think you you kind of mentioned it the you know, you're hearing agentic workloads and really people using 10x of what we're at right now because they have you know 100 agents doing XYZ things at kind of all times. Now I was just at some conference, they're like, Oh, yeah, the the compute need's gonna go up like a thousand times for every individual. You know, how's that gonna actually happen?
SPEAKER_00So Yeah, and how yeah, how is that sustainable?
SPEAKER_01Sustainable from a yeah, cost model and uh, you know, it it definitely will change, I think, the the shape of the compute, you know, moving forward. Uh, it starts unfolding, which is not you know that far away, I think.
SPEAKER_00Well, if there was one key takeaway you could leave our audience with today, what would that be?
SPEAKER_01I I'm going back to the the beginning point I had that you know, as you're rolling out these models, whether they're Frontier Lab, on hyperscale providers, Azure, you know, AWS, vertex, whatever, uh, whether you're bringing them on-prem, starting from the very beginning, putting in a system for visibility across all the different models and guardrails, if you can, you know, putting in a system for guardrails so that you do have the option to turn off the faucet if you need to. That's my biggest, I'd say, tip right now from what I've I've been seeing.
SPEAKER_00So okay, great. Well, thank you so much for coming on the show and sharing your insights with us today.
SPEAKER_01Sure. Thanks, Amanda. Good questions.
SPEAKER_00And thank you to our audience. If you have any questions or comments, please leave them below and I'll try to respond back as soon as possible. Have a wonderful week.