Edge of Excellence: Empowering People to Shape the Future

Cloud Costs Out of Control? How to Reduce AWS & Azure Spend | Cloud Governance

iuvo LLC Season 1 Episode 20

Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.

0:00 | 54:17

Cloud costs rarely get out of control because of one big decision. More often, they're the result of small decisions that accumulate: forgotten resources, over-provisioned environments, inefficient architectures, unclear ownership, and a lack of governance.

In this episode of Edge of Excellence, Jessica and Bryon are joined by Brian O'Neill, Principal Consultant at iuvo, for a practical conversation about cloud cost management, AWS and Azure spend, and cloud governance.

Brian explains why the ease and flexibility of the cloud can also make it surprisingly easy to spend money without realizing it. From forgotten development environments and oversized databases to running workloads 24/7 that don't need to be, the conversation explores some of the most common sources of unnecessary cloud spend.

The episode also dives into right-sizing, scaling, tagging, account and subscription structures, Infrastructure as Code, budgets, alerts, and governance guardrails, and how organizations can implement these practices without slowing down the teams that depend on cloud resources.

You'll hear a real-world example of how identifying unused and over-provisioned resources helped a client save roughly $50,000, as well as why some workloads may actually be more cost-effective to run outside of the cloud.  

Brian's central message is simple: you need to know what exists in your cloud environment, what it's doing, who owns it, and why you're paying for it. Effective cloud governance isn't about putting up barriers, it's about creating guardrails that give teams freedom while keeping costs, security, and the overall environment under control. 

If your organization is managing AWS or Azure and wondering whether you're spending more than you should, this episode offers practical ideas for where to look and how to start.

Topics covered include:

  •  AWS and Azure cost optimization 
  •  Cloud cost management 
  •  Cloud governance 
  •  Right-sizing cloud resources 
  •  Cloud architecture 
  •  Lift-and-shift migrations 
  •  AWS tagging and Azure subscriptions 
  •  Shadow IT 
  •  Infrastructure as Code 
  •  Cloud budgets and alerts 
  •  Cloud security 
  •  Cloud scalability 
  •  DevOps and governance 
  •  When to bring in cloud consulting expertise 
  •  When cloud may not be the right choice for a workload 

Learn more about iuvo

SPEAKER_02

This is The Edge of Excellence, empowering people to shape the future. Let's inspire, innovate, and explore together. Welcome to the Edge of Excellence. I'm Justin Forge, and today we're talking about cloud cost and governance.

SPEAKER_03

I'm Brian Beinelman, and joining us today is Brian O'Neill, principal consultant at IUVO. Brian worked closely with organizations on cloud architecture, AWS, and Azure environments, automation, governance, monitoring, and helping teams build cloud environments that are secure, scalable, and very important, easy to manage. Brian, welcome to the Edge of Excellence. Thank you.

SPEAKER_02

Brian, before we get into cloud cost and governance, can you tell us a little bit about yourself and your role and the kind of work that you do with clients? Sure.

SPEAKER_00

As Brian said, I'm a principal consultant here at IUVO and I've been here for almost 14 years now. I've been working professionally in information technology for over 35 years. So I've hit on a lot of different technologies, um, services. Um, I've been working with cloud platforms specifically for about the last 12 years, and I currently work with clients on various things like automation, infrastructure as code, controls of both data and uh infrastructure governance. And I currently hold certifications in both AWS and Azure infrastructures.

SPEAKER_02

Awesome. And I actually remember you were actually the first person at Iubo that I interviewed for a case study when I was really, really green. And I was so intimidated. And I remember you mentioning Kubernetes, and I just like my whole brain went blank. Like, what is he even talking about? Is this even English?

SPEAKER_00

It's not an unusual reaction to that term.

SPEAKER_02

I just like you are just like very much a core part of my memories of my early days here, just being so impressed with like, wow, I am in a much different role than I used to be. Um, but I love I love that memory and I love thinking about Kubernetes and now want to hear it now that I can explain what it is, but at least it's more familiar to me. It doesn't feel as abstract and strange. Um, so I want to dive right in. When a company says our cloud costs are getting out of control, what is usually happening beneath the surface?

SPEAKER_00

Um, usually one of the first things that comes to mind is that uh they they really forget about how the cloud really works. Um, you know, to put it simply, um, you know, you're really renting servers and services, uh, but you don't have you don't have the physical hands-on of it. It's out there, but it is, as they say, somebody else's computer. And the more you use it, the more you pay. And some people don't understand that even though you're not necessarily using it, you have it allocated, it's costing you money. Um, one of the challenges is that they they make the cloud very easy to use. So you can go onto the web the website and create a virtual machine or some other service. Um, you can use the command line, you can use automation tools. All of that makes it really easy to spend money. Problem is the person doing that isn't the one paying the bill, and therefore uh it makes it easy to create resources and even forget about them. You know, you can create a bunch of resources to test something, and then you don't necessarily clean up when you're done, and they just sit there and they actually continue accruing costs, and then eventually somebody down the lines looks at the bill and says, What is this? Why are we paying for it?

unknown

Right.

SPEAKER_00

Um in addition, um, the cloud can be more efficient cost-wise than running your own systems, but it really depends on the workloads you use on it. A lot of companies think, you know what, we don't want to have infrastructure in-house, we don't want to pay, you know, buy servers and uh, you know, deal with the the upfront costs of doing that. So they say, let's go into the cloud. But they don't really give a lot of thought to how they're actually going to implement themselves in the cloud. Sometimes it's just like, oh, we'll just build a virtual machine just like the server we would have bought, and we'll do everything in there. But that running a virtual machine 24 by seven, that's beefy enough to do certain tests, is actually extremely expensive, especially if you're looking at doing things like artificial intelligence and LOS, where you have to also pay for uh a GPU as part of that. And while it may be cheaper in the short term, if you do that over the long term, you're very likely going to spend more money than you would have if you had just bought the equipment up front and and had it in-house. So the it really depends on the workloads you're trying to utilize in the cloud, as well as you know what your what it is you want to put in the cloud, and is it capable of using services that the cloud platforms provide more directly as opposed to I'm just gonna create a VM and do it myself, where farming it out to the different services can actually cost you less, but it's not how people are used to thinking it in developing a platform and managing it. They're used to I build it on the server, I want to just build it on a server in the cloud. But that often isn't the the most efficient way of doing it.

SPEAKER_03

Yeah, Brian, I think what you're some of the things you're describing is around uh some of the lift and shift technology like mindset, like I've got the server, I'm just gonna virtualize it and put it in the cloud.

SPEAKER_00

And yes, and and we we we've done that for clients, and it wasn't necessarily the best way to do it, but they they for whatever reason they needed to get out quick and they do it. But after we get them in there, and they're they're operational at least, then we start looking at well, how can we adapt to the cloud platform natively as opposed to just recreating what you would would in-house or at a data center and trying to do the same thing in the cloud?

SPEAKER_03

Yeah, and Jess, this this goes back to your original comment about Brian, like uh you know, using containers instead of uh like in inside of Kubernetes, right? So these are some of the techniques that are used. So, but uh that's why we're talking to Brian, because he knows how all this works.

SPEAKER_02

Yeah, exactly. So, what do you think are some of the most common areas of cloud spend that organizations tend to overlook?

SPEAKER_00

Aside from the fact that resources are are easily forgotten. You know, I created a development environment, I ran my tests, and I simply forgot to delete it and its accruing costs there. Um, often uh in other cases, it's over provisioned. Um, it's easy for you know a manager to say, we can't have a slow website. So, how do you make sure you don't build a slow website in the cloud? Well, you make sure that you allocate enough resources to do at least the minimum, but also at least the maximum kind of load that you expect. Because if you have a lot of visitors to your website and you've got a database and everything else, it can all slow down as it gets busier. Um so without having like a real you know without having a great concept of uh what resources you necessarily need, especially if this is a brand new website, it's never seen production level traffic. But we know we're gonna build it out there, but we don't know what it's gonna need. Again, you tend to over provision it to make sure how much how much CPU do you allocate to it, how much RAM do you allocate to it? You don't necessarily know what that should be. And so the the normal reaction is either go small and find out it's not enough, and then you have to scale it up to you know more expensive resources, or you over provision it and make sure that you don't have that problem in the first place. Now, from a uh a customer or end user standpoint, you over provision, you don't want the customer to have a bad experience, especially even on day one.

SPEAKER_02

Right.

SPEAKER_00

Um but it may turn out that you've overprovisioned significantly, and you could be paying two, three, four times what you should just because of that decision. And so between underestimating and overestimating, you need to look at what the results are and make adjustments to it, and it can take you some time to really sort that out. Um, but once you get an idea of well, these are the resources we really need, then you can do the plan to right size the environment for your application, for your website, whatever that happens to be, um, and make sure that you're not you know just paying just because you don't want the site to ever be slow. Um and design decisions go in as as well. You can design an application to be able to scale. And there's two ways, two ways of scaling that that that we call there's scale up, which means make a bigger resource, run the website on a bigger resource. Usually that means we have to shut down or or or do something that could be uh there could be some downtime involved with that, or you scale out, and by scaling out, we simply allocate more re more identical resources. But it really depends on the architecture of the application that it understands that that can happen. If I have a website that has to maintain state of users um accessing it, and I have multiple web servers, I have to make I have to have some way of making sure the users are always on the same server or that that session information is shared. That has to be included in my architecture, but it's a lot faster and simpler to scale out so as the load of the website goes up, we add more resources behind it. And we and we distributed that load, you know, as evenly as possible. And we can even, if we have that ability, we can even scale out with larger resources too, if we have to. And then once that once we've kind of dealt with our our peak moments, say you make you make an announcement and thousands of people start hitting your website, right? We can scale up, we can even anticipate that scale up in advance, or we can do it automatically. And then when that has died down, we scale in and remove those extra resources, and that way we're only paying for them for the duration in which we actually needed them. And then we go back to kind of our our steady state resource where we have we have our expected costs.

SPEAKER_02

This seems like it requires someone or a team to really be paying attention, watching, and monitoring that. Is that difficult for small businesses? Like, is that typical that they would have someone kind of monitoring this?

SPEAKER_00

Um, well, again, it depends on who who you have on staff who's doing this. Right. You know, a small business, you might have a technical resource who knows how this stuff works. You might have a technical resource who understands the minimum in order to get this running. That that can be uh a bit of a challenge there. Um, what it mostly demands is communication from all facets of the organization. If marketing is going to make an announcement that is gonna drive traffic to the application, marketing needs to communicate that to the technical people to be prepared for it so that they aren't scrambling after the after it's already started to make adjustments.

SPEAKER_02

No, that makes a lot of sense and it's a very proactive approach. So I'm curious what makes AWS and Azure so powerful, but also so easy to lose track of?

SPEAKER_00

Again, it's I think it's uh the ease of use. Um, you can create an account in AWS or Azure, and you you can often get free credits to to help learn it. Um and then then you you know, with a couple button pushes, you can create a virtual machine, and then you can log into it, you create a website, and so forth. Very easy to use for the most part. Understanding what what you've done can can be a little harder. But again, the the fact that you did that, and then maybe you you created that account to do your own little learning session, use your free credits, you left it running. Those free credits expire. Uh you had to have given a credit card just to open the account. All of a sudden you get a couple hundred dollars on your credit card the next month.

SPEAKER_02

Surprise.

SPEAKER_00

And it's like, well, why did I get that? You forgot to shut down the system the resources that you created as part of your test. Um and to to kind of bring security into it. Uh, I've there I have just so many horror stories of people setting up these like temporary learning accounts and then forgetting all about them. Now, maybe they did shut down the resources and and there was nothing left, but the account still exists and they didn't necessarily secure it very well. Um, you know, with with AWS that you know early on this was a challenge in in how do I make make it so it's more secure. Um, they they've addressed that. But so many stories of people is like all of a sudden I get had a $3,000 bill from AWS. It's like, why did that happen? Oh, well, they found out their their account was compromised. And someone went in and started allocating resources, typically with the purpose of trying to compromise other accounts or or say run like password crackers or or even spin up a GPU instance and uh start hacking away. Um, and that you know, all of a sudden you can go from a zero dollar bill to multiple thousands of dollars. I I remember one case about fifteen thousand dollars on on a student who obviously would have really no way to pay that. Wow. Um, so you know making sure you're aware of your resources and keeping track of them is a big key. Um, even you know, even in you know, you know, larger environments, um, it it wasn't all that unusual that someone needs to set up a test copy of the production website. So they allocate those resources and they allocate the same size resources production has, even though they're just doing a simple test and don't necessarily need them. And again, they forget about them. And depending on what your normal AWS bill is, it can be hard to notice that hey, this this is higher, but maybe that's normal. And it, you know, sometimes it can go months before someone asks the question.

SPEAKER_02

Right.

SPEAKER_00

And you know, why why is this bill still high? And it, yep, there's two copies of production out there, one of which is actually being used, the other is just sitting there doing nothing but costing us money.

SPEAKER_02

I want to I want to flex here, see if I can throw at a technical term. Would this be would this be a case of like shadow IT where you just have things going on that you're not aware of? Or is that not a correct use of that term?

SPEAKER_00

I'm not sure because you know that's that's a term I tend not to use.

SPEAKER_02

Okay.

SPEAKER_00

Um, you know, I'm an IT person, so I'm usually in there, so there's usually nothing that goes on that I don't know about. But I think with shadow IT, um, and this gets into some of the uh concepts of what's called development operations or DevOps, where you enable more people to do IT functions, like creating resources, but though those people aren't as knowledgeable about what they're necessarily doing. They're given some instruction, perhaps, but they are allowed to do it outside of the direct purview necessarily of IT who's trying to keep an eye on it.

SPEAKER_02

Right.

SPEAKER_00

And so IT has to use other tools to make, you know, see, okay, this person's done this, this person's done that, and keep track of everything that's going on. Um, you know, even if it is just I have to click around and make, you know, look at everything that's been allocated. Um but the you know, enabling those people and then and reminding them as a you need to clean this up when you are done.

SPEAKER_02

Yeah.

SPEAKER_00

Um other ways we can kind of control that shadow IT is that we make sure that we set certain guardrails up in in ways that so if somebody does allocate a bunch of resources, but it's like a development environment, they're doing development and testing in it, but they go home at the at the end of the day. Right, they're not gonna use it after 5 p.m. or or whatever time. We can shut that stuff down. Most of the costs are on the runtime of virtual machines and so forth. When they're not needed, there's no need to run them. A lot of people are kind of afraid about, well, if I shut it down, I gotta spin it back up, and you know, maybe it won't come back up, or you know, you know, problems like that, which are really rare. But those of us have been in IT for a long time remember that it used to be a little more common.

SPEAKER_01

Yeah.

SPEAKER_00

But so, you know, they there's a reluctancy to do that, especially, you know, like if I come in in the morning and I need to get working, I don't want to have to wait for the environment to spin back up.

SPEAKER_02

Yeah.

SPEAKER_00

But in reality, you know, in the simplest cases, it's really a minute or two to get it back and running. And the cost savings of it having running it only eight hours of a day instead of 24 are extremely significant.

SPEAKER_02

Yeah.

SPEAKER_00

So um those are ways where we can help prevent runaway costs because we've delegated the ability to other people. Um, and you just make sure that we're not spending more than we really needed to. Um and then, you know, other other ways of kind of detecting resources we didn't expect can be implemented. Um, you know, again, for that shadow IT, you know, always be informed of what's in your environment. So we can uh we can have applications running within the cloud that scans all the resources in the cloud and says, hey, this wasn't here before. And it doesn't bear bear the information that we're we require to have on it. And therefore, we need to either correct that or get rid of it. Um, and then make sure whoever did create it knows really how they should be creating it so that we know who created it and so forth.

SPEAKER_03

You know, Brian makes me think a little bit. I mean, we I know we have some tools that we've used to you know assess people's environments. Well, here's what your, you know, here's where your costs are. Uh there's even an example I could think of that we did uh for a customer where they acquired a customer uh and then they go, wow, can you look at this environment? And and we found like we just like there was like 50,000 or some crazy amount of money that we're just in that very scenario, right? Where we we can look at these environments, yeah. We have the but between the tools and know-how, I guess. So there's a there's various ways that this can be accomplished, right?

SPEAKER_00

Like, you know, from from uh Yeah, and and and we may we may touch on this a bit later, but I I I remember that situation actually, and we came in and we said, do you realize you have all of these databases? You've got this database and this database, and this is a copy of that uh for for your development environment, but they're all giant databases. You've allocated them as if they're production. And we looked and they weren't weren't being utilized. They had more storage than needed to be allocated, more RAM, more CPU. And, you know, one, they didn't realize they had so many copies. So we were able to get rid of some of them right away. Uh, because again, they had forgotten about them. And those that that they did need, we were able to recommend how to right size them for the the development purposes. And the uh I I think you said it, the savings is like $50,000 a month.

unknown

Wow.

SPEAKER_00

Yeah.

SPEAKER_02

So those numbers are rather impressive. If someone is listening and they think that they would like to try to reduce cloud cost, is there something that they should understand first? Like what would be a good starting point?

SPEAKER_00

The first thing is accounting for everything that you actually have in the environment. And there's a variety of ways to do it, and it really depends on how big that environment is and how many people have the ability to create resources and so forth. Um generally what the organization organization sees at the top level is the bill. You know, obviously that's telling me you how much you pay, but uh they generally break it down by service, but that's usually not enough. That's not detailed enough to tell you, well, okay, we're running 20 virtual machines. It doesn't tell us what they were doing, or you know, and or who, for that matter, ran them. So there's a variety of ways we can kind of split that information up. Uh, one thing I and I know AWS supports this, is that you can tag resources for billing purposes, and you can generate a billing report that says, okay, of the bill, what of the resources have this particular tag so that we can see what the cost of that particular resource is. So you can say all the resources for department A needs to be tagged with department A and Department B, same department B. And we can run a run a uh billing report that says, okay, tell me the cost of everything for department A or department B. Uh we can segregate things even further through the use of uh AWS calls it organizations uh and accounts. The account term gets a little confusing with AWS. In Azure, it's subscriptions, and so uh you create a single organization, but you can have multiple subscriptions or accounts, depending which platform it is, and the costs are allocated to those subscriptions. Now they can all bubble up to a master bill that gets paid, but you can see by account or by subscription what those specific costs were. So I can give Department A their own subscription or their own account, and they can do whatever they want in it. And then there's a bill that says this is how much Department A costs. Okay. And it's 100% independent of Department B in terms of resources. They, you know, they can't like cause problems with each other because they're independent domains of control, but also we get the that billing separately broken down by the account or subscription, uh, so that we can say, hey, department A's bill has jumped three times in the past month. Well, department B is steady. Why is that? Did we expect that? So forth.

SPEAKER_02

So does that help as far as if you are deciding that you want to move into the cloud, kind of establishing that from the the very beginning so that you have more visibility and it's easier to kind of see where discrepancies are? Is this tagging that you're talking about?

SPEAKER_00

Yeah. So so tagging is something that can be adopted later, but it is best if we can get it in there early. Uh, the account structure or you know, with with Azure, you're kind of forced into it because you you create your Azure organization and then you have it, you have to get a subscription anyways.

SPEAKER_02

Okay.

SPEAKER_00

And so you at least have one subscription, and then you can just add more. But if you want to go in that direction and you want to segregate the resources in the cost costing, uh, you really want to think about that in advance and be prepared for it uh so that you can create the structure around it first. Otherwise, you have to do some finagling to get it in there. One one of the chief problems is that when you create a resource within a subscription or within an account in AWS, you cannot, at least last I knew, move that from one to the other.

SPEAKER_02

Okay.

SPEAKER_00

So if you have a single single account in AWS and you have two department A and department B in there, well, they're they're allocating resources and they're all together. And then now you want to separate them, you actually have to tell one of the departments you need to recreate it all, start over in a way, yeah, and migrate all your stuff and then delete them from the original uh situation. I believe in Azure they do allow moving between subscriptions now, but there's still control domains and stuff involved as well that you need to set up just to make sure that people only have access to the resources they're supposed to have access to. And the with the simplest way is to do that at the organizational level, at the subscription or at the at the account in AWS, and then you can have finer grain controls underneath that. But you want you getting that set up with the you know, in advance, if you think you will ever possibly want to do it, you should just do it. Even if you only do it for one instance, rather than down the road say, hey, we got eight departments now and they're all in the same account, and we need to split that up. That just that just gets much more complicated and there's downtime and everything we have to deal with.

SPEAKER_02

Now, are there risks to optimizing too aggressively or cutting the wrong things?

SPEAKER_00

Uh well, so especially when you're dealing with forgotten resources, there's always the risk that you are going to um delete a resource that you didn't think was important, but it turned out it was like the linchpin of the whole thing and everything collapses. Um, we did have a client one time, this wasn't in the cloud, it was pre before their move, but they swore up and down the server wasn't actually being utilized. And then one day their website ran extremely slowly. Like it was taking 10 seconds before we'd even respond and begin loading. And they're like, Yeah, we're looking at all the servers and everything's fine. But then I took a look and there was this other server, the one they swore wasn't being used, and the application had crashed. Oh boy. And then we restarted the application, and instantly the website started working again. And it turned out that there was a dependency, but whoever had written the dependency into the website had left the company previously, and so no one actually knew it was still in there. So you always run the risk that, hey, yeah, we didn't think this is being used, so we shut it off. And it turns out it is, and then hopefully you can just turn it back on, or you know, or you have to like recreate it or work around it or whatever. Right. Um in terms of like downsizing, scaling down, uh, you could always pass that point of performance impact. You know, you scale down too far. Unfortunately, the the the fix to that is scale it back up again. Um so so the the risks aren't huge depending on how we're trying to fix it. Um, but we always want to make sure that we have a backup plan of whatever it is we do, especially when we're dealing with a resource that no one's quite sure about. So if we can simply stop it and then then determine that there's it wasn't necessarily, and then we can get rid of it, that's great. Depends on the resource how easily you can do that. You can stop a virtual machine, and then you can always just start it up again. But there's other resources that they they're either there or they're not, and you know, so you have to get rid of it. It's like if you get rid of it, turns out it was important, you have to not only recreate it, but you gotta you gotta recreate all the access controls or whatever else that went went as part of configuring it as well.

SPEAKER_03

So so Brian makes me think, is this I don't know if I this this ties in, but with the DevOps methodology, if you start that from the from scratch, where it's just going, I'm gonna you can go easily click on a web and go add resource and do this whole thing. And but if you start with kind of uh kind of maybe it's like like the cloud launch pad methodology where you're like, we're gonna build this via in infrastructure as code and things happen, then you can can you describe that a little bit more, maybe where the where the pros on on that kind of thing, that thinking is so inf infrastructure as code for anybody who's not familiar with it, is being able to create your cloud infrastructure in a in a programmatic way.

SPEAKER_00

So you literally write text code that describes what your environment should look like, all the configuration parameters and so forth. And then you apply that to the cloud, and those resources get created. And they're created according to how you describe them. And there are various advantages to that. One is that you know you can document your environment in the code. You don't know it's as I you know, I you don't have to write a separate document that says, hey, I've got all these resources, and this does this, and this does the you can annotate the code with comments and say what each one did. So as you as you add a resource, you put in this is what it does. And one of the one of the great things is it it does it really fast. So it's a lot faster than having to click around and create this resource and create this resource and so forth, or even you know, command line, this command creates this and this command creates that. It can all be done at once. And the tools are pretty smart about if there's a dependency of one on the other, this one gets created first, and so forth. And if we need to create multiple copies of an environment, we can simply copy that code, change a couple parameters to say, okay, do it in this account or or give it this IP space or these names to distinguish it, with just a few edits, and then run the command again, and poof, now we have a duplicate of the environment. Exactly the same aside from whatever little edits we make. Um there are tools, there are in fact like Web GUI tools that you can use to do this so that you know people who aren't familiar with the command line or whatever can push literally push a button and say, give me a copy of this environment for me to use. And you know, boom, it gets done, and they're given the and they get the information to how to connect to it and begin utilizing it. So so the the the infrastructure's code you know is actually used within that tool. You have to write the code for the tool to use. Um, but then it provides a nice simple go do it button for people, you know, making making it even simpler for them. And we can provide the policy enforcement of what we want for those resources. Uh, we haven't talked about tagging much yet, and I I I know we want to touch on that. Tagging is a way of attaching an annotation to a resource so that we can identify it. And we like to have we we have a set of tags we like to use all the time. One is who created it, because that isn't necessarily easy to find after the fact, but we can create a tag that says creator Brian O'Neill, creation date, when it was actually created, or sometimes we use a review date where we, you know, we run a report that says we need to review this to make sure we actually still need it, or maybe need to update it, or whatever. Uh, so we can put dates in there. Uh, what service is it providing? You know, so we have an application called, you know, our our favorite app. We can have a tag in there that says this is part of our favorite app. So we know that if something happens to this resource, we're affecting our favorite app. Um and also, you know, we we all sometimes put in a tag that says, well, what tool do we use to create it? Because even in a single environment, we might use a few different tools. So there you could use control tower, or you can use Terraform, or you can use Ansible. All of these are capable of creating resources um in an infrastructure as code way. Um, but if you know we create a resource with Terraform, well, we can't do something to it with Ansible easily. You know, so we want to make sure we use this the same tool for that resource. You know, Terraform, if we create it with Terraform, we want to destroy it with Terraform. Destroying being kind of a strong term for getting rid of it. But that's actually what it's called in Terraform.

SPEAKER_02

What would the bill be the biggest warning sign for businesses to look for that a cloud environment needs attention, or are there other warning signs that businesses should be aware of?

SPEAKER_00

From like a business management standpoint, there's usually nothing that they receive other than the bill. Um unless there's somebody really keeping an eye on things and saying, hey, wait a second, we had a lot of resources. You might want to expect a bigger bill. Um, unfortunately, the bill is sometimes the point at which it's too late. Interesting. To get the bill, you are liable to pay that bill. Right. And and according to the terms of payment. And so if suddenly you go from what you thought would be a $500 bill to a $5,000 bill, and you know, that's not in your budget to pay it. Well, you're still liable for it.

SPEAKER_02

Right.

SPEAKER_00

So there are tools available where we can set uh typically they're called budgets, and it's just like an electronic um value that we say we do not expect to spend more than this amount of money in this particular account or subscription. So say we set that to $500, and we are typically under that $500 limit, and we don't hear anything, and that's fine. But if the predicted value before the end of the month looks like it's going to exceed that because something has increased in cost, then it sends out alerts through various means, typically an email, uh, but it can it can do texts and and so forth. But it sends out a notification or an alert that tells you, hey, look, this account is going to is predicted to exceed your budget by this time, you get an earlier indication that maybe someone needs to go and take a look and say, well, why is that happening?

SPEAKER_02

Yeah.

SPEAKER_00

And it may only be a small amount because we had a little extra traffic, and maybe that's going to be a continuous thing and we up the budget on it. But if it tells you on the first of the month that you're going to exceed your budget on the second day, something is obviously wrong and needs to be fixed right away.

SPEAKER_02

Yeah.

SPEAKER_00

Yeah. So we typically use that. And it's interesting because we sometimes get, you know, we we we get some alerts sometimes, you know, but it's like, oh, you're gonna you're you're gonna exceed your budget, and then we look at it, and it turns out we'll actually end up exceeding it by 10 cents. So it's fine. No need for at least we have that safeguard in there. Uh, you know, so that you know, maybe we have to increase the budget by 10 cents. I don't know. Right. But you know, we we at least get notified sooner rather than later, because like I said, once the bill comes, it's too late. Right. Uh you you're still liable for whatever charges have accrued at that point, but it's better to stop it if it's a thousand than it is five thousand at the end of the month. So you you you can you can you you're able to react faster to that. You can also set um budgetary guardrails that say we absolutely cannot exceed $500 for this account. And you set and say, well, what should happen if if we do? And one of the things you can do is it shut it all down. You know, that is that is used often with like development environments where, or at least testing environments that aren't as vital to the, you know, it's not your production environment. No one on the outside's gonna see it, but it's running our QA tests. And let's say our QA tests were unusually intensive early in the month, and we absolutely don't want to exceed the budget, and we're willing to push off work into the next budget, right? You know, we can say, okay, we're just shutting it down. And then that that that stops most of the billing. It doesn't necessarily stop all of the billing because the way the cloud works, you know, there's some continuous back end costs, whether your virtual machine is running or not, but they're usually a fraction of of the of the primary cost. But we can stop those primary costs, and therefore our bill should still fit within our the actual expected budget. Uh, and we don't have to like constantly say, oh, we need to, we need to, we need to give that department a little more money this month, you know, playing playing shell games with with with your actual money.

SPEAKER_02

No, that's huge. Um I'm curious when you think it makes sense for an organization to bring in outside cloud expertise.

SPEAKER_00

It's easy to say you should bring it in on day one.

SPEAKER_02

Yeah.

SPEAKER_00

The in reality that that's typically not done. Um typically companies move into the cloud and then say, what are we doing? Uh once you start to start to lose what is actually going on in the cloud, you need to bring somebody in sooner than later. If if you're not sure where your costs actually are or whether your costs are what they should be, and but you're not sure, well, how how can how can we reduce our costs because we're not familiar with how the cloud actually works? You know, back to that, you know, we said early on, it's easy to just lift and shift and create a bunch of VMs and do the work, and that's not necessarily the best way. And you find out that's costing us way too much money. Right. But you don't, you know, you don't know how to adapt your application to work in the cloud, you should be bringing in a resource who is very familiar with the cloud, hopefully also familiar with your your your business and your application usage, and can work with you to re re-architect, resize, um, you know, all these different ways that we may be able to use the cloud more effectively. Or there's the possibility that we have to point out maybe you shouldn't be using the cloud for it. There are there are industries where where the cloud is not cost effective at all. Uh, we deal with a lot of electronic design automation companies. These are CAD users who design chips. And with the, you know, if you follow the industry, chip die sizes are getting smaller, the individual components of a chip are getting smaller, which means you can fit more on a chip, which means there's a lot more design and calculation that has to go in to, you know, they run simulations to see does this work? Do we exceed specifications or do we run into issues with uh RF interference perhaps on data paths? You know, there's a lot of information that gets calculated as part of this. And as more and more, if the chips get more complicated, there are more and more complic, you know, calculations. The storage requirements that are necessary behind it and the computational requirements are very high. And they typically run these simulations 24 by seven because you know they take a long time to run and they Also, they have to run regression tests and and all sorts of things. So they're submitting jobs into the cloud, which get executed, but those servers are running 24 by 7 and they're beefy servers, and that's where you start paying more for the cloud than you could have in-house, you know, or even at a data center. And you know, it depends on the size of the operation. But we've seen in most of the cases that doing those workloads in the cloud has is just too expensive. Even if you have just a few servers on site to do the work, it tends to cost you a lot less than doing it in the same thing in the cloud because you can't you can't realize savings of downsizing, you can't realize savings of um uh shutting down because they're constantly in use.

SPEAKER_01

Right.

SPEAKER_00

You know we may be able to do some scaling, but you know, scaling, you know, what are your minimums? You're running these some of these jobs 24 by seven, we have to have a bunch of systems to run those jobs on.

SPEAKER_02

Now we're getting close to the end of our session. I can't believe that the time is flying the way that it is. Brian, if there was, you know, one thing that you would hope that listeners or viewers would take away from this conversation when it comes to cloud cost, what would that be for you?

SPEAKER_00

Um I think the main the main thing is, you know, to make sure that you're aware of and in control of your costs is to make sure you have a governance system in place to make sure we're not overallocating. We don't want to be in the way of the users who need to use the resources. We don't want to have them have to have them request a resource and there be a delay and they're sitting idle because there's nothing for them to run their tests on. We don't want to, we don't want to be a blocker from a management point of view. We don't want to be a blocker to them getting their work done because we put, you know, we we're we're too tightly controlling it. But we can at least have guardrails in place to make sure that things are being created with the information we want to be in there so that we can identify it, you know, when we go and say, well, what is this, what is this VM for? What is this database for? And we're not guessing it and then you know shutting it down and breaking something if we're not careful. Um having the policies and the governance in place to allow the users to do what they need to do with as minimal delay as possible, but be able to make sure we are keeping within our budgets and our our purview of control that we know what everything is in there for, and that it's legitimate and you know, is in fact being used. I think that's that's the main thing.

SPEAKER_02

Yeah.

SPEAKER_00

You know, so you know, if there's anything in the cloud and we don't know what it does, we have failed in being able to control our costs because yes, it might be a legitimate resource, but it's costing us money and we don't know why.

SPEAKER_02

Yeah. 100%. Brian Weilman, I'm just realizing I'm with the two Brians uh right now. Is there anything else that you would like to touch on before I ask our closing question?

SPEAKER_03

No, um I would I think Brian said it really well. I would like to say that, you know, it Brian has a nickname in Ivo Doc, you know, because uh and and you know, just the way you can explain uh these very complex concepts, yeah, yeah, uh show that demonstrate your knowledge. And I think our our listeners really appreciate this because it's it can be overwhelming. Uh it seems simple at first, but then uh usually this is how the the chaos comes in that we come in and we we turn uh chaos to clarity. And so it's easy to get in that space. So with people like you explaining it, I think uh hopefully they can get some insight that there's a there's a there is a better way.

SPEAKER_02

No, definitely. So, Brian O'Neill, it's your first time on Edge of Excellence. We like to ask our first-time guest to share something that's interesting about yourself that might surprise people.

SPEAKER_00

Yeah, no, there's a few things. Um since Brian mentioned my nickname of doc, that actually got shortened from professor, which was a name I was given by my teachers in fifth grade. Mainly because I I knew so much that I could tell them some things. Wow. So they nicknamed me Professor. Um, but um I I've done a variety of things on the outside that uh are you know are not IT related. Right now, I'm very involved in local theater. Um, not on stage myself, at least not yet. Uh, but I do tech, stage, uh, front of house. Uh I helped form a 501c3 organization uh to help support uh the arts in our public schools. Amazing with with a focus on the drama departments. Um and I for acting, I did work as a background actor once on an Apple TV show called Defending Jacob with Chris Evans. Oh, and my scene, it turns out that you can just see the back of my head for a second or two. I'll look at it. My entire family actually got to participate in that show. My oldest daughter was in the same yeah, in similar, you know, related scenes as a photographer. I was I was a new sound engineer, which you know really wasn't all that far off of what I things I've done in the past.

SPEAKER_02

Oh, cool.

SPEAKER_00

And my wife and other daughter were in a uh a school scene where my daughter got you know a good several seconds of focused camera on her. Uh as they're like panning the audience, she was kind of like the focal point of the camera.

SPEAKER_02

Oh, cool.

SPEAKER_00

So we we got to participate in that. That was very interesting. Uh, haven't haven't really had time to do you know anything similar, but you know, what a fun like family experience, too. To get to it was very interesting because there's a lot of sitting around and doing nothing.

SPEAKER_01

Yeah.

SPEAKER_00

Because they got to set up scenes and they do it from different angles and and so forth. It you know, I I've been on TV sets before because I I used to run a website, you know, I'm retired from it, but I used to run a uh science fiction news website. Oh, cool, and so I got to go to various uh TV sets, and I see I've seen uh them actually filming, and for like a 15-second scene, it took 20 minutes.

SPEAKER_02

Oh, yeah.

SPEAKER_00

You know, it's like figuring out where, okay, we're gonna put the camera here for this shot, and then we're gonna, you know, but we need to move the lights now, um, and and so forth. But seeing the process and you know what the sets are actually like in a lot of cases, you know, they're a lot more real than you would think in some cases. So, so I yeah, I've I've it was very enjoyable to actually be able to experience how that work actually gets done. Yeah, which a lot of people don't understand.

SPEAKER_02

A hundred percent. A hundred percent.

SPEAKER_00

It's really cool.

SPEAKER_03

Yeah, that is amazing. Uh, I didn't know I'm I know a lot about you because we've worked together for a long time. I did not know the the fifth grade professor thing. That is that is all that is great. Um well, Brian, this is for listeners who are hearing this and thinking, this sounds a lot like our cloud environment.

SPEAKER_00

What is the next best step? It's easy enough to go to our website, uh iuvotech.com, i U V-O-T-E-C-H dot com. And right at the top on the right, there's a schedule a consultation button. If you're on a mobile device, I think you get to open the little menu, but it but it's in there. And you fill out the form, and that will you know get Brian and and our salespeople involved, and we'll we'll come talk to you and see well, what is your situation? Um, is it really scary or do you just think it's scary? Um, and then you know, we can help you make it not scary.

unknown

Cool.

SPEAKER_02

I love that. Brian, thank you so much for joining us today. This brings us to the end of today's episode of Edge of Excellence. Today's conversation is a reminder that cloud cost problems are rarely just about the bill, they're about visibility, ownership, architecture, governance, and making sure cloud environments are actually supporting the business.

SPEAKER_03

And if your organization is trying to understand AWS or Azure Spend, improve cloud governance, or build a more secure and repeatable cloud foundation, visit iuvotech.com to learn more about ivo's cloud consulting services. And if you like this podcast and you want to and you want to support us, uh click the like button, the share button. Uh the algorithms like that. It'll help your friends. And thanks for listening. And until next time, stay curious.

SPEAKER_02

Thank you for tuning in to the edge of excellence. We hope today's insights empower you to shape your future and rise to your full potential. Let's continue to grow, innovate, and lead, pushing the boundaries of excellence.