Microsoft Teams Insider
Microsoft Teams discussions with industry experts sharing their thoughts and insights with Tom Arbuthnot of Empowering.Cloud. Podcast not affiliated, associated with, or endorsed by Microsoft.
Microsoft Teams Insider
How To Troubleshoot Microsoft Teams with Call Quality Dashboard (CQD), Hands-on with Microsoft
Use Left/Right to seek, Home/End to jump to start or end. Hold shift to jump forward or backward.
Victor Guzman, Senior Technical Program Manager at Microsoft, and Siunie Sutjahjo, Principal Product Manager, Microsoft Teams Meeting, walk through the exact quality review process from Microsoft's Get to Green programme, giving every Microsoft Teams admin the tools to diagnose and resolve call quality issues.
• Victor shares the seven-step quality and reliability checklist that resolved 99% of escalated customer issues
• A full walkthrough of the CQD QER 5.3 templates, covering usage summaries, media health dashboards, transport reports and VPN analysis
• How to identify whether issues stem from proxy configurations, VPN tunnelling, DNS resolution or network congestion
• Using weekly metrics and drill-down reports to pinpoint problems by country, subnet, hour and individual meeting
• The per-user search experience for troubleshooting specific complaints and meeting IDs
• How to present CQD data to network and security teams to drive action
Thanks to Barco, this episode's sponsor, for their continued support of Empowering.Cloud.
Victor Guzman: Looking at it high level, start from there, make sure you, you cover all the steps on the checklist. That will take you almost to the finish line, and then start to, to dig into users or particular, locations, particular subnets, and that will get you through the rest. That's really how we, we work with this.
Tom Arbuthnot: Hi, and welcome back to the Teams Insider Podcast. This week we have a deep dive on Microsoft Teams call quality and how to diagnose it. We look at the Microsoft CQD and QER Power BI reports, and Victor from Microsoft takes us through exactly how to drive those reports, where to look, how to understand what's going on in your environment. Many thanks to Victor and Siunie for joining the show. Really appreciate them joining and sharing with the community. And also many thanks to Barco, who are the sponsor of this show. Really appreciate all their support to the community.
Victor Guzman: On with the show.
Tom Arbuthnot: Hey, everybody. Welcome back to the show. A few shows ago we talked about some of the advances in CQD, and one of the conversations came out of that was there's tons of value to unlock in call quality dashboard and, and digging into the data. And Victor promised to come back and give us a bit more of a tour. So we're gonna go deep in this one on, on how to use CQD to f- find and, and understand issues. We'll just do intros. Siunie, do you just wanna give a intro about you and your role?
Siunie Sutjahjo: Hi. Thanks, Tom, for having, Us. Yeah, my name is Siunie Sutjahjo. I am, I've been in Microsoft, for a while, and some of you might already know me, and I had the oc- opportunity to grow alongside a lot of the technology that we now take it for granted. In this case, it's real-time communication, within Microsoft, which is the experience behind Microsoft Teams and Copilot Au- Audio/Video, and also, Windows, Cloud Windows. So a big part of what I do is, related to supportability and debuggability, using telemetry diagnostic and increasingly the AI to help us understand what's happening when, an experience doesn't work the way, it should. So, today we are, like what you said, Tom, we are going to focus on, CQD tool and how we always partner with the customer escalation team, helping our customer listen to what you guys feedback, our customers' feedback and, make an improvement on the tools itself. In addition to that, together with Victor and, also this team is we, we actually help them, we help our customers to ef- effectively and efficiently use our tools. With that, like, I guess Victor?
Tom Arbuthnot: Yeah, Victor, give us a, give us a little intro. And also, you're, you're on the front line with this, with the actual customers, aren't you? So I don't think there's anybody better to take us through it.
Victor Guzman: Yeah. So my, my role is I'm a, I'm a Technical Program Manager. But my team focuses more on, on customer-facing engagements. We focus on escalations and, just, just working with customers. And, and for a long time we ran the Get to Green Program, in Microsoft. So that was a program where we actually reached out to customers that were having quality issues. We, reviewed their quality with them and make sure they get. They were in a good path to, to get to best quality possible. So my team is actually the, the ones that created the CQD templates, the QER templates. I, I was not the one that created most of them, but I, I worked and collaborated on, on the templates for, for a long time. So for, for I would say like five years, my job was looking at those templates and making sure that customers, got helped.
Tom Arbuthnot: Yeah. That Get to Green program is, f- fairly well known in our space for like, yeah, bad things happening, pull in the, pull in the A team to help us understand it, and that was, you and your team.
Victor Guzman: And, and actually this, this presentation that I'm sharing is, is the deliverable for the Get to Green program. This is. This was the first, like, our, our main, presentation that we started, like, when, whenever we got an escalation or we got a customer to work with, this is a presentation that we created. This is the Teams quality review, so some of you might already have seen this, for, for many it might be, the first time they, they look at it. And, and really we started with the quality checklist. We reviewed the health of the tenant using the, the same QR templates that, that customers have access to, and then we created a list of, of, recommendations based on our findings. But honestly, we, we started with this checklist, which is just seven steps that, I, I think 99% of the issues that we worked on were solved by just following this checklist. So it was the. We faced some specific escalations from customers where they had some very, very specific pain points and very specific issues, but for most of the customers that we helped, this, this is really what we did to, to resolve the issues. And I will say this will get you 99% of the way there, and that there might be some, some additional things that you have to resolve to really, to really solve the issues, but this is, this is what gets you to a healthy-
Tom Arbuthnot: And this is great to understand 'cause this is not a one and done. Like, things change over time, networks change, of- you have office moves and different bits and pieces, so understanding how to get into these reports and understand where the issues are is, is, is great 'cause y- you wanna be doing this on some kind of cadence ideally.
Victor Guzman: Yeah, and, and also user behavior changed a lot. Like, when, when we started doing this, everyone was in the office. No one had, had a use for video. Most of the users were just using audio, and they were sitting in the same place in an internal network that was managed. Then we got the, the big work from home after COVID. People were working from home. No issues with UDP, but a lot of issues with VPN, for example. So we started seeing different behaviors there, more usage for video with hybrid meetings, more things like that. And then we got another change where people returned to the office and, and they were having better experience at home than in the office because the office was never prepared for the amount of traffic for video and all that stuff. So I, I think every time the user behavior changes, maybe n- right now it's more going back to the office still and more movement towards that, less hybrid work. So you need to keep an eye on it and, and make sure that you periodically review this and, and, and you're on top of it. Yep. So really, I, I will go through just these seven things in the checklist at the start, and then I'm going to jump to the CQ templates and show you how we do the review and how we validate that these things are, are being, addressed in the checklist. Because that's, that's really most of the work that we did during, during our QA engagement. But really the first thing is we want to make sure that the ports and protocols are open. We don't use a lot of, ports, and we don't use a lot of subnets, but we want to make sure that traffic is allowed there. And we did a lot of consolidation, so now we're, we're down to three subnets. Make sure you check the, the latest on the article. But it, it's just three subnets where all our media services are hosted at. And really one, important port, or a set of important ports for UDP, one for TCP, four for three for TCP, thirty-four seven eight to thirty-four eighty-one for UDP. Some people think that those are mapped still to, to media. It doesn't work that way all the time, so in many cases, we just send media, all media types to thirty-four seventy-eight. In some configurations, we do send it to different ports. So depending on how you're configured, you might see different behavior in ports, but just make sure that this is open to the correct subnets. The second one is bypassing proxies and deep packet inspection. We have a lot of customers that send all the traffic through proxies. Proxies are not sized for real-time media. They enqueue things, they do deep packet inspection, they break open encryption, and honestly, they won't be able to do a lot with your real-time media be-besides just delaying it. They won't be able to decrypt it. They won't be able to, to do anything. The deep packet inspection is just going to de-delay the packets, and we see a lot of quality degradation when you're using proxies. And the other thing is that your proxies are going to get saturated, because of the amount of traffic that is sent on during a video conference, for example. They're not sized for that The third one is split tunneling. We don't want you sending media traffic through a VPN tunnel because that, that's, that basically breaks the benefit of using UDP instead of TCP. We want to use UDP as much as possible because it's just you send and forget. We don't have any controls. We don't have retransmissions. We don't have that. Either the traffic arrives or it doesn't arrive, and the protocol will build with missing packets. If you move to TCP, then you have retransmissions, which makes it worse because now we have to wait for those packets to arrive to process the, the subsequent packets. So it's, it's really bad. It really degrades the media, and we have s- implemented some ways of dealing with that, but, but really you don't want to use, TCP. And putting it in a tunnel is just going to mean that we are going to basically package all that UDP traffic into TCP traffic, send to a specific location, and then send out. You're going to add latency. You're going to add jitter. No real benefit from security. And this is just for the auto traffic. Y- you can still do VPN for, for your, transmissions, for your UD- for your TCP traffic, for your web traffic. That doesn't matter. It's just for the media traffic Fourth one is local DNS resolution. We have a lot of customers where they, they have offices in different places. They resolve DNS centrally. So for example, they have users in Brazil, they have users in France, they have users in China, and all of them resolve, DNS names from the US. So you end up going all the way to the, US for services, because we use DNS to detect the services closer to you. There are several types of things that we do, but that we actually use, geo-based DNS, name resolving to try to load balance the services and, and get you to a correct location. So if you resolve DNS from a specific location, your meetings might be hosted far away from the user, so you don't want that. You want to minimize, the, the latency to where the, the meetings are being hosted. Take the shortest path to the internet. Just send the traffic as quickly as possible. We have customers that pack all the traffic to a central location, even when they have, local egress in home offices because they, they were used to seeing the benefit of just putting everything in a central location before sending it out to the internet. You don't want this for media. You want media to, to egress as, as quickly as possible. Some customers think that their networks are great and that, just, just sending the traffic to, to a central location in the US for users in, in Latin America or something like that is going to be better than sending it outside. It's not. Microsoft has one of the best networks in the world, and we treat the traffic for audio with preference. So you want the traffic to hit the Microsoft network as quickly as possible. You, you don't want it to wait.
Tom Arbuthnot: Mm-hmm. It's not really going on the big, bad internet for very long. It's going to your provider and then probably jumping straight onto the Microsoft network in most cases.
Victor Guzman: E- exactly, and that, that's what you want. You want your traffic to egress to the internet as quickly as possible. Go to a, to a provider that is paired directly with Microsoft, because that means that our traffic is going to hop. Probably it will have one or two hops in, in the public internet, and that, those hops are really you to your provider, your provider to us. Yeah. And that's it. Your traffic is not going with all the public internet traffic intermixed. It, it's usually not like that, unless you're working with, with a small provider that doesn't have a direct peering, connection. So yeah, definitely send it as quickly as possible to the internet. We'll take over from there. We'll make sure that traffic reaches the correct destination as quickly as possible. Mm-hmm Quality of service, only if needed. Really quality of service helps a lot if there's congestion or there's the possibility for congestion. It's not going to do anything if there's no congestion. It's not going to do anything for the user sitting at home. It's just going to help the user sitting on your network if there's congestion. We kind of backed away from, just suggesting quality of service for everyone because we found a lot of customers, had issues with conf- path configuration for quality of service, where instead of helping, it was actually blocking traffic even though there was bandwidth available.
Tom Arbuthnot: Yeah, they, they would have, they would have a, Why
Victor Guzman: Get into that-
Tom Arbuthnot: A enhanced stream and then a drop rule after they maxed out their enhanced stream. So it'd be like, it'd be great till it's awful.
Victor Guzman: Exactly. That, that's what we saw, like customers just sending, a fixed bandwidth even though they have more, and then just dropping packets because we were-
Tom Arbuthnot: Yeah. It ma- it made, it made more sense during the, the, the Skype server days where everything was on network and the WAN was back hauling to the MCU. There was some logic there, but now we're all going to the cloud. There's no QoS on the internet either, so you really, it just is your internal, probably local network if you're breaking out locally on the internet as well.
Victor Guzman: No, i- if your WAN links are saturated or you have, potentially saturated during portions of the. Yeah, definitely, we want to deploy quality of service because it will help in those situation. It, it's just if you have a good network, deploying quality of service is not going to do anything for you unless there's congestion. Mm-hmm. So, so you need to, to make sure that it's properly configured if you want to use it. And finally, the, the final step, which is really important, is exclude the Teams processes from antivirus and DLP scanning. A lot of these, antivirus and DLP scanners are just interrupting the operation. They see a lot of traffic going through the network, they want to scan it, and we see, especially with, with some providers, we saw a, a, a really, really big impact both in performance and quality. We have worked with a lot of the antivirus vendors to make sure this doesn't happen, so it's, it's less important now. But it, it, it's important to exclude the Teams CXC from, from antivirus and, and anti-malware because it's, it's over-fronting, right? So it, it's just going to impact, your performance if, if you don't do it. And, and we have a lot of, of guidance on this. There's a Teams security guide that you can use if, if you want to see why we consider this to be, a, a good trade-off from performance versus security. You're not going to lose a, a, a lot of security by inspecting media traffic, for example. There's, there's nothing to gain there. You're, you're only opening the, the encryption looking at media traffic that you still cannot use. So yeah, do take a look at the security guide if, if you want to, to understand why that, that's probably not, not impacting your security posture.
Siunie Sutjahjo: Mm-hmm
Victor Guzman: Any- anything,
Siunie Sutjahjo: Else, Tom, on, on the checklist? I don't know
Tom Arbuthnot: No, no, I think that's great. I think that's a good. That's already a good kind of start of a 10. These are the fundamentals. I guess the interesting thing is, is how do you identify some of those issues, which I guess is where the, the reporting comes in.
Victor Guzman: Yeah. Perfect. So let me switch over to the, to the templates, and this is the latest version of the QR templates. So this is. I'm, I'm using 5.3. Make sure you're using, five, Really five or later, you, you want to make sure you're using the latest version of, of the QR templates because we did a lot of changes with the classifiers. So if you're using old versions, you, you don't get the benefit of that. So yeah, make sure you, you have the latest one.
Tom Arbuthnot: To get everybody up to, get everybody up to speed here so they can see basic reporting in the Teams admin center, but this is a separate set of reports they can download, connect to the cloud back end, and kind of pull and slice the data in different ways.
Siunie Sutjahjo: And in addition to that, there are some reports that, like, provide details reporting on the event, meetings held, and then also VDI, which is not available, in the admin center at this moment, and if you are, if you have to troubleshoot, for those two features for your tenancy, then, this is the, the place that you should try to dig in.
Victor Guzman: And the other important thing is that this, this, QER are, templates are templates for a reason, right? This is just what we use to troubleshoot. This is the same template that we use when, when we run the Get to Green program and these templates implement the feedback that we receive from the customers. So we have. You- you're going to see here that we have 45 different pages of reports that we have created over time. Probably most customers are not going to use all of them, and I, and I'll walk through the main ones that I use in, in a second. But really they are templates. So we expect customers to get into this, edit the templates, modify the visualizations to fit their needs and fit the look and feel that they want, make sure that they tweak them to show the, the things in, in a way that they understand better. So it's, it's just a template that you can modify. We distribute it that way intentionally.
Siunie Sutjahjo: That's right, Victor. Just to remind people, this is the reason why we are heading towards the Power BI direction, is because, like, some of the customer actually come to us and say, like"Hey, we don't care weekend, so how can we remove reporting on the weekend because impacting, like, the overall, our quality percentage?" Well, now you have your own report. You can actually customize it, and you can even not just personalize the, the report itself, but you can actually customize based on your, based on, like, what you, you think it's, important to your customer, or to your, tenancy.
Tom Arbuthnot: Yeah. I, I, I've seen a variant of that where, like, we only look after North America, so the rest of the world, we can't do much about, so that's not of interest to us to report on for our score, that kind of thing
Siunie Sutjahjo: Yeah. And then, like, you can actually, like, s-, for example, Microsoft has admin that's managing, like, North America, which is different than the admin for Europe, and they can customize it to just, like"Hey, I just w- I am, I'm taking care of Europe, so, I don't need, like, the whole world to do it, take a look at the, whole, quality." And, you can keep it and save it, and then when you come back, it goes back to where, where, your preference is.
Victor Guzman: So yeah, we have a lot of customers that, that, edit these templates and, and just focus it on a, on a single subnet for, for a specific admins to look at, or a single office, or the single country, or y- y- you can do a lot of things with the filters. So it, it, it's just a- It, it's a great way to customize the templates what you're, what you're comfortable with. But really when, when we start looking at, we use the vanilla templates, I always start looking at this particular, report here, which will give you a pretty good indication of how you're using Teams and media in your environment. So it has a lot of information about the usage. So I, I chose a tenant that it's kind of smallish. We're talking about seven thousand, students per day, so it's not a big tenant. But the-- but it's significant enough so that we can see some, some of the data. It actually gives you, information about how media is distributed, so I always take a look at this because i-if I know that the customer is, is mostly using audio, then I will focus my review in audio. If, if I see, for example, this customer is mostly audio with only ten percent of the sessions having video, but they're losing a lot of screen sharing, so probably the screen share is something that I have to look at in, in more details. The other breakdown I like is this one, type of network that they're using. Most of the customer-- Most of the users for this customer are using Wi-Fi, which is very common across all customers. You have wired thirty-three percent of the connections, and then a lot of mobile users. Twenty-five percent of the users are mobile for this customer. And also, they're not using a lot of, peer-to-peer. So ninety-four percent conference, but five percent peer-to-peer. Usually, the breakdown that we see for most tenants is around eighty to eighty-five percent conferences and twenty to, fifteen percent, peer-to-peer. So this is, this is on the, on the lower side. And then you can see the trend over time, things like that, just, just because it's important to see if, if your usage is increasing, maybe your network is stressing because, because your users are using it more, things like that. So it's important to look, look at it that way. The other thing that I like to see here is the meetings hosted by country. This give you a good indication whether you're resolving DNS correctly or not. So if, if you're talking about the checklist, this is one indication that, that things might not be working as expected. So for example, if, if your tenant, is located in, in the US, and you see this distribution with France, Ireland, and Netherlands being the first three and the United States being the fourth- That means that there's something wrong with the configuration. Either your DNS is not working correctly or your geolocation is not correct.
Tom Arbuthnot: It should usually reflect your populations of users, right? Shouldn't it? Obviously, it depends on who schedules and who joins first and that kind of thing. But like broadly speaking, if your majority of users are in Germany, you would expect the majority of meetings to be in Germany.
Victor Guzman: Exactly, or at least in EMEA, right? For example, this customer, majority of the users are in Egypt in this case, and it's located in France, which is correct. That's the same region, right? So this is the first one I look at and I just take note of things here like how much is Wi-Fi, how much is wire, how much is conference, all that stuff because that will give me insights into the next reports. The second report I look at is this one, which is the main dashboard for health and this will give you a high-level view of how media is behaving for your tenant. So this is, for me, this is also a report that admins should be checking every day or at least on a weekly basis because this is what will give you the health, overall health of your tenant. But really we have a breakdown here Poor audio rate, poor video rate, sharing rate, drop failure rate, setup failure rate, and the feedback rate. The first f-, five, years here actually show you, objective measures, while the last one is just user perception. So all of them are, are important, but depending on, on the region and depending on the user's willingness to provide feedback, this, this last one is going to be less useful or more useful. For example, users in Japan are really, really strict, so it's very common to see this poor feedback rate being higher than the, the problems, than the problem rate. There are users in some companies that are not willing to provide feedback at all. So like, for example, this one, th- this is a, a bad one, right? They have one or two surveys per, per day, like five here, two here, no surveys here, and they, they're not giving bad feedback, but the number of surveys is really, really low. So it's, it's not going to be very useful because your users are not answering the, the s- And,
Tom Arbuthnot: And that's something you can educate your organization on as well, 'cause often they think that those surveys are just going off to Microsoft. They're like"Well, who's gonna pay attention to my call in amongst three hundred and fifty million users?" But actually if, if IT can teach the business"Oh, no, we can see that those reports and what's happening, there's a reason to fill them in" that might help.
Victor Guzman: Exactly, and, and that's one of the things that when we do the catering, we, we stress that out. Like you want the feedback, but you don't just want the feedback, you want the users to know that you're seeing that feedback and that you're actioning on them. So one of the things that we usually suggest is go, go to the feedback, report, look at users that are giving the poorest feedback, and reach out to them because you will learn a lot when you reach out to them. And, and when you solve the issues, they're going to be happy, they're going to continue to provide feedback, and you're going to improve the, the quality, we had some, some cases where the feedback was really, really bad and it was just because of the device they were using, the, the headset or, or something like that, or, or they had a really bad problem that was not related to Teams at all. In some cases it was just cabling. They needed to change the cables because they, they were, damaged. The, the connectors in, in one case it was just the connector that was damaged, in, when they, the user was plugging in their machine. That solved it. That user was happy, and yeah, they, they solved the issue. They was a VIP user, so they were giving poor feedback and they were not just giving poor feedback here, they were giving poor feedback in meetings and impacting, leadership meetings with their poor quality. So it, it's important to look at that.
Siunie Sutjahjo: Yeah, I just wanna add it, like, talking to the customer, some of the admin actually use this feedback to actually track. So whenever there's user that's giving one, star or something, they actually, get back to them. Like, and, and they're responsible to actually make sure that, like, anybody who get back to, anybody who report, like, one star, the problem get addressed and resolved, and that's, like, their commitment from their, admin team. And that way it's actually, they're also using, like"Hey, there's, like, one person who's reporting this."they're, like, actually finding that one person who's complaining this room is too hot, for example, is actually impacting, like, other people in the same room as well. They just quietly have not- Yeah,
Tom Arbuthnot: Just hadn't given the feedback.
Siunie Sutjahjo: Yes. So they're, like, not, like, not willing to give feedback.
Tom Arbuthnot: Yeah, everyone, everyone in the Barcelona office has rubbish Wi-Fi, but just one person has the diligence to keep reporting that
Siunie Sutjahjo: Kind of thing. Yes, yes. And then, like, they, they figure out, like, the problem, and then they're wondering, like"Oh my God" like"How come, like, nobody is screaming?" Because it's, like, it's, it's impacting 100 people in this building, but only one person. So those are, like, I also encourage our, our tenant admin to kinda encourage their users, maybe, like, give them a Kit Kat bar or something, if they do a report.
Victor Guzman: Now, if, if we go through to this, and, and I chose this tenant because of, of exactly this configuration, you can see that they're borderline healthy in audio, kind of elevated for poor video rate. You want this to really be below 3%. That's what we see for a global company. Some customers might be more stringent and think that"Hey, we, 3% is too bad for us. We want to shoot for 2%." That's pretty good. Even 1% might be a good target for, for a managed network or where users are mostly located in, in places that have very good connectivity. But for global customers, we use a 3%, target So anything below 3%, we're going to mark as poor. This one is a little bit above 3.36%, so a little bit on the high side for all of you. Higher for video, 4.98%. And then poor sharing rate, it, it's almost getting to a red state, so definitely something we need to look at because as, as we saw on the first page, they are heavy, users of, of sharing, so definitely something important There are two metrics here, drop failure rate, how many calls drop in the middle with no, no particular reason. This is really high, 4%. Again, we want to shoot for 3%. And the highest one is this one because 3.58% is, is already in the red because our target is less than 1% of the calls should fail to, to connect. So this set of failure rate is just how many calls are failing totally to connect. The user is not even able to use the product. For this tenant in particular is, is, is very high. It, it's almost 3.5%, so three times, more than three times as much as we want. And you can see that the trend is all over the place, so this, this is not a new issue. Many times we see that these things are, are, on a, on a trend, like either it started on a particular date and that might indicate a network change, or they, they have been trending up for a while, which might have to do with users. So these graphs tell a story. In this case, it, they are all over the place. It's just jumping, but it seems to be like, I, I think they are having this problem for a while. It, it doesn't seem to be, going away. Now, this one, for example, is interesting. Here it's almost 6%. It usually is not that bad, but they are always above the threshold. So definitely something we, we need to look at for, for this tenant, right? And we do have reports for media setup and media reliability that, that we can use to, to just look at particular networks where this is happening, particular machines, particular users. I'm not going to, to jump into those, but you can, jump into each one of them, like for some media setup, media reliability, and then we have breakdowns for health, audio health, video health, sharing health. So you can jump to those, and, and look at the details. But usually after this, I, I, what I do, I, is I start looking at the checklist. One of the things that we're talking about is just making sure that you can reach the network, that you have the right ports open, and that you are bypassing proxies and, and Deepak inspection. And the transport report is the one that we can use to, to get that information. Let me open it here.
Siunie Sutjahjo: So while it's loading, I just wanna mention, so at the top, the graph is actually showing yellow, red, green, yellow, and red, right? So you can see like for the share, for, for one of them is, it's 5.88, and it's very close to red. So take a look at those, chart really, really well because even though it's currently yellow, but you are very at the edge of getting to red. So you might wanna jump in and try to tackle those kinda issues. Anything that's like almost borderline to red or it's already in red. And even when it's green, if it's almost borderline to yellow, those are like the reason why we put like those kind of visualization Instead of just yellow, red, and green
Victor Guzman: Yeah, and, and again, those targets are, are customizable. So if you feel like 3% is not, not what you want for, for yellow, you, you can customize it. We see customers that have more stringent targets. Again, for our global company, our users all over the world, users working in countries that don't have a, a good connectivity, we think that a three is, is achievable. That's why we chose that. But for if, if, if you know that you can, yeah, that for you two is, is, is a good threshold, you can definitely use that. Mm-hmm Now, i-immediately jumping on this one, I start seeing some, some things, right? We're talking about the, the, making sure that you could reach internet. This is not the case here. Only twenty-three percent of the traffic is actually going through UDP, with most of the traffic actually using compound TCP, which is the worst kind of TCP that you can use. That means that it's going through the proxy. So this is using HTTPS proxy, not even just TCP, it's HTTPS. And, and you can see here, like, remember we were talking about a, a six percent, rate, for all media or actually for, for the worst media, like five point thirty-eight, something like that. The inbound rate and the outbound rate for compound TCP users i-is actually way worse. It's seven point twen-, twenty-seven percent, seven point twenty-four for the outbound. So this is, this is where, you will see most of the, the users, having, having issues, and you can see why. Like the jitter nine hundred and seventy-three, average network jitter max, average network jitter thirty-seven. You can see packet loss, round trip just- Big peaks here. So definitely this is where the majority of the users are. This is, this is the main issue that the customer is-
Tom Arbuthnot: Yeah, it's really nice to see how, how well it's visualized here. Like when, once you've got the data in, it's like, well, there's some obvious things to jump to that are, this is not right, and this is fairly typical of CQD, like we're going to be engaging with the, the network team. The- there might be local teams for the LAN and different teams for the WAN network, and might be, as you said, desktop team for the antivirus or it could be the security team. So with this kind of project, so I'm sure you've seen this a lot in the Get Screen, you're, you're involved with multiple different, stakeholders in the organization.
Victor Guzman: Yep, and, and this is something that, yeah, you need to work with security, you need to work with, the network admins, you need to work across the organization to make this change. But usually with the data, it's really easy to convince them, like"Hey, look at this. We, we know why you're being impacted. It's, it's not. We-- Let's just start with this." And for example, here you have BASN name, associated with that traffic, and you can see the, the worst performing. You have, TE data and Vodafone data in Egypt, and those are the two, that are affected, and with the first one being the worst, right? The other thing is that, as, as you see here, this, this gives you a lot of information about how many users are connected, how many streams are going there, what, what subnets in, in my network are, are in use. So all this information should let you, go to the admins, understand what's happening and, and resolve it. Let, let me talk about this for a second. Just, these subnets, they're really useful, but many times, it's not as useful as if you have the network names. So we encourage all the customers to import the building data into CQD. We know it's a lot of work. We know it takes some time, and it, it's really hard to get full coverage. But as. If, if you start doing it, you really see a benefit when using the reports because you can see the names of the networks and not just the ten eighty two three two zero, which doesn't tell anyone anything beyond the, the, the network admins So definitely something you, you want to do. So this is one of the issues that, that we see for this customer. Let's go to estimated VPN because that, that was the next step on the checklist. So we covered step one, ensure the right ports or protocols are open, and now step two, bypass proxy. Let's look at the, at, at, at the VPN.
Siunie Sutjahjo: I wanna add, to from the earlier report as well. So the impact to the TCP is not always the same for different modality, whether it's audio or video. So it's also good to check, like if your, if your TCP is impacting a lot more audio and your tenant is actually used a lot more audio, then better hurry to, to take a look and fix it. And then some team actually don't, focus on using video that much. So if the TCP usage impacting the video, maybe, you wanna get, to solve the problem, but not like as in a hurry as if it's like impacting both audio and video or audio only.
Tom Arbuthnot: It's interesting to know if it's, it's a cultural decision or a bandwidth decision. Like lots of people will back off video if they think they've got dodgy internet. They're kind of trained to be like"Oh, okay, I'll drop video. I won't do video." So maybe they want to, and they, they're not doing it because the network can't sustain it.
Victor Guzman: One, one of the important things with, TCP is that we have different types of TCP. So you will see TCP Turn, you will see TCP Multi, and you will see Compound TCP. Compound TCP is the worst kind of TCP because that means that it's going through a proxy. Turn TCP is the, the next worst, and then multi, multi-TURN TCP, it, it's a little bit better. But always, TCP your, your quality over TCP is always going to be less than the quality you get over UDP with more bandwidth usage. So you want to move away from TCP as, as much as possible. And really the only way to do that is open the ports and make sure that UDP is allowed. That's really it Okay, so jumping back to, to this one, this is the estimated VPN, usage, and we have two reports for VPN. We have mapped VPN, and we have estimated VPN. I always just start with estimated VPN because mapped VPN requires the admins to m-- to at least upload, and mark the, the VPN subnets on the building file. Since most of the admins are not doing that, we, we, we cannot use it, confidently. But just, just looking at the, the estimated VPN, which really what it's looking at is just the subnet mask, and if it sees a subnet mask that matches the IP of the user, so basically a subnet of one, then it, it marks it as a potential VPN. So this is potential VPN. We're, we're just trying to guess if you're connecting for a VPN or not, but it's pretty accurate. And w- for this customer, we immediately see something. Look at the poor rates for users where the VPN, estimated VPN is one. So we're talking about not-- users not on VPN having a six percent inbound rate and a six, point fifty-five percent outbound rate. For users on the VPN, it's twelve percent and twelve point five. The sell rate for the users not on VPN is, is zero point three one, which is below the target. For the users on the VPN is two point twenty-six. So definitely we see a big impact of VPN for this user. It's not that they have a lot of VPN sessions per day, but the users that are on those, VPN connections are behaving really badly. They're having a lot of, issues, and probably they're impacting other users. So definitely something you want to take a look at. We want to make sure that you bypass the VPN as possible, as much as possible. We have seen instances of users having all their, their VPN concentrators in the US, having a user in India connecting over VPN to the US, talking to another user in India that it's also connecting over VPN to the US and have a latencies of two thousand, three thousand milliseconds. So it's really impactful. Even when the VPN is working correctly, you're, you're going to see big increases in, in jitter and big increases in latency. If the users are in the same, country as the VPN concentrator, if your VPN concentrator is sized properly, you might not see a lot of impact. We have customers that decided that they want to continue using the VPN, and that's fine as long as your VPN is sized properly and your concentrators are really, really close to your users mm-hmm. But for most customers, you, you want to just get away of VPN. And, and not just your VPN, also, for example, any type of third-party, security services, Zscaler, Cisco Umbrella, any, any type of those, of, of, of those services that you're using, do not send traffic to them. Do not send the media traffic to them. Send all your web traffic, that's fine, but you don't want to send the media traffic because we have seen a lot of instances where we are seeing the gradation of, of, audio and video because you're sending all your traffic to those providers, and they have to egress from a single, egress point. So it's, it's, it's not ideal. It's only going to increase the latency, and again, there's really, really low benefit that you get from sending traffic that way. Now, local DNS resolution, we already talk about that one in, in the first report, but I, I really like, the way it's shown in this one. This is one of my favorite reports, which is the weekly metrics, because th- this will give you raw access to, to the metrics over time. Jitter, latency, packet loss, just your basic network metrics is really good when you want to have a conversation with, with the networking guys. This is a report that I use the most, and it gives you also a lot of insight of where your users are, are coming from, how they're connecting. It helps when you're troubleshooting and you see quality for a particular. Maybe even a particular firewall is having issues. This will show you that because you can see the public IP, how it's behaving, what, what the, what's the score for, for traffic going across that particular network So it, it's a really, really nice way of seeing things, and gives you this nice view over time. So y-you can see, for example, this customer, almost all days they, they have issues. You have the breakdown by country, so you can actually select, for example, where the users are coming from. In this case, T-Data. You can see, where those conferences are being hosted and see if it matches that particular location where they s- that the ASN is located, and then, see the public networks that they're using. So you can see a breakdown, see if you have, some, some particular subnets being affected or not, and see all the metrics here. So for example, in this one, when we, when we go this way, you can see that again, we, we have some days where things are, are really bad, or most of days things are really bad. There are some weeks here where they, they actually didn't have any problems even though they have a, a higher amount of usage. They're all coming from, from a single IP, so it doesn't seem like they have multiple firewalls, load balancing or anything. In many cases, we might see, like, five different firewalls coming from a single subnet and two of them having issues and the other ones not, or, or one of them having issues. That will tell you that maybe there's one firewall not working. But the other thing that I like on this report is that you can actually do a drill down and see a particular week in detail So you can just click on the arrow, click on the bar that you want, and this will take us to, to the next level. So you will be able to see day by day on that one. So you can see if it-- if this is something that's happening every day, if it's something that happens during a particular day. And then you can actually go even deeper and, and go to a, to a hourly view. And
Tom Arbuthnot: That was the things, things we used, we used to see back in the day was things like, you know, there'd be a site to head office backup run over the WAN or something. So you'd see every day at X o'clock it would get saturated. And i-if you can find that kind of correlation is a particular day of the week or a particular time of the day, that's interesting.
Victor Guzman: Y-y-- and, and we see that a lot right now, right? When, when people are going back to the office, like, everyone is in the office on Wednesdays, but they're not in the office on, on Fridays, and you see behaviors like this, where, okay, it's Mondays are really bad. Let's, let's look at Monday by hour. And, and you can start to see patterns there where maybe there is some congestion, maybe there is not. But it, it's, it's really nice to see this in a, in a breakdown by hour, by day. It also let you, see if, if some particular users on some regions are having problems. Like, for example, maybe you see a lot of issues during, nighttime for the US because your users from India are connecting to the US.
Tom Arbuthnot: Oh, yeah. Very interesting. Yeah. Issues, lots of issues out of hours. Yeah. Bad routing. That's a good tip.
Victor Guzman: So you can see, for example, this, and it doesn't seem to be load related. You, you have path quality all day long right here. But it lets you see this and, and see how this, packet loss is increasing. And you can do the same with any of the graphs. So it's just a nice way of, of looking at different days, different times, different hours, and start to, to see what's happening. You can even go down to specific, meetings here. So if you, if you ke-click on the hours, it will actually take you to the, to the conference level view, where you can actually see the conference IDs and see how many conferences were affected during that time. So it's. I, I really like this report. It will tell you a lot of information about, saturation. If you are having issues with a particular egress, you have a lot of information that will help you, know if, if you're really seeing the shortest path or not. Like, for example, there were only three meetings here. One of them had thirty-one percent packet loss. The other one, a max packet loss, twenty-seven percent max packet loss. And this one was really, really bad. It had ninety-ni- ninety-eight percent packet loss at some point and fifty-six average. So definitely something that you can look at, take a note of this, ID for the conference, go look at the details, see where the users were connecting from, all that stuff, and that, that really takes you to the M- to where you want. But, but really these are the reports that we use to, to see if you're meeting with those, those things in the checklist. I, I think with these reports you will have a good idea of where you're, you're in a good path, to, to complying with the checklist or not. And as I said, if, if you comply with the checklist, you're going to get 99% of the way to, to resolution of any quality issues. This is our, our search, search experience, which is the, the report that many times you use when you're troubleshooting, specific issues. So remember we mentioned that, that you want to look for a specific meeting. This is where you will come with that meeting ID, or that conference ID, search it here and find the details so you can see what, what particular thing happened during that one. And you can search by UPN. I- in this case, I don't have access to, to the, to the UPN for the user, but customers do have access to, to, to their own UPNs. We as Microsoft won- don't have- This is
Tom Arbuthnot: Quite powerful now 'cause it can be, you know, Sarah said the all hands last week was bad. Was it w- was the all hands bad or was it bad for her? We can look up her as a user and see experience. We can also look at the, the meeting or the event. So it's not just about global packet loss trends. We can get into specifics of individuals. Again, providing the customer enables the PII and providing the right person has the RBAC to get all the access to all that good stuff. But like if the admin has the permission, they can go in and see the actual user's experience.
Victor Guzman: And the other thing that you can do here is search for a particular user, a particular subnet, a particular public subnet even, filter by, by a conference ID or a particular meeting ID and then do a zoom in. Like for example, let's do one here. You can do a drill through to the meeting health detail report and that will give you a lot of information just, just for that meeting in particular. Why was that? Who was the user that dropped? What was the packet loss for the users? Where were they connecting from? All that stuff. So we have a lot of information here that is really relevant.
Tom Arbuthnot: Yeah. It's helpful if you get those scenarios as well where someone, someone reports they had a bad meeting but their connectivity looks fine. Actually it was the person presenting had the bad connectivity, but their perception was they had a bad connection because the presenting party had a bad connection. So this can help you deduce what actually went wrong rather than kind of anecdotal reports.
Victor Guzman: Exactly, and that's why we have this dominant participant, right? Yeah, I think we talked a little bit about that on, on the other call, where the new classifier, it's actually taking into account who the dominant participant was in the meeting and if they have, issues, they c- N- th- those, metrics actually reflect on the other participants so that we know that everyone was affected.
Siunie Sutjahjo: Yeah, so that's, the, the things that, like, we are trying to make it easier for our admin, but with the intelligent classifier, we actually help point out who's the dominant speakers and if that dominant speaker has certain kind of issue. But one thing that, like, like, in the past when looking into, like, some. E- even internally, we have all hands, and then I was curious, and, I can actually see, what is actually a certain kind of, like, issue, with a certain kind of user being kinda, who's actually a presenter, but actually having the issue. And apparently, time-wise, it was not, like, ideal. That person was dropping off their kids in school, so taking the, the call and presenting on mobile is not necessarily what we suggested, but I could actually see it over there and kinda, okay, it's our internal, and we can actually see both, participants' and attendee kind of information. And, another thing that's interesting when we were looking into some of our customer issues is, like"Hey, they are not leveraging ECDN in certain places." So you can actually also see based on, like, looking into, the, event, held through, like, hey, we just have this event held. What is going on with, like, the folks in, say, like, for example, in California. And then we realize, well, that's, because they, we. It's because they're not leveraging the, ECDN, and a certain kind of port is still actually blocked, for that kind of, area. So those are, like, helpful, for, like, a quick kinda way of tying this to the, user feedback. You could actually try to figure out, like, hey, this user giving, like, a rating of one, how many times they'd like you, this person has been actually, and how long has this person been, like, in pain, right? So.
Tom Arbuthnot: Yeah. Well, th- th- this gives the IT admins the tooling to try and find out, 'cause at that, very often you get the feedback, you're like"Well, where do I go?" And Victor, this has been a great session to give those IT admins or service owners an idea of where to start. So just to re-up again, they can go and get these reports from download. Microsoft. Com, the QER and the CTD reports, and then populate them and, and they'll see these and more, in, in the report pack, but these are the great place to start. Mm-hmm.
Victor Guzman: I, I will start with the reports that we reviewed today. This one is also pretty good as, as we mentioned, like for example, this one, this is one of the cases where one user had a problem and it impacted all of them. So you can quickly see this here So it, it, it is also a good example just, for that scenario. But yeah, it's, it's just looking at it high level, start from there, make sure you, you cover all the steps on the checklist. That will take you almost to the finish line, and then start to, to dig into users or particular, locations, particular subnets, and that will get you to rest. That's, that's really how we, we work with this. And, and it's an iteration. You won't get 100% of things resolved first, and as you resolve things, more things are going to bubble up to the top. So perhaps right now it's compound TCP, but when you change to UDP, maybe there's something else there that you need to resolve.
Siunie Sutjahjo: Yeah, and, CQD has, like, the whole data for, like, a whole year, so you don't have to upload the whole year. But then if you wanna see the improvement that you'd, of the quality that you're changing the TCP to UDP, then you can track it and say like"Oh, it's really, really, like, impacting." Another thing that I wanna point out is at the top there is a help, s-, for us to help to get better, so, do, provide feedback. So, we actually, take, the, the, the, the feedback and, your input seriously. And it's not just, like, about, the, the template itself, but if you see, like, a. You would like to see a certain kind of metrics that you deem very important in your daily, troubleshooting, let us know because, for example, we, we actually realized that like, oh, on VDI, the optimize and op- unoptimized kind of metrics is very important. In the past we nev- we don't have it, out, but then we add it, and we added, like, a different flavor of, like, like, metrics as well. And all of this, based on, like, you guys' feedback on, like, debugging and, like, figuring out, like"Hey, we need this information to, to realize that we are not having the VDI client updated or the Teams client updated" things like that. I think, what's missing is, like, from the list is the client version that, but we have the report for it. It's actually pretty important to always, check your, tenancy client version, make sure that you up- your te- And then actually up- has, like, the updated, like, client. You'll be amazed because I have deal with, like, some big enterprise company that has two years old, Yeah,
Tom Arbuthnot: Blocking the client updates is crazy. With this, I, I don't, I don't understand how you really don't, not just from a quality point of view, but from a security point of view these days, so,
Victor Guzman: A- and I will say, we had it in the checklist before. That was big part of the checklist. We removed it when we released the, the new Teams version because we're really good at blocking all clients.
Tom Arbuthnot: Yeah, yeah, blocking all clients and, and, and it's pretty, pretty robust in updating these days unless IT's gone to a real extensive- trying to block it
Victor Guzman: So we still take a look at that, but we don't, we don't pay too much attention to it because you actually get blocked now if, if you're on a really old client, so. But yeah, we, with the old Teams, we had a lot of that.
Tom Arbuthnot: Vixa, this is an amazing tool. Thanks so much. Really important product. Thank you. Any, any s-, Team service owners, network teams, anybody in the larger enterprise. So yeah, and, and I think, Victor, you've helped. Like, it can be quite intimidating. There's a lot there, so if you just start with those key reports, and as you said, pick something. Like, like pick a thing and go and talk to the network team or go and talk to your peers, you can have quite a big impact. It's one of those areas where it's a very measurable positive impact back to the business to be like"Look, we saw this issue and we've improved it by percentage points, and we've got the personal feedback of the people who've raised the ticket and, and had the issue, and we've resolved it." So it's one of those ones where you can have some meaningful wins if you look into the data.
Siunie Sutjahjo: Right, and actually,
Victor Guzman: And some of them are quick wins.
Siunie Sutjahjo: Yeah. And I wanna add one more thing. I know that, like, people sometimes feel, like, intimidated, and they feel like it's a little bit too complex. But I think with, the amount of, like, interaction with other department, like whether it's a network admin or the security admin, you have to provide the evidence. Like, if you're going to court, you need to provide the document that's, like, full of details. Like-
Tom Arbuthnot: It's, it's never, it's never the network team until you can prove it. We all know that.
Siunie Sutjahjo: But that's, that's the reason why we are equipping our, our, our tenant admin with this much of information so you can use it to kinda talk with, like, credibility. Because if it's just like we are saying, like"Network is bad"then like what? They will say like"Prove it." Then- Yeah,
Tom Arbuthnot: Yeah, yeah, yeah.
Siunie Sutjahjo: Then, like, is it packet loss? Is it, like, wait and see?
Tom Arbuthnot: And, and, and I, I, I s- I see the point of view from the network team because they're the common denominator of every network and everything, and like, like, you know, like, they get, they get called out on every single ticket. So often they appreciate when you're engaging with them with real data to be like"Well, now look, here's, here's the actual data, here's the IPs, here's the-" And they like real data. The level of detail. We're talking their language.
Victor Guzman: Having a short conversation.
Tom Arbuthnot: Yeah, exactly.
Siunie Sutjahjo: And they actually wanna know, is it inbound? Is it outbound? Or like. Because, like, they quiz you with all of this information before, like- Yeah. They, they are convinced that they are, they need to do something, right? Those are, like, the reason why we actually provide, like, the numbers and the volume, right? Like"Hey, these are, like, impacting thousand users" or impacting just one users, right? Yeah. Yeah. Anyway, really excited about like this. Hopefully you guys try it and, use, CQD, 5.3, that's the latest one, and provide us feedback. We'd love to hear from you.
Tom Arbuthnot: Awesome. Thanks both. Appreciate your time.