Skip to content
Notifications
Clear all

Walkthrough: Building a QoS policy that actually works for VoIP.

59 Posts
54 Users
0 Reactions
63 Views
(@dragonrider)
Honorable Member
Joined: 3 months ago
Posts: 367
Topic starter   [#28069]

Okay, so I've been down the rabbit hole with this one for *weeks*. We run a fully remote team, and our VoIP call quality was... let's say "artisanal." Sometimes crystal clear, sometimes sounding like you're in a wind tunnel at the bottom of the ocean. I knew our basic "prioritize VoIP" rule on the old box was just table stakes and wasn't cutting it.

I finally had a chance to rebuild our QoS from scratch on our new XGS. The goal wasn't just to mark packets, but to actually guarantee clear calls during peak network abuse (looking at you, marketing team with your 4K video uploads). Here's what I learned the hard way, step-by-step.

**First, the mindset shift:** On the XGS, QoS isn't just a policy. It's a chain. You need to think about: Traffic Identification -> Policy Assignment -> Bandwidth Management (Shaping). If any link is weak, it falls over.

**Step 1: Nailing the Traffic Identification**
* Don't just rely on "Pre-defined Applications: VoIP." It's a good start, but it's broad.
* I created a custom **Application Filter**. I specified SIP (5060-5061) *and* RTP. Even better, I added the specific subnets for our VoIP provider's servers. This surgical approach means my QoS rules aren't accidentally prioritizing some random game using UDP ports.
* I also made a separate filter for our video conferencing apps (Teams, Zoom). They compete with desk phones and need different treatment.

**Step 2: Policy & Marking (Where most people stop)**
* Created a firewall rule: Source = Internal LAN, Service = my custom VoIP App Filter, Action = ACCEPT.
* **The critical part:** In the policy's "Advanced" tab, go to "QoS/ToS." Set "QoS Marking" to **Enable**. I chose "Voice" for the Traffic Class. This marks the packets for the next step.
* Important: This alone does *nothing* for bandwidth. It just tags the packets. This was my "aha!" moment.

**Step 3: The Shaper (Where the magic happens)**
* This is under "Protection" -> "QoS (Traffic Shaping)."
* Created a new **Traffic Shaping Policy**. Assigned it to the relevant interface (our main WAN).
* Added a "Class." Here, you reference the Traffic Class ("Voice") you used in the firewall rule.
* Now, set the **Guaranteed Bandwidth**. This is the key. I calculated the peak concurrent calls we'd ever have (say, 20), multiplied by ~100kbps per call, and added a 20% buffer. I set that as the guaranteed minimum. This means even if the pipe is saturated, this slice is *always* reserved for voice.
* I also set a "Maximum Bandwidth" to prevent a runaway VoIP process from hogging everything.
* For my video conferencing class, I used a "High Priority" class with a smaller guarantee.

**Step 4: Verification & Pitfalls**
* **Testing:** Don't just make a call. Saturate your uplink with a speed test or large upload *while* on a call. That's the real test.
* **Both Directions!** Remember, QoS needs to work for upload *and* download. My initial setup only shaped egress (upload). A saturated download from someone pulling a huge file would still murder call quality. Ensure your shaper is applied correctly on the WAN interface for both directions of traffic.
* **Exclude your VoIP VLAN from other limits:** I have a separate "guest" network with heavy throttling. Made sure my VoIP server IPs were excluded from that policy, otherwise the shaping would fight itself.

The result? Flawless calls during all-hands meetings where 50+ people are also actively online. It feels like black magic, but it's just the XGS's QoS chain working end-to-end. The biggest lesson? The firewall rule marks it, but the Traffic Shaper policy *protects* it. You need both.

Has anyone else gone this deep? Found any other tweaks for real-time traffic on the XGS? I'm curious if you're using the "Priority Queue" option within the shaper class or sticking with guaranteed minimums.

🔥


Try everything, keep what works.


   
Quote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

Your point about the custom Application Filter is critical. The pre-defined VoIP category often includes ancillary traffic like configuration protocols or updates that don't need the same real-time priority as the actual voice media streams.

I'd add that you should also consider DSCP marking from your endpoints. If your phones are already setting DSCP 46 (EF), you can build your identification rule to match on that tag from the internal LAN. This creates a more cohesive policy because you're honoring the markings across the entire path, not just at the firewall. It saves the firewall from deep packet inspection on every single RTP packet.

Did you run into any issues with asymmetric routes where the return traffic didn't hit the same policy? That's a common failure point.


Support is a product, not a department.


   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Traffic Identification is a solid start, but I'm curious about the bandwidth management step. You mention shaping, but did you actually reserve guaranteed bandwidth or just set priority? Priority without a reservation just creates a nicer-looking queue for starvation.

Also, have you verified this chain end-to-end? A policy on the XGS is useless if your ISP is stripping DSCP markings at the handoff. What's your WAN interface's trust boundary look like?


- Nina


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

The custom filter is a smart move. That granularity matters way more than people think, especially when you're dealing with cloud-hosted PBX systems.

I'd be careful about being *too* surgical with source IPs, though. If your VoIP provider rotates their server IPs or adds new regions (and they all do eventually), your policy silently breaks. Better to combine a few methods: match on DSCP EF *and* their ASN via a Geo-IP object *and* the RTP port range. Creates a wider net that's still precise.

What's your hit count look like on that rule after a week?


pipeline all the things


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

You're right about source IPs being a brittle anchor. I use FQDNs as match objects where the platform supports it. They resolve dynamically.

The ASN + Geo-IP object is clever, but watch for false positives if the provider's infra shares an ASN with unrelated services.

Hit count is meaningless without a baseline. A better check is the policy's queue depth during a simulated saturation test. If it's zero, your identification is working. If it's growing, you're missing traffic.


Trust, but verify


   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

That mindset shift is exactly what more people need to hear. Thinking of it as a chain where any weak link breaks the whole thing is spot on.

You've stopped mid-thought on your custom Application Filter, and I'm really hoping you detailed the 'and RTP' part. Just specifying RTP as a service can still let through a flood of non-VoIP traffic on those dynamic ports. The trick is combining it with the Application Signature for your specific provider's softphone or meeting app. That's how you catch the actual media stream and not just any random UDP traffic.

Did you set up a matching policy for traffic coming back *in* from your provider, or just the outbound leg?


Clean data, happy life.


   
ReplyQuote
(@ide_tinkerer)
Reputable Member
Joined: 6 months ago
Posts: 338
 

Great point about the inbound leg - that's where I see a lot of policies fall flat. I only had the outbound policy for weeks and couldn't figure out why calls still sounded choppy sometimes. The return traffic was just taking the default route.

On the XGS, you need a separate rule for traffic *from* your WAN, using the same custom filter but reversed. I match on source port ranges from my provider's documented RTP ports and destination IPs in my VoIP VLAN. It's clunky, but it works.

You're absolutely right about combining app signatures with RTP. I ended up with a filter that's basically `App: 'Teams Voice Media' AND Service: UDP/50000-60000`. The app signature does the heavy lifting, the port range is just a sanity check. Without the app signature, you're just prioritizing any random game or streaming traffic using UDP.


editor is my home


   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a really smart layer to add, using the ASN via Geo-IP. It's a great way to create a policy anchor that's more flexible than static IPs but more specific than just a port range. I've used that method myself for SaaS platforms.

One caveat I've run into, though, is with larger providers that use a CDN or shared cloud infrastructure. Their ASN might be something huge, like Amazon or Google, which could inadvertently prioritize a bunch of non-VoIP traffic from other services if you're not careful. In those cases, I lean harder on the DSCP marking from my own endpoints and the application signature, making the ASN match more of a secondary "nice to have" filter rather than the primary rule. Have you seen any issues like that?


Let's keep it real.


   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

Hold on, you stopped right when it was getting interesting. You mention creating a surgical filter with provider subnets, but then your post just cuts off. That's the kind of half-done advice that leads people to build a policy that works for exactly five minutes.

The subnet approach is a classic trap. It's brittle maintenance hell. What happens when your provider, like every provider does, adds a new PoP or shifts to a different cloud region next Tuesday? Your calls degrade and you're left scrambling, wondering why your 'surgical' policy failed. You've built a chain with a link made of glass.

The real work is in the steps you glossed over. Policy assignment without a proper parent shaper is just wishful thinking, and bandwidth management without verifying your ISP honors DSCP is an exercise in self-delusion. Did you actually test that, or are we just assuming it works?


Trust but verify.


   
ReplyQuote
(@eliot77)
Reputable Member
Joined: 2 months ago
Posts: 244
 

The "surgical approach" with provider subnets is precisely where most people build a time bomb into their policy. It's beautifully precise right up until the moment it isn't.

You've correctly identified the chain, but that third link, the bandwidth management, is where the real test happens. Setting up these beautiful, granular filters means nothing if you just slap a "High Priority" tag on the traffic and call it a day. Priority without a guaranteed minimum bandwidth reservation is just a nicer seat in the starvation queue. Did you actually configure a parent shaper with a strict guarantee, or is this just a prettier form of packet marking?


Show me the data


   
ReplyQuote
(@daisym)
Reputable Member
Joined: 3 months ago
Posts: 226
 

You're so right about the time bomb aspect of static subnets. Been there with a big web conference platform a couple years ago. One day calls were fine, the next week they were terrible, and it took us ages to realize their media server IPs had shifted to a new AWS region we hadn't whitelisted.

Your point about the parent shaper is the real kicker, though. It's like building a beautiful VIP lane on a highway that has no guaranteed asphalt. I've found the most success with a two-tier shaper: a parent guaranteeing a minimum pipe for the whole VoIP class, and then child policies for priority within that. That way even if identification isn't perfect, the traffic that does make it through has a real floor.

Have you actually had an ISP that stripped DSCP? I've mostly seen them leave it alone on business plans, but I'm paranoid enough to run a few test calls with a packet capture at the WAN interface anyway.



   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Hit counts are vanity metrics. They don't tell you if you're catching all the traffic or just most of it.

Focus on queue depth and latency under load. If your parent shaper has a guarantee and its queue is empty, your identification is solid. If it's building, you're missing packets. That's the only benchmark that matters.

And yes, combining DSCP, ASN, and ports is the right approach, but prioritize DSCP from your own gear first. You control that, you can trust it.


Metrics don't lie.


   
ReplyQuote
(@data_pipeline_rookie_43)
Honorable Member
Joined: 5 months ago
Posts: 365
 

That surgical filter with the specific provider subnets makes a lot of sense for precision. But reading through the later comments, I'm getting worried about the maintenance part. How do you keep those subnet lists updated when your provider changes things? Is it a manual process of checking their docs, or is there a way to automate that on the XGS?


rookie


   
ReplyQuote
(@data_diver_42)
Honorable Member
Joined: 7 months ago
Posts: 400
 

Yeah, the surgical approach with provider subnets is powerful for immediate precision, but it becomes an operational headache fast. It's a high-fidelity filter with a short shelf life.

What saved me was using a Geo-IP filter on the provider's ASN instead of static IPs. It's a good middle ground that can adapt as they add new subnets within their network block. But like others mentioned, you gotta verify their ASN isn't some massive cloud one that'll catch unrelated traffic.

My question is, how do you plan to maintain that subnet list? Manual updates from their docs, or are you scripting something?


Data is the new oil - but it's usually crude.


   
ReplyQuote
(@data_pipeline_ops)
Reputable Member
Joined: 6 months ago
Posts: 176
 

>watch for false positives if the provider's infra shares an ASN with unrelated services.

That's a solid warning. I've seen it happen with a service that shared Microsoft's ASN. The policy suddenly started prioritizing random Azure storage traffic along with Teams media.

The FQDN trick is interesting. Do you have to worry about TTL on those DNS resolutions, or does the firewall handle that automatically?


PipelinePadawan


   
ReplyQuote
Page 1 / 4