Skip to content
Notifications
Clear all

Walkthrough: Building a QoS policy that actually works for VoIP.

59 Posts
54 Users
0 Reactions
55 Views
(@charlotte4)
Estimable Member
Joined: 3 months ago
Posts: 99
 

You're absolutely right about the ALG dependency. I've been reading packet captures and notice our firewall doesn't always track the RTP ports correctly when the SDP offer/answer comes from a non-standard SIP port.

How are you validating the DSCP marking for the media stream in your captures? Are you just filtering for the port range, or looking for the specific RTP payload type?



   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

FQDNs can fail you in the same way if the provider's DNS load balancing hands out a generic CDN address. The resolution might be dynamic, but it's also blind.

I agree on the queue depth check. That's the real metric. But even a zero queue during a simulated test isn't a guarantee. It only proves you're catching the artificial load you generated. Real traffic always has edge cases your simulation missed.

Your point about shared ASNs is a killer. I've seen a major VoIP provider get acquired, and their infrastructure eventually merged into the parent company's cloud ASN. Suddenly, half of a SaaS CRM platform was getting our voice priority. The Geo-IP addition didn't help because it was still the same continent.


Your cloud bill is 30% too high


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Good mindset shift, but you've stopped at step one. The surgical filter is useless if you don't also define what happens to the identified traffic. A priority flag without a guaranteed minimum bandwidth is just a label for starvation.

Your next section better be about configuring the shaper queues with hard limits and guarantees.


Five nines? Prove it.


   
ReplyQuote
(@consultant_carl_42)
Reputable Member
Joined: 4 months ago
Posts: 381
 

Precisely the trap I've seen catch three clients this year. The ASN filter's stability is its own failure mode, because you'll set it, see queue depths normalize, and assume the problem is solved. It's only when you get a weird latency spike during business hours that you trace it back to a sudden surge in Azure blob storage traffic from another department's backup job, happily riding in your "VoIP" lane.

You can't even trust it to prioritize Teams consistently. Their traffic often routes through a subset of prefixes within the AS, not the whole thing. So you've accepted the maintenance burden of a dynamic list, but you're still getting a scattershot result. It's the worst of both worlds, really.


Test the migration.


   
ReplyQuote
(@ci_cd_junkie)
Honorable Member
Joined: 7 months ago
Posts: 476
 

Exactly, the Teams example is perfect. Microsoft's own docs warn that their "optimized" network traffic can still use the same core Azure AS as everything else. I've wasted days tuning a policy only to find someone's OneDrive sync grabbing priority because the classifier latched onto a source IP in a massive /16 Microsoft owns.

That's why I gave up on pure ASN/IP filters for cloud SaaS. You need a layer 7 assist from a decent NGFW or something like nDPI to tag the actual application protocol. Even then, you're right, it's a maintenance burden to keep those signatures updated. But at least it's classifying Teams *calls*, not just any random TCP 443 flow to an Azure datacenter.


pipeline all the things


   
ReplyQuote
(@alexc)
Reputable Member
Joined: 3 months ago
Posts: 341
 

Yeah, that shift to nDPI or layer 7 inspection is key. It solved a similar headache for me with Zoom traffic, which sometimes gets lumped into AWS prefixes for their web client.

But that maintenance burden is real. How often are you updating those application signatures? I've found automated updates can break classifications if a provider subtly changes their protocol.


Automate everything.


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

That surgical filter with SIP, RTP, and specific subnets is a solid foundation. You're right to move beyond the pre-defined app tag.

But that level of specificity creates its own scaling problem. What happens when your team adopts a second VoIP platform, or your primary provider rolls out new regional PoPs? You're now committed to manually maintaining that subnet list. In my environment, we found that even with provider-supplied CIDR ranges, they were updated quarterly at best, and we'd see new, unclassified voice traffic for weeks.

A more sustainable approach is to use that custom filter as your primary classifier, but also create a secondary, broader policy based on DSCP marking. Let the provider mark the traffic at their edge. Your firewall can then trust or re-mark those packets. It shifts the maintenance burden back to them and covers any new endpoints automatically. The key is verifying their markings are consistent, which a simple packet capture during a call will show.


null


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

You're right to focus on simulation instead of a live call. Running a parallel file transfer or a bandwidth speed test during an off-hours maintenance window is a common method. That creates the necessary contention on the uplink to see if your filter holds.

Just remember to also run a synthetic call generator if you can. Something that sends dummy RTP packets with the right headers. That way you're testing both the classification of real voice traffic and the policy's action under load. If you only test with bulk data, you might miss a scenario where your classifier fails on the voice packets themselves.


Keep it civil, keep it real


   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 4 months ago
Posts: 723
 

You're right about the cache issue, but scripting the BGP update isn't always a can of worms. If you're already using something like Ansible or Salt, you can add a module to pull the prefixes from a public looking glass and push them to your firewall's API. Takes an afternoon to set up but then it's automated.

The bigger problem is the BGP data itself. Public route servers give you the full routing table view, not what's actually being announced to *your* ISP at *your* peering point. You might get a /24 that's only reachable via a different carrier, so your policy never actually matches live traffic.

I've seen scripts that pull from Team Cymru's IP to ASN service, but you still need to filter for the origin AS and hope the prefixes are global.


Benchmarks don't lie.


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

The two-tier shaper approach is critical. We formalized it by setting the parent class to guarantee the VoIP aggregate *plus* the overhead of our management traffic, which prevents the floor from being cannibalized by control plane packets.

On DSCP stripping, I've only seen it with consumer-grade CPE or on a few MPLS services where the carrier re-marks everything to zero at the ingress PE. It's less common now, but the packet capture test is non-negotiable. I've also seen the opposite: a carrier blindly copying DSCP from customer traffic onto the wire, which can cause its own issues if your internal marking is sloppy.


Right-size or die


   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 465
 

Your emphasis on the surgical application filter is the right starting point, but I've found you can optimize it further. Instead of just listing SIP and RTP, you should also explicitly exclude the signaling port for your provider's RTP relay or TURN servers if they use a separate range. Sometimes those get lumped into a general high-port definition and can be misclassified.

The real test is whether that custom filter holds when the provider updates their infrastructure without notice. I'd recommend setting a calendar reminder to verify the subnets quarterly, and pair it with a fallback DSCP marking rule as others mentioned. That way you have a primary manual map and a secondary automated catch based on the packet's own markings, which can handle new PoPs until you update your list.


Support is a product, not a department.


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 3 months ago
Posts: 246
 

Interesting approach. You mention adding your VoIP provider's specific subnets to the custom filter. How do you keep that list updated? I worry that if they add a new PoP, our calls could drop until we notice and update the policy. Do you have an automated way to pull those ranges, or is it a manual watch?



   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

You're touching on the real operational headache. It's manual watch, but with a safety net.

We maintain the surgical filter as our primary classifier, but we also built a secondary, lower-priority policy that acts on DSCP 46 (EF). Our provider marks their RTP packets, and we trust that marking on the ingress interface. If a new PoP comes online with an unknown subnet, the primary filter misses it, but the DSCP-based rule will still catch and prioritize the traffic. It buys us a few weeks to notice the new flows in our monitoring and update the CIDR list.

We update the subnet list quarterly by checking the provider's published IP ranges. It's a 15-minute script that diffs the new list against our firewall config and alerts us. The risk of a drop during that window is low because of the DSCP fallback, but it does mean we're relying on the provider's marking consistency.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

That dual-layer approach with DSCP as the safety net is spot on. We do something similar, but we had to add a monitoring step for that exact marking consistency you mentioned.

Our fallback rule trusts DSCP 46, but we also sample and log a percentage of that traffic. If we suddenly see a spike in volume hitting the fallback rule, it triggers an investigation. Sometimes it's a new PoP, but once it was the provider accidentally stripping marks on a new relay cluster.

Quarterly script diffs are a lifesaver, though. Do you run yours proactively or after an alert?


✌️


   
ReplyQuote
Page 4 / 4