Skip to content
Notifications
Clear all

What is the best way to manage 50+ SRX devices without a pricey manager?

6 Posts
6 Users
0 Reactions
18 Views
(@ethans)
Reputable Member
Joined: 2 months ago
Posts: 241
Topic starter   [#23551]

We're scaling up our SRX deployment past 50 firewalls. The official Juniper management solutions look powerful but are way over budget for this project.

I need a central way to push configs, maybe collect logs, and do basic compliance checks. Has anyone built a reliable system using free/open-source tools? Thinking Ansible, maybe a custom script with PyEZ, or a lightweight monitoring setup. What's actually working in production without becoming a full-time job to maintain?



   
Quote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

I manage a 120 SRX fleet for a national retail chain. We run a hybrid system: PyEZ for config pushes and Ansible for orchestration, with Graylog for logs.

* **Deployment and integration effort:** PyEZ takes about 2-3 days to get a reliable script framework going. The real time is building your config templates, which took our team two weeks.
* **Where it clearly wins:** Cost is nearly zero, just engineering hours. It's also flexible; you can make it do exactly what you need, like our staged config validation that checks for syntax errors before commit.
* **Where it breaks:** This isn't a live compliance dashboard. You have to schedule compliance checks as jobs, and reporting is only as good as the scripts you write. It will become a part-time job (about 2-3 hours/week in my shop) to maintain and adjust for new requirements.
* **Real pricing and hidden cost:** The tools are free, but the hidden cost is your team's Python/Ansible skill. If you don't have that, the learning curve adds 4-6 weeks. We also spun up a dedicated, small VM ($40/mo) to host the scripts and scheduler.

I'd recommend the PyEZ+Ansible path if you have in-house automation skills and your needs are config-and-audit focused. If you need real-time alerting or a compliance dashboard out of the box, it's the wrong choice. Tell us if you have a dedicated network automation person and whether you need live alerting from logs.


Ask me about hidden egress costs.


   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 2 months ago
Posts: 350
 

Agreed on the hidden skill cost. That's the real price tag.

You mentioned $40/mo for a scheduler VM. At 50+ devices, don't forget the cost of a backup instance. A single point of failure for config pushes can ruin your day. Run two small VMs or at least have a documented rebuild playbook ready to go.


Show me the bill


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

You're on the right track with Ansible and PyEZ. The key to avoiding a maintenance nightmare is to separate the logic from the data entirely.

Build your playbooks or scripts to be completely data-driven, using variable files or a simple inventory database. That way, pushing a new policy or updating a DNS server for all 50 devices is a one-line change in a YAML file, not a rewrite of your automation. The tooling becomes a stable framework, and the config data is what you manage.

For compliance, you'll need to run periodic audits. Schedule a daily PyEZ script that pulls specific config sections from each device, converts them to a structured format (like JSON), and dumps them into a directory. You can then use even simple diff tools against a gold master to see drift. It's not real-time, but it's effective and low-overhead once built.

Graylog or a similar ELK stack for logs is almost mandatory at that scale; trying to manage 50 individual syslog feeds is unworkable. Just factor the VM and storage costs for that into your total.


Every dollar counts.


   
ReplyQuote
(@brianw)
Reputable Member
Joined: 3 months ago
Posts: 242
 

A key point that often gets overlooked in these build-your-own scenarios is the operational cost of failure. While the automation tools themselves are free, a misapplied config push across 50 devices can create an outage that far exceeds the price of a commercial manager.

If you go the Ansible/PyEZ route, you must bake in a multi-stage validation and rollback mechanism. Our playbooks commit to a candidate configuration, then run a series of operational command checks (e.g., `show interfaces terse`, `show security policies hit-count`) before a final confirmed commit. This adds script complexity but is non-negotiable for production.

For log collection, don't build a parser. Use syslog-ng to forward everything to a cloud provider's logging service (like GCP Cloud Logging or Azure Monitor). The ingestion cost is trivial compared to engineering time spent maintaining a Graylog or ELK stack, and you get decent querying out of the box.


Spreadsheets or it didn't happen.


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

You're looking in the right direction with Ansible and PyEZ. The most common pitfall I see is teams treating their automation scripts like a one-off config editor, which quickly becomes unmanageable.

Invest your initial time in creating a solid data structure. Define all your device properties, security policies, and interface settings in YAML or JSON files that are completely separate from the playbooks. Then your Ansible roles or PyEZ scripts just become a generic engine that applies that data. Updating a NAT rule across the fleet becomes a five-minute edit in one file, not a risky hunt through dozens of scripts.

For the logs and compliance part, keep it simple. Point all device syslog to a central server, even a basic rsyslog VM, and use scheduled PyEZ jobs to pull config snippets for comparison. It won't be a live dashboard, but a daily diff report emailed to you catches 99% of the drift.


catdad


   
ReplyQuote