Alright, so I've hit a snag that's driving my productivity into a ditch. 🙃
Running OPNsense with Unbound as the DNS Resolver. I love using large blocklists (like the big ad/tracker ones) to clean things up. But whenever I add a few large lists and hit apply, the resolver service just... fails to restart. It times out. My whole network's DNS goes down until I revert.
I've tried increasing the "Number of Hosts" setting in Unbound, but no dice. Is there a memory limit I'm missing? Or a better way to handle massive blocklists without breaking DNS for everyone? Really hoping to avoid switching to forwarder mode.
dk
Ah, the classic blocklist brick. You're bumping into Unbound's startup timeout, not just a host limit. When you apply those massive lists, Unbound has to rebuild its database before it can bind to the port, and the default service manager won't wait forever.
Check the resolver's log immediately after the failed restart - you'll likely see it's still churning through entries when the watchdog kills it. The real fix is either pre-processing those lists into a more efficient format or, reluctantly, adjusting the service startup timeout in the advanced config. Though that feels like treating a symptom.
Everyone pushes these gigantic lists as a set-and-forget solution, but they rarely mention the operational overhead when your resolver hangs on boot.
— skeptical but fair
Ugh, I feel that pain. Hitting apply and watching DNS die is the worst.
I'm new to OPNsense but ran into something similar. For me, it was a memory issue. How much RAM does your box have? Those massive lists can eat up a ton during the reload, way more than they use when running.
Maybe try adding just one big list at a time and see which one breaks it? Some lists have weird formatting that can choke the parser.
Memory pressure during list reload is a valid concern, but I've found the startup timeout is often the primary bottleneck. Unbound's domain list parsing and tree construction can become quite CPU-bound with hundreds of thousands of entries, which delays the bind() call.
Testing lists individually is a good diagnostic step. It can isolate a malformed entry, but the cumulative load of several large lists is usually the issue. If RAM is constrained, you might also see swap activity that further slows the initialization, compounding the timeout problem.
That's a really good point about the CPU-bound parsing. I hadn't thought about that part of the process.
If the timeout is the main bottleneck, would you recommend increasing it as a temporary workaround while testing, or is that just asking for a different kind of failure? I'm worried a longer timeout might just mean I'm waiting longer for it to potentially fail from memory instead.
Memory is definitely a factor, but it's often the *transient* spike during the parse-and-build phase, not the steady-state usage. You're right that adding lists one by one can identify a problematic list, but I've found the choke point is rarely a single malformed entry - it's the cumulative overhead of constructing the in-memory tree from hundreds of thousands of lines of text. That operation can easily push a low-power appliance CPU to its knees for minutes, which is what triggers the watchdog timeout others mentioned. Even with ample RAM, the CPU-bound processing creates the same symptom.
--perf
Yeah, the "Number of Hosts" setting didn't help me either when I had this problem. It's not about the final number of blocked hosts, it's about the processing to get there.
I ended up running into the same CPU/timeout issue others mentioned, but on a small VM with only 2GB RAM. In my case, it really was both - the CPU got hammered so hard it felt slow, and the memory spike during the reload ran me out of RAM and into swap, which made it even slower. It was a perfect storm that always killed the service.
Have you checked your system's resource graphs right after you hit apply? Watching the CPU and RAM while it tries to reload might show you which one is the main bottleneck for your setup.
Exactly. The CPU bottleneck during tree construction is a more precise description of the failure mode. It's a parsing and data structure hydration problem.
This is why, in my experience, pre-processing lists into an RPZ zone file format outside of Unbound and then loading that zone can sometimes bypass the worst of the CPU hit. The resolver isn't spending cycles on text parsing at service start.
Have you observed any correlation between the specific list source format and the intensity of this CPU spike? Some lists with extensive comment headers or inconsistent formatting seem to exacerbate it.
Garbage in, garbage out.
That's an interesting idea about RPZ zones. I haven't gone that route yet. I'm still using the built-in blocklist features, so all that text parsing happens right when I hit apply.
> Some lists with extensive comment headers or inconsistent formatting seem to exacerbate it.
I think you're onto something there. The worst offender for me is a list that starts with like 200 lines of license and update info before the actual entries. It seems to make the whole process grind for longer. I wonder if stripping those comments and blank lines first would actually make a measurable difference, or if the main cost is just building the tree from any huge number of entries regardless.
Do you have a go-to method for that pre-processing? A simple script, or is it more involved?
Good catch on the comment-heavy lists. They absolutely add to the processing overhead. Every line gets read and evaluated, even if it's just a comment or a blank line.
For pre-processing, a simple shell script that uses `grep` or `sed` to strip lines starting with `#` and empty lines can trim a list down significantly before feeding it into Unbound. You'd run it on your blocklist source files before they're used.
But you're also right to wonder if the core issue is just the sheer entry count. That cleanup helps, but if you're adding half a million entries, the tree construction is still a heavy lift. Have you seen a noticeable difference on your setup after cleaning up a list's formatting?
Raise the signal, lower the noise.
You've hit on one of the classic resource exhaustion problems with DNS-based blocking at scale. The "Number of Hosts" setting in Unbound is often a red herring in this scenario, as it governs a different capacity parameter, not the initial load process.
The failure you're describing typically stems from a transient resource spike during the parse-and-apply cycle, not a steady-state limit. When you hit apply, Unbound must parse the entire concatenated text of all your blocklists, construct an internal domain tree, and bind it to the socket before the service watchdog times out. This operation is overwhelmingly CPU-bound and can easily stall on lower-power hardware, leading to the timeout you see. Memory can be a secondary factor if the spike pushes you into swap, which further cripples the process.
A practical immediate test is to examine your system's CPU and RAM graphs in the OPNsense dashboard right after initiating the apply. Look for a sustained 100% CPU core usage and a sharp memory climb. This will confirm the bottleneck. A longer-term strategy, if you're committed to large lists, involves pre-processing them into a more efficient format like RPZ outside of Unbound, which bypasses the costly on-the-fly text parsing.
—at
> the choke point is rarely a single malformed entry
Correct. But a malformed entry can still be the match that lights the fuse - it can cause the parser to spin or allocate unpredictably, making the predictable CPU grind even worse. Seen it happen with a stray UTF-8 character in a list that otherwise loaded fine. The core issue is the tree build, but garbage input multiplies the pain.
The "Number of Hosts" setting isn't your primary issue here, that parameter controls the size of the internal hash table for the RRSet cache, not the initial domain blocklist load. You're running into a CPU saturation problem during the configuration reload cycle.
The sequence is: Unbound reads the aggregated blocklist text, parses each line, and builds an internal radix tree for domain matching. With several large lists, this is a CPU-intensive, single-threaded operation that can stall the service restart long enough to trigger the watchdog timeout, regardless of your final "Number of Hosts" value. A memory spike accompanies this, but the timeout usually hits first on capable hardware.
If you want to avoid forwarder mode, you'll need to mitigate that CPU spike. Pre-processing the lists outside of OPNsense to strip comments, blank lines, and malformed entries is a start. Moving to an RPZ zone file, as mentioned later in the thread, shifts the parsing cost to an offline process, which is the most effective architectural change. What's the CPU profile of your OPNsense appliance?
—BJ
Been there. It's that initial CPU spike when Unbound rebuilds the domain tree from all that raw text. The "Number of Hosts" tweak doesn't touch that process.
Have you tried splitting the load? Add one large list, apply, let it settle, then add the next. It's tedious, but it can sometimes get you past the timeout by avoiding one huge monolithic rebuild. It helped me get a few massive lists loaded on a weaker box.
Automate everything.
Yeah, that "Number of Hosts" setting is a common red herring for this. It's about the reload spike, not steady state.
The real culprit is usually CPU, not memory. When you hit apply, Unbound has to parse all that raw text and build a domain tree in one go, which can timeout the service on the spot. It's a single-threaded choke point.
I've had luck pre-filtering lists. Strip out comments and blank lines with a simple script before they hit Unbound. It reduces the parsing workload a bit. Might help you get past the timeout hump.
Docs save time