Bug #16937
openUnbound cannot reload due to SSL/TLS certificate file permissions
100%
Description
Please refer to following forum post: https://forum.netgate.com/topic/200922/unbound-dns-resolver-periodically-non-responsive
When Unbound is restarted (by pfBlockerNG for example), and this option is selected, Unbound starts to produce the following error and DNS stops resolving:
unbound 1669 [1669:4] error: Error in SSL_CTX use_PrivateKey_file crypto error:8000000D:system library::Permission denied
I have unchecked the option for now. See attached screenshots/log entries
Files
GT Updated by Georgiy Tyutyunnik 2 months ago
can't reproduce with stock config on 26.03.1, unbound config verified to match the forum thread
can you add the full config sections for the DNSBL, unbound and the server certificate, if possible?
please also add the output of the command
ls -lah /var/unbound/
RS Updated by RED SKULL about 2 months ago
Georgiy Tyutyunnik wrote in #note-1:
can't reproduce with stock config on 26.03.1, unbound config verified to match the forum thread
can you add the full config sections for the DNSBL, unbound and the server certificate, if possible?
please also add the output of the command
ls -lah /var/unbound/
total 163 MB
drwxr-xr-x 8 unbound unbound 35B Jul 31 22:46 .
drwxr-xr-x 27 root wheel 27B Aug 15 2025 ..
-rw-r--r-- 1 root unbound 1.3K Jul 31 22:45 access_lists.conf
drwxr-xr-x 2 unbound unbound 2B Aug 15 2025 conf.d
dr-xr-xr-x 9 root wheel 512B Jul 31 21:59 dev
-rw-r--r-- 1 root unbound 0B Jul 31 22:45 dhcpleases_entries.conf
-rw-r--r-- 1 root unbound 3.3K Jul 31 21:58 dnsbl_cert.pem
-rw-r--r-- 1 root unbound 0B Jul 31 22:45 domainoverrides.conf
-rw-r--r-- 1 root unbound 488B Jul 31 22:45 host_entries.conf
drwxr-xr-x 2 root unbound 3B Oct 3 2025 leases
drwxr-xr-x 4 root wheel 86B Jul 31 21:57 lib
-rw-r--r-- 1 root unbound 1.4K Oct 3 2025 pfb_dnsbl_lighty.conf
-rw-r--r-- 1 root unbound 197B Jul 31 19:47 pfb_dnsbl.safesearch.conf
-rw-r--r-- 1 root unbound 691M Jul 31 20:26 pfb_py_data.txt
-rw-r--r-- 1 unbound unbound 8.0K Jul 31 22:44 pfb_py_dnsbl.sqlite
-rw-r--r-- 1 root unbound 1.6M Jul 31 21:58 pfb_py_hsts.txt
-rw-r--r-- 1 unbound unbound 12K Jul 31 22:46 pfb_py_resolver.sqlite
-rw-r--r-- 1 root unbound 25K Oct 3 2025 pfb_py_ss.txt
-rw-r--r-- 1 root unbound 5.1K Jul 20 13:20 pfb_py_whitelist.txt
-rw-r--r-- 1 root unbound 5.5K Jul 31 21:58 pfb_unbound_include.inc
-rw-r--r-- 1 root unbound 651B Jul 5 04:25 pfb_unbound.ini
-rw-r--r-- 1 root unbound 65K Jul 31 21:58 pfb_unbound.py
-rw-r--r-- 1 root unbound 300B Oct 3 2025 remotecontrol.conf
-rw-r--r-- 1 unbound unbound 1.2K Jul 31 22:46 root.key
-rw-r--r-- 1 root unbound 2.0K Jul 31 22:24 sslcert.crt
-rw------- 1 root unbound 306B Jul 31 22:24 sslcert.key
-rw------- 1 unbound unbound 2.4K Oct 3 2025 unbound_control.key
-rw-r----- 1 unbound unbound 1.5K Oct 3 2025 unbound_control.pem
-rw------- 1 unbound unbound 2.4K Oct 3 2025 unbound_server.key
-rw-r----- 1 unbound unbound 1.5K Oct 3 2025 unbound_server.pem
-rw-r--r-- 1 unbound unbound 3.3K Jul 31 22:45 unbound.conf
-rw-r--r-- 1 root unbound 3.9K Jul 19 04:14 unbound.conf.error
drwxr-xr-x 3 root unbound 3B Oct 3 2025 usr
drwxr-xr-x 3 root unbound 3B Oct 3 2025 var
PD Updated by P Davis about 2 months ago
- File Certificates.png Certificates.png added
- File DNSBL.png DNSBL.png added
- File Unbound.png Unbound.png added
Can you add the full config sections for the DNSBL, unbound and the server certificate, if possible?
See attached files, let me know if this is what you are looking for.
Also, here is the output of the command ls -lah /var/unbound/:
total 20 MB
drwxr-xr-x 8 unbound unbound 33B Aug 1 03:31 .
drwxr-xr-x 28 root wheel 28B Dec 1 2025 ..rw-r--r- 1 root unbound 343B Jul 23 05:09 access_lists.conf
drwxr-xr-x 2 unbound unbound 2B Dec 1 2025 conf.d
dr-xr-xr-x 8 root wheel 512B Jul 23 05:09 devrw-r--r- 1 root unbound 0B Jul 23 05:09 dhcpleases_entries.confrw-r--r- 1 root unbound 3.3K Jul 3 05:41 dnsbl_cert.pemrw-r--r- 1 root unbound 0B Jul 23 05:09 domainoverrides.confrw-r--r- 1 root unbound 523B Jul 23 05:09 host_entries.conf
drwxr-xr-x 2 root unbound 4B Nov 26 2024 leases
drwxr-xr-x 5 root wheel 96B May 28 05:26 librw-r--r- 1 root unbound 1.4K Aug 8 2025 pfb_dnsbl_lighty.confrw-r--r- 1 unbound unbound 8.0K Aug 1 03:28 pfb_py_cache.sqliterw-r--r- 1 root unbound 86M Aug 1 00:32 pfb_py_data.txtrw-r--r- 1 unbound unbound 8.0K Aug 1 03:28 pfb_py_dnsbl.sqliterw-r--r- 1 root unbound 1.6M Jul 3 05:41 pfb_py_hsts.txtrw-r--r- 1 unbound unbound 16K Aug 1 03:31 pfb_py_resolver.sqliterw-r--r- 1 root unbound 1.3K Dec 18 2024 pfb_py_whitelist.txtrw-r--r- 1 root unbound 5.5K Jul 3 05:41 pfb_unbound_include.incrw-r--r- 1 root unbound 356B Dec 19 2025 pfb_unbound.inirw-r--r- 1 root unbound 65K Jul 3 05:41 pfb_unbound.pyrw-r--r- 1 root unbound 300B Dec 31 2021 remotecontrol.confrw-r--r- 1 unbound unbound 1.2K Aug 1 00:32 root.keyrw-r--r- 1 root unbound 3.7K Jul 7 19:22 sslcert.crtrw------ 1 root unbound 1.7K Jul 7 19:22 sslcert.keyrw------ 1 unbound unbound 2.4K Dec 31 2021 unbound_control.keyrw-r---- 1 unbound unbound 1.3K Dec 31 2021 unbound_control.pemrw------ 1 unbound unbound 2.4K Dec 31 2021 unbound_server.keyrw-r---- 1 unbound unbound 1.3K Dec 31 2021 unbound_server.pemrw-r--r- 1 unbound unbound 2.2K Jul 23 05:09 unbound.confrw-r--r- 1 root unbound 2.6K Jul 6 12:35 unbound.conf.error
drwxr-xr-x 3 root unbound 3B Jul 1 2022 usr
drwxr-xr-x 3 root unbound 3B Feb 17 2023 var
JS Updated by John S 15 days ago
I am also experiencing this issue on pfSense Plus 26.07. I also use pfBlockerNG, but I also see this issue on a restart of pfSense.
With Respond to incoming SSL/TLS queries from local clients enabled, Unbound can fail with the same TLS private-key permission error described in this ticket.
If I disable Respond to incoming SSL/TLS queries from local clients, the problem goes away and DNS works normally.
On my system pfSense generates the incoming DNS-over-TLS files as:
-rw-r--r-- root unbound /var/unbound/sslcert.crt -rw------- root unbound /var/unbound/sslcert.key
These are also the same permissions shown by the original reporter.
The different permissions themselves make sense: the certificate is public, while the private key requires stronger protection.
However, pfSense explicitly sets the certificate to 0644 and the private key to 0600 in:
src/etc/inc/unbound.inc
https://github.com/pfsense/pfsense/blob/master/src/etc/inc/unbound.inc
pfSense also configures Unbound to run as:
username: "unbound"
With 0600 root:unbound, the running unbound user cannot read sslcert.key after privileges have been dropped. The unbound group ownership does not provide access because the group permission bits are ---.
This appears potentially significant with newer Unbound TLS reload behaviour.
Current upstream Unbound can reread changed tls-service-key / tls-service-pem during reload/fast_reload when those files are accessible. The upstream documentation also notes that if the TLS key has root-only permissions, changing the TLS credentials requires a restart rather than a reload:
https://github.com/NLnetLabs/unbound/blob/master/doc/unbound.conf.rst
This appears consistent with the reported error:
SSL_CTX use_PrivateKey_file Permission denied
pfSense core also contains an Unbound reload path which is executed as the unbound user, effectively:
su -m unbound -c '/usr/local/sbin/unbound-control ... reload'
But it's also possible pfBlockerNG is doing the reload.
This seems relevant because /var/unbound/sslcert.key is 0600 root:unbound and therefore cannot be read by the runtime unbound user.
The fact that disabling incoming SSL/TLS resolves the problem for me on 26.07 seems to point specifically toward this TLS path.
To assist further, I am trying to narrow down what in my configuration means I trigger this, but I do see that unbound can't read the key file during a reload.
JS Updated by John S 15 days ago
To update.
I have also identified a reproducible trigger for this on my system: a WAN IPv4 address change. I just need Respond to incoming SSL/TLS queries from local clients enabled, and a WAN IPv4 address change. This happens even on a fresh install.
In my case, after the WAN address changes, Unbound subsequently hits the same TLS private-key permission failure and DNS stops working.
JP Updated by Jim Pingle 15 days ago
I still haven't managed to replicate this here, but if it is related to the ownership of the cert/key, try this diff:
diff --git a/src/etc/inc/unbound.inc b/src/etc/inc/unbound.inc
index 7babe305ce..49285b9b0f 100644
--- a/src/etc/inc/unbound.inc
+++ b/src/etc/inc/unbound.inc
@@ -385,8 +385,12 @@ EOF;
// Write CA and Server Cert
file_put_contents($tlscert_path, $cert_chain);
+ chown($tlscert_path, 'unbound');
+ chgrp($tlscert_path, 'unbound');
chmod($tlscert_path, 0644);
file_put_contents($tlskey_path, base64_decode($cert['prv']));
+ chown($tlskey_path, 'unbound');
+ chgrp($tlskey_path, 'unbound');
chmod($tlskey_path, 0600);
// Add config for CA and Server Cert
You can install the System Patches package and then create an entry for that diff to apply the fix.
Save and apply the Unbound settings (or reboot) after applying.
JP Updated by Jim Pingle 13 days ago
- Project changed from pfSense Plus to pfSense
- Category changed from DNS Resolver to DNS Resolver
- Affected Plus Version deleted (
26.03.1)
JP Updated by Jim Pingle 13 days ago
- Has duplicate Bug #17069: DNS Resolver TLS key is written 0600 owned by root, unbound cannot read it after startup added
EK Updated by Emre K 13 days ago
Also seeing this on CE 2.9.0-RELEASE (unbound 1.25.2).
Same permissions as the others: /var/unbound/sslcert.key is 0600 root:unbound
while unbound runs as user unbound.
If it helps with the "can't reproduce" part, this fails on demand with the
resolver running and SSL/TLS enabled:
touch /var/unbound/sslcert.key
unbound-control -c /var/unbound/unbound.conf fast_reload
unbound[<pid>]: [<pid>:14] error: error for private key file: /sslcert.key
unbound[<pid>]: [<pid>:14] error: Error in SSL_CTX use_PrivateKey_file crypto error:8000000D:system library::Permission denied
unbound[<pid>]: [<pid>:14] fatal error: could not set up listen SSL_CTX
Test it on something you don't need. The process doesn't exit afterwards. It
keeps udp/53, tcp/53 and tcp/953 bound, answers nothing, and ignores SIGTERM.
kill -9 is the only thing that ends it. That's why the GUI Restart button just
spins, and a GUI reboot hangs too. Mine needed a power cycle, DNS was down
about 16 minutes.
Anything that changes the key's timestamp while unbound is running will do it.
On my box it's Kea rather than pfBlockerNG: every lease event runs kea2unbound,
which issues fast_reload. That's harmless normally, because a restart rewrites
the key and the new process reads it as root. It went wrong when my WAN port
flapped and two reconfigure runs overlapped, so one rewrote the key underneath
the unbound the other had just started. Three "bind: address already in use"
errors in the system log from that window.
EK Updated by Emre K 13 days ago · Edited
Applied the patch from #note-7 on CE 2.9.0 (unbound 1.25.2). It fixes the permission
error, but the patch alone isn't enough - there look to be a few separate things
tangled together here, so I've tried to keep them apart below.
1) The key permissions (what the patch fixes)
After patching, both files are unbound:unbound and su -m unbound -c cat
succeeds. That part works.
2) A second file unbound can't reach on reload
touch + fast_reload still kills unbound, now with a different error:
unbound[30938]: [30938:14] error: error in SSL_CTX verify crypto error:80000002:system library::No such file or directory
unbound[30938]: [30938:14] error: and additionally crypto error:10000080:BIO routines::no such file
unbound[30938]: [30938:14] error: and additionally crypto error:05880020:x509 certificate routines::BIO lib
unbound[30938]: [30938:14] fatal error: could not set up connect SSL_CTX
This is the outgoing context rather than the listening one. unbound.conf has
tls-cert-bundle: "/etc/ssl/cert.pem"
which is outside the chroot. At startup it loads fine, because that happens
before chroot. On a reload the process is already in /var/unbound and the path
doesn't resolve. unbound seems to rebuild both contexts in the same call, so a
change to the key's timestamp reaches both. The permission error was hiding
this one.
Same shape as the first problem: something that's readable at startup but not
afterwards. I haven't tested whether copying the bundle into the chroot helps.
3) unbound doesn't exit when this happens
unbound-control times out and never returns, and the process stays up. The
fast_reload thread is stuck:
30938 112181 unbound unbound/freload mi_switch+0xbc sleepq_catch_signals+0x27d
sleepq_wait_sig+0x9 sleep+0x1a2 umtxq_sleep+0x2cd __umtx_op_sem2_wait+0x4a3
sys_umtx_op+0x89 amd64_syscall+0x133 fast_syscall_common+0xf8
It keeps 53 and 953 bound and still answers cached queries, but recursion is
dead and new lookups time out, so it looks alive from the outside. SIGTERM does
nothing.
I'd guess this one is upstream rather than pfSense. A failed reload killing a
server that was running fine seems wrong on its own, and it doesn't even die
cleanly - if it just exited, the normal restart would pick it up.
4) The GUI can't recover it
kill -9 on the pid clears it in seconds. Process exits, ports release, DNS is
back after a Save+Apply.
As far as I can tell services_unbound_configure only sends TERM and waits 30s
on the restart path, with no escalation, so the GUI has no way out of this
state. That's the difference between a fast fix and the power cycle my first
incident needed, and it's probably why this gets reported as needing reboots.
So: 1 and 2 both look like pfSense writing files unbound can't get at once it's
running, 2 being the one still open. 3 looks upstream. 4 is what turns any of
it into an outage instead of a minor inconvenience.
For anyone hitting this now, disabling "Respond to incoming SSL/TLS queries from
local clients" is still the only thing that avoids it completely, and kill -9 is
the way out if you're already stuck.
A seemingly innocent enable "DoT setting" in the unbound package is hiding a landmine.
JP Updated by Jim Pingle 13 days ago · Edited
- Status changed from Incomplete to Confirmed
- Assignee set to Jim Pingle
- Target version set to CE-Next
- Plus Target Version set to 26.10
OK, using this sequence I was able to reproduce the error:
$ touch /var/unbound/sslcert.key
$ unbound-control -c /var/unbound/unbound.conf fast_reload
Try the following patch instead which appears to fix it here for me. Back out the other patch first as this contains the previous fix as well.
I have it copy the CA bundle into the chroot and reference it locally. After this change I can reload it without error.
diff --git a/src/etc/inc/unbound.inc b/src/etc/inc/unbound.inc
index 7babe305ce..c5ab65782d 100644
--- a/src/etc/inc/unbound.inc
+++ b/src/etc/inc/unbound.inc
@@ -365,7 +365,11 @@ EOF;
}
// TLS Configuration
- $tlsconfig = "tls-cert-bundle: \"/etc/ssl/cert.pem\"\n";
+ $cabundle = "{$g['unbound_chroot_path']}/cabundle.pem";
+ copy("/etc/ssl/cert.pem", $cabundle);
+ chown($cabundle, 'unbound');
+ chgrp($cabundle, 'unbound');
+ $tlsconfig = "tls-cert-bundle: \"{$cabundle}\"\n";
if (isset($unboundcfg['enablessl'])) {
$tlscert_path = "{$g['unbound_chroot_path']}/sslcert.crt";
@@ -385,8 +389,12 @@ EOF;
// Write CA and Server Cert
file_put_contents($tlscert_path, $cert_chain);
+ chown($tlscert_path, 'unbound');
+ chgrp($tlscert_path, 'unbound');
chmod($tlscert_path, 0644);
file_put_contents($tlskey_path, base64_decode($cert['prv']));
+ chown($tlskey_path, 'unbound');
+ chgrp($tlskey_path, 'unbound');
chmod($tlskey_path, 0600);
// Add config for CA and Server Cert
JP Updated by Jim Pingle 13 days ago
- Subject changed from Unbound aborts, and DNS stops resolving after restarted while using "Respond to incoming SSL/TLS queries from local clients" option to Unbound cannot reload due to SSL/TLS certificate file permissions
EK Updated by Emre K 13 days ago · Edited
Second patch applied on, TLS service re-enabled.
Confirmed fixed: after Save and Apply the key, cert and bundle are all unbound:unbound, the bundle is a copy of /etc/ssl/cert.pem and the config points at the chroot copy.
Then touch sslcert.key + unbound-control fast_reload, ok, twice, same PID, uptime kept climbing. Touching the bundle and reloading also works. Same box, same test that reliably killed it before, so this is the right fix.
Things I noticed reading the patch, (not sure if this patch will go into the production as is but I am pointing out anyway)
- The bundle copy's return value isn't checked. If it fails for any reason, Unbound refuses to start with the same "could not set up connect SSL_CTX" message this ticket was about, but at startup rather than on reload, and nothing in the logs points at the copy as the cause. You also get PHP warnings from the chown on a file that isn't there.
- The copy happens outside the TLS-enabled block. Every install (regardless of the unbound TLS settings) now writes a ~180 KB bundle into /var/unbound on every resolver config write, whether or not TLS is in use, and it's never removed.
- The key and the bundle are both written in place, truncate, then write. Unbound only re-reads on an mtime change and the window is tiny, but concurrent reconfigures are exactly how this bug surfaced in the first place (WAN events firing services_unbound_configure in parallel). A reload landing mid-write reads a partial file, and the result is the same fatal and the same hang. Narrow, but real, and only dangerous because the failure hangs instead of erroring.
- Nothing cleans up when TLS is disabled. sslcert.crt and sslcert.key stay on disk, and since the chown only runs inside the enabled block, a box that had TLS on and then off before this fix keeps a root-owned key indefinitely. Harmless until TLS is re-enabled, at which point the next Save fixes it, so residue rather than a live bug.
- The certificate itself can't be deleted while Unbound uses it, but the CA behind it can, the CA delete check doesn't consider certificates the CA has issued. Delete it and the next Save writes a leaf-only sslcert.crt; Unbound starts fine and serves an incomplete chain with no warning. Clients that validate the chain fail. The chroot bundle only refreshes when the resolver config is written, so it can lag /etc/ssl/cert.pem after a ca_root_nss update. In practice every restart regenerates the config, so this is close to theoretical.
This fixes the trigger, not the behaviour underneath. Anything that makes the key or bundle unreadable at reload time in future produces the same thing.
A process that's alive with 53 bound, answers only from cache, ignores SIGTERM, spins the Restart button and hangs the reboot.
- Reload failure being fatal in Unbound
- the fatal hanging rather than exiting
- pfSense restart path waiting 30 seconds after TERM and then giving up instead of escalating to KILL
- nothing serialising concurrent reconfigures.
The first two are I guess upstream (or not, I cannot be sure); I'd assume you'd want to raise those with yourselves.
JP Updated by Jim Pingle 13 days ago
Emre K wrote in #note-14:
- The bundle copy's return value isn't checked. If it fails for any reason, Unbound refuses to start with the same "could not set up connect SSL_CTX" message this ticket was about, but at startup rather than on reload, and nothing in the logs points at the copy as the cause. You also get PHP warnings from the chown on a file that isn't there.
It's just a proof of concept, though it should be OK as-is. If the source wasn't there, Unbound wouldn't start anyhow. If the file couldn't be copied for some other reason, there are larger problems than Unbound not starting (e.g. disk is out of space, which would likely already mean Unbound is broken in other ways).
- The copy happens outside the TLS-enabled block. Every install (regardless of the unbound TLS settings) now writes a ~180 KB bundle into /var/unbound on every resolver config write, whether or not TLS is in use, and it's never removed.
Yes, that's intentional. The CA bundle is not just for acting as a TLS server. Unbound needs the CAs for other contexts, like speaking to upstream TLS sources like external DNS over TLS servers. 180KB is inlikely to be consequential anyhow.
- The key and the bundle are both written in place, truncate, then write. Unbound only re-reads on an mtime change and the window is tiny, but concurrent reconfigures are exactly how this bug surfaced in the first place (WAN events firing services_unbound_configure in parallel). A reload landing mid-write reads a partial file, and the result is the same fatal and the same hang. Narrow, but real, and only dangerous because the failure hangs instead of erroring.
Unlikely to be a problem or it would have been hit long ago. Also ZFS is copy on write so in practice it's essentially identical to writing to a new file and then moving after.
- Nothing cleans up when TLS is disabled. sslcert.crt and sslcert.key stay on disk, and since the chown only runs inside the enabled block, a box that had TLS on and then off before this fix keeps a root-owned key indefinitely. Harmless until TLS is re-enabled, at which point the next Save fixes it, so residue rather than a live bug.
Doesn't really matter either, it could be cleaned up but it's not necessary. The files are not referenced when the feature isn't enabled so there is no need to change the owner at that time.
- The certificate itself can't be deleted while Unbound uses it, but the CA behind it can, the CA delete check doesn't consider certificates the CA has issued. Delete it and the next Save writes a leaf-only sslcert.crt; Unbound starts fine and serves an incomplete chain with no warning. Clients that validate the chain fail. The chroot bundle only refreshes when the resolver config is written, so it can lag /etc/ssl/cert.pem after a ca_root_nss update. In practice every restart regenerates the config, so this is close to theoretical.
That's a byproduct of how CAs work for pretty much anything currently, it's not unique to this.
This fixes the trigger, not the behaviour underneath. Anything that makes the key or bundle unreadable at reload time in future produces the same thing.
Four pieces to that, and the patch touches none of them
A process that's alive with 53 bound, answers only from cache, ignores SIGTERM, spins the Restart button and hangs the reboot.
- Reload failure being fatal in Unbound
- the fatal hanging rather than exiting
- pfSense restart path waiting 30 seconds after TERM and then giving up instead of escalating to KILL
- nothing serialising concurrent reconfigures.
The first two are I guess upstream (or not, I cannot be sure); I'd assume you'd want to raise those with yourselves.
None of those are the problem here, this issue is specifically for the certificate error. Some of those are deeper in unbound, architecture/design, or other areas of the code that could be investigated separately on their own dedicated Redmine issues. Each issue needs to be considered on its own, rather than lumping all of it together.
JP Updated by Jim Pingle 13 days ago
- File 16937.patch 16937.patch added
- Status changed from Confirmed to Feedback
- % Done changed from 0 to 100
Fixed in commit 5fa516864b65c045ff61b4ced9e8a7f5b278330d
Patch is attached but is the same as the diff above.