Repository navigation
lib_rtnl.c: uc_nl_request - if the socket assigned to 'sock' fails in some way there is no way to close and reopen it #422
Description
Activity
- changed the title
[-]lib_rtnl.c: uc_nl_request - it the socket assigned to 'sock' fails in some way there is no way to close and reopen it[/-][+]lib_rtnl.c: uc_nl_request - if the socket assigned to 'sock' fails in some way there is no way to close and reopen it[/+]on Aug 3, 2026 Any chance to capture an strace once in the broken state? Like letting the daemon continuously try to fetch the interface list, then - in failing state -
strace -pto it and see how the requests fail (socket level error or socket connected but not subscribed, ...).Not opposed to some sort of socket reconnect API either but I'd like to avoid just papering over an actual problem. At the very least I'd like to get the success state logic right to avoid caching a half-open broken socket connection.
I believe this is the snippet that's relevant (I can post more of course, but wanted to save you the trouble of wading though a lot of unimportant stuff):
[pid 10832] getsockopt(96, SOL_NETLINK, NETLINK_GET_STRICT_CHK, [0], [4]) = 0 [pid 10832] mmap2(NULL, 16384, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x77c09000 [pid 10832] sendmsg(96, {msg_name={sa_family=AF_NETLINK, nl_pid=0, nl_groups=00000000}, msg_namelen=12, msg_iov=[{iov_base=[{nlmsg_len=32, nlmsg_type=0x12 /* NLMSG_??? */, nlmsg_flags=NLM_F_REQUEST|NLM_F_ACK|0x300, nlmsg_seq=1785792563, nlmsg_pid=4196112}, "\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00"], iov_len=32}], msg_iovlen=1, msg_controllen=0, msg_flags=0}, 0) = 32 [pid 10832] mmap2(NULL, 20480, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x77c04000 [pid 10832] recvmsg(96, {msg_name={sa_family=AF_NETLINK, nl_pid=0, nl_groups=00000000}, msg_namelen=12, msg_iov=[{iov_base=[{nlmsg_len=44, nlmsg_type=NLMSG_ERROR, nlmsg_flags=0, nlmsg_seq=1785792564, nlmsg_pid=4196112}, {error=-EBUSY, msg=[{nlmsg_len=24, nlmsg_type=0x16 /* NLMSG_??? */, nlmsg_flags=NLM_F_REQUEST|NLM_F_ACK|0x300, nlmsg_seq=1785792564, nlmsg_pid=4196112}, "\x00\x00\x00\x00\x00\x00\x00\x00"]}], iov_len=16384}], msg_iovlen=1, msg_controllen=0, msg_flags=0}, 0) = 44 [pid 10832] munmap(0x77c04000, 20480) = 0 [pid 10832] munmap(0x77c09000, 16384) = 0 [pid 10832] getsockopt(96, SOL_NETLINK, NETLINK_GET_STRICT_CHK, [0], [4]) = 0 [pid 10832] mmap2(NULL, 16384, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x77c09000 [pid 10832] sendmsg(96, {msg_name={sa_family=AF_NETLINK, nl_pid=0, nl_groups=00000000}, msg_namelen=12, msg_iov=[{iov_base=[{nlmsg_len=24, nlmsg_type=0x16 /* NLMSG_??? */, nlmsg_flags=NLM_F_REQUEST|NLM_F_ACK|0x300, nlmsg_seq=1785792564, nlmsg_pid=4196112}, "\x00\x00\x00\x00\x00\x00\x00\x00"], iov_len=24}], msg_iovlen=1, msg_controllen=0, msg_flags=0}, 0) = 24 [pid 10832] mmap2(NULL, 20480, PROT_READ|PROT_WRITE, MAP_PRIVATE|MAP_ANONYMOUS, -1, 0) = 0x77c04000 [pid 10832] recvmsg(96, {msg_name={sa_family=AF_NETLINK, nl_pid=0, nl_groups=00000000}, msg_namelen=12, msg_iov=[{iov_base=[{nlmsg_len=52, nlmsg_type=NLMSG_ERROR, nlmsg_flags=0, nlmsg_seq=1785792563, nlmsg_pid=4196112}, {error=-EBUSY, msg=[{nlmsg_len=32, nlmsg_type=0x12 /* NLMSG_??? */, nlmsg_flags=NLM_F_REQUEST|NLM_F_ACK|0x300, nlmsg_seq=1785792563, nlmsg_pid=4196112}, "\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00"]}], iov_len=16384}], msg_iovlen=1, msg_controllen=0, msg_flags=0}, 0) = 52 [pid 10832] munmap(0x77c04000, 20480) = 0 [pid 10832] munmap(0x77c09000, 16384) = 0 [pid 10832] munmap(0x77bbe000, 8192) = 0I haven't tried to understand this, but I did see EBUSY showing up a couple of times in this trace.
I should add that the trace information immediately before and after the above looks entirely related to my application and not lib_rtnl.c
And another small piece of information. In longer testing we sometimes see these calls fail after working correctly for some time.
For the moment we're just closing the connection after use so we get a new 'sock' each time.
- added a commit that references this issue
on Oct 8, 2026
We see an occasional problem during our boot process where a daemon, which is using the uc_nl_request call in lib_rtnl.c, fails to get a valid reply - in our case we are reading an interface list. We think that the socket is opened correctly (~line 3477) but then fails. However, once in this failed state there is no way to reopen it as it is cached internally and will continue to fail. The only fix we've found is to restart our daemon.
We'd be happy to write a PR but are not sure what solution you'd prefer? We could do away with the socket cache entirely, or we could try to close it on any error.