| CVE |
Vendors |
Products |
Updated |
CVSS v3.1 |
| In the Linux kernel, the following vulnerability has been resolved:
posix-cpu-timers: Prevent freeing a timer which is queued on the expiry list
Kijo analyzed another race in the POSIX CPU timer code:
Commit bf635681c906 converted cpu_timer::firing from a tristate value to a
boolean. This lost the distinction between "not owned by the firing list"
and "still owned, but delivery was canceled". The resulting race is:
expiry handler timer_settime() timer_delete()
-------------- --------------- --------------
collect timer onto
private firing list
firing = true
observes firing = true
firing = false
return TIMER_RETRY
wait for handler
observes firing = false
finish deletion
unhash and free timer
resume list traversal
read freed elist.next
-> UAF
The firing bit is clearly the wrong indicator since that commit.
Check whether the timer is queued on the expiry list or not instead. If it
is queued clear the firing bit to prevent signal delivery as before and
return TIMER_RETRY so the caller unlocks the timer which allows the expiry
code to make progress and remove it from the list. |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: libipw: reject too-short association responses
libipw_handle_assoc_resp() reads the capability, status and aid fields
of the 30-byte association response prefix and then computes the
information element length as
stats->len - sizeof(*frame)
stats->len is a u16 and sizeof() has type size_t, so the subtraction is
evaluated as size_t and wraps instead of going negative. Truncating
that to the u16 length parameter of libipw_parse_info_param() turns a
frame shorter than the fixed fields into a length near 64 KiB, and the
parser then reads past the receive buffer.
Both the ipw2100 and ipw2200 management receive paths reach this
function having established only that the frame carries the generic
24-byte three-address header.
Reject the frame before any fixed field is touched.
Found by an AI-assisted review of length arithmetic in management frame
parsers. Verified with a KUnit case under Generic KASAN on arm64 under
QEMU; I do not have the hardware, so it is not tested on a real device. |
| In the Linux kernel, the following vulnerability has been resolved:
xfrm: add missing rcu_read_lock(), skb_dst_force() and dev_hold() for xfrm_trans_reinject()
syzbot reported a suspicious RCU usage warning in ip6_pkt_drop():
WARNING: suspicious RCU usage in ip6_pkt_drop
include/net/addrconf.h:389 suspicious rcu_dereference_check() usage!
Call Trace:
__in6_dev_get_safely include/net/addrconf.h:389 [inline]
ip6_pkt_drop+0x596/0x610 net/ipv6/route.c:4620
ip6_pkt_discard+0x1c/0x30 net/ipv6/route.c:4651
xfrm_trans_reinject+0x324/0x630 net/xfrm/xfrm_input.c:806
process_one_work kernel/workqueue.c:3322 [inline]
process_scheduled_works+0xa8e/0x14e0 kernel/workqueue.c:3405
worker_thread+0xa47/0xfb0 kernel/workqueue.c:3486
When commit 4f4920669d21 ("xfrm: Reinject transport-mode packets through
workqueue") converted xfrm_trans_reinject from a tasklet to a workqueue,
the reinjection loop ceased running in softirq context. Workqueue workers
run in process context where local_bh_disable() does not enter an RCU
read-side critical section under CONFIG_PREEMPT_RCU.
Because finish callbacks (such as ip6_rcv_finish) expect to run under an
RCU read lock (performing route lookups, l3mdev lookups, and accessing
RCU-protected data structures), invoking them in workqueue context without
rcu_read_lock() triggers RCU lockdep warnings.
Furthermore, packets queued to the workqueue via xfrm_trans_queue_net()
may carry non-refcounted (noref) dst entries (e.g. from ip_route_input_noref).
Additionally, on netdevice unregistration, dst_dev_put() replaces dst->dev
with blackhole_netdev, so dst entries do not keep skb->dev alive while
queued in the workqueue.
Fix these issues by:
1. Calling skb_dst_force(skb) in xfrm_trans_queue_net() while still in the
caller's RCU section to ensure dst is reference-counted before queuing.
2. Holding a reference on skb->dev via dev_hold()/dev_put() across workqueue
deferral so skb->dev remains valid during finish() callback processing.
3. Acquiring rcu_read_lock() around the finish callback invocation loop in
xfrm_trans_reinject(). |
| In the Linux kernel, the following vulnerability has been resolved:
hwmon: (w83791d) remove fan/pwm 4-5 sysfs group on remove
When the fan/pwm 4-5 pins are not used as GPIO, w83791d_probe()
creates the w83791d_group_fanpwm45 sysfs group on the I2C client
device.
The probe error path removes this group when a later initialization
step fails, but the normal remove path only removes w83791d_group.
As a result, the optional fan/pwm 4-5 sysfs files can remain after the
driver is unbound.
The callbacks associated with these files access the driver data,
which is devm allocated and released after driver unbind. Leaving the
sysfs files behind can therefore result in accesses to stale driver
data.
Remove w83791d_group_fanpwm45 during normal teardown as well.
This issue was found by manual code inspection. |
| In the Linux kernel, the following vulnerability has been resolved:
mips: select CONFIG_WEAK_REORDERING_BEYOND_LLSC from CONFIG_EYEQ
On I6500 CPU cores, lld and scd give no ordering guarantees (same as all
other instructions). To respect the assumption that arch_cmpxchg() is
fully ordered, we must inject sync instructions above and below our
lld/scd loops using the already in place WEAK_REORDERING_BEYOND_LLSC
infrastructure.
Otherwise, bad things can happen:
[ 34.054496] CPU 3 Unable to handle kernel paging request at virtual address 0000000000000000, epc == a80000080838e01c, ra == a80000080838dfc4
[ 34.054559] Oops[#1]:
[ 34.069561] CPU: 3 UID: 0 PID: 170 Comm: pipe_race Not tainted 7.2.0-rc6-01553-gb73c35220968-dirty #103 VOLUNTARY
[ 34.079932] Hardware name: Mobile EyeQ5 MP5 Evaluation board
[ 34.085592] $ 0 : 0000000000000000 0000000000000001 0000000000000000 0000000000000000
[ 34.093616] $ 4 : a800000808ee2618 000000000b7a879d 0000000000001000 0000000000000000
[ 34.101638] $ 8 : 0000000000e3f2c9 0000000000000000 a800000808a2a9f8 0000000000000000
[ 34.109660] $12 : a8000008139ffcd8 ffffffff84080018 a80000080837fae0 7878787878787878
[ 34.117682] $16 : a800000807e82940 0000000000001000 0000000000000000 0000000000000000
[ 34.125704] $20 : a800000802920e00 a8000008139ffdf8 a800000802649400 0000000000e3f2c9
[ 34.133726] $24 : 0000000000000006 00000001200406e0
[ 34.141783] $28 : a8000008139fc000 a8000008139ffd10 0000000000e3f2c8 a80000080838dfc4
[ 34.149837] epc : a80000080838e01c anon_pipe_read+0xd4/0x428
[ 34.155697] ra : a80000080838dfc4 anon_pipe_read+0x7c/0x428
[ 34.161549] Status: 140000e3 KX SX UX KERNEL EXL IE
[ 34.166551] Cause : 40800408 (ExcCode 02)
[ 34.170574] BadVA : 0000000000000000
[ 34.174161] PrId : 0001b028 (MIPS I6500)
[ 34.178183] Process pipe_race (pid: 170, threadinfo=000000005ca35720, task=00000000e1013890, tls=000000014ebbb780)
[ 34.188568] Stack : a800000802649400 0000000000000000 0000000000000000 a8000008139ffdd0
[ 34.196623] 0000000000000fba a800000808ee0000 0000000000000001 a8000008130c3e80
[ 34.204676] a8000008080d1280 a8000008139ffd58 a8000008139ffd58 1dbd2b22ea1dd500
[ 34.212729] a800000802649400 a800000808ee0000 ffffffffffffffea 0000000000000001
[ 34.220783] 0000000000001000 0000000000000000 00000001200ae518 ffffffffffffffff
[ 34.228836] 000000fffbe0e530 a80000080837edf4 000000fffbe0e530 0000000000000000
[ 34.236890] 0000000000000000 0000000000000000 000000014ebb55a0 0000000000001000
[ 34.244943] 0000000000000001 a800000802649400 0000000000000000 0000000000000000
[ 34.252996] 0000000000000000 0000400400000000 0000000000000000 1dbd2b22ea1dd500
[ 34.261049] 00000000140000e3 a800000802649400 a800000802649400 a800000808ee0000
[ 34.269103] ...
[ 34.271568] Call Trace:
[ 34.274026] [<a80000080838e01c>] anon_pipe_read+0xd4/0x428
[ 34.279533] [<a80000080837edf4>] vfs_read+0x25c/0x318
[ 34.284607] [<a80000080837faac>] ksys_read+0x104/0x138
[ 34.289763] [<a80000080802b9cc>] syscall_common+0x44/0x68
[ 34.295187]
[ 34.296689] Code: f84000cf 02209825 de020010 <dc420000> d8400004 02002825 0040f809 02802025 f84000c3
[ 34.306504]
[ 34.308099] ---[ end trace 0000000000000000 ]---
My initial reproducer was the xdp-tools test suite. A standalone
reproducer would be an lld/scd loop that, when the read is reordered by
the CPU, triggers a fault. We can achieve this from userspace by
stressing an anonymous pipe, which uses a mutex. Program used:
// SPDX-License-Identifier: GPL-2.0
// pipe_race.c - reproducer for MIPS LL/SC reordering vs fs/pipe.c
//
// Two userspace processes on an anonymous pipe:
// parent = writer: tight write() loop
// child = reader: tight read() loop
#define _GNU_SOURCE
#include <assert.h>
#include <errno.h>
#include <sched.h>
#include <signal.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/types.h>
#include
---truncated--- |
| In the Linux kernel, the following vulnerability has been resolved:
KVM: PPC: Book3S HV: fix use-after-free in kvmhv_emulate_tlbie_all_lpid()
kvmhv_emulate_tlbie_all_lpid() iterates the nested-guest IDR and drops
mmu_lock before calling kvmhv_emulate_tlbie_lpid(), but does not hold a
reference on the kvm_nested_guest pointer obtained from the IDR. A
concurrent vCPU issuing a single-LPID tlbie (is=2, ric=2) can race
through kvmhv_flush_nested() -> kvmhv_remove_nested() -> idr_remove /
--refcnt -> kvmhv_release_nested() -> kfree(gp) in that window, leaving
the iterating vCPU with a dangling pointer. The subsequent
mutex_lock(&gp->tlb_lock) and accesses to gp->shadow_pgtable,
gp->shadow_lpid and gp->l1_host all touch freed memory. The free path
is fully L1-controlled.
Fix this by incrementing gp->refcnt inside the loop before dropping
mmu_lock, mirroring what kvmhv_get_nested() does, and releasing the
reference with kvmhv_put_nested() after the per-guest work completes.
This is the same get/put discipline already used at every other
call site that drops mmu_lock while holding a nested-guest pointer. |
| In the Linux kernel, the following vulnerability has been resolved:
ntfs: protect runlist updates with the runlist lock
ntfs_non_resident_attr_shrink() calls runlist helpers that require the
runlist write lock, but did not hold it while freeing clusters and
truncating the runlist. Serialize those operations and the resident
conversion with the runlist lock.
ntfs_attr_map_cluster() can merge a newly allocated run before updating
mapping pairs. If the update fails, free the clusters and restore both
the in-memory runlist and on-disk mapping pairs from a saved runlist.
Mark the volume in error if either rollback step fails. |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/siw: Bound fragmented header copies by the remaining length
siw_get_hdr() can receive an extended DDP/RDMAP header across more than
one TCP callback. The first callback may receive most of the header,
while the next one still limits the copy to hdrlen - MIN_DDP_HDR instead
of the number of missing bytes. This makes the destination move past the
end of the header and overwrite the receive state, including
fpdu_part_rcvd. A later callback can then use a negative fpdu_part_rcvd
value as a copy offset, which creates an OOB write.
Use the number of header bytes already received when calculating the
next copy length. |
| In the Linux kernel, the following vulnerability has been resolved:
xfrm: save input state data before secpath resets
xfrm_input() stores the current xfrm_state in the skb secpath while it
continues receive-side processing. Some input paths can reset that secpath
before xfrm_input() has finished dereferencing the state.
Receive callback users such as VTI and XFRM interfaces can reset the
secpath. The VTI receive path does so before checking whether the packet
crosses network namespaces, while the XFRM interface path does so only for
cross-network-namespace packets. The XFRM_MAX_DEPTH error path can also
reset the secpath before the final drop callback reports the current
state's protocol.
If secpath_reset() drops the last state reference while the state is
concurrently deleted, xfrm_input() can still dereference the freed state
when selecting transport_finish() or reporting the drop callback protocol.
Save the state protocol on the stack while the state is still valid,
and use the already saved address family for transport_finish(). A larval
XFRM_STATE_ACQ state has no type, so retain nexthdr as its protocol. This
preserves the existing drop-path fallback while avoiding the post-reset
state dereferences without adding an extra state reference to every
received packet. |
| In the Linux kernel, the following vulnerability has been resolved:
wifi: libipw: reject too-short beacon and probe responses
libipw_process_probe_response() and the libipw_network_init() call it
makes assume the frame contains the full 36-byte beacon and probe
response prefix, but the ipw2100 and ipw2200 receive paths only
establish that a management frame carries the generic 24-byte
three-address header.
libipw_network_init() then computes the information element length as
stats->len - sizeof(*beacon)
stats->len is a u16 and sizeof() has type size_t, so the subtraction is
evaluated as size_t and wraps instead of going negative. Truncating
that to the u16 length parameter of libipw_parse_info_param() yields
65524 for a 24-byte beacon, and the parser then walks the receive
buffer as if it held almost 64 KiB of information elements, reading
past the allocation.
Reject the frame before any fixed field is touched.
Found by an AI-assisted review of length arithmetic in management frame
parsers. Verified with a KUnit case under Generic KASAN on arm64 under
QEMU; I do not have the hardware, so it is not tested on a real device. |
| In the Linux kernel, the following vulnerability has been resolved:
RDMA/rxe: insert mcg into mcg_tree only after rxe_mcast_add() succeeds
rxe_get_mcg() publishes a newly allocated multicast group in
rxe->mcg_tree before programming the backing Ethernet multicast address
with rxe_mcast_add(), which runs outside mcg_lock. A local userspace
RDMA client reaches this path with ATTACH_MCAST on a UD QP; if
rxe_mcast_add() then returns an error (for example -ENODEV when the
backing netdev has been removed, or a propagated dev_mc_add() error),
the unwind frees the published group without removing it from the tree.
A later lookup of the same MGID dereferences the freed struct rxe_mcg
from __rxe_lookup_mcg().
Fix this by keeping the new mcg private until rxe_mcast_add() succeeds.
Split the tree publication into __rxe_publish_mcg(), call rxe_mcast_add()
before taking the tree reference, and free the still-private mcg on
failure. Because the group is never visible in mcg_tree until the
multicast address is programmed, no concurrent caller can look it up or
attach a QP to a group that is about to be torn down, so the error path
needs no conditional unwind. If another caller publishes the same MGID
while the address is being programmed, the post-add re-check under
mcg_lock finds the winner; this caller then drops its private object and
balances its own rxe_mcast_add() with rxe_mcast_del() before returning
the winner.
Reproduced by forcing the rxe_mcast_add() error return under KASAN:
without the change the next attach to the same MGID reports a
slab-use-after-free in __rxe_lookup_mcg(); with it the forced failure
returns cleanly. A no-injection attach/detach regression, including a
two-QP shared join/leave and re-attach, stays KASAN- and leak-clean. |
| In the Linux kernel, the following vulnerability has been resolved:
net: lan743x: fix RX checksum use-after-free
lan743x_rx_process_buffer() adds each non-first receive buffer to the
head skb's frag_list. On the last descriptor, lan743x_rx_trim_skb()
linearizes the head and frees the fragment skb metadata.
The checksum-success path then writes ip_summed through the local skb
pointer, which still points to the final fragment. This causes a
use-after-free write when a packet spans more than one receive buffer.
Set ip_summed on the surviving head skb instead. Multi-buffer receive
can occur after a live MTU increase because existing ring entries keep
their old buffer size until they are replenished.
A KUnit test invoking lan743x_rx_process_buffer() with a two-buffer
packet produced a one-byte KASAN use-after-free write before this change.
The same test passed after the change. The driver object also builds
with W=1. This was not tested on physical LAN743x hardware. |
| In the Linux kernel, the following vulnerability has been resolved:
ipv6: xfrm: use full sockets in local error paths
xfrm6_local_rxpmtu() and xfrm6_local_error() dereference skb->sk as if it
always pointed at a full IPv6 socket.
That is not guaranteed. TCP SYN-ACK skbs can be owned by a
TCP_NEW_SYN_RECV request_sock while the output path itself is driven by the
full listener. If rerouting selects an IPv6 XFRM tunnel route with a lower
MTU, the local PMTU/error handling path can reach these callbacks with that
mini-socket still attached to the skb.
The callbacks then miscast the request socket as a full inet/IPv6 socket and
can read beyond the request_sock allocation when they access inet_sock or
ipv6_pinfo state.
Resolve the owner with skb_to_full_sk() in both callbacks and bail out when
no full socket is attached. This matches the surrounding XFRM IPv6 PMTU/error
logic, which already reasons about full sockets with skb_to_full_sk(). |
| In the Linux kernel, the following vulnerability has been resolved:
powerpc/iommu: Fix the overflow validation in iommu_tce_check_ioba
The commit b1af23d836f8 ("KVM: PPC: iommu: Unify TCE checking") unified
IOBA parameter checking across KVM and VFIO into iommu_tce_check_ioba().
While doing so, the passed in argument npages is ignored and constant
value '1' is used leaving out a possible overflow as the callers can
legitimately be using npages > 1 for H_STUFF_TCE or H_PUT_TCE_INDIRECT
cases.
Fix this by accounting for 'npages', checking for arithmetic overflow,
and verifying that the entire requested range (ioba - offset + npages)
does not exceed the table capacity 'size'. |
| In the Linux kernel, the following vulnerability has been resolved:
dma-buf/dma-fence: fix checking signaling bit for timeline and driver name v3
The patch "dma-buf: dma-fence: Fix potential NULL pointer dereference"
changed the check to test for the ops pointer instead of the signaled
bit to avoid a potential NULL dereference when the ops pointer has been
cleared.
The problem is now that the ops pointer is cleared only when neither the
release nor the wait callback is implemented and this isn't true for a lot
of dma_fence implementations yet. So those implementations lost the RCU
protection after signaling of the returned string resulting in potential
use after free.
Add the signaling check additional to the ops pointer check so that we
have both the protection against NULL dereference as well as the RCU
protection after signaling for the returned string.
v2: improve comments to note RCU protection and explain why we check
both signaling state and ops pointer
v3: some comment improvements suggested by Philip |
| In the Linux kernel, the following vulnerability has been resolved:
openvswitch: avoid reallocating confirmed conntrack labels
ovs_ct_get_conn_labels() adds the labels extension when a conntrack
entry does not have one. Confirmed conntracks can be read locklessly,
so adding an extension may reallocate and free the extension block
while another CPU accesses it.
Only add the extension for unconfirmed conntracks. A confirmed
conntrack without labels now fails the caller's label operation instead
of reallocating its extension storage. |
| In the Linux kernel, the following vulnerability has been resolved:
signal: Prevent exec() race
Hyunwoo debugged the following KASAN UAF splat:
BUG: KASAN: slab-use-after-free in __send_signal_locked+0xb27/0xba0
Write of size 8 at addr ffff888007ed80c8 by task poc/79
...
Call Trace:
__send_signal_locked+0xb27/0xba0
do_send_sig_info+0xa7/0x160
do_send_specific+0x76/0xa0
__x64_sys_tgkill+0x193/0x270
...
Allocated by task 80:
do_timer_create+0x1a4/0x1030
__x64_sys_timer_create+0x145/0x190
...
Freed by task 12:
kmem_cache_free_bulk+0x1f8/0x4a0
kvfree_rcu_bulk+0x14f/0x1c0
kfree_rcu_work+0x128/0x1a0
...
Last potentially related work creation:
kvfree_call_rcu+0x39/0x390
__flush_itimer_signals+0x211/0x320
flush_itimer_signals+0x47/0x90
begin_new_exec+0xa6b/0x28c0
It turned out that this happens with a non-leader exec() as Hyunwoo
explained:
de_thread() calls exchange_tids() before release_task(leader), so the
struct pid held by a SIGEV_THREAD_ID timer created against the leader's tid
now points to the thread which called execve(). pid_task() returns that
thread and lock_task_sighand() on it succeeds.
If the timer signal is blocked, its sigqueue stays queued on the leader's
task::pending. The next expiry of that timer can then run while
release_task() flushes the queue.
posixtimer_send_sigqueue() checks whether the sigqueue is already queued
with a plain list_empty(), which only reads list_head::next.
list_del_init() is not atomic and INIT_LIST_HEAD() stores list_head::next
before list_head::prev, so the check can pass in between. list_add_tail()
queues the entry on the task::pending of the live thread, and the
list_head::prev store from the flush then overwrites the list_head::prev
link that list_add_tail() has just set.
__flush_itimer_signals() does not undo that either. With list_head::prev
pointing at the entry itself, its list_del_init() only stores the same
values again, so the entry is not removed from the list. It is still there
after the last reference is dropped and the timer is freed by RCU, and the
list_add_tail() of a later tgkill() follows that list_head::prev into the
freed timer.
This problem surfaced with the recent commit which moved the sigqueue flush
out of the sighand lock held region.
Hyonwoo proposed to fix this by using list_del_init_careful(), but that
just papers over the problem. After some disucssions and various attempts
to solve it, Eric pointed out that there is no reason to flush
task::pending late in release_task() and it should be done in
exit_signals() already.
As nothing can collect and deliver signals which are queued in a dying
task's pending queue, there is no reason to delay it further.
But it has to be ensured that no signals can be queued into it after that
point. exit_signals() sets PF_EXITING in task::flags, which can be used as
an indicator for this.
Cure it by:
- Preventing signal queueing for task private signals (PIDTYPE_PID) when
the task has PF_EXITING set in __send_signal_locked() and in
posixtimer_send_sigqueue().
- Protecting the unlocked setting of PF_EXITING in exit_signals() for the
task group empty and the group exit case with sighand lock
- Flushing task::pending signals right there.
Optimize that by moving the whole pending list to an on-stack list head
under sighand lock and free the signals without the lock held.
There has been quite some discussion about the lockless flush and the
non-leader exec case on weakly ordered systems. The problem is that a third
party which tries to send a posix timer signal relies on the PID lookup to
find the target task and that lookup might result in the new leader when
the signal was originaly directed to the old leader. In case that the
signal was queued on the old leader then the lockless flush raised a
concern over the following situation:
old_leader new_leader third party
A: flush_list() // list_del_in
---truncated--- |
| In the Linux kernel, the following vulnerability has been resolved:
smb: client: validate absolute native symlink targets before NT fixups
With symlinkroot unset, an absolute target is copied without conversion
to an NT drive path. Later code still assumes an NT prefix is present
when modifying the target and calculating the print name length.
For "/ab", this causes two failures: sym[5] and path[5] are written
past their allocations, and plen -= 2 * poff subtracts an assumed
8-byte prefix from a 6-byte UTF-16 target, wrapping u16 plen to 65534.
That underflow causes another overflow: memcpy() copies 65534 bytes
into a 24-byte buffer. A user with write access to a mounted share
can trigger these bugs with default settings.
Validate the NT drive prefix, including an ASCII drive letter, before
accessing fixed offsets or subtracting the prefix length. |
| In the Linux kernel, the following vulnerability has been resolved:
netfilter: flowtable: hold reference on ct until flow is released
nf_ct_put() releases the ct->ext area inmediately, the rcu typesafe
semantics also allow to refer to the wrong conntrack from the flowtable
datapath. Hold reference on ct until flow is released after rcu grace
period.
Add rcu_barrier() on module exit path, to ensure pending flow entries
are release before module goes away. |
| In the Linux kernel, the following vulnerability has been resolved:
smb: client: fix use-after-free of iface in cifs_try_adding_channels()
cifs_try_adding_channels() iterates ses->iface_list with
list_for_each_entry_safe_from(), which captures the next entry
(niface) under iface_lock. The loop body then drops iface_lock for
the whole duration of cifs_ses_add_channel().
A concurrent interface refresh (SMB3_request_interfaces() ->
parse_server_interfaces()) marks all ifaces inactive and removes and
frees any that are not re-advertised via list_del() + kref_put(),
where release_iface() is a bare kfree(). Since niface typically has
no channel holding a reference, the list reference is its last and it
can be freed inside the unlocked window. On continue, the iterator
advance step then dereferences niface->iface_head.next, and the loop
body reads iface->rdma_capable/is_active, both on freed memory.
Fix this by never keeping an unreferenced list pointer across the
unlocked window. Each channel attempt now re-scans the list from the
head under iface_lock, takes a kref on the selected candidate, and
passes only that referenced candidate to cifs_ses_add_channel().
weight_fulfilled still tracks selection progress, so restarting the
scan preserves the original weighted distribution and the
weight_fulfilled-before-kref_put ordering on the failure path.
Add a per-pass attempts cap so a flapping interface refresh cannot
keep the inner loop spinning within a single tries increment. |