mirror of
https://git.kernel.org/pub/scm/linux/kernel/git/stable/linux.git
synced 2026-09-22 09:34:56 +02:00
75fa9caeb8aaba19c2463dee0b0a1e09d39c04af
1481010
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
75fa9caeb8 |
ipv6: mcast: fix delay calculation in igmp6_join_group()
When joining a multicast group, if a report work is already pending
(e.g. scheduled by a query or a previous join), igmp6_join_group()
cancels the delayed work and recalculates the delay:
if (cancel_delayed_work(&ma->mca_work)) {
refcount_dec(&ma->mca_refcnt);
delay = ma->mca_work.timer.expires - jiffies;
}
Unlike igmp6_group_queried(), igmp6_join_group() did not check
if delay >= interval. This leads to two issues:
1. If the timer has already expired (timer.expires <= jiffies), the
stale expiry is reused by mod_delayed_work(), causing the second
unsolicited report to fire on the very next tick without a
randomized delay.
2. If the timer was originally armed by a query with a large
maximum response delay, delay could exceed
unsolicited_report_interval(ma->idev).
Fix this by initializing delay to unsolicited_report_interval(ma->idev)
and re-randomizing it with get_random_u32_below(interval) when
delay >= interval, mirroring the logic in igmp6_group_queried().
Fixes:
|
||
|
|
c073d1b070 |
ipv6: mcast: use copy-on-write RCU updates in ip6_mc_source()
pmc->sflist is read locklessly under rcu_read_lock() by
inet6_mc_check() during packet reception in the UDP and RAW
multicast receive paths.
ip6_mc_source() mutated psl->sl_addr and psl->sl_count in-place
when adding or removing a source filter. Additionally, when expanding
the filter buffer, newpsl was published via rcu_assign_pointer()
before writing the new source into the array.
Because 16-byte struct in6_addr writes are not atomic and array
shifting is not synchronized with RCU readers, concurrent readers in
inet6_mc_check() could read torn IPv6 addresses or observe
duplicated/missed source entries.
Fix this by switching ip6_mc_source() to copy-on-write RCU updates:
allocate and fully populate newpsl before publishing it via
rcu_assign_pointer(), and reclaim the old filter via kfree_rcu(),
matching ip6_mc_msfilter().
Also remove the now unused IP6_SFBLOCK macro.
Fixes:
|
||
|
|
93b4923984 |
ipv6: mcast: fix RCU list diversion in ip6_mc_del1_src()
When removing a source filter whose count reaches zero, ip6_mc_del1_src()
unlinks psf from pmc->mca_sources. If the filter was previously active,
the code moved psf directly into pmc->mca_tomb by updating psf->sf_next.
Because pmc->mca_sources is traversed locklessly under RCU (e.g. by
ipv6_chk_mcast_addr()), mutating psf->sf_next before a grace period
elapses diverts concurrent readers to the tombstone list. Consequently,
readers miss remaining active sources in pmc->mca_sources and improperly
examine deleted tombstone entries.
Fix this by allocating a new tombstone node for pmc->mca_tomb (as done
in sf_setstate()) and retiring the original psf via kfree_rcu().
Fixes:
|
||
|
|
d8d4d1cf40 |
ppp: ppp_synctty: simplify tty disc_data access
Apply the same simplification as the preceding ppp_async change.
Fixes:
|
||
|
|
9feb069e5e |
ppp: ppp_async: simplify tty disc_data access
tty_ldisc_hangup() invokes the hangup callback while holding only a read lock on tty->ldisc_sem, so it can run concurrently with other line discipline callbacks. This currently forces async PPP to maintain separate lifetime protection around tty->disc_data. Line discipline close is called under the write lock during hangup processing. Remove the hangup callback and rely on close for teardown, as done for SLIP by commit |
||
|
|
2987ee196c |
igmp: convert struct ip_sf_list to RCU
Commit |
||
|
|
1ea9fff22b |
Merge tag 'for-net-2026-08-31' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth
Luiz Augusto von Dentz says: ==================== bluetooth pull request for net: Core: - hci_core: Fix race condition during device registration - L2CAP: fix chan mode for LE_CONN_REQ + EXT_FLOWCTL pchan - L2CAP: fix out-of-bounds write in l2cap_ecred_connect - L2CAP: clear FLAG_DEFER_SETUP only for same PID/PSM Drivers: - hci_mrvl: Fix wrong return value check of wait_on_bit_timeout() - btintel_pcie: Clear automask on spurious interrupts - btintel: validate version TLV value lengths - btintel: bound firmware ID by TLV length - btintel: propagate version TLV parsing errors * tag 'for-net-2026-08-31' of git://git.kernel.org/pub/scm/linux/kernel/git/bluetooth/bluetooth: Bluetooth: hci_mrvl: Fix wrong return value check of wait_on_bit_timeout() Bluetooth: L2CAP: clear FLAG_DEFER_SETUP only for same PID/PSM Bluetooth: L2CAP: fix out-of-bounds write in l2cap_ecred_connect Bluetooth: L2CAP: fix chan mode for LE_CONN_REQ + EXT_FLOWCTL pchan Bluetooth: hci_core: Fix race condition during device registration Bluetooth: btintel: propagate version TLV parsing errors Bluetooth: btintel: bound firmware ID by TLV length Bluetooth: btintel: validate version TLV value lengths Bluetooth: btintel_pcie: Clear automask on spurious interrupts ==================== Link: https://patch.msgid.link/20260831181837.946230-1-luiz.dentz@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> |
||
|
|
fee1065570 |
net/sched: cls_flower: validate mask pointer after nla_next()
fl_set_enc_opt() iterates the key's nested tunnel-option attributes with nla_for_each_attr() while advancing a single mask pointer via nla_next() at the bottom of each loop, so the mask cursor is driven by the number of key attributes rather than by the mask's own attributes. The nla_ok() added by commit |
||
|
|
9d0f206eb3 |
Merge branch 'vsock-validate-packet-sources-after-bound-lookup-fallback'
Daehyeon Ko says: ==================== vsock: validate packet sources after bound lookup fallback Both virtio and VMCI look up connected sockets by the full tuple before falling back to a destination-only bound lookup. The fallback can select a non-listening socket without validating the packet source. V2 covered only the virtio path. Following Stefano's review, this series moves the source and transport validation into a documented AF_VSOCK helper and uses it for both virtio and VMCI. The VMCI patch checks both its bottom-half and deferred workqueue receive paths. V4 preserves VMCI's existing RST behavior when source validation fails. The reset is addressed from the received packet so that a bound but non-listening or concurrently closed socket still notifies the sender, without directing the reset to a connected socket's stored peer. The v3 regression was reproduced in three x86_64 KASAN boots: a REQUEST to a bound but non-listening socket returned VMCI_ERROR_NO_ACCESS but no RST arrived within one second. With v4, the sending context received the expected RST in all three boots. The original VMCI source-validation oracle also passed in three v4 boots: a matched RST reset the pending socket while a mismatched-context RST left it pending. No KASAN report occurred. Patch 1 is unchanged from v3 (identical stable patch-id) and carries Bobby's Reviewed-by for that revision. Its v3 validation covered the cross-UID injection oracle, local CID aliases, selected VSOCK selftests, and W=1 changed-object builds under allmodconfig and allyesconfig. The current-tree guest-CID vhost probe could not be rerun because the test user lacks access to /dev/vhost-vsock. ==================== Link: https://patch.msgid.link/20260826003929.966160-1-4ncienth@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> |
||
|
|
ad9a7da3fa |
vsock/vmci: validate packet source for connected sockets
vmci_transport_recv_stream_cb() looks up sockets first by the full source
and destination tuple, then by destination only in the bound table. The
fallback can select a non-listening socket without checking whether the
packet came from its stored peer.
This was reproduced with two VMCI contexts. A RST from the context not
stored in a TCP_SYN_SENT socket reset that socket after it was selected by
the destination-only lookup.
VMCI can process notification packets in bottom-half context when the
socket is not owned by user context, or defer packets to a workqueue. Use
vsock_check_source() after taking the socket lock in the bottom-half path,
and recheck after lock_sock() in the workqueue path. Listening sockets
continue to accept packets from any source.
Reply with a RST addressed from the received packet before dropping a
source that fails validation. This preserves the existing reset behavior
for bound non-listening and concurrently closed sockets without directing
the reset to a connected socket's stored peer.
Fixes:
|
||
|
|
dee44f41f2 |
vsock/virtio: validate packet source for connected sockets
virtio_transport_recv_pkt() looks up sockets first by the full source and
destination tuple, then by destination only in the bound table. The
fallback is needed for listening and connecting sockets, but sockets remain
in the bound table after connect(), so it can also return a non-listening
socket.
The fallback does not validate the source address. In TCP_SYN_SENT, a
RESPONSE from an unrelated source can transition the victim socket to
TCP_ESTABLISHED while its stored remote address remains unchanged.
Subsequent RW packets from that source are delivered through the same
destination-only fallback.
This was reproduced with capability-empty processes under different UIDs.
The attacker discovered the target tuple through unprivileged AF_VSOCK
sock_diag and caused the victim socket to read 16 attacker-chosen bytes;
the intended peer-side socket read 0 of those 16 bytes.
Add vsock_check_source() to validate the transport, source port and source
CID against the peer stored in a non-listening socket. The local transport
is the CID exception because its packets are generated internally with
VMADDR_CID_LOCAL as their source, including connections using CID aliases.
Use the helper after lock_sock() in the virtio receive path.
Fixes:
|
||
|
|
dc0df5a0c6 |
page_pool: keep frag_offset aligned for odd-sized requests
page_pool_alloc_frag_netmem() rounds the requested fragment size with
size = ALIGN(size, dma_get_cache_alignment());
dma_get_cache_alignment() returns 1 unless the architecture defines
ARCH_DMA_MINALIGN, which DMA-coherent architectures such as x86 do not.
There the ALIGN() is a no-op and pool->frag_offset advances by the raw,
unrounded size.
A single caller asking for an odd size then leaves frag_offset misaligned
for every fragment carved out of that page afterwards. The pool is shared,
so the damage is not confined to the caller that caused it.
The per-cpu system_page_pool used by generic XDP hits this.
skb_pp_cow_data() allocates its fragments with the raw packet length:
size = min_t(u32, len, PAGE_SIZE);
truesize = size;
page = page_pool_dev_alloc(pool, &page_off, &truesize);
leaving frag_offset odd for whatever is carved out of that page next. Its
own head allocation is already aligned -- SKB_HEAD_ALIGN(size) plus the
XDP_PACKET_HEADROOM its callers pass -- so it is a later user of the shared
pool that pays: page_pool_dev_alloc_va() returns a misaligned buffer,
napi_build_skb() installs it as skb->head, and skb_shinfo(skb) ==
skb->head + skb->end is misaligned with it.
skb_shinfo()->dataref is a 4-byte atomic_t at offset 0x20, so the
atomic_inc() in __skb_clone() straddles a cache line. On x86 with split
lock detection -- fatal for kernel split locks by default -- this panics
the machine:
Oops: Split lock detected
RIP: 0010:skb_clone+0x154/0x1e0
Call Trace:
<IRQ>
raw_local_deliver+0x1ed/0x2c0
ip_protocol_deliver_rcu+0x54/0x1c0
ip_local_deliver_finish+0x85/0x100
ip_local_deliver+0x67/0x100
__netif_receive_skb_one_core+0x85/0xa0
process_backlog+0x87/0x130
Reproduced by attaching any generic-mode XDP program to loopback and
opening a RAW IPPROTO_UDP socket, which makes raw_local_deliver() clone
every locally delivered UDP packet; ordinary DNS traffic then triggers it,
roughly once per 2500 clones. Observed on 6.12.101 and 7.1.8.
Tracing page_pool_alloc_frag_netmem() over one such run shows the
amplification -- two odd-sized requests, nine misaligned offsets:
requested size & 7: 0: 17035 5: 1 7: 1
frag_offset & 7: 0: 17028 3: 1 4: 1 5: 1 6: 1 7: 5
and skb_pp_cow_data() returning heads that were aligned on entry:
head 0xffff8f4c86aeac00 -> 0xffff8f4c53a9a9c4 (&7=4)
head 0xffff8f4d6a8a42c0 -> 0xffff8f4c4f7b7a45 (&7=5)
Round the fragment size up to at least the alignment struct skb_shared_info
requires, so fragments are always suitably aligned for the objects callers
build on them. Architectures needing a larger DMA alignment keep it.
This also makes the remainder computed in page_pool_alloc_netmem(),
*size = max_size - *offset;
aligned, since max_size is a power of two -- which fixes the matching
misalignment of skb->end.
Verified with a controlled A/B under QEMU/KVM: same tree, same config,
same compiler, same rootfs and identical traffic, differing only by this
patch. A SEC("xdp.frags") XDP_PASS program on lo plus UDP datagrams
larger than max_head_size drives skb_pp_cow_data()'s fragment loop, which
passes raw packet lengths to the pool. Measured at the return of
skb_pp_cow_data():
unpatched patched
skb_pp_cow_data calls 40800 40800
misaligned skb->head 1120 0
dataref at line offset >60 80 0
The last row counts the accesses that actually fault:
skb_shinfo()->dataref sits at head+end+0x20 and is a 4-byte atomic, so
`lock incl` splits a 64-byte cache line only when that address lands at
offset 61..63. All 80 occurrences were at offset 61; the panic reported
above was at offset 62. Eliminating the misalignment removes every one
of them.
Same class of bug as commit
|
||
|
|
7b120a7719 |
selftests: tc-testing: add u32 node ID pool exhaustion test
Add a tdc test case that fills the u32 node ID space with 4095 auto-generated handles, then attempts to add a 4096th. On the fixed kernel the 4096th filter is rejected with ENOSPC (exit 2). On the unfixed kernel it silently succeeds with a duplicate handle. The setup pipes the 4095 add commands directly into `tc -b -` inside a single bash -c (matching the existing test id 1234 pattern), avoiding any temp file. Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com> Link: https://patch.msgid.link/20260825081052.133898-2-jhs@mojatatu.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> |
||
|
|
d7e7e98d23 |
net/sched: cls_u32: fix duplicate handle when node ID pool is exhausted
gen_new_kid() falls back to returning max (htid | 0xFFF) when both
idr_alloc_u32() ranges are full, instead of reporting an error.
u32_change() trusts that value and inserts a new knode with a handle
that is already live in the hash table, breaking handle uniqueness
within the table's node ID space.
The handle was never reserved in ht->handle_idr, so every later error
path that does idr_remove(&ht->handle_idr, handle) removes the
reservation of a different, live knode, which is then reused — one
failed add compounds into further duplicates.
The 4095 limit is per (table, bucket) — ht->handle_idr is per hash
table and the range is derived from htid (bucketid), so a table with
divisor 256 can legitimately hold 256*4095 knodes.
The sibling helper gen_new_htid() has the same silent in-band failure:
it returns 0 when the tp_c handle pool (1..0x7FF) is full, and
u32_init() publishes the root hash table with handle 0 without
checking. Two root tables with handle 0 alias in u32_lookup_ht(),
allowing cross-tcf_proto knode add/lookup/delete. Add the same
exhaustion check that the divisor path already has.
Return an error so u32_change() fails with ENOSPC/ENOMEM when the
node ID space is exhausted, and so u32_init() fails with -ENOMEM
when the hash table ID space is exhausted. The extack message
distinguishes pool exhaustion (-ENOSPC) from a transient allocation
failure (-ENOMEM).
Conditions to recreate the bug:
- CONFIG_NET_SCHED=y, CONFIG_CLS_U32=y (or =m with module loaded)
- Create a clsact qdisc on a device, then add 4095 u32 filters with
auto-generated handles to fill the node ID space for the root hash
table (single bucket). The 4096th auto-handle filter add triggers
the duplicate handle (fh 800::fff reused). Reachable at Level 2
(unshare -Urn, namespace-local CAP_NET_ADMIN).
- For gen_new_htid: create 2047 u32 proto entries on the same block
to fill the tp_c handle pool, then create one more. The root table
gets handle 0 and aliases with other handle-0 root tables.
Fixes:
|
||
|
|
f05f0d85cd |
Merge branch 'fix-to-possible-skb-leak-due-to-race-condtion-in-tx-path'
Selvamani Rajagopal says: ==================== Fix to possible skb leak due to race condtion in tx path Now the traffic is handled in threaded IRQ, and the disable_traffic flag is checked before handling the data, new race condition is exposed, in which buffer may leak, if threaded IRQ interrupts the trasmit path midway. With this change, disable_traffic and waiting_tx_skb pointer are protected by spin lock/unlock pair. This is highlighted in Sashiko review https://netdev-ai.bots.linux.dev/sashiko/#/patchset/20260611-level-trigger-v5-0-4533a9e85ce2%40onsemi.com Also on buffer overrun condition, probably due to loss of SPI data chunks, receive path doesn't see the expected data chunk with end_valid bit set. As a result, driver keeps adding data chunks to the skb before running out of space and kernel panic is seen. With this change, before adding data to the skb, if there is no space, skb is freed and driver starts looking for new frame by looking for a data chunk with start_valid bit set. [ 705.405490] skbuff: skb_over_panic: text:ffffffd2eb72a264 len:1600 put:64 head:ffffff804e5cdc40 data:ffffff804e5cdc80 tail:0x680 end:0x640 dev:eth1 [ 705.405569] ------------[ cut here ]------------ [ 705.405575] kernel BUG at net/core/skbuff.c:214! [ 705.405589] Internal error: Oops - BUG: 00000000f2000800 [#1] SMP [ 6703.427690] Call trace: [ 705.925157] skb_panic+0x58/0x68 (P) [ 705.928726] skb_put+0x74/0x80 [ 705.931772] oa_tc6_update_rx_skb+0x44/0x98 [oa_tc6_mod] [ 705.937084] oa_tc6_macphy_threaded_irq+0x3f4/0x900 [oa_tc6_mod] [ 705.943084] irq_thread_fn+0x34/0xb8 [ 705.946654] irq_thread+0x1a0/0x300 [ 705.950134] kthread+0x138/0x150 [ 705.953356] ret_from_fork+0x10/0x20 ==================== Link: https://patch.msgid.link/20260824-fix-race-condition-and-crash-v7-0-4323279b18f2@onsemi.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> |
||
|
|
3cc2aa96b9 |
net: ethernet: oa_tc6: Fix for the wrong data type
Inadvertently bool data type is used where int is supposed to
be used. This might turn a negative error code into true or
false and sign of the return code would be lost.
Fixes:
|
||
|
|
349c366365 |
net: ethernet: oa_tc6: Disable tx queues on fatal error
Previously, TX queue interface was stopped when
disable_traffic flag was set, which would indicate fatal
error. It is more appropriate to disable the queue as,
unless driver is unloaded and reloaded, there is no recovery
after disable_traffic is set.
Queues may be re-enabled inadvertently by other layers.
Intention of disable_traffic is only to stop the traffic
from flowing on fatal error.
Fixes:
|
||
|
|
172c974113 |
net: ethernet: oa_tc6: Improve the error recovery
When oversubscribed traffic causes lot of buffer overflow errors,
probably due to loss of data chunks, driver fails to find a
data chunk with end_valid bit set, before it runs out of sk buffer
space. As a result, assert is seen during skb_put.
Now, check is made if skb buffer has enough tailroom for the
incoming data before accepting. If there is no room, current
frame is abandoned and it will start looking for a data chunk
with start_valid bit, that is a new frame.
SK buffer allocation error is considered as recoverable error.
rx_buf_overflow flag is too specific and no longer the only
condition this flag is used for. Therefore it is renamed as
wait_until_start_valid. This is more appropriate as this flag
is used to look for the next data chunk with SV bit set, after
failures like buffer overflow, buffer allocation failure, skb pointer
validity besides buffer overflow error.
Not writing to status0 if it reads 0.
Fixes:
|
||
|
|
5443d9c4f5 |
net: ethernet: oa_tc6: Protect skb pointer used by two different kernel instances
Threaded IRQ uses waiting_tx_skb. Transmit path also uses this pointer
without any mutual exclusion protection. As a result, it might leak skb
buffer, particularly if threaded IRQ sets disable_traffic true after
start_xmit already checked and found that disable_traffic being false,
if they happen to run on different cores.
On fatal error, where disable_traffic is set, transmit function drops the
packet and return NETDEV_TX_OK. Due to this change, skb_linearize call
is moved up to the beginning of the transmit function.
Since skb buffer may be freed from different contexts, dev_kfree_skb_any
is used to free skb buffer now, replacing one of the kfree_skb call.
oa_tc6_exit disables the irq before setting disable_traffic true.
Fixes:
|
||
|
|
fa5acd038e |
net/iucv: fix the recvmsg window update
iucv_sock_recvmsg() sends the HiperSockets-only AF_IUCV_FLAG_WIN without testing the transport, so on a classic z/VM socket iucv_send_ctrl() sizes the skb through a NULL iucv->hs_dev. SO_MSGLIMIT accepts 1, so msglimit / 2 is zero and one recvmsg() on its own socket is enough for an unprivileged process to take a spurious disconnect. It also calls iucv_send_ctrl() under spin_lock_bh(&message_q.lock), which allocates GFP_KERNEL inside a section the code treats as atomic. Sending outside that lock lets two recvmsg() reach afiucv_hs_send() at once, where msg_recv is sampled for the advertised window and subtracted after dev_queue_xmit() -- and sendmsg reaches that counter under lock_sock() while recvmsg holds no socket lock, so both can subtract the same value, the counter goes negative and the credit reaches the peer twice. Test the transport, claim the credit with atomic_xchg() after the last error exit and hand it back if the transmit fails, and send once the lock is dropped. Fixes: |
||
|
|
2deb76c21b |
Bluetooth: hci_mrvl: Fix wrong return value check of wait_on_bit_timeout()
wait_on_bit_timeout() returns 0 if the bit was cleared, -EINTR if the
process received a signal and the mode permitted wake up on that signal,
or -EAGAIN if the timeout elapsed. It never returns 1.
Hence the check "err == 1" in mrvl_load_firmware() is dead code: when
the waiting task is interrupted by a signal (-EINTR), the code falls
into the "else if (err)" branch and misreports it as "Firmware request
timeout" with -ETIMEDOUT instead of propagating -EINTR.
Fix this by testing for -EINTR so that an interrupted firmware load is
properly detected and reported.
Fixes:
|
||
|
|
0d77683237 |
Bluetooth: L2CAP: clear FLAG_DEFER_SETUP only for same PID/PSM
l2cap_ecred_defer_connect() clears FLAG_DEFER_SETUP also for channels
with different PID/PSM, which will not be added to the same
ECRED_CONN_REQ in any case. Consequently, only one ECRED connection
group can work at a time although it appears intended they would be
separate for each PID/PSM combination.
Fix by clearing FLAG_DEFER_SETUP only for the connections that could be
added in the request. Retain test_bit(FLAG_DEFER_SETUP) before calling
get_peer_pid as it may be NULL otherwise.
Fixes:
|
||
|
|
56c2b5831d |
Bluetooth: L2CAP: fix out-of-bounds write in l2cap_ecred_connect
l2cap_chan_connect() tries to ensure there are no more than
L2CAP_ECRED_CONN_SCID_MAX pending ECRED channels, so they fit in the
same L2CAP_ECRED_CONN_REQ that l2cap_ecred_connect() constructs.
However, the check only counts deferred channels. If 6 L2CAP sockets
are connected at the same time in order DDDDND (D=deferred,
N=non-deferred), the last can bump the total to max+1. It results to
one __le16 written out of bounds of the scid array, and an invalid
ECRED_CONN_REQ being sent.
Fix by leaving room for the non-deferred pending ECRED channels in the
counting in l2cap_chan_connect(), so the limit can't be exceeded.
Move counting under same critical section where the channel is added.
Although race conditions involving this appear unreachable, it's easier
to see.
Also add WARN_ON_ONCE check in l2cap_ecred_defer_connect() to make this
less brittle.
Fixes:
|
||
|
|
4ef05db5b0 |
Bluetooth: L2CAP: fix chan mode for LE_CONN_REQ + EXT_FLOWCTL pchan
l2cap_new_connection() sets default value of channel mode to match the
parent channel. l2cap_le_connect_req() left this at the default, and
created L2CAP_MODE_EXT_FLOWCTL channels if listening pchan has that
mode. This causes FLAG_DEFER_SETUP channels to reply to
L2CAP_LE_CONN_REQ with L2CAP_ECRED_CONN_RSP, which is incorrect.
It can also result to stack OOB write (of l2cap_alloc_cid determined
values) in l2cap_ecred_rsp_defer(), as l2cap_le_connect_req() does not
limit maximum number of deferred channels or check for duplicate ident.
Fix by setting chan->mode correctly in l2cap_le_connect_req().
Also check channel mode in l2cap_ecred_rsp_defer(), and do WARN_ON_ONCE
instead of OOB write to make it less brittle.
Fixes:
|
||
|
|
57938bbdb9 |
Bluetooth: hci_core: Fix race condition during device registration
In hci_register_dev(), the power_on work item is queued to
hdev->req_workqueue before initializing hdev->adv_monitors_idr and
registering the MSFT extension via msft_register(). For devices marked with
quirks such as HCI_QUIRK_RAW_DEVICE, the HCI_UNCONFIGURED flag is set on
the device. When the power_on work item runs concurrently on another CPU,
hci_power_on() detects that the device is unconfigured and immediately
invokes hci_dev_do_close(), which calls msft_do_close().
Concurrently, msft_register() allocates the msft structure and exposes it
to hdev->msft_data prior to calling mutex_init(&msft->filter_lock). If
msft_do_close() executes while hdev->msft_data is already assigned but the
mutex has not yet been initialized, mutex_lock(&msft->filter_lock) operates
on an uninitialized mutex, triggering a DEBUG_LOCKS warning:
DEBUG_LOCKS_WARN_ON(lock->magic != lock)
WARNING: kernel/locking/mutex.c:625 at __mutex_lock_common
kernel/locking/mutex.c:625 [inline]
WARNING: kernel/locking/mutex.c:625 at __mutex_lock+0x12d8/0x1550
kernel/locking/mutex.c:821
...
Call Trace:
<TASK>
msft_do_close+0x308/0x7b0 net/bluetooth/msft.c:693
hci_dev_close_sync+0x86b/0x10a0 net/bluetooth/hci_sync.c:5522
hci_dev_do_close net/bluetooth/hci_core.c:499 [inline]
hci_power_on+0x32c/0x750 net/bluetooth/hci_core.c:937
process_one_work kernel/workqueue.c:3322 [inline]
process_scheduled_works+0xa8e/0x14e0 kernel/workqueue.c:3405
worker_thread+0x92d/0xe10 kernel/workqueue.c:3486
kthread+0x388/0x470 kernel/kthread.c:436
ret_from_fork+0x514/0xb70 arch/x86/kernel/process.c:158
ret_from_fork_asm+0x1a/0x30 arch/x86/entry/entry_64.S:245
</TASK>
Fix this by moving the queue_work() call in hci_register_dev() to after
idr_init(&hdev->adv_monitors_idr) and msft_register(hdev) so that device
structures and extensions are fully initialized before asynchronous tasks
can access them. Additionally, assign hdev->msft_data in msft_register()
only after mutex_init(&msft->filter_lock) has completed.
Fixes:
|
||
|
|
3a74624b5d |
Bluetooth: btintel: propagate version TLV parsing errors
btintel_read_version_tlv() ignores the parser return value, so setup continues with partially initialized version data after a malformed TLV causes parsing to stop. Return the parser error to the caller so an invalid response fails setup instead of being treated as successful. Keep this behavioral change separate from the bounds checks so it can be reverted independently if an existing controller sends malformed data. Signed-off-by: Laxman Acharya Padhya <acharyalaxman8848@gmail.com> Reviewed-by: Ali Ahmet Memis <ali@iusegentoo.com> Tested-by: Kiran K <kiran.k@intel.com> Signed-off-by: Luiz Augusto von Dentz <luiz.von.dentz@intel.com> |
||
|
|
ac8aa9e0ec |
Bluetooth: btintel: bound firmware ID by TLV length
The firmware ID is treated as a NUL-terminated string even though the
TLV length is its only boundary. If the value does not contain a NUL
terminator, snprintf() can read beyond the received response.
Limit the conversion to the advertised TLV value length.
Fixes:
|
||
|
|
a086c08929 |
Bluetooth: btintel: validate version TLV value lengths
btintel_parse_version_tlv() verifies that a complete TLV is present in
the response, but it does not ensure that the value is long enough for
the specific TLV type. A short value can therefore cause an
out-of-bounds read through get_unaligned_le16(), get_unaligned_le32(),
or memcpy().
Reject values shorter than the minimum required by each known TLV type.
Also reject responses that do not contain the Command Complete Status
field.
Fixes:
|
||
|
|
ea2ee8b222 |
Bluetooth: btintel_pcie: Clear automask on spurious interrupts
On spurious interrupt where the TX and RX causes are not set, driver was
not clearing the auto mask which can block all the interrupts. Driver
needs to clear the automask even if no causes are set.
Fixes:
|
||
|
|
1376afc766 |
octeontx2-af: fix CN20K default MCAM rule removal on port cleanup
npc_mcam_free_all_entries() disables every MCAM entry mapped to a
port before freeing it. On CN20K, that also disables the default
broadcast, multicast, promiscuous, and unicast rules, which causes
packet drops when all rules are removed per port.
Only disable and free non-default entries. Leave CN20K default rules
enabled when freeing the remaining port entries.
Fixes:
|
||
|
|
a8455260b2 |
ipvlan: unregister upper devices outside pnodes_lock
syzbot reported the following circular locking dependency:
xs->mutex -> netdev lock -> pnodes_lock -> net->xdp.lock -> xs->mutex
The pnodes_lock -> net->xdp.lock edge is recorded when
ipvlan_device_event(NETDEV_UNREGISTER) calls unregister_netdevice_many()
while holding pnodes_lock. A nested NETDEV_UNREGISTER notification for
an IPvlan device enters xsk_notifier(), which acquires net->xdp.lock.
Keep pnodes_lock only while marking the upper devices as dying, removing
them from port->ipvlans, and queueing them for unregistration. Once the
devices have been detached from the protected list, release pnodes_lock
before unregister_netdevice_many() invokes notifier callbacks.
The port remains alive across unregistration because
ipvlan_device_event() holds the reference acquired by ipvlan_port_get().
The dying flag prevents a concurrent ->dellink() callback from deleting a
queued device again.
Fixes:
|
||
|
|
4aa61c88b4 |
vxlan: mdb: Fix use-after-free in vxlan_mdb_remote_src_del()
vxlan_mdb_is_valid_source(), which validates MDBE_ATTR_SOURCE and every
MDBE_ATTR_SRC_LIST member, accepts the all-zeros address.
A source list is only accepted on a (*, G) entry, whose source is the
all-zeros address, and for each member of the list an (S, G) entry is
derived from it by substituting the source. Entries are keyed by a plain
memcmp() of struct vxlan_mdb_entry_key, so if MDBE_ATTR_SOURCE is present
and holds the all-zeros address and the source list holds it as well, the
derived (S, G) key is byte-identical to the (*, G) key and resolves to the
same entry. Omitting MDBE_ATTR_SOURCE is not equivalent, as the key is
then left with a zero address family.
vxlan_mdb_remote_src_del() removes the forwarding entry of a source before
freeing the source entry:
vxlan_mdb_remote_src_fwd_del(vxlan, group, remote, &ent->addr);
vxlan_mdb_remote_src_entry_del(ent);
With the keys aliased, the first call deletes the remote of the entry that
owns 'ent' instead of a separate (S, G) entry, and frees 'ent'. The second
call then runs on the freed entry, and its hlist_del() reads ->pprev and
->next out of it and writes through them.
Adding the (*, G) entry with NLM_F_REPLACE and no source list marks the
all-zeros source for deletion and reaches this from the sweep at the end
of vxlan_mdb_remote_srcs_replace().
BUG: KASAN: slab-use-after-free in __vxlan_mdb_add+0x1cd/0xd70
Read of size 8 at addr ffff888102852500 by task poc/84
__vxlan_mdb_add+0x1cd/0xd70
vxlan_mdb_add+0xc0/0x140
rtnl_mdb_add+0x157/0x2a0
rtnetlink_rcv_msg+0x207/0x5a0
Allocated by task 84:
__kmalloc_cache_noprof+0x153/0x360
vxlan_mdb_remote_srcs_add+0x2eb/0x440
__vxlan_mdb_add+0x803/0xd70
Freed by task 84:
kfree+0x14c/0x3b0
vxlan_mdb_remote_del+0x129/0x1a0
__vxlan_mdb_del+0x4f/0xe0
vxlan_mdb_remote_src_fwd_del.isra.0+0x162/0x1b0
__vxlan_mdb_add+0x1c5/0xd70
The MDB operations are netns-scoped, so an unprivileged user can perform
them in a new user and network namespace.
Reject the all-zeros address in vxlan_mdb_is_valid_source(), which covers
both call sites. A (*, G) entry is expressed by omitting the source, so
nothing legitimate is refused.
Discovered by XBOW, triaged by Baul Lee <baul.lee@xbow.com>
Fixes:
|
||
|
|
6cfc1b90cb |
sctp: validate chunk length in the inqueue parser
SCTP chunks always include a four-byte generic header, but sctp_inq_pop() currently accepts shorter declared lengths. A zero-length chunk leaves chunk_end at the current header. When ASCONF is covered by the association's SCTP-AUTH policy, sctp_assoc_bh_rcv() can continue before the state machine performs its normal chunk-length check. sctp_inq_pop() then returns the same malformed chunk repeatedly and the receive softirq can lock up. A remote SCTP peer can trigger this after establishing an association on a kernel built with CONFIG_IP_SCTP and configured with net.sctp.addip_enable=1 and net.sctp.auth_enable=1. The reproducer did not require application credentials, a shared SCTP AUTH key, or net.sctp.addip_noauth_enable=1. On commit |
||
|
|
ac8d6b28d4 |
net: amd-xgbe: discard rx packets with bad FCS
amd-xgbe driver currently sets the MAC_RCR.DCRCC bit whenever
RX is enabled. This disables hardware FCS validation, causing packets
with bad FCS to be accepted unconditionally.
This change unsets DCRCC so that packets with bad FCS will be dropped,
in-line with typical behaviours of many other network controllers.
Tests:
- Verified that packets with bad FCS are now dropped.
- Verified that receiving packets with bad FCS will increment the
`rx_crc_errors` counter.
Fixes:
|
||
|
|
ac08d183da |
raw: annotate disconnect-side IPv4 match writers
raw_v4_match() reads inet_daddr, inet_rcv_saddr and sk_bound_dev_if locklessly under RCU. Bind and connect writers are annotated, but __udp_disconnect() still clears the same fields using plain stores. Commit |
||
|
|
2cb0b0b1ed |
sctp: fix soft lockup from unpadded ASCONF-ACK parameter iteration
sctp_verify_asconf() walks ASCONF-ACK parameters with
sctp_walk_params(), which advances by SCTP_PAD4(length), while the
consumer sctp_get_asconf_response() iterates the same parameters
advancing by the raw length, without padding. A single odd-length
parameter desynchronises the two walks and makes the consumer
interpret attacker-controlled bytes at a misaligned offset.
When those bytes yield a length of zero, the while loop over
asconf_ack_len makes no progress, spinning forever in softirq
context, and the watchdog reports a soft lockup. All reads stay
within the received skb, so the lockup is a pure remote denial of
service. A remote peer can trigger it with a crafted ASCONF-ACK on
an ADD-IP enabled association with an outstanding ASCONF (RFC 5061
section 4.1.2 requires the chunk to be authenticated, but the
predefined empty key id 0 allows the peer to compute the same
association HMAC from publicly exchanged parameters, so the gate
does not help).
The SCTP_PARAM_ERR_CAUSE case of sctp_verify_asconf() also performs
no length check, letting a parameter without a complete error
header reach the consumer, which reads errhdr.cause past the end of
the parameter, an out-of-bounds read.
Reject SCTP_PARAM_ERR_CAUSE parameters shorter than
sizeof(struct sctp_addip_param) + sizeof(struct sctp_errhdr) at the
verifier, and advance the consumer iterator with the same padding
rule as the verifier to keep the two walks in lockstep. The verifier
change guarantees a complete error header in every ERR_CAUSE
parameter the consumer can see, so the consumer's asconf_ack_len
check is dropped and it returns err_param->cause directly. The
consumer padding fix is still required because odd lengths remain
valid for SCTP_PARAM_ERR_CAUSE per RFC 5061.
The issue was found by ZeroHive, a vulnerability hunting agent at
Tencent Yunding Lab.
Fixes:
|
||
|
|
2188569e7e |
sctp: fix a TOCTOU race in SCTP_CMD_TIMER_START
The SCTP_CMD_TIMER_START handler checks timer_pending() before calling
timer_reduce(). The timer can expire and detach between these operations,
causing timer_reduce() to rearm the timer without taking the association
reference required for the newly armed timer.
The timer callback later unconditionally drops its association reference,
which can leave the association reference count unbalanced and result in
use-after-free during association teardown.
Use the return value of timer_reduce() to determine whether the timer was
actually armed. Take the association reference only when timer_reduce()
successfully starts a new timer, closing the race between checking the
timer state and rearming it.
This issue was reported by Nico Yip (@_cyeaa_) working with TrendAI Zero
Day Initiative.
Fixes:
|
||
|
|
b84cc38f3f |
Merge branch 'tcp-fix-use-after-free-in-do_tcp_getsockopt'
Cen Zhang says: ==================== tcp: fix use-after-free in do_tcp_getsockopt() do_tcp_getsockopt() has two lockless reads of icsk_ca_ops. Since BPF struct_ops congestion control made icsk_ca_ops point to dynamically allocated memory, a concurrent setsockopt(TCP_CONGESTION) can replace the pointer and free the old object while either reader is using it. Patch 1 fixes the TCP_CONGESTION path by copying ca_ops->name to a stack buffer while holding rcu_read_lock(). It also uses READ_ONCE() for the lockless load and annotates the relevant icsk_ca_ops stores with WRITE_ONCE(). Patch 2 fixes the TCP_CC_INFO path by keeping the READ_ONCE() load, ca_ops->get_info lookup, and call inside an RCU read-side critical section. ==================== Link: https://patch.msgid.link/cover.1787870710.git.blbllhy@gmail.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> |
||
|
|
385e474086 |
tcp: fix use-after-free in do_tcp_getsockopt(TCP_CC_INFO)
do_tcp_getsockopt() reads icsk->icsk_ca_ops and dereferences the
get_info function pointer without rcu_read_lock(). With BPF struct_ops
congestion control, ca_ops can point to dynamically allocated memory
that is freed concurrently, resulting in a use-after-free when the
kernel dereferences or calls through the stale pointer.
BUG: KASAN: slab-use-after-free in do_tcp_getsockopt+0x2037/0x23e0
Read of size 8 at addr ffff888013701258 by task exploit/149
do_tcp_getsockopt+0x2037/0x23e0 (net/ipv4/tcp.c:4564)
tcp_getsockopt+0x91/0xf0
__sys_getsockopt+0xf7/0x170
Fix this by wrapping the ca_ops load and get_info call within
rcu_read_lock()/rcu_read_unlock(), and using READ_ONCE() to load
the icsk_ca_ops pointer.
Fixes:
|
||
|
|
5271b79b7a |
tcp: fix use-after-free in do_tcp_getsockopt(TCP_CONGESTION)
do_tcp_getsockopt() reads icsk->icsk_ca_ops->name without holding rcu_read_lock(). Since commit |
||
|
|
7e1d6caa9c |
Merge branch 'net-sched-fix-remaining-actions-notification-accounting-issues'
Victor Nogueira says:
====================
net/sched: Fix remaining actions notification accounting issues
Commit
|
||
|
|
251367a0a3 |
net/sched: act_api: fix skb sizing and action leak on reoffload delete
tcf_reoffload_del_notify_msg() sizes the RTM_DELACTION skb with tcf_action_fill_size(action) alone. Unlike every other notification path it never wraps that in tcf_action_full_attrs_size(), so the nlmsg_put() header, struct tcamsg and the TCA_ACT_TAB nest that tca_get_fill() emits - 24 bytes on x86_64 - are not budgeted. As long as the single action stays well under NLMSG_GOODSIZE the floor in alloc_skb() hides this, but once its fill size crosses NLMSG_GOODSIZE the allocation is exactly 24 bytes short and tca_get_fill() runs out of tailroom. That is now easy to reach for an offloadable act_pedit with a large tcfp_nkeys, which commit |
||
|
|
e9ca46ebc3 |
net/sched: act_api: size the RTM_GETACTION reply from the actions
tca_action_gd() already walks every requested action and accumulates
attr_size += tcf_action_fill_size(act), then wraps the result in
tcf_action_full_attrs_size(). For RTM_DELACTION that value is handed to
tcf_del_notify_msg(), which allocates max(attr_size, NLMSG_GOODSIZE). For
RTM_GETACTION it is silently discarded and tcf_get_notify() allocates a
fixed NLMSG_GOODSIZE skb instead.
Any action whose dump exceeds that fixed budget therefore cannot be read
back. For example, act_pedit overruns the budget with 32 actions of four
munge keys each, act_police with 32 policers once the optional
rate/peakrate/result/avrate attributes are present
Fix this by passing attr_size through and allocate the reply the way the
add and delete paths do.
Note on exposure: RTM_GETACTION is the only one of the three action
commands that is not capability checked - tc_ctl_action() requires
CAP_NET_ADMIN for RTM_NEWACTION and RTM_DELACTION only - so this turns a
fixed NLMSG_GOODSIZE reply into a user sized allocation on an
unprivileged path. It is bounded by TCA_ACT_MAX_PRIO actions per
request, and tca_action_gd() does not reject duplicate indices, so a
single large action can be requested 32 times; an act_bpf program near
BPF_MAXINSNS is about 32KB of dump, or roughly 1MB for one request.
Creating such an action still requires CAP_NET_ADMIN, and the add and
delete paths have sized their skbs this way since the Fixes commit.
Should this ever need bounding, GFP_KERNEL_ACCOUNT would charge the
reply to the caller's memcg.
Fixes:
|
||
|
|
13eb543ceb |
net/sched: act_api: budget all shared attributes in notify skbs
tcf_action_shared_attrs_size() is supposed to return an upper bound on the
netlink attributes every action dump emits outside of TCA_ACT_OPTIONS, so
that tcf_add_notify_msg(), tcf_del_notify_msg() and friends can allocate
an skb large enough for the reply. It has fallen behind the dump path and
is now an underestimate for every single action.
Attributes, such as, TCA_ACT_IN_HW_COUNT and TCA_STATS_BASIC_HW are
emitted unconditionally and never accounted for. TCA_STATS_PKT64,
TCA_ACT_USED_HW_STATS, TCA_STATS_RATE_EST, TCA_STATS_RATE_EST64 require
specific conditions, but are also not accounted for.
Fix the issue by budgeting all of them so that we have a legitimate
upper bound. Even tough for of them require specific conditions, they
are cheap so, to avoid overcomplicating, we opted to account for them
unconditionally as well to account for a real worst case scenario.
Fixes:
|
||
|
|
28a57fb2c5 |
net: iptunnel: fix stale transport header during tunnel decapsulation
Syzbot reported a crash in qdisc_pkt_len_segs_init() caused by a stale transport_header offset after tunnel decapsulation. BUG: unable to handle page fault for address: ffffed102091a42e Oops: Oops: 0000 [#1] SMP KASAN NOPTI CPU: 0 UID: 0 PID: 340 Comm: qdisc_uaf_repro Not tainted 7.2.0-rc4-00061-g248951ddc14d #256 PREEMPT(full) Hardware name: QEMU Standard PC (Q35 + ICH9, 2009), BIOS 1.16.3-debian-1.16.3-2 04/01/2014 RIP: 0010:__asan_load2 <IRQ> qdisc_pkt_len_segs_init (net/core/dev.c:4145) __dev_queue_xmit (net/core/dev.c:4787) br_dev_queue_push_xmit (net/bridge/br_forward.c:53) br_handle_frame_finish (net/bridge/br_input.c:229) br_handle_frame (net/bridge/br_input.c:315) __netif_receive_skb_core.constprop.0 (net/core/dev.c:6099) __netif_receive_skb_list_core (net/core/dev.c:6287) netif_receive_skb_list_internal (net/core/dev.c:6445) napi_complete_done (net/core/dev.c:6813) gro_cell_poll (net/core/gro_cells.c:74) __napi_poll (net/core/dev.c:7735) net_rx_action (net/core/dev.c:7798 net/core/dev.c:7955) handle_softirqs (kernel/softirq.c:622) do_softirq (kernel/softirq.c:523 kernel/softirq.c:510 ) __local_bh_enable_ip (kernel/softirq.c:450) tun_get_user (drivers/net/tun.c:1986 (discriminator 1)) tun_chr_write_iter (drivers/net/tun.c:2032) The issue is completely latent until qdisc read transport header in commit |
||
|
|
b873bf8141 |
Merge branch 'net-mlx5e-prevent-stale-xsk-buffer-release-on-refill-retries'
Jerome Tollet says: ==================== net/mlx5e: Prevent stale XSK buffer release on refill retries Prevent duplicate XSK buffer release when a deferred RX refill fails and the same WQE is retried. Patch 1 fixes legacy cyclic RQ. It is unchanged from v3 and retains Dragos' Reviewed-by tag. Patch 2 fixes the analogous striding-RQ MPWQE path. Following Dragos' review, it now fills skip_release_bitmap in the common error path of mlx5e_xsk_alloc_rx_mpwqe(), consistently with mlx5e_alloc_rx_mpwqe(). Targeted fault injection covered both an early allocation failure and a partial 8-of-16-buffer unwind. With three consecutive failures for one MPWQE, the original 16 XSK buffers were released only once, retries saw a full bitmap, and a later successful allocation cleared it. A clean 20-second AF_XDP zero-copy pressure run exercised 1,575,262 buffer allocation failures without invalid descriptors, WQE errors, or kernel warnings. ==================== Link: https://patch.msgid.link/20260824141645.23700-1-jtollet@cisco.com Signed-off-by: Jakub Kicinski <kuba@kernel.org> |
||
|
|
63811edf51 |
net/mlx5e: Prevent stale XSK buffer release on MPWQE refill retry
With AF_XDP on a striding RQ, mlx5e defers releasing XSK buffers until
an MPWQE is refilled. If XSK allocation then returns -ENOMEM,
actual_wq_head is not advanced and a later NAPI poll retries the same
WQE.
mlx5e_free_rx_mpwqe() leaves each released slot marked as releasable. On
retry it can therefore call xsk_buff_free() again through stale pointers
after the frames have returned to the XSK pool and been reallocated.
Set all skip_release_bitmap bits in the common error path of
mlx5e_xsk_alloc_rx_mpwqe(). This matches mlx5e_alloc_rx_mpwqe(). A
successful allocation already clears the bitmap after replacing every
buffer, so retries become idempotent without changing the success path.
Fault injection forced three consecutive failures for one selected MPWQE.
Both an early allocation failure and a partial 8-of-16-buffer unwind
released the original 16 XSK buffers only once. Each error left a full
bitmap, the following NAPI retry skipped the release, and a later
successful allocation cleared it. A 20-second AF_XDP zero-copy pressure
run exercised 1,575,262 buffer allocation failures without invalid
descriptors, WQE errors, or kernel warnings.
Fixes:
|
||
|
|
e01620844c |
net/mlx5e: Prevent stale XSK buffer release on refill retry
When an XDP redirect to an AF_XDP socket fails because its RX ring is
full, the XSK core frees the buffer. During the subsequent batched refill
of a legacy cyclic RQ, mlx5e also releases the WQE's XSK buffer before
allocating a replacement. If that refill succeeds only partially, a WQE
left without a replacement retains its old buffer pointer.
The buffer can meanwhile be allocated to another WQE. A later refill
retry can then free the live buffer through the stale pointer and publish
the same UMEM frame twice.
Mark the WQE as released immediately after the driver-side free. The flag
is already cleared when a replacement buffer is assigned, so refill
retries no longer release stale pointers.
The failure is silent and produces no kernel warning or splat. A
standalone legacy cyclic-RQ zero-copy libxsk reproducer, using 64-byte UDP
traffic offered at 12 Mpps, detected it: stock stopped after 2,854,914
packets in 4.094 seconds, with 4,542 xdp_rx_ring_full events and 64
ownership/double-publication errors. With this change it processed
356,904,225 packets in 30 seconds despite 571,405 xdp_rx_ring_full events,
with no ownership or data errors.
Fixes:
|
||
|
|
a5d946466a |
net: stmmac: fix dma mapping leak in stmmac_tso_xmit()
In stmmac_tso_xmit(), if the DMA mapping of an skb fragment fails, the
frame is dropped but the DMA mappings already created for the linear
part and for the fragments mapped before the failure are never
unmapped, leaking DMA mappings.
Fix the leak by walking back over the descriptors used by the frame and
releasing each of them with stmmac_free_tx_buffer(). Moreover, release
the descriptors with stmmac_release_tx_desc() unmapping the DMA buffers.
Fixes:
|
||
|
|
18666c73af |
tcp: use GFP_ATOMIC in tcp_send_active_reset()
tcp_send_active_reset() can be called from contexts where gfp_any()
(in tcp_disconnect()) or sk->sk_allocation (in __tcp_close() and
mptcp_do_fastclose()) evaluates to GFP_KERNEL, which includes
__GFP_FS and __GFP_DIRECT_RECLAIM.
Allocating with GFP_KERNEL while holding the socket lock (sk_lock) creates
a lockdep dependency:
sk_lock -> fs_reclaim
This causes false-positive lockdep circular locking warnings with storage
subsystems (such as nvme-tcp) that acquire socket locks in block I/O paths
and invoke tcp_disconnect() or close sockets upon teardown:
set->srcu -> sk_lock -> fs_reclaim -> elevator_lock -> set->srcu
Active resets are small RST packet headers that should never
enter direct reclaim or block while holding socket locks.
Use sk_gfp_mask(sk, GFP_ATOMIC | __GFP_NOWARN) inside tcp_send_active_reset()
and remove its priority argument. This preserves __GFP_MEMALLOC access
for SOCK_MEMALLOC sockets, suppresses allocation failure warnings,
and aligns with other control packet allocations (e.g. tcp_send_fin(),
__tcp_send_ack(), tcp_xmit_probe_skb()).
Fixes:
|