Commit Graph
1472403 Commits
Author SHA1 Message Date
Shiji Yang c0ef04232f net: ethernet: mtk_wed: increase WED v2 WDMA RESV_BUFF to 0x80
Change WDMA RESV_BUFF from 0x40 to 0x80 to avoid CDM TX FIFO overflow.
Without this patch mt7986 and mt7981 may have WDMA TX hang issue. This
patch was pulled from mtk-openwrt-feeds GPL open source project.

Link: https://github.com/mediatek/mtk-openwrt-feeds/commit/07c87502e854b68b48544d101b6fe17ec059b97b
Signed-off-by: Shiji Yang <yangshiji66@outlook.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Acked-by: Lorenzo Bianconi <lorenzo@kernel.org>
Link: https://patch.msgid.link/OSZPR01MB779537889255E2F606E47EABBCA52@OSZPR01MB7795.jpnprd01.prod.outlook.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-24 11:10:54 -07:00
Giuseppe Piscitelli 7cbfb18094 net/sched: sch_cake: fix autorate reconfiguration throttling
CAKE's autorate-ingress path intends to limit shaper reconfiguration to
once per 250 ms, but last_reconfig_time is only checked and never updated.
Since the field stays zero, every qualifying capacity-estimate window can
call cake_reconfigure(), causing avoidable rate churn and scheduler work
under bursty traffic.

Store the current timestamp when autorate actually reconfigures the qdisc
so the guard enforces the intended interval.

Fixes: 7298de9cd7 ("sch_cake: Add ingress mode")
Signed-off-by: Giuseppe Piscitelli <ooonea@gmail.com>
Acked-by: Toke Høiland-Jørgensen <toke@toke.dk>
Link: https://patch.msgid.link/20260820154503.892214-1-ooonea@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:51:40 -07:00
Jian Shen 11efd7963d net: hibmcge: fix page_pool DMA direction mismatch
The driver memsets the RX buffer page head to zero before
submitting it to hardware, then calls dma_sync_single_for_device()
with DMA_TO_DEVICE.  This sync direction does not match the pool
dma_dir which is DMA_FROM_DEVICE, violating the DMA API contract
that the sync direction must match the mapping direction.

On swiotlb platforms the mismatch can cause incorrect bounce-buffer
behaviour, and CONFIG_DMA_API_DEBUG emits a warning.

Switch the page_pool dma_dir to DMA_BIDIRECTIONAL so that the
CPU-to-device memset sync becomes legal.

Fixes: c305959175 ("net: hibmcge: add support for pagepool on rx")
Signed-off-by: Jian Shen <shenjian15@huawei.com>
Signed-off-by: Jijie Shao <shaojijie@huawei.com>
Link: https://patch.msgid.link/20260820124346.4097115-1-shaojijie@huawei.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:49:13 -07:00
Eric Dumazet 137b8ae233 net_sched: sch_fq: fix pacing delay underflow with pacing offload
When pacing offload is enabled (q->offload_horizon > 0),
FQ can dequeue packets early (now < f->time_next_packet).

In this case, the drift calculation (now - f->time_next_packet)
underflows to a large unsigned value.

min(len/2, now - f->time_next_packet) then evaluates to len/2,
incorrectly halving the pacing delay for the next packet.

Fix this by only applying drift compensation if now > f->time_next_packet.

This bug was triggered when flow_max_rate was set on the qdisc
or for non EDT packets (packets with a zero skb->tstamp).

Fixes: f26080d470 ("net_sched: sch_fq: add the ability to offload pacing")
Reported-by: Willem de Bruijn <willemb@google.com>
Closes: https://lore.kernel.org/netdev/CANn89iK6O7ujR9zCJzd04MNLQoDi3mA+HWsR-hgQWYzLS3gZfw@mail.gmail.com/
Signed-off-by: Eric Dumazet <edumazet@google.com>
Signed-off-by: Willem de Bruijn <willemb@google.com>
Link: https://patch.msgid.link/20260820120706.1995449-1-willemdebruijn.kernel@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:47:53 -07:00
Rong Zhang 039f248a6c net: page_pool: Remove zone/policy GFP flags when allocating XArray entries
Net drivers request GFP flags according to both the current context and
the device constraints, but the XArray entry itself is by no mean used
by the device. Passing though device constraints to XArray allocation is
a bug and will be warned and fixed up by slab, e.g.:

    Unexpected gfp: 0x4 (GFP_DMA32). Fixing up to gfp: 0x82820 (GFP_ATOMIC|__GFP_NOWARN|__GFP_NOMEMALLOC). Fix your code!
    CPU: 2 UID: 0 PID: 1071629 Comm: kworker/u80:1 Not tainted 7.2.0-rc7+ #1 PREEMPT(lazy)
    Hardware name: LENOVO 21Q4/LNVNB161216, BIOS PXCN27WW 10/20/2025
    Workqueue: mt76 mt792x_pm_wake_work [mt792x_lib]
    Call Trace:
     <TASK>
     dump_stack_lvl+0x6e/0x90
     kmalloc_fix_flags+0x4d/0x6a
     refill_objects+0x10a/0x330
     __pcs_replace_empty_main+0x292/0x5c0
     kmem_cache_alloc_lru_noprof+0x4c2/0x680
     ? __xas_nomem+0x3a/0x120
     __xas_nomem+0x3a/0x120
     __xa_alloc+0xd4/0x190
     page_pool_dma_map+0xef/0x400
     __page_pool_alloc_netmems_slow+0xed/0x480
     ? lock_release+0x280/0x490
     page_pool_alloc_frag_netmem+0xe0/0x3a0
     page_pool_alloc_frag+0xe/0x20
     mt76_dma_rx_fill_buf+0x1f6/0x580 [mt76]
     mt76_dma_rx_reset+0x1cf/0x230 [mt76]
     mt792x_wpdma_reset+0x183/0x1b0 [mt792x_lib]
     mt792x_wpdma_reinit_cond+0x5e/0xa0 [mt792x_lib]
     mt792xe_mcu_drv_pmctrl+0x28/0x60 [mt792x_lib]
     mt792x_mcu_drv_pmctrl+0x3e/0x90 [mt792x_lib]
     mt792x_pm_wake_work+0x2d/0x1d0 [mt792x_lib]
     ? process_one_work+0x20e/0x600
     process_one_work+0x230/0x600
     ? process_one_work+0x256/0x600
     worker_thread+0x1ec/0x3c0
     ? rescuer_thread+0x610/0x610
     kthread+0xf2/0x130
     ? kthread_affine_node+0x140/0x140
     ret_from_fork+0x2a5/0x380
     ? kthread_affine_node+0x140/0x140
     ret_from_fork_asm+0x11/0x20
     </TASK>

Currently mt76 and stmmac may allocate page pool pages with GFP_DMA32.

Fix it by removing zone/policy GFP flags when allocating XArray entries.
This is inspired by commit 96d5780880 ("iommu/dma: Use the gfp
parameter in __iommu_dma_alloc_noncontiguous()").

Fixes: ee62ce7a1d ("page_pool: Track DMA-mapped pages and unmap them when destroying the pool")
Signed-off-by: Rong Zhang <i@rong.moe>
Reviewed-by: Toke Høiland-Jørgensen <toke@redhat.com>
Link: https://patch.msgid.link/20260821-page-pool-xa-drop-dma32-v1-1-6eab295c3478@rong.moe
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:28:14 -07:00
Jakub Kicinski e16d750a90 Merge branch 'net-smc-fix-out-of-bounds-and-use-after-free-in-smc-rv2-llc-processing'
Yehyeong Lee says:

====================
net/smc: fix out-of-bounds and use-after-free in SMC-Rv2 LLC processing

Patch 1 fixes a use-after-free of the LLC queue entry in
smc_llc_srv_add_link(), patch 2 bounds the peer's rkey counts, and patch 3
carries the tail of an oversized v2 message in the queue entry so that both
readers are bounded by what arrived. All three are tagged for stable: a
tree that takes 1 and 2 without 3 still deletes rkeys read from whatever an
earlier message left in the shared receive buffer.
====================

Link: https://patch.msgid.link/20260819023306.644849-1-yhlee@isslab.korea.ac.kr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:25:55 -07:00
Yehyeong Lee 8d3c1ab82c net/smc: carry oversized SMC-Rv2 LLC messages in the queue entry
smc_llc_rmt_delete_rkey() and smc_llc_save_add_link_rkeys() read the part
of a v2 message that does not fit into the 44-byte union smc_llc_msg, and
both bound themselves by the size of the buffer it landed in, not by what
arrived. On a link with a shared v2 receive buffer a 44-byte
DELETE_RKEY_V2 declaring 255 rkeys reaches rkey[9..254] in whatever an
earlier message left in lgr->wr_rx_buf_v2, and passes each of them to
smc_rtoken_delete(). One of those 255 matched a registered rtoken and
deleted it. An ADD_LINK on such a link installs up to 255 rtokens from
the same bytes.

Copy the tail into the queue entry, so its length is the length of the
message that arrived, and declare the rkeys that fit inline as a member of
the union instead of reaching them through a cast. The same
DELETE_RKEY_V2 now processes the 9 rkeys it carries. The copy is limited
to the longest tail the two functions can read, so the peer does not pick
the size of the entry.

The bound the previous patch placed on links without a shared v2 receive
buffer is no longer needed.

Fixes: 27ef6a9981 ("net/smc: support SMC-R V2 for rdma devices with max_recv_sge equals to 1")
Cc: stable@vger.kernel.org
Suggested-by: D. Wythe <alibuda@linux.alibaba.com>
Reviewed-by: Sidraya Jayagond <sidraya@linux.ibm.com>
Signed-off-by: Yehyeong Lee <yhlee@isslab.korea.ac.kr>
Link: https://patch.msgid.link/20260819023306.644849-4-yhlee@isslab.korea.ac.kr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:25:52 -07:00
Yehyeong Lee 2d1e7c5aaa net/smc: bound the peer rkey counts in SMC-Rv2 LLC messages
On a link whose device has max_recv_sge == 1 there is no shared v2 receive
buffer, and smc_llc_save_add_link_rkeys() takes the v2 extension from 44
bytes past the start of the queue entry's inline message:

  ext = (struct smc_llc_msg_add_link_v2_ext *)(llc_msg + SMC_WR_TX_SIZE);

The entry is a 72-byte allocation and the extension starts at offset 68, so
ext->num_rkeys at offset 94 is already past it. This happens on every
SMC-Rv2 link addition, whatever the peer sends:

  [    2.490065] BUG: KASAN: slab-out-of-bounds in smc_llc_save_add_link_rkeys+0x333/0x350
  [    2.490431] Read of size 2 at addr ffff8880056406de by task smctest/106
  [    2.490709]
  [    2.490792] CPU: 0 UID: 0 PID: 106 Comm: smctest Not tainted 7.2.0-rc5-p1-g77a5d9d9c99f #32 PREEMPT(lazy)
  [    2.490795] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
  [    2.490798] Call Trace:
  [    2.490803]  <TASK>
  [    2.490805]  dump_stack_lvl+0x53/0x70
  [    2.490810]  print_report+0xd0/0x630
  [    2.490828]  ? __pfx__raw_spin_lock_irqsave+0x10/0x10
  [    2.490832]  ? smc_llc_save_add_link_rkeys+0x333/0x350
  [    2.490834]  kasan_report+0xce/0x100
  [    2.490836]  ? smc_llc_save_add_link_rkeys+0x333/0x350
  [    2.490837]  smc_llc_save_add_link_rkeys+0x333/0x350
  [    2.490839]  ? smcr_buf_map_lgr+0x1bf/0x2b0
  [    2.490844]  smc_llc_cli_add_link+0xca7/0x1e80
  [    2.490848]  ? smc_llc_wait+0x355/0x810
  [    2.490850]  ? __pfx_smc_llc_wait+0x10/0x10
  [    2.490851]  ? __pfx_smc_llc_cli_add_link+0x10/0x10
  [    2.490853]  ? __pfx_autoremove_wake_function+0x10/0x10
  [    2.490863]  __smc_connect+0x3f5c/0x4980
  [    2.490873]  ? __pfx_kernel_connect+0x10/0x10
  [    2.490888]  ? __pfx___smc_connect+0x10/0x10
  [    2.490891]  ? release_sock+0x148/0x1d0
  [    2.490894]  smc_connect+0x42c/0x580
  [    2.490896]  __sys_connect+0xfc/0x130
  [    2.490898]  ? __pfx___sys_connect+0x10/0x10
  [    2.490900]  ? handle_mm_fault+0x1a1/0x430
  [    2.490908]  __x64_sys_connect+0x6d/0xb0
  [    2.490909]  ? fpregs_assert_state_consistent+0x56/0xe0
  [    2.490917]  do_syscall_64+0xf9/0x540
  [    2.490921]  entry_SYSCALL_64_after_hwframe+0x77/0x7f
  [    2.490924] RIP: 0033:0x421bb4
  [    2.490927] Code: ff f7 d8 64 89 01 48 83 c8 ff c3 66 2e 0f 1f 84 00 00 00 00 00 90 f3 0f 1e fa 80 3d ad 34 09 00 00 74 13 b8 2a 00 00 00 0f 05 <48> 3d 00 f0 ff ff 77 4c c3 0f 1f 00 55 48 89 e5 48 83 ec 10 89 55
  [    2.490929] RSP: 002b:00007ffd473b01a8 EFLAGS: 00000202 ORIG_RAX: 000000000000002a
  [    2.490935] RAX: ffffffffffffffda RBX: 0000000000000000 RCX: 0000000000421bb4
  [    2.490936] RDX: 0000000000000010 RSI: 00007ffd473b01d0 RDI: 0000000000000003
  [    2.490937] RBP: 0000000000003930 R08: 0000000000000004 R09: 0000000000000000
  [    2.490938] R10: 00007ffd473b0f98 R11: 0000000000000202 R12: 0000000000000006
  [    2.490939] R13: 00007ffd473b0f87 R14: 0000000000000003 R15: 00007ffd473b0f90
  [    2.490940]  </TASK>
  [    2.490941]
  [    2.499545] Allocated by task 44:
  [    2.499693]  kasan_save_stack+0x33/0x60
  [    2.499860]  kasan_save_track+0x14/0x30
  [    2.500026]  __kasan_kmalloc+0x8f/0xa0
  [    2.500190]  __kmalloc_cache_noprof+0x158/0x370
  [    2.500393]  smc_llc_enqueue+0x72/0x560
  [    2.500559]  smc_wr_rx_tasklet_fn+0x474/0xa80
  [    2.500747]  tasklet_action_common+0x20f/0x8a0
  [    2.500945]  handle_softirqs+0x18e/0x590
  [    2.501115]  do_softirq+0x3b/0x60
  [    2.501266]  __local_bh_enable_ip+0x61/0x70
  [    2.501446]  __alloc_skb+0x732/0x890
  [    2.501604]  rxe_init_packet+0x16b/0x4f0
  [    2.501783]  prepare_ack_packet+0xb8/0x830
  [    2.501962]  rxe_receiver+0x495/0x96e0
  [    2.502125]  do_work+0x144/0x470
  [    2.502269]  process_one_work+0x633/0x1030
  [    2.502450]  worker_thread+0x45b/0xd10
  [    2.502617]  kthread+0x2c6/0x3b0
  [    2.502762]  ret_from_fork+0x36e/0x5a0
  [    2.502925]  ret_from_fork_asm+0x1a/0x30
  [    2.503103]
  [    2.503177] The buggy address belongs to the object at ffff888005640680
  [    2.503177]  which belongs to the cache kmalloc-96 of size 96
  [    2.503692] The buggy address is located 22 bytes to the right of
  [    2.503692]  allocated 72-byte region [ffff888005640680, ffff8880056406c8)
  [    2.504227]
  [    2.504300] The buggy address belongs to the physical page:
  [    2.504535] page: refcount:0 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x5640
  [    2.504865] flags: 0x100000000000000(node=0|zone=1)
  [    2.505076] page_type: f5(slab)
  [    2.505221] raw: 0100000000000000 ffff888001041280 dead000000000122 0000000000000000
  [    2.505544] raw: 0000000000000000 0000000000200020 00000000f5000000 0000000000000000
  [    2.505867] page dumped because: kasan: bad access detected
  [    2.506102]
  [    2.506176] Memory state around the buggy address:
  [    2.506380]  ffff888005640580: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc
  [    2.506683]  ffff888005640600: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc
  [    2.506987] >ffff888005640680: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc
  [    2.507291]                                                     ^
  [    2.507548]  ffff888005640700: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc
  [    2.507850]  ffff888005640780: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc

Whatever that read finds then bounds the ext->rt[] loop, so a peer that
declares 255 rkeys reads much further. smc_llc_rmt_delete_rkey() has the
same shape for llcv2->rkey[].

Bound both loops by the buffer they read from, and skip the extension
altogether when there is no shared v2 receive buffer. The extension
does arrive on the link, but smc_llc_enqueue() copies only
sizeof(union smc_llc_msg) into the queue entry, so what that code read
past the 44 inline bytes was heap and not peer data.

Fixes: 27ef6a9981 ("net/smc: support SMC-R V2 for rdma devices with max_recv_sge equals to 1")
Cc: stable@vger.kernel.org
Reviewed-by: Sidraya Jayagond <sidraya@linux.ibm.com>
Signed-off-by: Yehyeong Lee <yhlee@isslab.korea.ac.kr>
Link: https://patch.msgid.link/20260819023306.644849-3-yhlee@isslab.korea.ac.kr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:25:51 -07:00
Yehyeong Lee a42a459ef0 net/smc: fix use-after-free of the LLC qentry in smc_llc_srv_add_link()
smc_llc_srv_add_link() keeps add_llc pointing into the queue entry:

  add_llc = &qentry->msg.add_link;			smc_llc.c:1482
  ...
  smc_llc_save_add_link_info(link_new, add_llc);	smc_llc.c:1494
  smc_llc_flow_qentry_del(&lgr->llc_flow_lcl);		smc_llc.c:1495
  ...
  u8 *llc_msg = smc_link_shared_v2_rxbuf(link) ?
	(u8 *)lgr->wr_rx_buf_v2 : (u8 *)add_llc;	smc_llc.c:1504
  smc_llc_save_add_link_rkeys(link, link_new, llc_msg);	smc_llc.c:1506

smc_llc_flow_qentry_del() kfree()s the entry, so on a link without a shared
v2 receive buffer the pointer handed to smc_llc_save_add_link_rkeys() is
already freed. Before the Fixes: commit that branch always used
lgr->wr_rx_buf_v2 and add_llc was not used after the free.

Reproduced on an unpatched tree over rxe, with KASAN, kasan_multi_shot
and a link forced to max_recv_sge == 1: the entry is freed and read by
the same call, and the freeing frame is smc_llc_srv_add_link() itself.

  [    2.523161] BUG: KASAN: slab-use-after-free in smc_llc_save_add_link_rkeys+0x333/0x350
  [    2.523499] Read of size 2 at addr ffff8880052194de by task kworker/0:1/11
  [    2.523789]
  [    2.523862] CPU: 0 UID: 0 PID: 11 Comm: kworker/0:1 Not tainted 7.2.0-rc5-p0-g2c9dd296545d #35 PREEMPT(lazy)
  [    2.523865] Hardware name: QEMU Ubuntu 24.04 PC v2 (i440FX + PIIX, arch_caps fix, 1996), BIOS 1.16.3-debian-1.16.3-2 04/01/2014
  [    2.523866] Workqueue: smc_hs_wq smc_listen_work
  [    2.523869] Call Trace:
  [    2.523870]  <TASK>
  [    2.523871]  dump_stack_lvl+0x53/0x70
  [    2.523872]  print_report+0xd0/0x630
  [    2.523874]  ? __pfx__raw_spin_lock_irqsave+0x10/0x10
  [    2.523876]  ? smc_llc_save_add_link_rkeys+0x333/0x350
  [    2.523878]  kasan_report+0xce/0x100
  [    2.523879]  ? smc_llc_save_add_link_rkeys+0x333/0x350
  [    2.523881]  smc_llc_save_add_link_rkeys+0x333/0x350
  [    2.523883]  ? smcr_buf_reg_lgr+0x2a4/0x660
  [    2.523885]  smc_llc_srv_add_link+0xaa2/0x1e50
  [    2.523888]  ? _printk+0xba/0xf0
  [    2.523897]  ? __pfx_smc_llc_srv_add_link+0x10/0x10
  [    2.523899]  ? down_write+0xb0/0x130
  [    2.523903]  ? __pfx_down_write+0x10/0x10
  [    2.523905]  smc_listen_work+0x489e/0x4d00
  [    2.523907]  ? kmem_cache_free+0x1c6/0x3a0
  [    2.523911]  ? __pfx_smc_listen_work+0x10/0x10
  [    2.523913]  ? release_sock+0x148/0x1d0
  [    2.523915]  ? smc_tcp_listen_work+0xb4f/0xfc0
  [    2.523917]  ? _raw_spin_lock_irq+0x80/0xe0
  [    2.523918]  ? __pfx__raw_spin_lock_irq+0x10/0x10
  [    2.523920]  process_one_work+0x633/0x1030
  [    2.523922]  ? assign_work+0x11d/0x370
  [    2.523924]  worker_thread+0x45b/0xd10
  [    2.523926]  ? __pfx_worker_thread+0x10/0x10
  [    2.523928]  ? __pfx_worker_thread+0x10/0x10
  [    2.523929]  kthread+0x2c6/0x3b0
  [    2.523931]  ? recalc_sigpending+0x15c/0x1e0
  [    2.523934]  ? __pfx_kthread+0x10/0x10
  [    2.523935]  ret_from_fork+0x36e/0x5a0
  [    2.523937]  ? __pfx_ret_from_fork+0x10/0x10
  [    2.523938]  ? __switch_to+0x572/0xdd0
  [    2.523943]  ? __pfx_kthread+0x10/0x10
  [    2.523944]  ret_from_fork_asm+0x1a/0x30
  [    2.523947]  </TASK>
  [    2.523948]
  [    2.531253] Allocated by task 48:
  [    2.531399]  kasan_save_stack+0x33/0x60
  [    2.531570]  kasan_save_track+0x14/0x30
  [    2.531737]  __kasan_kmalloc+0x8f/0xa0
  [    2.531905]  __kmalloc_cache_noprof+0x158/0x370
  [    2.532100]  smc_llc_enqueue+0x72/0x560
  [    2.532268]  smc_wr_rx_tasklet_fn+0x474/0xa80
  [    2.532491]  tasklet_action_common+0x20f/0x8a0
  [    2.532714]  handle_softirqs+0x18e/0x590
  [    2.532886]  do_softirq+0x3b/0x60
  [    2.533036]  __local_bh_enable_ip+0x61/0x70
  [    2.533221]  __alloc_skb+0x732/0x890
  [    2.533384]  rxe_init_packet+0x16b/0x4f0
  [    2.533567]  prepare_ack_packet+0xb8/0x830
  [    2.533760]  rxe_receiver+0x495/0x96e0
  [    2.533933]  do_work+0x144/0x470
  [    2.534078]  process_one_work+0x633/0x1030
  [    2.534257]  worker_thread+0x45b/0xd10
  [    2.534424]  kthread+0x2c6/0x3b0
  [    2.534569]  ret_from_fork+0x36e/0x5a0
  [    2.534737]  ret_from_fork_asm+0x1a/0x30
  [    2.534907]
  [    2.534980] Freed by task 11:
  [    2.535112]  kasan_save_stack+0x33/0x60
  [    2.535279]  kasan_save_track+0x14/0x30
  [    2.535444]  kasan_save_free_info+0x3b/0x60
  [    2.535625]  __kasan_slab_free+0x43/0x70
  [    2.535798]  kfree+0x121/0x380
  [    2.535935]  smc_llc_srv_add_link+0x9a8/0x1e50
  [    2.536128]  smc_listen_work+0x489e/0x4d00
  [    2.536305]  process_one_work+0x633/0x1030
  [    2.536482]  worker_thread+0x45b/0xd10
  [    2.536652]  kthread+0x2c6/0x3b0
  [    2.536794]  ret_from_fork+0x36e/0x5a0
  [    2.536958]  ret_from_fork_asm+0x1a/0x30
  [    2.537133]
  [    2.537205] The buggy address belongs to the object at ffff888005219480
  [    2.537205]  which belongs to the cache kmalloc-96 of size 96
  [    2.537719] The buggy address is located 94 bytes inside of
  [    2.537719]  freed 96-byte region [ffff888005219480, ffff8880052194e0)
  [    2.538216]
  [    2.538289] The buggy address belongs to the physical page:
  [    2.538524] page: refcount:0 mapcount:0 mapping:0000000000000000 index:0x0 pfn:0x5219
  [    2.538857] flags: 0x100000000000000(node=0|zone=1)
  [    2.539066] page_type: f5(slab)
  [    2.539210] raw: 0100000000000000 ffff888001041280 dead000000000122 0000000000000000
  [    2.539534] raw: 0000000000000000 0000000000200020 00000000f5000000 0000000000000000
  [    2.539863] page dumped because: kasan: bad access detected
  [    2.540098]
  [    2.540170] Memory state around the buggy address:
  [    2.540379]  ffff888005219380: fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc
  [    2.540684]  ffff888005219400: fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc fc
  [    2.540988] >ffff888005219480: fa fb fb fb fb fb fb fb fb fb fb fb fc fc fc fc
  [    2.541291]                                                     ^
  [    2.541548]  ffff888005219500: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc
  [    2.541857]  ffff888005219580: 00 00 00 00 00 00 00 00 00 fc fc fc fc fc fc fc

The offset is past the 72-byte queue entry because the out-of-bounds read
fixed by the next patch is on the same line; what this patch removes is the
free at smc_llc_srv_add_link+0x9a8 happening before the read at +0xaa2.

Detach the entry instead of freeing it there, and free it at the single
exit label. The reject path has to detach as well, otherwise it would be
freed twice.

This changes only the lifetime of the entry. The same read still runs past
its end until the next two patches bound it, so a backport wants all three.

Fixes: 27ef6a9981 ("net/smc: support SMC-R V2 for rdma devices with max_recv_sge equals to 1")
Cc: stable@vger.kernel.org
Reviewed-by: Sidraya Jayagond <sidraya@linux.ibm.com>
Signed-off-by: Yehyeong Lee <yhlee@isslab.korea.ac.kr>
Reviewed-by: Breno Leitao <leitao@debian.org>
Link: https://patch.msgid.link/20260819023306.644849-2-yhlee@isslab.korea.ac.kr
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:25:51 -07:00
Jiawen Wu f05516dd7b net: libwx: fix concurrent bitmap overwrite in PTP setup
In wx_ptp_set_timestamp_mode(), the driver copies the global `wx->flags`
bitmap to a local variable, modifies the PTP-related bits, and then writes
the entire bitmap back using memcpy().

This Read-Copy-Update pattern is unsafe and introduces a critical race
condition. Other asynchronous contexts (such as Tx timeout routines or
GPIO IRQ handlers) update individual bits in `wx->flags` concurrently
using atomic bitops like set_bit() or clear_bit(). The memcpy() write-back
can silently overwrite and drop these concurrent changes, potentially
causing the driver to miss critical module reset or PCIe recovery requests.

Fix this by removing the local bitmap copy. Instead, evaluate the intended
PTP flag states locally and apply them directly to `wx->flags` using
atomic set_bit() and clear_bit() operations only after the hardware is
successfully configured.

Fixes: 06e75161b9 ("net: wangxun: Add support for PTP clock")
Signed-off-by: Jiawen Wu <jiawenwu@trustnetic.com>
Reviewed-by: Vadim Fedorenko <vadim.fedorenko@linux.dev>
Link: https://patch.msgid.link/6C7EC12D69217315+20260818074721.45536-1-jiawenwu@trustnetic.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:24:18 -07:00
Vaibhav Nagare 06aa3d2632 qede: Fix NULL pointer dereference in TPA fragment processing
Under memory pressure, the qede driver encounters NULL pointer
dereferences when processing TPA continuation fragments.

Commit 8a8633978b ("qede: Add build_skb() support.") accidentally
dropped the assignment of tpa_info->buffer.data in qede_tpa_start().

When memory pressure causes an SKB allocation failure in qede_tpa_start(),
the driver sets tpa_start_fail = true and attempts to recycle the physical
page later in qede_tpa_end() via qede_reuse_page(). However, because
buffer.data was left uninitialized (NULL), qede_reuse_page() pushes a
"ghost" BD (valid DMA mapping but NULL data pointer) back into the
active Rx ring.

The next time the hardware uses this ring slot, it passes a NULL page
to qede_fill_frag_skb(), causing a kernel panic.

Example crash from production system:
 BUG: unable to handle kernel NULL pointer dereference at 0x8
 RIP: qede_fill_frag_skb+0x96/0x430 [qede]
 Call Trace:
   qede_rx_int+0xb06/0x1de0
   qede_poll+0x2f4/0x6c0
   __napi_poll+0x2d/0x130

Fix the root cause by restoring the tpa_info->buffer.data assignment
in qede_tpa_start(), ensuring valid pages are correctly tracked and
recycled. Additionally, update the stale comment for
struct qede_agg_info::buffer to reflect its current usage.

Fixes: 8a8633978b ("qede: Add build_skb() support.")
Cc: stable@vger.kernel.org
Signed-off-by: Vaibhav Nagare <vnagare@redhat.com>
Link: https://patch.msgid.link/20260818073309.2266072-1-vnagare@redhat.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:23:02 -07:00
Jiawen Wu 7bf29145d7 net: txgbe: fix MISC interrupt unmasking in non-MSI-X mode and device shutdown
In txgbe_misc_irq_thread_fn(), the driver unmasks the miscellaneous
interrupt at the end of the handler using TXGBE_INTR_MISC(wx) (which
resolves to BIT(wx->num_q_vectors)). While this is correct for MSI-X
mode, it is incorrect for legacy INTx or single MSI modes.

Due to hardware behavior, the WX_PX_MISC_IVAR register is completely
ignored by the hardware when MSI-X is disabled. In non-MSI-X mode, the
hardware forcibly merges all interrupt causes (both Queue and MISC) into
a single bit: BIT(0) of the interrupt register.

Unconditionally unmasking TXGBE_INTR_MISC(wx) (e.g., BIT(1)) in non-MSI-X
mode means the actual MISC interrupt bit (BIT(0)) is not unmasked
promptly at the end of the MISC thread. Instead, it remains masked until
NAPI completes its polling and unmasks the shared BIT(0). This delays the
assertion of subsequent MISC interrupts, preventing timely handling of
events like link state changes.

Fix this by explicitly checking `pdev->msix_enabled` and falling back
to BIT(0) as the interrupt mask for the MISC cause when MSI-X is disabled.

Additionally, unconditionally unmasking the interrupt at the end of the
thread introduces a race condition during device teardown. Guarding the
wx_intr_enable() call with a check for the WX_STATE_DOWN bit, to prevent
re-arming the interrupt during device shutdown.

Fixes: e37546ad1f ("net: wangxun: revert the adjustment of the IRQ vector sequence")
Signed-off-by: Jiawen Wu <jiawenwu@trustnetic.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/56A53978B83EEDE9+20260818023026.6631-1-jiawenwu@trustnetic.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:22:15 -07:00
Jakub Kicinski dac0c3fa97 Merge branch 'net-sparx5-misc-fixes-for-sparx5-and-lan969x'
Daniel Machon says:

====================
net: sparx5: misc fixes for sparx5 and lan969x

This series fixes various issues in the sparx5 driver, which also
serves lan969x.

Details are in the individual commit descriptions.
====================

Link: https://patch.msgid.link/20260817-misc-fixes-sparx5-lan969x-v3-0-c7c7fef723a8@microchip.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:21:29 -07:00
Daniel Machon b62793a7ba net: sparx5: fix sleep in atomic context in MAC table access
sparx5_set_rx_mode() runs with netif_addr_lock_bh held and iterates
dev->mc via __dev_mc_sync(), which per address calls sparx5_mc_sync() /
sparx5_mc_unsync() -> sparx5_mact_learn() / sparx5_mact_forget().  These
take sparx5->lock, a mutex, and then poll the MAC access command
register with readx_poll_timeout(). A mutex may block, which is not
allowed from atomic context.

Convert the driver to the new .ndo_set_rx_mode_async callback introduced
in commit 3554b4345d ("net: introduce ndo_set_rx_mode_async and
netdev_rx_mode_work"). The async callback is invoked from process
context, so the mutex and sleeping completion poll can remain.

Observed with CONFIG_PROVE_LOCKING, CONFIG_DEBUG_SPINLOCK,
CONFIG_DEBUG_MUTEXES and CONFIG_DEBUG_ATOMIC_SLEEP enabled:

  BUG: sleeping function called from invalid context at kernel/locking/mutex.c:591
  in_atomic(): 1, irqs_disabled(): 0, non_block: 0, pid: 217, name: ip
  preempt_count: 201, expected: 0
  Call trace:
   __might_resched+0x144/0x248
   __might_sleep+0x48/0x7c
   __mutex_lock+0x74/0x850
   mutex_lock_nested+0x24/0x30
   sparx5_mact_learn+0x78/0x100
   sparx5_mc_sync+0x40/0x54
   __hw_addr_sync_dev+0xc4/0x170
   sparx5_set_rx_mode+0x4c/0x58
   __dev_set_rx_mode+0x64/0xa4
   __dev_open+0x1ec/0x26c

Fixes: d6fce51419 ("net: sparx5: add switching support")
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20260817-misc-fixes-sparx5-lan969x-v3-2-c7c7fef723a8@microchip.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:21:26 -07:00
Daniel Machon b7adcc56fd net: microchip: vcap: use port number instead of netdev name for debugfs
sparx5_vcap_init() runs before sparx5_register_netdevs() in probe, and
its debugfs setup calls vcap_port_debugfs() for every port using
netdev_name(ndev) as the debugfs file name. At that point the netdevs
have only been allocated, not registered, so dev->name still holds the
"eth%d" template and netdev_name() returns "(unnamed net_device)".
Every port tries to create the same file under vcaps/, producing a
flood of warnings at boot:

  debugfs: '(unnamed net_device)' already exists in 'vcaps'
  debugfs: '(unnamed net_device)' already exists in 'vcaps'
  ...

Add vcap_port_debugfs_portno(), a variant of vcap_port_debugfs() that
takes the port's stable hardware port number and uses "p%u" as the
debugfs file name instead of netdev_name(ndev). This makes the file
name independent of registration order; the file still stores and
later dereferences the netdev itself, same as before. sparx5 already
reports the same "p%d" string via ndo_get_phys_port_name(), so the
debugfs name now matches that.

Only sparx5 (and lan969x, which shares this code) is switched to the
new function. lan966x keeps calling vcap_port_debugfs() unchanged, so
this fix does not rename any of its existing debugfs files.

Fixes: b8909aad5b ("net: sparx5: move netdev and notifier block registration to probe")
Signed-off-by: Daniel Machon <daniel.machon@microchip.com>
Link: https://patch.msgid.link/20260817-misc-fixes-sparx5-lan969x-v3-1-c7c7fef723a8@microchip.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:21:26 -07:00
Qing Ming 498386b6d4 gtp: serialize PDP context updates
PDP contexts can be deleted through GTP_CMD_DELPDP or while the GTP
network device is being unregistered. The latter is serialized by RTNL,
but the generic-netlink delete path only holds RCU.

Running both paths concurrently can therefore make both paths delete the
same PDP context. The issue was found through static analysis and
reproduced on a KASAN-enabled kernel by a simple two-thread program
racing GTP_CMD_DELPDP against RTM_DELLINK:

  Oops: general protection fault, probably for non-canonical address
  KASAN: maybe wild-memory-access in range
         [0xdead000000000120-0xdead000000000127]
  RIP: gtp_genl_del_pdp+0x1c1/0x420 [gtp]
  RBP: dead000000000122

The second deletion dereferenced the poisoned hlist pprev pointer.

Serialize gtp_pdp_add(), gtp_genl_del_pdp(), and gtp_dellink() with a
shared mutex. Keep the mutex held until the final use of a PDP context in
the NEWPDP path, and keep the RCU read-side section around the complete
PDP context use in the DELPDP path.

Fixes: 459aa660eb ("gtp: add initial driver for datapath of GPRS Tunneling Protocol (GTP-U)")
Cc: stable@vger.kernel.org
Signed-off-by: Qing Ming <a0yami@mailbox.org>
Link: https://patch.msgid.link/20260818150000.7670-1-a0yami@mailbox.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:17:12 -07:00
Weiming Shi 71283aaa6c xdp: fix zero-copy frame layout
xdp_convert_zc_to_xdp_frame() clones an XSK packet into an order-0 page
and advertises PAGE_SIZE as its frame size.  It allows the copied frame
to occupy the page tail needed by skb_shared_info and records zero
headroom even when metadata separates the frame header from packet data.
An AF_XDP zero-copy packet redirected through cpumap can therefore make
the skb overlap skb_shared_info or place it beyond the allocated page.

Limit the copied layout to SKB_WITH_OVERHEAD(PAGE_SIZE) and include the
metadata length in frame headroom.  Redirect callers already handle a
NULL conversion result.

BUG: KASAN: slab-out-of-bounds in skb_gro_receive
Write of size 4 at addr ffff88800cf37004 by task cpumap/1/map:1/146
Call Trace:
 skb_gro_receive (net/core/gro.c:174)
 udp_gro_receive (net/ipv4/udp_offload.c:812)
 inet_gro_receive (net/ipv4/af_inet.c:1539)
 dev_gro_receive (net/core/gro.c:515)
 gro_receive_skb (net/core/gro.c:633)
 cpu_map_kthread_run (kernel/bpf/cpumap.c:395)
 kthread (kernel/kthread.c:436)
 ret_from_fork (arch/x86/kernel/process.c:164)
 ret_from_fork_asm (arch/x86/entry/entry_64.S:255)
Kernel panic - not syncing: KASAN: panic_on_warn set ...

Fixes: b0d1beeff2 ("xdp: implement convert_to_xdp_frame for MEM_TYPE_ZERO_COPY")
Cc: stable@vger.kernel.org
Reported-by: Xiang Mei <xmei5@asu.edu>
Signed-off-by: Weiming Shi <bestswngs@gmail.com>
Link: https://patch.msgid.link/20260818154516.793517-1-bestswngs@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:10:48 -07:00
Jakub Kicinski 7c3a1a3481 Merge branch 'net-ntb_netdev-fix-rx-statistics-accounting'
Koichiro Den says:

====================
net: ntb_netdev: Fix RX statistics accounting

This series addresses Jakub's comment on ntb_netdev RX statistics
accounting:
https://lore.kernel.org/r/20260818092938.4121c220@kernel.org/

It fixes the double counting and also the related packet/byte accounting
on RX refill failure.
====================

Link: https://patch.msgid.link/20260819172539.1450821-1-den@valinux.co.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:06:10 -07:00
Koichiro Den 31ded341c3 net: ntb_netdev: Count packets dropped on RX refill failure
When replacement skb allocation fails, ntb_netdev drops a packet that
was received successfully and requeues the original buffer. The drop is
counted, but rx_packets and rx_bytes are not.

Count every good packet before allocating its replacement.

Fixes: d2121faf13 ("NTB: ntb_netdev: Preserve RX queue depth on allocation failure")
Cc: stable@vger.kernel.org
Signed-off-by: Koichiro Den <den@valinux.co.jp>
Link: https://patch.msgid.link/20260819172539.1450821-3-den@valinux.co.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:06:08 -07:00
Koichiro Den 82e15be2d8 net: ntb_netdev: Avoid double-accounting netif_rx() drops
netif_rx() already accounts packets it drops in the core rx_dropped
counter. ntb_netdev counts them again as both errors and drops.

Leave netif_rx() drops to the core. Count the packet and bytes
unconditionally since it was received successfully by the driver.

Fixes: 548c237c0a ("net: Add support for NTB virtual ethernet device")
Cc: stable@vger.kernel.org
Suggested-by: Jakub Kicinski <kuba@kernel.org>
Signed-off-by: Koichiro Den <den@valinux.co.jp>
Link: https://patch.msgid.link/20260819172539.1450821-2-den@valinux.co.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:06:08 -07:00
Yong Wang 870a9e42ec tcp: clamp route advmss to TCP_MIN_MSS
tcp_select_initial_window() assumes that callers never pass an MSS
smaller than 1, but route-derived advmss values can violate that
assumption.

A too-small explicit RTAX_ADVMSS is one way to get there, but it is not
the only one. The same divide-by-zero can also be reached through the
"default advmss" path when RTAX_ADVMSS is left at 0 and the effective
advmss is later driven down by route MTU and min_adv_mss.

Introduce a tcp_dst_advmss() helper that clamps route advmss to
TCP_MIN_MSS before TCP consumes it, and use it in the TCP paths that
derive advmss from dst metrics. This keeps the effective MSS from
dropping to zero before tcp_select_initial_window() rounds the receive
window.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Reported-by: Vega <vega@nebusec.ai>
Signed-off-by: Yong Wang <edragain@163.com>
Signed-off-by: Ren Wei <weir@nebusec.ai>
Link: https://patch.msgid.link/251eaf8277fa7c66364c9815c5da01662d269181.1787074852.git.edragain@163.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:05:19 -07:00
Anton Danilov 6776efe4a5 ipip: fix skb leak in collect_md mode when metadata_dst allocation fails
In collect_md mode ipip_tunnel_rcv() returns 0 without freeing the skb
when ip_tun_rx_dst() fails to allocate the metadata_dst. ipip_rcv() and
mplsip_rcv() are registered as xfrm_tunnel handlers, so tunnel4_rcv()
and tunnelmpls4_rcv() read the zero return as "the packet has been
consumed" and do not free it either. The skb is leaked.

The other tunnel drivers all dispose of the packet at this point:
ip6_tunnel.c jumps to its drop label, ip_gre.c and ip6_gre.c return
PACKET_REJECT, which makes gre_rcv() free the skb. Only ipip returns 0.

Jump to the existing drop label instead. It frees the skb and still
returns 0, so the packet keeps being reported as consumed, which is what
we want here: the outer header has already been pulled, and neither the
remaining handlers nor an ICMP unreachable have any use for it.

Triggering this needs an ipip or mplsip tunnel in collect_md mode and an
atomic allocation failure, which is why it has gone unnoticed.

Fixes: cfc7381b30 ("ip_tunnel: add collect_md mode to IPIP tunnel")
Cc: stable@vger.kernel.org
Signed-off-by: Anton Danilov <littlesmilingcloud@gmail.com>
Reviewed-by: Fernando Fernandez Mancera <fmancera@suse.de>
Link: https://patch.msgid.link/20260819104338.432631-2-littlesmilingcloud@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:03:10 -07:00
Jamal Hadi Salim 1beb81947e net/sched: account classifier filter allocations to memcg
Allocations in the tc classifier *_change() paths (filter objects,
per-CPU counters, and per-filter aux data) use plain GFP_KERNEL without
__GFP_ACCOUNT, allowing unprivileged users to pin kernel memory outside
memcg charging. The shared tcf_exts_init_ex() action array allocation in
cls_api.c was also uncharged; this patch closes it along with the
per-classifier filter-object/percpu/aux allocations that remain
unaccounted.

Add GFP_KERNEL_ACCOUNT to:
- the shared tcf_exts_init_ex() action array (cls_api.c), common to every
  filter of every classifier (32 pointers, 256 bytes);
- the filter-object, per-CPU-counter, and per-filter aux allocations in
  cls_basic, cls_bpf, cls_cgroup, cls_flow, cls_flower, cls_fw,
  cls_matchall, cls_route and cls_u32;
- the u32_init_knode() replace-path knode allocation (cls_u32.c), which
  allocates the same struct tc_u_knode + sel.keys on every replace of an
  existing knode and was missed by the create-path-only conversion.

Also fix the cls_basic error path: basic_change() inserts fnew into the
IDR before allocating the per-CPU counter. If alloc_percpu() fails the
errout path kfree'd fnew without idr_remove, leaving a dangling pointer
in the IDR. With GFP_KERNEL_ACCOUNT the percpu alloc becomes failable
on demand (memcg at memory.max), making the dead path attacker-reachable
and burning the handle permanently. Add the idr_remove on the percpu
failure path, matching the basic_set_parms failure-path pattern.

Note: vega@nebusec.ai provided a poc for basic_cls, but it was easy to
extend to the other classifiers.

Conditions to recreate the bug:
- CONFIG_NET_SCHED, CONFIG_NET_CLS_* (the classifier being used),
  CONFIG_NET_CLS_ACT, CONFIG_MEMCG, CONFIG_USER_NS, CONFIG_NET_NS.
- Unprivileged user in a fresh user+network namespace (unshare -Urn),
  or root with CAP_NET_ADMIN.
- Create a large number of tc filters (e.g. tc filter add dev lo
  ingress ... <classifier> ...) while watching a memcg-limited cgroup:
  system slab grows far faster than memory.current, pinning kernel
  memory outside memcg charging.

Fixes: 0da974f4f3 ("[NET]: Conversions from kmalloc+memset to k(z|c)alloc.")
Reported-by: vega@nebusec.ai
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Breno Leitao <leitao@debian.org>
Link: https://patch.msgid.link/20260819143733.57538-1-jhs@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 13:02:32 -07:00
Jakub Kicinski e755276c9f Merge tag 'batadv-net-pullrequest-20260821' of https://git.open-mesh.org/batadv
Simon Wunderlich says:

====================
Here are a few batman-adv bugfixes:

 - fix stale receive device on merged fragments, by Zhiling Zou

the others are written by Sven Eckelmann:

 - bla: fix potential CRC corruption issues (2 patches)

 - dat: avoid unaligned fault in IP extraction

 - dat: atomically update mac addresses

 - mcast: fix TX priority extraction for BATADV_FORW_MCAST

 - mcast: fix skb sharing and linearization (2 patches)

 - bla: fix freeing of claims on meshif deletion

* tag 'batadv-net-pullrequest-20260821' of https://git.open-mesh.org/batadv:
  batman-adv: bla: fix freeing of claims on meshif deletion
  batman-adv: mcast: linearize skbuff for packet generation
  batman-adv: mcast: ensure unshared skb for multicast packets
  batman-adv: fix TX priority extraction for BATADV_FORW_MCAST
  batman-adv: dat: atomically update mac addresses
  batman-adv: dat: avoid unaligned fault in IP extraction
  batman-adv: bla: prevent CRC corruptions after claim flush
  batman-adv: bla: avoid CRC corruption due to parallel claim add
  batman-adv: fix stale receive device on merged fragments
====================

Link: https://patch.msgid.link/20260821094813.201800-1-sw@simonwunderlich.de
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 12:42:58 -07:00
Wei Fang 777dbc9914 ptp: netc: fix period truncation and potential divide-by-zero in PEROUT
The max_period bound in net_timer_enable_perout() was computed as:

  max_period = (u64)NETC_TMR_DEFAULT_FIPER + integral_period;

which exceeds U32_MAX when integral_period > 0 (e.g. 0x100000002 for
the default 333333333 Hz clock). A period_ns that passes this check but
exceeds U32_MAX is then silently truncated when stored into the u32
struct netc_pp::period field.

A truncated value of zero can reach netc_timer_set_perout_alarm(), where
the local u32 period variable would also be 0, causing a divide-by-zero
in roundup_u64(delta, period) whenever the stime < min_time branch is
taken (which always happens for a start time of {0, 0}).

Additionally, netc_timer_enable_periodic_pulse() and
netc_timer_enable_fiper() both compute:

  fiper = pp->period - integral_period;

A zero pp->period results in an unsigned wraparound to 0xFFFFFFFD,
mis-programming the FIPER hardware register.

Fix all three issues by capping max_period at NETC_TMR_DEFAULT_FIPER
(0xFFFFFFFF). This ensures that any period_ns passing the range check
fits in a u32 without truncation, so the stored value is always valid
and non-zero. The accepted range is reduced by integral_period ns
(typically only a few nanoseconds), which is negligible in practice.

Fixes: 671e266835 ("ptp: netc: add periodic pulse output support")
Signed-off-by: Wei Fang <wei.fang@nxp.com>
Reviewed-by: Abel Vesa <abel.vesa@oss.qualcomm.com>
Link: https://patch.msgid.link/20260821032449.1235065-1-wei.fang@oss.nxp.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 12:34:39 -07:00
Thomas Walsh a70859cf31 bnxt_en: Gate TPH enablement behind BNXT_SUPPORTS_QUEUE_API check
In bnxt_request_irq(), pcie_enable_tph() is called unconditionally to
enable PCIe TPH when setting up interrupts.

If the NIC hardware or firmware capabilities do not support queue ops,
attempting to enable TPH during bnxt_request_irq() is unnecessary.

As a result a flood of "RX queue restart failed: err=-95"  messages is
seen upon boot.

Older NICs (pre-Thor / BCM57414) do not support TPH or queue management.
TPH requires queue management to restart the queue.  NICs that support
queue management (with updated FW) all support TPH.

Gate the call to pcie_enable_tph() and setting of bp->tph_mode
behind BNXT_SUPPORTS_QUEUE_API(bp) to ensure TPH is only initialized
on devices capable of supporting queue ops. This prevents a guaranteed
-EOPNOTSUPP error from occurring due to NULL operations.

Fixes: c214410c47 ("bnxt_en: Add TPH support in BNXT driver")
Suggested-by: Michal Schmidt <mschmidt@redhat.com>
Signed-off-by: Thomas Walsh <thwalsh@redhat.com>
Reviewed-by: Michael Chan <michael.chan@broadcom.com>
Reviewed-by: Pavan Chebbi <pavan.chebbi@broadcom.com>
Link: https://patch.msgid.link/20260820220544.1240879-1-thwalsh@redhat.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 12:33:46 -07:00
Guenter Roeck 622d698df4 bnxt_en: Fix call to hardware monitoring event handler
The first parameter of hwmon_notify_event() is supposed to be the hardware
monitoring device. The bnxt driver calls it with the platform device as
first parameter instead. This API break results in undefined behavior and
may result in a crash.

Pass the hardware monitoring device as parameter instead to fix the
problem.

Fixes: a19b480145 ("bnxt_en: Event handler for Thermal event")
Signed-off-by: Guenter Roeck <linux@roeck-us.net>
Reviewed-by: Kalesh AP <kalesh-anakkur.purayil@broadcom.com>
Reviewed-by: Vadim Fedorenko <vadim.fedorenko@linux.dev>
Link: https://patch.msgid.link/20260821044512.663941-1-linux@roeck-us.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-22 12:32:59 -07:00
Joe Damato 4e15e89faa net: bnxt: ring the doorbell when SW USO exits early
When a burst of packets is handed down to the driver, the driver defers
the doorbell to the end by setting txr->kick_pending = 1. The normal TX
path handles this, but the SW USO path can miss it if it returns
early.

If bnxt_sw_udp_gso_xmit runs but returns early with NETDEV_TX_BUSY and
txr->kick_pending was previously set to 1, then the TX queue can
stall because the driver wrote some BDs but never wrote the doorbell.
The device won't know to do the TX which would generate the completion
that would wake the queue back up.

Simplify bnxt_sw_udp_gso_xmit to set txr->kick_pending in its success
case and check the flag on return. The added check after
bnxt_sw_udp_gso_xmit returns ensures that any pending doorbells are
written handling both successful USO and any early returns, which
prevents the TX queue stall mentioned above.

This TX queue stall was observed on a production system with a netdev TX
watchdog informing about the queue stall.

Fixes: cc5d90667d ("net: bnxt: Implement software USO")
Cc: stable@vger.kernel.org
Signed-off-by: Joe Damato <joe@dama.to>
Reviewed-by: Michael Chan <michael.chan@broadcom.com>
Link: https://patch.msgid.link/20260819233213.3673149-1-joe@dama.to
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-21 11:10:57 -07:00
Mehrdad Afshari 746fc0787f net: usb: cdc_ncm: add Apple MacBook Pro USB product ID 0x1902
The cdc_devs[] quirk table special-cases the Mac CDC-NCM private
interface personality only for USB product ID 0x1905.
Some MacBook Pro models (e.g. M1 Max) connected over a
USB4/Thunderbolt 3/4 cable to a host whose Thunderbolt controller
lacks PCIe tunneling support (no NHI function, USB4-only mode)
present themselves with product ID 0x1902 instead,
using the same descriptor layout as 0x1905: a Communications
control interface with zero endpoints (no interrupt/status endpoint)
paired with a CDC Data interface, at
interface numbers 0 and 2.

Because 0x1902 is unmatched, these devices fall through to the
generic cdc_ncm_info driver_info, which sets FLAG_LINK_INTR and
therefore requires an interrupt endpoint on the control interface.
Apple's private NCM interface never provides one, so cdc_ncm_bind()
fails outright:

  cdc_ncm 2-1:1.0: bind() failure
  cdc_ncm 2-1:1.2: bind() failure

and no network device is created, breaking Ethernet-over-USB4
between the Mac and any USB4 host lacking Thunderbolt PCIe
tunneling.

Add matching entries for 0x1902 alongside the existing 0x1905
ones, reusing apple_private_interface_info as with the other Mac
ID.

Signed-off-by: Mehrdad Afshari <mehrdad@signeen.com>
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 14:47:29 -07:00
Jakub Kicinski e264c04870 Merge branch 'bridge-vxlan-fix-reading-neigh-ha-without-synchronization'
Nikolay Aleksandrov says:

====================
bridge/vxlan: fix reading neigh ha without synchronization

Neigh ha address must be read using the seqlock to get a stable snapshot.
Both the bridge and vxlan read it directly and can see partial updates.
I reproduced both issues with running neigh updates and exercising these
paths in parallel and saw partial addresses, e.g. updating between
neigh A: 02:00:00:00:00:00 neigh B: fe:ff:ff:ff:ff:ff was able to observe
02:00:ff:ff:ff:ff and fe:ff:00:00:00:00 in packets. Noticed this initially
in the bridge, then checked vxlan and its arp/neigh_reduce functions have
the same bug, route_shortcircuit is doing the right thing already.
====================

Link: https://patch.msgid.link/20260818150756.890025-1-razor@blackwall.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 14:23:08 -07:00
Nikolay Aleksandrov b824059a67 vxlan: fix reading neigh ha
Currently arp/neigh_reduce read neigh ha directly which can lead to
partial reads while the neigh is being updated. Use neigh_ha_snapshot to
take a stable snapshot of the address similar to route_shortcircuit which
already does the right thing.

Fixes: e4f67addf1 ("add DOVE extensions for VXLAN")
Fixes: f564f45c45 ("vxlan: add ipv6 proxy support")
Signed-off-by: Nikolay Aleksandrov <razor@blackwall.org>
Reviewed-by: Petr Machata <petrm@nvidia.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260818150756.890025-3-razor@blackwall.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 14:23:06 -07:00
Nikolay Aleksandrov 57549ab907 net: bridge: arp/nd proxy: fix reading neigh ha
Currently neigh ha address is read directly, but that can result in
torn/partial reads if the neigh is being updated. Use neigh_ha_snapshot
to take a stable snapshot of the address.

Fixes: 057658cb33 ("bridge: suppress arp pkts on BR_NEIGH_SUPPRESS ports")
Fixes: ed842faeb2 ("bridge: suppress nd pkts on BR_NEIGH_SUPPRESS ports")
Signed-off-by: Nikolay Aleksandrov <razor@blackwall.org>
Reviewed-by: Petr Machata <petrm@nvidia.com>
Reviewed-by: Ido Schimmel <idosch@nvidia.com>
Link: https://patch.msgid.link/20260818150756.890025-2-razor@blackwall.org
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 14:23:06 -07:00
Mahanta Jambigi 5ee0ceddc7 net/smc: free pending qentry in smc_llc_flow_stop() before memset
smc_llc_flow_stop() resets a flow struct with a blind memset:

	spin_lock_bh(&lgr->llc_flow_lock);
	memset(flow, 0, sizeof(*flow));
	flow->type = SMC_LLC_FLOW_NONE;
	spin_unlock_bh(&lgr->llc_flow_lock);

If flow->qentry is non-NULL at this point the pointer is overwritten without the
allocation being freed, leaking one kmalloc object.

A late-arriving duplicate CONFIRM_LINK or ADD_LINK_CONT message can set
flow->qentry after the legitimate message has been consumed by the waiter via
smc_llc_flow_qentry_clr() (which NULLs the pointer but leaves flow->type
non-zero) but before the flow completes and smc_llc_flow_stop() runs.  In that
window the duplicate is stashed into flow->qentry, and then lost when
smc_llc_flow_stop() zeros the struct.

Call smc_llc_flow_qentry_del() inside the lock before the memset.
smc_llc_flow_qentry_del() already checks flow->qentry before freeing, so the
normal case where no entry is pending is a no-op.

Fixes: 555da9af82 ("net/smc: add event-based llc_flow framework")
Reviewed-by: Hidayath Khan <hidayath@linux.ibm.com>
Signed-off-by: Mahanta Jambigi <mjambigi@linux.ibm.com>
Link: https://patch.msgid.link/20260818073943.1108383-1-mjambigi@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 14:12:16 -07:00
Mahanta Jambigi 036322025d net/smc: free stashed qentry before overwrite in REQ_ADD_LINK to ADD_LINK transition
When smc_llc_event_handler() transitions the local LLC flow from
SMC_LLC_FLOW_REQ_ADD_LINK to SMC_LLC_FLOW_ADD_LINK on arrival of an ADD_LINK
request, it calls smc_llc_flow_qentry_set() unconditionally:

	if (lgr->llc_flow_lcl.type == SMC_LLC_FLOW_REQ_ADD_LINK) {
		lgr->llc_flow_lcl.type = SMC_LLC_FLOW_ADD_LINK;
		smc_llc_flow_qentry_set(&lgr->llc_flow_lcl, qentry);
		...
	}

A CONFIRM_LINK or ADD_LINK_CONT arriving while flow->type is
SMC_LLC_FLOW_REQ_ADD_LINK is stashed into flow->qentry via the
SMC_LLC_CONFIRM_LINK / SMC_LLC_ADD_LINK_CONT handler (which stores into
flow->qentry for any non-NONE flow type).  When the subsequent ADD_LINK
arrives, the REQ_ADD_LINK branch overwrites flow->qentry with the new pointer
without first freeing the stashed allocation, leaking one kmalloc object.

The stashed entry has no consumer: smc_llc_wait() is only called from
llc_add_link_work, which is not yet scheduled while the flow type remains
REQ_ADD_LINK.  No waiter is sleeping on llc_msg_waiter at this point.
It is safe to unconditionally free any stashed qentry before
the overwrite.

Call smc_llc_flow_qentry_del() before smc_llc_flow_qentry_set() in the
REQ_ADD_LINK branch.  smc_llc_flow_qentry_del() already checks flow->qentry
before freeing, so the normal path where no entry is stashed is a no-op.

Fixes: b4ba4652b3 ("net/smc: extend LLC layer for SMC-Rv2")
Reviewed-by: Hidayath Khan <hidayath@linux.ibm.com>
Signed-off-by: Mahanta Jambigi <mjambigi@linux.ibm.com>
Link: https://patch.msgid.link/20260818073107.466506-1-mjambigi@linux.ibm.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 14:12:11 -07:00
Christian Marangi f2849b1fd0 net: phylink: correctly validate returned PCS in phylink_inband_caps
In phylink_inband_caps(), the PCS returned by mac_select_pcs is only
checked if NULL but mac_select_pcs can also return an error pointer.

This can cause a kernel panic as phylink_pcs_inband_caps() only checks
if passed PCS is not NULL and directly dereference ops from the phylink_pcs
struct.

Use the IS_ERR_OR_NULL macro to address both case where the returned
PCS can be NULL or an error pointer and prevent a kernel panic.

Cc: stable@vger.kernel.org
Fixes: df874f9e52 ("net: phylink: add pcs_inband_caps() method")
Signed-off-by: Christian Marangi <ansuelsmth@gmail.com>
Link: https://patch.msgid.link/20260817213009.13924-1-ansuelsmth@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 14:07:43 -07:00
Jakub Kicinski 52ffb39e71 Merge branch 'net-ntb_netdev-fix-tx-completion-and-error-handling'
Koichiro Den says:

====================
net: ntb_netdev: Fix TX completion and error handling

This small series fixes several TX buffer ownership and queue handling
bugs in ntb_netdev and ntb_transport.

Patch 4 first appeared in my "NTB: Add direct TX/RX using PCI endpoint
DMA" series. Sashiko later reported the same pre-existing leak while
reviewing another series, so I moved the fix here. See:
https://lore.kernel.org/r/20260815032932.151F11F000E9@smtp.kernel.org/

The ntb_transport fixes affect buffer ownership and queue handling in
ntb_netdev, the only in-tree ntb_transport_client implementation, so the
patches need to go in together. I am targeting the net tree for the
series.
====================

Link: https://patch.msgid.link/20260817053519.4135287-1-den@valinux.co.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:43:27 -07:00
Koichiro Den a4f2387db6 NTB: ntb_transport: Reject oversized TX buffers
ntb_process_tx() handles an oversized buffer by calling tx_handler()
with a NULL data pointer and returning success. ntb_netdev therefore
neither frees the skb in its completion callback nor takes its enqueue
error path, leaking it.

Reject oversized buffers in ntb_transport_tx_enqueue() before acquiring
a queue entry and return -EMSGSIZE. The caller retains ownership of the
buffer, and the preceding netdev patch frees the skb when enqueue
returns this permanent error.

Fixes: fce8a7bb5b ("PCI-Express Non-Transparent Bridge Support")
Cc: stable@vger.kernel.org
Signed-off-by: Koichiro Den <den@valinux.co.jp>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Link: https://patch.msgid.link/20260817053519.4135287-5-den@valinux.co.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:43:24 -07:00
Koichiro Den 873ce713fe NTB: ntb_transport: Fail TX enqueue when the QP link is down
Commit f195a1a6fe ("ntb: Drop packets when qp link is down") meant to
make ntb_transport_tx_enqueue() drop packets submitted while the QP link
is down, but it only returns 0 without consuming the packet. Zero means
success by this function's contract, so ntb_netdev reports NETDEV_TX_OK
and forgets the skb: nothing queued it, nothing frees it, and it leaks,
one skb for every transmit racing a link-down.

Return -ENOLINK instead, restoring the contract that a non-zero return
leaves the buffer owned by the caller. With the preceding patch,
ntb_netdev frees the skb on non-retryable enqueue failures and returns
NETDEV_TX_OK, so a packet racing with link-down is dropped without leaking
or entering a busy retry loop.

Fixes: f195a1a6fe ("ntb: Drop packets when qp link is down")
Cc: stable@vger.kernel.org
Signed-off-by: Koichiro Den <den@valinux.co.jp>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Link: https://patch.msgid.link/20260817053519.4135287-4-den@valinux.co.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:43:24 -07:00
Koichiro Den 8aaa47351d net: ntb_netdev: Fix TX busy and drop handling
Currently, ntb_netdev returns NETDEV_TX_BUSY for every enqueue error. It
also increments the drop and error counters while leaving the skb owned
by the qdisc, and may return BUSY with the subqueue still awake.
Retrying a permanent error cannot succeed either.

The unconditional BUSY return and premature accounting date back to the
initial driver. The error-path queue stop was later removed without
changing that return value. The current flow-control code includes a
resource check, but ntb_netdev does not honor its result before enqueue.

Honor the resource check before enqueue. For -EAGAIN and -EBUSY, stop
the subqueue, arm the existing reaper timer, and return BUSY without
touching the skb. For other errors, free the skb, increment tx_dropped,
and return NETDEV_TX_OK.

Fixes: 548c237c0a ("net: Add support for NTB virtual ethernet device")
Fixes: d723485cb4 ("ntb_netdev: remove tx timeout")
Fixes: e74bfeedad ("NTB: Add flow control to the ntb_netdev")
Cc: stable@vger.kernel.org
Signed-off-by: Koichiro Den <den@valinux.co.jp>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Link: https://patch.msgid.link/20260817053519.4135287-3-den@valinux.co.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:43:24 -07:00
Koichiro Den 2564963972 NTB: ntb_transport: Recycle TX entries before client callbacks
ntb_tx_copy_callback() invokes the client callback before returning the
entry to tx_free_q. The callback may wake a stopped client queue, only
for the next enqueue to find no local entry and return -EBUSY. The window
is narrow, but the retry is unnecessary.

Save the callback data and length, then return the entry to tx_free_q
before invoking the client. A completion callback then means both the
client buffer and transport entry are ready for reuse.

Fixes: fce8a7bb5b ("PCI-Express Non-Transparent Bridge Support")
Cc: stable@vger.kernel.org
Signed-off-by: Koichiro Den <den@valinux.co.jp>
Reviewed-by: Dave Jiang <dave.jiang@intel.com>
Link: https://patch.msgid.link/20260817053519.4135287-2-den@valinux.co.jp
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:43:24 -07:00
Stefan Wahren 8197c18005 docs: oa-tc6-framework: Fix link to specification
Current link for 10BASE-T1x MAC-PHY Serial Interface Specification
doesn't work - it returns 404. Update the link to the working one.

Signed-off-by: Stefan Wahren <wahrenst@gmx.net>
Acked-by: Randy Dunlap <rdunlap@infradead.org>
Link: https://patch.msgid.link/20260818135958.17311-1-wahrenst@gmx.net
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:31:54 -07:00
Hyunwoo Kim 03a9d10ecf sctp: drop a chunk if its transport was removed
sctp_rcv() resolves the transport once per packet and leaves it in
chunk->transport. The lookup reference, or the one sctp_add_backlog() takes
if the socket is owned by userspace, keeps it around until the chunk has
been processed.

An authenticated ASCONF DEL-IP can remove it in the meantime.
sctp_assoc_rm_peer() takes the transport out of the association and calls
sctp_transport_free(), which tags it dead and drops the reference the
association held. There is a window on both paths: the packet can sit on
the socket backlog, and on the direct path the lookup completes before
bh_lock_sock().

The DATA chunk in that packet puts the removed transport back into
asoc->peer.last_data_from. Once the packet is done that reference goes
away and the transport is freed by RCU, so the next delayed SACK carries
the pointer into the SACK chunk and sctp_outq_select_transport() reads the
freed transport's state.

Drop the chunk in sctp_inq_push(), next to the existing rcvr->dead check.
Both paths reach it with the association's socket lock held. The peer
retransmits it.

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Cc: stable@vger.kernel.org
Signed-off-by: Hyunwoo Kim <imv4bel@gmail.com>
Acked-by: Xin Long <lucien.xin@gmail.com>
Link: https://patch.msgid.link/aoUJHQmxL0LFIMCw@v4bel
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:31:09 -07:00
Triet Hoang 25b863cd6d tools: ynl: handle calloc failure in ynl_ntf_parse
Check the return value of calloc() before dereferencing the allocated
response structure in ynl_ntf_parse().

Signed-off-by: Triet Hoang <triet.hoang.dev@gmail.com>
Link: https://patch.msgid.link/20260818132739.469624-1-triet.hoang.dev@gmail.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:29:36 -07:00
Jamal Hadi Salim 4c660ee8c8 net: sched: fix 32-bit backlog wrap in gred, bfifo and plug enqueue
gred_enqueue(), bfifo_enqueue() and plug_enqueue() admit a packet when the
current backlog plus the packet length fits within the queue limit:

  sch->qstats.backlog + qdisc_pkt_len(skb) <= sch->limit (gred default VQ)
  gred_backlog+qdisc_pkt_len(skb) <= q->limit  (gred configured VQ)
  sch->qstats.backlog + qdisc_pkt_len(skb) <= sch->limit (bfifo)
  sch->qstats.backlog + skb->len <= q->limit             (plug)

sch->qstats.backlog and q->backlog are u32, and qdisc_pkt_len()/skb->len
are unsigned int, so all sums are computed in 32 bits and wrap at 2^32.
Once the true backlog exceeds 4 GiB the wrapped sum becomes small and
admission keeps succeeding, so the queue grows without bound and the kernel
can be driven to OOM.

Promote the sums to u64 so admission stops once the true backlog exceeds
the limit.  The limit is u32, so the bounded queue stays below 2^32 and
the stored u32 backlog never wraps.

The bug can only be reproduced as root (albeit with ridiculous setup):
 attach a gred (or bfifo/plug) qdisc with a limit near 4 GiB,
 leaving the default VQ unconfigured (for gred), and drive >4 GiB of
 queued traffic (e.g. via a size table / stab to inflate qdisc_pkt_len,
 or sustained high-rate traffic). The u32 backlog+len sum wraps at 2^32,
 admission keeps succeeding, and the queue grows unboundedly to OOM.

Fixes: a3eb95f891 ("net_sched: gred: add TCA_GRED_LIMIT attribute")
Reported-by: vega@nebusec.ai
Tested-by: Victor Nogueira <victor@mojatatu.com>
Signed-off-by: Jamal Hadi Salim <jhs@mojatatu.com>
Reviewed-by: Simon Horman <horms@kernel.org>
Link: https://patch.msgid.link/20260818095927.15901-1-jhs@mojatatu.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:28:45 -07:00
Mina Almasry d9c56501c7 net: tcp: block mixing readable and unreadable frags
Protect tcp_sendmsg_locked() from mistakenly mixing readable and
unreadable page fragments in the same SKB.

Check that the devmem binding matches the existing SKB's readability.
If a mismatch is detected, avoid collapsing and create a new segment.

Fixes: bd61848900 ("net: devmem: Implement TX path")
Suggested-by: Eric Dumazet <edumazet@google.com>
Cc: Pavel Begunkov <asml.silence@gmail.com>
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Bobby Eshleman <bobbyeshleman@gmail.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
Link: https://patch.msgid.link/20260814191336.187243-2-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:22:14 -07:00
Mina Almasry 68d8c65326 net: core: propagate unreadable flag in skb_zerocopy
skb_zerocopy() fails to propagate the unreadable flag when copying
unreadable fragments, causing target skbs to appear as readable memory.

This patch fixes the flag propagation. Additionally, it returns -EFAULT
if readable fragments are mixed with unreadable fragments during
extraction, and returns -EFAULT in openvswitch queue_userspace_packet().

Fixes: 65249feb6b ("net: add support for skbs with unreadable frags")
Cc: Stanislav Fomichev <sdf@fomichev.me>
Cc: Bobby Eshleman <bobbyeshleman@gmail.com>
Cc: Florian Westphal <fw@strlen.de>
Cc: Aaron Conole <aconole@redhat.com>
Cc: Eelco Chaudron <echaudro@redhat.com>
Cc: Willem de Bruijn <willemb@google.com>
Signed-off-by: Mina Almasry <almasrymina@google.com>
Reviewed-by: Pavel Begunkov <asml.silence@gmail.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Link: https://patch.msgid.link/20260814191336.187243-1-almasrymina@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:22:14 -07:00
Kyle Zeng 992cc9f94c net/packet: defer vmalloc TX_RING free until skbs finish
AF_PACKET TX_RING skbs keep a raw pointer to their ring frame. The skb
page references preserve page-backed ring blocks after pg_vec is freed,
but they do not preserve a vmalloc mapping.

tpacket_destruct_skb() currently drops the pending reference before
writing the timestamp and TP_STATUS_AVAILABLE to the frame. Move the
decrement after those stores. The smp_wmb() in __packet_set_status()
orders the frame stores before the decrement.

Also recheck pending TX frames under pg_vec_lock before non-closing
ring replacement, so a racing send cannot add a pending skb between
the initial check and the ring swap.

Ring allocation can produce a mixture of page-backed and vmalloc-backed
blocks. Allocate deferred-work storage during TX ring setup when the
first vmalloc-backed block is encountered, and keep its pointer in the
pg_vec allocation header. If allocation fails, return -ENOMEM from ring
setup. On socket close, a non-NULL pointer identifies a vmalloc-backed
vector without a scan. If TX skbs remain, defer the whole vector to
system_long_wq.

After pg_vec is detached, a late destructor can skip the pending
decrement. Use socket write-memory accounting as the deferred lifetime
gate instead: an skb remains charged through its final sock_wfree(),
after all ring-frame accesses. The delayed work retains a socket
reference and reschedules itself until no TX skbs remain.

Move pending_refcnt release to packet_sock_destruct() so late skb
destructors and deferred cleanup can safely use it after
packet_release(). Page-backed teardown remains synchronous, and no lock
is added to the TX completion hot path.

Fixes: b013840810 ("packet: use percpu mmap tx frame pending refcount")
Cc: stable@vger.kernel.org
Link: https://lore.kernel.org/netdev/20260721015824.45829-1-kylebot@openai.com/
Suggested-by: Eric Dumazet <edumazet@google.com>
Suggested-by: Willem de Bruijn <willemdebruijn.kernel@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Signed-off-by: Kyle Zeng <kylebot@openai.com>
Link: https://patch.msgid.link/20260816235646.76500-1-kylebot@openai.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:18:46 -07:00
Jakub Kicinski 51a26bb843 Merge branch 'ionic_rcq_shared' of https://github.com/abhijitG-xlnx/linux
Abhijit Gangurde says:

====================
Extend the net/ionic firmware identity structure to expose
the rcq_sign_bit field from the RDMA LIF identity.

* 'ionic_rcq_shared' of https://github.com/abhijitG-xlnx/linux:
  net: ionic: Fetch RCQ sign bit from firmware
====================

Link: https://patch.msgid.link/
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:07:01 -07:00
Eric Dumazet 447cbe95eb vlan: fix skb_under_panic and races when toggling HW VLAN offload
Toggling hardware VLAN TX offload (NETIF_F_HW_VLAN_CTAG_TX or
NETIF_F_HW_VLAN_STAG_TX) on a lower device invokes vlan_transfer_features(),
which dynamically changed vlandev->hard_header_len.

This causes two issues:
1. Lockless TX paths (e.g. packet_snd in af_packet.c, ip6_finish_output2)
   read dev->hard_header_len without holding RTNL lock. Mutating
   hard_header_len dynamically under RTNL creates a data race where upper
   layers reserve insufficient headroom based on a stale hard_header_len,
   resulting in skb_under_panic when vlan_dev_hard_header() is called.
2. In addition, vlan_transfer_features() updated hard_header_len without
   updating header_ops, causing a mismatch between allocated headroom
   and header creation.

Always setting dev->hard_header_len = real_dev->hard_header_len and
dev->needed_headroom = real_dev->needed_headroom + VLAN_HLEN unconditionally
ensures:
- dev->hard_header_len remains 100% static and immutable at real_dev->hard_header_len,
  eliminating all dynamic runtime updates and data races on hard_header_len.
- Upper layers allocating skbs via LL_RESERVED_SPACE() will always reserve
  sufficient headroom for software VLAN tag insertion (real_dev->hard_header_len +
  real_dev->needed_headroom + VLAN_HLEN).
- vlandev inherits real_dev->needed_tailroom so underlying trailer/padding/ICV
  requirements are honored.
- AF_PACKET SOCK_RAW network header offsets remain correctly aligned at
  real_dev->hard_header_len.
- vlan_header_ops is used unconditionally.

Note to stable teams: Make sure to backport these commits:

e16e960d55 ("ipvlan: inherit needed_headroom and needed_tailroom from phy_dev")
cef51860be ("macvlan: inherit needed_headroom and needed_tailroom from lowerdev")

Fixes: 1da177e4c3 ("Linux-2.6.12-rc2")
Reported-by: Tangxin Xie <xietangxin@h-partners.com>
Closes: https://lore.kernel.org/netdev/99d678ae-c7b2-4b44-b534-b8320679deb3@h-partners.com/
Cc: <stable@vger.kernel.org> # 3.19: e16e960d55: ipvlan: inherit needed_headroom and needed_tailroom from phy_dev
Cc: <stable@vger.kernel.org> # 3.19: cef51860be: macvlan: inherit needed_headroom and needed_tailroom from lowerdev
Cc: <stable@vger.kernel.org> # 3.19
Signed-off-by: Eric Dumazet <edumazet@google.com>
Link: https://patch.msgid.link/20260811085246.2267779-1-edumazet@google.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:05:43 -07:00
Zhan Xusheng d302d7109f ethtool: remove unused __ETHTOOL_LINK_MODE_MASK_NWORDS
From: Zhan Xusheng <zhanxusheng@xiaomi.com>

Added by commit f625aa9be8 ("ethtool: provide link mode information with
LINKMODES_GET request") and never used.  The same count is computed as
__ETHTOOL_LINK_MODE_MASK_NU32 in net/ethtool/ioctl.c.

Signed-off-by: Zhan Xusheng <zhanxusheng@xiaomi.com>
Link: https://patch.msgid.link/20260818023704.125721-1-zhanxusheng@xiaomi.com
Signed-off-by: Jakub Kicinski <kuba@kernel.org>
2026-08-20 13:04:44 -07:00