GHSA-P5XX-69G5-8564
Vulnerability from github – Published: 2026-07-24 18:31 – Updated: 2026-07-27 06:30In the Linux kernel, the following vulnerability has been resolved:
net/mlx5e: xsk: Fix unlocked writing to ICOSQ
During napi poll, when the affinity changes and there's still XSK work to be done, we trigger an ICOSQ interrupt on the new CPU. However, this triggering on the ICOSQ is done unprotected.
There are 2 such races:
A) mlx5e_trigger_irq() is called while mlx5e_xsk_alloc_rx_mpwqe() is running from a different CPU due to affinity change. This can happen because IRQ triggering is done after napi_complete_done(). At this point the NAPI can be scheduled on a different CPU. Like this:
CPU A (old affinity, NAPI tail) CPU B (new affinity, fresh NAPI) ------------------------------- -------------------------------- napi_complete_done() clears SCHED mlx5e_cq_arm(...) napi_schedule_prep() sets SCHED mlx5e_napi_poll() mlx5e_xsk_alloc_rx_mpwqe() mlx5e_icosq_sync_lock() // noop memcpy 640 B UMR body advance sq->pc by 10 mlx5e_trigger_irq(&c->icosq) wqe_info[pi] = {NOP, 1} mlx5e_post_nop() advances sq->pc
B) mlx5e_trigger_irq() is called on the ICOSQ when mlx5e_trigger_napi_icosq() is running.
The obvious fix would be to lock the ICOSQ. But ICOSQ has an optimized locking scheme that doesn't work for this scenario. Kick the async ICOSQ instead which is always locked.
This issue was noticed in the wild with the following splat:
netdevice: ge-0-0-1: Bad OP in ICOSQ CQE: 0xd WARNING: drivers/net/ethernet/mellanox/mlx5/core/en_rx.c:826 [...] [...] Call Trace: mlx5e_napi_poll+0x11d/0x7f0 [mlx5_core] __napi_poll+0x30/0x200 ? skb_defer_free_flush+0x9c/0xc0 net_rx_action+0x2fe/0x3f0 handle_softirqs+0xd8/0x340 __irq_exit_rcu+0xbc/0xe0 common_interrupt+0x85/0xa0 asm_common_interrupt+0x26/0x40 [...] ---[ end trace 0000000000000000 ]--- mlx5_core 0000:08:00.0 ge-0-0-1: Error cqe on cqn 0x548, ci 0x2022, qn 0x8f4, opcode 0xd, syndrome 0x2, vendor syndrome 0x68 00000000: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000010: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000020: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000030: 00 00 00 00 01 00 68 02 01 00 08 f4 de 14 59 d2 WQE DUMP: WQ size 16384 WQ cur size 0, WQE index 0x1e14, len: 64 00000000: 00 00 00 01 d9 ed 80 02 00 00 00 01 d9 ed 90 02 00000010: 00 00 00 01 d9 ed a0 02 00 00 00 01 d9 ed b0 02 00000020: 00 00 00 01 d9 ed c0 02 00 00 00 01 d9 ed d0 02 00000030: 00 00 00 01 d9 ed e0 02 00 00 00 01 d9 ed f0 02 mlx5_core 0000:08:00.0 ge-0-0-1: Error cqe on cqn 0x548, ci 0x2023, qn 0x8f4, opcode 0xd, syndrome 0x5, vendor syndrome 0xf9 00000000: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000010: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000020: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00000030: 00 00 00 00 01 00 f9 05 01 00 08 f4 de 15 cf d2
{
"affected": [],
"aliases": [
"CVE-2026-64210"
],
"database_specific": {
"cwe_ids": [],
"github_reviewed": false,
"github_reviewed_at": null,
"nvd_published_at": "2026-07-24T16:16:48Z",
"severity": "HIGH"
},
"details": "In the Linux kernel, the following vulnerability has been resolved:\n\nnet/mlx5e: xsk: Fix unlocked writing to ICOSQ\n\nDuring napi poll, when the affinity changes and there\u0027s still XSK work\nto be done, we trigger an ICOSQ interrupt on the new CPU. However, this\ntriggering on the ICOSQ is done unprotected.\n\nThere are 2 such races:\n\nA) mlx5e_trigger_irq() is called while mlx5e_xsk_alloc_rx_mpwqe() is\nrunning from a different CPU due to affinity change. This can happen\nbecause IRQ triggering is done after napi_complete_done(). At this point\nthe NAPI can be scheduled on a different CPU. Like this:\n\n CPU A (old affinity, NAPI tail) CPU B (new affinity, fresh NAPI)\n ------------------------------- --------------------------------\n napi_complete_done() clears SCHED\n mlx5e_cq_arm(...)\n napi_schedule_prep() sets SCHED\n mlx5e_napi_poll()\n mlx5e_xsk_alloc_rx_mpwqe()\n mlx5e_icosq_sync_lock() // noop\n memcpy 640 B UMR body\n advance sq-\u003epc by 10\n mlx5e_trigger_irq(\u0026c-\u003eicosq)\n wqe_info[pi] = {NOP, 1}\n mlx5e_post_nop() advances sq-\u003epc\n\nB) mlx5e_trigger_irq() is called on the ICOSQ when\nmlx5e_trigger_napi_icosq() is running.\n\nThe obvious fix would be to lock the ICOSQ. But ICOSQ has an optimized\nlocking scheme that doesn\u0027t work for this scenario. Kick the async ICOSQ\ninstead which is always locked.\n\nThis issue was noticed in the wild with the following splat:\n\n netdevice: ge-0-0-1: Bad OP in ICOSQ CQE: 0xd\n WARNING: drivers/net/ethernet/mellanox/mlx5/core/en_rx.c:826 [...]\n [...]\n Call Trace:\n \u003cIRQ\u003e\n mlx5e_napi_poll+0x11d/0x7f0 [mlx5_core]\n __napi_poll+0x30/0x200\n ? skb_defer_free_flush+0x9c/0xc0\n net_rx_action+0x2fe/0x3f0\n handle_softirqs+0xd8/0x340\n __irq_exit_rcu+0xbc/0xe0\n common_interrupt+0x85/0xa0\n \u003c/IRQ\u003e\n \u003cTASK\u003e\n asm_common_interrupt+0x26/0x40\n [...]\n ---[ end trace 0000000000000000 ]---\n mlx5_core 0000:08:00.0 ge-0-0-1: Error cqe on cqn 0x548, ci 0x2022, qn 0x8f4,\n opcode 0xd, syndrome 0x2, vendor syndrome 0x68\n 00000000: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n 00000010: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n 00000020: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n 00000030: 00 00 00 00 01 00 68 02 01 00 08 f4 de 14 59 d2\n WQE DUMP: WQ size 16384 WQ cur size 0, WQE index 0x1e14, len: 64\n 00000000: 00 00 00 01 d9 ed 80 02 00 00 00 01 d9 ed 90 02\n 00000010: 00 00 00 01 d9 ed a0 02 00 00 00 01 d9 ed b0 02\n 00000020: 00 00 00 01 d9 ed c0 02 00 00 00 01 d9 ed d0 02\n 00000030: 00 00 00 01 d9 ed e0 02 00 00 00 01 d9 ed f0 02\n mlx5_core 0000:08:00.0 ge-0-0-1: Error cqe on cqn 0x548, ci 0x2023, qn 0x8f4,\n opcode 0xd, syndrome 0x5, vendor syndrome 0xf9\n 00000000: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n 00000010: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n 00000020: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00\n 00000030: 00 00 00 00 01 00 f9 05 01 00 08 f4 de 15 cf d2",
"id": "GHSA-p5xx-69g5-8564",
"modified": "2026-07-27T06:30:30Z",
"published": "2026-07-24T18:31:26Z",
"references": [
{
"type": "ADVISORY",
"url": "https://nvd.nist.gov/vuln/detail/CVE-2026-64210"
},
{
"type": "WEB",
"url": "https://git.kernel.org/stable/c/8d3b91e7d81000d295cd914d4d9d6f860252e2bf"
},
{
"type": "WEB",
"url": "https://git.kernel.org/stable/c/c326f9c68921e2f14dfcecb2f6b4216313d50248"
}
],
"schema_version": "1.4.0",
"severity": [
{
"score": "CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:N/A:H",
"type": "CVSS_V3"
}
]
}
Sightings
| Author | Source | Type | Date | Other |
|---|
Nomenclature
- Seen: The vulnerability was mentioned, discussed, or observed by the user.
- Confirmed: The vulnerability has been validated from an analyst's perspective.
- Published Proof of Concept: A public proof of concept is available for this vulnerability.
- Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
- Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
- Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
- Not confirmed: The user expressed doubt about the validity of the vulnerability.
- Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.