GHSA-PH3R-5JFG-F84F

Vulnerability from github – Published: 2026-10-06 00:02 – Updated: 2026-10-06 00:02
VLAI
Summary
vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core
Details

Affected

  • Ecosystem / package: pip / vllm
  • Affected versions: vLLM ≤ 0.25.1 (confirmed on 0.25.1, commit 752a3a504485). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches.

Summary

vLLM's default multimodal cache (mm_processor_cache_type="lru") mirrors state across two processes: the frontend (P0) holds only metadata (MultiModalProcessorSenderCache) while the engine core (P1) holds the real payload (MultiModalReceiverCache). The design invariant is that get_and_update() runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer "is this cached in P1?" without talking to P1.

That invariant breaks when a request is rejected after P0 has rendered and hashed the multimodal input (populating the P0 cache) but before P1 receives the item — for example, an oversized chat prompt rejected on max_model_len after rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends None instead of the payload, and P1 — which has nothing cached — trips assert mm_item is not None, f"Expected a cached item for {mm_hash=}".

This is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration.

Affected code

Links pinned to the confirmed commit 752a3a504485 (v0.25.1):

The P1 receiver sink — the assertion that fires when P1 receives None for a hash it never cached:

# vllm/multimodal/cache.py Lines 651-663
    @override
    def get_and_update_item(
        self,
        mm_item: MultiModalKwargsItem | None,
        mm_hash: str,
    ) -> MultiModalKwargsItem:
        if (cached_item := self._cache.get(mm_hash)) is not None:
            return cached_item

        assert mm_item is not None, f"Expected a cached item for {mm_hash=}"

        self._cache[mm_hash] = mm_item
        return mm_item

The P0 sender — on a hit it drops the payload (returns None) and, on a miss during render, unconditionally commits the metadata entry with no rollback tied to admission:

# vllm/multimodal/cache.py Lines 409-422
    @override
    def get_and_update_item(
        self,
        mm_item: MultiModalProcessorCacheInItem,
        mm_hash: str,
    ) -> MultiModalProcessorCacheOutItem:
        if (cached_item := self._cache.get(mm_hash)) is not None:
            return None, cached_item.prompt_updates

        assert mm_item is not None, f"Expected a cached item for {mm_hash=}"

        self._cache[mm_hash] = MultiModalProcessorCacheItemMetadata(*mm_item)

        return mm_item

Impact

A remote client submitting multimodal requests can poison a cache identity — render a media item successfully, then have that request rejected — so a later request reusing the same media hash fails on the P1 receiver assertion. This is an availability failure against a shared serving instance. No code execution, memory corruption, or data disclosure is claimed.

On this revision the failure is scoped as a request-level preprocessing error (the engine core catches around preprocess_add_request); public reports show the same assertion cascading into further engine-loop assertions on other revisions. It applies to multimodal models running the default mirrored lru cache.

Suggested Fix

Two complementary changes:

  1. Make the mirrored commit atomic with admission — insert into the P0 sender cache only after the request has passed all admission checks (length, limits) and P1 has acknowledged the item, or roll back the P0 insert on rejection.
  2. Defense in depth — convert the P1 receiver assert mm_item is not None into a checked, request-scoped error (fetch-on-miss from P0) so a desync degrades a single request rather than asserting in the engine loop.

The core of the rollback half: wrap the post-render length check so a ValueError rejection discards the P0 entries the render just committed, before re-raising. Add a discard_sender_cache_item() on the processor cache (no-op default, pop on the sender) and a Renderer.discard_mm_cache_entries() that walks a rendered request's mm_hashes:

# vllm/entrypoints/openai/chat_completion/serving.py (_create_chat_completion)
-            max_tokens = get_max_tokens(
-                max_model_len,
-                ...,
-                truncate_prompt_tokens=request.truncate_prompt_tokens,
-            )
+            try:
+                max_tokens = get_max_tokens(
+                    max_model_len,
+                    ...,
+                    truncate_prompt_tokens=request.truncate_prompt_tokens,
+                )
+            except ValueError:
+                for rendered_input in engine_inputs:
+                    if mm_hashes := rendered_input.get("mm_hashes"):
+                        self.renderer.discard_mm_cache_entries(mm_hashes)
+                raise
# vllm/multimodal/cache.py (MultiModalProcessorSenderCache)
+    @override
+    def discard_sender_cache_item(self, mm_hash: str) -> None:
+        self._cache.pop(mm_hash, None)

This closes the max_model_len rejection path; because any other rejection-after-render path reopens the same window, pairing it with the defense-in-depth change above (making the P1 assert a checked, request-scoped error) is recommended.

Credit

Reported by: Patch the Planet (Trail of Bits + OpenAI collaboration)

This vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.


Proposed fix: a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51897

Show details on source website

{
  "affected": [
    {
      "package": {
        "ecosystem": "PyPI",
        "name": "vllm"
      },
      "ranges": [
        {
          "events": [
            {
              "introduced": "0"
            },
            {
              "fixed": "0.28.0"
            }
          ],
          "type": "ECOSYSTEM"
        }
      ]
    }
  ],
  "aliases": [
    "CVE-2026-105753"
  ],
  "database_specific": {
    "cwe_ids": [
      "CWE-617"
    ],
    "github_reviewed": true,
    "github_reviewed_at": "2026-10-06T00:02:05Z",
    "nvd_published_at": null,
    "severity": "MODERATE"
  },
  "details": "## Affected\n\n- **Ecosystem / package:** pip / `vllm`\n- **Affected versions:** vLLM \u2264 0.25.1 (confirmed on 0.25.1, commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34)). The lower bound predates 0.25.1; maintainers can confirm how far back the mirrored sender/receiver cache protocol reaches.\n\n## Summary\n\nvLLM\u0027s default multimodal cache (`mm_processor_cache_type=\"lru\"`) mirrors state across two processes: the frontend (P0) holds only metadata (`MultiModalProcessorSenderCache`) while the engine core (P1) holds the real payload (`MultiModalReceiverCache`). The design invariant is that `get_and_update()` runs on P0 and P1 in lockstep for every request, so eviction order stays mirrored and P0 can answer \"is this cached in P1?\" without talking to P1.\n\nThat invariant breaks when a request is **rejected after P0 has rendered and hashed the multimodal input** (populating the P0 cache) **but before P1 receives the item** \u2014 for example, an oversized chat prompt rejected on `max_model_len` *after* rendering. P0 now believes the media is cached while P1 never got it. A later request reusing the same media hash gets a P0 hit, so P0 sends `None` instead of the payload, and P1 \u2014 which has nothing cached \u2014 trips `assert mm_item is not None, f\"Expected a cached item for {mm_hash=}\"`.\n\nThis is a remotely reachable, request-controlled cache-mirroring desync on the standard multimodal inference path. It requires only the default cache configuration.\n\n## Affected code\n\nLinks pinned to the confirmed commit [`752a3a504485`](https://github.com/vllm-project/vllm/tree/752a3a504485790a2e8491cacbb35c137339ad34) (v0.25.1):\n\n- **P0 metadata cache** \u2014 `MultiModalProcessorSenderCache` at [`vllm/multimodal/cache.py#L379`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L379); `get_and_update_item` at [`#L410-L421`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L410-L421), commit assertion at [`#L418`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L418).\n- **P1 payload cache (the actual sink)** \u2014 `MultiModalReceiverCache` at [`vllm/multimodal/cache.py#L630`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L630); `get_and_update_item` at [`#L652-L663`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L652-L663), with the failing `assert mm_item is not None, f\"Expected a cached item for {mm_hash=}\"` at [`#L660`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/cache.py#L660).\n- **Default `mm_processor_cache_type = \"lru\"`** at [`vllm/config/multimodal.py#L132`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/config/multimodal.py#L132); dispatch in [`vllm/multimodal/registry.py#L294-L307`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/registry.py#L294-L307) (sender) and [`#L322-L331`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/registry.py#L322-L331) (receiver). The `processor_only`, disabled-caching, and `shm` paths are not affected.\n- **Rejection-after-render window** \u2014 rendering happens before length validation in [`vllm/entrypoints/openai/chat_completion/serving.py#L206-L231`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/openai/chat_completion/serving.py#L206-L231) (`render_chat_request`), and the length check raises after the render in [`vllm/entrypoints/serve/utils/api_utils.py#L171-L189`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/entrypoints/serve/utils/api_utils.py#L171-L189).\n- **Cache-commit call site** \u2014 [`vllm/multimodal/processing/processor.py#L1347`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/multimodal/processing/processor.py#L1347) (`_merge_mm_kwargs`) commits the P0 sender entry during render.\n- The separate stale-order eviction hang is already fixed via [`vllm/utils/cache.py#L120-L121`](https://github.com/vllm-project/vllm/blob/752a3a504485790a2e8491cacbb35c137339ad34/vllm/utils/cache.py#L120-L121) (`LRUCache.touch()` guarding `if key in self:`) \u2014 a different mechanism that does not touch the sender/receiver commit-ordering protocol and does not remediate this assertion.\n\nThe P1 receiver sink \u2014 the assertion that fires when P1 receives `None` for a hash it never cached:\n\n```python\n# vllm/multimodal/cache.py Lines 651-663\n    @override\n    def get_and_update_item(\n        self,\n        mm_item: MultiModalKwargsItem | None,\n        mm_hash: str,\n    ) -\u003e MultiModalKwargsItem:\n        if (cached_item := self._cache.get(mm_hash)) is not None:\n            return cached_item\n\n        assert mm_item is not None, f\"Expected a cached item for {mm_hash=}\"\n\n        self._cache[mm_hash] = mm_item\n        return mm_item\n```\n\nThe P0 sender \u2014 on a hit it drops the payload (returns `None`) and, on a miss during render, unconditionally commits the metadata entry with no rollback tied to admission:\n\n```python\n# vllm/multimodal/cache.py Lines 409-422\n    @override\n    def get_and_update_item(\n        self,\n        mm_item: MultiModalProcessorCacheInItem,\n        mm_hash: str,\n    ) -\u003e MultiModalProcessorCacheOutItem:\n        if (cached_item := self._cache.get(mm_hash)) is not None:\n            return None, cached_item.prompt_updates\n\n        assert mm_item is not None, f\"Expected a cached item for {mm_hash=}\"\n\n        self._cache[mm_hash] = MultiModalProcessorCacheItemMetadata(*mm_item)\n\n        return mm_item\n```\n\n## Impact\n\nA remote client submitting multimodal requests can poison a cache identity \u2014 render a media item successfully, then have that request rejected \u2014 so a **later** request reusing the same media hash fails on the P1 receiver assertion. This is an availability failure against a shared serving instance. No code execution, memory corruption, or data disclosure is claimed.\n\nOn this revision the failure is scoped as a request-level preprocessing error (the engine core catches around `preprocess_add_request`); public reports show the same assertion cascading into further engine-loop assertions on other revisions. It applies to multimodal models running the default mirrored `lru` cache.\n\n\n## Suggested Fix\n\nTwo complementary changes:\n\n1. **Make the mirrored commit atomic with admission** \u2014 insert into the P0 sender cache only after the request has passed all admission checks (length, limits) and P1 has acknowledged the item, or roll back the P0 insert on rejection.\n2. **Defense in depth** \u2014 convert the P1 receiver `assert mm_item is not None` into a checked, request-scoped error (fetch-on-miss from P0) so a desync degrades a single request rather than asserting in the engine loop.\n\nThe core of the rollback half: wrap the post-render length check so a `ValueError` rejection discards the P0 entries the render just committed, before re-raising. Add a `discard_sender_cache_item()` on the processor cache (no-op default, `pop` on the sender) and a `Renderer.discard_mm_cache_entries()` that walks a rendered request\u0027s `mm_hashes`:\n\n```python\n# vllm/entrypoints/openai/chat_completion/serving.py (_create_chat_completion)\n-            max_tokens = get_max_tokens(\n-                max_model_len,\n-                ...,\n-                truncate_prompt_tokens=request.truncate_prompt_tokens,\n-            )\n+            try:\n+                max_tokens = get_max_tokens(\n+                    max_model_len,\n+                    ...,\n+                    truncate_prompt_tokens=request.truncate_prompt_tokens,\n+                )\n+            except ValueError:\n+                for rendered_input in engine_inputs:\n+                    if mm_hashes := rendered_input.get(\"mm_hashes\"):\n+                        self.renderer.discard_mm_cache_entries(mm_hashes)\n+                raise\n```\n\n```python\n# vllm/multimodal/cache.py (MultiModalProcessorSenderCache)\n+    @override\n+    def discard_sender_cache_item(self, mm_hash: str) -\u003e None:\n+        self._cache.pop(mm_hash, None)\n```\n\nThis closes the `max_model_len` rejection path; because any other rejection-after-render path reopens the same window, pairing it with the defense-in-depth change above (making the P1 `assert` a checked, request-scoped error) is recommended.\n\n## Credit\n\n**Reported by:** Patch the Planet (Trail of Bits + OpenAI collaboration)\n\nThis vulnerability was discovered using GPT-5.5-Cyber as part of the Patch the Planet security initiative.\n\n---\n\n**Proposed fix:** a fix for this issue is proposed in a public pull request: https://github.com/vllm-project/vllm/pull/51897",
  "id": "GHSA-ph3r-5jfg-f84f",
  "modified": "2026-10-06T00:02:06Z",
  "published": "2026-10-06T00:02:05Z",
  "references": [
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/security/advisories/GHSA-ph3r-5jfg-f84f"
    },
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/pull/46747"
    },
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/pull/51897"
    },
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/commit/396204230423b7cc6798300926b8fa30190d26a9"
    },
    {
      "type": "PACKAGE",
      "url": "https://github.com/vllm-project/vllm"
    },
    {
      "type": "WEB",
      "url": "https://github.com/vllm-project/vllm/releases/tag/v0.28.0"
    }
  ],
  "schema_version": "1.4.0",
  "severity": [
    {
      "score": "CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H",
      "type": "CVSS_V3"
    }
  ],
  "summary": "vLLM: Mirrored multimodal IPC caches desync after a rejected request \u2014 a later request reusing the same media hash trips a receiver assertion in the engine core"
}



Log in or create an account to share your comment.




Tags
Taxonomy of the tags.


Loading…

Loading…

Loading…

Forecast uses a logistic model when the trend is rising, or an exponential decay model when the trend is falling. Fitted via linearized least squares.

Sightings

Author Source Type Date Other

Nomenclature

  • Seen: The vulnerability was mentioned, discussed, or observed by the user.
  • Confirmed: The vulnerability has been validated from an analyst's perspective.
  • Published Proof of Concept: A public proof of concept is available for this vulnerability.
  • Exploited: The vulnerability was observed as exploited by the user who reported the sighting.
  • Patched: The vulnerability was observed as successfully patched by the user who reported the sighting.
  • Not exploited: The vulnerability was not observed as exploited by the user who reported the sighting.
  • Not confirmed: The user expressed doubt about the validity of the vulnerability.
  • Not patched: The vulnerability was not observed as successfully patched by the user who reported the sighting.

Loading…

Loading…

Loading…

Related by attack behaviour

Vulnerabilities whose description is nearest to this one in the vector space of the CIRCL/vulnerability-attack-technique-biencoder model. This is a similarity search over the bi-encoder space (plain cosine), not a classification, and it has no measured accuracy.


Loading…