CVE-2026-105752
Received Received - Intake

Information Disclosure in vLLM via Cache Salt Bypass

Vulnerability report for CVE-2026-105752, including description, CVSS score, EPSS score, affected products, exploitability, helpful resources, and attack-flow context.

Publication date: 2026-10-05

Last updated on: 2026-10-05

Assigner: GitHub, Inc.

Description

vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cache_salt value, placing the continuation prefix in the global unsalted cache namespace even when the caller enabled salting. On deployments with prefix caching enabled, which is the default, an authenticated tenant who can reconstruct a victim's low-entropy post-tool history can submit the same continuation and use the cached_tokens_per_turn count to determine whether the prefix was previously processed, defeating the intended tenant isolation of salted prefix caching. This issue is fixed in version 0.30.0.

CVSS Scores

EPSS Scores

Probability:
Percentile:

Meta Information

Published
2026-10-05
Last Modified
2026-10-05
Generated
2026-10-06
AI Q&A
2026-10-06
EPSS Evaluated
N/A
NVD
EUVD

Affected Vendors & Products

Showing 1 associated CPE
Vendor Product Version / Range
vllm-project vllm < 0.30.0

Helpful Resources

Exploitability

CWE
CWE Icon
KEV
KEV Icon
CWE ID Description
CWE-524 The code uses a cache that contains sensitive information, but the cache can be read by an actor outside of the intended control sphere.
CWE-200 The product exposes sensitive information to an actor that is not explicitly authorized to have access to that information.

Attack-Flow Graph

AI Quick Actions

Instant insights powered by AI
Executive Summary

This vulnerability in vLLM before version 0.30.0 allows an authenticated user to exploit prefix caching to infer whether a victim tenant processed a specific low-entropy continuation. The issue occurs because the system rebuilds engine input without preserving the cache_salt value, placing continuations in an unsalted global cache. This defeats tenant isolation intended by salted prefix caching.

Detection Guidance

This vulnerability requires checking if your vLLM deployment is running a version prior to 0.30.0 and if prefix caching is enabled. Inspect the vLLM version with 'pip show vllm' or 'vllm --version'. Verify prefix caching settings in your configuration files or environment variables. No specific commands are provided for detection beyond version and configuration checks.

Impact Analysis

If you are an authenticated user on a shared vLLM deployment with prefix caching enabled, an attacker could reconstruct your past interactions and determine if you previously processed certain continuations. This could expose information about your usage patterns or interactions with the system.

Compliance Impact

This vulnerability could potentially impact compliance with GDPR and HIPAA by undermining tenant isolation in multi-tenant environments. If an authenticated tenant exploits this flaw, they may infer whether a victim's data was processed through cached tokens, which could expose sensitive information or processing history. This may violate data protection principles under GDPR (e.g., integrity and confidentiality) and HIPAA (e.g., access controls and audit requirements).

Mitigation Strategies

Upgrade vLLM to version 0.30.0 or later to address the prefix caching issue and ensure tenant isolation is maintained.

Chat Assistant

Ask questions about this CVE
Hi! I’m here to help you understand CVE-2026-105752. Ask me anything about the vulnerability, its impact, or mitigation strategies.
0/70

EPSS Chart