Listen to this Post
CVE-2026-102994 is a high-severity vulnerability in the pypdf library, a pure-Python PDF manipulation toolkit. The flaw resides in the parsing logic for indirect objects, a fundamental PDF structure used to reference objects by number and generation. When pypdf reads a PDF, it must locate indirect object headers, which are typically followed by whitespace to delimit the object number and generation number. The vulnerable code, specifically the `read_until_whitespace` function in `pypdf/_reader.py` and pypdf/generic/_base.py, scans the input stream until it encounters a whitespace character. An attacker can craft a malicious PDF where an indirect object header contains an extremely long sequence of digits without any terminating whitespace. When pypdf attempts to parse this malformed header, the `read_until_whitespace` routine enters an effectively unbounded loop, consuming significant CPU time and memory as it scans the entire malicious token. This results in long runtimes and can ultimately lead to application unavailability, constituting a denial-of-service (DoS) condition. The vulnerability is classified under CWE-400 (Uncontrolled Resource Consumption) and CWE-407 (Inefficient Algorithmic Complexity). The issue affects all versions of pypdf prior to 6.18.0. The CVSS score is 8.7, reflecting the high impact on availability. An attacker can exploit this by supplying a crafted PDF to any application that uses pypdf to process PDF files, such as a web service that allows PDF uploads or a document processing pipeline. The attack does not require authentication or user interaction, making it remotely exploitable. The core problem is that the parser lacks a reasonable upper bound on the length of these tokens, allowing a single malformed object to exhaust system resources. The fix introduced in pypdf 6.18.0 limits the allowed length of indirect object tokens, preventing the excessive scanning. This vulnerability highlights the importance of robust input validation and resource limits when parsing untrusted file formats.
DailyCVE Form:
Platform: pypdf
Version: < 6.18.0
Vulnerability: Uncontrolled Resource Consumption
Severity: High
date: Sep 7, 2026
Prediction: Sep 7, 2026
What Undercode Say:
pip show pypdf | grep Version
import pypdf from pypdf import PdfReader Vulnerable code path (simplified) A PDF with an indirect object header like "1234567890... (no whitespace)" will cause read_until_whitespace to scan excessively.
Exploit: (Educational Purposes!)
Educational example: crafting a PDF with a long indirect object token This is for understanding the vulnerability only. malicious_pdf_content = b"""%PDF-1.4 1 0 obj << /Type /Catalog /Pages 2 0 R >> endobj 2 0 obj << /Type /Pages /Kids [3 0 R] /Count 1 >> endobj 3 0 obj << /Type /Page /Parent 2 0 R /MediaBox [0 0 612 792] >> endobj """ + b"9" 1000000 + b" 0 obj\n<< >>\nendobj\n" The long sequence of '9's without whitespace will trigger the vulnerability when pypdf parses the indirect object header.
Protection: from this CVE
Upgrade to pypdf version 6.18.0 or later. If immediate upgrade is not possible, apply the changes from PR 4055 manually. Implement resource limits and input validation to reject PDFs with excessively long object tokens.
Impact:
Denial of Service (DoS) due to excessive CPU and memory consumption, leading to application unavailability.
🎯Let’s Practice Exploiting & Learn Patching For Free:
🎓 Live Courses & Certifications:
Join Undercode Academy for Verified Certifications
🚀 Request a Custom Project:
Secure, high-velocity infrastructure and disruptive technological engineering. Contact our engineering team for high-tier development and proprietary systems:
[email protected]
💎 Smart Architecture | 🛡️ Secure by Design | ⭐ Trusted by Thousands
Sources:
Reported By: github.com
Extra Source Hub:
Undercode

