<?xml version='1.0' encoding='UTF-8'?>
<?xml-stylesheet href="/static/style.xsl" type="text/xsl"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" version="2.0">
  <channel>
    <title>Most recent entries from pysec</title>
    <link>https://127.0.0.1:10001</link>
    <description>Contains only the most 10 recent entries.</description>
    <docs>http://www.rssboard.org/rss-specification</docs>
    <generator>python-feedgen</generator>
    <language>en</language>
    <lastBuildDate>Sat, 03 Oct 2026 23:16:31 +0000</lastBuildDate>
    <item>
      <title>pysec-2025-19</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2025-19</link>
      <description>picklescan before 0.0.22 only considers standard pickle file extensions in the scope for its vulnerability scan. An attacker could craft a malicious model that uses Pickle and include a malicious pickle file with a non-standard file extension. Because the malicious pickle file inclusion is not considered as part of the scope of picklescan, the file would pass security checks and appear to be safe, when it could instead prove to be problematic.</description>
      <content:encoded>picklescan before 0.0.22 only considers standard pickle file extensions in the scope for its vulnerability scan. An attacker could craft a malicious model that uses Pickle and include a malicious pickle file with a non-standard file extension. Because the malicious pickle file inclusion is not considered as part of the scope of picklescan, the file would pass security checks and appear to be safe, when it could instead prove to be problematic.</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2025-19</guid>
      <pubDate>Mon, 03 Mar 2025 19:15:34 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2024-115</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2024-115</link>
      <description>A vulnerability in the GraphCypherQAChain class of langchain-ai/langchain-community version 0.2.5 allows for SQL injection through prompt injection. This vulnerability can lead to unauthorized data manipulation, data exfiltration, denial of service (DoS) by deleting all data, breaches in multi-tenant security environments, and data integrity issues. Attackers can create, update, or delete nodes and relationships without proper authorization, extract sensitive data, disrupt services, access data across different tenants, and compromise the integrity of the database.</description>
      <content:encoded>A vulnerability in the GraphCypherQAChain class of langchain-ai/langchain-community version 0.2.5 allows for SQL injection through prompt injection. This vulnerability can lead to unauthorized data manipulation, data exfiltration, denial of service (DoS) by deleting all data, breaches in multi-tenant security environments, and data integrity issues. Attackers can create, update, or delete nodes and relationships without proper authorization, extract sensitive data, disrupt services, access data across different tenants, and compromise the integrity of the database.</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2024-115</guid>
      <pubDate>Tue, 05 Nov 2024 16:04:14 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2026-2946</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2026-2946</link>
      <description>## Summary

The approval system in PraisonAI Agents caches tool approval decisions by tool name only, not by invocation arguments. Once a user approves `execute_command` for any command (e.g., `ls -la`), all subsequent `execute_command` calls in that execution context bypass the approval prompt entirely. Combined with `os.environ.copy()` passing all process environment variables to subprocesses, this allows an LLM agent (potentially via prompt injection) to silently exfiltrate API keys and credentials without further user consent.

## Details

The `require_approval` decorator in `src/praisonai-agents/praisonaiagents/approval/__init__.py:176-178` checks approval status by tool name only:

```python
@wraps(func)
def wrapper(*args, **kwargs):
    if is_already_approved(tool_name):   # line 177 — checks only tool_name
        return func(*args, **kwargs)     # line 178 — bypasses ALL approval
```

The `mark_approved` function in `registry.py:144-147` stores only the tool name string:

```python
def mark_approved(self, tool_name: str) -&gt; None:
    approved = self._approved_context.get(set())
    approved.add(tool_name)              # stores "execute_command", not args
    self._approved_context.set(approved)
```

The approval context is never cleared during agent execution — `clear_approved()` exists (`registry.py:152`) but is never called in the agent's tool execution path (`agent/tool_execution.py`).

Meanwhile, the `ConsoleBackend` UI at `backends.py:95-96` misleads the user:

```python
return Confirm.ask(
    f"Do you want to execute this {request.risk_level} risk tool?",
    # "this" implies per-invocation approval
)
```

The UI displays the specific command arguments (lines 81-85), creating a reasonable expectation that the user is approving only that specific invocation.

Additionally, `shell_tools.py:77` passes the full process environment to every subprocess:

```python
process_env = os.environ.copy()  # includes OPENAI_API_KEY, etc.
```

There is no command filtering, blocklist, or environment variable sanitization in the shell tools module.

## PoC

```python
from praisonaiagents import Agent
from praisonaiagents.tools.shell_tools import execute_command

# Step 1: Create agent with shell tool
agent = Agent(
    name="worker",
    instructions="You are a helpful assistant.",
    tools=[execute_command]
)

# Step 2: Agent requests benign command — user sees Rich panel:
#   Function: execute_command
#   Risk Level: CRITICAL
#   Arguments:
#     command: ls -la
#   "Do you want to execute this critical risk tool?" [y/N]
# User approves → mark_approved("execute_command") is called

# Step 3: All subsequent execute_command calls bypass approval silently:
# execute_command(command="env")
#   → returns ALL environment variables (OPENAI_API_KEY, AWS_SECRET_ACCESS_KEY, etc.)
#   → NO approval prompt shown

# Step 4: Targeted extraction also bypasses approval:
# execute_command(command="printenv OPENAI_API_KEY")
#   → returns the specific API key
#   → NO approval prompt shown

# Verification: check the approval cache
from praisonaiagents.approval import is_already_approved
# After approving "ls -la":
# is_already_approved("execute_command") → True
# Any execute_command call now returns immediately at __init__.py:177-178
```

## Impact

- **Secret exfiltration**: An LLM agent (or one subjected to prompt injection) can dump all process environment variables after a single benign command approval. Common secrets include `OPENAI_API_KEY`, `AWS_SECRET_ACCESS_KEY`, `DATABASE_URL`, and any other credentials passed via environment.
- **Misleading consent UI**: The console prompt displays specific arguments and uses language ("this tool") that implies per-invocation consent, but the system grants session-wide blanket approval.
- **No expiration or scope**: The approval cache uses a `ContextVar` that persists for the entire agent execution context with no timeout, no command-count limit, and no clearing between tool calls.
- **No environment filtering**: `os.environ.copy()` passes every environment variable to subprocesses without filtering sensitive patterns.

## Recommended Fix

1. **Per-invocation approval for critical tools** — store a hash of `(tool_name, arguments)` instead of just `tool_name`, or require re-approval for each invocation of critical-risk tools:

```python
# In registry.py — change mark_approved/is_already_approved:
import hashlib, json

def mark_approved(self, tool_name: str, arguments: dict = None) -&gt; None:
    approved = self._approved_context.get(set())
    risk = self._risk_levels.get(tool_name)
    if risk == "critical" and arguments:
        key = f"{tool_name}:{hashlib.sha256(json.dumps(arguments, sort_keys=True).encode()).hexdigest()}"
    else:
        key = tool_name
    approved.add(key)
    self._approved_context.set(approved)

def is_already_approved(self, tool_name: str, arguments: dict = None) -&gt; bool:
    approved = self._approved_context.get(set())
    risk = self._risk_levels.get(tool_name)
    if risk == "critical" and arguments:
        key = f"{tool_name}:{hashlib.sha256(json.dumps(arguments, sort_keys=True).encode()).hexdigest()}"
        return key in approved
    return tool_name in approved
```

2. **Filter environment variables** in `shell_tools.py`:

```python
SENSITIVE_PATTERNS = ('_KEY', '_SECRET', '_TOKEN', '_PASSWORD', '_CREDENTIAL')

process_env = {
    k: v for k, v in os.environ.items()
    if not any(p in k.upper() for p in SENSITIVE_PATTERNS)
}
if env:
    process_env.update(env)
```</description>
      <content:encoded>## Summary

The approval system in PraisonAI Agents caches tool approval decisions by tool name only, not by invocation arguments. Once a user approves `execute_command` for any command (e.g., `ls -la`), all subsequent `execute_command` calls in that execution context bypass the approval prompt entirely. Combined with `os.environ.copy()` passing all process environment variables to subprocesses, this allows an LLM agent (potentially via prompt injection) to silently exfiltrate API keys and credentials without further user consent.

## Details

The `require_approval` decorator in `src/praisonai-agents/praisonaiagents/approval/__init__.py:176-178` checks approval status by tool name only:

```python
@wraps(func)
def wrapper(*args, **kwargs):
    if is_already_approved(tool_name):   # line 177 — checks only tool_name
        return func(*args, **kwargs)     # line 178 — bypasses ALL approval
```

The `mark_approved` function in `registry.py:144-147` stores only the tool name string:

```python
def mark_approved(self, tool_name: str) -&gt; None:
    approved = self._approved_context.get(set())
    approved.add(tool_name)              # stores "execute_command", not args
    self._approved_context.set(approved)
```

The approval context is never cleared during agent execution — `clear_approved()` exists (`registry.py:152`) but is never called in the agent's tool execution path (`agent/tool_execution.py`).

Meanwhile, the `ConsoleBackend` UI at `backends.py:95-96` misleads the user:

```python
return Confirm.ask(
    f"Do you want to execute this {request.risk_level} risk tool?",
    # "this" implies per-invocation approval
)
```

The UI displays the specific command arguments (lines 81-85), creating a reasonable expectation that the user is approving only that specific invocation.

Additionally, `shell_tools.py:77` passes the full process environment to every subprocess:

```python
process_env = os.environ.copy()  # includes OPENAI_API_KEY, etc.
```

There is no command filtering, blocklist, or environment variable sanitization in the shell tools module.

## PoC

```python
from praisonaiagents import Agent
from praisonaiagents.tools.shell_tools import execute_command

# Step 1: Create agent with shell tool
agent = Agent(
    name="worker",
    instructions="You are a helpful assistant.",
    tools=[execute_command]
)

# Step 2: Agent requests benign command — user sees Rich panel:
#   Function: execute_command
#   Risk Level: CRITICAL
#   Arguments:
#     command: ls -la
#   "Do you want to execute this critical risk tool?" [y/N]
# User approves → mark_approved("execute_command") is called

# Step 3: All subsequent execute_command calls bypass approval silently:
# execute_command(command="env")
#   → returns ALL environment variables (OPENAI_API_KEY, AWS_SECRET_ACCESS_KEY, etc.)
#   → NO approval prompt shown

# Step 4: Targeted extraction also bypasses approval:
# execute_command(command="printenv OPENAI_API_KEY")
#   → returns the specific API key
#   → NO approval prompt shown

# Verification: check the approval cache
from praisonaiagents.approval import is_already_approved
# After approving "ls -la":
# is_already_approved("execute_command") → True
# Any execute_command call now returns immediately at __init__.py:177-178
```

## Impact

- **Secret exfiltration**: An LLM agent (or one subjected to prompt injection) can dump all process environment variables after a single benign command approval. Common secrets include `OPENAI_API_KEY`, `AWS_SECRET_ACCESS_KEY`, `DATABASE_URL`, and any other credentials passed via environment.
- **Misleading consent UI**: The console prompt displays specific arguments and uses language ("this tool") that implies per-invocation consent, but the system grants session-wide blanket approval.
- **No expiration or scope**: The approval cache uses a `ContextVar` that persists for the entire agent execution context with no timeout, no command-count limit, and no clearing between tool calls.
- **No environment filtering**: `os.environ.copy()` passes every environment variable to subprocesses without filtering sensitive patterns.

## Recommended Fix

1. **Per-invocation approval for critical tools** — store a hash of `(tool_name, arguments)` instead of just `tool_name`, or require re-approval for each invocation of critical-risk tools:

```python
# In registry.py — change mark_approved/is_already_approved:
import hashlib, json

def mark_approved(self, tool_name: str, arguments: dict = None) -&gt; None:
    approved = self._approved_context.get(set())
    risk = self._risk_levels.get(tool_name)
    if risk == "critical" and arguments:
        key = f"{tool_name}:{hashlib.sha256(json.dumps(arguments, sort_keys=True).encode()).hexdigest()}"
    else:
        key = tool_name
    approved.add(key)
    self._approved_context.set(approved)

def is_already_approved(self, tool_name: str, arguments: dict = None) -&gt; bool:
    approved = self._approved_context.get(set())
    risk = self._risk_levels.get(tool_name)
    if risk == "critical" and arguments:
        key = f"{tool_name}:{hashlib.sha256(json.dumps(arguments, sort_keys=True).encode()).hexdigest()}"
        return key in approved
    return tool_name in approved
```

2. **Filter environment variables** in `shell_tools.py`:

```python
SENSITIVE_PATTERNS = ('_KEY', '_SECRET', '_TOKEN', '_PASSWORD', '_CREDENTIAL')

process_env = {
    k: v for k, v in os.environ.items()
    if not any(p in k.upper() for p in SENSITIVE_PATTERNS)
}
if env:
    process_env.update(env)
```</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2026-2946</guid>
      <pubDate>Mon, 13 Jul 2026 14:36:52 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2025-102</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2025-102</link>
      <description>Local File Inclusion in dagster._grpc.impl.get_notebook_data in Dagster 1.10.14 allows attackers with access to the gRPC server to read arbitrary files by supplying path traversal sequences in the notebook_path field of ExternalNotebookData requests, bypassing the intended extension-based check.</description>
      <content:encoded>Local File Inclusion in dagster._grpc.impl.get_notebook_data in Dagster 1.10.14 allows attackers with access to the gRPC server to read arbitrary files by supplying path traversal sequences in the notebook_path field of ExternalNotebookData requests, bypassing the intended extension-based check.</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2025-102</guid>
      <pubDate>Tue, 22 Jul 2025 17:15:33 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2026-4182</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2026-4182</link>
      <description>### Impact

Denial of service via memory exhaustion. Affects all callers who streamed compressed responses relying on the chunk size — explicit (`iter_bytes(chunk_size=...)`) or the default — to bound memory. The decoder ignored that bound, so a chunk could be far larger than requested and a single compressed response could overflow memory.

```python
import gzip, zapros

# Server returns ~1 GiB of zeros gzip-compressed to ~1 MiB,
# with header: Content-Encoding: gzip
bomb = gzip.compress(b"\0" * 1_000_000_000)  # ~1 MiB on the wire

with zapros.stream("GET", "https://malicious.example/bomb") as response:
    # Caller asks for 8 KiB chunks, expecting bounded memory:
    for chunk in response.iter_bytes(chunk_size=8192):
        ...  # first `chunk` is ~1 GiB, not 8 KiB -&gt; memory exhaustion
```

### Patches

Upgrade to `0.14.0` or later. The decoders now bound the output of each decompression step to the requested `chunk_size`: gzip/deflate via `zlib`'s `max_length` + `unconsumed_tail`, brotli via `output_buffer_limit`, and zstd via a bounded `stream_writer`. Peak memory during streaming decode is now proportional to `chunk_size` for all supported encodings.

### Workarounds

For unpatched versions:
- Read the still-compressed body with `Response.iter_raw()` / `Response.async_iter_raw()`, which bypass the built-in decoders, and decompress it yourself with an explicit output-size bound (e.g. `zlib`'s `max_length`), aborting once a configured limit is exceeded.
- Where feasible, send `Accept-Encoding: identity` to disable response compression so bodies are not decompressed client-side.
- Avoid decoding response bodies from untrusted servers.</description>
      <content:encoded>### Impact

Denial of service via memory exhaustion. Affects all callers who streamed compressed responses relying on the chunk size — explicit (`iter_bytes(chunk_size=...)`) or the default — to bound memory. The decoder ignored that bound, so a chunk could be far larger than requested and a single compressed response could overflow memory.

```python
import gzip, zapros

# Server returns ~1 GiB of zeros gzip-compressed to ~1 MiB,
# with header: Content-Encoding: gzip
bomb = gzip.compress(b"\0" * 1_000_000_000)  # ~1 MiB on the wire

with zapros.stream("GET", "https://malicious.example/bomb") as response:
    # Caller asks for 8 KiB chunks, expecting bounded memory:
    for chunk in response.iter_bytes(chunk_size=8192):
        ...  # first `chunk` is ~1 GiB, not 8 KiB -&gt; memory exhaustion
```

### Patches

Upgrade to `0.14.0` or later. The decoders now bound the output of each decompression step to the requested `chunk_size`: gzip/deflate via `zlib`'s `max_length` + `unconsumed_tail`, brotli via `output_buffer_limit`, and zstd via a bounded `stream_writer`. Peak memory during streaming decode is now proportional to `chunk_size` for all supported encodings.

### Workarounds

For unpatched versions:
- Read the still-compressed body with `Response.iter_raw()` / `Response.async_iter_raw()`, which bypass the built-in decoders, and decompress it yourself with an explicit output-size bound (e.g. `zlib`'s `max_length`), aborting once a configured limit is exceeded.
- Where feasible, send `Accept-Encoding: identity` to disable response compression so bodies are not decompressed client-side.
- Avoid decoding response bodies from untrusted servers.</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2026-4182</guid>
      <pubDate>Thu, 01 Oct 2026 16:38:38 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2026-4181</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2026-4181</link>
      <description>### Impact

**Who is impacted**:
  - Any application using Zapros to make HTTP requests to untrusted servers
  - Applications that follow redirects to attacker-controlled hosts

**Attack vector**:
  - A malicious HTTP server returns a response with many chained content encodings. When the client attempts to decode, it creates a deeply nested decompression chain consuming excessive resources.

### Patches

Fixed in version 0.14.0.

  The fix adds a hardcoded limit of **5** `Content-Encoding` layers. Responses exceeding this limit raise `DecodingError`.

### Workarounds

Add middleware that checks for a malicious Content-Encoding header.

```python
from typing import cast

from zapros import (
    AsyncBaseHandler,
    AsyncBaseMiddleware,
    BaseHandler,
    BaseMiddleware,
    Client,
    DecodingError,
    Request,
    Response,
)

MAX_DECODE_LAYERS = 5


class ContentEncodingCheckMiddleware(BaseMiddleware, AsyncBaseMiddleware):
    def __init__(
        self,
        next_handler: BaseHandler | AsyncBaseHandler,
        *,
        max_layers: int = MAX_DECODE_LAYERS,
    ) -&gt; None:
        self.next = cast(BaseHandler, next_handler)
        self.async_next = cast(AsyncBaseHandler, next_handler)
        self._max_layers = max_layers

    def _check(self, response: Response) -&gt; None:
        encoding_header = response.headers.get("Content-Encoding")
        if not encoding_header:
            return

        layers = [enc.strip().lower() for enc in encoding_header.split(",") if enc.strip()]
        if len(layers) &gt; self._max_layers:
            raise DecodingError(f"Too many Content-Encoding layers ({len(layers)}), maximum is {self._max_layers}")

    def handle(self, request: Request) -&gt; Response:
        response = self.next.handle(request)
        self._check(response)
        return response

    async def ahandle(self, request: Request) -&gt; Response:
        response = await self.async_next.ahandle(request)
        self._check(response)
        return response


with Client().wrap_with_middleware(lambda next: ContentEncodingCheckMiddleware(next)) as client:
    ...
```</description>
      <content:encoded>### Impact

**Who is impacted**:
  - Any application using Zapros to make HTTP requests to untrusted servers
  - Applications that follow redirects to attacker-controlled hosts

**Attack vector**:
  - A malicious HTTP server returns a response with many chained content encodings. When the client attempts to decode, it creates a deeply nested decompression chain consuming excessive resources.

### Patches

Fixed in version 0.14.0.

  The fix adds a hardcoded limit of **5** `Content-Encoding` layers. Responses exceeding this limit raise `DecodingError`.

### Workarounds

Add middleware that checks for a malicious Content-Encoding header.

```python
from typing import cast

from zapros import (
    AsyncBaseHandler,
    AsyncBaseMiddleware,
    BaseHandler,
    BaseMiddleware,
    Client,
    DecodingError,
    Request,
    Response,
)

MAX_DECODE_LAYERS = 5


class ContentEncodingCheckMiddleware(BaseMiddleware, AsyncBaseMiddleware):
    def __init__(
        self,
        next_handler: BaseHandler | AsyncBaseHandler,
        *,
        max_layers: int = MAX_DECODE_LAYERS,
    ) -&gt; None:
        self.next = cast(BaseHandler, next_handler)
        self.async_next = cast(AsyncBaseHandler, next_handler)
        self._max_layers = max_layers

    def _check(self, response: Response) -&gt; None:
        encoding_header = response.headers.get("Content-Encoding")
        if not encoding_header:
            return

        layers = [enc.strip().lower() for enc in encoding_header.split(",") if enc.strip()]
        if len(layers) &gt; self._max_layers:
            raise DecodingError(f"Too many Content-Encoding layers ({len(layers)}), maximum is {self._max_layers}")

    def handle(self, request: Request) -&gt; Response:
        response = self.next.handle(request)
        self._check(response)
        return response

    async def ahandle(self, request: Request) -&gt; Response:
        response = await self.async_next.ahandle(request)
        self._check(response)
        return response


with Client().wrap_with_middleware(lambda next: ContentEncodingCheckMiddleware(next)) as client:
    ...
```</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2026-4181</guid>
      <pubDate>Thu, 01 Oct 2026 16:38:38 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2026-4180</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2026-4180</link>
      <description>### Impact

wlc could send an unscoped API token to an unintended server when run inside a directory tree containing attacker-controlled project configuration.

If `.weblate`, `.weblate.ini`, or `weblate.ini` defines an API url, and the user supplies a token with `WLC_KEY` or `--key` without also pinning the URL, wlc would send the token to the project-configured URL.

Impacted users are those running wlc in untrusted repositories, pull request checkouts, or directories with untrusted ancestor configuration while using `WLC_KEY` or `--key`.

### Patches

The issue is patched in wlc 2.0.1 via https://github.com/WeblateOrg/wlc/pull/1500.

The fix rejects unscoped keys when the API URL comes from automatically discovered project configuration:

- `WLC_KEY` now requires `WLC_URL`.
- `--key` now requires `--url`.
- URL-scoped keys in the `[keys]` configuration section remain supported.

Users should upgrade to wlc 2.0.1 or newer.

### Workarounds

Without upgrading, users can avoid the issue by explicitly pinning the API URL whenever using an unscoped key:

`WLC_URL=https://hosted.weblate.org/api/ WLC_KEY=... wlc ...`

or:

`wlc --url https://hosted.weblate.org/api/ --key ... ...`

Alternatively, use URL-scoped keys in the [keys] section instead of WLC_KEY or --key, and avoid running wlc with secrets in untrusted checkouts.

- The issue was independently reported by [type5afe](https://hackerone.com/type5afe) and [visionx7](https://hackerone.com/visionx7) using HackerOne.</description>
      <content:encoded>### Impact

wlc could send an unscoped API token to an unintended server when run inside a directory tree containing attacker-controlled project configuration.

If `.weblate`, `.weblate.ini`, or `weblate.ini` defines an API url, and the user supplies a token with `WLC_KEY` or `--key` without also pinning the URL, wlc would send the token to the project-configured URL.

Impacted users are those running wlc in untrusted repositories, pull request checkouts, or directories with untrusted ancestor configuration while using `WLC_KEY` or `--key`.

### Patches

The issue is patched in wlc 2.0.1 via https://github.com/WeblateOrg/wlc/pull/1500.

The fix rejects unscoped keys when the API URL comes from automatically discovered project configuration:

- `WLC_KEY` now requires `WLC_URL`.
- `--key` now requires `--url`.
- URL-scoped keys in the `[keys]` configuration section remain supported.

Users should upgrade to wlc 2.0.1 or newer.

### Workarounds

Without upgrading, users can avoid the issue by explicitly pinning the API URL whenever using an unscoped key:

`WLC_URL=https://hosted.weblate.org/api/ WLC_KEY=... wlc ...`

or:

`wlc --url https://hosted.weblate.org/api/ --key ... ...`

Alternatively, use URL-scoped keys in the [keys] section instead of WLC_KEY or --key, and avoid running wlc with secrets in untrusted checkouts.

- The issue was independently reported by [type5afe](https://hackerone.com/type5afe) and [visionx7](https://hackerone.com/visionx7) using HackerOne.</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2026-4180</guid>
      <pubDate>Thu, 01 Oct 2026 16:38:34 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2026-4179</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2026-4179</link>
      <description>### Summary
The audio decode-duration guard (`max_duration_s`, env `VLLM_MAX_AUDIO_DECODE_DURATION_S`, default 600s) that protects against audio decompression-bomb DoS is wired into **only** the speech-to-text path (`/v1/audio/transcriptions`). The **chat** audio path (`/v1/chat/completions`, `input_audio` content parts) calls the same decoder with **no** limit, so an **unauthenticated** client can submit a few-KB compressed audio file that expands to multiple GB of float32 PCM at decode time, OOM-killing the worker. This is a distinct sibling of **CVE-2026-5497** (video frame-count bomb, `VideoMediaIO.load_base64`) and **GHSA-pq5c-rjhq-qp7p** (image) in the same media subsystem.

Verified against `main` at HEAD `d78650c` (2026-06-16); applicable to the latest release v0.23.0.

### Details
The guard rejects long audio *during* decode (before allocation), implemented in `vllm/multimodal/media/audio.py`:
- `load_audio_pyav` — metadata reject (~82-98) and live sample-count reject (~129-136)
- `load_audio_soundfile` — frames reject (~165-174)

All are gated on `if max_duration_s is not None`.

It is passed in exactly **one** place — the transcription serving layer:
```python
# .../speech_to_text/base/serving.py:~170-174
load_audio(buf, sr=..., max_duration_s=self.max_audio_decode_duration_s)
#   self.max_audio_decode_duration_s = envs.VLLM_MAX_AUDIO_DECODE_DURATION_S  (default 600)
```

The chat path never threads it:
```python
# vllm/multimodal/media/audio.py:237-238
def load_bytes(self, data: bytes) -&gt; tuple[npt.NDArray, float]:
    return load_audio(BytesIO(data), sr=None)   # no max_duration_s -&gt; every guard above is skipped
```

Unauthenticated reachability chain (chat):
`parse_input_audio` (`chat_utils.py`) -&gt; `parse_audio` -&gt; `connector.fetch_audio` -&gt; `AudioMediaIO._load_data_url` -&gt; `load_base64` -&gt; `load_bytes` -&gt; `load_audio(..., sr=None)`. The connector never passes `max_duration_s`, and inline `data:` URLs need no HTTP fetch (so `VLLM_AUDIO_FETCH_TIMEOUT` does not bound them). The OpenAI-compatible server has no auth by default (auth only when `--api-key` / `VLLM_API_KEY` is set).

### Impact
Unauthenticated remote denial of service (availability) via memory amplification on a default-no-auth endpoint, on any deployment serving an audio-capable model. Same class and impact as the sibling CVE-2026-5497 (video). CWE-770 / CWE-409.

### Fix
A fix was introduced in this MR: https://github.com/vllm-project/vllm/pull/45908</description>
      <content:encoded>### Summary
The audio decode-duration guard (`max_duration_s`, env `VLLM_MAX_AUDIO_DECODE_DURATION_S`, default 600s) that protects against audio decompression-bomb DoS is wired into **only** the speech-to-text path (`/v1/audio/transcriptions`). The **chat** audio path (`/v1/chat/completions`, `input_audio` content parts) calls the same decoder with **no** limit, so an **unauthenticated** client can submit a few-KB compressed audio file that expands to multiple GB of float32 PCM at decode time, OOM-killing the worker. This is a distinct sibling of **CVE-2026-5497** (video frame-count bomb, `VideoMediaIO.load_base64`) and **GHSA-pq5c-rjhq-qp7p** (image) in the same media subsystem.

Verified against `main` at HEAD `d78650c` (2026-06-16); applicable to the latest release v0.23.0.

### Details
The guard rejects long audio *during* decode (before allocation), implemented in `vllm/multimodal/media/audio.py`:
- `load_audio_pyav` — metadata reject (~82-98) and live sample-count reject (~129-136)
- `load_audio_soundfile` — frames reject (~165-174)

All are gated on `if max_duration_s is not None`.

It is passed in exactly **one** place — the transcription serving layer:
```python
# .../speech_to_text/base/serving.py:~170-174
load_audio(buf, sr=..., max_duration_s=self.max_audio_decode_duration_s)
#   self.max_audio_decode_duration_s = envs.VLLM_MAX_AUDIO_DECODE_DURATION_S  (default 600)
```

The chat path never threads it:
```python
# vllm/multimodal/media/audio.py:237-238
def load_bytes(self, data: bytes) -&gt; tuple[npt.NDArray, float]:
    return load_audio(BytesIO(data), sr=None)   # no max_duration_s -&gt; every guard above is skipped
```

Unauthenticated reachability chain (chat):
`parse_input_audio` (`chat_utils.py`) -&gt; `parse_audio` -&gt; `connector.fetch_audio` -&gt; `AudioMediaIO._load_data_url` -&gt; `load_base64` -&gt; `load_bytes` -&gt; `load_audio(..., sr=None)`. The connector never passes `max_duration_s`, and inline `data:` URLs need no HTTP fetch (so `VLLM_AUDIO_FETCH_TIMEOUT` does not bound them). The OpenAI-compatible server has no auth by default (auth only when `--api-key` / `VLLM_API_KEY` is set).

### Impact
Unauthenticated remote denial of service (availability) via memory amplification on a default-no-auth endpoint, on any deployment serving an audio-capable model. Same class and impact as the sibling CVE-2026-5497 (video). CWE-770 / CWE-409.

### Fix
A fix was introduced in this MR: https://github.com/vllm-project/vllm/pull/45908</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2026-4179</guid>
      <pubDate>Thu, 01 Oct 2026 16:38:32 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2026-4178</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2026-4178</link>
      <description>## Summary

Current vLLM `main` lets an inference request choose the PyNvVideoCodec GPU video decoder through `media_io_kwargs.video.video_backend`, but engine GPU memory reservation is computed only from static startup configuration and `VLLM_VIDEO_LOADER_BACKEND`. If the server starts with the default OpenCV/software backend and no `--mm-ipc-gpu-memory-gb` budget, a client can still route a video request into the PyNvVideoCodec path after startup, causing frontend CUDA-context, decoder-surface, and decoded-frame GPU allocations that were not carved out of the engine KV-cache budget.

## Technical Details

The vulnerable boundary is the split between request-time media decoding choices in the API server and startup-time memory budgeting in the engine worker. Request bodies for Chat Completions and Responses expose `media_io_kwargs`, and those values are forwarded to the shared media connector. For video inputs, `MediaConnector.fetch_video()` copies `self.media_io_kwargs["video"]` into `video_io_kwargs`, only setting a model-derived backend when `video_backend` is absent. `VideoMediaIO.__init__()` then consumes `video_backend` from those kwargs and loads that backend from `VIDEO_LOADER_REGISTRY`.

The relevant request-side source path is:

```python
video_io_kwargs = dict(self.media_io_kwargs.get("video", {}))
if "video_backend" not in video_io_kwargs and (
    video_backend := get_video_loader_backend_for_processor(video_processor)
):
    video_io_kwargs["video_backend"] = video_backend
video_io = VideoMediaIO(image_io, **video_io_kwargs)
```

```python
video_loader_backend = (
    kwargs.pop("video_backend", None) or envs.VLLM_VIDEO_LOADER_BACKEND
)
self.video_loader = VIDEO_LOADER_REGISTRY.load(video_loader_backend)
```

`VideoBackend.load_bytes()` then dispatches `backend == "pynvvideocodec"` into `decode_frames_pynvvideocodec()`, which constructs a PyNvVideoCodec decoder, creates or uses a CUDA stream, reads stream metadata, decodes selected frames on the GPU, and copies those frames into pinned host memory. The new frontend GPU memory pool accounts only for raw decoded frame bytes when a pool exists; it does not make request-time backend selection safe when no startup reservation was made.

The engine-side reservation code makes its decision from static model config and environment only:

```python
def _uses_pynvvideocodec_video_backend(mm_config) -&gt; bool:
    video_kwargs = mm_config.media_io_kwargs.get("video", {})
    video_loader_backend = (
        video_kwargs.get("video_backend") or envs.VLLM_VIDEO_LOADER_BACKEND
    )
    codec_backend = video_kwargs.get("backend")
    return (
        video_loader_backend == PYNVVIDEOCODEC_VIDEO_BACKEND
        or codec_backend == PYNVVIDEOCODEC_VIDEO_BACKEND
    )
```

```python
decoder_reserved_bytes = (
    num_api_servers * per_server_decoder_bytes
    if self._uses_pynvvideocodec_video_backend(mm_config)
    else 0
)
reserved_bytes = raw_frame_reserved_bytes + decoder_reserved_bytes
if reserved_bytes &lt;= 0:
    return available_kv_cache_memory_bytes
```

With default static video configuration, `mm_config.media_io_kwargs["video"]` does not name PyNvVideoCodec and `VLLM_VIDEO_LOADER_BACKEND` defaults to OpenCV/software decoding. The worker therefore reserves no PyNv decoder/CUDA-context bytes. A later request can still set `media_io_kwargs.video.video_backend="pynvvideocodec"` and reach the GPU decoder path because that runtime field is intentionally honored by `VideoMediaIO`.

## PoV

An ordinary multimodal inference request can carry the backend override in the request body:

```json
{
  "model": "served-vlm",
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": "summarize this clip"},
        {"type": "video_url", "video_url": {"url": "data:video/mp4;base64,&lt;small-mp4&gt;"}}
      ]
    }
  ],
  "media_io_kwargs": {
    "video": {
      "video_backend": "pynvvideocodec"
    }
  }
}
```

The following bounded source-level check confirms the code path without allocating GPU memory:

```bash
git clone --filter=blob:none https://github.com/vllm-project/vllm.git
cd vllm
git checkout ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a
python3 check_pynv_backend_reservation.py --repo .
```

## PoC

The bounded check validates current source markers, simulates the exact static reservation predicate, and compares vulnerable and negative-control configurations. Key output:

```json
{
  "vulnerable": true,
  "head": "ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a",
  "reservation_simulation": {
    "env_video_loader_backend": "opencv",
    "request_selects_pynv_after_startup": true,
    "vulnerable_static_reserved_bytes": 0,
    "negative_control_static_pynv_reserved_bytes": 2066953011,
    "raw_frame_only_control_reserved_bytes": 268435456,
    "unreserved_decoder_bytes_when_only_request_selects_pynv": 2066953011
  }
}
```

The negative control is important: when PyNvVideoCodec is selected statically, the worker reserves `2066953011` bytes per API process for decoder surfaces plus CUDA context. The vulnerable case reserves `0` bytes for the same decoder overhead because PyNvVideoCodec is selected only by the later request. A second control with static OpenCV plus `mm_ipc_gpu_memory_gb=0.25` reserves only the raw-frame semaphore budget and still does not reserve PyNv decoder/CUDA-context bytes.

## Impact

An attacker who can submit video requests to a vLLM deployment with PyNvVideoCodec available can force frontend GPU decoding even when the engine did not reserve memory for that decoder during startup. On high-utilization serving deployments, the unreserved CUDA context, retained decoder surfaces, and decoded-frame allocations can reduce or exhaust GPU memory that the engine assumed was available for weights, activations, or KV cache, causing request failures, worker crashes, or service-level denial of service.

Suggested severity is Medium with conservative CVSS v3.1 `CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H` (6.5). If a deployment exposes the affected API without authentication, `PR:N` would raise the deployment-specific score. Suggested weaknesses are `CWE-770` (Allocation of Resources Without Limits or Throttling) and `CWE-400` (Uncontrolled Resource Consumption). This should not be rated Low because the affected resource is shared GPU memory in the serving path and the code already treats the PyNv decoder/CUDA-context footprint as large enough to reserve at startup when statically configured.

Limitations: exploitation requires a GPU deployment where PyNvVideoCodec is installed and usable, and the request must reach a video-capable model/path. The issue does not claim code execution, data disclosure, or SSRF.

## Suggested Fix

Do not allow untrusted request fields to select a GPU decoder that was not included in startup memory reservation. The simplest fix is to reject request-level `media_io_kwargs.video.video_backend="pynvvideocodec"` unless the static server configuration already selected PyNvVideoCodec and reserved its decoder/CUDA-context budget.

If dynamic backend selection remains supported, split software and GPU decoder policies: allow request selection among CPU/software decoders only, require an explicit operator allowlist for GPU decoders, and include every request-selectable GPU decoder in the startup reservation predicate. Add regression coverage for static OpenCV startup config plus request-level PyNvVideoCodec override, and preserve the negative control where static PyNvVideoCodec configuration reserves decoder/CUDA-context bytes.

## Affected Package/Versions

Package: `vllm` from `vllm-project/vllm`.

Confirmed affected: current `main` at `ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a`.

Introduced by: `af16446bf39de047ab57649c933063cf1cbf1e50`, `Vram semaphore infra (#44465)`, committed 2026-06-26T17:32:51-07:00.

Release status checked: `git tag --contains af16446bf` returned no release tags in the fresh checkout. GitHub repository metadata reported latest published release `v0.23.0` published 2026-06-15T05:27:20Z; the local `v0.24.0` tag also does not contain the introducing commit. The affected range should therefore be current `main` builds containing `af16446bf` until fixed, rather than a confirmed released-version range.

## Advisory History

Public vLLM advisories checked included audio decompression-bomb DoS, unbounded `video/jpeg` frame-count DoS, MediaConnector SSRF, video processing RCE, multimodal embedding DoS/RCE, GGUF GPU memory exposure, multimodal hashing, and other request-parameter DoS classes. None matched request-selected PyNvVideoCodec or the static VRAM reservation mismatch.

Prior local/private vLLM report families checked included request-level `media_io_kwargs` reopening `video/jpeg` frame fanout, GLM video metadata amplification, and audio media decode duration-limit bypass. Those reports share the request-level media kwargs boundary, but they target CPU/media decode limits or model metadata amplification. This report targets a different privileged asset and fix surface: GPU decoder selection after engine startup memory reservation.

Focused GitHub issue/PR searches for `pynvvideocodec`, `mm_ipc_gpu_memory`, `video_backend media_io_kwargs`, `Vram semaphore infra`, and `frontend multimodal GPU decoding` found the PyNvVideoCodec zero-copy RFC, an old do-not-review prototype, merged PR `#44465`, and an unrelated TorchCodec backend PR. No public issue or PR described this security boundary.

## Appendix: Bounded Source-Level Check

```python
#!/usr/bin/env python3
from __future__ import annotations

import argparse
import json
import re
import subprocess
from pathlib import Path

MIB = 1024 * 1024
GIB = 1024 * MIB

def read(repo: Path, rel: str) -&gt; str:
    return (repo / rel).read_text(encoding="utf-8")

def const_int(source: str, name: str) -&gt; int:
    expr = re.search(rf"^{name}\s*=\s*(.+)$", source, flags=re.MULTILINE).group(1).strip()
    if expr == "128 * MiB_bytes":
        return 128 * MIB
    if expr == "int(1.8 * 1024 * MiB_bytes)":
        return int(1.8 * 1024 * MIB)
    if expr == "1":
        return 1
    raise AssertionError(expr)

def uses_pynv_static(static_media_io_kwargs: dict[str, dict[str, str]], env_backend: str) -&gt; bool:
    video_kwargs = static_media_io_kwargs.get("video", {})
    video_loader_backend = video_kwargs.get("video_backend") or env_backend
    codec_backend = video_kwargs.get("backend")
    return video_loader_backend == "pynvvideocodec" or codec_backend == "pynvvideocodec"

def reserve_bytes(static_media_io_kwargs, env_backend, mm_ipc_gpu_memory_gb, decoder_bytes, cuda_context_bytes, retained_decoders):
    raw_frame_reserved_bytes = int(mm_ipc_gpu_memory_gb * GIB)
    per_server_decoder_bytes = decoder_bytes * retained_decoders + cuda_context_bytes
    decoder_reserved_bytes = per_server_decoder_bytes if uses_pynv_static(static_media_io_kwargs, env_backend) else 0
    return raw_frame_reserved_bytes + decoder_reserved_bytes

parser = argparse.ArgumentParser()
parser.add_argument("--repo", required=True, type=Path)
repo = parser.parse_args().repo.resolve()

media_video = read(repo, "vllm/multimodal/media/video.py")
connector = read(repo, "vllm/multimodal/media/connector.py")
chat_protocol = read(repo, "vllm/entrypoints/openai/chat_completion/protocol.py")
responses_protocol = read(repo, "vllm/entrypoints/openai/responses/protocol.py")
gpu_worker = read(repo, "vllm/v1/worker/gpu_worker.py")
video_core = read(repo, "vllm/multimodal/video.py")

assert "media_io_kwargs: dict[str, dict[str, Any]] | None = Field(" in chat_protocol
assert "media_io_kwargs: dict[str, dict[str, Any]] | None = Field(" in responses_protocol
assert 'video_io_kwargs = dict(self.media_io_kwargs.get("video", {}))' in connector
assert 'if "video_backend" not in video_io_kwargs and (' in connector
assert 'kwargs.pop("video_backend", None) or envs.VLLM_VIDEO_LOADER_BACKEND' in media_video
assert "elif backend == PYNVVIDEOCODEC_VIDEO_BACKEND:" in video_core
assert 'video_kwargs = mm_config.media_io_kwargs.get("video", {})' in gpu_worker

decoder_bytes = const_int(video_core, "PYNVVIDEOCODEC_DECODER_GPU_MEMORY_BYTES")
retained_decoders = const_int(video_core, "PYNVVIDEOCODEC_MAX_RETAINED_DECODERS")
cuda_context_bytes = const_int(video_core, "PYNVVIDEOCODEC_CUDA_CONTEXT_BYTES")
per_server_decoder_bytes = decoder_bytes * retained_decoders + cuda_context_bytes

vulnerable_static_reserved = reserve_bytes({}, "opencv", 0.0, decoder_bytes, cuda_context_bytes, retained_decoders)
negative_control_reserved = reserve_bytes({"video": {"video_backend": "pynvvideocodec"}}, "opencv", 0.0, decoder_bytes, cuda_context_bytes, retained_decoders)
raw_frame_only_control = reserve_bytes({}, "opencv", 0.25, decoder_bytes, cuda_context_bytes, retained_decoders)

head = subprocess.check_output(["git", "-C", str(repo), "rev-parse", "HEAD"], text=True).strip()
print(json.dumps({
    "head": head,
    "vulnerable": vulnerable_static_reserved == 0 and negative_control_reserved == per_server_decoder_bytes,
    "reservation_simulation": {
        "env_video_loader_backend": "opencv",
        "request_selects_pynv_after_startup": True,
        "vulnerable_static_reserved_bytes": vulnerable_static_reserved,
        "negative_control_static_pynv_reserved_bytes": negative_control_reserved,
        "raw_frame_only_control_reserved_bytes": raw_frame_only_control,
        "unreserved_decoder_bytes_when_only_request_selects_pynv": per_server_decoder_bytes,
    },
}, indent=2, sort_keys=True))
```</description>
      <content:encoded>## Summary

Current vLLM `main` lets an inference request choose the PyNvVideoCodec GPU video decoder through `media_io_kwargs.video.video_backend`, but engine GPU memory reservation is computed only from static startup configuration and `VLLM_VIDEO_LOADER_BACKEND`. If the server starts with the default OpenCV/software backend and no `--mm-ipc-gpu-memory-gb` budget, a client can still route a video request into the PyNvVideoCodec path after startup, causing frontend CUDA-context, decoder-surface, and decoded-frame GPU allocations that were not carved out of the engine KV-cache budget.

## Technical Details

The vulnerable boundary is the split between request-time media decoding choices in the API server and startup-time memory budgeting in the engine worker. Request bodies for Chat Completions and Responses expose `media_io_kwargs`, and those values are forwarded to the shared media connector. For video inputs, `MediaConnector.fetch_video()` copies `self.media_io_kwargs["video"]` into `video_io_kwargs`, only setting a model-derived backend when `video_backend` is absent. `VideoMediaIO.__init__()` then consumes `video_backend` from those kwargs and loads that backend from `VIDEO_LOADER_REGISTRY`.

The relevant request-side source path is:

```python
video_io_kwargs = dict(self.media_io_kwargs.get("video", {}))
if "video_backend" not in video_io_kwargs and (
    video_backend := get_video_loader_backend_for_processor(video_processor)
):
    video_io_kwargs["video_backend"] = video_backend
video_io = VideoMediaIO(image_io, **video_io_kwargs)
```

```python
video_loader_backend = (
    kwargs.pop("video_backend", None) or envs.VLLM_VIDEO_LOADER_BACKEND
)
self.video_loader = VIDEO_LOADER_REGISTRY.load(video_loader_backend)
```

`VideoBackend.load_bytes()` then dispatches `backend == "pynvvideocodec"` into `decode_frames_pynvvideocodec()`, which constructs a PyNvVideoCodec decoder, creates or uses a CUDA stream, reads stream metadata, decodes selected frames on the GPU, and copies those frames into pinned host memory. The new frontend GPU memory pool accounts only for raw decoded frame bytes when a pool exists; it does not make request-time backend selection safe when no startup reservation was made.

The engine-side reservation code makes its decision from static model config and environment only:

```python
def _uses_pynvvideocodec_video_backend(mm_config) -&gt; bool:
    video_kwargs = mm_config.media_io_kwargs.get("video", {})
    video_loader_backend = (
        video_kwargs.get("video_backend") or envs.VLLM_VIDEO_LOADER_BACKEND
    )
    codec_backend = video_kwargs.get("backend")
    return (
        video_loader_backend == PYNVVIDEOCODEC_VIDEO_BACKEND
        or codec_backend == PYNVVIDEOCODEC_VIDEO_BACKEND
    )
```

```python
decoder_reserved_bytes = (
    num_api_servers * per_server_decoder_bytes
    if self._uses_pynvvideocodec_video_backend(mm_config)
    else 0
)
reserved_bytes = raw_frame_reserved_bytes + decoder_reserved_bytes
if reserved_bytes &lt;= 0:
    return available_kv_cache_memory_bytes
```

With default static video configuration, `mm_config.media_io_kwargs["video"]` does not name PyNvVideoCodec and `VLLM_VIDEO_LOADER_BACKEND` defaults to OpenCV/software decoding. The worker therefore reserves no PyNv decoder/CUDA-context bytes. A later request can still set `media_io_kwargs.video.video_backend="pynvvideocodec"` and reach the GPU decoder path because that runtime field is intentionally honored by `VideoMediaIO`.

## PoV

An ordinary multimodal inference request can carry the backend override in the request body:

```json
{
  "model": "served-vlm",
  "messages": [
    {
      "role": "user",
      "content": [
        {"type": "text", "text": "summarize this clip"},
        {"type": "video_url", "video_url": {"url": "data:video/mp4;base64,&lt;small-mp4&gt;"}}
      ]
    }
  ],
  "media_io_kwargs": {
    "video": {
      "video_backend": "pynvvideocodec"
    }
  }
}
```

The following bounded source-level check confirms the code path without allocating GPU memory:

```bash
git clone --filter=blob:none https://github.com/vllm-project/vllm.git
cd vllm
git checkout ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a
python3 check_pynv_backend_reservation.py --repo .
```

## PoC

The bounded check validates current source markers, simulates the exact static reservation predicate, and compares vulnerable and negative-control configurations. Key output:

```json
{
  "vulnerable": true,
  "head": "ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a",
  "reservation_simulation": {
    "env_video_loader_backend": "opencv",
    "request_selects_pynv_after_startup": true,
    "vulnerable_static_reserved_bytes": 0,
    "negative_control_static_pynv_reserved_bytes": 2066953011,
    "raw_frame_only_control_reserved_bytes": 268435456,
    "unreserved_decoder_bytes_when_only_request_selects_pynv": 2066953011
  }
}
```

The negative control is important: when PyNvVideoCodec is selected statically, the worker reserves `2066953011` bytes per API process for decoder surfaces plus CUDA context. The vulnerable case reserves `0` bytes for the same decoder overhead because PyNvVideoCodec is selected only by the later request. A second control with static OpenCV plus `mm_ipc_gpu_memory_gb=0.25` reserves only the raw-frame semaphore budget and still does not reserve PyNv decoder/CUDA-context bytes.

## Impact

An attacker who can submit video requests to a vLLM deployment with PyNvVideoCodec available can force frontend GPU decoding even when the engine did not reserve memory for that decoder during startup. On high-utilization serving deployments, the unreserved CUDA context, retained decoder surfaces, and decoded-frame allocations can reduce or exhaust GPU memory that the engine assumed was available for weights, activations, or KV cache, causing request failures, worker crashes, or service-level denial of service.

Suggested severity is Medium with conservative CVSS v3.1 `CVSS:3.1/AV:N/AC:L/PR:L/UI:N/S:U/C:N/I:N/A:H` (6.5). If a deployment exposes the affected API without authentication, `PR:N` would raise the deployment-specific score. Suggested weaknesses are `CWE-770` (Allocation of Resources Without Limits or Throttling) and `CWE-400` (Uncontrolled Resource Consumption). This should not be rated Low because the affected resource is shared GPU memory in the serving path and the code already treats the PyNv decoder/CUDA-context footprint as large enough to reserve at startup when statically configured.

Limitations: exploitation requires a GPU deployment where PyNvVideoCodec is installed and usable, and the request must reach a video-capable model/path. The issue does not claim code execution, data disclosure, or SSRF.

## Suggested Fix

Do not allow untrusted request fields to select a GPU decoder that was not included in startup memory reservation. The simplest fix is to reject request-level `media_io_kwargs.video.video_backend="pynvvideocodec"` unless the static server configuration already selected PyNvVideoCodec and reserved its decoder/CUDA-context budget.

If dynamic backend selection remains supported, split software and GPU decoder policies: allow request selection among CPU/software decoders only, require an explicit operator allowlist for GPU decoders, and include every request-selectable GPU decoder in the startup reservation predicate. Add regression coverage for static OpenCV startup config plus request-level PyNvVideoCodec override, and preserve the negative control where static PyNvVideoCodec configuration reserves decoder/CUDA-context bytes.

## Affected Package/Versions

Package: `vllm` from `vllm-project/vllm`.

Confirmed affected: current `main` at `ddd3855a28a561a5bb54d380c6e6b8b1e883cc4a`.

Introduced by: `af16446bf39de047ab57649c933063cf1cbf1e50`, `Vram semaphore infra (#44465)`, committed 2026-06-26T17:32:51-07:00.

Release status checked: `git tag --contains af16446bf` returned no release tags in the fresh checkout. GitHub repository metadata reported latest published release `v0.23.0` published 2026-06-15T05:27:20Z; the local `v0.24.0` tag also does not contain the introducing commit. The affected range should therefore be current `main` builds containing `af16446bf` until fixed, rather than a confirmed released-version range.

## Advisory History

Public vLLM advisories checked included audio decompression-bomb DoS, unbounded `video/jpeg` frame-count DoS, MediaConnector SSRF, video processing RCE, multimodal embedding DoS/RCE, GGUF GPU memory exposure, multimodal hashing, and other request-parameter DoS classes. None matched request-selected PyNvVideoCodec or the static VRAM reservation mismatch.

Prior local/private vLLM report families checked included request-level `media_io_kwargs` reopening `video/jpeg` frame fanout, GLM video metadata amplification, and audio media decode duration-limit bypass. Those reports share the request-level media kwargs boundary, but they target CPU/media decode limits or model metadata amplification. This report targets a different privileged asset and fix surface: GPU decoder selection after engine startup memory reservation.

Focused GitHub issue/PR searches for `pynvvideocodec`, `mm_ipc_gpu_memory`, `video_backend media_io_kwargs`, `Vram semaphore infra`, and `frontend multimodal GPU decoding` found the PyNvVideoCodec zero-copy RFC, an old do-not-review prototype, merged PR `#44465`, and an unrelated TorchCodec backend PR. No public issue or PR described this security boundary.

## Appendix: Bounded Source-Level Check

```python
#!/usr/bin/env python3
from __future__ import annotations

import argparse
import json
import re
import subprocess
from pathlib import Path

MIB = 1024 * 1024
GIB = 1024 * MIB

def read(repo: Path, rel: str) -&gt; str:
    return (repo / rel).read_text(encoding="utf-8")

def const_int(source: str, name: str) -&gt; int:
    expr = re.search(rf"^{name}\s*=\s*(.+)$", source, flags=re.MULTILINE).group(1).strip()
    if expr == "128 * MiB_bytes":
        return 128 * MIB
    if expr == "int(1.8 * 1024 * MiB_bytes)":
        return int(1.8 * 1024 * MIB)
    if expr == "1":
        return 1
    raise AssertionError(expr)

def uses_pynv_static(static_media_io_kwargs: dict[str, dict[str, str]], env_backend: str) -&gt; bool:
    video_kwargs = static_media_io_kwargs.get("video", {})
    video_loader_backend = video_kwargs.get("video_backend") or env_backend
    codec_backend = video_kwargs.get("backend")
    return video_loader_backend == "pynvvideocodec" or codec_backend == "pynvvideocodec"

def reserve_bytes(static_media_io_kwargs, env_backend, mm_ipc_gpu_memory_gb, decoder_bytes, cuda_context_bytes, retained_decoders):
    raw_frame_reserved_bytes = int(mm_ipc_gpu_memory_gb * GIB)
    per_server_decoder_bytes = decoder_bytes * retained_decoders + cuda_context_bytes
    decoder_reserved_bytes = per_server_decoder_bytes if uses_pynv_static(static_media_io_kwargs, env_backend) else 0
    return raw_frame_reserved_bytes + decoder_reserved_bytes

parser = argparse.ArgumentParser()
parser.add_argument("--repo", required=True, type=Path)
repo = parser.parse_args().repo.resolve()

media_video = read(repo, "vllm/multimodal/media/video.py")
connector = read(repo, "vllm/multimodal/media/connector.py")
chat_protocol = read(repo, "vllm/entrypoints/openai/chat_completion/protocol.py")
responses_protocol = read(repo, "vllm/entrypoints/openai/responses/protocol.py")
gpu_worker = read(repo, "vllm/v1/worker/gpu_worker.py")
video_core = read(repo, "vllm/multimodal/video.py")

assert "media_io_kwargs: dict[str, dict[str, Any]] | None = Field(" in chat_protocol
assert "media_io_kwargs: dict[str, dict[str, Any]] | None = Field(" in responses_protocol
assert 'video_io_kwargs = dict(self.media_io_kwargs.get("video", {}))' in connector
assert 'if "video_backend" not in video_io_kwargs and (' in connector
assert 'kwargs.pop("video_backend", None) or envs.VLLM_VIDEO_LOADER_BACKEND' in media_video
assert "elif backend == PYNVVIDEOCODEC_VIDEO_BACKEND:" in video_core
assert 'video_kwargs = mm_config.media_io_kwargs.get("video", {})' in gpu_worker

decoder_bytes = const_int(video_core, "PYNVVIDEOCODEC_DECODER_GPU_MEMORY_BYTES")
retained_decoders = const_int(video_core, "PYNVVIDEOCODEC_MAX_RETAINED_DECODERS")
cuda_context_bytes = const_int(video_core, "PYNVVIDEOCODEC_CUDA_CONTEXT_BYTES")
per_server_decoder_bytes = decoder_bytes * retained_decoders + cuda_context_bytes

vulnerable_static_reserved = reserve_bytes({}, "opencv", 0.0, decoder_bytes, cuda_context_bytes, retained_decoders)
negative_control_reserved = reserve_bytes({"video": {"video_backend": "pynvvideocodec"}}, "opencv", 0.0, decoder_bytes, cuda_context_bytes, retained_decoders)
raw_frame_only_control = reserve_bytes({}, "opencv", 0.25, decoder_bytes, cuda_context_bytes, retained_decoders)

head = subprocess.check_output(["git", "-C", str(repo), "rev-parse", "HEAD"], text=True).strip()
print(json.dumps({
    "head": head,
    "vulnerable": vulnerable_static_reserved == 0 and negative_control_reserved == per_server_decoder_bytes,
    "reservation_simulation": {
        "env_video_loader_backend": "opencv",
        "request_selects_pynv_after_startup": True,
        "vulnerable_static_reserved_bytes": vulnerable_static_reserved,
        "negative_control_static_pynv_reserved_bytes": negative_control_reserved,
        "raw_frame_only_control_reserved_bytes": raw_frame_only_control,
        "unreserved_decoder_bytes_when_only_request_selects_pynv": per_server_decoder_bytes,
    },
}, indent=2, sort_keys=True))
```</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2026-4178</guid>
      <pubDate>Thu, 01 Oct 2026 16:38:32 +0000</pubDate>
    </item>
    <item>
      <title>pysec-2026-4177</title>
      <link>https://127.0.0.1:10001/vuln/pysec-2026-4177</link>
      <description>## Impact

urllib3's [streaming API](https://urllib3.readthedocs.io/en/2.7.0/advanced-usage.html#streaming-and-i-o) is designed for efficiently handling large HTTP responses by reading the content in chunks, rather than loading the entire response body into memory at once. When decoding a [chunked-transfer-encoded response](https://httpwg.org/specs/rfc9112.html#chunked.encoding), this API reads each chunk's size field by buffering until it sees `\n` or EOF.

A malicious HTTP server can return `Transfer-Encoding: chunked` and then send a very long run of bytes without any newline, causing the streaming client to buffer that entire run before urllib3 can reject the chunk size as invalid and allocate more memory than intended.

To fix the issue, we'll reject chunk-size fields larger than 65536 bytes, as already done in the non-streaming case, which is currently handled by the Python standard library. Note that this only applies to reading the chunk-size field, not the chunk data, which is already handled correctly. Thus, there shouldn't be any impact for non-malicious servers.

## Affected usages

Applications and libraries using urllib3 versions earlier than 2.8.0 may be affected when streaming a chunked response from untrusted sources. Specifically, this affects the [read_chunked()](https://urllib3.readthedocs.io/en/stable/reference/urllib3.response.html#urllib3.response.BaseHTTPResponse.read_chunked) and [stream()](https://urllib3.readthedocs.io/en/stable/reference/urllib3.response.html#urllib3.response.BaseHTTPResponse.stream) methods of the HTTPResponse object.

This also affects requests streaming API, which uses urllib3 under the hood.

## Remediation

Upgrade to urllib3 version 2.8.0 or later, where chunk-size fields larger than 65536 will be rejected. If upgrading is not immediately possible, consider reading the response at once using the [read()](https://urllib3.readthedocs.io/en/stable/reference/urllib3.response.html#urllib3.response.BaseHTTPResponse.read) method.</description>
      <content:encoded>## Impact

urllib3's [streaming API](https://urllib3.readthedocs.io/en/2.7.0/advanced-usage.html#streaming-and-i-o) is designed for efficiently handling large HTTP responses by reading the content in chunks, rather than loading the entire response body into memory at once. When decoding a [chunked-transfer-encoded response](https://httpwg.org/specs/rfc9112.html#chunked.encoding), this API reads each chunk's size field by buffering until it sees `\n` or EOF.

A malicious HTTP server can return `Transfer-Encoding: chunked` and then send a very long run of bytes without any newline, causing the streaming client to buffer that entire run before urllib3 can reject the chunk size as invalid and allocate more memory than intended.

To fix the issue, we'll reject chunk-size fields larger than 65536 bytes, as already done in the non-streaming case, which is currently handled by the Python standard library. Note that this only applies to reading the chunk-size field, not the chunk data, which is already handled correctly. Thus, there shouldn't be any impact for non-malicious servers.

## Affected usages

Applications and libraries using urllib3 versions earlier than 2.8.0 may be affected when streaming a chunked response from untrusted sources. Specifically, this affects the [read_chunked()](https://urllib3.readthedocs.io/en/stable/reference/urllib3.response.html#urllib3.response.BaseHTTPResponse.read_chunked) and [stream()](https://urllib3.readthedocs.io/en/stable/reference/urllib3.response.html#urllib3.response.BaseHTTPResponse.stream) methods of the HTTPResponse object.

This also affects requests streaming API, which uses urllib3 under the hood.

## Remediation

Upgrade to urllib3 version 2.8.0 or later, where chunk-size fields larger than 65536 will be rejected. If upgrading is not immediately possible, consider reading the response at once using the [read()](https://urllib3.readthedocs.io/en/stable/reference/urllib3.response.html#urllib3.response.BaseHTTPResponse.read) method.</content:encoded>
      <guid isPermaLink="false">https://127.0.0.1:10001/vuln/pysec-2026-4177</guid>
      <pubDate>Thu, 01 Oct 2026 16:38:41 +0000</pubDate>
    </item>
  </channel>
</rss>
