Skip to content

NVDEC cache silently mis-crops videos that share coded size but not display size #1704

Description

@YashJain14

The default nvdec backend caches CUvideodecoder objects keyed on coded width/height (plus codec / chroma / bit depth / surface format), not on CUVIDEOFORMAT.display_area.

Two H.264 videos with different display heights that NVDEC treats as the same coded size therefore share a cache entry. The second video is decoded with the first video’s crop. There is no exception, and shape / frame count stay correct. The damage is a destroyed band of rows (usually the top).

set_nvdec_cache_capacity(0) makes the pixels match a fresh-process decode (same tensor hash). set_cuda_backend("ffmpeg") is also clean (it does not use this cache).

ffprobe still reports coded_height == height on the files below (532 vs 530). The collision is inside NVDEC / CUVIDEOFORMAT, not in the container metadata.

Minimal repro

Needs one GPU, ffmpeg, and torchcodec with NVDEC. Each scenario must run in a fresh process (the cache is process-global).

ffmpeg -y -f lavfi -i testsrc2=s=1280x720:d=4 -vf scale=1280:532 \
  -c:v libx264 -pix_fmt yuv420p -an primer.mp4
ffmpeg -y -f lavfi -i testsrc2=s=1280x720:d=4 -vf scale=1280:530 \
  -c:v libx264 -pix_fmt yuv420p -an victim.mp4
import torch
from torchcodec.decoders import VideoDecoder, set_cuda_backend

def nvdec_range(path, lo, hi):
    with set_cuda_backend("nvdec"):
        dec = VideoDecoder(path, device="cuda", seek_mode="exact", dimension_order="NCHW")
    frames = dec.get_frames_in_range(lo, hi).data
    torch.cuda.synchronize()
    out = frames.cpu().contiguous()
    del dec
    return out

ref = VideoDecoder("victim.mp4", device="cpu", seek_mode="exact", dimension_order="NCHW")
ref = ref.get_frames_in_range(12, 44).data.contiguous()

_ = nvdec_range("primer.mp4", 0, 32)   # populate the cache
gpu = nvdec_range("victim.mp4", 12, 44)

ad = (gpu.to(torch.int16) - ref.to(torch.int16)).abs().float()
print(tuple(gpu.shape), "max", int(ad.max()), "mae", float(ad.mean()),
      "top4_mae", float(ad[:, :, :4, :].mean()))

What we get on 0.16.0+cu130 (fresh process per row):

scenario max Δ vs CPU MAE top-4-row MAE
victim.mp4 alone 2 0.40 0.39
primer.mp4 then victim.mp4, default cache 255 5.38 136
same order, set_nvdec_cache_capacity(0) 2 0.40 0.39
same order, set_cuda_backend("ffmpeg") 2 0.40 0.39

Shape is always (32, 3, 530, 1280). Clean NVDEC vs CPU is the usual CSC slack (max Δ 2). The collide row is not slack. cache-off and ffmpeg match the victim.mp4-alone tensor hash.

Why

In 0.16.0, NVDECCache::CacheKey sets width / height from coded_width / coded_height and does not store display_area:

https://github.com/meta-pytorch/torchcodec/blob/v0.16.0/src/torchcodec/_core/NVDECCache.h

Any mixed-resolution corpus that lands two display sizes in the same coded bucket will hit this in a long-lived process.

Related: #944 #1176

Workaround

set_nvdec_cache_capacity(0) is correct, but it forces a full NVDEC create/destroy on every video. That is cheap on one quiet GPU and expensive when many processes open decoders at once (create/destroy does not scale with GPU count on a node).

Versions

torchcodec: 0.16.0+cu130
PyTorch version: 2.12.1+cu130
Is debug build: False
CUDA used to build PyTorch: 13.0
OS: Ubuntu 24.04.4 LTS (x86_64)
Libc version: glibc-2.39
Python version: 3.11.15 (64-bit runtime)
Python platform: Linux-5.15.0-170-generic-x86_64-with-glibc2.39
Is CUDA available: True
CUDA runtime version: 13.0.88
GPU models and configuration:
  GPU 0–7: NVIDIA H200
Nvidia driver version: 595.71.05

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions