The default nvdec backend caches CUvideodecoder objects keyed on coded width/height (plus codec / chroma / bit depth / surface format), not on CUVIDEOFORMAT.display_area.
Two H.264 videos with different display heights that NVDEC treats as the same coded size therefore share a cache entry. The second video is decoded with the first video’s crop. There is no exception, and shape / frame count stay correct. The damage is a destroyed band of rows (usually the top).
set_nvdec_cache_capacity(0) makes the pixels match a fresh-process decode (same tensor hash). set_cuda_backend("ffmpeg") is also clean (it does not use this cache).
ffprobe still reports coded_height == height on the files below (532 vs 530). The collision is inside NVDEC / CUVIDEOFORMAT, not in the container metadata.
Minimal repro
Needs one GPU, ffmpeg, and torchcodec with NVDEC. Each scenario must run in a fresh process (the cache is process-global).
ffmpeg -y -f lavfi -i testsrc2=s=1280x720:d=4 -vf scale=1280:532 \
-c:v libx264 -pix_fmt yuv420p -an primer.mp4
ffmpeg -y -f lavfi -i testsrc2=s=1280x720:d=4 -vf scale=1280:530 \
-c:v libx264 -pix_fmt yuv420p -an victim.mp4
import torch
from torchcodec.decoders import VideoDecoder, set_cuda_backend
def nvdec_range(path, lo, hi):
with set_cuda_backend("nvdec"):
dec = VideoDecoder(path, device="cuda", seek_mode="exact", dimension_order="NCHW")
frames = dec.get_frames_in_range(lo, hi).data
torch.cuda.synchronize()
out = frames.cpu().contiguous()
del dec
return out
ref = VideoDecoder("victim.mp4", device="cpu", seek_mode="exact", dimension_order="NCHW")
ref = ref.get_frames_in_range(12, 44).data.contiguous()
_ = nvdec_range("primer.mp4", 0, 32) # populate the cache
gpu = nvdec_range("victim.mp4", 12, 44)
ad = (gpu.to(torch.int16) - ref.to(torch.int16)).abs().float()
print(tuple(gpu.shape), "max", int(ad.max()), "mae", float(ad.mean()),
"top4_mae", float(ad[:, :, :4, :].mean()))
What we get on 0.16.0+cu130 (fresh process per row):
| scenario |
max Δ vs CPU |
MAE |
top-4-row MAE |
victim.mp4 alone |
2 |
0.40 |
0.39 |
primer.mp4 then victim.mp4, default cache |
255 |
5.38 |
136 |
same order, set_nvdec_cache_capacity(0) |
2 |
0.40 |
0.39 |
same order, set_cuda_backend("ffmpeg") |
2 |
0.40 |
0.39 |
Shape is always (32, 3, 530, 1280). Clean NVDEC vs CPU is the usual CSC slack (max Δ 2). The collide row is not slack. cache-off and ffmpeg match the victim.mp4-alone tensor hash.
Why
In 0.16.0, NVDECCache::CacheKey sets width / height from coded_width / coded_height and does not store display_area:
https://github.com/meta-pytorch/torchcodec/blob/v0.16.0/src/torchcodec/_core/NVDECCache.h
Any mixed-resolution corpus that lands two display sizes in the same coded bucket will hit this in a long-lived process.
Related: #944 #1176
Workaround
set_nvdec_cache_capacity(0) is correct, but it forces a full NVDEC create/destroy on every video. That is cheap on one quiet GPU and expensive when many processes open decoders at once (create/destroy does not scale with GPU count on a node).
Versions
torchcodec: 0.16.0+cu130
PyTorch version: 2.12.1+cu130
Is debug build: False
CUDA used to build PyTorch: 13.0
OS: Ubuntu 24.04.4 LTS (x86_64)
Libc version: glibc-2.39
Python version: 3.11.15 (64-bit runtime)
Python platform: Linux-5.15.0-170-generic-x86_64-with-glibc2.39
Is CUDA available: True
CUDA runtime version: 13.0.88
GPU models and configuration:
GPU 0–7: NVIDIA H200
Nvidia driver version: 595.71.05
The default
nvdecbackend cachesCUvideodecoderobjects keyed on coded width/height (plus codec / chroma / bit depth / surface format), not onCUVIDEOFORMAT.display_area.Two H.264 videos with different display heights that NVDEC treats as the same coded size therefore share a cache entry. The second video is decoded with the first video’s crop. There is no exception, and
shape/ frame count stay correct. The damage is a destroyed band of rows (usually the top).set_nvdec_cache_capacity(0)makes the pixels match a fresh-process decode (same tensor hash).set_cuda_backend("ffmpeg")is also clean (it does not use this cache).ffprobe still reports
coded_height == heighton the files below (532 vs 530). The collision is inside NVDEC /CUVIDEOFORMAT, not in the container metadata.Minimal repro
Needs one GPU, ffmpeg, and torchcodec with NVDEC. Each scenario must run in a fresh process (the cache is process-global).
What we get on 0.16.0+cu130 (fresh process per row):
victim.mp4aloneprimer.mp4thenvictim.mp4, default cacheset_nvdec_cache_capacity(0)set_cuda_backend("ffmpeg")Shape is always
(32, 3, 530, 1280). Clean NVDEC vs CPU is the usual CSC slack (max Δ 2). The collide row is not slack. cache-off andffmpegmatch thevictim.mp4-alone tensor hash.Why
In 0.16.0,
NVDECCache::CacheKeysetswidth/heightfromcoded_width/coded_heightand does not storedisplay_area:https://github.com/meta-pytorch/torchcodec/blob/v0.16.0/src/torchcodec/_core/NVDECCache.h
Any mixed-resolution corpus that lands two display sizes in the same coded bucket will hit this in a long-lived process.
Related: #944 #1176
Workaround
set_nvdec_cache_capacity(0)is correct, but it forces a full NVDEC create/destroy on every video. That is cheap on one quiet GPU and expensive when many processes open decoders at once (create/destroy does not scale with GPU count on a node).Versions