Skip to content

Exposing Scan and metadata in Blocks API #1655

Description

@NicolasHug

Claude's plan - just opening this so I don't forget (I don't trust myself with files)

Demuxer.scan() and metadata for the Blocks API

Context

Demuxer.seek(seconds) is, structurally, VideoDecoder's seek_mode="approximate":
it converts seconds to a pts and hands that to avformat_seek_file
(Demuxer.cpp:93-106), the identical pair of steps SingleStreamDecoder takes
for get_frame_played_at() (SingleStreamDecoder.cpp:858 → :1360 → :1483
→ :1504). This is pinned by
test_seek_matches_video_decoder_approximate_get_frame_played_at.

Because the seek is a pure seconds → pts conversion, the only thing
seek_mode="exact" adds for a timestamp lookup is a scanned,
presentation-order keyframe index
: it snaps the target back to the preceding
keyframe's pts before seeking (SingleStreamDecoder.cpp:1490-1502), working
around FFmpeg resolving seeks against decode timestamps
(https://trac.ffmpeg.org/ticket/11137). Without it FFmpeg can land on a keyframe
displayed after the target, leaving the frames in between permanently
unreachable — test_seek_to_non_keyframe_can_land_past_target shows this on
h265_video.mp4 (keyframes at 0.0/0.2/0.4/0.6/0.8; seeking to 0.3 lands on 0.4).
This reproduces on variable-frame-rate files too; it is the entire difference
between the modes for timestamp lookups, not an incidental quirk of one asset.

Verified: snapping the target to the preceding keyframe pts in Python and
then calling Demuxer.seek() reproduces seek_mode="exact" output on all 10
targets of h265_video.mp4. A scan is the missing piece, and nothing else is.

Two standing TODOs are really the same design question — _demuxer.py:52
(should the blocks offer exact seeking?) and __init__.py:32 ("we probably need
a way to expose metadata … that would avoid using VideoDecoder?"). Some
metadata is only knowable after a scan
, so the two must be designed together.

The metadata design question

VideoStreamMetadata (_core/_metadata.py) has three tiers:

  1. _from_header fields — cheap, possibly wrong.
  2. _from_content fields — only populated by a scan.
  3. Computed fields — num_frames, average_fps, duration_seconds,
    begin_stream_seconds, end_stream_seconds — whose value silently depends
    on whether a scan happened (Metadata.cpp:16-131, all keyed on SeekMode).

Tier 3 is defensible in VideoDecoder because the seek mode is fixed at
construction: num_frames means one thing for the object's whole life. It is
not defensible in the blocks
, where there is no seek mode and a scan is
something the caller may do later — a computed field would change meaning
mid-life. That is the worst form of the ambiguity, and exactly the kind of
implicit behaviour the blocks exist to avoid.

Recommendation: make provenance structural, not a suffix

  • Demuxer.metadata → header-only, available at construction, free.
    Precisely VideoStreamMetadata minus the computed fields and minus the
    _from_content fields.
  • Demuxer.scan() → StreamIndex, which is the content-derived metadata:
    the per-frame arrays plus the aggregates derived from them.

No field ever means two different things. Holding a StreamIndex means your
numbers are exact by construction; holding only metadata means the
_from_header suffix is right there telling you they are guesses.

discard_first_keyframe.mp4 makes this concrete: the header says 30 frames, the
demuxer yields 30 packets, and the pipeline produces 25. Both
metadata.num_frames_from_header == 30 and len(scan()) == 25 are correct, and
both are usefully named.

Two corollaries, treated as rules:

  • scan() must not write back into demuxer.metadata. Tempting, and
    morally what VideoDecoder does, but it reintroduces "same attribute,
    different meaning depending on history".
  • Keep the _from_header suffix on the four ambiguous fields
    (num_frames_from_header, average_fps_from_header,
    duration_seconds_from_header, begin_stream_seconds_from_header) even
    though everything in that object is from the header and the suffix looks
    redundant. Someone arriving from VideoDecoder will read a bare
    demuxer.metadata.num_frames as the exact one. Drop it only where
    VideoDecoder also carries none (width, height, codec, pixel_format,
    rotation, bit_rate, pixel_aspect_ratio, color_*) — no new vocabulary.

API

demuxer = Demuxer(path)

demuxer.metadata                        # free, header-only
  .width / .height / .codec / .pixel_format
  .rotation / .pixel_aspect_ratio / .bit_rate / .stream_index
  .color_space / .color_primaries / .color_transfer_characteristic
  .num_frames_from_header
  .average_fps_from_header
  .duration_seconds_from_header
  .begin_stream_seconds_from_header

index = demuxer.scan()                  # one full demux pass, no decoding
  len(index)                            # exact frame count
  index.pts_seconds                     # float64 [N], presentation order
  index.duration_seconds                # float64 [N]
  index.is_key_frame                    # bool    [N]
  index.begin_stream_seconds            # pts_seconds[0]
  index.end_stream_seconds              # pts_seconds[-1] + duration_seconds[-1]
  index.average_fps                     # from content
  index.index_at(seconds)               # -> int
  index.keyframe_seconds_at_or_before(seconds)   # -> float

Exact seeking becomes one visible extra call, leaving seek() a single
primitive meaning "go to this timestamp":

demuxer.seek(index.keyframe_seconds_at_or_before(t))   # exact
packet_decoder.reset()
# ... then decode forward, dropping frames until pts + duration > t

demuxer.seek(t)                                        # approximate, unchanged

Chosen over seek(t, index=index) (couples the primitive to the index type) and
over scan() flipping hidden state on the Demuxer (no index object to sample
with, and no way back to the cheap seek).

Use cases unlocked

  1. Exact-mode seeking, opt-in and explicitly priced. The motivating case.
  2. Index-based access, which the blocks have none of today: index_at()
    plus pts_seconds[i] gives seconds↔index both ways, enough to build
    get_frame_at(i) or a clip sampler on top of the blocks.
  3. Honest stream metadata without a VideoDecoder — closes __init__.py:32,
    with the header/content split visible in the types.
  4. Scan once, decode many — the index is plain tensors, cacheable to disk and
    reusable across epochs and dataloader workers.
  5. (Follow-up) Feed VideoDecoder(custom_frame_mappings=...). That path
    exists (_video_decoder.py:545, consumed at SingleStreamDecoder.cpp:361)
    but today needs an ffprobe subprocess — see
    examples/decoding/custom_frame_mappings.py. A native scan replaces it.

Implementation

The two pieces are independent (the scan op returns the timebase itself, so
StreamIndex does not need Demuxer.metadata) and can land in either order.

1. Metadata

Reuse, don't duplicate. SingleStreamDecoder.cpp:116-212 is pure
AVStream → StreamMetadata population that touches nothing else on the
decoder. Factor it into Metadata.cpp as a free function, e.g.
StreamMetadata stream_metadata_from_av_stream(const AVStream*, const AVFormatContext*),
and call it from both SingleStreamDecoder and the new op. StreamMetadata
already lives in Metadata.h:24.

  • custom_ops.cpp: _blocks_demuxer_metadata(Tensor demuxer) -> str returning
    the same JSON shape the existing metadata ops use (custom_ops.cpp:~1189), so
    the Python side can parse it the way _metadata.py:249 already does.
  • _blocks/_metadata.py: a lean StreamMetadata dataclass with only the fields
    above. Do not reuse _core/_metadata.VideoStreamMetadata — its computed
    fields are plain dataclass fields, so we would have to fill them with None
    and mislead. (If we want provenance in the type name too, HeaderMetadata is
    the alternative; Demuxer.metadata reads fine either way.)
  • Demuxer.metadata as a cached property.

2. scan()

C++ — Demuxer.{h,cpp}, modelled on
SingleStreamDecoder::scan_file_and_update_metadata_and_index
(SingleStreamDecoder.cpp:271-359) reduced to the active stream:

struct StreamIndex {
  std::vector<int64_t> pts;        // timebase units, presentation order
  std::vector<int64_t> duration;
  std::vector<bool> is_key_frame;
};
StreamIndex scan();
  • Loop read_next_packet(...) (already shared, Demuxer.h:21) to AVERROR_EOF.
  • Skip AV_PKT_FLAG_DISCARD packets, as the existing scan does
    (SingleStreamDecoder.cpp:292). Load-bearing, not cosmetic: on
    test/resources/discard_first_keyframe.mp4 the demuxer yields 30 packets
    but the pipeline produces 25 frames, and seek_mode="exact" reports
    num_frames == 25. Counting packets would disagree with both.
  • Record get_pts_or_dts(packet), packet->duration, flags & AV_PKT_FLAG_KEY.
  • Sort by pts: packets arrive in decode order and presentation order is the
    whole point (cf. sort_all_frames, SingleStreamDecoder.cpp:232).
  • Rewind via avformat_seek_file, reusing get_seek_error_message()
    (Demuxer.h:26) so an unseekable source gives the same "does not support
    seeking" message seek() gives.
  • Hold no state: scan() returns the index and caches nothing.

Ops: _blocks_demuxer_scan(Tensor(a!) demuxer) -> (Tensor, Tensor, Tensor, int, int)
— pts int64 [N], duration int64 [N], is_key_frame bool [N], timebase
num/den. Schema next to the other _blocks_* defs (custom_ops.cpp:77-92),
impl next to _blocks_demuxer_seek (:1511), name in
_ffmpeg_op_names.py:32-42, binding in _ffmpeg_ops.py:99. Returning raw
timebase units plus the timebase (rather than seconds) keeps the
custom_frame_mappings follow-up free, since that path is
int64-in-timebase-units.

Python — _blocks/_demuxer.py: @dataclass class StreamIndex with the
fields above, __len__, and the two lookups via torch.searchsorted.
keyframe_seconds_at_or_before returns pts_seconds[0] when the target
precedes the first keyframe, matching
get_key_frame_index_for_pts_using_scanned_index returning -1 and exact mode
falling through (SingleStreamDecoder.cpp:1497-1501). The seconds → pts round
trip is lossless (seconds_to_closest_pts is round(s*den/num),
pts_to_seconds is pts*num/den), so a pts_seconds value fed back into
seek() lands on that exact pts.

scan()'s docstring must say it reads the whole stream and leaves the demuxer
at the start
, so any PacketDecoder built from it needs reset() — the same
contract seek() has.

Export StreamIndex and StreamMetadata from _blocks/__init__.py; retire the
__init__.py:32 and _demuxer.py:52 TODOs.

3. Docs — examples/decoding/blocks.py

Extend the Seeking section (:171-194) with exact seeking via the index, and
add a short metadata section making the header/scan split explicit.

Pre-existing bug worth fixing while there: :192 uses pts_seconds >= seconds
to find the "target frame", which on nasa_13013.mp4 at 2.5s returns 2.5025
when the frame actually playing at 2.5s starts at 2.4691. Correct criterion is
pts + duration > seconds.

Tests — test/test_decoders.py, class TestBlocks

New # ===== scanning ===== section after the seeking one:

  1. test_scan_matches_video_decoder_index — over the trio the seek tests use
    (NASA_VIDEO, H265_VIDEO, TEST_SRC_2_720P_MPEG4): len(index) equals
    VideoDecoder(seek_mode="exact").metadata.num_frames; index.pts_seconds
    equals every frame's pts; index.is_key_frame.nonzero() equals
    _get_key_frame_indices().
  2. test_scan_skips_discarded_packets — DISCARD_FIRST_KEYFRAME_VIDEO gives
    len(index) == 25, the frame count the pipeline produces, not 30.
  3. test_scan_seek_matches_video_decoder_exact — the payoff. Same shape as the
    existing approximate test but seeking via keyframe_seconds_at_or_before()
    and comparing against seek_mode="exact". Must include H265_VIDEO, where
    the plain seek overshoots and this must not.
  4. test_scan_rewinds_demuxer — decoding after a scan() yields the whole
    stream from the start; scan() after a seek() gives the same index.
  5. test_scan_on_non_seekable_source_raises — mirror
    test_seek_on_non_seekable_source_raises (:4556) and its mkfifo setup.
  6. test_metadata_matches_video_decoder — every Demuxer.metadata field equals
    the corresponding VideoDecoder(...).metadata one, over _ALL_VIDEOS.
  7. test_metadata_is_header_only — on DISCARD_FIRST_KEYFRAME_VIDEO,
    metadata.num_frames_from_header is the header's 30 while len(scan()) is
    25, and a scan() does not mutate demuxer.metadata.

Note when asserting durations: blocks frames carry the frame duration
(get_duration(av_frame), custom_ops.cpp:894) whereas the scan carries the
packet duration; compare index.duration_seconds against VideoDecoder's
values, not against DecodedFrame.duration_seconds.

Verification

/home/nicolashug/miniconda3/envs/codeccuda/bin/pip install -e . --no-build-isolation

pytest test/test_decoders.py -k "TestBlocks and (scan or metadata)" -v
pytest test/test_decoders.py -k "TestBlocks" -q      # no regressions
python examples/decoding/blocks.py                    # example still runs

Manual end-to-end check that exact mode is genuinely reproduced:

index = Demuxer("test/resources/h265_video.mp4").scan()
# plain seek(0.3) -> 0.4 ; via the index -> 0.3, matching seek_mode="exact"

Explicitly out of scope

  • VideoDecoder(custom_frame_mappings=index) — worth doing next (it would
    delete the ffprobe subprocess from
    examples/decoding/custom_frame_mappings.py) but it touches VideoDecoder's
    public signature and _read_custom_frame_mappings (_video_decoder.py:545).
    The op above already returns the timebase-unit tensors that path needs.
  • The IndexError: Invalid frame index=145 ... must be less than 120 that
    VideoDecoder(seek_mode="approximate").get_frames_played_at() raises on
    variable-frame-rate files. Found while investigating this; a pre-existing
    VideoDecoder bug with no coverage today, unrelated to the blocks.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions