Claude's plan - just opening this so I don't forget (I don't trust myself with files)
Demuxer.scan() and metadata for the Blocks API
Context
Demuxer.seek(seconds) is, structurally, VideoDecoder's seek_mode="approximate":
it converts seconds to a pts and hands that to avformat_seek_file
(Demuxer.cpp:93-106), the identical pair of steps SingleStreamDecoder takes
for get_frame_played_at() (SingleStreamDecoder.cpp:858 → :1360 → :1483
→ :1504). This is pinned by
test_seek_matches_video_decoder_approximate_get_frame_played_at.
Because the seek is a pure seconds → pts conversion, the only thing
seek_mode="exact" adds for a timestamp lookup is a scanned,
presentation-order keyframe index: it snaps the target back to the preceding
keyframe's pts before seeking (SingleStreamDecoder.cpp:1490-1502), working
around FFmpeg resolving seeks against decode timestamps
(https://trac.ffmpeg.org/ticket/11137). Without it FFmpeg can land on a keyframe
displayed after the target, leaving the frames in between permanently
unreachable — test_seek_to_non_keyframe_can_land_past_target shows this on
h265_video.mp4 (keyframes at 0.0/0.2/0.4/0.6/0.8; seeking to 0.3 lands on 0.4).
This reproduces on variable-frame-rate files too; it is the entire difference
between the modes for timestamp lookups, not an incidental quirk of one asset.
Verified: snapping the target to the preceding keyframe pts in Python and
then calling Demuxer.seek() reproduces seek_mode="exact" output on all 10
targets of h265_video.mp4. A scan is the missing piece, and nothing else is.
Two standing TODOs are really the same design question — _demuxer.py:52
(should the blocks offer exact seeking?) and __init__.py:32 ("we probably need
a way to expose metadata … that would avoid using VideoDecoder?"). Some
metadata is only knowable after a scan, so the two must be designed together.
The metadata design question
VideoStreamMetadata (_core/_metadata.py) has three tiers:
_from_header fields — cheap, possibly wrong.
_from_content fields — only populated by a scan.
- Computed fields —
num_frames, average_fps, duration_seconds,
begin_stream_seconds, end_stream_seconds — whose value silently depends
on whether a scan happened (Metadata.cpp:16-131, all keyed on SeekMode).
Tier 3 is defensible in VideoDecoder because the seek mode is fixed at
construction: num_frames means one thing for the object's whole life. It is
not defensible in the blocks, where there is no seek mode and a scan is
something the caller may do later — a computed field would change meaning
mid-life. That is the worst form of the ambiguity, and exactly the kind of
implicit behaviour the blocks exist to avoid.
Recommendation: make provenance structural, not a suffix
Demuxer.metadata → header-only, available at construction, free.
Precisely VideoStreamMetadata minus the computed fields and minus the
_from_content fields.
Demuxer.scan() → StreamIndex, which is the content-derived metadata:
the per-frame arrays plus the aggregates derived from them.
No field ever means two different things. Holding a StreamIndex means your
numbers are exact by construction; holding only metadata means the
_from_header suffix is right there telling you they are guesses.
discard_first_keyframe.mp4 makes this concrete: the header says 30 frames, the
demuxer yields 30 packets, and the pipeline produces 25. Both
metadata.num_frames_from_header == 30 and len(scan()) == 25 are correct, and
both are usefully named.
Two corollaries, treated as rules:
scan() must not write back into demuxer.metadata. Tempting, and
morally what VideoDecoder does, but it reintroduces "same attribute,
different meaning depending on history".
- Keep the
_from_header suffix on the four ambiguous fields
(num_frames_from_header, average_fps_from_header,
duration_seconds_from_header, begin_stream_seconds_from_header) even
though everything in that object is from the header and the suffix looks
redundant. Someone arriving from VideoDecoder will read a bare
demuxer.metadata.num_frames as the exact one. Drop it only where
VideoDecoder also carries none (width, height, codec, pixel_format,
rotation, bit_rate, pixel_aspect_ratio, color_*) — no new vocabulary.
API
demuxer = Demuxer(path)
demuxer.metadata # free, header-only
.width / .height / .codec / .pixel_format
.rotation / .pixel_aspect_ratio / .bit_rate / .stream_index
.color_space / .color_primaries / .color_transfer_characteristic
.num_frames_from_header
.average_fps_from_header
.duration_seconds_from_header
.begin_stream_seconds_from_header
index = demuxer.scan() # one full demux pass, no decoding
len(index) # exact frame count
index.pts_seconds # float64 [N], presentation order
index.duration_seconds # float64 [N]
index.is_key_frame # bool [N]
index.begin_stream_seconds # pts_seconds[0]
index.end_stream_seconds # pts_seconds[-1] + duration_seconds[-1]
index.average_fps # from content
index.index_at(seconds) # -> int
index.keyframe_seconds_at_or_before(seconds) # -> float
Exact seeking becomes one visible extra call, leaving seek() a single
primitive meaning "go to this timestamp":
demuxer.seek(index.keyframe_seconds_at_or_before(t)) # exact
packet_decoder.reset()
# ... then decode forward, dropping frames until pts + duration > t
demuxer.seek(t) # approximate, unchanged
Chosen over seek(t, index=index) (couples the primitive to the index type) and
over scan() flipping hidden state on the Demuxer (no index object to sample
with, and no way back to the cheap seek).
Use cases unlocked
- Exact-mode seeking, opt-in and explicitly priced. The motivating case.
- Index-based access, which the blocks have none of today:
index_at()
plus pts_seconds[i] gives seconds↔index both ways, enough to build
get_frame_at(i) or a clip sampler on top of the blocks.
- Honest stream metadata without a
VideoDecoder — closes __init__.py:32,
with the header/content split visible in the types.
- Scan once, decode many — the index is plain tensors, cacheable to disk and
reusable across epochs and dataloader workers.
- (Follow-up) Feed
VideoDecoder(custom_frame_mappings=...). That path
exists (_video_decoder.py:545, consumed at SingleStreamDecoder.cpp:361)
but today needs an ffprobe subprocess — see
examples/decoding/custom_frame_mappings.py. A native scan replaces it.
Implementation
The two pieces are independent (the scan op returns the timebase itself, so
StreamIndex does not need Demuxer.metadata) and can land in either order.
1. Metadata
Reuse, don't duplicate. SingleStreamDecoder.cpp:116-212 is pure
AVStream → StreamMetadata population that touches nothing else on the
decoder. Factor it into Metadata.cpp as a free function, e.g.
StreamMetadata stream_metadata_from_av_stream(const AVStream*, const AVFormatContext*),
and call it from both SingleStreamDecoder and the new op. StreamMetadata
already lives in Metadata.h:24.
custom_ops.cpp: _blocks_demuxer_metadata(Tensor demuxer) -> str returning
the same JSON shape the existing metadata ops use (custom_ops.cpp:~1189), so
the Python side can parse it the way _metadata.py:249 already does.
_blocks/_metadata.py: a lean StreamMetadata dataclass with only the fields
above. Do not reuse _core/_metadata.VideoStreamMetadata — its computed
fields are plain dataclass fields, so we would have to fill them with None
and mislead. (If we want provenance in the type name too, HeaderMetadata is
the alternative; Demuxer.metadata reads fine either way.)
Demuxer.metadata as a cached property.
2. scan()
C++ — Demuxer.{h,cpp}, modelled on
SingleStreamDecoder::scan_file_and_update_metadata_and_index
(SingleStreamDecoder.cpp:271-359) reduced to the active stream:
struct StreamIndex {
std::vector<int64_t> pts; // timebase units, presentation order
std::vector<int64_t> duration;
std::vector<bool> is_key_frame;
};
StreamIndex scan();
- Loop
read_next_packet(...) (already shared, Demuxer.h:21) to AVERROR_EOF.
- Skip
AV_PKT_FLAG_DISCARD packets, as the existing scan does
(SingleStreamDecoder.cpp:292). Load-bearing, not cosmetic: on
test/resources/discard_first_keyframe.mp4 the demuxer yields 30 packets
but the pipeline produces 25 frames, and seek_mode="exact" reports
num_frames == 25. Counting packets would disagree with both.
- Record
get_pts_or_dts(packet), packet->duration, flags & AV_PKT_FLAG_KEY.
- Sort by pts: packets arrive in decode order and presentation order is the
whole point (cf. sort_all_frames, SingleStreamDecoder.cpp:232).
- Rewind via
avformat_seek_file, reusing get_seek_error_message()
(Demuxer.h:26) so an unseekable source gives the same "does not support
seeking" message seek() gives.
- Hold no state:
scan() returns the index and caches nothing.
Ops: _blocks_demuxer_scan(Tensor(a!) demuxer) -> (Tensor, Tensor, Tensor, int, int)
— pts int64 [N], duration int64 [N], is_key_frame bool [N], timebase
num/den. Schema next to the other _blocks_* defs (custom_ops.cpp:77-92),
impl next to _blocks_demuxer_seek (:1511), name in
_ffmpeg_op_names.py:32-42, binding in _ffmpeg_ops.py:99. Returning raw
timebase units plus the timebase (rather than seconds) keeps the
custom_frame_mappings follow-up free, since that path is
int64-in-timebase-units.
Python — _blocks/_demuxer.py: @dataclass class StreamIndex with the
fields above, __len__, and the two lookups via torch.searchsorted.
keyframe_seconds_at_or_before returns pts_seconds[0] when the target
precedes the first keyframe, matching
get_key_frame_index_for_pts_using_scanned_index returning -1 and exact mode
falling through (SingleStreamDecoder.cpp:1497-1501). The seconds → pts round
trip is lossless (seconds_to_closest_pts is round(s*den/num),
pts_to_seconds is pts*num/den), so a pts_seconds value fed back into
seek() lands on that exact pts.
scan()'s docstring must say it reads the whole stream and leaves the demuxer
at the start, so any PacketDecoder built from it needs reset() — the same
contract seek() has.
Export StreamIndex and StreamMetadata from _blocks/__init__.py; retire the
__init__.py:32 and _demuxer.py:52 TODOs.
3. Docs — examples/decoding/blocks.py
Extend the Seeking section (:171-194) with exact seeking via the index, and
add a short metadata section making the header/scan split explicit.
Pre-existing bug worth fixing while there: :192 uses pts_seconds >= seconds
to find the "target frame", which on nasa_13013.mp4 at 2.5s returns 2.5025
when the frame actually playing at 2.5s starts at 2.4691. Correct criterion is
pts + duration > seconds.
Tests — test/test_decoders.py, class TestBlocks
New # ===== scanning ===== section after the seeking one:
test_scan_matches_video_decoder_index — over the trio the seek tests use
(NASA_VIDEO, H265_VIDEO, TEST_SRC_2_720P_MPEG4): len(index) equals
VideoDecoder(seek_mode="exact").metadata.num_frames; index.pts_seconds
equals every frame's pts; index.is_key_frame.nonzero() equals
_get_key_frame_indices().
test_scan_skips_discarded_packets — DISCARD_FIRST_KEYFRAME_VIDEO gives
len(index) == 25, the frame count the pipeline produces, not 30.
test_scan_seek_matches_video_decoder_exact — the payoff. Same shape as the
existing approximate test but seeking via keyframe_seconds_at_or_before()
and comparing against seek_mode="exact". Must include H265_VIDEO, where
the plain seek overshoots and this must not.
test_scan_rewinds_demuxer — decoding after a scan() yields the whole
stream from the start; scan() after a seek() gives the same index.
test_scan_on_non_seekable_source_raises — mirror
test_seek_on_non_seekable_source_raises (:4556) and its mkfifo setup.
test_metadata_matches_video_decoder — every Demuxer.metadata field equals
the corresponding VideoDecoder(...).metadata one, over _ALL_VIDEOS.
test_metadata_is_header_only — on DISCARD_FIRST_KEYFRAME_VIDEO,
metadata.num_frames_from_header is the header's 30 while len(scan()) is
25, and a scan() does not mutate demuxer.metadata.
Note when asserting durations: blocks frames carry the frame duration
(get_duration(av_frame), custom_ops.cpp:894) whereas the scan carries the
packet duration; compare index.duration_seconds against VideoDecoder's
values, not against DecodedFrame.duration_seconds.
Verification
/home/nicolashug/miniconda3/envs/codeccuda/bin/pip install -e . --no-build-isolation
pytest test/test_decoders.py -k "TestBlocks and (scan or metadata)" -v
pytest test/test_decoders.py -k "TestBlocks" -q # no regressions
python examples/decoding/blocks.py # example still runs
Manual end-to-end check that exact mode is genuinely reproduced:
index = Demuxer("test/resources/h265_video.mp4").scan()
# plain seek(0.3) -> 0.4 ; via the index -> 0.3, matching seek_mode="exact"
Explicitly out of scope
VideoDecoder(custom_frame_mappings=index) — worth doing next (it would
delete the ffprobe subprocess from
examples/decoding/custom_frame_mappings.py) but it touches VideoDecoder's
public signature and _read_custom_frame_mappings (_video_decoder.py:545).
The op above already returns the timebase-unit tensors that path needs.
- The
IndexError: Invalid frame index=145 ... must be less than 120 that
VideoDecoder(seek_mode="approximate").get_frames_played_at() raises on
variable-frame-rate files. Found while investigating this; a pre-existing
VideoDecoder bug with no coverage today, unrelated to the blocks.
Claude's plan - just opening this so I don't forget (I don't trust myself with files)
Demuxer.scan()and metadata for the Blocks APIContext
Demuxer.seek(seconds)is, structurally,VideoDecoder'sseek_mode="approximate":it converts
secondsto a pts and hands that toavformat_seek_file(
Demuxer.cpp:93-106), the identical pair of stepsSingleStreamDecodertakesfor
get_frame_played_at()(SingleStreamDecoder.cpp:858→:1360→:1483→
:1504). This is pinned bytest_seek_matches_video_decoder_approximate_get_frame_played_at.Because the seek is a pure
seconds → ptsconversion, the only thingseek_mode="exact"adds for a timestamp lookup is a scanned,presentation-order keyframe index: it snaps the target back to the preceding
keyframe's pts before seeking (
SingleStreamDecoder.cpp:1490-1502), workingaround FFmpeg resolving seeks against decode timestamps
(https://trac.ffmpeg.org/ticket/11137). Without it FFmpeg can land on a keyframe
displayed after the target, leaving the frames in between permanently
unreachable —
test_seek_to_non_keyframe_can_land_past_targetshows this onh265_video.mp4(keyframes at 0.0/0.2/0.4/0.6/0.8; seeking to 0.3 lands on 0.4).This reproduces on variable-frame-rate files too; it is the entire difference
between the modes for timestamp lookups, not an incidental quirk of one asset.
Verified: snapping the target to the preceding keyframe pts in Python and
then calling
Demuxer.seek()reproducesseek_mode="exact"output on all 10targets of
h265_video.mp4. A scan is the missing piece, and nothing else is.Two standing TODOs are really the same design question —
_demuxer.py:52(should the blocks offer exact seeking?) and
__init__.py:32("we probably needa way to expose metadata … that would avoid using VideoDecoder?"). Some
metadata is only knowable after a scan, so the two must be designed together.
The metadata design question
VideoStreamMetadata(_core/_metadata.py) has three tiers:_from_headerfields — cheap, possibly wrong._from_contentfields — only populated by a scan.num_frames,average_fps,duration_seconds,begin_stream_seconds,end_stream_seconds— whose value silently dependson whether a scan happened (
Metadata.cpp:16-131, all keyed onSeekMode).Tier 3 is defensible in
VideoDecoderbecause the seek mode is fixed atconstruction:
num_framesmeans one thing for the object's whole life. It isnot defensible in the blocks, where there is no seek mode and a scan is
something the caller may do later — a computed field would change meaning
mid-life. That is the worst form of the ambiguity, and exactly the kind of
implicit behaviour the blocks exist to avoid.
Recommendation: make provenance structural, not a suffix
Demuxer.metadata→ header-only, available at construction, free.Precisely
VideoStreamMetadataminus the computed fields and minus the_from_contentfields.Demuxer.scan()→StreamIndex, which is the content-derived metadata:the per-frame arrays plus the aggregates derived from them.
No field ever means two different things. Holding a
StreamIndexmeans yournumbers are exact by construction; holding only
metadatameans the_from_headersuffix is right there telling you they are guesses.discard_first_keyframe.mp4makes this concrete: the header says 30 frames, thedemuxer yields 30 packets, and the pipeline produces 25. Both
metadata.num_frames_from_header == 30andlen(scan()) == 25are correct, andboth are usefully named.
Two corollaries, treated as rules:
scan()must not write back intodemuxer.metadata. Tempting, andmorally what
VideoDecoderdoes, but it reintroduces "same attribute,different meaning depending on history".
_from_headersuffix on the four ambiguous fields(
num_frames_from_header,average_fps_from_header,duration_seconds_from_header,begin_stream_seconds_from_header) eventhough everything in that object is from the header and the suffix looks
redundant. Someone arriving from
VideoDecoderwill read a baredemuxer.metadata.num_framesas the exact one. Drop it only whereVideoDecoderalso carries none (width,height,codec,pixel_format,rotation,bit_rate,pixel_aspect_ratio,color_*) — no new vocabulary.API
Exact seeking becomes one visible extra call, leaving
seek()a singleprimitive meaning "go to this timestamp":
Chosen over
seek(t, index=index)(couples the primitive to the index type) andover
scan()flipping hidden state on the Demuxer (no index object to samplewith, and no way back to the cheap seek).
Use cases unlocked
index_at()plus
pts_seconds[i]gives seconds↔index both ways, enough to buildget_frame_at(i)or a clip sampler on top of the blocks.VideoDecoder— closes__init__.py:32,with the header/content split visible in the types.
reusable across epochs and dataloader workers.
VideoDecoder(custom_frame_mappings=...). That pathexists (
_video_decoder.py:545, consumed atSingleStreamDecoder.cpp:361)but today needs an
ffprobesubprocess — seeexamples/decoding/custom_frame_mappings.py. A native scan replaces it.Implementation
The two pieces are independent (the scan op returns the timebase itself, so
StreamIndexdoes not needDemuxer.metadata) and can land in either order.1. Metadata
Reuse, don't duplicate.
SingleStreamDecoder.cpp:116-212is pureAVStream→StreamMetadatapopulation that touches nothing else on thedecoder. Factor it into
Metadata.cppas a free function, e.g.StreamMetadata stream_metadata_from_av_stream(const AVStream*, const AVFormatContext*),and call it from both
SingleStreamDecoderand the new op.StreamMetadataalready lives in
Metadata.h:24.custom_ops.cpp:_blocks_demuxer_metadata(Tensor demuxer) -> strreturningthe same JSON shape the existing metadata ops use (
custom_ops.cpp:~1189), sothe Python side can parse it the way
_metadata.py:249already does._blocks/_metadata.py: a leanStreamMetadatadataclass with only the fieldsabove. Do not reuse
_core/_metadata.VideoStreamMetadata— its computedfields are plain dataclass fields, so we would have to fill them with
Noneand mislead. (If we want provenance in the type name too,
HeaderMetadataisthe alternative;
Demuxer.metadatareads fine either way.)Demuxer.metadataas a cached property.2.
scan()C++ —
Demuxer.{h,cpp}, modelled onSingleStreamDecoder::scan_file_and_update_metadata_and_index(
SingleStreamDecoder.cpp:271-359) reduced to the active stream:read_next_packet(...)(already shared,Demuxer.h:21) toAVERROR_EOF.AV_PKT_FLAG_DISCARDpackets, as the existing scan does(
SingleStreamDecoder.cpp:292). Load-bearing, not cosmetic: ontest/resources/discard_first_keyframe.mp4the demuxer yields 30 packetsbut the pipeline produces 25 frames, and
seek_mode="exact"reportsnum_frames == 25. Counting packets would disagree with both.get_pts_or_dts(packet),packet->duration,flags & AV_PKT_FLAG_KEY.whole point (cf.
sort_all_frames,SingleStreamDecoder.cpp:232).avformat_seek_file, reusingget_seek_error_message()(
Demuxer.h:26) so an unseekable source gives the same "does not supportseeking" message
seek()gives.scan()returns the index and caches nothing.Ops:
_blocks_demuxer_scan(Tensor(a!) demuxer) -> (Tensor, Tensor, Tensor, int, int)—
ptsint64[N],durationint64[N],is_key_framebool[N], timebasenum/den. Schema next to the other_blocks_*defs (custom_ops.cpp:77-92),impl next to
_blocks_demuxer_seek(:1511), name in_ffmpeg_op_names.py:32-42, binding in_ffmpeg_ops.py:99. Returning rawtimebase units plus the timebase (rather than seconds) keeps the
custom_frame_mappingsfollow-up free, since that path isint64-in-timebase-units.
Python —
_blocks/_demuxer.py:@dataclass class StreamIndexwith thefields above,
__len__, and the two lookups viatorch.searchsorted.keyframe_seconds_at_or_beforereturnspts_seconds[0]when the targetprecedes the first keyframe, matching
get_key_frame_index_for_pts_using_scanned_indexreturning-1and exact modefalling through (
SingleStreamDecoder.cpp:1497-1501). Theseconds → ptsroundtrip is lossless (
seconds_to_closest_ptsisround(s*den/num),pts_to_secondsispts*num/den), so apts_secondsvalue fed back intoseek()lands on that exact pts.scan()'s docstring must say it reads the whole stream and leaves the demuxerat the start, so any
PacketDecoderbuilt from it needsreset()— the samecontract
seek()has.Export
StreamIndexandStreamMetadatafrom_blocks/__init__.py; retire the__init__.py:32and_demuxer.py:52TODOs.3. Docs —
examples/decoding/blocks.pyExtend the Seeking section (
:171-194) with exact seeking via the index, andadd a short metadata section making the header/scan split explicit.
Pre-existing bug worth fixing while there:
:192usespts_seconds >= secondsto find the "target frame", which on
nasa_13013.mp4at 2.5s returns 2.5025when the frame actually playing at 2.5s starts at 2.4691. Correct criterion is
pts + duration > seconds.Tests —
test/test_decoders.py,class TestBlocksNew
# ===== scanning =====section after the seeking one:test_scan_matches_video_decoder_index— over the trio the seek tests use(
NASA_VIDEO,H265_VIDEO,TEST_SRC_2_720P_MPEG4):len(index)equalsVideoDecoder(seek_mode="exact").metadata.num_frames;index.pts_secondsequals every frame's pts;
index.is_key_frame.nonzero()equals_get_key_frame_indices().test_scan_skips_discarded_packets—DISCARD_FIRST_KEYFRAME_VIDEOgiveslen(index) == 25, the frame count the pipeline produces, not 30.test_scan_seek_matches_video_decoder_exact— the payoff. Same shape as theexisting approximate test but seeking via
keyframe_seconds_at_or_before()and comparing against
seek_mode="exact". Must includeH265_VIDEO, wherethe plain seek overshoots and this must not.
test_scan_rewinds_demuxer— decoding after ascan()yields the wholestream from the start;
scan()after aseek()gives the same index.test_scan_on_non_seekable_source_raises— mirrortest_seek_on_non_seekable_source_raises(:4556) and itsmkfifosetup.test_metadata_matches_video_decoder— everyDemuxer.metadatafield equalsthe corresponding
VideoDecoder(...).metadataone, over_ALL_VIDEOS.test_metadata_is_header_only— onDISCARD_FIRST_KEYFRAME_VIDEO,metadata.num_frames_from_headeris the header's 30 whilelen(scan())is25, and a
scan()does not mutatedemuxer.metadata.Note when asserting durations: blocks frames carry the frame duration
(
get_duration(av_frame),custom_ops.cpp:894) whereas the scan carries thepacket duration; compare
index.duration_secondsagainstVideoDecoder'svalues, not against
DecodedFrame.duration_seconds.Verification
Manual end-to-end check that exact mode is genuinely reproduced:
Explicitly out of scope
VideoDecoder(custom_frame_mappings=index)— worth doing next (it woulddelete the
ffprobesubprocess fromexamples/decoding/custom_frame_mappings.py) but it touchesVideoDecoder'spublic signature and
_read_custom_frame_mappings(_video_decoder.py:545).The op above already returns the timebase-unit tensors that path needs.
IndexError: Invalid frame index=145 ... must be less than 120thatVideoDecoder(seek_mode="approximate").get_frames_played_at()raises onvariable-frame-rate files. Found while investigating this; a pre-existing
VideoDecoderbug with no coverage today, unrelated to the blocks.