Skip to content

Captions per segment recording 41 - #28

Open
antobinary wants to merge 6 commits into
merge-v4.0.x-release-into-v4.1.x-developfrom
captions-per-segment-recording-41
Open

antobinary wants to merge 6 commits into
merge-v4.0.x-release-into-v4.1.x-developfrom
captions-per-segment-recording-41

Conversation

@antobinary

Copy link
Copy Markdown
Owner

What does this PR do?

Records live captions per utterance instead of as one shared text per language, so the caption track of a recording keeps several speakers apart and names who said what. It also fixes the other caption-recording defects found while reproducing bigbluebutton#19700, and adds regression tests for all of them.

Before: akka-apps kept one running transcript per locale for the whole meeting and recorded every interim result as an EditCaptionHistoryEvent (a character-index edit of that shared text). Whenever two people spoke the same language, each interim result of one speaker committed the other speaker's partial sentence and appended the whole running sentence again, with no separator and no speaker. Example from the dev server, two overlapping speakers:

the birch canoegluethe birch
canoe slidglue thethe birch
canoe slid onglue the sheetthe

After (same scenario, caption_en-US.vtt of the published recording):

00:00:08.604 --> 00:00:12.704
Alice: the birch canoe slid on the
smooth planks

00:00:09.212 --> 00:00:13.412
Bob: glue the sheet to the dark blue
background

00:00:11.529 --> 00:00:15.329
Alice: it is easy to tell the depth of
a well

00:00:11.831 --> 00:00:15.731
Bob: these days a chicken leg is a
rare dish

Both recordings hold 37 caption events: the new event costs no more events than the old one.

How

  • akka-apps records CaptionUpdatedEvent (module CAPTION) from what the live view already has: one event per update of a caption segment (captionId, userId, locale, captionType, the whole text, isFinal), from TranscriptUpdatedEvtMsg (audio transcription) and CaptionSubmitTranscriptEvtMsg (typed / plugin captions). This is the "properly designed event" Recording "TranscriptUpdatedEvent" is useless and spammy. bigbluebutton/bigbluebutton#19701 asked for. AudioCaptions, EditCaptionHistoryEvtMsg, the dead TranscriptUpdatedRecordEvent and the unused transcript.words/lines settings are removed.
  • gen_webvtt replays each segment on its own (per-code-unit timestamps from the first appearance of each character), merges the cues of a locale in time order and names the speaker whenever it changes. Late interim updates after a segment's final are ignored. Recording pauses are applied per segment. Recordings with the legacy EditCaptionHistoryEvent take the old code path: output is byte-identical for all 60 legacy recordings on the test server.
  • Speaker names follow the chat anonymization settings (anonymize_chat, anonymize_chat_moderators, meta_bbb-anonymize-chat*) through a shared BigBlueButton.generate_webvtt helper used by the captions step and by the presentation, video and screenshare formats.
  • The client no longer sends a debounced interim result after the final one.

Other defects fixed on the way

Defect Also on Effect
Typed captions (captionSubmitTranscript, e.g. from caption plugins) were never recorded 3.0, 4.0 stored for the live view only
getRecordingTextTracks links captions_<lang>.vtt, the recording wrote caption_<lang>.vtt 3.0, 4.0 every live track link returned 404
Client: a pending 200 ms debounced interim was sent after the final result 3.0, 4.0 last word(s) cut from the live caption and the recording (4 of 14 real Chrome utterances; e.g. " well" lost)
rap-caption-inbox.rb wrote to props['presentation_dir'], a key that never existed (b88b70d, 2019) 3.0, 4.0 every putRecordingTextTrack upload raised ENOENT, the handler crash-looped and stayed failed, later uploads piled up
The inbox looked for integration scripts in scripts/captions/, packages install scripts/caption/presentation (983751c, 2019) 3.0, 4.0 uploaded tracks never reached the playback
caption/presentation lacked require 'json'; its sort discarded the result 3.0, 4.0 failed as soon as it ran

Closes Issue(s)

Closes bigbluebutton#19700
Closes bigbluebutton#19701 (the event is replaced by CaptionUpdatedEvent)
Closes bigbluebutton#19178 (already fixed by the 4.0 locale validation; this adds the regression test)
Closes bigbluebutton#19699 (cannot happen on 4.x: one bn-document__<meetingId> pad, no _cc_ pads; verified in the notes + captions scenario)

Motivation

bigbluebutton#19700 asks for a new caption recording design because the Flash-era EditCaptionHistoryEvent cannot represent several streams per language. The live side already stores captions per utterance and per user (caption table). Recording exactly that keeps the recording consistent with what participants saw and removes a meeting-long, quadratically growing text buffer from akka-apps.

How to test

Regression tests (all fail on the current code, pass with this PR):

# Playwright, against a server with this branch (needs allowOverrideClientSettingsOnCreateCall=true:
# the spec turns on audioCaptions and enableApolloDevTools per meeting)
cd bigbluebutton-tests/playwright
npx playwright test captions

# gen_webvtt unit tests, where python3-lxml and python3-pyicu are installed (any BBB server)
python3 -m unittest discover -s record-and-playback/core/test/gen_webvtt -v

Manual: enable public.app.audioCaptions.enabled, record a meeting where two people with Chrome speak English at the same time, end it, then check the caption track in the presentation playback, or fetch it via getRecordingTextTracks.

Upgrade note: live tracks of recordings processed before this change are still named caption_<lang>.vtt in /var/bigbluebutton/captions/<recordID>/. A one-off rename (in new-features.md) makes their getRecordingTextTracks links work. Rebuilding is not needed.

Commits

Verification

Tested on a 4.1 dev server (4.1.0-alpha.1 packages, components built from this branch: akka-apps, html5 client, record-and-playback).

Playwright captions spec, A/B (the same spec, only the build switched):

Test stock this PR
Two speakers in the same language are recorded as separate labelled utterances FAIL PASS
Speaker labels follow the chat anonymization of the recording FAIL PASS
Transcripts in two languages are recorded as separate tracks FAIL (link 404 only; the content was right) PASS
Typed captions are recorded FAIL PASS
An uploaded caption track is published next to the live captions FAIL PASS
A late interim result does not truncate the final transcript FAIL PASS
Transcripts with an invalid locale are rejected (bigbluebutton#19178 guard) PASS PASS

On stock, the recorded-track tests fail first on the 404 of getRecordingTextTracks. With only the record-and-playback part of this PR on a stock akka-apps and client, they fail on the content itself, e.g. "the birch canoe slid on the smooth planks" should appear exactly once in the recorded captions: thegluethe birchglue thethe birch canoe.... With this PR all 7 passed against the pushed head; the six tests other than the upload one also passed in 4 earlier runs, the upload test in 1 earlier run.

  • gen_webvtt unit tests: 11/11. The new script writes byte-identical output to the old one for all 60 legacy recordings on the server.
  • Reproduction harness: 16 synthetic scenarios on this branch, covering 1 to 3 speakers, overlap, interjection, two languages, typed captions, pause/resume, recording started late, leaving audio mid-utterance, notes plus captions and long sessions. Every utterance appears exactly once, under its own speaker. Speech before the recording started or during a pause is left out. On stock the same scenarios grew 163 → 761 characters (two speakers) and 470 → 3923 (three long overlaps).
  • Real Google Chrome Web Speech, two users through virtual microphones, 7 utterances each in turn-taking and overlapping speech: 14/14 utterances recorded once under the right speaker. On stock the overlap round was garbled, and one utterance lost its last word.
  • The presentation, video and screenshare process steps produce labelled, anonymized tracks (Alice: ..., Viewer 1: ...).
  • Caption uploads: on stock the upload handler crash-looped. With this PR, a de-DE upload is published next to the live en-US track for a recording made with this PR, and for one processed before it.
  • Regression: the recording, audio, options/audioProcessing* and parameters suites, @ci, on this branch: 150 passed, 5 failed. 3 are @flaky and fail identically on stock. The other 2 pass alone on both builds; they failed in the full run on a test-results collision with a parallel run.
  • Static: html5 tsc and Playwright tsc + eslint clean. akka-apps compiles; its test tree does not compile on this line already (stale fixtures), so there is no Scala unit test.

More

  • Based on the 4.0.x-release → 4.1.x-develop forward-merge (merge-v4.0.x-release-into-v4.1.x-develop, 6b9cf3d), which brings the locale validation and client changes these files build on. Until that merge lands, the diff against v4.1.x-develop includes it; this compare shows only this series.
  • The same per-locale model, client debounce and upload bugs exist on 3.0 and 4.0 (identical files). The upload and debounce fixes are small and backportable on their own. The event change is a recording format change and fits 4.1.

Found while working on this, not changed here:

  • CaptionUpdatedEvent is written for every interim result, with the whole segment text. Recordings have the same number of caption events as before, at 20–27% more bytes (two overlapping speakers: 12.8 → 15.5 kB, one speaker with 12 sentences: 35.6 → 45.3 kB). Recording only the first and final update per segment would shrink it, at the cost of per-word cue timing.
  • On 4.1/LiveKit the "must be in voice" check in UpdateTranscriptPubMsgHdlr never drops a transcript: closing the audio modal still yields a voice user, and "Leave audio" leaves user_voice.joined=true, muted=true.
  • caption.captionId is a global primary key, and the upsert ignores a captionId already used by another user, in any meeting. Real client ids are <userId>-<ms>, so this needs a buggy or hostile client. The spec keeps its ids unique.
  • caption/presentation's branch for playbacks without captions.json writes to '' and crashes. Only pre-2.x playback formats reach it.
  • Uploaded tracks of kind subtitles are listed by the API but not shown by the presentation playback (existing TODO in caption/presentation).
  • A speaker label is added after line wrapping, so a labelled first line can exceed 32 characters.
  • record-and-playback/deploy.sh (developer script) deletes core/scripts and restores only the presentation format, which removes the packaged video/screenshare scripts.

…#19700)

Playwright (bigbluebutton-tests/playwright/captions): a stand-in for the Web
Speech API is installed before the client loads, so the real webspeech
transcription path (debounce, diff, captionSubmitText, akka, recording) runs
in headless Chromium and the spec feeds it the recognizer results. It covers:

- two speakers in one language recorded as separate utterances, each under
  its speaker's label (akka keeps one text buffer per locale and re-appends
  and glues the other speaker's words at every turn)
- transcripts in two languages recorded as separate tracks
- typed captions (captionSubmitTranscript) reaching the recording (they are
  stored for the live view only)
- a pending debounced interim result sent after the final result, which
  truncates the caption (client race, seen with Chrome)
- invalid locales rejected before anything is recorded (bigbluebutton#19178)

The recorded tracks are read back the way integrations do, through
getRecordingTextTracks. core/apolloProbe.ts gains typedVarMutation for
actions with Int/Boolean arguments.

gen_webvtt (record-and-playback/core/test/gen_webvtt): unittest fixtures run
the generator the way the recording process does. Recordings with the legacy
EditCaptionHistoryEvent (single speaker, pause/resume, empty or unsafe locale
bigbluebutton#19178, AUDIO-CAPTIONS events bigbluebutton#19701) keep working; the CaptionUpdatedEvent
fixtures describe the per-segment format this series introduces (labels on
speaker change, interleaved long utterances, recording paused or started
late, an utterance without a final result, a late interim after the final,
names supplied by the recording scripts) and fail until it lands.

On the current code the recorded-caption tests also fail because
getRecordingTextTracks links <kind>_<lang>.vtt while the recording writes
caption_<lang>.vtt. Caption ids are a global primary key, so the spec keeps
them unique per meeting.

Run the generator tests where python3-lxml and python3-pyicu are installed:
  python3 -m unittest discover -s record-and-playback/core/test/gen_webvtt
…h the speaker

gen_webvtt understands CaptionUpdatedEvent, the per-segment caption event
akka-apps records from this series on: one event per update of a caption
segment (an utterance of a speaker, or a typed caption) carrying captionId,
userId, locale, captionType, the whole text and isFinal. Every segment is
replayed on its own, with each code unit timestamped when it first appeared,
so two people speaking the same language no longer share (and garble) one
text buffer. The cues of a locale are merged in time order into one track,
and a cue names its speaker ("Alice: ...") whenever the speaker differs from
the previous cue's; overlapping cues are legal WebVTT and players stack them.
A non-final update that arrives after a segment's final one is ignored, and
recording pauses are applied per segment as before.

Recordings with the legacy EditCaptionHistoryEvent go through the old code
path unchanged: the new script produces byte-identical files for all 60
legacy recordings on the dev server and for the legacy fixtures.

Speaker labels use the participant names anonymized like the chat
(anonymize_chat / anonymize_chat_moderators, meta_bbb-anonymize-chat*). The
anonymization lookup moves into Events.anonymize_settings /
participant_name_map, shared by the chat and the new
BigBlueButton.generate_webvtt helper, which passes the names to gen_webvtt
(--speaker-names) for the captions step and the presentation, video and
screenshare formats alike.

The captions step now stores the tracks as captions_<lang>.vtt, the
<kind>_<lang>.vtt name that getRecordingTextTracks links to and that uploaded
tracks already use: the generated tracks were never downloadable through the
API (404).

The Playwright spec gains a test that the labels follow the chat
anonymization of the recording.
…locale

AudioCaptions kept a single running transcript per locale for the whole
meeting and recorded every interim result as an EditCaptionHistoryEvent, a
character-index edit of that shared text. With two people speaking the same
language, each interim result of one speaker changed the buffer's current
transcriptId, so the other speaker's whole running sentence was committed and
appended again, without separator or speaker: 163 characters of overlapping
speech became 761 in the recording, three overlapped pairs of long sentences
3923 for 470. The text also grew for the whole meeting, in memory and in the
recording.

The recorder now writes CaptionUpdatedEvent from what the live view already
has, one segment per captionId:
- TranscriptUpdatedEvtMsg (audio transcription, every interim and the final
  result; isFinal = result)
- CaptionSubmitTranscriptEvtMsg (typed captions and plugin-submitted
  captions, which were stored for the live view only and never recorded)
with captionId, userId, locale, captionType, the whole text and isFinal.
That is the event bigbluebutton#19701 asked for: replace-not-append semantics per
captionId, distinguishable segments, the speaker, no duplicate of another
event.

AudioCaptions, EditCaptionHistoryEvtMsg, EditCaptionHistoryRecordEvent, the
commented-out TranscriptUpdatedRecordEvent and the unused transcript.words /
transcript.lines settings go away. gen_webvtt keeps reading
EditCaptionHistoryEvent for older recordings.
The webspeech recognizer reports the final result of an utterance right
after its last interim result. Interim results go through a 200 ms trailing
debounce, final results are sent directly, so a pending interim update was
sent after the final one and replaced it: the caption shown to everyone (and
the recording) lost its last word(s). Seen with Chrome in 4 of 14 real
utterances; in one of them " well" disappeared from "it is easy to tell the
depth over well" (final at +174 ms, stale interim at +201 ms).

debounce() gains cancel(), which drops a pending trailing call; the final
transcript cancels it before it is sent, and so does unmounting the
recognizer component.
…layback

putRecordingTextTrack accepted a track, but nothing published it:

- rap-caption-inbox.rb wrote "#{props['presentation_dir']}/<id>/captions.json"
  (b88b70d, 2019). bigbluebutton.yml has never had presentation_dir, so
  the path was "/<id>/captions.json" and every upload raised Errno::ENOENT.
  Only InvalidCaptionError was rescued: the handler died, systemd restarted
  it, the restart picked the same file up again from the inbox, and after
  five crashes the service stayed failed, so later uploads piled up
  unprocessed. The block also overwrote the playback's captions.json with
  the one new track and copied it over caption_en-US.vtt whatever its
  language. It goes away: the caption integration scripts do that job.
- The handler looked for the integration scripts in scripts/captions/, but
  the presentation one has always been installed as
  scripts/caption/presentation (983751c, 2019): it never ran.
- That script, once run, failed on a missing require 'json', and its sort
  of the track list discarded its result.

Unexpected errors are now logged and the upload stays in the inbox for the
next start, instead of taking the service down. For recordings processed
before the caption files were named captions_<lang>.vtt, the integration
falls back to caption_<lang>.vtt for the live track.

The Playwright spec gains a test that uploads a track for a recording with
live captions and finds both in the presentation playback and in
getRecordingTextTracks.
…ds fixed

What's new in 4.1: per-utterance caption recording with speaker names
(following the chat anonymization settings), typed captions in recordings,
downloadable live tracks from getRecordingTextTracks, uploaded tracks in the
presentation playback, CaptionUpdatedEvent in events.xml, and a one-off
rename for the live tracks of recordings processed before the upgrade.
@github-actions

github-actions Bot commented Oct 9, 2026

Copy link
Copy Markdown

🚨 Automated tests failed

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant