Repository navigation
Captions per segment recording 41 - #28
Open
antobinary wants to merge 6 commits into
Open
antobinary wants to merge 6 commits into
antobinary wants to merge 6 commits into
Conversation
…#19700) Playwright (bigbluebutton-tests/playwright/captions): a stand-in for the Web Speech API is installed before the client loads, so the real webspeech transcription path (debounce, diff, captionSubmitText, akka, recording) runs in headless Chromium and the spec feeds it the recognizer results. It covers: - two speakers in one language recorded as separate utterances, each under its speaker's label (akka keeps one text buffer per locale and re-appends and glues the other speaker's words at every turn) - transcripts in two languages recorded as separate tracks - typed captions (captionSubmitTranscript) reaching the recording (they are stored for the live view only) - a pending debounced interim result sent after the final result, which truncates the caption (client race, seen with Chrome) - invalid locales rejected before anything is recorded (bigbluebutton#19178) The recorded tracks are read back the way integrations do, through getRecordingTextTracks. core/apolloProbe.ts gains typedVarMutation for actions with Int/Boolean arguments. gen_webvtt (record-and-playback/core/test/gen_webvtt): unittest fixtures run the generator the way the recording process does. Recordings with the legacy EditCaptionHistoryEvent (single speaker, pause/resume, empty or unsafe locale bigbluebutton#19178, AUDIO-CAPTIONS events bigbluebutton#19701) keep working; the CaptionUpdatedEvent fixtures describe the per-segment format this series introduces (labels on speaker change, interleaved long utterances, recording paused or started late, an utterance without a final result, a late interim after the final, names supplied by the recording scripts) and fail until it lands. On the current code the recorded-caption tests also fail because getRecordingTextTracks links <kind>_<lang>.vtt while the recording writes caption_<lang>.vtt. Caption ids are a global primary key, so the spec keeps them unique per meeting. Run the generator tests where python3-lxml and python3-pyicu are installed: python3 -m unittest discover -s record-and-playback/core/test/gen_webvtt
…h the speaker
gen_webvtt understands CaptionUpdatedEvent, the per-segment caption event
akka-apps records from this series on: one event per update of a caption
segment (an utterance of a speaker, or a typed caption) carrying captionId,
userId, locale, captionType, the whole text and isFinal. Every segment is
replayed on its own, with each code unit timestamped when it first appeared,
so two people speaking the same language no longer share (and garble) one
text buffer. The cues of a locale are merged in time order into one track,
and a cue names its speaker ("Alice: ...") whenever the speaker differs from
the previous cue's; overlapping cues are legal WebVTT and players stack them.
A non-final update that arrives after a segment's final one is ignored, and
recording pauses are applied per segment as before.
Recordings with the legacy EditCaptionHistoryEvent go through the old code
path unchanged: the new script produces byte-identical files for all 60
legacy recordings on the dev server and for the legacy fixtures.
Speaker labels use the participant names anonymized like the chat
(anonymize_chat / anonymize_chat_moderators, meta_bbb-anonymize-chat*). The
anonymization lookup moves into Events.anonymize_settings /
participant_name_map, shared by the chat and the new
BigBlueButton.generate_webvtt helper, which passes the names to gen_webvtt
(--speaker-names) for the captions step and the presentation, video and
screenshare formats alike.
The captions step now stores the tracks as captions_<lang>.vtt, the
<kind>_<lang>.vtt name that getRecordingTextTracks links to and that uploaded
tracks already use: the generated tracks were never downloadable through the
API (404).
The Playwright spec gains a test that the labels follow the chat
anonymization of the recording.
…locale AudioCaptions kept a single running transcript per locale for the whole meeting and recorded every interim result as an EditCaptionHistoryEvent, a character-index edit of that shared text. With two people speaking the same language, each interim result of one speaker changed the buffer's current transcriptId, so the other speaker's whole running sentence was committed and appended again, without separator or speaker: 163 characters of overlapping speech became 761 in the recording, three overlapped pairs of long sentences 3923 for 470. The text also grew for the whole meeting, in memory and in the recording. The recorder now writes CaptionUpdatedEvent from what the live view already has, one segment per captionId: - TranscriptUpdatedEvtMsg (audio transcription, every interim and the final result; isFinal = result) - CaptionSubmitTranscriptEvtMsg (typed captions and plugin-submitted captions, which were stored for the live view only and never recorded) with captionId, userId, locale, captionType, the whole text and isFinal. That is the event bigbluebutton#19701 asked for: replace-not-append semantics per captionId, distinguishable segments, the speaker, no duplicate of another event. AudioCaptions, EditCaptionHistoryEvtMsg, EditCaptionHistoryRecordEvent, the commented-out TranscriptUpdatedRecordEvent and the unused transcript.words / transcript.lines settings go away. gen_webvtt keeps reading EditCaptionHistoryEvent for older recordings.
The webspeech recognizer reports the final result of an utterance right after its last interim result. Interim results go through a 200 ms trailing debounce, final results are sent directly, so a pending interim update was sent after the final one and replaced it: the caption shown to everyone (and the recording) lost its last word(s). Seen with Chrome in 4 of 14 real utterances; in one of them " well" disappeared from "it is easy to tell the depth over well" (final at +174 ms, stale interim at +201 ms). debounce() gains cancel(), which drops a pending trailing call; the final transcript cancels it before it is sent, and so does unmounting the recognizer component.
…layback
putRecordingTextTrack accepted a track, but nothing published it:
- rap-caption-inbox.rb wrote "#{props['presentation_dir']}/<id>/captions.json"
(b88b70d, 2019). bigbluebutton.yml has never had presentation_dir, so
the path was "/<id>/captions.json" and every upload raised Errno::ENOENT.
Only InvalidCaptionError was rescued: the handler died, systemd restarted
it, the restart picked the same file up again from the inbox, and after
five crashes the service stayed failed, so later uploads piled up
unprocessed. The block also overwrote the playback's captions.json with
the one new track and copied it over caption_en-US.vtt whatever its
language. It goes away: the caption integration scripts do that job.
- The handler looked for the integration scripts in scripts/captions/, but
the presentation one has always been installed as
scripts/caption/presentation (983751c, 2019): it never ran.
- That script, once run, failed on a missing require 'json', and its sort
of the track list discarded its result.
Unexpected errors are now logged and the upload stays in the inbox for the
next start, instead of taking the service down. For recordings processed
before the caption files were named captions_<lang>.vtt, the integration
falls back to caption_<lang>.vtt for the live track.
The Playwright spec gains a test that uploads a track for a recording with
live captions and finds both in the presentation playback and in
getRecordingTextTracks.
…ds fixed What's new in 4.1: per-utterance caption recording with speaker names (following the chat anonymization settings), typed captions in recordings, downloadable live tracks from getRecordingTextTracks, uploaded tracks in the presentation playback, CaptionUpdatedEvent in events.xml, and a one-off rename for the live tracks of recordings processed before the upgrade.
🚨 Automated tests failed |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Records live captions per utterance instead of as one shared text per language, so the caption track of a recording keeps several speakers apart and names who said what. It also fixes the other caption-recording defects found while reproducing bigbluebutton#19700, and adds regression tests for all of them.
Before: akka-apps kept one running transcript per locale for the whole meeting and recorded every interim result as an
EditCaptionHistoryEvent(a character-index edit of that shared text). Whenever two people spoke the same language, each interim result of one speaker committed the other speaker's partial sentence and appended the whole running sentence again, with no separator and no speaker. Example from the dev server, two overlapping speakers:After (same scenario,
caption_en-US.vttof the published recording):Both recordings hold 37 caption events: the new event costs no more events than the old one.
How
CaptionUpdatedEvent(moduleCAPTION) from what the live view already has: one event per update of a caption segment (captionId,userId,locale,captionType, the wholetext,isFinal), fromTranscriptUpdatedEvtMsg(audio transcription) andCaptionSubmitTranscriptEvtMsg(typed / plugin captions). This is the "properly designed event" Recording "TranscriptUpdatedEvent" is useless and spammy. bigbluebutton/bigbluebutton#19701 asked for.AudioCaptions,EditCaptionHistoryEvtMsg, the deadTranscriptUpdatedRecordEventand the unusedtranscript.words/linessettings are removed.gen_webvttreplays each segment on its own (per-code-unit timestamps from the first appearance of each character), merges the cues of a locale in time order and names the speaker whenever it changes. Late interim updates after a segment's final are ignored. Recording pauses are applied per segment. Recordings with the legacyEditCaptionHistoryEventtake the old code path: output is byte-identical for all 60 legacy recordings on the test server.anonymize_chat,anonymize_chat_moderators,meta_bbb-anonymize-chat*) through a sharedBigBlueButton.generate_webvtthelper used by the captions step and by the presentation, video and screenshare formats.Other defects fixed on the way
captionSubmitTranscript, e.g. from caption plugins) were never recordedgetRecordingTextTrackslinkscaptions_<lang>.vtt, the recording wrotecaption_<lang>.vtt" well"lost)rap-caption-inbox.rbwrote toprops['presentation_dir'], a key that never existed (b88b70d, 2019)putRecordingTextTrackupload raisedENOENT, the handler crash-looped and stayed failed, later uploads piled upscripts/captions/, packages installscripts/caption/presentation(983751c, 2019)caption/presentationlackedrequire 'json'; its sort discarded the resultCloses Issue(s)
Closes bigbluebutton#19700
Closes bigbluebutton#19701 (the event is replaced by
CaptionUpdatedEvent)Closes bigbluebutton#19178 (already fixed by the 4.0 locale validation; this adds the regression test)
Closes bigbluebutton#19699 (cannot happen on 4.x: one
bn-document__<meetingId>pad, no_cc_pads; verified in the notes + captions scenario)Motivation
bigbluebutton#19700 asks for a new caption recording design because the Flash-era
EditCaptionHistoryEventcannot represent several streams per language. The live side already stores captions per utterance and per user (captiontable). Recording exactly that keeps the recording consistent with what participants saw and removes a meeting-long, quadratically growing text buffer from akka-apps.How to test
Regression tests (all fail on the current code, pass with this PR):
Manual: enable
public.app.audioCaptions.enabled, record a meeting where two people with Chrome speak English at the same time, end it, then check the caption track in the presentation playback, or fetch it viagetRecordingTextTracks.Upgrade note: live tracks of recordings processed before this change are still named
caption_<lang>.vttin/var/bigbluebutton/captions/<recordID>/. A one-off rename (innew-features.md) makes theirgetRecordingTextTrackslinks work. Rebuilding is not needed.Commits
23eb9e5ad0test: regression tests for live captions in recordings (Tracking issue for automatic live captioning recording problems bigbluebutton/bigbluebutton#19700)6ee4f447b5feat(record-and-playback): caption tracks per utterance, labelled with the speakerc8151e5859feat(akka-apps): record captions per segment instead of one text per locale233f3abf23fix(html5): send no interim transcript after the final one76887c293cfix(record-and-playback): uploaded caption tracks never reached the playbackf0a3c4712fdocs: recorded captions name the speaker, caption track API and uploads fixedVerification
Tested on a 4.1 dev server (4.1.0-alpha.1 packages, components built from this branch: akka-apps, html5 client, record-and-playback).
Playwright
captionsspec, A/B (the same spec, only the build switched):On stock, the recorded-track tests fail first on the 404 of
getRecordingTextTracks. With only the record-and-playback part of this PR on a stock akka-apps and client, they fail on the content itself, e.g."the birch canoe slid on the smooth planks" should appear exactly once in the recorded captions: thegluethe birchglue thethe birch canoe.... With this PR all 7 passed against the pushed head; the six tests other than the upload one also passed in 4 earlier runs, the upload test in 1 earlier run.gen_webvttunit tests: 11/11. The new script writes byte-identical output to the old one for all 60 legacy recordings on the server.Alice: ...,Viewer 1: ...).de-DEupload is published next to the liveen-UStrack for a recording made with this PR, and for one processed before it.recording,audio,options/audioProcessing*andparameterssuites,@ci, on this branch: 150 passed, 5 failed. 3 are@flakyand fail identically on stock. The other 2 pass alone on both builds; they failed in the full run on atest-resultscollision with a parallel run.tscand Playwrighttsc+ eslint clean. akka-apps compiles; its test tree does not compile on this line already (stale fixtures), so there is no Scala unit test.More
merge-v4.0.x-release-into-v4.1.x-develop, 6b9cf3d), which brings the locale validation and client changes these files build on. Until that merge lands, the diff againstv4.1.x-developincludes it; this compare shows only this series.Found while working on this, not changed here:
CaptionUpdatedEventis written for every interim result, with the whole segment text. Recordings have the same number of caption events as before, at 20–27% more bytes (two overlapping speakers: 12.8 → 15.5 kB, one speaker with 12 sentences: 35.6 → 45.3 kB). Recording only the first and final update per segment would shrink it, at the cost of per-word cue timing.UpdateTranscriptPubMsgHdlrnever drops a transcript: closing the audio modal still yields a voice user, and "Leave audio" leavesuser_voice.joined=true, muted=true.caption.captionIdis a global primary key, and the upsert ignores a captionId already used by another user, in any meeting. Real client ids are<userId>-<ms>, so this needs a buggy or hostile client. The spec keeps its ids unique.caption/presentation's branch for playbacks withoutcaptions.jsonwrites to''and crashes. Only pre-2.x playback formats reach it.subtitlesare listed by the API but not shown by the presentation playback (existing TODO incaption/presentation).record-and-playback/deploy.sh(developer script) deletescore/scriptsand restores only the presentation format, which removes the packaged video/screenshare scripts.