ReCAP's approach to speaker diarization in panel discussions

When four people share a desk and a microphone, the broadcast looks tidy on screen but the audio underneath is anything but. Hosts interrupt each other, guests pause mid-thought to consult notes, a panel moderator in Sydney welcomes a remote expert joining from Adelaide, and somewhere in the background a studio audience reacts in waves. Turning that mess into a clean, searchable record of who spoke when is the job of speaker diarization, and panel discussion formats are the hardest test case in mainstream broadcasting.

ReCAP, an EU-funded research initiative hosted at recap-project.com, has spent the past two years building broadcast-grade video analysis tools that include this very task. The project pulls together partners working on face recognition, logo detection, video quality monitoring and near-duplicate detection, and speaker diarization sits as a core metadata layer underneath all of them. For Australian broadcasters facing the same problem on programmes like the ABC's Q&A, Network 10's The Project, or SBS's multilingual current affairs, the project's pipeline offers a practical reference point for what is achievable in 2025.

Why panel discussions break simple voice tools

Conventional voice activity detection was designed for two-state audio: someone is talking, or nobody is. That works for a single newsreader behind a desk, but a four-person panel produces a third state that most off-the-shelf systems cannot handle cleanly: two people are talking at once. Add the laughter, applause, and overlapping interjections that come with a live audience, and a single timeline quickly becomes a tangled waveform that no simple energy threshold can resolve.

ReCAP's team treats each panel as a unique acoustic scene. The first pass segments the audio into short windows and asks the simplest possible question for each: is this window dominated by a single voice, by multiple voices, or by non-speech noise such as applause or studio music? That tri-state classification is the foundation everything else sits on, and getting this layer right reduces downstream errors across the whole diarization chain.

A second complication is technical: panel discussions rarely use a tidy multi-microphone setup. Guests share lavalier mics, moderators roam with handheld units, and remote contributors join over IP codecs that introduce compression artefacts. In Australian studios, where productions often travel between Sydney, Melbourne, and sometimes a regional hub in Brisbane, the equipment changes from week to week, and so does the room acoustics.

Layered audio analysis inside ReCAP

Once the audio is segmented into single-speaker and overlapping windows, ReCAP runs a speaker embedding model on the single-speaker segments to produce a vector representation of each voice. These embeddings, also called speaker representations or voiceprints in the literature, encode the acoustic signature of a voice in a form that can be compared mathematically. Two segments whose embeddings are close together in this multi-dimensional space are likely to come from the same person; segments that are far apart are likely to be different speakers.

Clustering these embeddings produces the initial speaker labels. The project uses an agglomerative clustering approach with a calibrated similarity threshold, refined over several passes. Each pass reassigns segments to the closest cluster and recalculates the cluster centroid, gradually tightening the boundaries between voices. A fifth panellist who only speaks for thirty seconds near the end of the show is hard to lock in this way, so ReCAP adds a final reconciliation step that uses both acoustic similarity and visual cues from face detection to confirm the identity of low-speech contributors.

Visual cues matter because, in panel discussion footage, the camera usually shows whoever is speaking. ReCAP's face detection module identifies each visible face in every frame and tracks the face over time, producing a face track for each on-screen identity. The system then correlates the timing of those face tracks with the timing of the audio clusters. If an unidentified audio cluster coincides with a face track that the system has not yet labelled, the two are linked. This cross-modal matching is what lets the system recover the identity of a quiet panellist whose face appears only briefly between cuts.

Australian content shapes the model

The ReCAP consortium is European, but the team has gone out of its way to include non-European training material, and that includes Australian broadcast content. The reason is plain: Australian English has acoustic properties that differ from British and American baselines. The broad Australian accent has distinctive vowel shifts, and the general and cultivated accents each carry their own rhythmic patterns. A model trained only on northern-hemisphere news readers would stumble on a panel that mixes a Melbourne-based journalist with a Perth-based editor and a remote Indigenous affairs correspondent.

Multilingual programming complicates the picture further. SBS, Australia's multicultural broadcaster, regularly produces panel discussion content in English, Mandarin, Vietnamese, Italian, and Arabic, sometimes within the same half-hour show. ReCAP's diarization module is language-agnostic at the embedding level: the speaker embeddings describe voice quality rather than the words being spoken, so the same pipeline can separate a Mandarin-speaking guest from an English-speaking host without needing language-specific acoustic models. The same approach transfers cleanly to Australian multicultural formats.

Parliamentary broadcasts add another layer. The Australian House of Representatives in Canberra streams every sitting day, and committee hearings often run as panel-style discussions between MPs, departmental secretaries, and external witnesses. For the parliamentary record office and for media monitors, accurate speaker diarization on these streams is not a convenience but a compliance requirement. A misattributed quote can have legal consequences under Australian defamation law.

What it means for broadcasters and archives

For an Australian production team, the practical output of speaker diarization is a structured metadata file that travels with the video. That file maps every spoken segment to a speaker label, with timestamps, confidence scores, and links to the matching face tracks. Once that file exists, a number of downstream workflows become cheaper and more reliable, and for media organizations weighing how to fund specialist coverage, a paid channel setup is one route to making premium archive access self-sustaining.

Ways Australian media teams are already using this kind of metadata:

The same metadata also helps with content reuse. Australian broadcasters sit on vast tape libraries from the 1970s and 1980s, and modernising those archives is an active project at the National Film and Sound Archive in Canberra. Even when only the audio survives in watchable quality, speaker diarization on the soundtrack can recover who was speaking, which is often enough to make a clip findable.

Where the metadata goes next

ReCAP's broader architecture treats speaker diarization as one node in a larger content-analysis graph. The same graph contains the face recognition module, the logo and on-screen text recognition module, the video quality monitor, and the near-duplicate detection module. Once all of these are running on the same piece of content, the resulting metadata can be cross-referenced in ways that a single module cannot.

Cross-referencing scenarios worth noting:

This is also where the project's own research notes become relevant. The team's welke-slots-hebben-hoge-volatiliteit page covers variability in the data layer in terms that map closely to what ReCAP sees in long-form panel content: not every minute of a two-hour broadcast is equally important, and the system needs to know which segments to spend its analysis budget on.

The practical next step for any Australian broadcaster is small and concrete: pick one regular panel programme, run the existing ReCAP pipeline over the last twenty episodes, and compare the system's speaker labels against the official transcripts. That single comparison will reveal where the model breaks on local accents, where studio acoustics are degrading the embeddings, and how the face tracks line up with the audio clusters. The result is a calibration baseline that any production team can use to decide whether the technology is ready for daily use or whether a few more months of fine-tuning are warranted.