How ReCAP Counts On-Screen Text Overlays in Broadcast Video
On-screen text is a constant feature of modern television and online video. News tickers, scoreboards, lower-thirds, captions, emergency notices, sponsor graphics and programme titles all add information to the picture. For viewers, these elements may seem temporary and easy to ignore. For broadcasters and media libraries, however, they are valuable signals that need to be detected, classified and counted accurately.
ReCAP addresses this problem through real-time content analysis and processing. Its technology is designed to examine broadcast-quality video, identify meaningful visual elements and create metadata that can support production, monitoring and archive workflows. Text overlays are one part of this wider analysis, alongside face and logo recognition, quality assessment and duplicate-content detection.
A reliable overlay counter can answer practical questions: how many graphics appeared during a programme, when a sponsor message was shown, whether a required warning stayed on screen for the correct duration, or how frequently a broadcaster used a particular visual template. The ReCAP project website describes the research initiative and its broader focus on making complex video collections easier to analyse.
This capability has clear relevance in Australia, where television, streaming and live sport operate across large distances and varied production environments. A news bulletin from Sydney, an AFL match in Melbourne, a regional weather update in Queensland and a live event from Perth may all use different graphic conventions. Automated analysis can provide a consistent way to find and measure those overlays at scale.
What counts as an on-screen overlay
An on-screen text overlay is visual information placed over the underlying video rather than recorded as part of the original scene. It may be a single word, a block of text, a moving ticker or a complete graphic package containing text, branding and imagery. Examples include a presenter’s name, a breaking-news banner, a time-and-score display, a translation, a location label or a content classification notice.
The number of overlays depends on the counting rule. A system may count every separate appearance, every continuous period on screen, every text region or every event involving a particular category. For example, a scoreboard that remains visible for 90 minutes might count as one continuous overlay event, while a production team could instead measure each refreshed score update as a separate change.
This distinction matters because text can appear in layers. A sports broadcast may show a scoreboard, a player identification graphic and a sponsor message at the same time. A television news segment may combine a permanent station logo with a lower-third, a live-location label and a scrolling headline. A useful analysis system should report these elements separately when the workflow requires it, rather than treating the entire frame as one undifferentiated graphic.
How automated detection works
The first stage is usually scene and frame analysis. Video is sampled at a suitable rate, and algorithms compare neighbouring frames to locate regions that remain visually distinct from the background or change according to a recognisable graphic pattern. Colour contrast, edges, transparency, position and movement can all help distinguish a text overlay from objects that naturally occur in the scene.
Optical character recognition can then identify the words or characters inside the detected region. OCR performance is affected by font size, animation, compression, language, background movement and transparency. Broadcast graphics are often cleaner than text filmed inside a scene, but they can still slide, fade, flash or appear briefly. ReCAP’s real-time processing focus is relevant here because the system must analyse incoming material quickly enough to support live or near-live use.
A strong workflow combines several signals rather than relying on OCR alone. Text recognition can confirm that a region contains writing, while temporal analysis determines when it appeared and disappeared. Template matching can identify a recurring lower-third design, and logo detection can connect the graphic with a broadcaster, advertiser or programme. This produces richer metadata, such as the overlay type, screen position, wording, duration and confidence score.
Counting also benefits from tracking. If the same banner moves across the screen or changes one line of text, the analysis should decide whether it is one event with updated content or several separate overlays. Tracking reduces duplicate counts caused by minor frame-to-frame variation and supports a more useful timeline for editors, compliance teams and archive managers.
Where the results support Australian media
Australian broadcasters can use overlay analytics during live production, post-broadcast review and long-term asset management. In a Sydney newsroom, a producer could search a live or recorded bulletin for all lower-thirds that mention a particular suburb. A sports production team covering the AFL in Melbourne or rugby league in Brisbane could verify the timing of sponsor graphics and compare their actual screen presence with the planned schedule.
The same approach can help regional and remote operations. A station serving communities across Western Australia, the Northern Territory or far north Queensland may have fewer staff available to review every minute of output manually. Automated counts can flag missing captions, unexpected banners or repeated graphics for a human operator to inspect. This is especially useful when programmes are distributed across television, catch-up services and social clips.
Australian media organisations also face a mixed distribution market. Traditional free-to-air services coexist with subscription platforms, broadcaster video-on-demand libraries, FAST channels and short-form publishing. The same programme may be repackaged with different warnings, sponsorship messages or promotional labels. Overlay metadata makes those differences searchable and can help asset managers understand which version of a video they are handling.
The method can support accessibility and regulatory review as well. Captions and audio-description notices may be detected as recurring graphic elements, while classification labels and audience warnings can be checked for presence and duration. The analysis does not replace editorial judgement or legal review, but it can reduce the amount of material that must be inspected manually.
| Overlay category | Typical example | Useful measurement | Possible workflow |
|---|---|---|---|
| Lower-third | Presenter name and role | Start time, duration, wording | News archive search |
| Sports graphic | Score, clock or player statistic | Frequency and screen position | Live production review |
| Ticker | Continuous news or market feed | Text changes and total runtime | Compliance monitoring |
| Sponsor message | Brand strap or promotional panel | Number of appearances and duration | Campaign verification |
| Warning label | Classification or emergency notice | Presence, timing and persistence | Editorial assurance |
| Caption or translation | Dialogue or language support | Coverage and timing | Accessibility review |
| Programme branding | Channel logo or segment title | Continuity across episodes | Asset management |
Counting accuracy and difficult video conditions
Counting text overlays is harder when graphics are semi-transparent, animated or placed over visually similar colours. A white headline on a pale sky, for example, may be less distinct than a high-contrast banner in a studio. Rapid camera movement, flashing lights and compression artefacts can create false detections. A watermark that remains on screen throughout a programme may also be treated differently from a temporary overlay.
Australian viewing conditions add practical variation. Content may arrive from high-quality studio feeds, outside broadcasts, mobile uplinks or older archive formats. Large events can involve multiple production suppliers, each with different graphics packages and standards. A system designed for broadcast-grade material still needs configurable thresholds and quality checks so that uncertain detections are clearly identified.
Accuracy should therefore be measured at several levels. Detection accuracy asks whether an overlay was found at all. OCR accuracy concerns the words that were transcribed. Temporal accuracy checks when the event began and ended, while classification accuracy asks whether the system correctly identified a scoreboard, sponsor graphic, caption or warning. Counting accuracy combines these factors into a result that can be used operationally.
Human review remains valuable for edge cases. A confidence score can send ambiguous frames to an operator instead of presenting every machine result as certain. Corrections made during review can also help refine templates and improve future analysis. In this model, automation handles volume and consistency while specialists deal with unusual designs, culturally specific references and editorial context.
Privacy, rights and responsible metadata
Text overlay analysis generally focuses on graphic regions, but video analysis often operates alongside face and logo recognition. That combination creates a detailed record of what appeared in a programme and when. Organisations using such metadata should define retention periods, access permissions and appropriate safeguards, particularly when footage contains identifiable people or sensitive events.
The Australian Privacy Act 1988 is relevant when personal information is collected, stored or used in a way that relates to identifiable individuals. A lower-third naming a private person, a location label connected with a vulnerable event or a screen capture retained for quality review may require careful handling. Public interest broadcasting does not remove the need for responsible governance, and organisations should consider the purpose and proportionality of each analysis task.
Rights management also matters. Sponsor graphics, channel branding and programme footage may be subject to contractual restrictions. Counting an advertisement for campaign reporting is different from extracting and redistributing the creative asset. Metadata should record what is needed for the workflow without creating unnecessary copies of protected content. Clear audit trails can help demonstrate how results were produced and who accessed them.
For Australian teams working with an EU-funded research project, international data practices may also be part of procurement and collaboration discussions. Technical documentation should explain where processing occurs, how video is secured and whether results are transferred between jurisdictions. These questions are especially important for broadcasters managing unreleased programmes, sports rights or sensitive news material.
Turning overlay events into useful metadata
The most valuable output is more than a total count. A searchable record might include the programme identifier, timecode, overlay category, recognised text, screen coordinates, duration, associated logo and confidence level. Editors can then find every instance of a sponsor message, while archive staff can locate a segment by the wording of its lower-third rather than watching the whole file.
Real-time alerts can extend this value during production. If a required emergency notice does not appear, a sponsor graphic runs for too long or an unexpected phrase is detected, an operator can be notified while the programme is still live. The threshold for an alert should be chosen carefully: too many warnings create fatigue, while an overly strict system may miss important events.
For asset management, counts can support deduplication and version comparison. Two files may contain the same underlying programme but differ in captions, classification labels, advertising or promotional straps. Overlay timelines help identify those differences quickly. They can also support highlight creation, compliance evidence, searchable research collections and faster preparation of clips for digital platforms.
The result is a practical bridge between computer vision and everyday broadcast work. A machine can inspect thousands of hours consistently, while producers and archivists use the resulting metadata to make informed decisions. The quality of that process depends on clear counting definitions, reliable timecodes, transparent confidence measures and a review path for uncertain detections.
The key point to remember is that counting on-screen text is really about understanding when, where and why broadcast graphics appear. ReCAP’s approach places overlay detection within a broader real-time analysis system, turning transient words and banners into structured evidence that Australian media teams can search, verify and use.