How ReCAP Tags Artificial Lighting Changes in Video Scenes

Video lighting rarely stays perfectly stable. A studio may switch from a presenter’s key light to a cooler interview setup, an outdoor match can move under floodlights at dusk, or a replay may introduce a different colour temperature from the live feed. For broadcasters and media libraries, these transitions contain useful information about how a scene was produced and how it should be searched later.

ReCAP approaches this problem as part of real-time content analysis and processing. Its wider platform is designed to extract structured metadata from broadcast-quality video, helping production teams and media asset managers work with large volumes of footage. The ReCAP project provides the broader context for these capabilities, including analysis, monitoring and content-management applications.

The important distinction is between a lighting change and a general image change. A new camera angle, a fast cut or a moving crowd can alter many pixels without changing the illumination of the scene. ReCAP-style analysis therefore combines visual measurements, temporal patterns and scene context before assigning a tag such as artificial-light transition, LED colour shift or low-light production setup.

Reading Light As A Temporal Signal

The first stage is to examine how brightness and colour develop across consecutive frames. Rather than treating every pixel equally, the system can calculate features such as average luminance, contrast, colour temperature, saturation and the proportion of highlights or shadows. These measurements create a compact description of the scene’s illumination over time.

A gradual fall in luminance across the whole frame may indicate sunset, a dimmed studio, or a camera exposure adjustment. A sudden increase in brightness concentrated around faces and foreground objects may suggest that a spotlight has been switched on. Colour channels are equally informative: a move towards blue or green can reveal a changed LED fixture, while a warmer cast may come from tungsten lighting or a practical lamp entering the shot.

Temporal analysis helps separate lighting events from ordinary motion. A camera pan changes the composition, but its global brightness pattern may remain relatively stable. A lighting transition affects a broad region or the entire image and often persists for several seconds. The detector can therefore look for a sustained change point rather than reacting to a single frame.

Shot boundaries must be considered at the same time. A cut from a daylight exterior to an artificially lit studio will produce a dramatic difference, yet the event is not necessarily a lighting change within one scene. By combining shot-boundary detection with illumination tracking, the system can tag the new scene’s lighting conditions without confusing an editorial cut with a fixture adjustment.

Distinguishing Artificial Illumination From Natural Change

Artificial lighting is not identified by brightness alone. The analysis considers whether the visual pattern is consistent with controllable production equipment. Stable, localised highlights, repeated colour shifts and changes that follow a stage or venue layout can all support an artificial-light interpretation. A broad change in colour across the sky, by contrast, is more likely to be natural illumination.

The detector can compare different parts of the image to improve this judgement. If the floor, players and advertising boards brighten together while the sky remains comparatively unchanged, floodlights are a plausible explanation. If the entire frame gradually becomes warmer as the sun lowers, the event may be tagged as a natural-light transition. Shadows can provide another clue: artificial sources often create a new direction or sharper edge, while cloud cover tends to soften illumination across a wider area.

Moving subjects create a difficult edge case. A person walking beneath a spotlight may cause a bright patch to travel through the image, even though the lighting system itself has not changed. Object tracking, face detection and region-level comparisons help prevent this from becoming a false scene tag. A large illuminated area that remains fixed relative to the set is more significant than a highlight attached to a moving shirt or vehicle.

Australian production conditions make this distinction especially practical. A cricket broadcast in Melbourne can move from bright afternoon sunlight to stadium floodlights within one session, while a night AFL fixture in Sydney may combine strong field lighting with changing LED advertising panels. In Queensland, where daylight-saving schedules differ from New South Wales and Victoria, broadcast systems also need to rely on visual evidence rather than assuming that a particular clock time implies a particular lighting state.

Turning Visual Evidence Into Reliable Tags

Once a possible event is detected, the system can assign several metadata fields instead of one vague label. These may include the timecode of the change, duration, direction of brightness movement, dominant colour shift, affected image area and confidence score. A record might describe a scene as “artificial lighting introduced, cool white, gradual transition” rather than simply “dark” or “bright”.

Confidence is important because broadcast footage contains many conditions that resemble illumination changes. Camera iris adjustments, auto white balance, compression, screen reflections and rolling exposure can all distort image statistics. A robust pipeline compares multiple signals before creating a high-confidence tag, then retains lower-confidence events for review or downstream filtering.

The tagging layer can also distinguish between a lighting state and a lighting transition. “Artificially lit interior” describes the scene after analysis, whereas “warm-to-cool light change” describes what happened during the sequence. Both are useful: an editor may search for all interviews recorded under cool studio light, while a quality operator may need to find every abrupt change that could indicate a faulty fixture or an inconsistent camera feed.

ReCAP’s wider technical direction supports this kind of structured interpretation. The project’s project objectives describe a focus on automated metadata extraction, quality monitoring and analysis tasks that can operate within media workflows. Lighting tags can therefore sit alongside face, logo, duplication and quality metadata rather than existing as an isolated computer-vision result.

Video evidence Likely interpretation Useful metadata tag Main risk of error
Whole frame brightens and the change persists A light source has been raised or switched on Artificial light increase Camera exposure adjustment
Image becomes cooler around faces and set surfaces LED or colour-temperature change Cool artificial lighting Auto white balance
Foreground brightens while sky remains stable Floodlight or local fixture activation Localised artificial illumination Moving spotlight or reflection
Frame gradually warms during an exterior shot Natural sunset or changing daylight Natural-light transition Delayed camera colour correction
Abrupt difference at a shot boundary New scene with different lighting Lighting state after cut Incorrectly tagging a within-scene event
Repeated pulsing across several frames Flicker, faulty fixture or capture issue Lighting instability LED display or compression artefact

Supporting Broadcast And Asset Workflows

In a live control room, lighting tags can help operators find the beginning of a technical change without watching an entire programme. A confidence threshold may trigger an alert for a sudden flicker, while a slower transition can be written quietly into the event log. The same metadata can support post-production, where an editor needs to locate all scenes with matching illumination before assembling a package.

For media asset management, searchable lighting information adds another route through an archive. A producer looking for a night-time interview, a stage performance with coloured LEDs or footage recorded in a dim venue can filter by lighting state, colour family or transition type. This is particularly valuable when filenames and production notes are inconsistent, a familiar issue in large collections assembled from outside crews, agencies and live feeds.

Australian broadcasters manage a wide range of locations, from major studio facilities in Sydney and Melbourne to regional venues and outdoor events. A lighting-aware index can help compare footage captured under very different conditions, including bright coastal exteriors, enclosed sports arenas and remote productions with limited crew access. It can also support faster review of material from long-running events, where a lighting change may mark a meaningful production phase.

The tags are most useful when they remain interoperable with existing production systems. Timecodes should align with the source media, event descriptions should use consistent vocabulary, and confidence values should be preserved when metadata is exported. A human operator can then correct an uncertain label without losing the original visual evidence or the reason the system created the tag.

Handling Difficult Scenes And False Positives

Reflections are one of the main sources of confusion. A moving car window, glossy studio floor or large video wall can produce a broad brightness change that resembles a light being switched on. Region masks and surface-aware models can reduce the effect, especially when the suspected change appears only on reflective material rather than on faces, backgrounds and fixed objects.

LED displays create another challenge in Australian sports and entertainment venues. Advertising ribbons, scoreboard animations and perimeter screens can flash or change colour while the physical lighting remains constant. A detector should examine whether the change is confined to known display regions. When it is, the event may deserve a screen-change tag rather than an artificial-illumination tag.

Low bitrate feeds and camera processing can also produce misleading results. Noise in dark scenes may make a small brightness variation appear significant, while automatic gain control can brighten an entire image without any change to the venue lighting. Quality-monitoring signals, including blockiness, noise level and exposure behaviour, help the system lower confidence when the source itself is unstable.

Human review remains useful for unusual cases, especially during early deployment. Reviewers can inspect representative clips, correct tags and identify recurring patterns at a particular venue or with a particular camera model. Those corrections can refine thresholds and training data without turning the workflow into a fully manual process. The aim is to make review selective: people examine ambiguous events while routine lighting states are processed automatically.

From Detection To Useful Metadata

A scene tag becomes valuable when it describes an event in language that another system can use. A practical record could include the start and end time, scene identifier, lighting category, colour direction, spatial coverage, confidence and supporting quality indicators. It might also link the event to a detected logo, face or programme segment, allowing users to search across several dimensions at once.

This layered approach avoids making a single visual judgement carry too much meaning. “Artificial light detected” can be an initial machine observation; “studio interview under cool LED lighting” can be a richer interpretation after scene classification and context checks. Keeping these levels separate makes the metadata easier to audit and safer to reuse in broadcast automation, archive search and quality reporting.

For a live production, the immediate benefit is situational awareness: a sudden, persistent change can be identified near its time of occurrence. For an archive, the benefit is discoverability: a large collection can be filtered by illumination conditions without relying on someone to remember which reel contains a particular look. Across both settings, time-aligned metadata gives editors and operators a practical way to move from visual evidence to an operational decision.

The next step is to run a representative sample containing daylight exteriors, studio interviews, sports floodlights, LED displays and camera exposure changes, then compare the generated lighting tags with a frame-accurate human annotation set.