Redundancy and Backup Connectivity for OB Vans and Live Events

Every serious live production has a backup. The transmissions that fail on air are rarely the ones that had no second path. They are the ones where a second path existed and turned out to share something with the first: the same mast, the same operator, the same generator, or the same engineer who had to notice and act.

A truck at a showground carries two connections, and both modems are looking at the same mast three kilometres away because it is the only site serving that valley. When maintenance takes that mast down at four in the afternoon, the production has one connection. It always had one connection. The second was a duplicate rather than an alternative.

This article covers what redundancy has to mean when the output is live: which failures actually happen at events, why some backups do not take over, what a failover looks like on air, and what to test before the day rather than during it.

Live changes what failover has to achieve

In most of IT, failover measured in seconds is a good result. A branch office that switches to a backup circuit in twenty seconds has had a good day, because users retry, sessions re-establish and the working day continues. Business continuity is built around recovering quickly from an interruption.

Live output has no equivalent tolerance. Twenty seconds of black during a goal is the incident, and no recovery afterwards undoes it. The recovery window is zero, which changes the design goal: the objective is continuity through the failure rather than recovery after it. That single difference explains most of what follows, and it is why a redundancy design borrowed from an office network usually disappoints the first time it is called on.

It also changes what counts as a success. A backup that restores the feed in eight seconds is a functioning backup and a failed production. The measure is how much of the failure the audience sees.

Redundancy that shares a failure domain is duplication

A second path adds protection only against failures the first path does not share. Most disappointing backups fail this test somewhere, and the shared component is usually invisible from the truck.

Two SIMs from the same operator are one network with two subscriptions. Two operators can still land on the same physical mast, because rural sites are frequently shared between operators, and a site outage takes both. Two connections from different operators can meet again in the same backhaul or the same fibre duct on the way out of the area. Two modems in one encoder are one device. Everything at the venue depends on one generator or one distro. And a design that needs a person to notice and switch depends on that person being free at the moment of failure, which is exactly when they are least likely to be.

The practical test is to name the layer each path shares with the other and decide whether that shared layer is acceptable. Perfect independence is rarely achievable at a temporary site. Knowing precisely where the paths converge is achievable, and it is what separates a redundancy design from a hopeful one.

This is the real argument for network diversity across several carriers in one encoder, covered in the batch 1 article on multiple carrier SIMs in one encoder. Weconnect provides non-steered access across 700+ carrier partnerships in 195+ countries, so an encoder selects the networks actually performing at that location rather than a fixed pair chosen in advance, which keeps the paths from collapsing onto one operator when conditions change.

What actually fails at a live event

Redundancy budgets get spent on the failures people imagine rather than the ones that happen. The list below is closer to the operational reality at temporary sites.

Power comes first, more often than connectivity engineers expect. Generators, distro, a cable pulled by a forklift, or a UPS that was never load tested against the actual rack. Everything downstream inherits it.

Human error and patching come second. A cable moved during a reset, a setting changed during troubleshooting, a modem left on a test configuration from the morning rig.

Congestion is third and is specific to events. Public network capacity at a venue collapses when the crowd arrives, and the uplink direction suffers first, which is the direction contribution depends on. That mechanism is covered in detail in the article on remote and live sports production.

Then come mast and backhaul faults, weather (rain fade on satellite, wind on temporary antenna mounts), and hardware. Hardware failure is last on the list and first in most people’s planning.

Two shapes of failover, and only one of them is invisible

Backup arrangements behave in one of two ways when the primary path degrades, and the difference is what the audience sees.

Continuous degradation is what bonding across diverse networks provides. Every path carries part of the load, so losing one reduces available headroom instead of triggering an event. There is no switch, no reacquisition and no gap. The picture gets closer to its floor and stays on air, which is why the encoder ladder should be built with deliberate headroom rather than sized to the capacity measured during a quiet rig.

A standby switch behaves differently. The second path sits idle, something detects the failure, and the service moves. Every stage costs time: detection, modem attach, session establishment, decoder resynchronisation. Seconds is a good outcome, and seconds are visible. Automatic failover is far better than a manual switch here, because it removes the slowest component in the chain, which is a person realising what has happened.

The practical design uses both. Continuous degradation for the programme feed, where any gap is unacceptable, and standby paths for services that tolerate a brief interruption. Mixing bearer types strengthens it further, since cellular and satellite fail for unrelated reasons, as set out where bonded cellular and satellite are compared.

The return path is the redundancy nobody budgets for

Talkback consumes almost no bandwidth and ends more productions than bandwidth ever does. A director who cannot talk to camera operators has lost the show while every video feed is still arriving perfectly.

The exposure comes from carrying comms on the same link as the video, which means comms inherits every failure the video path has. Giving the comms circuit its own bearer, on a different network from the contribution feeds, is inexpensive insurance because the bandwidth involved is trivial. The same applies to programme return and cues, which the on-site crew depends on to do their jobs at all.

The far end is half the chain

A venue can be designed to a high standard and still sit behind a single point of failure at the receiving end. The decoder, the ingest path, the teleport and the studio’s own connection are all part of the same chain, and a production protected at one end and exposed at the other is protected on paper.

This matters more in remote production, where the gallery is somewhere else and the link is part of the production chain rather than a delivery route for a finished programme. When the venue holds only cameras and a small crew, there is no finished version of the event sitting anywhere to fall back on, and every element between the camera and the gallery carries the full consequence of a failure.

Designing the layers

The design is easier to hold in your head layer by layer. Each row below is a place where a live production loses its broadcast connectivity, and the redundancy that genuinely covers that layer rather than appearing to.

Layer What fails Practical redundancy
Power Generator, distro, a pulled cable Second supply, UPS load tested against the real rack
Access network Mast outage, crowd congestion, coverage hole Multiple carriers in one bonded encoder, non-steered selection
Bearer type A failure mode common to all cellular paths A satellite path bonded alongside, or held as standby
Encoder Device or software failure Second encoder, split camera feeds across both
Comms and return Shares the video path and inherits its failure Separate bearer on a different network, minimal bandwidth
Receiving end Decoder, ingest, teleport, studio connection Redundant receive path, tested end to end
People Nobody free to notice or act mid-event Automatic failover, alarms routed to someone who can act

Test the failure, not the setup

A successful rig test proves the primary path works. It proves nothing about what happens when the primary path stops, which is the only scenario the redundancy exists for.

Rehearse the failure explicitly. Pull a modem during the run-through and watch what the picture does. Kill the primary path and time how long the switch takes end to end, including decoder resync at the far end. Confirm that the alarm reaches a human who is in a position to act, rather than an inbox nobody is watching during the event. Where the schedule allows, do the capacity check at a comparable time on a comparable event, because a measurement taken during a quiet morning tells you almost nothing about conditions at the moment of highest demand.

Then write down what the crew does when it happens, because the fallback that exists only in the head of the engineer who built it stops existing the week they are covering a different event.

Across a fixture list, redundancy becomes an inventory problem

A production company covering a season is at a different venue every week, often in a different country, for eight or nine months. That is the normal shape of sports broadcasting connectivity, and it turns redundancy from a per-event design into a question of whether every kit set going out of the door is still configured the way it was intended.

Central visibility is what makes that manageable: seeing which SIMs carried traffic at the weekend and which quietly carried none, spotting the modem that has been failing over silently for three fixtures, and confirming the spare path in the flight case is live before it is needed. Weconnect makes every SIM visible from one platform across every country on the calendar, which turns a failed backup into something noticed on Monday rather than during the next event.

Frequently Asked Questions

What does redundant connectivity mean for a live broadcast?

It means a second path that fails for different reasons from the first. Two connections that share a mast, an operator, a backhaul route or a power supply protect against fewer failures than they appear to, because a single fault takes both. Useful redundancy is defined by what the paths do not have in common, and the first step in any design is naming the layer where they converge.

Is bonding the same as having a backup connection?

Bonding provides a stronger form of protection than a standby link for the programme feed. Every connection carries part of the load, so losing one reduces available headroom instead of triggering a switch, and there is no gap while a backup takes over. A standby link still has a place for services that tolerate a brief interruption, and for adding a different bearer type such as satellite that fails for unrelated reasons.

How fast does failover need to be for live television?

Fast enough to be invisible, which in practice means the programme feed should degrade rather than switch. A standby switch involves detection, modem attach, session establishment and decoder resynchronisation, so seconds is a realistic best case and seconds are visible on air. Automatic failover is substantially better than a manual switch because it removes the slowest step, which is a person noticing during the busiest moment of the event.

What is the most common cause of connectivity failure at a live event?

Power and human error account for more lost transmissions than network faults, followed by crowd congestion at the venue. Public network capacity drops sharply once the audience arrives, and the uplink direction is affected first, which is the direction contribution feeds depend on. Hardware failure is the scenario most often planned for and among the least frequent in practice.

Do you need a separate connection for talkback and comms?

On any production where a comms failure would end the show, yes. Talkback uses very little bandwidth, so a dedicated bearer on a different network is inexpensive, and it removes the risk of comms inheriting a failure in the video path. Programme return and cues deserve the same treatment, since the on-site crew depends on them to work at all.

Next steps

Weconnect provides the broadcast connectivity layer of that design: non-steered multi-network SIMs across 700+ carrier partnerships in 195+ countries so paths do not collapse onto one operator, automatic failover, consumption-based data that follows the fixture list, and every SIM visible from one platform. Tell us which venues have caused problems, what the productions carry beyond the programme feed and how much notice you get, and we will map coverage per location and build the redundancy around it. Challenge us with your connectivity requirements. Direct response within 4 business hours.

Share