Was it us or them? Answering “the live froze for me” with evidence
A viewer writes in after your live event. The video froze, and for a few minutes they couldn’t hear the host.
Support forwards the ticket to engineering, and the question is simple:
Was it us, or was it them?
For many live-video teams, answering that question takes far longer than it should.
Why the question is hard
By the time the ticket arrives, the moment is gone. The event ended an hour or a day ago, and nobody can simply go back and watch exactly what that viewer experienced.
Server-side logs can tell you a lot about what happened inside the media infrastructure. What they often don’t tell you is what one particular viewer actually received, and how their device behaved while receiving it.
Room-level dashboards have a similar problem. They can tell you that quality was healthy overall. But the complaint is not about the room overall. It is about one person.
So the investigation turns into a process of elimination. The video team checks the stream, the app team checks the client, someone looks at the host’s setup, and eventually the answer may become: “Probably your Wi-Fi.” Sometimes that is correct. Sometimes nobody really knows.
What that uncertainty costs
One support ticket can pull several people into an investigation, because every team first has to rule itself out.
Support is then left with two bad options: apologise for something that may not have been your fault, or blame the customer’s connection without enough evidence.
There is another group you never hear from at all: viewers who had a bad experience and simply left. That matters, because one person’s terrible session can disappear almost completely inside a healthy room average. We saw exactly that in production.
Compare people, not just averages
Instead of looking at the affected viewer in isolation, compare them with everyone else who was receiving the same media at the same moment.
If one viewer says the host’s voice broke up at 8:14, dozens of other people may have been listening to the same host at 8:14. What happened to them gives you a reference point, and that usually lets you narrow the problem down to a few broad categories:
- The source or stream had a problem. Many viewers receiving the same host deteriorated at the same time.
- The viewer’s connection had a problem. One viewer deteriorated while other people receiving the same streams remained healthy.
- The viewer’s device had a problem. Network delivery looked healthy, but device-side signals showed that the phone could no longer keep up.
From there, additional signals can narrow the cause further: a broadcaster’s upload reaching its limit, an unstable viewer connection, thermal throttling, decoding pressure, or something else. Different causes mean different fixes, and very different replies to the customer.
Two real cases
Both examples below come from production rooms and have been anonymised. Interestingly, nobody filed a support ticket in either case. We found them by looking for the participants with the worst experience. That matters because the people with the worst experience are not always the ones who complain.
The room that looked mildly unhealthy
Nineteen viewers watched three people on camera for 49 minutes. Across the room, average video packet loss was around 1%. That number is not especially alarming. Looking only at the aggregate, you might assume the room had a small quality problem shared by everyone.
It didn’t.
One viewer, whom we’ll call Maya, experienced around 9% video packet loss and repeated periods of blocky or stalled video for roughly 14 of those 49 minutes. At the same moments, the other viewers watching the same cameras were largely healthy. Twelve of the other eighteen viewers had no bad stretches at all.
Remove Maya from the calculation and the room average falls from roughly 1% to 0.17%. One person’s bad session had been spread across the whole room by the average. The comparison strongly pointed to a viewer-side network problem rather than a shared stream or platform issue.
The host whose voice broke up
In another room, two people were on camera and 37 viewers were watching. About twenty minutes into the session, one host’s audio deteriorated for 36 of the 37 viewers at almost the same moment. The co-host remained clear.
That pattern immediately rules out many viewer-side explanations.
The broadcaster’s own telemetry filled in the rest. A few seconds before viewers noticed the degradation, the host’s phone reported that its upload had reached the available connection limit, and it dropped to a much lower video quality to compensate. The device itself was not under significant computational pressure.
Taken together, the evidence pointed to the host’s upstream network as the source of the problem. A shared server-side failure became much less likely because the co-host remained healthy, and individual viewer connections could not explain why almost everyone deteriorated at the same moment.
Same symptom to the viewer. Completely different cause.
If you want the more technical versions of these investigations, including the signals and edge cases we had to rule out, see One viewer was losing 9% of her video while the room average looked fine and Stream, network or phone? Three real live-stream incidents, taken apart. The second article also includes a third case: a phone that overheated and began discarding a large part of the video it received.
Questions to ask your team this week
None of these require Rewitness.
- When somebody reports a bad session, can you tell reasonably quickly whether the problem was shared or isolated to that participant?
- Do you collect enough telemetry from the viewers who did not complain to compare them with the person who did?
- Can you line up what different viewers experienced from the same host at the same point in time?
- When support tells somebody “it was your connection”, what evidence is that conclusion based on?
If the answer to some of these is “no” or “we’d have to investigate”, that is not unusual. Most monitoring systems naturally grow around infrastructure, servers and room-level metrics. Those are important, but they answer a different question.
What we built
We’re building Rewitness around this idea. A small SDK in your web or React Native app records the media, network and device signals needed to reconstruct what was happening to each participant during a live session. Every participant becomes another witness to the same event.
Rewitness aligns those accounts on a single timeline and compares them using a deterministic diagnostic engine. The diagnosis is based on explicit rules and observable evidence, so every conclusion can be traced back to the signals that produced it. The goal is to answer two questions: where did the problem start, and what evidence supports that conclusion?
When a complaint comes in, you can identify the participant and the approximate moment and see the diagnosis for that part of the session. If you want to investigate further, you can move through the reconstructed timeline and see who was on stage, what different viewers were receiving, which participants were struggling and what signals changed at the same time.
There are limits. Rewitness does not currently read your media server’s internal logs, so some server-side failures have to be inferred from their effect across participants. It can also only observe applications that include the SDK.
We think being explicit about those limits matters, because the point is not to produce another confident guess. The point is to show enough evidence that your team can verify the answer themselves.