One viewer was losing 9% of her video while the room average looked fine
One of 19 viewers in a live video room lost 9% of her incoming video packets, which is severe. Averaged over the room, loss came to 1.1%, which reads as a mild problem spread across everyone. Take her out and the room average falls to 0.17%. The server was fine, and one person’s connection was the problem. Here is how we checked, and the traps we had to rule out first. The data comes from production and is anonymised.
The room
A few people were on camera or sharing a screen in a long, open room on a web client, and everyone else watched. We looked at a 49-minute stretch that starts when one viewer joined. We’ll call her Maya. During that stretch, 19 viewers, Maya included, sent WebRTC receive statistics, and three people were publishing video.
Nobody filed a ticket about Maya. We picked her because her data was the worst in the room. The complaint she would have written is easy to guess: the video kept going blocky and stalling, on and off, for most of the time she was there. That is the normal case: most viewers who have a bad session never file a ticket, they just leave. If Maya had complained, the answer would have been short: her own connection, not the server. Because she didn’t, the room average was the only thing anyone would have seen.
What the room-level view says
Most teams start from the room. A typical dashboard aggregates loss, round-trip time and connection quality over a room or an event.
Seen that way, this room looked a little unwell. Pooled across every viewer, inbound video loss for the stretch was 1.1%: mild, and easy to read as a small problem everyone shared. Look at the whole two hours after Maya joined and it gets worse: 11 viewers had some kind of trouble, and for 9 of them it was flagged as network trouble. When most of the people with problems have network problems during the same evening, the natural reading is a shared cause, such as an overloaded SFU or a capacity limit. The natural next step is to open the media server’s dashboards.
The weak point is the word “same”. Over two hours, nearly everyone in a room has some network wobble: a few seconds of poor connection-quality readings, a brief loss spike, a laptop changing Wi-Fi access points. Put nine sets of short, unrelated blips inside one long window and they look like one event. The viewers who had no trouble at all are missing from that count, and they are the ones who could show the server was fine.
What the comparison at the same moment says
The better question is narrower. When Maya’s video degraded, what were the other people watching the same publisher receiving at that moment?
For every viewer and every incoming track, we record whether media was arriving cleanly, arriving with more than 5% loss, or frozen, with timestamps. Across the camera and screen-share feeds she watched, Maya had 100 bad stretches in the 49 minutes, from about 4 to 48 seconds each (median 10 seconds). Many overlapped across feeds, so in wall-clock time her picture was bad for about 14 of the 49 minutes. For the 36 stretches of 15 seconds or longer, we checked who else was receiving the same feed at the midpoint. Between 8 and 13 other viewers were, and in 32 of those 36 cases every one of them was receiving it cleanly. Of the 18 other viewers, 12 had no bad stretches at all, five had five seconds or less, and one had about three and a half minutes.
Timeline grid with 19 horizontal lanes, one per viewer, across 49 one-minute columns. Cells are shaded where that viewer's inbound video packet loss exceeded 2%. Maya's lane is shaded in 36 of 49 minutes, spread across the whole period; one other lane is shaded in 11 minutes; the other 17 lanes have no shaded cells.
The 36 minutes in the figure and the 14 minutes above measure different things. Per minute, Maya’s loss was above 2% in 36 minutes, above 5% in 26 and above 10% in 21. The 14 minutes count only the stretches where a feed stayed above 5% loss or froze.
Then we checked the raw counters, because the per-track record is derived from them. We took cumulative packet counters from each viewer’s WebRTC stats and split them into series per client instance, track and connection, so that reconnects would not look like counter resets. Then we summed the deltas over the stretch.
| Maya | The other 18 viewers | |
|---|---|---|
| Inbound video packet loss | 9.2% | 0.00–0.13% each; one at 1.3% |
| Minutes with loss above 2% | 36 of 49 | none for 17 of them |
| Median round-trip time (browser estimate) | 500 ms | 50–200 ms |
| Downlink estimate (browser) | ~2 Mbps | mostly ~10 Mbps |
A low downlink estimate on its own proves little: two other viewers also reported about 2 Mbps and lost almost nothing.
Splitting by camera makes it clear. Maya lost 11%, 9% and 5% of packets from the three people on camera. The 13 to 17 other viewers receiving those same three streams lost 0.13–0.21% pooled. If one publisher’s uplink had been the problem, other viewers would have seen that publisher’s stream degrade too. If the SFU had been overloaded, the loss would have spread across many viewers. Instead it followed one receiver across every sender, which points at the path between the SFU and Maya: her access network, her Wi-Fi, or her ISP. We cannot tell those three apart from the data we have. The SFU itself rated her connection “poor” in about a third of its readings.
Paired bar chart, one pair per publisher. Maya's inbound video packet loss from publishers A, B and C is 11.0%, 9.3% and 5.1%; the pooled loss for all other viewers of the same publishers is 0.16%, 0.13% and 0.21%, barely visible on the same axis.
Checking the usual traps
Two traps mattered here. Browsers throttle hidden tabs, and a throttled tab can look like frozen video. Maya’s tab was hidden for 111 of 2,940 seconds, and packet loss is a transport counter that tab throttling does not produce; several viewers kept their tabs hidden far longer and lost close to nothing. Stats collectors also often drop healthy samples to save bandwidth, and ours does, so an average over stored rows would be an average of bad moments. We used counter deltas instead, which that filtering does not affect.
The others did not apply. Maya’s data arrived at the same cadence as everyone else’s, so this was not a collection stall that made a client look dark. Upload lag was 7 to 11 seconds for nearly every viewer, so clock differences were a few seconds at most, well below the per-minute buckets the verdict rests on. And each publisher had 13 to 17 other receivers, enough to say “everyone else was fine”.
The general lesson
A complaint is a claim about one seat. Before you open the SFU dashboard or the publisher’s logs, compare that seat with the rest of the room at the same moment, for the same tracks. A problem that follows one receiver across every publisher is their downlink; one that follows a publisher across every receiver is that publisher’s uplink.
Room averages hide this in both directions. A single bad seat gets diluted into a mild number for the whole room, as it did here: 9.2% for one viewer became 1.1% for everyone. And a real outage that hit a quarter of the audience can disappear inside a healthy-looking mean. Any check for “many people at once” has to use a window about as long as the symptom, or it will turn unrelated blips into a shared cause.
This does not need special tooling: per-viewer receive stats with timestamps you can line up, and a join on track and time. The hard parts are the traps listed above.
About the tooling
We build Rewitness, which records what each participant in a realtime session received and compares viewers against each other at the same instant. Its complaint view turns this comparison into a one-line answer, and for this case the answer is “their own connection”. It has limits: it infers server-side causes from what viewers received, since it does not see the media server’s logs, and it needs its SDK in the client.