1. Home
  2. Blog
  3. Methodology

How We Test Streaming Video Quality

A review is only as good as the method behind it. Here’s exactly how we measure streaming quality — the hardware, the metrics and the conditions — so our scores actually mean something.

“It looked good to me” is not a test. If we’re going to hand out grades, you deserve to know how we reach them. This page is our methodology in full — the metrics we measure, the rig we measure them on, and the conditions we measure under. Nothing here is proprietary; the whole point is that you could repeat it.

The Four Metrics That Matter

Perceived quality comes down to four measurable things. Everything else is marketing.

MetricWhat It MeasuresWhy It Matters
Start-up timeSeconds from pressing play to first frameThe first impression; long waits feel broken
Delivered resolutionActual pixels on screen vs the label“HD” and “4K” labels often overstate reality
Buffering ratioTime spent stalled as a share of playbackThe single biggest driver of frustration
StabilityQuality drops and re-buffers over a full titleA stream is only as good as its worst minute

The Test Rig

We test on a consistent, mid-range setup rather than a high-end showpiece, because that’s closer to what most people actually use: a current mid-tier Android phone, a mainstream 4K TV, and a wired reference connection for the baseline. Using the same hardware every time is what makes scores comparable across reviews.

Testing Under Real Conditions

The number that matters isn’t the best case — it’s the typical one. So we test across a spread of conditions rather than a single lucky run:

  • Off-peak vs peak: weekday afternoons and Saturday-night prime time, because congestion changes everything.
  • Broadband vs mobile: a stable wired baseline plus a throttled connection simulating real 4G.
  • Cold vs warm: first play of the session versus later, since caching flatters repeat runs.
One test proves nothing

Any single stream can be lucky or unlucky. We run each scenario multiple times across different days and report the typical result, not the best one we saw. A cherry-picked demo is exactly what we’re trying not to publish.

How Metrics Become a Score

Each metric is scored against a fixed scale, then combined with a consistent weighting — buffering and stability count for more than a second or two of start-up delay, because that’s what viewers actually feel. The same weighting is applied to every product, so a 7 in one review means the same as a 7 in another.

Repeatability and Re-Testing

Streaming performance drifts as services change infrastructure and libraries. We date-stamp every result and re-test periodically, updating scores when reality moves. A review here is a snapshot with a date on it, not a permanent verdict.

Methodology FAQ

Why test on mid-range hardware instead of the best available?
Because most people watch on mid-range devices. Testing on a flagship would flatter every service and misrepresent the experience the majority actually get.
How do you measure the “real” resolution?
We compare the delivered stream against the advertised label, watching for services that badge an upscaled or lower-bitrate stream as HD or 4K when the detail isn’t truly there.
How often are scores updated?
We re-test periodically and whenever a service makes a major change. Every published score carries the date it reflects.
Can I reproduce your tests?
That’s the goal. The metrics and conditions here are standard enough that anyone with a stopwatch, a data meter and patience can sanity-check our numbers.
JR
Jack Rogers

Jack has covered consumer streaming apps and Android security since 2021. He tests every app he reviews on dedicated hardware with isolated accounts, never his daily driver. Reach him via the contact page.

Keep reading

More original guides from the same testing desk.