A closer look at Backscatter's strobe test | By Retra UWT
Backscatter recently published a large comparison of twenty underwater strobes, testing brightness, beam pattern, recycle time, color temperature, weight, power consistency, and flash duration. It's a real amount of lab work, and we think it has genuine value — we also think several choices in how the data was collected and presented need adjusting before the charts can be called objective. This page walks through why, alongside our own measurements and a re-analysis of theirs.
All test data discussed below originates from Backscatter's article, "The Best Underwater Strobes Tested 2026" . This page is published by Retra UWT, manufacturer of the Retra Maxi and Retra Pro Max (II) strobes referenced throughout, including in the additional measurements below that we carried out ourselves. We're not a neutral third party here, any more than a retailer testing its own products is — see the conflict of interest section further down for more on that.
Jump to: Data presentation, Inconsistent measurements, What's in and left out, Conflict of interest, Where this leaves you.
How the data is presented
Even where the underlying measurements are sound, several charting choices distort how those numbers read at a glance — and nearly all of them push in the same direction: making differences at the bright, powerful end of the field look bigger and cleaner than the data actually supports.
1. Guide number is plotted on the wrong kind of axis
This is a worrying issue with a couple of charts in the comparison, and it's fixable with math rather than opinion. Guide number isn't a ruler — it works more like the Richter scale for earthquakes, or decibels for sound. Each "stop" is a doubling of actual brightness, not a fixed step. Going from GN 8 to GN 11 doubles the light. Going from GN 32 to GN 45 also doubles the light — even though a jump from 32 to 45 looks much bigger on paper than one from 8 to 11.
Picture ranking earthquakes by drawing a bar as long as the raw magnitude number: a magnitude 8 quake gets a bar twice as long as a magnitude 4 quake. But a magnitude 8 quake isn't twice as strong as a magnitude 4 — it's roughly 10,000 times stronger. Nobody charts earthquakes that way, because the raw number isn't the thing that matters; the exponential relationship is. Guide number works the same way: what matters to a photographer is how many stops apart two strobes are, not how many raw GN units apart they are.
On Backscatter's linear-axis chart, strobes clustered at the high-GN end look dramatically more separated from each other than they really are, while genuine differences at the low-GN end get compressed into looking almost irrelevant. So the GN comparison chart in Backscatter's report effectively shows the linear increase in brightness rather than the actual difference in F-stops, which are much more intuitive for every photographer. The fix is a logarithmic GN axis, so equal spacing always means an equal brightness ratio, wherever you look on the chart.
Guide number: stated vs. measured, on a corrected log axis
Put side by side with Backscatter's original linear-axis version, the difference is immediately obvious: strobes that looked dramatically separated at the bright end of their chart now sit much closer together, and strobes that looked nearly identical at the dim end turn out to be a full stop or more apart. Once the axis stops exaggerating the high end and compressing the low end, the actual spread of brightness across this field of strobes looks a lot less dramatic than Backscatter's chart implies.
From the guide number chart above, one more comparison is worth pulling out: how much each strobe's guide number changes between air and underwater. That jump isn't just a curiosity — since water can only concentrate light that isn't already being spread by a strobe's internal optics or a domed front element, a strobe that gains a lot of brightness underwater is telling you something about how narrow its beam becomes once submerged.
Brightness gained underwater, by strobe
The recycle time chart also need a similar axis adjustment treatment as the GN comparison chart.
Recycle time vs. brightness, on a corrected axis
There's also a distinction worth raising about what this chart measures. Backscatter records recycle time two ways: time to full flash readiness, and time to any usable flash at reduced output. Only the full-flash number is charted. By several independent accounts, some strobes — including Backscatter's own — recycle to full power quickly but are comparatively slow to fire again at any reduced power setting, a pattern Backscatter's own review text touches on in passing without charting it. Leaving the any-flash number out of the comparison means a reader can't see that behavior at a glance, even though it's arguably more relevant to burst shooting than full-power recycle time is.
More broadly, a lab number for recycle time only tells part of the story. A practical, qualitative test — shooting real subjects back to back and simply noting whether the strobe keeps up — can surface real-world recycle behavior that a single stopwatch figure misses. Dave Hicks' long-term review of the Retra Maxi is a good example of that approach, and it's the kind of test we'd like to see paired with the lab numbers here, not instead of them.
2. The beam viewer's layout invites a comparison it can't deliver
Backscatter's side-by-side beam comparison shoots every strobe at full power, with exposure held constant per shutter-speed group — except that different shutter speeds were deliberately used for strobes with longer flash durations, specifically to show each strobe's "max potential power output." That's reasonable if the goal is showcasing peak output, but it undercuts the idea that the viewer offers an apples-to-apples brightness comparison, since the effective exposure isn't held constant across every strobe being compared. A reliable side-by-side view needs every strobe's center brightness calibrated to the same value first, and that step doesn't appear anywhere in Backscatter's stated method — so the slider format ends up implying a fair, at-a-glance comparison it isn't actually equipped to make.
3. Alphabetical order works against comparison
Every chart in the article sorts strobes alphabetically. That's a defensible default, but for charts whose whole purpose is ranking strobes against each other, it works against the reader — you have to scan the whole chart to find where any given strobe lands relative to its competitors. Sorting by measured value, as we've done throughout this page, turns each chart into a ranking a reader can use in seconds instead of a lookup table they have to search.
4. Combining both weights onto one chart hides the more useful comparison
Backscatter plots weight in air and weight underwater together on a single chart, sharing one scale. Air weights span a much wider range than underwater weights do, so putting both on the same axis compresses the underwater figures — which is exactly the number that matters most for buoyancy — into a narrow band where real differences between strobes are hard to read. A strobe that's heavy in air isn't necessarily heavy underwater — buoyant housings and battery placement affect the two numbers differently — and a shared scale makes it harder to see, at a glance, which strobes lose or gain the most weight relative to each other once submerged. Splitting the two into separate charts, as below, gives each its own scale and makes those differences legible.
Weight in air
Weight underwater
5. No error bars, almost anywhere
Outside of a passing mention that color temperature readings varied by "100–150 Kelvin, and in a few cases more," the article reports no variance, confidence intervals, or trial counts for guide number, recycle time, or flash duration. We're told "multiple readings are taken to ensure accuracy" for GN, but not how many or how consistent they were. For a test whose whole premise is that only careful, controlled measurement can be trusted over manufacturer specs, the absence of uncertainty reporting is a real gap — a single outlier reading is presented on a chart with no way for a reader to know how stable it is. The charts above use estimated error bars wherever we had a defensible basis for one: light-meter accuracy for guide number, and stated scale resolution for weight.
Inconsistent measurements
Beyond how the data is charted, a few results in the underlying measurements themselves gave us pause — either because we couldn't reproduce them, or because Backscatter's own numbers seem to contradict each other.
1. Color temperature discrepancies
We tested two strobes ourselves — the Retra Maxi and the Backscatter Hybrid Flash (HF-1) — using a Minolta Color Meter III-F, one meter from the strobe, with 5 samples taken at each power level rather than Backscatter's 3. That's a less advanced instrument than Backscatter's Sekonic C-7000, so we don't expect our absolute Kelvin values to match theirs exactly — but to compare relative differences between strobes and across power levels, it should be more than adequate.
Backscatter's own article explains that most strobes run warmest at full power and cool down as output is reduced. Their measured number for the Maxi breaks that pattern: 7168 K at maximum power, also the coolest reading they report for it. Our own measurements show the opposite and more conventional shape — the Maxi runs warmer at high output and cools down as power comes down, the trend their own text says is typical. Given that temperature generally warms up as GN increases, Backscatter's own claim of the coolest Maxi reading occurring at its highest power setting doesn't add up.
Color temperature vs. output: our own measurements
A grayscale card shot at matching power levels supports the same conclusion. Below, the HF-1 is compared against the Maxi at full power on the left, and a Retra Pro Max II is compared against the same Maxi at full power on the right — both pairs shot identically, with the same GN in the center of the strobe.
Full power comparison: HF-1 vs. Maxi
Full power comparison: Pro Max II vs. Maxi
After equalizing brightness between frames, we pulled CCT values directly from the gray patches: 6410 K for the Maxi, 5330 K for the Pro Max II, and 6096 K for the HF-1. Looking only at relative differences between the color temperatures, these values line up well with our meter readings. The photos also make a point the numbers alone can understate: the visible difference between the HF-1 and the Maxi is much smaller than the difference between the Pro Max II and the Maxi — which is consistent with our measured curves, but easier to see at a glance in the images than in a chart. It is obvious that the difference between the color temperature of the Maxi and the HF-1 is no where near the 688 K stated by Backscatter.
2. The Isotta RED64 doesn't add up either
The RED64 measures as the single highest guide-number strobe in Backscatter's entire test, yet in their own beam viewer it visibly appears less bright than several lower-GN strobes, and several photographers we've spoken to who actually shoot it don't consider it a standout performer.
We raised this directly with Backscatter. Their reply attributed the discrepancy to the meter reading being taken at a 1/60 second shutter speed, implying the RED64's longer flash duration explained the inflated number. But Backscatter's own published flash-duration test gives the RED64 a duration of roughly 1/231 second — already faster than the 1/200 second shutter speed used for the underwater beam photo. The full flash should have been captured within that exposure. Their own numbers don't support their own explanation.
This matters beyond one strobe: if the RED64's visual ranking can diverge this sharply from its measured GN, and the explanation offered for it doesn't hold up, there's no reason to assume no other strobe in the set has a similar, just less obvious, discrepancy.
What's in the test, and what's conspicuously not
1. Beam angle was left unquantified
Backscatter is upfront about this one: "we purposely did not try to call out what specification a beam angle should be from our testing… we think the best comparison is to just look at the photo." That's a defensible position for a genuinely ambiguous metric, but it's also the only major spec in the entire article that isn't reduced to a number.
Beam angle and falloff are measurable — a goniophotometric setup takes light-meter readings at fixed angles off-axis and plots relative intensity falloff by degree (here is an example). It's substantially more data collection per strobe than a single center-beam GN reading, which is plausibly why it wasn't attempted here. Whatever the reason it was left out, the underwater-brightness-gain data above shows Backscatter's own strobes among the most beam-narrowing in the entire test — exactly the kind of result a beam-angle chart would have made hard to miss.
2. Power-level range gets less attention than spacing
Backscatter's "consistency of power levels" chart focuses on how evenly spaced each click stop is, which is a reasonable thing to measure. But total usable range — the gap between a strobe's brightest and dimmest setting — often matters more in practice, and isn't given its own comparable metric. A strobe with slightly uneven click spacing but a much wider range can be more practically flexible (closer macro work at the bottom, more headroom in bright ambient light at the top) than one with perfectly even clicks and a narrow range, and the current chart doesn't let a reader see that trade-off at a glance.
3. A lot of what matters to buyers isn't tested at all
This is a rigorous look at a handful of optical and electrical specifications; it is not a complete buying guide, and it shouldn't be weighed as one. Durability, long-term dependability, serviceability, price, video-light capability, battery type and battery life, and flood-alert systems are all things a buyer might reasonably weight as heavily as guide number or color temperature, and none of them appear in this test. That's a legitimate scope limitation, not a flaw in what was measured, but it's worth keeping in view when deciding how much this test alone should influence a purchase.
A retailer testing what it sells
Backscatter is a retailer that sells nearly every strobe in this comparison, including its own branded strobes tested directly against third-party competitors. That alone doesn't invalidate the data, and to their credit, their own products appear to have been run through the same protocol as everyone else's, with no sign of the underlying lab work itself being rigged.
But a review where the publisher financially benefits from the sale of the products it's grading is not a neutral, disinterested study by default, and that conflict is never acknowledged anywhere in the article. Given that several of the presentation issues above happen to work in favor of clarity at the high-output end of the market, where Backscatter's own strobes are positioned, this is exactly the kind of test that calls for more rigor and more transparency about incentives, not less, before it's marketed as "the most comprehensive strobe review ever."
The same standard applies to us. This page is published by Retra UWT, and the Retra Maxi and Retra Pro Max (II) appear throughout it — including in the color temperature, guide number, brightness-gained, and recycle-time measurements we carried out ourselves, several of which show the Maxi in a favorable light. We think the specific findings here hold up on their own merits, and we've tried to describe what the data shows rather than argue for a conclusion. But we're a strobe manufacturer writing about a comparison that includes our own product, which is the same basic conflict we're describing above, and we'd rather say so directly than have a reader work it out on their own.
Good study, needs adjustment
None of this means the underlying lab work is worthless. The raw measurements likely have real value, and Backscatter is transparent, to their credit, about several aspects of their method that other outlets don't bother explaining at all. But "here is data we chose how to visualize, and didn't error-bar" is a different thing from "here is an objective scientific comparison," and the specific choices above, taken together, consistently make the differences between strobes look bigger and cleaner than the underlying data actually supports.
If you're using Backscatter's test to shop for a strobe: read the guide number chart in stops, not raw numbers. Treat the beam viewer as a rough visual guide rather than a calibrated brightness measurement. Don't put too much weight on differences of a few hundred Kelvin in color temperature or a fraction of a stop in guide number — by Backscatter's own admission, that's within their measurement noise. And remember this test covers a handful of optical specs, not everything that makes a strobe worth owning.
You fired the only strobe on this page we didn't put an error bar on.
Full marks.

