You already bought it.
Maybe it was a set of deep wheels — the ones with the tall rims. Maybe new tires, a tighter jacket, ceramic bearings, or five psi less than you used to run. Somewhere behind the purchase there was a review, or a number, or a friend who swore by it. Eight watts. Twenty watts. Faster.
So you fitted it and you went riding, and the bike felt different. It usually does. But feeling faster and being faster are not the same thing, and nothing since has told you which one you got.
The question is no longer whether it is worth buying. You own it. The question is whether it is doing anything for you, on your roads, at the speeds you actually ride.
The evidence exists. It is just not yours.
The industry does test this properly. Wind tunnels, velodromes, controlled protocols, repeat runs. The numbers that come out of that work are real.
They are also not about you. A tunnel result is produced at a chosen yaw angle — one fixed angle of wind hitting the bike — at a speed the test selected, on a rider who is not you, in conditions that do not exist outdoors. It tells you what a component did in a tunnel. Whether it does the same thing on your roads, at your speeds, in your position, on the ride you actually do on a Sunday, is a different question, and nobody has answered it for you.
You cannot answer it with a before-and-after either. You changed the wheels — but you also had a tailwind, you were fresher, and you took the flatter way home. Any one of those is worth more than the wheels.
Your rides are already the experiment
Here is what makes this solvable: you have already done the testing.
Every ride you have uploaded is a record of how much power it took to move you at a given speed, in real conditions, on real roads. The data was never the missing piece.
What was missing was the controls. RideIQ Lab supplies them — you tell it what you were running, and it compares like with like: the same rider, matched rides, measured against each other rather than against a manufacturer’s claim.

Lower watts at the same speed
The method is the one a tunnel uses, applied to a road.
Lab sorts your rides into speed bins and asks a single question: at this speed, how many watts did it take? Two setups, the same speed, different power required — the one needing fewer watts is the more efficient setup. That is the entire comparison.
What comes back is a verdict in watts. Skinsuit came out about 28 W ahead. That is a tight one-piece against a normal jacket, and 28 watts is a large gap — enough to feel on every climb. Not a rating out of five, not “feels fast” — a number, in the unit that decides how hard your ride is, taken from your own riding.
Underneath it sits the evidence chart: the speed bins, the filtered samples, both curves plotted against each other. You can look at the result rather than take it on trust.
Not every second of a ride is evidence
A ride is mostly not a test. You freewheel down descents, soft-pedal through junctions, stop at lights and sprint for the odd sign. None of that says anything about your wheels.
So before any comparison happens, the samples are filtered. Coasting comes out. So does anything where you were barely turning the pedals — and, just as importantly, anything where you were producing far more power than that speed should need. A sprint distorts an efficiency curve exactly as badly as a freewheel does, only in the opposite direction.
How that filtering works depends on which comparison you are running. The long-run one, built from months of tagged rides, applies the test behind everything else in RideIQ. Not is this a lot of watts, but is this what you normally produce at this speed? Samples close to your own normal count in full. Samples further out count for less. Samples well outside it do not count at all.
A head-to-head A/B works differently, because the two rides you picked may predate your current baseline entirely. There it judges the riding itself: sustained pedalling rather than coasting, flat ground rather than climbs and descents, and speeds high enough that drag is actually the thing being measured.
What survives is the part of the ride that was actually riding. That is the only part that can tell you anything about your equipment.

It tells you when it is not sure
This is the part that matters most, and it is the part almost every equipment claim leaves out.
Every verdict carries its confidence rating and the sample it was built from — how many matched rides, across how many speed bins — with a note in plain language about how far to trust it.
If the two rides never met at a common speed, there is no verdict at all. The report says it does not have enough to compare, rather than reaching for an answer anyway.
It also audits its own fairness. Both sides pass through the same filters, so it tracks how much of each ride survived them and tells you when the two differ — a tenth of one ride set aside against a third of the other. That gap is worth seeing, because a comparison where one side was filtered far harder than the other is not really comparing like with like.
That honesty is the whole point. A tool that always sounds certain is a tool you cannot use, because you have no way of telling which of its answers to believe.
Worth separating two things that sound alike, though. The rating describes how much data sits behind the curve. It does not describe how carefully you set the test up — and that part is yours.
How to run a comparison worth trusting
What ruins an equipment comparison is not a small sample. It is everything that changed between the two rides.
You will need power data. The comparison works by measuring the watts required to hold a given speed, so a power meter is what makes an equipment verdict possible. Lab’s heart-rate and cadence charts work without one; the equipment verdicts do not.
Ride the same route, and pick a flat one. Not a similar route — the same one, so corners and junctions cancel out instead of having to be averaged away. Flat matters more than you would expect: anything on a noticeable gradient is set aside before the comparison starts, because a climb tells you about your legs and a descent tells you about gravity. Neither tells you about your wheels.
Ride them close together. Back to back on the same morning is ideal. Wind is the largest uncontrolled force acting on you, and it changes over hours, not minutes.
Change one thing. Same bike, same position, same clothing. Swap the wheels or swap the skinsuit, not both.
Ride both at a similar pace. The comparison works by matching speeds across the two rides and asking what each one cost in watts. Attack one lap and soft-pedal the other and there is less common ground left to compare.
Keep it steady, and keep the speed up. Sustained pedalling is what counts. Coasting, junctions and anything slower than about 15 km/h drops out before the comparison starts, so a short continuous effort beats a long stop-start ride.
It does not need to be long. Any ride from the last six months can go into a test, and rides as short as a kilometre are eligible provided the riding in them is usable. More gives the comparison more to work with — but you do not have to find a spare afternoon to run one.
Do that, and two rides is a real test. The reason scattered rides need volume is that they are averaging away conditions nobody controlled. Control those conditions yourself, and you do not need the volume to do it for you.
Quick tests, and the long game
Some questions settle in a morning. A jacket against a skinsuit produces a gap so large that one properly controlled pair of rides will show it plainly.
Others do not. The difference between two chains, or two sets of bearings, or five psi either side of your usual, can be small enough to sit inside the variation of even a well-run comparison.
So Lab also lets you assign equipment to rides as you go. Tag what you were running, keep riding, and the comparison accumulates in the background. Over a season, with dozens of rides on each setup, a difference too small to see in a morning becomes visible — and the confidence rating climbs with the evidence.
Controlled A/B for the big questions. Long-run tagging for the small ones.
Which bike, which wheels, for the ride that matters
Plenty of riders end up with more than one of everything. A summer bike and a winter one. A shallow set of wheels and a deep set. Two pairs of tires, because the first pair was fine but the second was on offer.
For them the question changes shape. It is no longer whether a purchase was worth making — that money is spent either way. It is which one to be on when the ride actually matters. The club hill climb, the hundred miles you have been building towards, the day you finally go for the segment.
This is what long-run tagging is really for. Assign each setup as you ride it, carry on riding normally, and over a season Lab builds a power-at-speed picture of each one from your own roads. Not a manufacturer’s claim about a wheel, but what that wheel cost you, at the speeds you ride, on the surfaces you ride them on.
By the time the important ride comes round, the choice is a number rather than a hunch. And occasionally the answer is that it makes no measurable difference — which is worth knowing too, because it means you can take whichever bike you enjoy riding more.
What this means if you review equipment
Most equipment reviews are, in fairness, subjective. A reviewer rides the wheels, the wheels feel quick, and the review reports that they feel quick. That is not dishonesty — without access to a tunnel there has not been much else on offer.
Testing on real roads changes what a reviewer is able to say, and it fits how reviewing actually works. You rarely get a season with a product. You can nearly always get one morning and one loop.
Ride the control, swap the part, ride the same loop again, and report the watts with the sample and the conditions attached. It will not match a tunnel for precision. It will beat “feels fast” by a distance — and unlike a tunnel result, the reader can repeat it on their own roads.
What a wind tunnel still does better
Worth being straight about the trade.
A tunnel controls everything — yaw, speed, temperature, position, repeat runs on demand. That control is exactly why it can resolve differences too fine for road testing to see. If you need a component’s drag to a tight tolerance, that is still where you go.
Lab trades that control for realism. It cannot hold the wind still. What it can do is answer a question the tunnel never asks: what did this equipment do for you, on your roads, at the speeds you actually ride, across the rides you actually did.
Both are real measurements. They answer different questions.
Stop guessing
You do not need a wind tunnel, a velodrome or a testing budget to find out whether the thing you bought did anything at all.
You need the rides you were already going to do, and something honest enough to tell you what they showed — including when what they showed was nothing.
You bought it on somebody else’s number. You can check it against your own.
Did that new gear make you faster? Now you can find out.
Find out what your equipment is actually worth — on your roads.
Free for 30 days, no obligation. Subscribe whenever you like — the rides you have already brought in stay yours either way.
New to RideIQ? See everything the app does →