A causal-inference case study, worked end to end
Did the UFC’s anti-doping era change how fights play out?
In July 2015, the UFC did something combat sports almost never does: it brought in independent drug testing, on a simple theory. Cleaner fighters, a cleaner sport. This project checks whether that actually shows up in the fights themselves, using real fight data and the fact that two rival promotions never made the same move.
Short answer: no detectable change. The estimate is +2.2 pp on finish rate, with a margin of error of ±5.8 pp either way, a result this page will show can't be told apart from ordinary noise.
- UFC (tested)
- 1,513 fighters
- Bellator (never tested)
- 1,890 fighters
Round 1
Why not just compare before and after
The obvious approach doesn’t work: just compare UFC fights before and after July 2015. The sport changed in a hundred ways over those years that have nothing to do with drug testing, so any shift you found could be explained by almost anything else.
The fix is to also watch a promotion that never made the switch. Bellator never adopted USADA testing, so tracking Bellator’s own trend over the same years gives an estimate of what the UFC would probably have looked like without the policy change. The gap between what actually happened to the UFC and that projected path is the estimated effect.
Illustrative only, not built from this project’s real data. The real result is below.
Round 2
So, did anything change
Comparing 1,513 UFC fighters against 1,890 Bellator fighters around the July 2015 switch-on, on how often fights end in a finish rather than going to the judges: an estimated shift of +2.2 pp, give or take ±5.8 pp. That margin of error is wider than the shift itself, which is the statistical way of saying this number could easily be zero.
Two reasons to hold this loosely. First, this comparison only has one promotion on each side, so the true uncertainty is probably wider than the margin of error above suggests. Second, the fighters competing in each promotion changed a lot over this period, and not in matching ways (women’s fights grew into a much bigger share of UFC cards than of Bellator’s, and women’s fights finish less often), which this specific estimate doesn’t yet correct for.
Round 3
Could this just be noise
A fair question: with data this messy, could a similar-looking result show up even if USADA changed nothing? To check, this project reran the exact same comparison at 21 other dates, pretending the switch had happened then instead (7 of those didn’t produce a usable result and were dropped, leaving 14 fake dates to compare against).
The real date’s estimate, +1.7 pp, is not unusual at all next to the fake ones. Formally: a p-value of 0.867. In plain terms, most of the fake dates produced a result just as big as the real one, so this method genuinely can't tell them apart.
Two honest limits on this check. With only 14 comparison dates, it’s a blunt instrument: it could easily miss a real effect that’s moderate in size, not just a large one. And one of the strongest fake dates happens to line up with a different, real anti-doping policy change made by a state athletic commission around the same time, a reminder that even the comparison dates aren’t perfectly clean.
Round 4
Can this method be fooled
Any statistical comparison can be nudged toward a preferred answer by a subtle mistake in how the comparison group is chosen. Rather than just warn about that, this project deliberately makes that exact mistake, on real data, and shows what it does to the result.
PFL bought Bellator outright in November 2023, about four months after PFL brought in USADA testing of its own. Bellator itself was never tested, but once PFL owns both promotions, is Bellator still a fair, independent comparison group for judging PFL? Two judges score the same fight data and disagree.
Judge A: keeps Bellator as a comparison group forever
- Estimated effect
- -25.5 pp
- Margin of error
- ±10.2 pp
- Likely range
- [-45.4, -5.5] pp
- Quarters counted
- switch quarter, +1, +2, +3, +4
Statistically significant
Judge B: stops counting Bellator once PFL owns it
- Estimated effect
- -21.3 pp
- Margin of error
- ±11.2 pp
- Likely range
- [-43.3, 0.7] pp
- Quarters counted
- switch quarter, +1
Not statistically significant
The two judges don’t disagree because they see the underlying fights differently: on every stretch of time both can measure, they land on the exact same number. What actually splits them is reach. Judge A keeps leaning on Bellator data from well after the sale, some of it down to a handful of fights, and that thin extra evidence is what pulls the naive number, and its verdict, in a different direction.
Neither number here is a genuine finding about PFL’s own anti-doping program; this is a controlled demonstration of a specification mistake, not a real estimate. And even Judge B’s more careful verdict isn’t rock solid: re-running the same numbers under 30 different random starting points, Judge A’s call holds in 30 of 30 tries, but Judge B’s holds in only 24 of 30. A result that shaky deserves a second look before anyone treats it as settled.
This comparison is about which statistical setup can be trusted, not about PFL, Bellator, or any fighter’s actual doping status.
What this does, and doesn’t, say
Every number on this page is a comparison across thousands of fights in two or three promotions. Nothing here says or implies that any individual fighter, on either side of any comparison, used or did not use a banned substance. If a real policy effect exists at all, it could come from any number of things that have nothing to do with any one fighter’s choices: matchmaking, training, judging norms, or the roster itself changing. This project is scoped to what data at that scale can actually support, and stops there.
The bigger idea here isn’t really about MMA. Any time a policy, a product change, or a new rule gets credit for improving something, the honest next question is: compared to what? This project is one worked example of answering that question carefully, including showing exactly how the answer can be manipulated when the comparison isn’t chosen with care.