FAIR DECATHLON MODEL
PART VIII
The B Parameter
Where Should Zero Points Begin?
10,720 lower-range performances | A lower-tail audit of the zero-point parameter
Rafal Snoch | for Decathlon2000.com | August 2026
|
THE QUESTION LEFT BY PART VII Part VII ended with an awkward question. The ten event distributions remained remarkably coherent through most of the observed decathlon scale, even after every complete performance of 7000 points or more was removed. Only deep in the lower tail did that coherence begin to weaken sharply. By R7800 the 110 m hurdles had already separated strongly from the other nine; at R10400 several events were moving apart at once. That left one question unresolved: how much of the final divergence belongs to the athletes themselves, and how much belongs to B, the parameter that determines where a performance reaches zero points? Answering it requires something the upper calibration could never provide on its own: enough real performances close to the bottom of the scoring curves. The first clue, however, had appeared much earlier. |
Part III produced two forty-season versions of FDM. FDM-40N, the natural model, allowed the exponent C to take whatever value followed from the data. FDM-40C was a control experiment: whenever necessary, B was moved so that C remained inside the traditional interval 1 < C < 2. Both versions were forced through the same upper calibration structure.
Most of those changes looked harmless around ordinary high-level decathlon performances.
The javelin was the most striking example.
FDM-40N: P = 34.363170 (M - 7)^0.797549
FDM-40C: P = 11.980302 (M + 6)^1.006822
The zero threshold had moved from +7 m to -6 m.
That sounds enormous. Yet around the calibrated high-level range the two curves were almost indistinguishable. Part III already noted that the visible cost of the constraint was displaced toward weaker performances and was most obvious in the lower part of the javelin curve.
The numerical comparison makes the problem clearer.
|
Javelin |
FDM-40C |
FDM-40N |
|
10 m |
195 |
82 |
|
12 m |
219 |
124 |
|
15 m |
256 |
180 |
|
20 m |
318 |
265 |
|
30 m |
441 |
418 |
|
40 m |
565 |
558 |
|
48.96 m |
676 |
676 |
|
50 m |
689 |
690 |
|
60 m |
813 |
815 |
|
65.82 m |
885 |
885 |
|
70 m |
937 |
935 |
At 50-70 metres, the choice of B is almost invisible. At 10 or 20 metres, it defines two completely different scoring systems.
That is the central identification problem of this article. If two upper anchors can support B = -6 and B = +7 with almost no visible difference around normal high-level performances, the upper anchors cannot tell us where B belongs. They can determine the upper scale very tightly while leaving the basement surprisingly free.
Part III had therefore already shown that B could move.
Part VIII asks where it should move.
The FDM equations retain the standard combined-events form.
Running events: P = A(B - M)^C, for M < B
Field events: P = A(M - B)^C, for M > B
A valid performance on the non-scoring side of B receives zero points. B is therefore a mathematical zero threshold, not a definition of whether a performance is legitimate. A runner can finish a race slower than B. A thrower can record a valid throw shorter than B. The mark remains valid; the scoring function simply assigns it zero.
Changing B alone would obviously change the entire scoring scale. That is not what is tested here.
For every candidate B, A and C are recalculated while the established upper FDM relationship is held fixed. The same Part III upper reference performances remain the structural anchors. B is allowed to move, while A and C reshape the curve so that those anchor marks retain exactly the same calibration targets.
This produces a highly asymmetric effect.
• Near 700-900 points, the curves barely move.
• Around 300-400 points, the difference becomes visible.
• Near 0-200 points, it can become enormous.
Preliminary lower-tail tests confirmed that the freedom exposed by the Part III javelin example was not unique. Moving B while preserving the established upper calibration could substantially reshape weak performances with very little effect on the upper scale. Those exploratory tests also suggested that forcing every event into a universal 1 < C < 2 interval would not survive contact with the lower population.
The final audit uses the much larger database developed for Parts VI and VII: exactly 10,720 complete performances below 7000 OFF points, ranging from 1314 to 6999. Every record contains ten valid senior-specification event marks. The unified master of 27,741 complete performances remains available as a wider control population.
The larger sample changes the nature of the problem. A few spectacular weak performances no longer have to carry the argument alone. The bottom can be examined as a distribution.
The observed extreme tail is nevertheless useful as a sporting reality check.
|
Event |
Weakest observed |
11th weakest |
21st weakest |
Mean weakest 21 |
|
100 m |
18.30 |
14.86 |
14.53 |
15.13 |
|
LJ |
2.92 |
3.86 |
4.05 |
3.75 |
|
SP |
4.84 |
5.69 |
5.94 |
5.62 |
|
HJ |
1.10 |
1.25 |
1.29 |
1.24 |
|
400 m |
96.78 |
78.28 |
74.39 |
79.18 |
|
110H |
48.00 |
28.81 |
27.31 |
30.34 |
|
DT |
6.09 |
12.44 |
13.92 |
12.06 |
|
PV |
1.20 |
1.50 |
1.60 |
1.44 |
|
JT |
8.33 |
14.16 |
16.15 |
13.92 |
|
1500 m |
9:05.20 |
7:45.86 |
7:28.18 |
7:57.68 |
The purpose of this table is not to set B equal to the weakest mark, the 21st weakest mark or any other arbitrary order statistic. It tells us where real decathlon performances actually begin to populate the basement, but it does not by itself decide what the scoring threshold should be.
There is no single numerical criterion from which B can be read directly.
The B audit uses a common framework but not a mechanical optimiser.
For each event, candidate B values are inserted into the scoring equation and A and C are re-solved so that the established upper FDM geometry remains fixed. The candidate curve is then examined through the ordinary lower distribution, deeper stress percentiles, the empirical extreme tail, the number of valid marks pushed to the zero side, and the event residual relative to the level implied by the other nine events.
The main residual comparison concentrates on the P60-P90 region of the 10,720-performance event distributions, with P95 and deeper percentiles used as stress tests rather than automatic fitting targets. The same candidate is also checked against the upper scale so that an apparent lower-tail improvement cannot hide an elite distortion.
That last diagnostic requires caution.
A residual is not a pure measurement of scoring error.
Parts V and VI already showed that lower-level decathletes do not deteriorate symmetrically across ten events. Some abilities survive surprisingly well as the total falls. Others collapse earlier. The event residual therefore mixes at least two things: the geometry of the scoring function and the real structure of the athlete population.
The audit consequently does not ask for the B that drives every lower residual to zero. It asks for a B that improves the lower geometry, remains sportingly interpretable, avoids creating unnecessary empty scoring space or excessive zero-side cases, and does not damage the established upper scale.
Residuals are diagnostics, not verdicts.
That distinction becomes decisive once the events are examined one by one.
The 100 m produces a broad stable region rather than one sharp optimum. Pure lower-tail residual minimisation prefers values closer to 15 seconds, and the numerical differences among neighbouring candidates are small through most of the distribution.
Part VI, however, already indicated a positive sprint population effect: at low overall decathlon levels, flat sprinting often survives relatively well. Forcing that positive residual completely away would risk using B to erase a real characteristic of lower-level athletes.
Between approximately 16.0 and 16.6 seconds, the upper and middle portions of the curve are almost unchanged. The differences appear mainly in the final fractions of the extreme tail. Within that stable interval, the lower value also avoids giving the very slowest isolated marks too much influence over the curve.
I therefore adopt B = 16.0 s. The exponent falls from 2.465 to approximately 1.752.
Long jump is an important negative result. The inherited 220 cm threshold sits below even the weakest recovered valid mark of 2.92 m. Moving B downward can improve some mechanical lower-tail diagnostics, but only by placing the mathematical zero still farther from the observed population. In that region B starts to behave mainly as a hidden curvature control rather than a plausible zero threshold.
Moving it upward was also tested. A 250 cm version barely disturbed the top, but it did not solve a meaningful lower-tail problem either.
There is therefore no convincing reason to move B. Long jump remains at B = 220 cm.
Shot put provides one of the clearest cases against the inherited threshold. A zero at 1.50 m leaves a very large region of positive scoring space that essentially no complete senior decathlete uses. The lower-tail scans improve sharply as B approaches the actual bottom of the population.
The numerical optimum is broad rather than sharp. Values around 4.70-4.80 m behave almost identically: the main P60-P90 RMSE is virtually flat, while the deepest tail changes only gradually. The weakest observed valid mark is 4.84 m.
The centre of that stable region is B = 4.75 m. The resulting curve is deliberately concave, with C approximately 0.748. That is the geometry required to remove the unused basement while preserving the upper anchors.
The inherited high-jump threshold of 75 cm lies far below the observed lower tail. A purely numerical fit continues to improve as B is raised well above one metre, but high jump has a known positive population effect: lower-level decathletes can retain relatively good jumping ability even when several other events deteriorate badly.
Chasing a zero residual would therefore move B too far into the population. A more conservative threshold also has a simple empirical interpretation: the weakest valid high jump recovered in the 10,720-performance database is 1.10 m.
I therefore set B = 110 cm. The resulting C is approximately 1.028, making the practical high-jump curve almost linear through much of the observed range.
The 400 m moves in the opposite direction from the field events. The inherited 82-second zero produces a strongly convex curve with C = 2.368. Sensitivity tests show that a lower B can preserve the established upper relationship while producing a much cleaner continuation into the lower distribution.
The event also exhibits a mild negative population effect at lower levels. A mechanically attractive higher-B solution would therefore partly compensate for real athlete structure rather than improve the scoring function itself.
The audit settles on B = 77.0 s. The resulting exponent is approximately 1.999: for practical purposes, an almost exactly quadratic curve.
The hurdles cannot be treated as an ordinary smooth lower-tail problem. A very slow 100 m is still recognisably the same task performed badly. A very slow 110 m hurdles can represent something else: the athlete has reached a technical barrier at which clearing ten 106.7 cm hurdles becomes a fundamentally different exercise.
The senior database contains valid times far slower than 25 seconds, including some spectacular extremes. Those marks are real and remain in the data. But they form a very thin extreme tail rather than a normal continuation of the central distribution.
Lower-hurdle U18 and U20 evidence provides useful context. Technical breakdown exists there too, including occasional very slow races, but comparable 25-second-plus performances occur much deeper in the overall ability distribution. The senior 106.7 cm hurdle height moves this breakdown upward: athletes who remain competent decathletes in other events can reach a point where the hurdles themselves become the dominant limitation.
This makes the event different from a simple sprint. The lower-tail evidence supports treating 24.9 seconds as a plausible boundary for meaningful scoring continuity, while recognising that slower races can still be valid completed performances.
I therefore adopt B = 24.9 s. The resulting C is approximately 1.995. That resemblance to the traditional 1 < C < 2 principle is interesting, but it is not the reason for the choice; the technical-threshold evidence is the independent argument.
Discus is a warning against selecting B from one attractive statistic. Mechanical residual minimisation can favour values near 1-2 m, and the Part III constrained experiment also happened to produce B = 1 m. Neither result has much sporting meaning as a zero threshold.
Part VI showed a clear lower-level discus trench. Once performances fall far enough, the structure of the athletes changes; the event does not simply continue as a perfectly smooth scaled-down version of elite discus. The audit therefore should not use B to flatten that population feature.
A conservative movement toward the empirical tail is sufficient. I adopt B = 7.0 m, producing a moderately concave curve with C approximately 0.854.
Pole vault provides perhaps the cleanest rejected experiment. A higher threshold looked plausible when judged from the observed lower tail alone. A later 100-125 cm sensitivity test showed the opposite: as B rose, the lower ladder deteriorated monotonically, additional valid marks crossed to zero and the upper scale paid a small but unnecessary cost.
Lowering B could mechanically reduce the strong negative pole-vault residual, but that would mainly compensate for a real population effect: at lower decathlon levels, pole vault deteriorates faster than many other events.
The inherited value therefore survives the audit. Pole vault remains at B = 100 cm.
Javelin is the event that first exposed the freedom of B, and it provides the clearest example of why a statistical optimum cannot be accepted blindly. As B rises from the inherited 7 m, the large positive lower-tail residual steadily shrinks. Mechanical RMSE continues improving well into the high teens.
But javelin has a strong positive population effect at low decathlon levels. Many weak decathletes remain relatively competent throwers. A positive residual is therefore partly real. Higher B values also push increasingly many valid throws onto the zero side and drive C far below 0.7. The curve begins to solve the statistical residual by removing too much meaningful lower scoring space.
The first stable region in which the most severe lower-tail distortion largely disappears is around 12 m. It is deliberately more conservative than the mechanical optimum and leaves the known positive population effect visible.
I therefore adopt B = 12.0 m, with C approximately 0.717.
Only now can the Part III javelin comparison be completed. The same upper scale has led us through three very different assumptions about the basement: the constrained Part III control at -6 m, the natural FDM-40N inheritance at 7 m, and the lower-tail audit at 12 m.
|
Javelin |
FDM-40C |
FDM-40N |
FDM-B |
|
10 m |
195 |
82 |
0 |
|
12 m |
219 |
124 |
0 |
|
15 m |
256 |
180 |
111 |
|
20 m |
318 |
265 |
225 |
|
30 m |
441 |
418 |
404 |
|
40 m |
565 |
558 |
554 |
|
48.96 m |
676 |
676 |
676 |
|
50 m |
689 |
690 |
690 |
|
60 m |
813 |
815 |
816 |
|
65.82 m |
885 |
885 |
885 |
|
70 m |
937 |
935 |
934 |
At 50-70 metres the three curves are still almost indistinguishable. At 10-20 metres they describe radically different scoring systems. That is why the lower data were needed in the first place.
The 1500 m is the strongest example of an optimiser answering the wrong question correctly. Lower-level decathletes are often relatively better at endurance running than their technical-event level would predict. The event therefore develops a strong positive residual.
If B is allowed to chase that residual too aggressively, the preferred zero threshold moves dramatically upward. Exploratory fits pushed the 1500 m B toward approximately 547-566 seconds, or roughly 9:07 to 9:26.
Mathematically, such a curve can fit the population. Sportingly, it is difficult to justify. The model would be rebuilding an event function around a known characteristic of lower-level decathletes rather than diagnosing the inherited threshold.
The inherited eight-minute boundary therefore survives. The 1500 m remains at B = 480 s. The fact that the weakest recovered performance is slower than nine minutes does not change that conclusion: valid completion and positive scoring are separate questions.
Only after all ten events have been examined individually can the final pattern be stated. Seven inherited thresholds move; three survive unchanged.
|
Event |
FDM-40N B |
Audited B |
Audited C |
Decision |
|
100 m |
18.0 s |
16.0 s |
1.751913 |
revised |
|
Long jump |
220 cm |
220 cm |
1.606329 |
unchanged |
|
Shot put |
1.50 m |
4.75 m |
0.748259 |
revised |
|
High jump |
75 cm |
110 cm |
1.028153 |
revised |
|
400 m |
82.0 s |
77.0 s |
1.999033 |
revised |
|
110H |
28.5 s |
24.9 s |
1.995062 |
revised |
|
Discus |
4.0 m |
7.0 m |
0.854170 |
revised |
|
Pole vault |
100 cm |
100 cm |
1.108810 |
unchanged |
|
Javelin |
7.0 m |
12.0 m |
0.716812 |
revised |
|
1500 m |
480 s |
480 s |
1.390723 |
unchanged |
That asymmetry matters. The purpose of the exercise was never to manufacture ten new numbers. An unchanged B is just as meaningful a result when alternative values have been tested and rejected.
There is also no surviving universal rule for C. The audited exponents range from approximately 0.717 to almost 2. The Part III experiment that constrained every event to 1 < C < 2 remains useful as a sensitivity test, but it is not supported as a universal parameter-selection principle.
Only at this point does another Part III echo become visible. Three of the constrained values, obtained a month earlier from a completely different question, now sit strikingly close to the independently audited thresholds.
|
Event |
FDM-40C B |
Lower-tail audit B |
|
100 m |
16.6 s |
16.0 s |
|
400 m |
77.0 s |
77.0 s |
|
110H |
24.9 s |
24.9 s |
This is not validation of FDM-40C. The same constrained experiment produced B = 1 m in discus and B = -6 m in javelin, neither of which survives the lower-tail audit. The safer interpretation is narrower: the upper geometry had already signalled some of the directions that the lower population later rediscovered independently.
For reproducible scoring, the following six-decimal coefficients define the audited model. These are the public coefficients: event points are calculated from them and then truncated to the integer below.
|
Event |
A |
B |
C |
|
100 m |
47.538494 |
16.0 |
1.751913 |
|
Long jump |
0.036021 |
220 |
1.606329 |
|
Shot put |
149.409375 |
4.75 |
0.748259 |
|
High jump |
7.830960 |
110 |
1.028153 |
|
400 m |
1.058283 |
77.0 |
1.999033 |
|
110H |
7.805308 |
24.9 |
1.995062 |
|
Discus |
37.000168 |
7.0 |
0.854170 |
|
Pole vault |
1.125696 |
100 |
1.108810 |
|
Javelin |
50.887984 |
12.0 |
0.716812 |
|
1500 m |
0.496015 |
480 |
1.390723 |
Long jump, high jump and pole vault use centimetres inside the formula. The throws use metres. Running events use seconds.
Track events: P = A(B - M)^C
Field events: P = A(M - B)^C
A mark on the non-scoring side of B receives zero. The calculated event score is truncated to the integer below.
The preliminary tests suggested that the answer should be no, but the completed ten-event audit still needs to be checked against familiar elite profiles.
Using the public six-decimal coefficients, the same five all-time profiles used in earlier FDM stress tests change only slightly:
|
Athlete |
OFF |
FDM-40N |
FDM-B |
Change |
|
Kevin Mayer (2018) |
9126 |
9109 |
9103 |
-6 |
|
Ashton Eaton (2015) |
9045 |
9144 |
9139 |
-5 |
|
Roman Sebrle (2001) |
9026 |
9052 |
9049 |
-3 |
|
Damian Warner (2021) |
9018 |
9133 |
9126 |
-7 |
|
Tomas Dvorak (1999) |
8994 |
9031 |
9022 |
-9 |
The entire B audit changes these complete elite totals by only 3-9 points. That is tiny compared with the scale of the intervention near the bottom and confirms the central geometric observation of the exercise: once the upper anchors are protected, substantial changes in B primarily alter the basement.
The independent 75th-position control tells the same qualitative story. It was not used as a calibration anchor in FDM-40N, and re-evaluating it after the B audit leaves the common middle level and event spread essentially unchanged. The model has therefore not been recalibrated to weak decathletes. The lower population has been used to inspect a part of the curve that the upper population could not identify.
Part VII gives us one final test that did not exist when the B project began. Its FDM-7000- experiment independently sorted all ten event distributions in the same 10,720-performance lower database. Two ranks, R650 and R3250, were used as calibration anchors for that separate experiment; six others were left unused. The bespoke FDM-7000- model remained extraordinarily compact through R5200, weakened at R7800 and deteriorated sharply at R10400.
We can now return to exactly the same raw marks and ask a different question. What happens when those marks are scored by the ordinary FDM-40N curves and by the newly audited FDM-B curves? No lower-rank recalibration is performed.
|
Rank |
OFF RMSE |
FDM-40N RMSE |
FDM-B RMSE |
|
R130 |
63.80 |
29.64 |
29.47 |
|
R325 |
66.76 |
30.76 |
30.59 |
|
R650 |
69.98 |
31.59 |
31.48 |
|
R1950 |
75.25 |
32.62 |
32.77 |
|
R3250 |
77.87 |
34.06 |
34.45 |
|
R5200 |
79.74 |
37.11 |
37.89 |
|
R7800 |
79.92 |
53.30 |
55.11 |
|
R10400 |
82.64 |
84.86 |
85.73 |
At first sight this may look disappointing. The complete ten-event RMSE does not improve substantially after B is corrected. At R7800 and R10400 it is even slightly larger.
But that is exactly why RMSE was never allowed to choose B mechanically. Part VII already showed that the deep loss of compactness is not a uniform ten-event failure. At R7800, nine events remained relatively close while the hurdles separated dramatically. By R10400 the hurdles were roughly 175 points below the common mean. Population structure was already dominating the picture.
Changing B should not erase that structure merely to produce a smaller headline statistic.
The event-by-event results above reinforce a contrast already visible in Part VI: at low levels, the hurdles become disproportionately weak, while endurance running often remains disproportionately strong. This suggests a useful diagnostic test of how much of the remaining divergence reflects the athletes rather than the scoring functions.
If 110H and 1500 m are treated as a single diagnostic package, H1500 = (P110H + P1500) / 2, while the other eight events remain separate, the comparison becomes a nine-unit RMSE instead of a ten-event RMSE.
|
Rank |
FDM-40N |
FDM-B |
|
R130 |
21.65 |
21.66 |
|
R325 |
21.74 |
21.75 |
|
R650 |
22.13 |
22.20 |
|
R1950 |
21.85 |
22.08 |
|
R3250 |
21.42 |
21.72 |
|
R5200 |
20.64 |
20.77 |
|
R7800 |
24.10 |
21.59 |
|
R10400 |
54.74 |
43.32 |
Through most of the distribution, nothing meaningful changes. At the deepest ranks, it does. At R7800 the package RMSE falls from 24.10 to 21.59. At R10400 it falls from 54.74 to 43.32.
That is a more informative result than forcing the raw ten-event RMSE downward. The B audit has not made lower-level decathletes artificially flat. It has reduced some of the scoring-curve distortion while leaving the strongest athlete-population effects visible.
The basement becomes cleaner. The people standing in it remain complicated.
The original purpose of FDM was to ask whether comparable performances in different decathlon events receive comparable nominal value. For several parts, B was largely invisible to that question. The upper data were sufficiently far from zero that inherited thresholds could remain in place without materially affecting the conclusions.
The lower database changed that. B turned out to be one of the most weakly identified parameters from the upper scale and one of the most powerful parameters near the bottom. The javelin makes this almost absurdly clear: B values separated by many metres can describe nearly the same high-level relationship while producing radically different values for weak throws.
That freedom cannot be resolved from elite performances alone. But neither can it be resolved by blindly minimising lower residuals. The final B values therefore come from a combination of fixed upper geometry, actual lower distributions, extreme-tail plausibility, residual structure, zero-side behaviour and sporting interpretation.
Seven inherited thresholds fail that audit. Three survive it. Several resulting exponents fall below 1. Two are almost exactly 2. One is almost exactly 1. There is no universal preferred curvature.
A historical comparison is also worth noting. The 1962 scoring tables used throwing formulas that were linear in the square root of the measured distance, with zero-point marks of 4.70 m in shot put, 12.81 m in discus and 14.02 m in javelin. The 1985 tables instead use progressive power functions with exponents above 1. In FDM-B, the corresponding exponents are 0.748, 0.854 and 0.717, with B values of 4.75 m, 7.00 m and 12.00 m respectively.
The function should describe the event. The event should not be forced to satisfy a preferred exponent.
I did not return to B because it promised an especially attractive extension of FDM, but because leaving it unexamined would have left the model incomplete.
The inherited thresholds had survived the earlier stages largely because the upper calibration did not require them to be questioned. Once the database reached far enough into the lower tail, that was no longer satisfactory. If FDM was to be completed on its own terms, the zero thresholds had to be examined rather than assumed.
That task is now finished.
The resulting curves are not claimed to reveal ten uniquely correct physical boundaries. In several events the evidence supports a region rather than one magic number. What the audit does provide is a complete model in which no inherited B remains merely because it was already there.
There is also one methodological choice that has remained unchanged throughout all eight parts of this project. FDM has never used specialist world records to determine the relative value of the ten events. Its reference points, validation tests and now its zero-threshold audit have all come from decathlon itself.
That was deliberate.
I will return to the reasons why in Part IX.