Record AR-CMP-001 · Comparative method
Two configurations,
one honest difference.
A single configuration always looks reasonable on its own. The useful question is never "is this good" but "is this better than the thing I would otherwise run, and at what cost". This arena holds two complete configurations against the same engine, the same seed and the same horizon, and reports every difference between them — including the ones that moved in the wrong direction.
Interactive demonstration — not a historical backtest
Configuration slots
A Baseline
default configurationB Challenger
default configurationSaved experiments
Stored in this browser only03VerdictAR-CMP-001·03
Comparative reading
Illustrative—
—
04Measure by measureAR-CMP-001·04
Every measure, both sides,
and the signed difference.
A change that improves expectancy while deepening drawdown has not made the system better; it has moved the system along a trade-off. The direction column states which way each difference actually points.
| Measure | A · baseline | B · challenger | Difference | Direction |
|---|
Figure A · Relative movement, every measure
IllustrativeEach measure moves on its own scale — a 0.3 R change in expectancy and a 0.3 percentage point change in drawdown are not comparable numbers. Each difference is therefore shown as a percentage of side A's own value, so the bars can be read against each other. Green is an improvement in that measure's own preferred direction, red a regression; measures with no preferred direction are drawn in slate because calling them either would be a judgement the model has not made.
—
Figure B · The trade-off plane
Expectancy against drawdown, both sides plotted on one surface. Up and to the right is more return for more decline; the useful question is not which point is higher but which direction the move between them travels, and whether that direction is one you can hold.
—
Figure C · Component profile
The eight weighted components of the preview score as a shape rather than a ranking. Two configurations can reach a similar total by pinching different axes, and a profile with one deep notch is a different proposition from one that is evenly moderate.
—
05Overlaid historiesAR-CMP-001·05
The same seed,
two different records.
Both records are generated from an identical random sequence, so the divergence between the two lines is structural rather than coincidental. Neither line is a forecast.
Figure 1 · Equity index, both sides
—
Figure 2 · Score components, both sides
—
Figure 3 · Per-parameter attribution
IllustrativeEvery differing parameter is moved from its A value to its B value on its own, against an otherwise unchanged A. The bar is the score movement that single rail produces by itself. What the individual movements do not account for is charted separately as interaction — the part of the result that exists only when the changes are made together.
—
Figure 4 · Verdict across many sequences
Same pairs, both sidesFigure 1 shows one sequence. One sequence can flatter either side. Here both configurations are run against the same set of seeded histories, pair by pair, and the question becomes how often B finishes ahead rather than whether it happened to.
—
Figure 5 · Regime stress
Eight regimesBoth configurations are re-scored under each of the eight regimes, with dispersion charged to slippage and spread and with correlation imposed rather than chosen. The width of a side's range is its regime dependence: a narrow range survives being wrong about the environment, a wide one does not.
—
| Regime | A score | B score | A permission | B permission |
|---|
06Parameter differenceAR-CMP-001·06
Exactly what changed
between the two.
Only the parameters that actually differ are listed. If this table is long, the comparison is not a controlled experiment — it is two different systems, and the difference in outcome cannot be attributed to any single decision.
| Parameter | Group | A | B | Drives |
|---|
—
What this comparison can and cannot settle
It can settle direction: which of two configurations the model rates higher, on which measures, and by how much. It can expose a trade-off that a single-configuration view hides — capture bought with tail, drawdown bought with frequency, expectancy bought with exposure.
It cannot settle whether either configuration is profitable. Both sides are illustrative models of relationships between design decisions. Neither reads a market, a broker feed or a trade journal. A comparison between two models is a comparison between two models.
Controlled comparison
- Change one parameter, or one small group of related parameters.
- Read the direction column, not only the score.
- Check what the improvement was paid for with.
- Keep the link — it is the record of the experiment.