PEG-Bench

When Prediction Leaderboards Can Fail

Benchmarking the Prediction-Execution Gap

Prediction rankings can reverse after costs, turnover, and risk.

Visual abstract / six-panel result

Leader-shift slopegraph showing LightGBM as the R0 top predictor and four Kairos-64 leader changes at clean-active R4
4 of 6active leader identities change at 3 bps
01

Real-world problem

A good offline score can answer the wrong decision question.

In plain language

A forecast is not the same thing as an action. Two models can predict the same direction, but one may change position much more often. Fees and slippage can erase that model's apparent advantage before the decision produces a useful outcome.

The problem is not limited to markets. It appears whenever a prediction is converted into a constrained action: a treatment choice, a recommendation, a resource allocation, or a control signal.

Formal question

Given the same forecasts, chronological split, forecast-to-action interface, cost schedule, and risk rules, does the R0 prediction leader remain the R4 execution-aware leader?

PEG-Bench treats this as an evaluation problem. It does not treat the market setting as evidence of a deployable trading strategy.

Prediction score -> constrained action -> realized utility

Offline metricWhat did the model predict?
Action interfaceWhat position or action did it produce?
Realistic outcomeWhat survived costs, risk, and stress?
02

Plain-language concept

The prediction-execution gap is the distance between a forecast and what survives a decision.

Prediction score -> constrained action -> realized utility

A

Forecast

A model assigns scores or predicts a future change. This is the familiar offline leaderboard.

B

Action

A policy turns the score into an exposure, recommendation, or control move with a frequency and size.

C

Outcome

The action meets fees, slippage, drawdown limits, and distribution shift. Utility is measured here.

R0

Prediction proxy: RankIC, directional accuracy, and errors.

R1

Costs: the fee schedule is applied to the same action path.

R2

Execution: turnover, slippage, and action activity become visible.

R3

Risk: drawdown and path-dependent utility are included.

R4

Stress: regimes, survival, and robustness test the conclusion.

03

Experiment design

One frozen contract makes the comparison inspectable.

The completed evidence slice is deliberately narrow. That narrowness lets a reader distinguish a measured result from a future benchmark target.

Frozen PEG-Bench contract and its reader-facing purpose
DimensionFrozen specificationWhy it matters
Data window2025-09-28 to 2025-12-26 UTC; BTCUSDT and ETHUSDT; Binance and OKX.Fixes the population before comparing methods.
Chronology60 / 15 / 15 train, validation, and test; 15 held-out test days.Prevents future information from entering training or tuning.
PanelsTime and volume views at horizons 5, 20, and 50.Shows whether a conclusion holds across six matched views.
Friction0 / 1 / 3 / 5 / 10 bps cost ladder; 3 bps headline report card.Tests whether small implementation costs change the ranking.
Evidence unitOne view-horizon panel with R0, R4, gap, risk, stress, and audit fields.Keeps every claim tied to a traceable result object.
Controlled diagnostic contract connecting a frozen 90-day six-panel data design, shared action interface, R0-to-R4 diagnostic, unified report card, and the bounded result that four of six clean-active leader identities change at 3 bps without establishing trading alpha
Controlled diagnostic contractFrozen data and interfaces lead to one traceable report card and a bounded selection conclusion.

Cost sensitivity

Costs are part of the evaluation object.

0 / 1 / 3 / 5 / 10 bps

Fixed fee ladder available in the normalized 90-day reports.

3 bps

Headline setting for the clean-active R4 comparison and displayed report card.

Sensitivity status

Other fee settings are released as supporting sensitivity evidence, not as separately rerun headline claims.

04

Mechanism interaction

Why can the ranking change?

The gap is not one hidden penalty. It is an interaction between the signal, the action rule, the path of positions, and the conditions under which those positions are evaluated.

Mechanisms that connect an offline score to an execution-aware outcome
MechanismPlain-language effectReport-card signal
Action frequencyA small forecast difference can trigger many more position changes.Turnover, fees, slippage
Position pathThe same average prediction can create a very different drawdown path.Sharpe, maximum drawdown, recovery
Regime shiftA rule that worked in one market condition can lose its ordering in another.Worst-regime utility, stress survival
Safety floorsNo-action or anchor rows can look strong on a utility floor without being active model winners.Clean-active eligibility rule

Toy illustration

Model A can have the higher RankIC but switch positions every bar. Model B can have a lower RankIC and hold a steadier position. Once a fee is charged for each switch, B may have the stronger R4 point estimate. This is the mechanism PEG-Bench measures; it is not a claim that either model is profitable in production.

05

Evidence results

The six-panel result is a change in identity, not a promise of profit.

LightGBM is the RankIC leader among populated prediction-first rows in all six panels. At 3 bps, the clean-active R4 identity changes in four panels. The confidence intervals below make clear where the R4 point estimates remain uncertain.

Official prediction-first

All normalized prediction-first rows under the frozen split. RR@1 uses deterministic stable sorting at k=1.

Clean-active R4

Active model identity after zero-action safety-floor rows and non-active anchors are excluded from the active comparison.

Diagnostic

Wrapper and formal-method rows inform sensitivity but do not replace the clean ranked leaderboard.

Target only

Method families without normalized 90-day evidence under the shared evaluator.

Six-panel comparison showing four active leader identity changes after 3 bps accounting
Four of six active selections change identity. LightGBM leads all R0 panels; Kairos-64 leads four clean-active R4 panels at 3 bps.
R0 RankIC leader and clean-active R4 leader for the six completed panels. R4 intervals are day-block 95% confidence intervals.
PanelR0 top predictorR0 RankICR4 top active methodR4 Sharpe at 3 bpsDay-block 95% CIResult
Time h5LightGBM0.145LightGBM-0.432[-0.630, -0.239]same; negative
Time h20LightGBM0.079Kairos-64+0.205[-0.137, +0.613]leader changed
Time h50LightGBM0.047Kairos-64-0.207[-0.529, +0.096]leader changed
Volume h5LightGBM0.274Kairos-64+0.069[-0.174, +0.325]leader changed
Volume h20LightGBM0.220LightGBM-0.072[-0.331, +0.012]same; negative
Volume h50LightGBM0.171Kairos-64+0.056[-0.256, +0.378]leader changed

Positive R4 point estimates have intervals crossing zero. The table establishes leader-identity changes and evidence boundaries, not statistically secure positive utility.

Protocol note

Official RR@1 and the clean-active identity audit answer different questions.

Official RR@1 is 1.0 across the six prediction-first panels when k=1 and deterministic stable sorting are used. It is a protocol-level gap diagnostic. The clean-active R4 audit asks whether the active R4 model identity changes after non-active safety-floor rows are excluded from the active model comparison. Neither value should be substituted for the other.

Panel bootstrap

Mean Ranking Gap: -0.541 with 95% CI [-0.852, -0.259]; winner-change rate 0.667 with 95% CI [0.333, 1.000]. These estimates use six panels and 1,000 panel resamples. They are not day-level inference.

Open the bootstrap table
Heatmap of clean-active R4 Sharpe across methods and time or volume panels
Supporting view. Utility mapPost-cost utility is uneven and often negative.
Scatter plots separating R0 RankIC, execution activity, and R4 Sharpe
Supporting view. Separate axesSignal, turnover, and utility describe different objects.
06

Limitations

What the released evidence does not establish.

Scope limits that qualify the main conclusion
LimitConsequenceCorrect reading
90-day evaluation windowHeadline comparisons use one frozen 90-day window.Persistence beyond this period and market regime remains untested.
Licensed market dataRaw L2 feeds and reconstructable tensors are not redistributed.Inspect normalized evidence and evaluator code; do not infer raw-data access.
Six populated panelsThe result is a completed evidence slice, not a universal market claim.Read the conclusion within the fixed dates, assets, venues, views, and horizons.
Confidence-interval uncertaintyEvery positive R4 point estimate has a day-block 95% confidence interval that crosses zero.The finding is a ranking-identity and evidence-boundary result, not statistically secure positive utility.
Clean-active eligibilityClean-active selection excludes zero-action safety-floor rows and non-active anchors.Compare identities only within this eligible active-method universe; keep official RR@1 separate.
Incomplete method panelSeveral target families have no normalized 90-day result rows.Target-only methods are explicitly not claimed as completed baselines.
Evaluation, not deploymentApproximate frictions and frozen replay do not reproduce production operations.Use the benchmark to audit evaluation conclusions, not to claim deployable returns or readiness.
RR@1 diagnosticRR@1=1.0 uses k=1 and deterministic stable sorting in the official universe.Keep it separate from the clean-active R4 identity result.

Completed evidence

Six-panel prediction-first comparison

LightGBM, Kairos, DeepLOB, and Ridge/HistGBM-style controls are reported where implemented under the shared evaluator.

Diagnostic evidence

Wrappers and formal-method slots

Execution wrappers and OG-PA/DPO rows are informative sensitivity evidence, outside the clean active leaderboard.

Not claimed

Full method-panel completion

Additional target families do not yet have normalized 90-day evidence under the same evaluator.

07

Engineering applications

Same gap, different stakes.

Illustrative transfer scenarios. These domains were not evaluated in the released benchmark.

AI coding

Prediction
Code-generation or patch-quality score
Action
Accepted patch and tool calls
Realized utility
Build, test, security, latency, and cost outcomes
Failure cost
Regression, unsafe change, or wasted review time

Finance

Prediction
Return forecast
Action
Sizing, rebalance, and order path
Realized utility
Fees, slippage, turnover, drawdown, and risk
Failure cost
Capital loss and instability

Healthcare

Prediction
Risk or treatment score
Action
Clinician or policy action
Realized utility
Contraindications, workflow capacity, delay, and patient outcomes
Failure cost
Missed or harmful intervention

Cybersecurity

Prediction
Threat score
Action
Block, isolate, or escalate action
Realized utility
False-positive load, response latency, and service continuity
Failure cost
Breach or business interruption

Energy and industrial control

Prediction
Demand or fault forecast
Action
Scheduling or control action
Realized utility
Stability, peak cost, reserve, and safety limits
Failure cost
Outage, damage, or unsafe operation

Engineering boundary: PEG-Bench is an evaluation protocol and audit artifact. It does not replace domain-specific validation, governance, or production risk review.

08

Materials

Every public claim has a corresponding inspection route.

The technical report explains the method and appendix. The repository routes below expose the normalized evidence, implementation, provenance, integrity manifest, and release boundary needed for an audit.

Reference record

PEG-Bench (2026). When Prediction Leaderboards Fail: Benchmarking the Prediction-Execution Gap. Technical report.

https://computational-decision-lab.github.io/peg-bench/

Public release boundary

Auditable does not mean unrestricted.

This public release includes the evaluator, schemas, split and stress manifests, normalized evidence, generated tables, figures, and inspection samples.

It does not redistribute raw vendor L2 feeds, tick-level trades, or full processed tensors derived from licensed data. Full raw-data reconstruction requires authorized access to the underlying market-data provider.

Read the release scope