Perception-XAlpha Lite treats academic literature as a source of testable mechanisms and research constraints, not as proof that an implementation has alpha. This document records:
- the claim supported by each reference;
- the corresponding public implementation;
- the boundary the implementation must not cross.
No empirical research result is included.
The separately generated User-Supplied Literature Registry tracks additional papers proposed during development. It distinguishes partial mappings, direct extension candidates, data-gated mechanisms, deferred architectures, and reviewed out-of-scope material. CI resolves every public-code mapping so a citation cannot silently be presented as an implementation.
| Method | Evidence from the literature | Public implementation | Explicit boundary |
|---|---|---|---|
| Probability of Backtest Overfitting | CSCV estimates the probability that selection among many backtests produces an in-sample winner that underperforms out of sample | pbo() |
PBO cannot repair contaminated data, invalid labels, or an understated trial count |
| Deflated Sharpe Ratio | DSR corrects apparent Sharpe evidence for multiple testing and non-normal returns | deflated_sharpe_ratio() |
All attempted candidates, including failures, must enter the trial ledger |
| Purged chronological validation | Overlapping outcome horizons can leak information across adjacent train/test samples | make_split() and purged walk-forward folds |
Purge length must cover the maximum label horizon |
| Counterfactual and placebo controls | A candidate should outperform mechanism-specific alternatives and a permutation distribution | Primary/Counter/Placebo evaluation | One lucky placebo draw is insufficient; the empirical distribution is required |
| Proper probability scoring | Strictly proper scores incentivize honest probabilistic forecasts | probability_metrics() |
AUC alone is insufficient; Brier, LogLoss, and calibration error are reported separately |
| Stationary bootstrap | Geometrically distributed circular blocks preserve weak time dependence under resampling | stationary_bootstrap_indices() and mean intervals |
Block length must be frozen and weak stationarity must be defensible |
| Reality Check | The best member of a searched family must be tested against the joint data-snooping null | white_reality_check() |
Global rejection does not identify an executable strategy |
| Step-down max-t | Joint resampling and studentized step-down tests control family-wise error more powerfully than single-step correction | romano_wolf_stepdown() |
Candidate family and benchmark must be frozen before testing |
| False discovery rate | BH controls FDR under its dependence conditions; BY adds a conservative arbitrary-dependence correction | benjamini_hochberg_qvalues() and benjamini_yekutieli_qvalues() |
FDR is complementary to, not a replacement for, family-wise inference |
- Bailey, D. H., Borwein, J. M., López de Prado, M., & Zhu, Q. J. (2015). The Probability of Backtest Overfitting. Journal of Computational Finance. SSRN · DOI
- Bailey, D. H., & López de Prado, M. (2014). The Deflated Sharpe Ratio: Correcting for Selection Bias, Backtest Overfitting and Non-Normality. Journal of Portfolio Management, 40(5), 94–107. SSRN · DOI
- Gneiting, T., & Raftery, A. E. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. Journal of the American Statistical Association, 102(477), 359–378. DOI
- Politis, D. N., & Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association, 89(428), 1303–1313. DOI
- White, H. (2000). A Reality Check for Data Snooping. Econometrica, 68(5), 1097–1126. DOI
- Romano, J. P., & Wolf, M. (2005). Stepwise Multiple Testing as Formalized Data Snooping. Econometrica, 73(4), 1237–1282. DOI
- Benjamini, Y., & Hochberg, Y. (1995). Controlling the False Discovery Rate: A Practical and Powerful Approach to Multiple Testing. Journal of the Royal Statistical Society: Series B, 57(1), 289–300. DOI
- Benjamini, Y., & Yekutieli, D. (2001). The Control of the False Discovery Rate in Multiple Testing under Dependency. Annals of Statistics, 29(4), 1165–1188. DOI
The complete method contract and command-line workflow are documented in EVIDENCE_LAB.md.
| Research question | Literature contribution | Public implementation | Deliberate simplification |
|---|---|---|---|
| How can a system explore formulaic factors while retaining interpretability? | AlphaForge uses a generative-predictive architecture for formulaic factor generation and combination | allowlisted JSON DSL, economic seeds, bounded mutation/crossover | No neural generator; every expression is statically inspectable |
| How can factor lineage remain reproducible? | Evolutionary and grammar-guided research motivates explicit generation histories | immutable factor ID, parent IDs, operator, generation, expression hash | No claim that lineage quality implies predictive quality |
| How should redundant candidates be handled? | Diverse factor sets reduce repeated variants of the same behavior | train-only behavioral-correlation pruning | Validation/shadow correlations never select the parent pool |
| How should proposal strategies be compared without turning an LLM into a factor judge? | AlphaBench separates alpha generation from empirical evaluation and compares structured search paradigms | preregistered CoE/ToT/EA schedule, common DSL/budget, train-only pairwise selection and per-arm audit | Paper-inspired orchestration only; reported AlphaBench performance is not reproduced and absolute zero-shot factor judgement is prohibited |
- Shi, H., Song, W., Zhang, X., Shi, J., Luo, C., Ao, X., Arian, H., & Seco, L. (2024). AlphaForge: A Framework to Mine and Dynamically Combine Formulaic Alpha Factors. arXiv:2406.18394
- Zhang, T., et al. (2020). AutoAlpha: An Efficient Hierarchical Evolutionary Algorithm for Mining Alpha Factors in Quantitative Investment. arXiv:2002.08245
- AlphaBench: A Benchmark for LLM-Driven Alpha Mining (2026). ICLR 2026. Conference paper
The public protocol and its interpretation boundary are documented in SEARCH_PROTOCOL.md.
Prediction error is not identical to downstream decision loss. The optional decision toolkit therefore fits a bounded Top-K pairwise objective after the factor definitions and directions are frozen.
| Method | Public implementation | Boundary |
|---|---|---|
| Decision-focused learning | fit_pairwise_topk_weights() |
This is a transparent pairwise surrogate, not an implementation of the full SPO optimizer |
| Weight sensitivity | fit_pairwise_weight_ensemble() |
Complete-block replicas measure instability; they are not a posterior distribution |
| Expected-outcome calibration | fit_ridge_score_model() |
Fitted only on the independent calibration block |
| Event-probability calibration | fit_logistic_probability_model() |
Ordinary and severe loss events are separate models |
| Reliability evaluation | probability_metrics() |
Validation/shadow may reject but never refit the calibrator |
- Elmachtoub, A. N., Liang, J. C. N., & McNellis, R. (2020). Decision Trees for Decision-Making under the Predict-then-Optimize Framework. Proceedings of Machine Learning Research, 119, 2858–2867. PMLR
- Gneiting, T., & Raftery, A. E. (2007). Strictly Proper Scoring Rules, Prediction, and Estimation. DOI
- Politis, D. N., & Romano, J. P. (1994). The Stationary Bootstrap. Journal of the American Statistical Association, 89(428), 1303–1313. DOI
| Mechanism | Literature claim | Public implementation | Boundary |
|---|---|---|---|
| Hidden Markov regime model | Time-series parameters may depend on an unobserved discrete-state Markov process | causal_two_state_hmm_filter() |
Forward filtering only; full-path Viterbi labels use future observations |
| Early-warning signals | Rising variance and autocorrelation may accompany critical slowing before some transitions | early_warning_features() |
A generic warning is not directional return alpha |
| Bayesian online change-point detection | Run-length posteriors support online inference of abrupt generative changes | bocpd_change_probability() |
Hazard assumptions and observation models must be frozen |
- Hamilton, J. D. (1989). A New Approach to the Economic Analysis of Nonstationary Time Series and the Business Cycle. Econometrica, 57(2), 357–384. DOI
- Scheffer, M., et al. (2009). Early-Warning Signals for Critical Transitions. Nature, 461, 53–59. DOI
- Adams, R. P., & MacKay, D. J. C. (2007). Bayesian Online Changepoint Detection. arXiv:0710.3742
| Mechanism | Literature claim | Public implementation | Boundary |
|---|---|---|---|
| Dynamic Mode Decomposition | DMD approximates dynamics through the eigendecomposition of a fitted linear operator and connects to Koopman analysis | causal_dmd_residual() |
Fixed past window; linear consistency and rank deficiency require audit |
| LPPLS | Reduced-parameter calibration can improve the numerical stability of log-periodic power-law fitting | simplified_lppls_features() |
Fixed grid, residuals, and t_c stability distribution; no exact crash-date forecast |
- Tu, J. H., Rowley, C. W., Luchtenburg, D. M., Brunton, S. L., & Kutz, J. N. (2014). On Dynamic Mode Decomposition: Theory and Applications. Journal of Computational Dynamics, 1(2), 391–421. arXiv:1312.0041
- Filimonov, V., & Sornette, D. (2013). A Stable and Robust Calibration Scheme of the Log-Periodic Power Law Model. Physica A, 392(17), 3698–3707. arXiv:1108.0099 · DOI
| Mechanism | Public implementation | Boundary |
|---|---|---|
| Black–Litterman view shrinkage | black_litterman_posterior() |
Stabilizes uncertain views; cannot create predictive content |
| Triple Barrier labeling | triple_barrier_labels() |
Uses future paths by design and belongs only in an offline label table |
| Hoeffding lower bound | hoeffding_lower_bound() |
Independence assumptions and effective sample size remain explicit limitations |
- Black, F., & Litterman, R. (1992). Global Portfolio Optimization. Financial Analysts Journal, 48(5), 28–43. CFA Institute · DOI
- López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley. Triple Barrier and purged validation are represented as offline research tools.
- Hoeffding, W. (1963). Probability Inequalities for Sums of Bounded Random Variables. Journal of the American Statistical Association, 58(301), 13–30. DOI
| Direction | Public-package boundary | Reference |
|---|---|---|
| Instrumented PCA | PIT characteristics and neutral portfolios are available; full IPCA estimation is not included | Kelly, Pruitt & Su, Characteristics Are Covariances (published article) |
| Conditional factor timing | Block-replica weights are available; validation-driven or online timing is forbidden | Factor Timing with Portfolio Characteristics (article) |
| Neural formula generation | The generator is bounded and symbolic; it does not reproduce AlphaForge's neural architecture | Shi et al., arXiv:2406.18394 |
| Gaussian-process ensembles | Probability calibration is included; Gaussian-process forecasting is not | Ensemble Gaussian Process Regression for Time Series Forecasting (arXiv:2212.01048) |
The project also maintains a machine-readable audit of submitted research directions, including CogAlpha, PRISM-VQ, Kronos, AI-Trader, heavy-tail HMMs, HSMM duration models, Deep LPPLS, Hawkes reflexivity, Koopman-based stochastic resilience, dependent concentration, and heavy-tail martingale bounds.
- Rendered registry
- Public JSON
- Source of truth:
research/literature_registry.json - Schema:
research/literature_registry.schema.json
The registry records a next falsifiable step for each relevant paper and retains explicit exclusions when no causal financial mechanism can be defended. It contains no empirical performance result.
- A cited paper does not validate this implementation on a new market or dataset.
- A mechanism primitive is not a production model.
- Better calibration is not a guarantee of profit.
- Lower turnover or fewer selections is not automatically alpha.
- Historical validation is not fresh forward evidence.
- This repository publishes no empirical factor result and defines no execution path.