
In high-frequency trading data, noise isn't the problem. Assumptions are.
7/24/2026
This post introduces the problem of brittle data preprocessing in high-frequency trading, where fixed strategies fail due to non-stationary market conditions. It reframes the problem as representation learning, highlighting research on learning limit order book representations (e.g., SimLOB). The post outlines Lambda's exploration into using a Bayesian framework to learn market structure directly from order book data. This learned structure is intended to drive adaptive preprocessing, distinguishing between noisy and novel data using statistical divergence. The contribution emphasizes the practical benefits for Lambda's customers: reduced training time, more reliable convergence, and more efficient GPU utilization by improving upstream data quality and regime awareness.