I'm currently renovating my website… please be patient…

projects→order-book-trading-signal

Looking for signal in the noise — and accepting what the data says

A live data pipeline and ML experiments to test whether Bitcoin's order book can predict short-term price moves.

golden magnifying glass over rice

Status: In progress · Stack: Python, Binance API, SQLite, Hetzner VPS, scikit-learn ⚠️

The problem

The problem

Can the order book (the live list of buy and sell orders waiting on an exchange) tell you where the price will go next? Many trading strategies assume it can, but testing it is harder than it sounds. Historical price data (candles, volumes, trades) is easy to find and often free. Historical order-book data isn’t: exchanges don’t keep it for you, and what you can buy is costly or incomplete. So if you want to test the idea properly, you first have to build the dataset yourself, one snapshot at a time.

I wanted to know for sure, with my own data and proper tests, before putting real money behind it: would a 1% move up or a 1% move down come first?

What I built

A data collector running 24/7 on my own server, taking a snapshot of the Bitcoin order book (BTC/USDT, 100 price levels) every minute, synced to the candle boundaries. On top of that data, a series of machine-learning experiments to see whether those snapshots predict which way the next 1% move goes.

How it works

  1. Collector on a Hetzner VPS fetches order-book depth from the Binance API every minute, aligned with the candles.
  2. SQLite in WAL mode, with a unique index on symbol + timestamp, so the collector never stores a snapshot twice and reads don’t block writes.
  3. ML experiments predict take-profit vs stop-loss outcomes from the snapshot features.
  4. Honest validation: walk-forward testing and market-regime analysis, so the model is always tested on data from after what it learned from.
  5. Result so far: minute-level snapshots hit a ceiling. Over time the predictions were close to random, with only a modest, unstable signal inside specific market regimes.
  6. Next: 5-second snapshots, to test whether faster data reveals what minute data can’t.

What it could do for you

Most data projects fail quietly: a model looks great on old data and falls apart in real use. I build the pipeline and the tests that tell you whether to trust it. Forecasting demand, detecting anomalies, scoring leads: before you bet on a model, you should know if the signal is real. Sometimes the most valuable answer is “not yet”, and it’s cheaper to learn that from a test than from a loss.

contact

Let's talk.

Need help with something? Looking for a professional who's reliable and believes in karma*?

Tell me about it here, or send me an email. I reply within two working days.

* which means value and integrity come first; reward follows.