XGBoost, Python, Streamlit, GitHub Actions
The first version was a bootcamp project: notebooks, a model, a number that looked good. I wanted to make it real, something that predicted Bitcoin, ran on its own, and that I could point at and say it works.
Rebuilding it properly is what changed the goal. The interesting problem stopped being the model and became the evaluation: how do you know a number is real? What I ended up with is less a prediction system than an apparatus for catching myself being wrong. It caught me.
Can a gradient-boosted model predict the direction of Bitcoin's 7-day return better than a naive baseline?
A pipeline runs unattended every day at 22:00 UTC: it fetches market, on-chain and macro data, validates it, retrains, predicts, logs the prediction, and commits the result. No manual step. A four-page dashboard exposes the model, its history and its documentation, and a separate weekly job watches for drift.
The model predicts 7-day return rather than price, deliberately: tree models cannot extrapolate beyond the range they were trained on, so asking one for a price in a market at all-time highs guarantees failure.
The published result was 76.7% direction accuracy against a 52.6% baseline. Two independent errors produced it.
Overlapping windows sold as independent. The results file holds 2,467 daily rows, but consecutive 7-day windows share six of their seven days. Counting those as independent predictions inflates the sample roughly sevenfold. The true independent count, using the same windowing as every figure below, is 367.
Target leakage, which was the dominant effect. The 7-day target is computed from a close seven days ahead. The walk-forward loop trained on every row up to the prediction point, but the six rows immediately before it have targets built from closes that fall after it. The model was being trained on the answer.
Evaluated properly, on 367 non-overlapping windows from June 2019 to June 2026:
| Variant | Direction accuracy | R² | Correlation |
|---|---|---|---|
| As built, with the leak | 73.3% | 0.571 | 0.774 |
| Leak-free | 48.5% | −0.308 | −0.030 |
| Naive, always predict up | 52.6% | — | — |
Leak-free, the model scores below chance and below simply guessing the more common direction. A correlation of −0.030 between prediction and outcome is, for practical purposes, none.
A second claim went with it. The app advertised 90.6% accuracy on high-confidence predictions, where confidence is assigned by the size of the predicted move. Evaluated leak-free, the largest-move tercile scores 44.0%, below the smallest-move tercile at 48.8%. Predicted move size carries no information about whether the prediction is right. The thresholds in the code were themselves calibrated on the leaked evaluation.
A third claim, that XGBoost beat the deep-learning models, also fails. LSTM scored 50.0% and GRU 54.1% on the same 74 independent windows, but those were evaluated leak-free from the start, while XGBoost was not. It was never a like-for-like comparison. The neural nets were reporting honest numbers the whole time; XGBoost only appeared to beat them because it was reading the answer.
The strongest evidence is the part I did not control. Since March 2026 the pipeline has logged every prediction before its outcome existed, with no opportunity to re-specify anything afterwards. As of today: 125 logged, 120 resolved, right 48.3% of the time. The live log and the leak-free re-evaluation were produced by completely different routes and agree within a point. Both say the same thing.
One caveat in the other direction, stated deliberately: even the 48.5% is optimistic. The hyperparameters were tuned over 100 trials scored on the whole history with no held-out period, and the 52 features were selected from 271 candidates over the full eight years. The defensible claim is no demonstrated edge, not that Bitcoin is unpredictable.
The evaluation was the deliverable and I built it last. Everything that went wrong sits downstream of that: leakage a stricter split would have caught immediately, an overlapping window scheme that inflated both the sample and the score, and feature selection and tuning run over data the model was then scored on.
The monitor was worse than the model. A weekly job existed precisely to catch this, and it reported "All Clear" with a green check every week from April to August while all three of its checks were broken. Two crashed on a date-format inconsistency left by an earlier manual repair. The third measured file modification time, which in CI only ever reflects when the repository was checked out, so it reported the tuning parameters as zero days old forever; they were 130 days old.
The date format was not the root cause. Every exception handler appended to an informational list rather than the alerting one, so a check that could not run was reported as the thing it checked passing. A monitor that cannot fail is not a monitor. It needs a test that breaks it on purpose and confirms it screams.
Separately, deliberately attacking the pipeline, asking what each step assumes and how it could fail silently, found 13 real issues, including market index values frozen for a week and a 17-day gap in Bitcoin's own price history. Both were invisible in the output and would have quietly corrupted anything computed downstream.
The pipeline still runs daily and still logs every prediction, and the figures above are read from that log. I kept it running deliberately after the retraction. A live record that disagrees with a published claim is the most useful thing this project produced, and switching it off would have destroyed the evidence.
The one page that used the model to simulate trading returns is unlinked rather than deleted. It was built on the inflated accuracy, and any strategy result derived from it would be fiction.
Re-derived 2026-09-01. 76.7% on 2,467 rows comes from the committed results/XGB_7d_walkforward_results.csv. 48.5%/−0.308/−0.030 on 367 windows, and the 52.6% naive baseline, come from the committed results/XGB_7d_walkforward_leakfree.csv. 73.3%/0.571/0.774, the as-built figure evaluated on those same 367 windows, is not itself in a committed file, scripts/evaluate_leakfree.py computes it for comparison but only persists the leak-free column, so reproducing it means re-running that script rather than reading a results file directly. The live log is fetched from the public repository on page load, not hard-coded: 125 logged, 120 resolved, 48.3%. LSTM and GRU figures come from results/dl_walkforward_summary.csv, using the non-overlapping columns, the same 74-window basis the comparison above uses (GRU's overlapping-window accuracy is 52.5%; its independent figure, the one that means something here, is 54.1%). The leak diagnosis, the 44.0% tercile figure, the monitor post-mortem and the 13 pipeline issues come from the project's own audit records.