Python, Backtesting, MEXC Futures, Flask
This started as a bot that followed trade signals from a Telegram channel. The goal was simple: make money. That version is not public and never will be, because it is tied to a real trading account.
Pulling out the part that stood on its own, rules-based strategies driven only by price, no external signals, no discretion, turned into the whole project. Somewhere in the rebuild the goal quietly changed. Making a strategy look profitable is easy. Knowing whether it is turned out to be the hard part, and the more interesting one. What I built stopped being a bot that makes money and became a machine that can tell me honestly whether a bot makes money.
This is what happened when I pointed that machine at itself. The headline result is negative. The live part is still running.
Does a mean-reversion strategy on Bollinger Bands have a real edge on BTC perpetual futures, once it pays the costs it actually incurs?
Five rules-based strategies share one indicator library and one backtest engine. Only one, BB Channel Rider, has been through full validation and runs live.
It trades BTC/USDT perpetuals on 15-minute candles at 10x leverage, committing 25% of capital per position. It enters when price reaches a Bollinger Band (20-period, 3.0 standard deviations) in the direction an EMA-150 trend filter allows. It manages the position with a stop-limit plus a market backstop, a stop that snaps toward breakeven and then trails, and a take-profit at the opposite band. When the target is hit with price still at the band it reverses immediately rather than sitting idle. After a stop-loss, that direction waits two candles before re-entering. The specification was frozen on 10 May 2026 and has not changed since.
It runs as a service on a cloud VPS with automatic restart, alongside a dashboard for backtesting any of the five strategies and watching live status. The exchange key is read-only and IP-locked, and paper mode never sends an order, it only reads prices.
This number has been wrong publicly twice, and each correction came from auditing my own audit.
The original published result was 290 trades, 63.8% win rate, +$10,348. An audit script had drifted from the deployed strategy on five separate counts, every one flattering, measuring a configuration that had never actually run. Corrected: 311 trades, 60.5% win rate, +$730.45 over a year of real 15-minute data. It reproduces exactly. It is also meaningless, and the reason is worth more than the number.
The backtest tested each candle's wick against a Bollinger Band computed from that same candle's own close, then booked the fill at that band. Both facts are known once the candle has closed, so nothing looks wrong on inspection. But the wick that triggers the trade happens during the candle, before its close exists. The price it fills at could not have been calculated at the moment it fills. You cannot place that order.
The live bot does something different. It builds bands from closed candles only, tests the forming candle against the last closed candle's band, and manages the position from live tick data rather than the completed candle's cumulative wick, a level fixed before the candle opens can rest as a limit order and genuinely fill there, and a stop set mid-candle can't be hit by a price that printed before it existed.
Correcting the fill assumption alone, holding everything else fixed: +$730.45 → −$569.84.
Then correcting the stop-loss threshold, which the strategy's own frozen spec says should recalculate every candle, and which the backtest had instead frozen at trade entry: −$569.84 → −$683.90, on 437 trades, 53.8% win rate, 72.8% maximum drawdown.
Two independent bugs, found six days apart, both making the backtest friendlier than the live bot could ever be. Gate 1 already imports the live strategy module directly, the fix from the first correction, so this wasn't drift between an audit script and the deployed bot. It was the simulation logic itself.
The sharpest version of the problem: the published model cannot be paper-traded at all. It needs a band that does not exist until the candle is over. That impossibility is the proof it was never a strategy, it was an artifact of how it was measured.
The full-year figure above includes the exact window the parameters were selected from, which flatters any evaluation run over it, in the direction of a loss this time, for the same reason it flattered a gain before. The honest test is the period after the 10 May freeze, which the parameter search never saw:
116 trades, 57.8% win rate, +$5.62, 18.1% maximum drawdown. Essentially breakeven.
There's a real gross signal in that window, with fees switched off, the same trades return +$396.13 instead of +$5.62. At 10x leverage a round trip costs roughly 1.2% of margin, which is most of what a thin edge like this one has to give.
Zero-fee, full year: +$70.40, barely positive. With the fees the bot actually pays: -$683.90. The gross pattern over the full period is thin to begin with, and costs take the rest.
Two separate evaluation bugs on this project now, on top of the earlier audit-script-drift correction. All three are the same mistake in different clothes: a research script keeping its own private copy of the strategy, which drifts from the deployed bot, and drifts in the flattering direction, because a result that looks good gets checked less carefully than one that looks bad.
The fix is structural rather than a matter of being more careful. Whatever evaluates a strategy must import that strategy, never re-implement it. Gate 1 does this now. The parameter grid search (16,200 combinations) still doesn't, it carries its own copy of the strategy logic, has never been reconciled with either correction, and I'm deliberately not citing its numbers here until it is.
Two limits remain, stated because they still flatter the numbers. Paper mode models fees but assumes instant fills and no slippage, measured as negligible at this order size, but not zero. And the backtest cannot tell a stop-limit fill from the market backstop, so stop exits are costed at the cheaper rate.
The bot has traded since 4 August 2026 on a $1,000 paper balance: 31 trades, -$59.00 net, currently holding a position. It moved from -1.9% to -5.9% in the space of a few weeks, which is itself the point. A sample this small tells you almost nothing on its own; it stays in the same small-single-digit range as the out-of-sample backtest rather than the old +73% story or a clear loss, and that's the most it's honest to say about it right now.
What would actually settle this: running the two most realistic evaluation models side by side, live, on an identical tick stream, so any difference between them is the rule and nothing else. At roughly one trade a day that needs about three months to mean anything. That's the next build, not something this page pretends already happened.
Re-derived 2026-09-01. The evaluation-model fix (same-candle → last-closed fill basis) was found and committed 2026-08-16 but never deployed or propagated to any published surface until this pass. The snap-threshold fix was found independently the same day this page was rewritten, while re-verifying the first fix before publishing it, re-deriving before publishing is what caught it. Both are now live in the code the dashboard actually runs; figures above come from research/gate1_audit.py importing that code directly, against a freshly fetched year of real 15-minute candles. The live-bot figures come from the bot's own public status endpoint, fetched on page load, not hard-coded.
The parameter grid search (Gate 2) is not re-run here, it predates both fixes above, has its own unreconciled copy of the strategy, and was already excluded from the CV/LinkedIn copy before either correction. Left out rather than published stale a third time.