Six forensic checks for the failure modes that turn a paper edge into a live loss. One independent verdict, in writing.
Start a Check → Read the findings ↓One person runs every check. One person signs the verdict. Named, and liable for it.
Illustrative summary, not a client report. Every check gets a verdict, a measured haircut where possible, and a fix. Read a full sample verdict (PDF).
Everything here is checkable:
pip install backtest-bias ·
The notebook that reproduces +3.0pp on public data ·
Published audit findings ·
A complete sample verdict (PDF)
All real, all from the last few weeks: two from public-data audits, two from my own pipeline. I publish what I catch in my own work for the same reason a good lab publishes its failed experiments.
A delivery-percentage signal showed t = 7.61, monotone across buckets. The field is only knowable after the close. Lagged one day: t = 0.66, ordering reversed.
Illiquid-stock reversal in India: +28.8%/yr gross at t = 11.65, beautifully monotone. Then you pay to trade it.
Popular free sources drop delisted names. Rebuilt with every dead company kept in the universe, one realistic strategy lost +4.0pp/yr of its claimed return. A public 99-name version reproduces +3.0pp, runnable on Kaggle, so you can check the method before you trust it.
Two fields in a widely-used panel are exactly 0.0 before 2022 rather than NaN, so notna() passes and any ranking built on them silently sorts "has the field" from "doesn't". Produced a fake t = 7.6 result before it was caught.
Every audit runs all six. Each gets a pass, flag or fail, the estimated performance haircut where measurable, and a fix.
Is your universe quietly missing the companies that died?
Does any signal use information that didn't exist on the trade date? Adjusted prices, restated fundamentals, index membership as of today.
Are commissions, slippage, impact and taxes at levels a real fill would actually pay?
How many configurations did you try, over how many years? Deflated metrics and a parameter-sensitivity map.
Is the edge distinguishable from luck at your sample size? Streaks and drawdowns checked against the distribution.
Could the market have absorbed your orders at those prices? Volume participation, limit fills, gap handling.
Trade list and equity curve, a platform export, or written rules precise enough to act on. Code welcome, never required.
Starts when payment and a complete folder are both in. Inside it, I reconstruct your test, rerun it independently, and cross-check it against my own panel. If something is missing I ask once and the clock pauses.
Pass, flag or fail on each check, the measured haircut where possible, and a prioritised fix list.
Not a grade for your ego. An engineering document for your next iteration. If a check fails in a way that needs deeper work than the fixed scope covers, I tell you what it would take before doing anything further. No surprise invoices.
None of these people are clients. They are researchers and founders who engaged with the method in the open, corrected it where it was wrong, and adopted it where it held. That is a more useful thing than a testimonial, and a harder one to manufacture.
Dr. Heather Dempsey, quantitative researcher, amended a published piece after an exchange about a discrepancy, and noted the change in the body of the post rather than editing quietly. Correcting a published piece in public is rarer than the catch that prompts it.
A simulation showing that ranking on the same window that selected a candidate can manufacture apparent mean reversion was published as a follow-up piece with named credit, and the disjoint-window recompute became stated practice. The useful test is whether a method travels beyond the person who proposed it.
On multiple testing under correlated search, we each corrected the other in the same thread: an equicorrelation figure that did not survive checking, and a threshold approximation of mine that ran high at practical sample sizes. Both corrections sit in the public exchange.
Two trading-platform founders have taken methodology suggestions from public threads into their own products, and one invited me in as an independent voice on data honesty for their users. The checks are not specific to my own book.
Everything above is public and linkable, and quoted with permission where a person is named.
The full six-check audit. Written verdict in 48 hours. For most backtests this is the only step needed. See a full sample verdict (PDF).
Guaranteed: the intake asks what you already believe is wrong, in numbers. If the verdict doesn't go beyond what you wrote, the fee comes back.
India: pay by Razorpay, instantly. International: pay $250 by card, invoice on request.
Start a Check →A Check found real problems and you want them repaired: the leak closed, the cost model rebuilt, the test rerun clean.
Scoped in writing after a Check. 7-day delivery window.
Full pipeline engagement: your data rebuilt point-in-time correct, the audit run at source level, the report your next strategies inherit.
Scoped in writing. 7-day delivery window.
48h turnaround Fixed price No code required NDA on request
I use AI heavily myself; it's why the price is a few hundred dollars and not a consultant's week. But three things don't come with the free prompt. The reference data: AI cannot conjure a survivorship-correct, point-in-time panel to check yours against, and mine took months to build. The adversarial stance: a model in the strategy owner's hands tends to validate the owner. And liability: a named human signs this verdict. In one recent week my own tooling produced four significant-looking results, up to t = 7.6, and every one was a bug. The catching is the product.
No. A trade list and an equity curve are enough for all six checks. The audit tests whether your test is trustworthy; it doesn't need to know why you trade what you trade. Code only enters at the Fix and Deep tiers, under an engagement letter with custody and deletion terms. NDA on request.
The six checks are method checks: they test how your backtest was built, so they work on equities, futures and options from any market. The data-level cross-check against my own reference panel is a different matter: it is deepest for Indian equities, partial for US indices, and elsewhere I verify your data's internal consistency rather than compare it against an independent panel. I say which kind of check produced each finding.
I'll tell you whether your test can be trusted. That's the part you cannot see from inside, and the part everything else depends on. The market grades the strategy.
Check tier only. Your intake asks what you already believe about each check: which fields might leak, what your costs really are, how many configurations you tried. That document is the baseline. If the verdict doesn't go beyond it, tell me and the fee comes back. Beyond means a named mechanism, a measured magnitude, or a fix that isn't in your answers. You hold both documents, so the comparison is checkable from your side, not just mine. No guarantee on Fix and Deep, which are scoped in writing before any money moves.
Then you get a verdict saying so, per check, and that document is worth something: an independent audit you can show a funder, a prop firm, or yourself at 2am before going live. A clean pass from a hostile reviewer is not the same as your own backtest saying it's fine.
Whatever you send stays on my machine, is never reused or redistributed, and is deleted on request. If you'd rather not send code at all, don't. A trade list and an equity curve are enough for all six checks.
Quants about to size real money on a backtest. Prop-challenge takers, where an audit costs less than a second attempt and tells you which one to make. Developers selling strategies who want independent verification to point at. Algo providers preparing for SEBI's PaRRVA regime, where a research report per algo is a standing requirement. And anyone whose equity curve looks a little too clean to be true.