Ayan Jain · Backtest Audit
Independent verification for quant strategies

Your backtest says 34% CAGR. Most of them are lying.

Six forensic checks for the failure modes that turn a paper edge into a live loss. One independent verdict, in writing.

Start a Check → Read the findings ↓

One person runs every check. One person signs the verdict. Named, and liable for it.

Audit verdictREPORT FORMAT
SurvivorshipPASS
Look-ahead / point-in-time
The signal uses a field only knowable after the close.
FAIL
Cost realism
The edge survives to 60bps; 30 was assumed.
FLAG
Overfitting / selectionFLAG
Statistical validityPASS
Fill feasibilityPASS

Illustrative summary, not a client report. Every check gets a verdict, a measured haircut where possible, and a fix. Read a full sample verdict (PDF).

3,852Stocks in my point-in-time Indian panel.
1,186No longer trade; most datasets silently drop them.
+4.0pp/yrMeasured bias cost on one realistic strategy.
48hFrom complete folder to written verdict.

Everything here is checkable: pip install backtest-bias · The notebook that reproduces +3.0pp on public data · Published audit findings · A complete sample verdict (PDF)

Receipts

Recent findings, with the verdicts they got

All real, all from the last few weeks: two from public-data audits, two from my own pipeline. I publish what I catch in my own work for the same reason a good lab publishes its failed experiments.

LOOK-AHEAD A t = 7.6 "edge" that was a data leak

A delivery-percentage signal showed t = 7.61, monotone across buckets. The field is only knowable after the close. Lagged one day: t = 0.66, ordering reversed.

t = 2, the usual significance bar same-day field: t = 7.61 7.61 Same-day field lagged one day: t = 0.66 0.66 Lagged one day
The t-statistic of the same signal, before and after removing the leak. Caught in my own pipeline by the checks this service runs.
COST REALISM A 28.8%/yr strategy that dies at its own cost

Illiquid-stock reversal in India: +28.8%/yr gross at t = 11.65, beautifully monotone. Then you pay to trade it.

0% gross: +28.8%/yr +28.8 Gross at 30bps round trip: +13.7%/yr +13.7 @30bps at 100bps round trip: minus 21.6%/yr −21.6 @100bps at 200bps round trip: minus 72.0%/yr −72.0 @200bps
Net %/yr at realistic round-trip cost for that liquidity bucket. The gross number was real. It just wasn't reachable.
SURVIVORSHIP The dead companies your data source deleted

Popular free sources drop delisted names. Rebuilt with every dead company kept in the universe, one realistic strategy lost +4.0pp/yr of its claimed return. A public 99-name version reproduces +3.0pp, runnable on Kaggle, so you can check the method before you trust it.

DATA DEFECT Fields that are zero, not missing

Two fields in a widely-used panel are exactly 0.0 before 2022 rather than NaN, so notna() passes and any ranking built on them silently sorts "has the field" from "doesn't". Produced a fake t = 7.6 result before it was caught.

Scope

The six checks

Every audit runs all six. Each gets a pass, flag or fail, the estimated performance haircut where measurable, and a fix.

01Survivorship

Is your universe quietly missing the companies that died?

02Look-ahead / point-in-time

Does any signal use information that didn't exist on the trade date? Adjusted prices, restated fundamentals, index membership as of today.

03Cost realism

Are commissions, slippage, impact and taxes at levels a real fill would actually pay?

04Overfitting / selection

How many configurations did you try, over how many years? Deflated metrics and a parameter-sensitivity map.

05Statistical validity

Is the edge distinguishable from luck at your sample size? Streaks and drawdowns checked against the distribution.

06Fill feasibility

Could the market have absorbed your orders at those prices? Volume participation, limit fills, gap handling.

Process

How it works

1 · Send the folder.

Trade list and equity curve, a platform export, or written rules precise enough to act on. Code welcome, never required.

2 · The 48h clock runs.

Starts when payment and a complete folder are both in. Inside it, I reconstruct your test, rerun it independently, and cross-check it against my own panel. If something is missing I ask once and the clock pauses.

3 · Written verdict.

Pass, flag or fail on each check, the measured haircut where possible, and a prioritised fix list.

Not a grade for your ego. An engineering document for your next iteration. If a check fails in a way that needs deeper work than the fixed scope covers, I tell you what it would take before doing anything further. No surprise invoices.

The record

Work that was checked in public

None of these people are clients. They are researchers and founders who engaged with the method in the open, corrected it where it was wrong, and adopted it where it held. That is a more useful thing than a testimonial, and a harder one to manufacture.

Variance-ratio testing

Dr. Heather Dempsey, quantitative researcher, amended a published piece after an exchange about a discrepancy, and noted the change in the body of the post rather than editing quietly. Correcting a published piece in public is rarer than the catch that prompts it.

Selection bias, adopted

A simulation showing that ranking on the same window that selected a candidate can manufacture apparent mean reversion was published as a follow-up piece with named credit, and the disjoint-window recompute became stated practice. The useful test is whether a method travels beyond the person who proposed it.

Effective sample size

On multiple testing under correlated search, we each corrected the other in the same thread: an equicorrelation figure that did not survive checking, and a threshold approximation of mine that ran high at practical sample sizes. Both corrections sit in the public exchange.

Platform founders

Two trading-platform founders have taken methodology suggestions from public threads into their own products, and one invited me in as an independent voice on data honesty for their users. The checks are not specific to my own book.

Everything above is public and linkable, and quoted with permission where a person is named.

Price

Start with a Check. Everything else is scoped from it.

Fix

$1,500

A Check found real problems and you want them repaired: the leak closed, the cost model rebuilt, the test rerun clean.

Scoped in writing after a Check. 7-day delivery window.

Deep

$6,000

Full pipeline engagement: your data rebuilt point-in-time correct, the audit run at source level, the report your next strategies inherit.

Scoped in writing. 7-day delivery window.

48h turnaround Fixed price No code required NDA on request

Questions

Questions people actually ask

Can't I just ask an AI to do this for free?

I use AI heavily myself; it's why the price is a few hundred dollars and not a consultant's week. But three things don't come with the free prompt. The reference data: AI cannot conjure a survivorship-correct, point-in-time panel to check yours against, and mine took months to build. The adversarial stance: a model in the strategy owner's hands tends to validate the owner. And liability: a named human signs this verdict. In one recent week my own tooling produced four significant-looking results, up to t = 7.6, and every one was a bug. The catching is the product.

Do I have to share my code or my alpha?

No. A trade list and an equity curve are enough for all six checks. The audit tests whether your test is trustworthy; it doesn't need to know why you trade what you trade. Code only enters at the Fix and Deep tiers, under an engagement letter with custody and deletion terms. NDA on request.

Which markets do you cover?

The six checks are method checks: they test how your backtest was built, so they work on equities, futures and options from any market. The data-level cross-check against my own reference panel is a different matter: it is deepest for Indian equities, partial for US indices, and elsewhere I verify your data's internal consistency rather than compare it against an independent panel. I say which kind of check produced each finding.

Will you tell me if my strategy is good?

I'll tell you whether your test can be trusted. That's the part you cannot see from inside, and the part everything else depends on. The market grades the strategy.

How does the guarantee work?

Check tier only. Your intake asks what you already believe about each check: which fields might leak, what your costs really are, how many configurations you tried. That document is the baseline. If the verdict doesn't go beyond it, tell me and the fee comes back. Beyond means a named mechanism, a measured magnitude, or a fix that isn't in your answers. You hold both documents, so the comparison is checkable from your side, not just mine. No guarantee on Fix and Deep, which are scoped in writing before any money moves.

What if you find nothing wrong?

Then you get a verdict saying so, per check, and that document is worth something: an independent audit you can show a funder, a prop firm, or yourself at 2am before going live. A clean pass from a hostile reviewer is not the same as your own backtest saying it's fine.

Custody

Confidentiality

Whatever you send stays on my machine, is never reused or redistributed, and is deleted on request. If you'd rather not send code at all, don't. A trade list and an equity curve are enough for all six checks.

Fit

Who this is for

Quants about to size real money on a backtest. Prop-challenge takers, where an audit costs less than a second attempt and tells you which one to make. Developers selling strategies who want independent verification to point at. Algo providers preparing for SEBI's PaRRVA regime, where a research report per algo is a standing requirement. And anyone whose equity curve looks a little too clean to be true.