Methodology

How the research is judged

Most trading results that look good are the best of many tries. These are the rules that stop us fooling ourselves, written down before any result is published.

Every idea is registered before it is tested

Each trial is written down before it sees the data it will be judged on: the hypothesis in words, the metric, the direction, the threshold, and how many 15-minute windows must pass before the answer counts. The record is frozen. Changing your mind is allowed, but the revised trial is registered as a new one and both count.

Until the registered number of windows has passed, the answer is “undecided”, however good the running figure looks.

Trying more ideas raises the bar

Test enough ideas and one will look good by chance. So the bar every result must clear rises with the number of trials ever registered, not just the ones that worked. The correction is Bonferroni-style: the chance of a false positive allowed across the whole family of trials is split between them, so each extra idea makes every future verdict harder to pass.

This is deliberately strict. Most of our trials are variations on a few ideas and are correlated, and Bonferroni assumes the worst case. We would rather miss a small real edge than publish a lucky one.

The sample size is windows, not trades

Thousands of decisions inside one 15-minute window share one outcome. They are one observation, not thousands. Counting them separately makes any result look far more certain than it is.

Standard errors are clustered

Contracts that close at the same moment move together, and so do days. Several crypto series close at the same time on the same underlying market. So the uncertainty on any result is computed with standard errors clustered by closing time and by day, rather than treating each contract as independent.

Fees come out per fill, at the real rate

Every result is net of the exact taker fee on every fill, rounded the way the exchange rounds it. Paper results are scored with a fill model, and results from real fills are labelled as such. The two are never mixed in one number.

Retiring a strategy runs the other way

A strategy that passed its bar can stop working. It is retired when even an optimistic reading of its recent results falls short, not after one bad fortnight.

Why there are no numbers on this page

Publishing a performance claim needs legal review first. When results do appear on this site, each will show its trial count, whether it came from paper or real fills, its fees, the number of windows and dates, and its clustered error. A number without those is not worth reading.