Methodology

Caliper Trading is a site with a simple objective: allow users to understand their trading strategies via the scientific method. A Caliper is a machining tool that measures materials to a high degree of precision, dating back to the 6th century BC. Our methodology can best be understood in light of the scientific method, which we lay out here. It is enforced by our system throughout.

Observe and Formulate an Idea

First a nucleus of an idea to test must exist in the mind of the user, which is ultimately based on observations of the markets and of world as a whole. This seed of an idea can be brought to the chat bot and interrogated. Some questions of this idea must be asked to ensure that it is grounded in some measurement, anomaly, prior research, or first principles, i.e. an argument pointing to some mechanism justifying why this idea may be true. Ideas without a grounding mechanism is likely the result of some noise, basically tantamount to astrology.

Research

One of the key mechanisms to fleshing out an idea is the research process, which is done in response to inquiry. Looking beyond the headline and abstract results, combing through papers will produce secondary findings, tables, footnotes and the like that may help in producing a real edge. A known result is reproduced on overlapping conditions before their data count as evidence.

Every idea must also account for the counterparty, the other side of each trade made. A credible hypothesis names a counterparty, explains why they are making the trade, and why you are not facing adverse selection. Trading against better information and execution is a normal affair for markets, and ideas must have justification to go against these better operators. "Nobody else has noticed" is a delusion; every edge decays, and publication of the edge marks the top of the decay curve.

Hypothesize

From the process of generating and researching ideas, the system will arrive at a collection of hypotheses to be tested. Each hypothesis will be falsifiable, i.e. if an experiment fails to meet a pre-registered criterion, it will be rejected.

Experiment

The closest thing financial trading has to proper scientific experiments are backtests. Like any experiment, backtests require significant design considerations in their setup, measurement, assumptions, and significance.

Backtests face common setup failures that we account for in the design process.

  • No look-ahead. All mechanisms look at the past, data is served in the simulation's present, and prices are computed for that date.
  • Costs are a grid. Fees and spread are priced under several sets of assumptions on both legs (buying and selling). Optimistic bounds are noted.
  • Makers and takers operate differently with respect to a hypothesis and are thus tested separately.
  • The counterparty price is built of their information: we investigate what that information is.
  • Nothing is silently ignored. Exclusions come with a written justification, and filtered outliers are reported.

Ultimately, what is produced for a backtest is code, which operates in an isolated sandbox without network access, working on a fixed snapshot of historical data, fully staged before computation begins, during which measurement occurs. What is produced from a backtest is statistical in nature, so care is taken in that regard too.

  • Events are counted, not observations.
  • If permutations of hypotheses are tested, significance is checked against permutation nulls.
  • Concentration is always reported: how much of returns come from a few flaky trades?
  • A few winning trades (<20) warrants only a few units of trust.
  • The following metrics are always reported: annualized return, a measure of risk adjusted return, maximum drawdown, a t-statistic, numbers of total and winning trades, and compound versus constant-size returns.
  • Unproven is different from disproven; if a test produces a small sample size, the honest output is "unproven", alongside the sample size needed to decide it.

Finally the underlying assumptions are also listed. No analysis knows your personal capital, access to markets, fees or execution speed. Each conclusion therefore lists the dependent operational assumptions, and their relative sensitivity. An edge that requires institutional information and high-frequency computation is different in many ways from an edge available to retail.

Test Hypotheses and Conclusions

Once all is said and done, the results of the backtest are compared to the falsifiability conditions set out by the hypotheses, resulting in a rejection or accepted result. Adjusting falsifiability criteria post hoc is a new hypothesis, and should be tested on different data. This discipline separates science from curve fitting.

This is Not Financial Advice

A hypothesis is only tested against returns in the past, under stated assumptions. We can only measure the past. We cannot predict the future, nor do we advocate for any future action: No forecasts, No price targets, No recommendations.

Report

Finally, a verdict is issued regarding a hypothesis, stating what did and did not work and why. The result is presented with the above metrics, and gives an honest analysis of the idea.