The specific rules GC8 trades on are proprietary. The process that produced them is not, and it is the part an allocator should judge us on.
A backtest is an argument, not evidence. It is trivially easy to produce a beautiful historical curve by searching hard enough through the same data, and nothing about the curve itself tells you whether you found an effect or fitted the noise.
So the research process is built around disconfirmation. Each stage exists to kill a candidate for a specific reason, and a strategy earns production by failing to die rather than by looking impressive. Most of what we test does not make it, and that is the process working.
The stages below are sequential and none is optional. A candidate that fails at any point stops there — which is the point of having the sequence at all. Our two strategies are marked at the stage each has reached: GC8 is live and trading, R34 is still in research.
An idea is written down as a testable rule before any data is run against it. If it cannot be specified, it cannot be tested.
The rule is built out and measured across years of market history, with the conditions under which it must not trade defined alongside the conditions under which it does.
Full historical simulation, costed. We look specifically for the ways it could be wrong: survivorship, parameter sensitivity, dependence on one unusually good stretch.
Re-tested on periods the research never touched. A result that only holds on the data used to build it is where most candidates end.
Behavior through the worst available conditions rather than the average ones. Average conditions are not the test.
Run against live markets recording real prices and real fills, with no capital at risk, until the record is long enough to check the model against.
R34Approved and trading. Orders placed, monitored and closed by software, with no manual override in normal operation.
GC8Live results reconciled against what the specification says should have happened. Divergence is investigated as a defect, whether it cost money or made it.
Daily price history across the tradable universe, stored and versioned so a result can be reproduced later against the same inputs. Bars are re-read rather than assumed final — a price captured during a session is not the price that session closed at.
Rules simple enough to be written down completely. Complexity is a cost: every additional parameter is another thing that can be fitted to noise and another thing that can fail silently in production.
Out-of-sample testing on periods the research never touched, and sensitivity analysis around every parameter. A result that depends on one exact setting is not a result.
The default assumption is that a promising backtest is overfitted until it survives data it has never seen. We look for the specific ways a result could be an artifact rather than for reasons to believe it.
Modeled explicitly, including the spread crossed on entry and exit and the cost of borrowing to short. A strategy whose edge disappears under realistic costs never had one.
How much capital a strategy can absorb before its own trading moves the prices it depends on. Capacity is measured as part of research, not discovered after launch.
Orders are placed, monitored and closed by software connected directly to the broker. Live fills are compared against the prices the research assumed, because the gap between them is where returns quietly disappear.
Live results are reconciled against what the specification says should have happened. Divergence is investigated as a defect — whether it cost money or made it.
Prospective investors receive a fuller methodology discussion under confidentiality, including validation results and the assumptions behind cost and capacity modeling.