Skip to content
GetProfitable
Search
Dictionary

Statistical power

The chance a test detects an effect that is genuinely there. Low power means your research mostly produces silence and flukes.

Power rises with sample size, with the size of the true edge, and with lower volatility of the observations. It falls as you tighten the significance-level. A power of 0.8 is the usual target, meaning you would find a real effect four times in five.

The underrated consequence of low power: when an underpowered study does return a significant result, that result is usually an overstatement, because only the luckiest samples clear the bar. This is why the first live year so often prints half the backtest's return even when the edge is real.

Example: to detect a Sharpe of 0.5 with 80% power at the 5% level you need roughly 32 years of annual observations, or about 8 years if you measure monthly and the returns are well behaved. Three years of backtest is not a test, it is a hint.

Related: type-ii-error, sample-size, minimum-backtest-length, sharpe-ratio

See it drawn

Original diagrams for the ideas on this page. Illustrative, not real market data.

The spread of outcomes behind an expectancyA histogram of forty trades: a tall block of small losses on the left, a low spread of larger wins on the right, and a line marking the average outcome.NUMBER OF TRADES051024 LOSSES, AVG −$20016 WINS, AVG +$600EXPECTANCY +$120−$400−$200$0+$200+$400+$600+$800PROFIT OR LOSS PER TRADEexpectancy = (40% × $600) − (60% × $200) = +$120 per trade
Expectancy: the average trade. Forty trades sorted by outcome: 24 small losses and 16 larger wins. Weighting each side by how often it happens gives the average result per trade, marked here by the dashed line at +$120.

Educational only, not advice. Spotted an error? Post in Site Feedback.