P-hacking rarely feels like cheating. It looks like dropping 2008 because it was unusual, switching from daily to weekly bars, adding a volatility filter, excluding one instrument that behaved oddly, and testing a one-sided rather than two-sided hypothesis. Each step is defensible alone; together they manufacture significance.
The defence is pre-registration, borrowed from clinical trials. Write down the universe, the sample period, the exact test, and the threshold before running anything. Then run it once. Any deviation is recorded as exploratory rather than confirmatory.
Example: a strategy returns p equal to 0.11 over 2004 to 2024. Trimming to 2010 to 2024 gives 0.04. Unless you had a reason to start in 2010 that was written before you saw the number, the honest reading is still 0.11.
Related: multiple-testing, data-snooping