Koraki · korakiwatch.com · status report

Six weeks in, no measurable edge

Koraki is a paper experiment testing whether mechanical trend rules add to or subtract from a stock-picking signal. Six books run over one identical set of analyst Top Picks and one identical price history; the only variable is the exit rule. This is a status report, not a result.

Closes through 5 Sep 2026 · 12 picks · 26 trading days

Where it stands

−0.76
mean relative
points vs SPY
6 / 12
picks ahead
of the index
5.49
standard
deviation
0.48
standard errors
from zero

The picks have returned −0.76 points against SPY on average. Six are ahead, six behind, and the median is −0.75 — so the average is not being pulled by any single name.

The number that matters is the fourth one. Dispersion is 5.49 points, roughly seven times the size of the average, which puts the mean 0.48 standard errors from zero. Anything under about two is indistinguishable from noise. This is a quarter of that.

The picks are not measurably different from the index in either direction. That is a null, not a negative — the sample is too small and too noisy to support a claim of underperformance any more than it supports one of outperformance.

The distribution

−8 −4 0 +4 +8 SW FTAI FWONK EQIX CL WMG PANW C IFF BKR NVDA META −7.47 −7.22 −6.91 −4.83 −2.89 −2.30 +0.79 +0.79 +1.99 +3.46 +5.74 +9.75 mean −0.76 index
Relative strength in points versus SPY, each pick measured from its own entry date. A 17-point spread around a mean of −0.76.

The shape is the finding. Nothing clusters near the average — the picks are spread across a 17-point band, and the two ends are populated by real moves in both directions. Six weeks of data has produced a wide, roughly balanced distribution centred on nothing in particular.

No single name is driving it

A month ago one pick controlled the entire result: removing it moved the mean from −0.79 to −0.03. That is no longer true, which is the clearest sign the sample is starting to do its job.

SampleMean relative
All 12 picks−0.76
Drop the worst (SW)−0.15
Drop the best (META)−1.71
Drop both extremes−1.14
Drop two from each end−1.24

Every version lands between −1.7 and −0.2. The conclusion does not depend on which picks you include, which is a meaningfully stronger position than a month ago even though the headline number barely moved.

What would change the picture

The pre-registered read is at 40 positions with confidence intervals clustered by entry date, because picks published on the same day are not independent observations. At 12 picks and 9 cohorts, the interval would be wide enough to contain almost any plausible truth.

The design commits to publishing a null. The threshold was written before any code existed: if the clustered 90% interval at 40 positions sits entirely within ±$150 per pick per $10k and contains zero, that is a real finding — the rules do not move the needle — and it gets published as one.

A test nobody will abandon is not a test. The other three stop conditions are thirty days without opening the page, fewer than fifteen completed positions by month twelve, and a corrupted log.

Two positions have completed so far. Both were stopped out, and in both cases the rule books lost less than simply holding would have — which is what a stop is for, and which is a separate question from whether the picks themselves are any good.