TokSeq documentation/Findings

Findings

What the platform confirmed, what it rejected and with which numbers, including the central negative result that direction is not predictable.

This page is the point of the platform. It lists what survived and what did not, with the measurement that decided each case. The rejections are longer than the confirmations, which is the expected shape of an honest research record.

The central negative result: direction

Short-horizon directional prediction was tested six independent ways and failed every time.

MeasurementResult
Out-of-fold evaluation over 730 daysAUC 0.52
First-stage systematic search0 edges over 15 003 experiments, best deflated Sharpe 0.329 against a 0.90 threshold, best score 36.42 against a threshold of 40
Paper desk journalDirectional hit rate 0.51
Second-generation canary sample0 significant results out of 24, verdict "no edge" on all
Conditional direction, restricted to declared states445 hypotheses, 0 survivors
External corroborationNine months of a documented public record of short-horizon directional trading, and published work on a large prediction-market dataset where one per cent of traders capture most of the profit and seven in ten lose

The decisive number is not the AUC, it is the comparison against the search. The best result of the directional family sat at 3.52 standard deviations of the Sharpe estimator, while the expected maximum of pure noise across fifteen thousand trials is about 3.96. The family's best finding was weaker than noise should have produced.

The structural diagnosis is straightforward. Tens of thousands of hypotheses, three assets whose daily returns correlate at 0.80 to 0.87, a forward window measured in weeks and horizons of five minutes to an hour, produce an effective sample of a few dozen daily blocks of a single underlying factor against a multiplicity of tens of thousands. Any effect in the range these searches surface, a few thousandths to a few hundredths of AUC, drowns in the correction.

The directional family is therefore frozen in code. Registering a new directional hypothesis raises an exception, and lifting the freeze requires an explicit flag intended for control runs.

Confirmed

FindingMeasurement
Volatility is predictableBenchmark model reaches AUC 0.73 against 0.69 for the trivial predictor, with an R-squared of 0.63 to 0.66 on log volatility
Regime change is detectableChange-AUC 0.695 in the registered live model, 0.7285 on a half-hour target in the research prototype
Heavy models earn their cost for regime, and only for regime87 wins out of 95 paired comparisons, a gap of +0.0717 in balanced accuracy, zero negative comparisons
Volatility risk premium exists and is conditionalImplied exceeds realized on 64 per cent of days for the first asset and 78 per cent for the second, with median premia of +4.0 and +12.9 volatility points
The premium is strongly state-dependentIn the richest implied-volatility tercile the premium rises to +8.2 and +16.2 points, on 78 and 92 per cent of days. In the cheapest tercile it disappears entirely, at −0.8 points on 47 per cent of days
Microstructure is an early signal for regime change, not for directionPooled incremental AUC +0.031, interval [+0.0159; +0.0480], on 195 995 observations
A cascade precursor reproducesCompression of taker-flow dispersion before liquidation cascades, one-sided p-values 0.0070, below 0.0001 and 0.0005 across three independently computed assets
Two anomaly signals qualify as a risk gateTaker-dispersion compression reaches AUC 0.671, interval [0.634; 0.709], with an increment of +0.099 over the core baseline at five minutes of lead time; a volume-burst signal reaches 0.640 with +0.082
A clock-boundary microstructure regularityExcess activity in the last seconds before quarter-hour boundaries, ratio 1.144, 1.117 and 1.126 across three assets against a control ratio of 1.00, driven by trade count rather than volume

The live record of the volatility forecasters, measured on the full journal rather than a recent window, is a hit rate of 0.692 to 0.708 for volatility expansion against a base rate of 0.48, and 0.643 to 0.654 for regime change against a base rate of 0.52.

Rejected, with the number that rejected it

ApproachVerdict
Averaging, voting and regime routing as ensemble methodsDead. Only the trained stacking combiner works
Committees beating their best member22.5 per cent of the time, median difference −0.0294
A regime router as a portfolio brainCeiling of the entire idea, oracle minus best constant, is +16.92 basis points with an interval crossing zero and p 0.087. The actual router delivered −2.12 basis points
The hybrid portfolio assemblyThe measured arm delivered +15.55 basis points, and a shuffled-label placebo delivered +13.22 with a 95th percentile of +34.48. Gate not passed
Meta-labelling of entriesNegative twice. The second attempt reached +2.02 basis points against a placebo 90th percentile of +2.31
Smart execution−0.078 basis points against a placebo of −0.079. Short-horizon price drift is unpredictable, with a negative out-of-sample R-squared
Delta-neutral funding carry−2.93 per cent annualized above the risk-free rate, interval [−3.61; −2.06]. The entire 45-cell parameter grid fits inside 0.78 percentage points against a shortfall of 2.5 to 3.3 points
Cross-venue basis arbitrageBasis flat at +4.5 to +6.2 basis points, with only one tradable leg
Crowding features for regime changeSignificantly harmful on the first asset, incremental AUC interval [−0.0426; −0.0192]
Microstructure in the production feature setNot adopted. The volatility increment interval crosses zero, and the regime increment of +0.0048 falls below the adoption threshold once any family correction is applied
Improved volatility models over the heterogeneous autoregressive benchmarkSix variants tried. None beats it at the daily horizon
Alternative regime definitions by change-point detection and cumulative-sum methods248 configurations, no change in screen behaviour
Rule-based strategies over declared states336 configurations, zero rejections of the null
News as a featureGate failed. No news feature reaches an AUC of 0.55 against volatility expansion, and every increment over the volatility baseline is negative. The baseline itself reaches 0.74 to 0.78
Large language models as a second opinionNo edge. Against a frequency baseline the best cloud model gains +0.059 in precision-at-one with an interval crossing zero, p 0.199. The live local run produced 10 914 forecasts and zero abstentions, and matched a constant
Spread as a level featureRejected by its own measurement. The daily 95th percentile equals the median to three decimals
A cascade detector as a threshold switchPooled onset AUC 0.4943 on 13 322 observations. Confirmed as a feature, refused as a switch
The one apparent edge of the first research generationWithdrawn. Forward evaluation gave median Sharpe ratios of −7.92, −7.94 and −0.98 across three assets

That last row is the one worth reading twice. A stacked committee cleared every gate of the first research generation and was the platform's only validated directional edge. Its forward window destroyed it, an erratum was issued, and the finding was retracted rather than defended.

Selection overfitting, measured

Two candidate groups that passed the significance gate were then tested for backtest overfitting. One reached an overfitting probability of 0.599 with 23 per cent positive forward periods; the other 0.393 with zero per cent. Both were refused. This is the reason the combinatorial cross-validation stage exists at all: passing the significance gate is necessary and nowhere near sufficient.

Forward record

Over five and a half weeks of paper forward testing with 152 agents, turnover reached 323 million dollars, gross profit was positive at about 25 000 dollars, and the net result after costs was a loss of about 369 000 dollars – minus 11.4 basis points on turnover, with fees consuming a quarter of deployed capital. Twenty-three agents out of 152 finished positive, and two had a confidence interval whose lower bound stayed above zero. Over that same window the three underlying assets rose 22.7, 29.5 and 41.4 per cent.

That is the honest shape of the problem. Cost is the dominant term, the beta of the window swamps the signal, and a naive reading of gross performance would have been off by an order of magnitude and the wrong sign.