Webull Hybrid Lab — Does Exhaustion + Divergence Additivity Survive the Options Pivot?NULL-HYPOTHESIS STUDY

Webull Hybrid Lab · Phase 1 (paper mode) · 2026-06-17 · ~9 min read · N = 0 closed option-trades at publication · pre-registered protocol


Abstract. The Webull Hybrid Lab is the options-rail successor to the Hybrid Lab — the integration of the Exhaustion and Divergence engines whose binary question was "does merging two weak-edge labs produce an additive edge?" It tests one pre-registered null hypothesis, stated in full here and never softened downstream: H0 — the exhaustion-plus-divergence additivity does not carry over from the binary venue to options; combining the two regime scores into one structure selection produces no calibrated, friction-surviving edge beyond what either lab achieves alone. Each component lab (Exhaustion, Divergence) independently expects to fail to reject its own null; this lab asks the sharper question of whether their combination clears break-even where the parts do not — the definition of additivity. Its /observe response echoes both exhaustion_score and divergence_score so the two-component HUD is always auditable. We reject H0 only on the §3 evidence; a failure to reject is the lab's expected and fully successful outcome. No real order is placed this phase; every leg is a dry_run=True broker echo.
Contents
  1. Lineage — the additivity question, on a new rail
  2. Method — classifier, selector, calibration, ergodic evaluation
  3. Null hypothesis & stop criterion
  4. What a failure to reject tells us
  5. Phase 2 — must-do / must-not
  6. References

1. Lineage — the additivity question, on a new rail

This lab descends from the Hybrid Lab (hybrid_engine.py), documented in the hybrid_v1 methodology whitepaper. That engine ran a six-stage pipeline — deterministic exhaustion prior, Kaufman-ER regime gate, a logit-space ensemble of five regime-gated features, a bounded vision-LLM fold, PAV isotonic calibration, and drawdown-braked ¼-Kelly sizing — to answer one question: whether merging two individually-weak signals (exhaustion regime, provider divergence) yields an edge larger than the sum of its parts. Its honesty commitments — an 88% logit clamp, a per-bucket Wilson lower bound, an openly-displayed identity calibrator below 25 samples — are inherited here wholesale.

The pivot's headline findings frame why additivity is the right question to re-ask on options. Finding #4 said both source regimes "contain regime information … [and] DO meaningfully shift implied-vol expectations" but produced no executable binary edge. Finding #5 adopted the project conclusion: stop monetising direction as a binary bet, and shape an options payoff around the predicted regime instead. The Hybrid Lab is where those two findings collide most interestingly — because the two component regimes make opposite volatility statements:

The exhaustion regime says "the move is overextended — expect vol contraction as it mean-reverts." The divergence regime says "the readers disagree — expect vol expansion as uncertainty resolves." When both fire at once, they are not redundant confirmations; they are a structured disagreement about the next regime.

That tension is exactly what makes the additivity question non-trivial. In the binary world the two signals could only ever vote on direction, so their vol disagreement was invisible — washed into a single bull-probability. On the options rail the disagreement becomes a selectable structure: the relative magnitude of exhaustion_score vs divergence_score tells the selector whether to lean short-vol (contraction wins), long-vol (expansion wins), or vol-neutral-directional (they cancel, leaving only the exhaustion fade direction). Findings #1 and #3 still bind: read real b off geometry, size off p_lower.

2. Method

2.1 The regime classifier — which forecast engine feeds it

The classifier is hybrid_engine.py, used unmodified as the forecast layer. It emits the §2.2 regime block with both exhaustion and divergence populated. Per the §2.3 contract, this lab's /observe response additionally echoes exhaustion_score and divergence_score as top-level fields so the page can render the two-component HUD — the user always sees the two raw inputs that drove the structure choice, never just the verdict. This is the hybrid_v1 "factors[] audit trail" discipline carried onto the options rail.

2.2 The strategy selector — the candidate set

The selector reads the relative magnitudes of the two scores and selects from the full defined-risk subset of the Periodic Table (WEBULL_PIVOT §3) — this is the only lab whose candidate set spans both the contraction and expansion families, because only the hybrid sees both signals:

Score relationshipVol readPrimary candidateAlternates
Exhaustion ≫ divergence, directionalcontraction winsbear_put_spread / bull_call_spread (debit)long_iron_butterfly
Divergence ≫ exhaustionexpansion winslong_stranglelong_straddle
Both high, comparable (structured disagreement)direction known, vol ambiguousbull_call_spread / bear_put_spread (vol-neutral defined-risk)long_call_calendar_spread
Both lowfloor — regime not armed

The "both high" row is the additivity test's sharp edge: it routes to a defined-risk vertical that monetises the exhaustion fade direction while staying roughly vol-neutral, on the thesis that the two signals agree a regime change is imminent even as they disagree on its vol sign. With defined_risk_only=true the selector never returns a naked short-vol leg even when the contraction read dominates — the Mertonian fat-tail discipline from the hybrid_v1 CHOP-BUSTER section binds here too.

2.3 The calibration pipeline — reuse, do not reinvent

Closed trades persist on the existing paper_trade_service.py row (no parallel DB) with the §2.4 columns. The additivity question makes per-structure bucketing indispensable: /api/cv/paper/calibration?strategy_id= lets us compare the hybrid's bull_call_spread bucket against the Edge and TRAIL labs' same-structure buckets, and the hybrid's long_strangle bucket against the Divergence lab's — the only way to tell whether combining the signals beat using either alone. Wilson 95% lower bound per bucket (Wilson 1927); exhaustion_score / divergence_score are persisted in legs_json context for post-hoc additivity analysis.

2.4 The ergodic Monte-Carlo evaluation (F7)

Break-even is p* = 1/(1+b) off each structure's geometry; sizing is fractional Kelly off p_lower (F7 README — never p_hat), with the hybrid_v1 drawdown brake and losing-streak damper carried forward. The additive ergodic Monte Carlo (Peters 2019) reports time-average growth. The hybrid lab is the most exposed to finding #2's trap: combining signals can lift ensemble EV (more "confirmations" feel like more edge) while the time-average growth path stalls or worse, because the combined selector trades more often and pays more friction. A positive EV with ≤ 0 growth rate is a fail.

3. Null hypothesis & stop criterion

H0 (pre-registered): exhaustion+divergence additivity does not carry over from binary to options — across every funded webull_strategy_id bucket, the hybrid's Wilson lower bound p_lower stays at or below the geometry break-even p*, and never exceeds the best single-component lab's p_lower for the same structure.

We reject H0 only if, over ≥ 100 closed paper option-trades in at least one webull_strategy_id bucket, the hybrid Wilson lower bound both exceeds break-even and exceeds the matching single-lab bucket (the additivity clause):

p_lowerhybrid(bucket) > p* = 1/(1+b)   AND   p_lowerhybrid(bucket) > max( p_loweredge/trail, p_lowerexhaustion, p_lowerdivergence ) (same structure)

The second condition is what makes this an additivity test rather than a fourth directional test. Beating break-even is necessary but not sufficient: if the hybrid clears p* only by routing to a structure that the Edge or Divergence lab already clears alone, then merging added nothing — the edge lived in one component, and H0 stands. Additivity is rejected only when the combination demonstrably outperforms its best part on a ≥ 100-trade bucket.

Why additivity is the hard bar. The binary hybrid lab's own whitepaper bounded the vision fold to weight 0.35 and clamped ensemble confidence at 88% precisely because stacking weak signals tends to stack their noise as fast as their signal. Two regimes that each fail to reject their own null can, when combined, easily produce a structure that trades more, pays more friction, and clears nothing. The "> best single component" clause guards against celebrating a coincidence.

4. What a failure to reject tells us

Per the lab's stance, a null-not-rejected result is a successful run. A sustained failure teaches:

The invalidating failure (vs. a valid null) is any dry_run=false broker response, any undefined-risk leg booked while defined_risk_only=true, or — hybrid-specific — an /observe response that omits exhaustion_score or divergence_score, which would mean the two-component HUD is rendering a verdict the operator cannot audit. All three are checked by the smoke-test sign-off.

5. Phase 2 — must-do / must-not

Must-do:

Must-not:

References

  1. Cox, J. C., Ross, S. A., & Rubinstein, M. (1979). Option Pricing: A Simplified Approach. Journal of Financial Economics, 7(3), 229–263. (Binomial valuation across the full defined-risk structure set the selector spans; alternative to Black & Scholes 1973.)
  2. Black, F., & Scholes, M. (1973). The Pricing of Options and Corporate Liabilities. Journal of Political Economy, 81(3), 637–654.
  3. Kelly, J. L. (1956). A New Interpretation of Information Rate. Bell System Technical Journal, 35(4), 917–926. (Payoff ratio b, break-even p* = 1/(1+b); drawdown-braked ¼-Kelly inherited from hybrid_v1.)
  4. Wilson, E. B. (1927). Probable Inference, the Law of Succession, and Statistical Inference. Journal of the American Statistical Association, 22(158), 209–212. (Per-bucket lower bound and the additivity comparison.)
  5. Peters, O. (2019). The Ergodicity Problem in Economics. Nature Physics, 15, 1216–1221. (Time-average vs ensemble growth — the F7 objective; especially sharp for the most-active lab.)
  6. Merton, R. C. (1969). Lifetime Portfolio Selection under Uncertainty: The Continuous-Time Case. Review of Economics and Statistics, 51(3), 247–257. (Defined-risk preference over short-vol fat tails, inherited from the hybrid_v1 CHOP-BUSTER analysis.)
  7. Taleb, N. N. (2007). The Black Swan: The Impact of the Highly Improbable. Random House. (Why stacked weak signals can stack noise; fragility of thin-sample additivity claims; optional.)
  8. Lineage: Binary Exhaustion Lab — hybrid_v1 Methodology Whitepaper; component labs: Webull Exhaustion and Webull Divergence; binary additivity framing: WEBULL_PIVOT_STUB_PROMPT.md §1.2 & §1.4, finding 4.