# Further Analysis: Hidden Markov Models and Market Structure **David Chan, Claude Sonnet 4.6** *AI-Symbiosis Research · April 2026* --- ## Overview Following Lab 4 (Baum HMM Regime Detector on BTC) and the KEGA gap analysis of Cappé-Moulines-Rydén, three results emerged from the intersection of theory and empirical observation. The first is proven. The second is a strong hypothesis. The third is an open conjecture with a testable experiment. --- ## Key Result 1 — The Candle Manifold (Proven) The observation vector $(fracChange, fracHigh, fracLow)$ does not live in flat $\mathbb{R}^3$. It lives on a **constrained submanifold** defined by: $$fracHigh \geq 0, \quad fracLow \geq 0, \quad fracHigh \geq fracChange, \quad fracLow \geq -fracChange$$ These are hard geometric constraints — not all of $\mathbb{R}^3$ is reachable by a valid candle. The emission space is a bounded, curved surface. Standard Gaussian HMM fits ellipsoids in flat $\mathbb{R}^3$, ignoring this curvature. The natural gradient upgrade (Baum-Sell manifold paper, Amari 1998) respects the geometry by replacing the flat M-step with: $$\theta_{t+1} = \theta_t + \eta \cdot G(\theta_t)^{-1} \nabla_\theta Q(\theta_t;\, \theta_t)$$ where $G(\theta)$ is the Fisher information matrix. This is Gap 2 from the KEGA analysis — the missing bridge between Baum 1970 and information geometry. The parameter space of any probabilistic automaton is already a manifold by construction (product of simplices). The flat EM ignores this. The natural gradient does not. **Implication for Lab 4:** The regime boundaries learned by standard Baum-Welch are hyperplanes in flat space. The true regime boundaries are geodesics on the candle manifold. The natural gradient version would find sharper, more stable regime separations. --- ## Key Result 2 — The Automaton-Manifold Correspondence (Strong Hypothesis) Observing that markets were modeled as a Game Theory Automaton (Fortuna project) and independently as an HMM (Lab 4), and that both arrived at the same architecture, raises a general question: > *Does every automaton system have an underlying manifold structure?* **What is proven:** - The parameter space of any probabilistic automaton is a Riemannian manifold under the Fisher information metric. This follows directly from Amari (1985) and is not in dispute. - The HMM is the canonical smooth relaxation of a deterministic finite automaton. Sharp state transitions become probabilistic; the automaton's skeleton becomes a manifold's flesh. **The hypothesis:** > Every system with hidden discrete states driving observable outputs has a natural manifold relaxation. The automaton is the discretization. The manifold is the continuous ground truth. | System | Automaton | Manifold Relaxation | |---|---|---| | Market regimes | Game Theory Automaton | Gaussian HMM on candle manifold | | Language | Formal grammar | Transformer attention manifold | | Brain states | Neural firing patterns | State-space model on neural manifold | | Protein folding | Conformational states | Energy landscape manifold | The direction **automaton → manifold** (smooth it) is always well-defined. The reverse direction **manifold → automaton** (discretize it) requires choosing a cell decomposition — and that choice is not canonical. This is precisely the model selection problem (Gap 5, KEGA): how many states? The answer depends on the resolution at which you discretize the manifold. **Status:** Proven for Markov chains and HMMs. Open for pushdown automata and beyond. Strong enough to publish as a conjecture with the proven cases as supporting evidence. --- ## Key Result 3 — Recursive Sub-Regimes (Conjecture + Testable Experiment) Lab 4 identifies three macro-regimes: Bull, Bear, Choppy. But within each macro-regime, sub-structure likely exists — a Bull regime contains Strong Bull, Weak Bull, and Topping Bull days. A Bear regime contains Panic Bear, Slow Bear, and Bottoming Bear days. **The conjecture:** > Market regimes have recursive structure. Each macro-regime contains distinct micro-regimes with their own emission statistics. The correct model is a Hierarchical HMM (HHMM), not a flat one. This corresponds to a multi-scale decomposition of the candle shape manifold — macro curvature separating Bull from Bear, micro curvature separating sub-states within each. **The experiment (Lab 4b):** ``` Step 1 — Use Lab 4 labels: Bull / Bear / Choppy per bar Step 2 — Extract each subset: df_bull, df_bear, df_choppy Step 3 — Fit child HMM(n=3) on each subset independently Step 4 — Compare log-likelihood: HHMM vs flat 9-state HMM Step 5 — Plot all 9 sub-regimes on the price chart ``` **Validation criterion:** A 3×3 HHMM has ~45 parameters. A flat 9-state HMM has ~99 parameters. If the HHMM matches flat performance at half the parameter count, the hierarchical structure is real and the manifold has genuine multi-scale geometry. **If validated:** The answer to "how many states?" is resolution-dependent — approximately $3^d$ states at depth $d$. The market manifold is fractal-like across scales. **If falsified:** The manifold is flat below the macro level — also an important result, constraining the model architecture for production systems. --- ## Summary | Result | Status | Next Step | |---|---|---| | 1. Candle manifold — natural gradient upgrade | Proven (math) | Implement in Lab 4b | | 2. Automaton-manifold correspondence | Strong hypothesis | Prove for input-output HMMs | | 3. Recursive sub-regimes — HHMM | Open conjecture | Run Lab 4b experiment | These three results form a coherent research arc: the observation space has manifold geometry (Result 1), all automaton systems share this structure (Result 2), and the manifold has recursive sub-structure at multiple scales (Result 3). Lab 4b tests Result 3 empirically while simultaneously providing a testbed for the natural gradient implementation of Result 1. --- ## References 1. Baum, L.E. et al. (1970). A maximization technique in the statistical analysis of probabilistic functions of Markov chains. *Ann. Math. Statist.*, 41(1), 164–171. 2. Baum, L.E. and Sell, G.R. Growth transformation for functions on manifolds. *Pacific J. Math.* 3. Amari, S. (1998). Natural gradient works efficiently in learning. *Neural Computation*, 10(2), 251–276. 4. Cappé, O., Moulines, E., and Rydén, T. (2005). *Inference in Hidden Markov Models.* Springer. 5. Fine, S., Singer, Y., and Tishby, N. (1998). The hierarchical hidden Markov model. *Machine Learning*, 32, 41–62. 6. Chan, D. (2026). Observing HMM Everywhere. *AI-Symbiosis Research.* 7. Chan, D. (2026). Analysis: Inference in Hidden Markov Models. *AI-Symbiosis Research.*