Analysis: Inference in Hidden Markov Models#
Cappé, Moulines & Rydén (2005)#
David Chan, Claude Sonnet 4.6 AI-Symbiosis Research · April 2026
📎 Further Analysis: Further Analysis PDF · Further Analysis MD
Page 1 — Chapter 2 & 5 Deep Dive#
Chapter 2: Filtering and Smoothing Recursions#
The chapter opens with a clean separation of three problems:
Definition 18 — The Three Inference Tasks:
- Smoothing: $\phi_{\nu,k|n}$ — distribution of $X_k$ given all $Y_0, \ldots, Y_n$. Best estimate, uses future data.
- Filtering: $\phi_{\nu,k|k}$ — distribution of $X_k$ given past $Y_0, \ldots, Y_k$. Real-time estimate.
- Prediction: $\phi_{\nu,k+p|k}$ — distribution of $X_{k+p}$ given data up to $k$. Forward-looking.
The book focuses on fixed-interval smoothing — $n$ is fixed, and you want all $\phi_{\nu,k|n}$ simultaneously. This is the offline training case.
The Forward-Backward Decomposition:
Any smoothed distribution factors as:
$$\phi_{\nu,k|n}(f) = L^{-1}_{\nu,n} \cdot \alpha_{\nu,k}(f \cdot \beta_{k|n})$$The forward measure $\alpha_{\nu,k}$ accumulates evidence left-to-right. The backward function $\beta_{k|n}$ accumulates evidence right-to-left. Multiply them and normalize — you get the full posterior.
Proposition 25 — Normalized Recursion (the one you implement):
$$c_{\nu,k} = \iint \phi_{\nu,k-1}(dx)\, Q(x,dx')\, g_k(x')$$$$\phi_{\nu,k}(f) = c^{-1}_{\nu,k} \int f(x') \int \phi_{\nu,k-1}(dx)\, Q(x,dx')\, g_k(x')$$Note: The scalars $c_{\nu,k}$ normalize at each timestep, preventing floating point underflow on long sequences. The total log-likelihood is simply $\sum_k \log c_{\nu,k}$ — it falls out of the forward pass for free.
Forward Smoothing Kernels (Definition 30):
$$F_{k|n}(x, A) = [\beta_{k|n}(x)]^{-1} \int_A Q(x, dx')\, g_{k+1}(x')\, \beta_{k+1|n}(x')$$This defines how probability mass flows forward through the smoother. The key structural insight: conditionally on $Y_{0:n}$, the time-reversed sequence $\bar{X}_k = X_{n-k}$ is itself a non-homogeneous Markov chain with these kernels as transitions. The backward pass is a time-reversed Markov chain.
Note: The book explicitly notes that filtering theory (continuous-time, Shiryaev 1966, Wonham 1965) and discrete-time HMMs evolved as two mostly independent fields. The continuous-time analogues involve stochastic differential equations — deliberately deferred.
Chapter 5: Maximum Likelihood Inference#
The Setup: Parameters $\theta = (\nu, Q, g)$ are unknown. You observe $Y_{0:n}$, want $\hat{\theta}_{MLE}$. Direct maximization of $\ell_n(\theta) = \log L_{\nu,n}(\theta)$ is hard — the hidden states make it a missing data problem.
Proposition 98 — The Fundamental EM Inequality:
$$\ell(\theta) - \ell(\theta') \geq Q(\theta;\, \theta') - Q(\theta';\, \theta')$$where
$$Q(\theta;\, \theta') = E_{\theta'}\left[\log f(X_{0:n}, Y_{0:n};\, \theta) \mid Y_{0:n}\right]$$Note: Maximizing $Q$ over $\theta$ is guaranteed to increase $\ell$. Every single EM iteration improves the model — you cannot go backward. This is Baum’s 1970 proof.
Fisher’s Identity (eq. 5.28):
$$\nabla_\theta \ell_n(\theta) = E_\theta\left[\nabla_\theta \log \nu(X_0) \mid Y_{0:n}\right] + \sum_{k=0}^{n} E_\theta\left[\nabla_\theta \log g_k(X_k) \mid Y_{0:n}\right] + \sum_{k=0}^{n-1} E_\theta\left[\nabla_\theta \log q(X_k, X_{k+1}) \mid Y_{0:n}\right]$$Note: The gradient of the log-likelihood decomposes into three smoothed expectations — initial state, emissions, transitions. Each is computable via forward-backward without automatic differentiation. Use this if you want gradient ascent instead of EM.
Louis’ Identity — Observed Information:
$$-\nabla^2_\theta \ell = -\nabla^2_\theta Q(\theta;\theta) + \text{Var}_\theta\left[\nabla_\theta \log f(X_{0:n}, Y_{0:n};\theta) \mid Y_{0:n}\right]$$Note: Observed information = curvature of $Q$ minus conditional variance of complete score. Gives you uncertainty estimates on parameters — useful for knowing how confident your regime estimates are.
Normal HMM — The Explicit EM Updates (Sec. 5.3):
$$\mu^*_i = \frac{\sum_k \phi_{k|n}(i)\, Y_k}{\sum_k \phi_{k|n}(i)}$$$$\upsilon^*_i = \frac{\sum_k \phi_{k|n}(i)\, Y_k^2}{\sum_k \phi_{k|n}(i)} - \left(\mu^*_i\right)^2$$$$q^*_{ij} = \frac{\sum_k \phi_{k-1:k|n}(i,j)}{\sum_k \phi_{k|n}(i)}$$Note: Every update is a weighted average, weighted by $\phi_{k|n}(i)$ — your current posterior belief about which regime was active at time $k$. If you were 90% sure you were in the “bull” regime on day $k$, that day’s return gets 90% weight in the bull mean estimate.
Gaussian Linear State-Space Extension (Sec. 5.4):
For continuous hidden states (e.g. hidden volatility), Kalman smoother replaces forward-backward:
$$A^* = \left[\sum_{k=0}^{n-1} C_{k,k+1|n} + \hat{X}_{k|n}\hat{X}^T_{k+1|n}\right]^T \left[\sum_{k=0}^{n-1} \Sigma_{k|n} + \hat{X}_{k|n}\hat{X}^T_{k|n}\right]^{-1}$$$$\Upsilon^*_R = \frac{1}{n} \sum_{k=0}^{n-1} \left\{ \left[\Sigma_{k+1|n} + \hat{X}_{k+1|n}\hat{X}^T_{k+1|n}\right] - A^* \left[C_{k,k+1|n} + \hat{X}_{k|n}\hat{X}^T_{k+1|n}\right] \right\}$$Note: Same EM structure, but now updating a state dynamics matrix $A$ and process noise covariance $\Upsilon_R$. The book notes Gaussian linear state-space models are the only important HMM subclass with tractable non-iterative estimators.
MLE Asymptotics (Ch. 6) — Three Guarantees:
As $n \to \infty$:
- $n^{-1}\ell_n(\theta) \to \ell(\theta)$ a.s. uniformly — log-likelihood converges to a continuous limit with unique maximum at $\theta^*$
- $n^{-1/2}\nabla_\theta \ell_n(\theta^*) \to \mathcal{N}(0, J(\theta^*))$ weakly — score is asymptotically normal
- $-n^{-1}\nabla^2_\theta \ell_n(\theta^*) \to J(\theta^*)$ a.s. — observed information converges to Fisher information
Note: With enough data, Baum-Welch finds the true parameters. Your regime estimates are statistically consistent.
Page 2 — KEGA Gap Analysis#
What is KEGA?#
KEGA (Knowledge Extension via Gap Analysis) is a methodology for deriving implied knowledge from coherent systems. Applied here: given the book’s internal structure, what theorems or extensions does it imply must exist — either explicitly deferred, structurally absent, or left as open problems?
Gap 1 — The Continuous-Time Bridge#
What the book says: “Filtering theory and hidden Markov models evolved as two mostly independent fields.” (Ch. 2, p.14)
The gap: The book proves forward-backward for discrete-time HMMs. Continuous-time filtering (Kushner-Stratonovich, Zakai equation) uses stochastic differential equations. These two frameworks solve the same problem using different mathematical tools and are not formally connected in this text.
Implied theorem: There exists a limiting procedure $\Delta t \to 0$ taking the discrete normalized recursion (Prop. 25) to the Zakai SDE for the unnormalized filter:
$$d\sigma_t(f) = \sigma_t(Lf)\, dt + \sigma_t(fh)\, dY_t$$The normalization constants $c_{\nu,k}$ should converge to the innovation process. The forward smoothing kernels should converge to the Rauch-Tung-Striebel backward SDE.
Why it matters for markets: High-frequency trading operates in continuous time. A unified discrete-to-continuous HMM filter would allow regime detection at tick level without discretization error.
Gap 2 — The Natural Gradient (Baum-Sell Manifold)#
What the book says: EM updates maximize $Q(\theta; \theta')$ over $\theta$. The parameter space is treated as flat Euclidean throughout.
The gap: The transition matrix $Q$ lives on a simplex manifold — each row sums to 1. The emission parameters $(\mu_i, \sigma_i)$ live on a curved statistical manifold with Fisher information metric $G(\theta)$. Standard EM ignores this curvature. Baum and Sell proved (in a paper referenced in the 1970 paper but never widely cited) that growth transformations generalize to functions on manifolds — but this book never uses it.
Implied algorithm: Replace the flat M-step with a natural gradient step:
$$\theta_{t+1} = \theta_t + \eta \cdot G(\theta_t)^{-1} \nabla_\theta Q(\theta_t;\, \theta_t)$$where $G(\theta)$ is the Fisher information matrix. This respects the geometry of the parameter space and is known to converge in fewer iterations (Amari, 1998).
The missing bridge: Baum 1970 → Baum-Sell manifold paper → Amari information geometry (1985). None of these cite each other. The unified treatment does not exist.
Gap 3 — Online / Streaming EM#
What the book says: Baum-Welch requires the full sequence $Y_{0:n}$ stored in memory — it’s a batch algorithm. Online variants are mentioned but deferred.
The gap: For real-time trading you cannot rerun Baum-Welch on the full history every bar. You need $\theta_t$ to update as new data arrives at $O(s^2)$ cost per step.
Implied algorithm: A stochastic approximation version of the M-step:
$$\theta_{t+1} = \theta_t + \gamma_t \left.\nabla_\theta Q(\theta_t;\, \theta_t)\right|_{\text{single observation}}$$where $\{\gamma_t\}$ are Robbins-Monro step sizes satisfying $\sum \gamma_t = \infty$, $\sum \gamma_t^2 < \infty$.
Open problem: Proving convergence of this recursion under mild market non-stationarity — where the true $\theta^*$ drifts slowly over time. Standard convergence proofs assume stationarity. This is the theorem that makes Lab 4 deployable in live trading.
Gap 4 — Non-Stationary Transition Matrix#
What the book says: The transition matrix $Q = (q_{ij})$ is fixed and time-homogeneous throughout.
The gap: Markets are non-stationary. The probability of transitioning from bull to bear in 2008 is not the same as in 2021. The book has no mechanism for time-varying $Q_k$.
Implied extension: Replace $q_{ij}$ with a covariate-driven function:
$$q_{ij}(k) = \text{softmax}(W \cdot z_k)_{ij}$$where $z_k$ is a vector of observable macro covariates (VIX, yield curve slope, credit spreads). This is an input-output HMM. The EM updates become:
$$q^*_{ij}(k) = f\!\left(\text{covariates}_k,\; \phi_{k-1:k|n}(i,j)\right)$$Open problem: Asymptotic theory for non-stationary $Q_k$. The current Ch. 6 consistency results assume stationarity. They do not apply when $Q_k$ varies. A new convergence theorem is needed.
Gap 5 — Model Selection: How Many Regimes?#
What the book says: Throughout the book, the number of hidden states $s$ is assumed known.
The gap: In practice you don’t know if markets have 2, 3, or 4 regimes. Standard AIC/BIC criteria don’t apply because the HMM likelihood is irregular at boundary cases — when two states merge, the Fisher information matrix is singular. The model is not identifiable at the boundary.
Implied theorem: A penalized likelihood criterion specific to HMMs:
$$\text{IC}(s) = -2\ell_n(\hat\theta_s) + \text{pen}(s, n)$$where $\text{pen}(s, n)$ must grow faster than $\log n$ to account for the irregular boundary. The correct rate is believed to be $O(s^2 \log n)$ but this has not been proved with full generality for continuous emission HMMs.
KEGA Summary Table#
| Gap | Type | Difficulty | Priority for Lab 4 |
|---|---|---|---|
| 1. Continuous-time bridge | Theoretical | Hard | Low |
| 2. Natural gradient on manifold | Algorithmic | Medium | High |
| 3. Online / streaming EM | Algorithmic | Medium | Critical |
| 4. Non-stationary transition matrix | Model extension | Medium | High |
| 5. Model selection ($s$ unknown) | Statistical | Hard | Medium |
The Next Theorem Worth Proving#
Gap 3 — a convergent online EM for HMMs under mild non-stationarity.
This is the one that makes Lab 4 deployable in live trading. The other four gaps are intellectually interesting but not blocking. Gap 3 is blocking: without it, the regime detector must retrain from scratch on the full history at each step, which is $O(n)$ per bar and infeasible at scale.
The proof strategy: extend the Cappé (2011) online EM results to a slowly time-varying parameter setting using a tracking argument — show that if $\|\theta^*_t - \theta^*_{t-1}\| \leq \epsilon$ for small $\epsilon$, the online EM estimate $\hat\theta_t$ stays within $O(\epsilon / \gamma_t)$ of $\theta^*_t$. This requires a uniform forgetting rate (Ch. 3 results) plus a persistence-of-excitation condition on the observations.
That theorem, if proved, closes the gap between the theory in this book and a production-grade regime detector.
References#
- Cappé, O., Moulines, E., and Rydén, T. (2005). Inference in Hidden Markov Models. Springer.
- Baum, L.E., Petrie, T., Soules, G., and Weiss, N. (1970). A maximization technique occurring in the statistical analysis of probabilistic functions of Markov chains. Annals of Mathematical Statistics, 41(1), 164–171.
- Baum, L.E. and Sell, G.R. Growth transformation for functions on manifolds. Pacific Journal of Mathematics.
- Amari, S. (1998). Natural gradient works efficiently in learning. Neural Computation, 10(2), 251–276.
- Cappé, O. (2011). Online EM algorithm for hidden Markov models. Journal of Computational and Graphical Statistics, 20(3), 728–749.
- Kushner, H.J. (1964). On the differential equations satisfied by conditional probability densities of Markov processes. SIAM Journal on Control, 2, 106–119.
- Chan, D. (2026). KEGA: Knowledge Extension via Gap Analysis. AI-Symbiosis Research.
- Chan, D. (2026). Observing HMM Everywhere. AI-Symbiosis Research.