Analysis: Inference in Hidden Markov Models
#

Cappé, Moulines & Rydén (2005)
#

David Chan, Claude Sonnet 4.6 AI-Symbiosis Research · April 2026

📄 Download PDF · Download MD

📎 Further Analysis: Further Analysis PDF · Further Analysis MD


Page 1 — Chapter 2 & 5 Deep Dive
#


Chapter 2: Filtering and Smoothing Recursions
#

The chapter opens with a clean separation of three problems:

Definition 18 — The Three Inference Tasks:

  • Smoothing: $\phi_{\nu,k|n}$ — distribution of $X_k$ given all $Y_0, \ldots, Y_n$. Best estimate, uses future data.
  • Filtering: $\phi_{\nu,k|k}$ — distribution of $X_k$ given past $Y_0, \ldots, Y_k$. Real-time estimate.
  • Prediction: $\phi_{\nu,k+p|k}$ — distribution of $X_{k+p}$ given data up to $k$. Forward-looking.

The book focuses on fixed-interval smoothing — $n$ is fixed, and you want all $\phi_{\nu,k|n}$ simultaneously. This is the offline training case.


The Forward-Backward Decomposition:

Any smoothed distribution factors as:

$$\phi_{\nu,k|n}(f) = L^{-1}_{\nu,n} \cdot \alpha_{\nu,k}(f \cdot \beta_{k|n})$$

The forward measure $\alpha_{\nu,k}$ accumulates evidence left-to-right. The backward function $\beta_{k|n}$ accumulates evidence right-to-left. Multiply them and normalize — you get the full posterior.


Proposition 25 — Normalized Recursion (the one you implement):

$$c_{\nu,k} = \iint \phi_{\nu,k-1}(dx)\, Q(x,dx')\, g_k(x')$$$$\phi_{\nu,k}(f) = c^{-1}_{\nu,k} \int f(x') \int \phi_{\nu,k-1}(dx)\, Q(x,dx')\, g_k(x')$$

Note: The scalars $c_{\nu,k}$ normalize at each timestep, preventing floating point underflow on long sequences. The total log-likelihood is simply $\sum_k \log c_{\nu,k}$ — it falls out of the forward pass for free.


Forward Smoothing Kernels (Definition 30):

$$F_{k|n}(x, A) = [\beta_{k|n}(x)]^{-1} \int_A Q(x, dx')\, g_{k+1}(x')\, \beta_{k+1|n}(x')$$

This defines how probability mass flows forward through the smoother. The key structural insight: conditionally on $Y_{0:n}$, the time-reversed sequence $\bar{X}_k = X_{n-k}$ is itself a non-homogeneous Markov chain with these kernels as transitions. The backward pass is a time-reversed Markov chain.

Note: The book explicitly notes that filtering theory (continuous-time, Shiryaev 1966, Wonham 1965) and discrete-time HMMs evolved as two mostly independent fields. The continuous-time analogues involve stochastic differential equations — deliberately deferred.


Chapter 5: Maximum Likelihood Inference
#

The Setup: Parameters $\theta = (\nu, Q, g)$ are unknown. You observe $Y_{0:n}$, want $\hat{\theta}_{MLE}$. Direct maximization of $\ell_n(\theta) = \log L_{\nu,n}(\theta)$ is hard — the hidden states make it a missing data problem.


Proposition 98 — The Fundamental EM Inequality:

$$\ell(\theta) - \ell(\theta') \geq Q(\theta;\, \theta') - Q(\theta';\, \theta')$$

where

$$Q(\theta;\, \theta') = E_{\theta'}\left[\log f(X_{0:n}, Y_{0:n};\, \theta) \mid Y_{0:n}\right]$$

Note: Maximizing $Q$ over $\theta$ is guaranteed to increase $\ell$. Every single EM iteration improves the model — you cannot go backward. This is Baum’s 1970 proof.


Fisher’s Identity (eq. 5.28):

$$\nabla_\theta \ell_n(\theta) = E_\theta\left[\nabla_\theta \log \nu(X_0) \mid Y_{0:n}\right] + \sum_{k=0}^{n} E_\theta\left[\nabla_\theta \log g_k(X_k) \mid Y_{0:n}\right] + \sum_{k=0}^{n-1} E_\theta\left[\nabla_\theta \log q(X_k, X_{k+1}) \mid Y_{0:n}\right]$$

Note: The gradient of the log-likelihood decomposes into three smoothed expectations — initial state, emissions, transitions. Each is computable via forward-backward without automatic differentiation. Use this if you want gradient ascent instead of EM.


Louis’ Identity — Observed Information:

$$-\nabla^2_\theta \ell = -\nabla^2_\theta Q(\theta;\theta) + \text{Var}_\theta\left[\nabla_\theta \log f(X_{0:n}, Y_{0:n};\theta) \mid Y_{0:n}\right]$$

Note: Observed information = curvature of $Q$ minus conditional variance of complete score. Gives you uncertainty estimates on parameters — useful for knowing how confident your regime estimates are.


Normal HMM — The Explicit EM Updates (Sec. 5.3):

$$\mu^*_i = \frac{\sum_k \phi_{k|n}(i)\, Y_k}{\sum_k \phi_{k|n}(i)}$$$$\upsilon^*_i = \frac{\sum_k \phi_{k|n}(i)\, Y_k^2}{\sum_k \phi_{k|n}(i)} - \left(\mu^*_i\right)^2$$$$q^*_{ij} = \frac{\sum_k \phi_{k-1:k|n}(i,j)}{\sum_k \phi_{k|n}(i)}$$

Note: Every update is a weighted average, weighted by $\phi_{k|n}(i)$ — your current posterior belief about which regime was active at time $k$. If you were 90% sure you were in the “bull” regime on day $k$, that day’s return gets 90% weight in the bull mean estimate.


Gaussian Linear State-Space Extension (Sec. 5.4):

For continuous hidden states (e.g. hidden volatility), Kalman smoother replaces forward-backward:

$$A^* = \left[\sum_{k=0}^{n-1} C_{k,k+1|n} + \hat{X}_{k|n}\hat{X}^T_{k+1|n}\right]^T \left[\sum_{k=0}^{n-1} \Sigma_{k|n} + \hat{X}_{k|n}\hat{X}^T_{k|n}\right]^{-1}$$$$\Upsilon^*_R = \frac{1}{n} \sum_{k=0}^{n-1} \left\{ \left[\Sigma_{k+1|n} + \hat{X}_{k+1|n}\hat{X}^T_{k+1|n}\right] - A^* \left[C_{k,k+1|n} + \hat{X}_{k|n}\hat{X}^T_{k+1|n}\right] \right\}$$

Note: Same EM structure, but now updating a state dynamics matrix $A$ and process noise covariance $\Upsilon_R$. The book notes Gaussian linear state-space models are the only important HMM subclass with tractable non-iterative estimators.


MLE Asymptotics (Ch. 6) — Three Guarantees:

As $n \to \infty$:

  1. $n^{-1}\ell_n(\theta) \to \ell(\theta)$ a.s. uniformly — log-likelihood converges to a continuous limit with unique maximum at $\theta^*$
  2. $n^{-1/2}\nabla_\theta \ell_n(\theta^*) \to \mathcal{N}(0, J(\theta^*))$ weakly — score is asymptotically normal
  3. $-n^{-1}\nabla^2_\theta \ell_n(\theta^*) \to J(\theta^*)$ a.s. — observed information converges to Fisher information

Note: With enough data, Baum-Welch finds the true parameters. Your regime estimates are statistically consistent.


Page 2 — KEGA Gap Analysis
#


What is KEGA?
#

KEGA (Knowledge Extension via Gap Analysis) is a methodology for deriving implied knowledge from coherent systems. Applied here: given the book’s internal structure, what theorems or extensions does it imply must exist — either explicitly deferred, structurally absent, or left as open problems?


Gap 1 — The Continuous-Time Bridge
#

What the book says: “Filtering theory and hidden Markov models evolved as two mostly independent fields.” (Ch. 2, p.14)

The gap: The book proves forward-backward for discrete-time HMMs. Continuous-time filtering (Kushner-Stratonovich, Zakai equation) uses stochastic differential equations. These two frameworks solve the same problem using different mathematical tools and are not formally connected in this text.

Implied theorem: There exists a limiting procedure $\Delta t \to 0$ taking the discrete normalized recursion (Prop. 25) to the Zakai SDE for the unnormalized filter:

$$d\sigma_t(f) = \sigma_t(Lf)\, dt + \sigma_t(fh)\, dY_t$$

The normalization constants $c_{\nu,k}$ should converge to the innovation process. The forward smoothing kernels should converge to the Rauch-Tung-Striebel backward SDE.

Why it matters for markets: High-frequency trading operates in continuous time. A unified discrete-to-continuous HMM filter would allow regime detection at tick level without discretization error.


Gap 2 — The Natural Gradient (Baum-Sell Manifold)
#

What the book says: EM updates maximize $Q(\theta; \theta')$ over $\theta$. The parameter space is treated as flat Euclidean throughout.

The gap: The transition matrix $Q$ lives on a simplex manifold — each row sums to 1. The emission parameters $(\mu_i, \sigma_i)$ live on a curved statistical manifold with Fisher information metric $G(\theta)$. Standard EM ignores this curvature. Baum and Sell proved (in a paper referenced in the 1970 paper but never widely cited) that growth transformations generalize to functions on manifolds — but this book never uses it.

Implied algorithm: Replace the flat M-step with a natural gradient step:

$$\theta_{t+1} = \theta_t + \eta \cdot G(\theta_t)^{-1} \nabla_\theta Q(\theta_t;\, \theta_t)$$

where $G(\theta)$ is the Fisher information matrix. This respects the geometry of the parameter space and is known to converge in fewer iterations (Amari, 1998).

The missing bridge: Baum 1970 → Baum-Sell manifold paper → Amari information geometry (1985). None of these cite each other. The unified treatment does not exist.


Gap 3 — Online / Streaming EM
#

What the book says: Baum-Welch requires the full sequence $Y_{0:n}$ stored in memory — it’s a batch algorithm. Online variants are mentioned but deferred.

The gap: For real-time trading you cannot rerun Baum-Welch on the full history every bar. You need $\theta_t$ to update as new data arrives at $O(s^2)$ cost per step.

Implied algorithm: A stochastic approximation version of the M-step:

$$\theta_{t+1} = \theta_t + \gamma_t \left.\nabla_\theta Q(\theta_t;\, \theta_t)\right|_{\text{single observation}}$$

where $\{\gamma_t\}$ are Robbins-Monro step sizes satisfying $\sum \gamma_t = \infty$, $\sum \gamma_t^2 < \infty$.

Open problem: Proving convergence of this recursion under mild market non-stationarity — where the true $\theta^*$ drifts slowly over time. Standard convergence proofs assume stationarity. This is the theorem that makes Lab 4 deployable in live trading.


Gap 4 — Non-Stationary Transition Matrix
#

What the book says: The transition matrix $Q = (q_{ij})$ is fixed and time-homogeneous throughout.

The gap: Markets are non-stationary. The probability of transitioning from bull to bear in 2008 is not the same as in 2021. The book has no mechanism for time-varying $Q_k$.

Implied extension: Replace $q_{ij}$ with a covariate-driven function:

$$q_{ij}(k) = \text{softmax}(W \cdot z_k)_{ij}$$

where $z_k$ is a vector of observable macro covariates (VIX, yield curve slope, credit spreads). This is an input-output HMM. The EM updates become:

$$q^*_{ij}(k) = f\!\left(\text{covariates}_k,\; \phi_{k-1:k|n}(i,j)\right)$$

Open problem: Asymptotic theory for non-stationary $Q_k$. The current Ch. 6 consistency results assume stationarity. They do not apply when $Q_k$ varies. A new convergence theorem is needed.


Gap 5 — Model Selection: How Many Regimes?
#

What the book says: Throughout the book, the number of hidden states $s$ is assumed known.

The gap: In practice you don’t know if markets have 2, 3, or 4 regimes. Standard AIC/BIC criteria don’t apply because the HMM likelihood is irregular at boundary cases — when two states merge, the Fisher information matrix is singular. The model is not identifiable at the boundary.

Implied theorem: A penalized likelihood criterion specific to HMMs:

$$\text{IC}(s) = -2\ell_n(\hat\theta_s) + \text{pen}(s, n)$$

where $\text{pen}(s, n)$ must grow faster than $\log n$ to account for the irregular boundary. The correct rate is believed to be $O(s^2 \log n)$ but this has not been proved with full generality for continuous emission HMMs.


KEGA Summary Table
#

GapTypeDifficultyPriority for Lab 4
1. Continuous-time bridgeTheoreticalHardLow
2. Natural gradient on manifoldAlgorithmicMediumHigh
3. Online / streaming EMAlgorithmicMediumCritical
4. Non-stationary transition matrixModel extensionMediumHigh
5. Model selection ($s$ unknown)StatisticalHardMedium

The Next Theorem Worth Proving
#

Gap 3 — a convergent online EM for HMMs under mild non-stationarity.

This is the one that makes Lab 4 deployable in live trading. The other four gaps are intellectually interesting but not blocking. Gap 3 is blocking: without it, the regime detector must retrain from scratch on the full history at each step, which is $O(n)$ per bar and infeasible at scale.

The proof strategy: extend the Cappé (2011) online EM results to a slowly time-varying parameter setting using a tracking argument — show that if $\|\theta^*_t - \theta^*_{t-1}\| \leq \epsilon$ for small $\epsilon$, the online EM estimate $\hat\theta_t$ stays within $O(\epsilon / \gamma_t)$ of $\theta^*_t$. This requires a uniform forgetting rate (Ch. 3 results) plus a persistence-of-excitation condition on the observations.

That theorem, if proved, closes the gap between the theory in this book and a production-grade regime detector.


References
#

  1. Cappé, O., Moulines, E., and Rydén, T. (2005). Inference in Hidden Markov Models. Springer.
  2. Baum, L.E., Petrie, T., Soules, G., and Weiss, N. (1970). A maximization technique occurring in the statistical analysis of probabilistic functions of Markov chains. Annals of Mathematical Statistics, 41(1), 164–171.
  3. Baum, L.E. and Sell, G.R. Growth transformation for functions on manifolds. Pacific Journal of Mathematics.
  4. Amari, S. (1998). Natural gradient works efficiently in learning. Neural Computation, 10(2), 251–276.
  5. Cappé, O. (2011). Online EM algorithm for hidden Markov models. Journal of Computational and Graphical Statistics, 20(3), 728–749.
  6. Kushner, H.J. (1964). On the differential equations satisfied by conditional probability densities of Markov processes. SIAM Journal on Control, 2, 106–119.
  7. Chan, D. (2026). KEGA: Knowledge Extension via Gap Analysis. AI-Symbiosis Research.
  8. Chan, D. (2026). Observing HMM Everywhere. AI-Symbiosis Research.