Le Cam's two-point method
\[\inf_{\hat\theta}\ \max_{j \in \{0, 1\}} \mathbb{E}_j\, d(\hat\theta, \theta_j) \ge \frac{s}{2}\big(1 - \lVert
P_0 - P_1 \rVert_{\mathrm{TV}}\big)\]
If \(d(\theta_0, \theta_1) \ge 2s\) and \(P_0, P_1\) are the corresponding distributions of the data.
- Gives
- A fast, clean minimax lower bound.
- Costs
- Two well-separated but statistically indistinguishable hypotheses.
- Wrong tool when
- You need dimension dependence, which takes many hypotheses.
Fano's inequality
\[\Pr(\hat V \ne V) \ge 1 - \frac{I(V; X) + \ln 2}{\ln M}\]
For \(V\) uniform over \(M\) hypotheses and observation \(X\).
- Gives
- Turns a large packing with small information into a minimax error.
- Costs
- \(M\) separated alternatives and control of their information or KL.
- Wrong tool when
- \(M = 2\); Le Cam is cleaner.
Pinsker and KL tensorization
\[\lVert P - Q \rVert_{\mathrm{TV}} \le \sqrt{\tfrac12 D_{\mathrm{KL}}(P \Vert Q)}, \qquad
D_{\mathrm{KL}}(P^{\otimes n} \Vert Q^{\otimes n}) = n\, D_{\mathrm{KL}}(P \Vert Q)\]
The second identity is for independent product observations.
- Gives
- Lets KL calculations feed Le Cam, and sets the critical separation as a function of \(n\).
- Costs
- Absolute continuity; independent observations for tensorization.
- Wrong tool when
- Observations are adaptive or interacting; use conditional chain rules instead.
Packing reduction
\[d(\theta_i, \theta_j) \ge 2s\ \ \forall i \ne j \;\Longrightarrow\; \text{error} < s \text{ gives an }
M\text{-way decoder}\]
Choose \(\{\theta_1, \dots, \theta_M\}\) separated in the metric you report.
- Gives
- A bridge from estimation to hypothesis testing.
- Costs
- Separation measured in the loss you actually report.
- Wrong tool when
- Parameters are separated but the predictions or risks are not.
Yao's minimax principle
\[\min_{\text{rand. } A}\ \max_{x}\ \mathbb{E}\, C(A, x) \;\ge\; \max_{\mu}\ \min_{\text{det. } a}\
\mathbb{E}_{x \sim \mu}\, C(a, x)\]
In finite settings.
- Gives
- Lower bounds for randomized algorithms from one hard input distribution and deterministic algorithms.
- Costs
- A correct minimax setup; care with infinite spaces and measurability.
- Wrong tool when
- Your lower bound already holds pointwise for randomized algorithms.