<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://shenmuxing.github.io/feed.xml" rel="self" type="application/atom+xml"/><link href="https://shenmuxing.github.io/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-06-29T01:15:44+00:00</updated><id>https://shenmuxing.github.io/feed.xml</id><subtitle>Personal academic website of Jingye Zhao, a Ph.D. student at Shanghai Jiao Tong University working on reinforcement learning theory and applications. </subtitle><entry><title type="html">Concentration Inequalities for Reinforcement Learning</title><link href="https://shenmuxing.github.io/blog/2026/concentration-inequalities-for-reinforcement-learning/" rel="alternate" type="text/html" title="Concentration Inequalities for Reinforcement Learning"/><published>2026-01-31T00:00:00+00:00</published><updated>2026-01-31T00:00:00+00:00</updated><id>https://shenmuxing.github.io/blog/2026/concentration-inequalities-for-reinforcement-learning</id><content type="html" xml:base="https://shenmuxing.github.io/blog/2026/concentration-inequalities-for-reinforcement-learning/"><![CDATA[<p>This note collects concentration inequalities that frequently appear in reinforcement learning analysis. It is meant as a working reference rather than a complete survey.</p> <h2 id="1-markov-inequality-and-chernoff-bound">1. Markov inequality and Chernoff bound</h2> <p>If \(X\) is a nonnegative random variable and \(a&gt;0\), then the probability that \(X\) is at least \(a\) is at most the expectation of \(X\) divided by \(a\):</p> \[\mathrm{P}(X \geq a) \leq \frac{\mathrm{E}(X)}{a}\] <p>In probability theory, Markov’s inequality gives an upper bound on the probability that a non-negative random variable is greater than or equal to some positive constant. Markov’s inequality is tight in the sense that for each chosen positive constant, there exists a random variable such that the inequality is in fact an equality.</p> <p>The generic Chernoff bound for a random variable \(X\) is attained by applying Markov’s inequality to \(e^{t X}\) (which is why it is sometimes called the exponential Markov or exponential moments bound). For positive \(t\) this gives a bound on the right tail of \(X\) in terms of its moment-generating function \(M(t)=\mathrm{E}\left(e^{t X}\right)\):</p> \[\mathrm{P}(X \geq a)=\mathrm{P}\left(e^{t X} \geq e^{t a}\right) \leq M(t) e^{-t a} \quad(t&gt;0)\] <p>Chernoff bound is an exponentially decreasing upper bound on the tail of a random variable based on its moment generating function.</p> <p>It is a sharper bound than the first- or second-moment-based tail bounds such as Markov’s inequality or <a href="https://en.wikipedia.org/wiki/Chebyshev%27s_inequality" title="Chebyshev's inequality">Chebyshev’s inequality</a>, which only yield power-law bounds on tail decay. However, when applied to sums the Chernoff bound requires the random variables to be independent, a condition that is not required by either Markov’s inequality or Chebyshev’s inequality.</p> <p>The Chernoff bound is related to the Bernstein inequalities. It is also used to prove Hoeffding’s inequality, Bennett’s inequality, and <a href="https://en.wikipedia.org/wiki/Doob_martingale#McDiarmid's_inequality" title="Doob martingale">McDiarmid’s inequality</a>.</p> <h2 id="2-hoeffdings-inequality">2. Hoeffding’s inequality</h2> <p>Suppose \(X_{1}, X_{2}, \ldots, X_{n}\) are a sequence of independent, identically distributed (i.i.d.) random variables with mean \(\mu\). Let \(\bar{X}_{n} = n^{-1}\sum_{i=1}^{n}X_{i}\). Suppose that \(X_{i} \in [b_{-}, b_{+}]\) with probability 1, then</p> \[P (\bar {X} _ {n} \geq \mu + \epsilon) \leq e ^ {- 2 n \epsilon^ {2} / (b _ {+} - b _ {-}) ^ {2}}.\] <p>Similarly,</p> \[P (\bar {X} _ {n} \leq \mu - \epsilon) \leq e ^ {- 2 n \epsilon^ {2} / (b _ {+} - b _ {-}) ^ {2}}.\] <p>The Chernoff bound implies that with probability \(1 - \delta\):</p> \[\bar {X} _ {n} - E X \leq (b _ {+} - b _ {-}) \sqrt {\ln (1 / \delta) / (2 n)}.\] <p>Hoeffding’s inequality is a special case of the Azuma–Hoeffding inequality and <a href="https://en.wikipedia.org/wiki/McDiarmid%27s_inequality" title="McDiarmid's inequality">McDiarmid’s inequality</a>. It is similar to the Chernoff bound, but tends to be less sharp, in particular when the variance of the random variables is small.</p> <h2 id="3-sub-gaussian-random-variable">3. Sub-Gaussian Random Variable</h2> <p>A random variable \(X\) is \(\sigma\)-subGaussian if for all \(\lambda \in \mathbb{R}\), it holds that \(\mathbb{E}[\exp (\lambda X)] \leq \exp (\lambda^2\sigma^2 / 2)\).</p> <p>One can show that a Gaussian random variable with zero mean and standard deviation \(\sigma\) is a \(\sigma\)-subGaussian random variable.</p> <p>The following theorem shows that the tails of a \(\sigma\)-subGaussian random variable decay approximately as fast as that of a Gaussian variable with zero mean and standard deviation \(\sigma\).</p> <p>The following lemma shows that the sum of independent sub-Gaussian variables is still sub-Gaussian.</p> <p><strong>Lemma 1.</strong> Suppose that \(X_{1}\) and \(X_{2}\) are independent and \(\sigma_{1}\) and \(\sigma_{2}\) subGaussian, respectively. Then for any \(c \in \mathbb{R}\), we have \(cX\) being \(\vert c \vert \sigma\)-subGaussian. We also have \(X_{1} + X_{2}\) being \(\sqrt{\sigma_1^2 + \sigma_2^2}\)-subGaussian.</p> <h2 id="4-hoeffding-azuma-inequality">4. Hoeffding-Azuma Inequality</h2> <p><strong>Definition 1 (Martingale Difference Sequence, M. D. S.).</strong> A sequence of random variables \(\{X_{n}\}_{n=1}^{\infty}\) is called a martingale difference sequence, or m. d. s., with respect to a filtration \(\{\mathcal{F}_{n}\}_{n=0}^{\infty}\) if for all \(n, \mathbb{E}[\vert X_{n}\vert]&lt;\infty, X_{n} \in \mathcal{F}_{n}\), and \(\mathbb{E}[X_{n} \mid \mathcal{F}_{n-1}]=0\) hold.</p> <p>As the name implies, the sum sequence of a martingale difference sequence is called a martingale.</p> <p>The Azuma-Hoeffding inequality provides an exponentially decaying tail bound for the sum of a martingale difference sequence:</p> <p>Suppose \(X_{1},\ldots ,X_{N}\) is a martingale difference sequence where each \(X_{i}\) is a \(\sigma_i\) sub-Gaussian. Then, for all \(\epsilon &gt;0\) and all positive integer \(K\):</p> \[P \left(\sum_ {i = 1} ^ {K} X _ {i} \geq \epsilon\right) \leq \exp \left(\frac {- \epsilon^ {2}}{2 \sum_ {i = 1} ^ {K} \sigma_ {i} ^ {2}}\right).\] <h2 id="5-bernsteins-inequality">5. Bernstein’s Inequality</h2> <p><strong>Lemma 2 (Bernstein’s inequality).</strong> Suppose \(X_{1},\ldots ,X_{n}\) are independent random variables. Let \(\bar{X}_n = n^{-1}\sum_{i = 1}^n X_i\), \(\mu = \mathbb{E}\bar{X}_n\), and \(Var(X_{i})\) denote the variance of \(X_{i}\). If \(X_{i} - EX_{i}\leq b\) for all \(i\), then</p> \[P (\bar {X} _ {n} \geq \mu + \epsilon) \leq \exp \left[ - \frac {n ^ {2} \epsilon^ {2}}{2 \sum_ {i = 1} ^ {n} \operatorname {V a r} (X _ {i}) + 2 n b \epsilon / 3} \right].\] <p>If all the variances are equal, the Bernstein inequality implies that, with probability at least \(1 - \delta\),</p> \[\bar {X} _ {n} - E X \leq \sqrt {2 \mathrm {V a r} (X) \ln (1 / \delta) / n} + \frac {2 b \ln (1 / \delta)}{3 n}.\] <p><strong>Lemma 3 (Bernstein’s Inequality for Martingales).</strong> Suppose \(X_{1}, X_{2} \ldots\) is a martingale difference sequence where \(\vert X_{i}\vert \leq M \in \mathbb{R}^{+}\) almost surely. Then for all positive \(t\) and \(n \in \mathbb{N}^{+}\), we have:</p> \[P \left(\sum_ {i = 1} ^ {n} X _ {i} \geq t\right) \leq \exp \left(- \frac {t ^ {2} / 2}{\sum_ {i = 1} ^ {n} \mathbb {E} \left[ X _ {i} ^ {2} \mid \mathcal{F}_{i-1} \right] + M t / 3}\right).\] <h2 id="6-concentration-for-discrete-distributions">6. Concentration for Discrete Distributions</h2> <p>Let \(z\) be a discrete random variable that takes values in \(\{1,\dots ,d\}\), distributed according to \(q\). We write \(q\) as a vector where \(\vec{q} = [\mathrm{Pr}(z = j)]_{j = 1}^{d}\). Assume we have \(N\) i. i. d samples, and that our empirical estimate of \(\vec{q}\) is \([\hat{q}]_j = \sum_{i = 1}^N\mathbf{1}[z_i = j] / N\).</p> <p>We have that \(\forall \epsilon &gt; 0\):</p> \[\Pr \left(\| \widehat {q} - \bar {q} \| _ {2} \geq 1 / \sqrt {N} + \epsilon\right) \leq e ^ {- N \epsilon^ {2}}.\] <p>which implies that:</p> \[\operatorname * {P r} \left(\| \widehat {q} - \overline \| _ {1} \geq \sqrt {d} (1 / \sqrt {N} + \epsilon)\right) \leq e ^ {- N \epsilon^ {2}}.\] <p>This result illustrates that the \(L_2\) bound is tighter than the \(L_1\) bound by roughly \(\sqrt{d}\). Geometrically, this stems from the fact that the \(L_2\) norm is Euclidean distance (straight line), whereas the \(L_1\) norm is Manhattan distance (grid path). Thus the \(L_1\) bound incurs an additional \(\sqrt{d}\) factor in the worst case.</p> <p>Intuitively, for any fixed vector, a higher-order norm yields a smaller value (i.e., tighter distance). We have the following relationship for \(x \in \mathbb{R}^{d}\):</p> \[\|x\|_{\infty} \leq\|x\|_{2} \leq\|x\|_{1} \leq \sqrt{d}\|x\|_{2} \leq d\|x\|_{\infty}.\] <h2 id="7-self-normalized-bound-for-vector-valued-martingales">7. Self-Normalized Bound for Vector-Valued Martingales</h2> <p>Let \(\{\varepsilon_i\}_{i=1}^{\infty}\) be a real-valued stochastic process with corresponding filtration \(\{\mathcal{F}_i\}_{i=1}^{\infty}\) such that \(\varepsilon_i\) is \(\mathcal{F}_i\) measurable, \(\mathbb{E}[\varepsilon_i \mid \mathcal{F}_{i-1}] = 0\) and \(\varepsilon_i\) is conditionally \(\sigma\)-subGaussian with \(\sigma \in \mathbb{R}^+\). Let \(\{X_i\}_{i=1}^{\infty}\) be a stochastic process with \(X_i \in \mathcal{H}\) (some Hilbert space) and \(X_i\) being \(\mathcal{F}_t\) measurable. Assume that a linear operator \(\Sigma: \mathcal{H} \to \mathcal{H}\) is positive definite, i.e., \(x^\top \Sigma x &gt; 0\) for any \(x \in \mathcal{H}\). For any \(t\), define the linear operator \(\Sigma_t = \Sigma_0 + \sum_{i=1}^{t} X_i X_i^\top\) (here \(xx^\top\) denotes outer-product in \(\mathcal{H}\)). With probability at least \(1 - \delta\), we have for all \(t \geq 1\):</p> \[\left\| \sum_ {i = 1} ^ {t} X _ {i} \varepsilon_ {i} \right\| _ {\Sigma_ {t} ^ {- 1}} ^ {2} \leq \sigma^ {2} \log \left(\frac {\det (\Sigma_ {t}) \det (\Sigma_ {0}) ^ {- 1}}{\delta^ {2}}\right).\] <h2 id="references">References</h2> <p>[1] Agarwal, A., Jiang, N., Kakade, S. M. &amp; Sun, W. (2022). <a href="https://rltheorybook.github.io/">Reinforcement Learning: Theory and Algorithms</a>. [2] Wikipedia. <a href="https://en.wikipedia.org/wiki/Markov%27s_inequality">Markov’s inequality</a>. [3] Wikipedia. <a href="https://en.wikipedia.org/wiki/Chernoff_bound">Chernoff bound</a>. [4] Wikipedia. <a href="https://en.wikipedia.org/wiki/Hoeffding%27s_inequality">Hoeffding’s inequality</a>. [5] Harin Lee. <a href="https://harinboy.github.io/posts/FreedmansInequality/">Freedman’s Inequality</a>.</p>]]></content><author><name></name></author><category term="theory"/><category term="machine-learning"/><category term="reinforcement-learning"/><category term="mathematics"/><category term="probability"/><summary type="html"><![CDATA[A summary of key concentration inequalities used in the analysis of reinforcement learning algorithms.]]></summary></entry></feed>