Capacity, Isostatic Jamming, and Full Replica Symmetry Breaking

Authors’ note

This manuscript is the result of a methodological experiment in using large language models to develop the rigorous mathematical theory of the negative perceptron, which has received significant attention in the physics literature as an accessible model for studying jamming and related phenomena.

With the exception of this note, the entire document is machine written. We have taken care to supervise this writing to ensure that the proofs meet a minimum standard of readability and that previous literature is appropriately cited. However, the resulting document still leaves much to be desired in terms of exposition and overall coherence. We are nonetheless releasing this manuscript on Hexagon to disseminate the results as quickly as possible, because we believe they may be of interest to the mathematics and physics communities.

We plan to take some time to write the results in this manuscript carefully for publication. While we focus on writing, we do not intend to produce further results for the negative perceptron. However, we invite readers to do so, because we believe there is much more to be done. Notably, it remains to establish the scaling solution of Franz, Parisi, Sevelev, Urbani, and Zamponi [ref-59] for the Parisi minimizers as κ↑κc\kappa\uparrow\kappa_{\mathrm{c}} and the predictions that follow from it. We welcome emails from anyone who obtains such results and will update the references accordingly. We also welcome suggestions concerning mathematical corrections, omitted citations, and attribution of ideas.

P.L. was partially supported by NSF grant DMS-2450004.

Notation

Shared notation. Both Parts use the following. GG is a standard Gaussian variable, ϕ\phi and Φ\Phi are its density and distribution function, R=ϕ/ΦR=\phi/\Phi, Rˉ(z)=z+R(z)\bar{R}(z)=z+R(z), V=RRˉV=R\bar{R}, m(k)=E(k−G)+m(k)=\mathbb{E}(k-G)_{+}, and m1(k)=E(k−G)+m_{1}(k)=\mathbb{E}(k-G)_{+}. An order parameter is a right-continuous nondecreasing γ:[0,1]→[0,1]\gamma:[0,1]\to[0,1] with γ=1\gamma=1 near one; μγ\mu_{\gamma} is the measure with distribution function γ\gamma, QγQ_{\gamma} (often written QQ) is the top of its support, δ=1−Qγ\delta=1-Q_{\gamma}, and λγ(q)=∫q1γ\lambda_{\gamma}(q)=\int_{q}^{1}\gamma; 1[q,1]\mathbf{1}_{[q,1]} is a replica-symmetric order parameter. U\mathcal{U} and Ur\mathcal{U}_{r} are the classes of order parameters, Pκ\mathcal{P}_{\kappa} is the Parisi functional, P∗(κ)\mathcal{P}_{*}(\kappa) its infimum (the Parisi value in Part I), PκRS(q)=Pκ(1[q,1])\mathcal{P}_{\kappa}^{RS}(q)=\mathcal{P}_{\kappa}(\mathbf{1}_{[q,1]}) Gardner’s replica-symmetric functional, uu the solution of the Parisi equation, XX the Parisi diffusion, DD, SS, and Hγ\mathcal{H}_{\gamma} the functions of the first-order conditions, Tm,sT_{m,s} one Gaussian step of the Parisi recursion, mjm_{j} the values of a step order parameter, K=(1−r)−1K=(1-r)^{-1} the constant of the class Ur\mathcal{U}_{r}, and b=∂xub=\partial_{x}u. The margins are the critical margin κc\kappa_{c}, Gardner’s prediction κRS\kappa_{\mathrm{RS}}, and the de Almeida–Thouless margin κd\kappa_{\mathrm{d}}; near α=2\alpha=2 we write α=2+e\alpha=2+e. The free energy is N−1log⁡ZN(U)N^{-1}\log Z_{N}(U) for a soft activation UU (Part I) and N−1log⁡VN(κ)N^{-1}\log V_{N}(\kappa) at a hard margin κ\kappa. In Part II the rescaled profile of γ\gamma is ζ(s)=γ(1−e2s)\zeta(s)=\gamma(1-e^{2}s).

The table lists the letters whose meaning differs between the two Parts. Many letters, for example τ\tau, ϑ\vartheta, ρ\rho, YY, ε\varepsilon, η\eta, and c0,c1,…c_{0}, c_{1}, \ldots, also carry local meanings that are fixed where they appear.

SymbolPart IPart II
ℓ\elllevel of the soft penalty −β(ℓ−x)2-\beta(\ell-x)^{2}; width of the top layer near α=2\alpha=2ℓ(s)\ell(s), ℓ0\ell_{0}: rescaled masses λγ/e2\lambda_{\gamma}/e^{2}
β\betaleaves of the cascade tree; strength of the soft penalty and of the source termβ=(1−q)/q\beta=\sqrt{(1-q)/q} for replica-symmetric order parameters; also δ/λ\delta/\lambda
JJJℓ(k)=Jk(ℓ−11[1−ℓ,ℓ1))J_{\ell}(k)=J^{k}(\ell^{-1}\mathbf{1}_{[1-\ell,\ell_{1})}); the level at which two replicas branchJkJ^{k}: the zero-temperature functional; J\mathcal{J}: the jamming coefficient
C0C_{0}−log⁡Φ(−1)-\log\Phi(-1); the growth constant 2L2L of an activationa growth constant for terminal data; the constant in the a priori bound s0≤C0e−2/3s_{0}\le C_{0}e^{-2/3}
UUthe activationU=UζU=U_{\zeta}, the solution in rescaled variables
AAa constant bounding an activation; the pattern matrixA(ζ)A(\zeta), the entropy term of the flat limit; a function A(k)A(k) of the margin
K,KK,\mathcal{K}∥U′′∥∞\lVert U''\rVert_{\infty}; the number of atoms of a cascade; an operator-norm constant, CopC_{\mathrm{op}}; K\mathcal{K} a cone or convex set; KNK_{N} a polyhedronKeK_{e}, a constant; K=γ/λK=\gamma/\lambda; K(d1,S)\mathcal{K}(d_{1},S), the compact class of profiles
S,sS,sS(q)=αλγ(q)2E ∂xxuγ(q,Xq)2S(q)=\alpha\lambda\gamma(q)^{2}\mathbb{E}\,\partial_{xx}u_{\gamma}(q,X_{q})^{2}; ss a sparsity levelSγS_{\gamma}, the same; the depth parameter SS in K(d1,S)\mathcal{K}(d_{1},S) and the bound S∗S_{*}; s=(1−q)/e2s=(1-q)/e^{2}, the rescaled overlap
mmmN(σ)m_{N}(\sigma), the minimal margin of σ\sigmam∗m_{*}, a bound on γ\gamma below the top of the support
ζ\zetaζl\zeta_{l}, conditional means of the response overlap; the limiting overlap lawthe rescaled profile
bbalso the rows ba=ga/Nb_{a}=g_{a}/\sqrt{N} and the top precision bb of a cascadeonly b=∂xub=\partial_{x}u

Table 1.

I Parisi formula, capacity, and isostaticity for the negative spherical perceptron

We study the spherical perceptron with Gaussian patterns at negative margins, a mean-field model of jamming in which the set of solutions is not convex. We prove the Parisi formula for its free energy for bounded-above C2C^{2} activations with bounded second derivative and either bounded first derivative or concavity, without any monotonicity assumption on the activation. For every constraint density α>2\alpha>2, we show that the normalized logarithm of the volume of solutions converges in probability to the Parisi value at every margin below a critical margin κc(α)<0\kappa_{\mathrm{c}}(\alpha)<0, that this value tends to −∞-\infty continuously as the margin increases to κc(α)\kappa_{\mathrm{c}}(\alpha), and that the largest achievable margin converges in probability to κc(α)\kappa_{\mathrm{c}}(\alpha). At sufficiently negative margins, Gardner’s replica-symmetric formula for the free energy is exact. For α\alpha slightly above two, the limit κc(α)\kappa_{\mathrm{c}}(\alpha) lies below Gardner’s prediction by at least a constant times (α−2)3(\alpha-2)^{3}; with Part II, this order is exact. At the optimal margin, with probability exponentially close to one, every optimal configuration has exactly NN contacts: the jamming point is isostatic. The empirical gap measures of optimal configurations, normalized by NN, converge weakly in probability to a deterministic measure with an atom of mass exactly one at zero. At a fixed margin below κc(α)\kappa_{\mathrm{c}}(\alpha), the gap statistics of all but an exponentially small fraction of the feasible set are given by the Parisi diffusion.

Introduction

The perceptron is the simplest model of a neural network and a basic example of a random constraint satisfaction problem. Given MM random patterns g1,…,gM∈RNg_{1},\ldots,g_{M}\in\mathbb{R}^{N} and a margin κ∈R\kappa\in\mathbb{R}, the spherical perceptron asks for a vector σ\sigma on the sphere of radius N\sqrt{N} such that

ga⋅σN≥κfor every a≤M.\frac{g_{a}\cdot\sigma}{\sqrt{N}}\ge\kappa\qquad\text{for every }a\le M.

Two questions are classical: how large the set of such vectors is, and what the largest achievable margin is when M/N→αM/N\to\alpha. Gardner and Derrida computed the answers with the replica method, under the assumption of replica symmetry [ref-62, ref-63]. In particular, Gardner’s formula predicts that the largest achievable margin converges to the solution κRS(α)\kappa_{\mathrm{RS}}(\alpha) of

α E(κRS−G)+2=1,\alpha\,\mathbb{E}\left(\kappa_{\mathrm{RS}}-G\right)_{+}^{2}=1,

where GG is a standard Gaussian variable.

For κ≥0\kappa\ge0 each constraint cuts out a geodesically convex spherical cap, and replica symmetry is expected to hold throughout the satisfiable phase. In this regime Gardner’s formula for the free energy was proved by [ref-115] and by [ref-128], Chapter 8, and Stojnic gave another proof of the formula for the capacity [ref-118]. For κ<0\kappa<0, which is the relevant regime when α>2\alpha>2, each constraint removes a cap smaller than a hemisphere, and the feasible set is a nonconvex intersection of the complements of such caps. Franz and Parisi proposed this negative spherical perceptron as the simplest model of jamming [ref-58]. Franz, Parisi, Sevelev, Urbani, and Zamponi argued that its satisfiability threshold belongs to the same universality class as the jamming transition of hard spheres in infinite dimension [ref-59]. The same model describes linear classification with a negative margin, a basic example of an overparametrized learning problem [ref-60, ref-90].

In the physical picture, replica symmetry is lost at negative margins before the satisfiability threshold is reached. The stability of the replica-symmetric solution of the perceptron was studied in [ref-23, ref-63]. Close to κ=0\kappa=0 replica symmetry is lost at a de Almeida–Thouless line [ref-45], through a continuous transition to full replica symmetry breaking in the sense of Parisi [ref-104, ref-105], and the whole jamming line lies in the broken phase [ref-58, ref-59]. The physics literature therefore predicts that Gardner’s formula for the critical margin is exact only for κ≥0\kappa\ge0. At the critical margin, jammed configurations are predicted to be isostatic: they are in contact with as many constraints as there are degrees of freedom. Franz and Parisi predicted that jamming is hypostatic in the convex regime κ>0\kappa> 0 and isostatic in the nonconvex regime κ<0\kappa< 0 [ref-58]; for the isostatic counting in jammed packings see [ref-2, ref-91]. The gaps and contact forces of jammed configurations are predicted to have power-law distributions with universal exponents [ref-29, ref-30, ref-58, ref-59, ref-107], related to marginal stability [ref-106, ref-132] and to the jamming of soft spheres [ref-95]. Annesi, Malatesta, and Zamponi computed the satisfiability threshold numerically from the equations of full replica symmetry breaking, and located the region of the phase diagram in which the overlap distribution of typical solutions has connected support [ref-3].

Rigorous results for negative margins have so far been partial. In [ref-118], Stojnic used Gordon’s comparison inequality [ref-66] to show that Gardner’s formula gives an upper bound on the capacity at every margin. He also proved a sharper bound by applying Gordon’s inequality to an exponential of the cost; its improvement over the replica-symmetric bound was shown numerically [ref-117]. El Alaoui and Sellke conjectured that the Parisi variational problem for the hard-margin perceptron, in the form used below, gives the free energy and the capacity [ref-54], Conjecture 1; they trace this conjecture to [ref-59]. Assuming that some Parisi minimizer γ\gamma has no overlap gap, and given oracle access to it, their algorithm constructs vectors of norm close to QγN\sqrt{Q_{\gamma}N} with a small vector of constraint violations, where Qγ<1Q_{\gamma} < 1 is the top of the support of γ\gamma [ref-54], Theorem 1; see also [ref-55]. Montanari, Zhong, and Zhou proved upper and lower bounds on the capacity that agree to leading order as κ→−∞\kappa\to-\infty, and studied its algorithmic tractability [ref-90]. Their upper bound [ref-90], Theorem 4.2 shows that Gardner’s prediction overestimates the capacity when α\alpha is large. Near α=2\alpha= 2, where κRS(α)\kappa_{\mathrm{RS}}(\alpha) is close to zero, the cited bounds do not quantify this separation. Stojnic later proposed a characterization of the capacity through fully lifted random duality theory, and computed the resulting values numerically at finite levels of a stationarized lifting [ref-119]. Huang, Sellke, and Sun characterized the empirical distributions of the margins that algorithms with dimension-free Lipschitz dependence on the Gaussian data can achieve [ref-73].

For the Ising perceptron, where σ∈{−1,1}N\sigma\in\{-1,1\}^{N}, the capacity predicted by Krauth and Mézard [ref-83] was studied by Ding and Sun and by Huang, who proved lower and upper bounds matching the prediction, subject in part to numerical conditions [ref-48, ref-72]. Xu proved a sharp threshold for the Ising perceptron with Bernoulli disorder [ref-133]. At small constraint densities, Talagrand proved that the free energy of the Ising perceptron is given by the replica-symmetric formula, for a class of activations that includes the one-sided threshold [ref-123], [ref-127], Chapter 2, [ref-128], Chapter 9; his treatment of the Shcherbina–Tirozzi model, whose spins are continuous, is in [ref-127], Chapter 3. Bolthausen, Nakajima, Sun, and Xu gave another proof of this replica-symmetric formula, for a larger class of activations [ref-20]. In the Bayes-optimal (teacher–student) setting, replica-symmetric formulas for perceptron-type generalized linear models were established by Barbier, Krzakala, Macris, Miolane, and Zdeborová [ref-13].

The Parisi formula for mean-field spin glasses was predicted in [ref-104, ref-105] (see also [ref-31, ref-88]) and proved by Guerra and Talagrand for the Sherrington–Kirkpatrick model [ref-68, ref-124]. Talagrand and W.-K. Chen proved the Crisanti–Sommers formula [ref-44] for spherical models, the latter through the Aizenman–Sims–Starr scheme [ref-1, ref-37, ref-125], and Panchenko proved the formula for general mixed pp-spin models, building on the ultrametricity theorem [ref-97, ref-98, ref-99]. The perceptron differs from these models in a basic way. Its Hamiltonian ∑aU(ga⋅σ/N)\sum_{a} U(g_{a}\cdot\sigma/\sqrt{N}) is a nonlinear function of Gaussian fields, and it is not itself a Gaussian process in σ\sigma. Structurally the model is closer to multi-species or bipartite spin glasses, with one species formed by the coordinates and the other by the constraints. For such models Guerra’s interpolation gives an upper bound when the interaction is convex [ref-14, ref-16, ref-100]. The matching lower bound by the Aizenman–Sims–Starr scheme requires synchronization of the overlaps of the species [ref-100, ref-101], and it needs no convexity, for Ising spins [ref-100] and for spherical spins [ref-16]; see also [ref-15, ref-17]. The interaction of the bipartite model, which the perceptron resembles, is not convex. For spherical bipartite models see [ref-6, ref-82, ref-122]. Mourrat proved that the solution of an infinite-dimensional Hamilton–Jacobi equation bounds the free energy of bipartite models from above, and conjectured that it is the limit [ref-92]; see also [ref-49]. Our upper bound follows the proof of the upper bound in [ref-92]. That work also contains a quantitative synchronization estimate for pairs of overlap arrays, which is an input to both bounds below. Mourrat later proved an upper bound of the same kind for vector spin glasses whose energy is a Gaussian field [ref-94]; the energy ∑aU(ga⋅σ/N)\sum_{a} U(g_{a}\cdot\sigma/\sqrt{N}) is not of this form. For nonconvex vector models, H.-B. Chen and Mourrat showed that the limit of the free energy, if it exists, is a critical value of a Parisi-type functional [ref-34], and H.-B. Chen proved the same for nonconvex multi-species models [ref-33]. H.-B. Chen, Issa, and Mourrat identified the free energy of nonconvex multi-species models with centered Ising spins [ref-36]. H.-B. Chen and Mourrat then identified the limit of the free energy of all multi-species spherical spin glasses, with no convexity assumption [ref-35]. Ho proved that the enriched free energy of vector spin glasses converges without convexity, so the critical-point representation of H.-B. Chen and Mourrat holds without assuming convergence [ref-69].

Our contribution is a Parisi formula for the free energy, an identification of the limit of the largest achievable margin, and a description of the gap statistics and contacts of optimal configurations for the negative spherical perceptron. We prove the Parisi formula for smooth activations, including activations that are neither monotone nor concave (Theorem 1.1). For α>2\alpha>2 we identify the exponential volume of the feasible set at every margin below the critical margin κc(α)\kappa_{\mathrm{c}}(\alpha) of the Parisi functional, and we prove that the largest achievable margin converges to κc(α)\kappa_{\mathrm{c}}(\alpha) (Theorem 1.2). For α\alpha slightly above two, we show that κc(α)\kappa_{\mathrm{c}}(\alpha) lies below Gardner’s prediction by at least a constant times (α−2)3(\alpha-2)^{3} (Theorem 1.4). Combined with the matching upper bound on κRS−κc\kappa_{\mathrm{RS}}-\kappa_{\mathrm{c}} (Part II, Theorem 1.1), this order is exact (Corollary 1.6).

Our main result on the geometry of solutions concerns the configurations of maximal margin. With probability exponentially close to one, every optimal configuration has exactly NN contacts. All configurations whose margin is close to optimal share a deterministic limiting distribution of gaps, and this distribution has an atom of mass exactly one at zero (Theorem 1.8). Thus, in the limit, no macroscopic set of constraints beyond the NN contacts accumulates near contact. This is a rigorous form of the predicted isostaticity of the jamming point recalled above. The link between the random model and the limiting gap law is a collapse of Parisi minimizers. As the margin increases to κc(α)\kappa_{c}(\alpha), the minimal value of the Parisi functional tends to −∞-\infty continuously, and the mass of every minimizer (defined in §1.1) tends to zero at a power rate (Theorem 1.3). This is a rigorous counterpart of the physical picture of jamming, in which the overlap of solutions tends to one and the weight of the overlaps below any fixed level tends to zero [ref-58].

Model and the Parisi functional

Fix α>0\alpha>0 and integers M=MNM=M_{N} with MN/N→αM_{N}/N\to\alpha. Let G=(gai)a≤M,i≤N\mathbf{G}=(g_{ai})_{a\leq M,i\leq N} have independent standard Gaussian entries, write gag_{a} for its aa-th row, and put ba=ga/Nb_{a}=g_{a}/\sqrt{N}. Let μN\mu_{N} be the uniform probability measure on

SN={σ∈RN:∥σ∥2=N}.S_{N}=\left\{\sigma\in\mathbb{R}^{N}:\lVert\sigma\rVert^{2}=N\right\}.

For σ∈SN\sigma\in S_{N}, the number ba⋅σb_{a}\cdot\sigma is the margin of the constraint aa at σ\sigma; for each fixed σ∈SN\sigma\in S_{N} it is a standard Gaussian variable. For an activation U:R→[−∞,∞)U:\mathbb{R}\to[-\infty,\infty) that is bounded above, define

ZN(U)=∫SNexp⁡{∑a=1MU(ba⋅σ)}μN(dσ),fN(U)=1Nlog⁡ZN(U),FN(U)=EfN(U).(1)\begin{aligned} Z_{N}(U)&=\int_{S_{N}}\exp\left\{\sum_{a=1}^{M}U(b_{a}\cdot\sigma)\right\}\mu_{N}(\mathrm{d}\sigma),\\ f_{N}(U)&=\frac{1}{N}\log Z_{N}(U),\qquad F_{N}(U)=\mathbb{E}f_{N}(U). \tag*{(1)} \end{aligned}

For a margin κ∈R\kappa\in\mathbb{R}, the feasible set, its volume, and the largest achievable margin are

SN(κ)={σ∈SN:ba⋅σ≥κ for all a≤M},VN(κ)=μN(SN(κ)),κN=max⁡σ∈SNmin⁡a≤Mba⋅σ.(2)\begin{aligned} S_{N}(\kappa)&=\left\{\sigma\in S_{N}:b_{a}\cdot\sigma\geq\kappa\ \text{for all }a\leq M\right\},\\ V_{N}(\kappa)&=\mu_{N}(S_{N}(\kappa)),\qquad \kappa_{N}=\max_{\sigma\in S_{N}}\min_{a\leq M}b_{a}\cdot\sigma. \tag*{(2)} \end{aligned}

We use the convention log⁡0=−∞\log0=-\infty. Thus VN(κ)=ZN(Hκ)V_{N}(\kappa)=Z_{N}(H_{\kappa}) for the hard constraint Hκ(x)=log⁡1{x≥κ}H_{\kappa}(x)=\log\mathbf{1}_{\{x\geq\kappa\}}. For σ∈SN\sigma\in S_{N} let mN(σ)=min⁡a≤Mba⋅σm_{N}(\sigma)=\min_{a\leq M}b_{a}\cdot\sigma, so that κN=max⁡SNmN\kappa_{N}=\max_{S_{N}}m_{N}. The gaps of σ\sigma are the numbers ba⋅σ−mN(σ)≥0b_{a}\cdot\sigma-m_{N}(\sigma)\geq0, and the empirical gap measure of σ\sigma is

νN,σ=1N∑a=1Mδba⋅σ−mN(σ),\nu_{N,\sigma}=\frac{1}{N}\sum_{a=1}^{M}\delta_{b_{a}\cdot\sigma-m_{N}(\sigma)},

a measure on [0,∞)[0,\infty) of total mass M/NM/N. A constraint with gap zero is called a contact of σ\sigma.

An order parameter is a nondecreasing, right-continuous function γ:[0,1]→[0,1]\gamma:[0,1]\to[0,1] that is identically one on [qˉ,1][\bar q,1] for some qˉ<1\bar q<1. We write U\mathcal U for the set of order parameters and

Qγ=inf⁡{q∈[0,1]:γ(q)=1}<1,λγ(q)=∫q1γ(t) dt.Q_{\gamma}=\inf\{q\in[0,1]:\gamma(q)=1\}<1,\qquad\lambda_{\gamma}(q)=\int_{q}^{1}\gamma(t)\,\mathrm{d}t.

Thus γ\gamma is the distribution function of a probability measure μγ\mu_{\gamma} on [0,Qγ][0,Q_{\gamma}], and QγQ_{\gamma} is the top of its support. We call λγ(0)\lambda_{\gamma}(0) the mass of γ\gamma. Given γ∈U\gamma\in\mathcal U and an activation UU, let uγ=uγ(q,x;U)u_{\gamma}=u_{\gamma}(q,x;U) solve the Parisi equation

∂quγ+12∂xxuγ+γ(q)2(∂xuγ)2=0,uγ(1,x;U)=U(x).(3)\partial_{q}u_{\gamma}+\frac{1}{2}\partial_{xx}u_{\gamma}+\frac{\gamma(q)}{2}(\partial_{x}u_{\gamma})^{2}=0,\qquad u_{\gamma}(1,x;U)=U(x). \tag*{(3)}

When γ\gamma is a step function, (3) is solved by an explicit backward recursion of Gaussian integrals, the Cole–Hopf transformation [ref-42, ref-71] on each step; for general γ\gamma, and for the terminal data used in this part, the solution is defined by approximation. Both are recalled in Chapter 2. For the Sherrington–Kirkpatrick model these constructions are in [ref-52, ref-67, ref-68, ref-105] and [ref-76]. The Crisanti–Sommers entropy [ref-44, ref-125] and the Parisi functional are

CS⁡(γ)=12∫0qˉdqλγ(q)+12log⁡(1−qˉ),PU(γ)=αuγ(0,0;U)+CS⁡(γ),P∗(U)=inf⁡γ∈UPU(γ),(4)\begin{aligned} \operatorname{CS}(\gamma)&=\frac{1}{2}\int_{0}^{\bar q}\frac{\mathrm{d}q}{\lambda_{\gamma}(q)}+\frac{1}{2}\log(1-\bar q),\\ \mathcal P_{U}(\gamma)&=\alpha u_{\gamma}(0,0;U)+\operatorname{CS}(\gamma),\qquad \mathcal P_{*}(U)=\inf_{\gamma\in\mathcal U}\mathcal P_{U}(\gamma), \tag*{(4)} \end{aligned}

where qˉ\bar q is any number in [Qγ,1)[Q_{\gamma},1). The value of CS⁡(γ)\operatorname{CS}(\gamma) does not depend on this choice, since λγ(q)=1−q\lambda_{\gamma}(q)=1-q on [Qγ,1][Q_{\gamma},1]. The first term in PU(γ)\mathcal P_{U}(\gamma) accounts for the constraints and the second for the spherical entropy.

For the hard constraint, let ϕ\phi and Φ\Phi denote the standard Gaussian density and distribution function. For γ∈U\gamma\in\mathcal U, qˉ∈[Qγ,1)\bar q\in[Q_{\gamma},1) and κ∈R\kappa\in\mathbb R, let uγ(q,x;κ)u_{\gamma}(q,x;\kappa) solve (3) on [0,qˉ]×R[0,\bar q]\times\mathbb R with the terminal datum

uγ(qˉ,x;κ)=log⁡P{x+1−qˉ G≥κ}=log⁡Φ(x−κ1−qˉ),(5)u_{\gamma}(\bar q,x;\kappa)=\log\mathbb P\{x+\sqrt{1-\bar q}\,G\geq\kappa\} =\log\Phi\left(\frac{x-\kappa}{\sqrt{1-\bar q}}\right), \tag*{(5)}

where GG is a standard Gaussian variable, and set

Pκ(γ)=αuγ(0,0;κ)+CS⁡(γ),P∗(κ)=inf⁡γ∈UPκ(γ),κc(α)=sup⁡{κ∈R:P∗(κ)>−∞}.(6)\begin{aligned} \mathcal P_{\kappa}(\gamma)&=\alpha u_{\gamma}(0,0;\kappa)+\operatorname{CS}(\gamma),\qquad \mathcal P_{*}(\kappa)=\inf_{\gamma\in\mathcal U}\mathcal P_{\kappa}(\gamma),\\ \kappa_{c}(\alpha)&=\sup\{\kappa\in\mathbb R:\mathcal P_{*}(\kappa)>-\infty\}. \tag*{(6)} \end{aligned}

Again the value does not depend on qˉ∈[Qγ,1)\bar q\in[Q_{\gamma},1). The terminal datum (5) is what one obtains by running the equation with U=HκU=H_{\kappa} from time 1 to time qˉ\bar q with γ=1\gamma=1, which avoids a singular terminal problem. This variational problem appears in the physics literature [ref-58, ref-59], and it was stated in the form (5)–(6) by El Alaoui and Sellke [ref-54]. For the terminal datum (5), they constructed a classical solution of (3), with bounds on its derivatives, and proved that its derivatives depend continuously on γ\gamma [ref-54]. We call P∗(κ)\mathcal{P}_{*}(\kappa) the Parisi value and κc(α)\kappa_{\mathrm{c}}(\alpha) the critical margin.

The replica-symmetric order parameters 1[q,1]\mathbf{1}_{[q,1]}, q∈[0,1)q\in[0,1), give Gardner’s replica-symmetric functional:

PκRS(q)=Pκ ⁣(1[q,1])=αElog⁡Φ ⁣(qG−κ1−q)+q2(1−q)+12log⁡(1−q).(7)\mathcal{P}_{\kappa}^{\mathrm{RS}}(q) = \mathcal{P}_{\kappa}\!\left(\mathbf{1}_{[q,1]}\right) = \alpha\mathbb{E}\log\Phi\!\left(\frac{\sqrt{q}G-\kappa}{\sqrt{1-q}}\right) + \frac{q}{2(1-q)} + \frac{1}{2}\log(1-q). \tag*{(7)}

Recall that κRS(α)\kappa_{\mathrm{RS}}(\alpha) is the unique solution of αE(κRS−G)+2=1\alpha\mathbb{E}(\kappa_{\mathrm{RS}}-G)_{+}^{2}=1. It is Gardner’s prediction for the limit of κN\kappa_{N} [ref-62], and κRS(α)<0\kappa_{\mathrm{RS}}(\alpha)<0 if and only if α>2\alpha>2, since EG−2=12\mathbb{E}G_{-}^{2}=\frac{1}{2}. The value α=2\alpha=2 at κ=0\kappa=0 is Cover’s capacity of the perceptron [ref-43].

Finally, for γ∈U\gamma\in\mathcal{U} and κ∈R\kappa\in\mathbb{R} we define the Parisi diffusion X=Xγ,κX=X^{\gamma,\kappa} as follows. On [0,Qγ][0,Q_{\gamma}] it is the unique strong solution of

dXq=γ(q) ∂xuγ(q,Xq;κ) dq+dBq,X0=0,(8)\mathrm{d}X_{q} = \gamma(q)\,\partial_{x}u_{\gamma}(q,X_{q};\kappa)\,\mathrm{d}q + \mathrm{d}B_{q}, \qquad X_{0}=0, \tag*{(8)}

for a standard Brownian motion BB. Conditionally on (Xq)q≤Qγ(X_{q})_{q\leq Q_{\gamma}}, the endpoint X1X_{1} has the law of XQγ+1−QγGX_{Q_{\gamma}}+\sqrt{1-Q_{\gamma}}G conditioned on the event {XQγ+1−QγG≥κ}\{X_{Q_{\gamma}}+\sqrt{1-Q_{\gamma}}G\geq\kappa\}, where GG is a standard Gaussian variable independent of BB. The equation (8) is well posed by Lemmas 2.6 and 2.7. For mixed pp-spin models this diffusion appears in the work of Auffinger and W.-K. Chen [ref-7], and for the terminal datum (5) it appears in [ref-54]. For step order parameters, XX observed at the breakpoints of γ\gamma and at time 1 is a Markov chain with explicit tilted Gaussian transitions (Lemma 9.3). The variable X1−κ≥0X_{1}-\kappa\geq0 plays the role of the gap of a typical constraint.

Main results

Our first result is the Parisi formula for smooth activations. We consider two classes of activations:

(A) U∈C2(R)U\in C^{2}(\mathbb{R}), sup⁡U<∞\sup U<\infty, and ∥U′∥∞+∥U′′∥∞<∞\lVert U'\rVert_{\infty}+\lVert U''\rVert_{\infty}<\infty;

(B) U∈C2(R)U\in C^{2}(\mathbb{R}) is concave, sup⁡U<∞\sup U<\infty, and ∥U′′∥∞<∞\lVert U''\rVert_{\infty}<\infty.

No sign condition is imposed on U′U' in class (A), and the activations in class (B) may tend to −∞-\infty quadratically.

Theorem 1.1. Let α>0\alpha>0 and let UU belong to class (A) or class (B). Then

lim⁡N→∞FN(U)=P∗(U),fN(U)⟶P∗(U)in probability.\lim_{N\to\infty}F_{N}(U)=\mathcal{P}_{*}(U), \qquad f_{N}(U)\longrightarrow\mathcal{P}_{*}(U) \quad\text{in probability}.

Moreover, the infimum defining P∗(U)\mathcal{P}_{*}(U) is attained, and every minimizer γ\gamma satisfies Qγ≤rUQ_{\gamma}\leq r_{U}, where rU<1r_{U}<1 is an explicit number depending only on α\alpha and UU, defined in (5.1). If UU belongs to class (A), then rU≤α∥U′∥∞2/(1+α∥U′∥∞2)r_{U}\leq\alpha\lVert U'\rVert_{\infty}^{2}/(1+\alpha\lVert U'\rVert_{\infty}^{2}).

Class (A) contains activations that are neither monotone nor concave. This generality is used for the gaps: adding a bounded source term to a smooth approximation of the hard wall destroys monotonicity and concavity. Class (B) contains smooth concave penalties comparable to −β(ℓ−x)+-\beta(\ell-x)_{+}, which we use to pass to hard constraints. The bound on QγQ_{\gamma} says that the minimizers for a smooth activation stay a fixed distance away from overlap one. For the hard wall this fails as the margin approaches the critical margin; see Theorem 1.3(iv).

For α>2\alpha>2 the critical margin is negative, and the feasible sets SN(κ)S_{N}(\kappa) with κ<0\kappa<0 are not convex. Our second result identifies their exponential volume and the largest achievable margin.

Theorem 1.2. Let α>2\alpha>2. Then −∞<κc(α)<0-\infty<\kappa_{\mathrm{c}}(\alpha)<0, and the following statements hold.

(i) For every κ<κc(α)\kappa<\kappa_{\mathrm{c}}(\alpha), N−1log⁡VN(κ)→P∗(κ)N^{-1}\log V_{N}(\kappa)\to\mathcal{P}_{*}(\kappa) in probability.

(ii) κN→κc(α)\kappa_{N}\to\kappa_{\mathrm{c}}(\alpha) in probability.

(iii) For every κ≥κc(α)\kappa\geq\kappa_{\mathrm{c}}(\alpha) and every C<∞C<\infty, P{N−1log⁡VN(κ)>−C}→0\mathbb{P}\{N^{-1}\log V_{N}(\kappa)>-C\}\to0.

Part (iii) states that the normalized log volume tends to −∞-\infty at the critical margin itself; it does not assert that SN(κc)S_{N}(\kappa_{\mathrm{c}}) is empty. The assumption α>2\alpha>2 enters only through the inequality κc(α)<0\kappa_{\mathrm{c}}(\alpha)<0; for every α>0\alpha>0, part (i) holds at the margins κ<min⁡{κc(α),0}\kappa<\min\{\kappa_{\mathrm{c}}(\alpha),0\} (see the end of §8.4). For α>2\alpha>2, parts (i) and (iii) confirm the free-energy part of a conjecture of El Alaoui and Sellke [ref-54] (Conjecture 1), and part (ii) confirms its capacity part in the form κN→κc(α)\kappa_{N}\to\kappa_{\mathrm{c}}(\alpha).

The proofs of Theorem 1.2 and of the gap theorems below rest on the following properties of the variational problem, which hold for every α>0\alpha>0. For κ∈R\kappa\in\mathbb{R} put Cκ=12(κ+κ2+4)C_{\kappa}=\frac{1}{2}(\kappa+\sqrt{\kappa^{2}+4}), and let C0=−log⁡Φ(−1)C_{0}=-\log\Phi(-1).

Theorem 1.3. Let α>0\alpha>0.

(i) If P∗(κ)>−∞\mathcal{P}_{*}(\kappa)>-\infty, then the infimum defining P∗(κ)\mathcal{P}_{*}(\kappa) is attained, every minimizer γ\gamma satisfies 1−Qγ>e2P∗(κ)1-Q_{\gamma}>e^{2\mathcal{P}_{*}(\kappa)}, and the set of minimizers is compact in L1([0,1])L^{1}([0,1]).

(ii) −∞<κc(α)≤κRS(α)<∞-\infty<\kappa_{\mathrm{c}}(\alpha)\leq\kappa_{\mathrm{RS}}(\alpha)<\infty. If α>2\alpha>2, then κc(α)<0\kappa_{\mathrm{c}}(\alpha)<0.

(iii) P∗\mathcal{P}_{*} is finite, nonincreasing, concave, and continuous on (−∞,κc)(-\infty,\kappa_{\mathrm{c}}), and P∗(κ)=−∞\mathcal{P}_{*}(\kappa)=-\infty for κ≥κc\kappa\geq\kappa_{\mathrm{c}}. More precisely, with C∗=CκcC_{*}=C_{\kappa_{\mathrm{c}}}, for all sufficiently small η>0\eta>0,

P∗(κc−η)≤12log⁡η+12log⁡(2αC∗)+12.\mathcal{P}_{*}(\kappa_{\mathrm{c}}-\eta)\leq\frac{1}{2}\log\eta+\frac{1}{2}\log(2\alpha C_{*})+\frac{1}{2}.

(iv) If P∗(κ)>−∞\mathcal{P}_{*}(\kappa)>-\infty and γ\gamma is a minimizer, then

P∗(κ)≥1+α2log⁡(1−Qγ)−α(C0+log⁡(1+Cκ)).\mathcal{P}_{*}(\kappa)\geq\frac{1+\alpha}{2}\log(1-Q_{\gamma})-\alpha\left(C_{0}+\log(1+C_{\kappa})\right).

Consequently there are C,η0>0C,\eta_{0}>0, depending only on α\alpha, such that for 0<η≤η00<\eta\leq\eta_{0} every minimizer γ\gamma of Pκc−η\mathcal{P}_{\kappa_{\mathrm{c}}-\eta} satisfies

1−Qγ≤Cη1/(1+α),λγ(0)≤Cη1/(2+2α).1-Q_{\gamma}\leq C\eta^{1/(1+\alpha)},\qquad\lambda_{\gamma}(0)\leq C\eta^{1/(2+2\alpha)}.

(v) There is κ0(α)<κc(α)\kappa_{0}(\alpha)<\kappa_{\mathrm{c}}(\alpha) such that, for every κ≤κ0(α)\kappa\leq\kappa_{0}(\alpha), every minimizer of Pκ\mathcal{P}_{\kappa} equals 1[q,1]\mathbf{1}_{[q,1]} for some q∈(0,1)q\in(0,1). In particular P∗(κ)=min⁡q∈[0,1)PκRS(q)\mathcal{P}_{*}(\kappa)=\min_{q\in[0,1)}\mathcal{P}_{\kappa}^{\mathrm{RS}}(q). If α≤1\alpha\leq1, the same holds for every κ<κc(α)\kappa<\kappa_{\mathrm{c}}(\alpha).

Part (iii) says that the jamming transition is continuous on the exponential scale: the Parisi value tends to −∞-\infty as the margin increases to κc\kappa_{\mathrm{c}}, and it does so at least logarithmically fast. Part (iv) shows that a very negative value pins the top of the support of every minimizer close to one. Together, (iii) and (iv) give the collapse of minimizers at the critical margin: their tops tend to one and their masses tend to zero, at power rates. This collapse is the input for the gap theorems. By part (v) and Theorem 1.2(i), for α>2\alpha>2 and κ≤κ0(α)\kappa\leq\kappa_{0}(\alpha),

1Nlog⁡VN(κ)⟶min⁡q∈[0,1)PκRS(q)in probability,\frac{1}{N}\log V_{N}(\kappa)\longrightarrow\min_{q\in[0,1)}\mathcal{P}_{\kappa}^{\mathrm{RS}}(q)\quad\text{in probability},

which is Gardner’s replica-symmetric formula for the free energy [ref-62]. Thus Gardner’s formula holds at sufficiently negative margins; the margin κ0(α)\kappa_{0}(\alpha) is not explicit. Theorem 1.3 is proved in §6.6, and part (v) is Corollary 6.9. Two structural facts behind it hold for every order parameter: an exact value identity (Lemma 6.2) and the universal slope bound 0<λγ(0) ∂xuγ(0,0;κ)<Cκ0<\lambda_{\gamma}(0)\,\partial_{x}u_{\gamma}(0,0;\kappa)<C_{\kappa} (Lemma 6.4).

Our next result concerns constraint densities slightly above two. There κRS(α)\kappa_{\mathrm{RS}}(\alpha) is close to zero, and Gardner’s prediction is closest to being correct. We write α=2+e\alpha=2+e.

Theorem 1.4. There are numerical constants e0>0e_{0}>0 and c>0c>0 such that, for 0<e<e00<e<e_{0},

κc(2+e)≤κRS(2+e)−ce3.(9)\kappa_{\mathrm{c}}(2+e)\leq\kappa_{\mathrm{RS}}(2+e)-ce^{3}. \tag*{(9)}

Theorem 1.4 is a statement about the variational problem alone, and it is proved in Chapter 7. Combined with Theorem 1.2(ii), it gives the following.

Corollary 1.5. Let e0e_{0} and cc be as in Theorem 1.4, let 0<e<e00<e<e_{0}, and put α=2+e\alpha=2+e. Then κN→κc(α)\kappa_{N}\to\kappa_{\mathrm{c}}(\alpha) in probability, and κc(α)≤κRS(α)−ce3\kappa_{\mathrm{c}}(\alpha)\leq\kappa_{\mathrm{RS}}(\alpha)-ce^{3}. In particular,

lim⁡N→∞P{κN≥κRS(α)−c2e3}=0.\lim_{N\to\infty}\mathbb{P}\left\{\kappa_{N}\geq\kappa_{\mathrm{RS}}(\alpha)-\frac{c}{2}e^{3}\right\}=0.

For large α\alpha, comparing the asymptotics as κ→−∞\kappa\to-\infty of the upper bound of Montanari, Zhong, and Zhou, recalled above, with the relation αE(κRS−G)+2=1\alpha\mathbb{E}(\kappa_{\mathrm{RS}}-G)_{+}^{2}=1 shows that, for all sufficiently large α\alpha, with probability tending to one no configuration has margin at least κRS(α)−1/∣κRS(α)∣\kappa_{\mathrm{RS}}(\alpha)-1/|\kappa_{\mathrm{RS}}(\alpha)|. Together with Theorem 1.2(ii), this gives κc(α)<κRS(α)\kappa_{\mathrm{c}}(\alpha)<\kappa_{\mathrm{RS}}(\alpha) for all sufficiently large α\alpha. Corollary 1.5 shows that the failure of Gardner’s formula extends to α\alpha close to two, where κRS(2+e)=−2π8e+O(e2)\kappa_{\mathrm{RS}}(2+e)=-\frac{\sqrt{2\pi}}{8}e+O(e^{2}) is itself close to zero (see eq:7.3). The proof of Theorem 1.4 evaluates the functional at an explicit order parameter with two support points near one.

In Part II we study the minimizers of Pκ\mathcal{P}_{\kappa} for α\alpha slightly above two. We show that the replica-symmetric solution loses its stability at a de Almeida–Thouless margin κd(α)<κc(α)\kappa_{\mathrm{d}}(\alpha)<\kappa_{\mathrm{c}}(\alpha), and that immediately beyond κd(α)\kappa_{\mathrm{d}}(\alpha) every minimizer exhibits full replica symmetry breaking: its Parisi measure is supported on a nondegenerate interval, with a positive continuous density in the interior and atoms at both endpoints (Part II, Theorem 1.2). We also determine the distances between the three margins κd<κc<κRS\kappa_{\mathrm{d}}<\kappa_{\mathrm{c}}<\kappa_{\mathrm{RS}} (Part II, Theorem 1.1). We now restate the results of that part that we use.

Let R=ϕ/ΦR=\phi/\Phi and V(z)=R(z)(z+R(z))V(z)=R(z)(z+R(z)), so that 0<V<10<V<1, and put

c0=∫RV(z)(1−V(z)) dz,θd=(2π)3/2128c02.c_{0}=\int_{\mathbb{R}}V(z)(1-V(z))\,\mathrm{d}z,\qquad \theta_{\mathrm{d}}=\frac{(2\pi)^{3/2}}{128c_{0}^{2}}.

By Part II, Lemma 3.5 and Lemma 3.3(c), for e>0e>0 sufficiently small and every κ<κRS(α)\kappa<\kappa_{\mathrm{RS}}(\alpha) the function PκRS\mathcal{P}_{\kappa}^{\mathrm{RS}} of (7) has a unique minimizer q∗(κ)q^{*}(\kappa) on [0,1)[0,1), and q∗(κ)∈(0,1)q^{*}(\kappa)\in(0,1). By Part II, Proposition 3.8, there is also a numerical constant c2>0c_{2}>0 such that exactly one margin κ∈(κRS(α)−c2,κRS(α))\kappa\in(\kappa_{\mathrm{RS}}(\alpha)-c_{2},\kappa_{\mathrm{RS}}(\alpha)) satisfies

αEV(Z∗)2=1,Z∗=q∗G−κ1−q∗,q∗=q∗(κ).\alpha\mathbb{E}V(Z^{*})^{2}=1,\qquad Z^{*}=\frac{\sqrt{q^{*}}G-\kappa}{\sqrt{1-q^{*}}},\qquad q^{*}=q^{*}(\kappa).

This margin is the de Almeida–Thouless margin κd(α)\kappa_{\mathrm{d}}(\alpha). The displayed condition says that the replicon eigenvalue of the replica-symmetric solution vanishes (see [ref-23, ref-59]); on (κd,κRS)(\kappa_{\mathrm{d}},\kappa_{\mathrm{RS}}) this eigenvalue is negative, and the replica-symmetric solution is unstable (Part II, Proposition 3.8). The ordering of thresholds (Part II, Theorem 1.1) states that there are numerical constants e0,c,C>0e_{0},c,C>0 such that, for 0<e<e00<e<e_{0} and α=2+e\alpha=2+e,

κRS(α)−κd(α)=θde2+O(e3log⁡(1/e)),ce3≤κRS(α)−κc(α)≤Ce3.(10)\kappa_{\mathrm{RS}}(\alpha)-\kappa_{\mathrm{d}}(\alpha) =\theta_{\mathrm{d}}e^{2}+O(e^{3}\log(1/e)),\qquad ce^{3}\leq\kappa_{\mathrm{RS}}(\alpha)-\kappa_{\mathrm{c}}(\alpha)\leq Ce^{3}. \tag*{(10)}

The lower bound on κRS−κc\kappa_{\mathrm{RS}}-\kappa_{\mathrm{c}} is Theorem 1.4, which Part II proves again through the zero-temperature functional (Part II, Theorem 3.13). The upper bound rests on a dual certificate for this zero-temperature version of the variational problem (Part II, Theorems 3.10 and 3.11). Zero-temperature Parisi functionals were introduced for the ground state energy of mixed pp-spin models by Auffinger and W.-K. Chen [ref-9], and for spherical models by [ref-40] and by [ref-78]; see also [ref-75]. Franz and Parisi described the jamming limit of the perceptron through the boundary condition that the replica equations take as the overlap tends to one [ref-58] (15). Part II also shows that for κ∈(κd(α),κc(α))\kappa\in(\kappa_{\mathrm{d}}(\alpha),\kappa_{\mathrm{c}}(\alpha)) no minimizer of Pκ\mathcal{P}_{\kappa} is replica symmetric (Part II, Corollary 3.14). Combining Theorems 1.2 and 1.3 with Part II, Theorem 1.1, and with this absence of replica-symmetric minimizers gives the following.

Corollary 1.6. Let e0,c,Ce_{0},c,C be as in (10), let 0<e<e00<e<e_{0}, and put α=2+e\alpha=2+e. Then the following hold.

(i) κN→κc(α)\kappa_{N}\to\kappa_{\mathrm{c}}(\alpha) in probability, ce3≤κRS(α)−κc(α)≤Ce3c e^{3}\leq\kappa_{\mathrm{RS}}(\alpha)-\kappa_{\mathrm{c}}(\alpha)\leq C e^{3}, and

κc(α)−κd(α)=θde2+O(e3log⁡(1/e)).(11)\kappa_{\mathrm{c}}(\alpha)-\kappa_{\mathrm{d}}(\alpha) =\theta_{\mathrm{d}}e^{2}+O\left(e^{3}\log(1/e)\right). \tag*{(11)}

(ii) For every κ∈(κd(α),κRS(α))\kappa\in(\kappa_{\mathrm{d}}(\alpha),\kappa_{\mathrm{RS}}(\alpha)), Gardner’s replica-symmetric formula min⁡q∈[0,1)PκRS(q)\min_{q\in[0,1)}\mathcal{P}_{\kappa}^{\mathrm{RS}}(q) is finite, and there is ε>0\varepsilon>0 such that

P{1Nlog⁡VN(κ)≤min⁡q∈[0,1)PκRS(q)−ε}⟶1.\mathbb{P}\left\{\frac{1}{N}\log V_{N}(\kappa) \leq\min_{q\in[0,1)}\mathcal{P}_{\kappa}^{\mathrm{RS}}(q)-\varepsilon\right\} \longrightarrow1.

Part (i) sharpens Corollary 1.5 to exact order: the largest achievable margin converges to a limit that lies below Gardner’s prediction by an amount of exact order (α−2)3(\alpha-2)^{3}, and above the de Almeida–Thouless margin by an amount of order (α−2)2(\alpha-2)^{2}. The thermodynamic limit is taken at fixed ee, and part (i) describes the subsequent behavior of the limit κc(2+e)\kappa_{\mathrm{c}}(2+e) as e↓0e\downarrow0. Part (ii) shows that Gardner’s formula for the free energy fails on the whole interval (κd(α),κRS(α))(\kappa_{\mathrm{d}}(\alpha),\kappa_{\mathrm{RS}}(\alpha)). By Theorem 1.3(v) it holds at sufficiently negative margins, and by Theorem 1.2(i) and Part II, Theorem 1.2(i), it holds at the margins in [κd(α)−c′e2,κd(α)][\kappa_{\mathrm{d}}(\alpha)-c'e^{2},\kappa_{\mathrm{d}}(\alpha)] for a numerical constant c′>0c'>0. The corollary follows directly from the results above. Part (i) is Theorem 1.2(ii) together with (10), since κc−κd=(κRS−κd)−(κRS−κc)\kappa_{\mathrm{c}}-\kappa_{\mathrm{d}}=(\kappa_{\mathrm{RS}}-\kappa_{\mathrm{d}})-(\kappa_{\mathrm{RS}}-\kappa_{\mathrm{c}}). For part (ii), the minimum of PκRS\mathcal{P}_{\kappa}^{\mathrm{RS}} is attained at q∗(κ)q^{*}(\kappa), so it is finite. Let κ∈(κd,κc)\kappa\in(\kappa_{\mathrm{d}},\kappa_{\mathrm{c}}). By Theorem 1.3(i),(iii), minimizers of Pκ\mathcal{P}_{\kappa} exist, and by Part II, Corollary 3.14, none of them is replica symmetric. Hence P∗(κ)<PκRS(q∗(κ))\mathcal{P}_{*}(\kappa)<\mathcal{P}_{\kappa}^{\mathrm{RS}}(q^{*}(\kappa)), and Theorem 1.2(i) gives the claim. For κ∈[κc,κRS)\kappa\in[\kappa_{\mathrm{c}},\kappa_{\mathrm{RS}}) the claim follows from Theorem 1.2(iii).

Part II also proves a lower bound on the Parisi value that matches Theorem 1.3(iii) up to the constant in front of the logarithm, through a comparison with a zero-temperature functional (Part II, Proposition 2.24): for every α>2\alpha>2 there are C,η0>0C,\eta_{0}>0, depending only on α\alpha, such that

αlog⁡η−C≤P∗(κc−η)≤12log⁡η+C(0<η≤η0).(12)\alpha\log\eta-C \leq\mathcal{P}_{*}(\kappa_{\mathrm{c}}-\eta) \leq\frac{1}{2}\log\eta+C \qquad(0<\eta\leq\eta_{0}). \tag*{(12)}

Together with Theorem 1.2(i), this logarithmic divergence gives the following.

Corollary 1.7. Let α>2\alpha>2. There are C,η0>0C,\eta_{0}>0, depending only on α\alpha, such that for every 0<η≤η00<\eta\leq\eta_{0},

P{αlog⁡η−C≤1Nlog⁡VN(κc−η)≤12log⁡η+C}⟶1.\mathbb{P}\left\{\alpha\log\eta-C\leq\frac{1}{N}\log V_{N}(\kappa_{\mathrm{c}}-\eta)\leq\frac{1}{2}\log\eta+C\right\}\longrightarrow1.

Equivalently, with probability tending to one, e−CNηαN≤VN(κc−η)≤eCNηN/2e^{-CN}\eta^{\alpha N}\leq V_{N}(\kappa_{\mathrm{c}}-\eta)\leq e^{CN}\eta^{N/2}. On the exponential scale the feasible volume thus vanishes polynomially in the distance η\eta to the critical margin. The upper bound uses only Theorems 1.2 and 1.3, and we do not claim that either exponent is optimal.

We next describe the configurations that nearly achieve the maximal margin, through their empirical gap measures νN,σ\nu_{N,\sigma}.

Theorem 1.8. Let α>2\alpha>2. There is a deterministic measure να\nu_{\alpha} on [0,∞)[0,\infty) with να([0,∞))=α\nu_{\alpha}([0,\infty))=\alpha and να({0})=1\nu_{\alpha}(\{0\})=1, whose positive part satisfies

να((0,t])≤αlog⁡3log⁡(2π/t)(0<t<2π),(13)\nu_{\alpha}((0,t])\leq\frac{\alpha\log3}{\log(\sqrt{2\pi}/t)} \qquad(0<t<\sqrt{2\pi}), \tag*{(13)}

such that the following statements hold.

(i) For every bounded Lipschitz ψ:[0,∞)→R\psi:[0,\infty)\to\mathbb{R} and every ε>0\varepsilon>0 there is ρ>0\rho>0 such that

P{sup⁡σ∈SN: mN(σ)≥κc−ρ∣∫ψ dνN,σ−∫ψ dνα∣≤ε}⟶1.\mathbb{P}\left\{ \sup_{\sigma\in S_{N}:\,m_{N}(\sigma)\geq\kappa_{\mathrm{c}}-\rho} \left|\int\psi\,\mathrm{d}\nu_{N,\sigma}-\int\psi\,\mathrm{d}\nu_{\alpha}\right| \leq\varepsilon \right\}\longrightarrow1.

(ii) For every c<(α−2)2/(2α)c<(\alpha-2)^{2}/(2\alpha) and all sufficiently large NN, with probability at least 1−e−cN1-e^{-cN} every maximizer σ∗\sigma_{*} of mNm_{N} has exactly NN contacts. For any choice of maximizers, νN,σ∗→να\nu_{N,\sigma_{*}}\to\nu_{\alpha} weakly in probability, and

1N#{a≤M:0<ba⋅σ∗−κN≤t}⟶να((0,t])in probability\frac{1}{N}\#\{a\leq M:0<b_{a}\cdot\sigma_{*}-\kappa_{N}\leq t\} \longrightarrow\nu_{\alpha}((0,t]) \quad\text{in probability}

at every continuity point t>0t>0 of να\nu_{\alpha}.

(iii) For every choice of margins k↑κc(α)k\uparrow\kappa_{\mathrm{c}}(\alpha) and minimizers γk\gamma_{k} of PkP_{k}, the measures α Law⁡(X1γk,k−k)\alpha\,\operatorname{Law}(X_{1}^{\gamma_{k},k}-k) converge weakly to να\nu_{\alpha}.

Here νN,σ∗→να\nu_{N,\sigma_{*}}\to\nu_{\alpha} weakly in probability means that ∫ψ dνN,σ∗→∫ψ dνα\int\psi\,\mathrm{d}\nu_{N,\sigma_{*}}\to\int\psi\,\mathrm{d}\nu_{\alpha} in probability for every bounded continuous ψ\psi, uniformly over the maximizers. The count NN in part (ii) is isostatic in the sense recalled above: it equals the N−1N-1 degrees of freedom on the sphere plus one for the optimized margin. The identity να({0})=1\nu_{\alpha}(\{0\})=1 and the bound (13) show that, in the limit, no additional macroscopic set of constraints has small positive gaps: the mass να((0,t])\nu_{\alpha}((0,t]) of the small positive gaps tends to zero as t↓0t\downarrow0, at least as fast as a constant times 1/log⁡(1/t)1/\log(1/t). Nearly optimal configurations need not have exactly NN contacts, but by part (i) their gap statistics are close to να\nu_{\alpha}, and by part (iii) the limiting gap law is identified through Parisi minimizers approaching the critical margin. The theorem does not identify the behavior of να\nu_{\alpha} near zero, where physics predicts a power law.

The last theorem identifies the gap statistics of typical feasible configurations at a fixed margin below the critical one.

Theorem 1.9. Let α>2\alpha> 2 and κ<κc(α)\kappa< \kappa_{\mathrm{c}}(\alpha), let γ\gamma be a minimizer of Pκ\mathcal{P}_{\kappa}, and let ψ:[0,∞)→R\psi: [0,\infty) \to\mathbb{R} be bounded and LL-Lipschitz with L>0L > 0. For every ε>0\varepsilon> 0, with probability tending to one, VN(κ)>0V_{N}(\kappa) > 0 and

μN({σ∈SN(κ):∣1N∑a=1Mψ(ba⋅σ−κ)−αEψ(X1γ,κ−κ)∣>ε})≤exp⁡{−ε2N8αL2λγ(0)}VN(κ).(14)\begin{aligned} &\mu_{N}\left(\left\{\sigma\in S_{N}(\kappa) : \left|\frac{1}{N}\sum_{a=1}^{M}\psi(b_{a}\cdot\sigma-\kappa)-\alpha\mathbb{E}\psi\left(X_{1}^{\gamma,\kappa}-\kappa\right)\right|>\varepsilon\right\}\right) \\ &\qquad\le\exp\left\{-\frac{\varepsilon^{2}N}{8\alpha L^{2}\lambda_{\gamma}(0)}\right\}V_{N}(\kappa). \tag*{(14)} \end{aligned}

Consequently, writing ⟨⋅⟩N,κ\langle\cdot\rangle_{N,\kappa} for expectation under the uniform probability measure on SN(κ)S_{N}(\kappa),

⟨1N∑a=1Mψ(ba⋅σ−κ)⟩N,κ⟶αEψ(X1γ,κ−κ)in probability,(15)\left\langle\frac{1}{N}\sum_{a=1}^{M}\psi(b_{a}\cdot\sigma-\kappa)\right\rangle_{N,\kappa} \longrightarrow\alpha\mathbb{E}\psi\left(X_{1}^{\gamma,\kappa}-\kappa\right) \qquad\text{in probability}, \tag*{(15)}

and all minimizers of Pκ\mathcal{P}_{\kappa} induce the same law of X1γ,κX_{1}^{\gamma,\kappa}.

Applied to finitely many test functions at once, the theorem shows that all but an exponentially small fraction of the feasible set has the gap statistics of the Parisi diffusion. The rate in (14) depends on the minimizer only through its mass λγ(0)\lambda_{\gamma}(0), and any minimizer may be used. We do not prove that the minimizers of Pκ\mathcal{P}_{\kappa} are unique, and none of our arguments requires it; the last assertion of the theorem shows that they all induce the same terminal law. For mixed pp-spin models, uniqueness follows from strict convexity of the Parisi functional [ref-7, ref-76], which is not known for Pκ\mathcal{P}_{\kappa}. For generic mixed pp-spin models, Auffinger and Jagannath expressed spin statistics through solutions of partial differential equations [ref-10].

Throughout, CC and cc denote positive constants. A numerical constant depends on nothing; otherwise the dependence of a constant is indicated in the statement where it appears, as in “depending only on α\alpha”. Their values may change from one occurrence to the next. Limits in probability refer to the joint law of the patterns and of any auxiliary randomness introduced in the proofs.

Ideas of the proof

The proof has three parts. Steps 1 and 2 below prove the Parisi formula for smooth activations. Steps 3 and 4 analyze the deterministic variational problem for the hard wall.

Steps 5 and 6 transfer these results to the hard constraints of the random model and to the gaps. Figure 1.1 shows how the steps depend on one another.

Dependency diagram for the six proof steps and their principal results

Figure 1.1. Selected dependencies among the steps of the proof. Each solid box names a step or chapter and, where there is one, the main result it proves; an arrow runs from a box to a box that uses it. The dashed boxes are the corollaries that also use results of Part II. Corollary 1.5 combines Theorems 1.2 and 1.4 and is not drawn, and neither are the analytic properties of the Parisi functional collected in Chapter 2, which are used throughout. Among them, the lemmas of §2.3 are derived from Part II, §§2.1–2.5 and Chapter A, which do not use Part I.

Step 1: The upper bound. The argument follows the proof of [ref-92], Theorem 4.1, a step in Mourrat’s upper bound for bipartite models; the opening of Chapter 3 compares our argument with his. We use a Guerra-type interpolation [ref-68] between the perceptron and a decoupled system on a Ruelle probability cascade [ref-19, ref-112]. Since PUP_U is continuous in the order parameter, it suffices to treat order parameters γh\gamma_h with finitely many atoms h0≤⋯≤hkh_0 \le\cdots\le h_k of equal mass. As the interpolation time tt runs from 0 to 1, the argument of each activation moves from a Gaussian field on the leaves of the cascade, with covariance hlh_l between two leaves that branch at level ll, to ba⋅σb_a \cdot\sigma. The coordinates are coupled to a second cascade field, whose levels p=(p0,…,pk)p=(p_0,\ldots,p_k) are free parameters. The derivative in tt of the interpolating free energy uNu_N is expressed through two overlaps: the spin overlap R12=σ1⋅σ2/NR_{12}=\sigma^1\cdot\sigma^2/N, and the response overlap X12=M−1∑aU′(Sa1)U′(Sa2)X_{12}=M^{-1}\sum_a U'(S_a^1)U'(S_a^2), where SalS_a^l is the interpolated field of constraint aa for replica ll. Overlaps of the derivatives of the activation enter in a similar way in Talagrand’s treatment of the perceptron [ref-127], Chapters 2 and 3, [ref-128], Chapter 8. Let JJ be the level at which the leaves of two replicas branch, wlw_l the probability that J=lJ=l, and rlr_l and ζl\zeta_l the conditional means of R12R_{12} and of αNX12\alpha_N X_{12} given J=lJ=l, where αN=M/N\alpha_N=M/N. Then

∂tuN=−12∑l=0kwl[(rl−hl)ζl+Cov⁡(R12,αNX12∣J=l)].\partial_t u_N=-\frac{1}{2}\sum_{l=0}^{k}w_l\left[(r_l-h_l)\zeta_l+\operatorname{Cov}\left(R_{12},\alpha_N X_{12}\mid J=l\right)\right].

Neither term has a sign. We control the products through the choice of pp. At a maximizer over pp of uN+12∑lwlhlplu_N+\frac{1}{2}\sum_l w_l h_l p_l, the first-order conditions, together with the monotonicity 0≤ζ0≤⋯≤ζk0\leq\zeta_0\leq\cdots\leq\zeta_k, make the total contribution of the products nonpositive. This monotonicity holds for every activation in class (A); we derive it from an exact change of measure on a cascade with marks. For Gaussian fields the corresponding monotonicity is [ref-128], Proposition 14.3.2 and [ref-92], Lemma 2.4. A maximum principle in tt alone then bounds the free energy by the maximum over pp at t=0t=0, up to the covariance terms. To control these, small Gaussian perturbations of the Hamiltonian, of the form used in [ref-92], Section 4, enforce approximate Ghirlanda–Guerra identities [ref-64] for the pair of overlap arrays at the points where the maximum principle is applied. Mourrat’s quantitative synchronization [ref-92], which builds on Panchenko’s synchronization mechanism [ref-100, ref-101], then makes the conditional variance of R12R_{12} small, and the Cauchy–Schwarz inequality makes the covariances small whatever the sign of U′U'. This is why no sign condition on U′U' is needed. At t=0t=0 the system splits into the row term αNuγh(0,0;U)\alpha_N u_{\gamma_h}(0,0;U) and a spherical term. An exact computation with Gaussian cascades bounds the spherical term plus 12∑lwlhlpl\frac{1}{2}\sum_l w_l h_l p_l, maximized over pp, by CS⁡(γh)+o(1)\operatorname{CS}(\gamma_h)+o(1); a computation of this kind for spherical spin glasses is in [ref-125]. This is carried out in Chapter 3.

Step 2: The lower bound. We use the Aizenman–Sims–Starr scheme [ref-1] in the spherical form of [ref-37]. For multi-species spherical models, Bates and Sohn combined W.-K. Chen’s cavity computation with Panchenko’s synchronization [ref-100, ref-101] to prove the lower bound [ref-16]. We follow the same route, with the response overlap in the role of the overlap of the second species. The scheme bounds the free energy from below by the increment of the expected log partition function when nn coordinates, and about αn\alpha n constraints, are added to a system with NN coordinates. We write the sphere of dimension N+nN+n exactly as a product of the sphere of dimension NN and a density for the nn new coordinates (compare [ref-47]). The increment then splits into a row term and a coordinate term. In the limit the row term is at least αn uζ(0,0;U)\alpha n\,u_\zeta(0,0;U), where ζ\zeta is the limiting overlap law. After a Gaussian interpolation, the coordinate term becomes a Gaussian integral over the new coordinates ε∈Rn\varepsilon\in\mathbb{R}^n. After a perturbation (see [ref-98] and [ref-100], (24)–(28)), the joint law of the spin overlap and the response overlap satisfies the Ghirlanda–Guerra identities. By synchronization and the characterization of such arrays [ref-49, ref-92, ref-97, ref-100, ref-101], it is a limit of Ruelle probability cascades with paired nondecreasing levels. On such a cascade the coordinate term is an explicit Gaussian integral, which for general paired levels is not the entropy of any order parameter.

Two identities of the system determine the result. First, a Gaussian integration by parts in the rows expresses the mean effective precision of the new coordinates through the two overlaps. In the limit, this precision identity shows that the smallest precision is at least one, so that the Gaussian integral is finite, and it cancels the constant terms. Second, the law of the patterns is invariant under rotations. Under a uniformly random rotation, the overlap ε1⋅ε2/n\varepsilon^{1}\cdot\varepsilon^{2}/n of the new coordinates of two replicas is close to their full overlap, with mean squared error at most 3/n3/n in the limit. Further, E∥ε∥2/n=1\mathbb{E}\lVert\varepsilon\rVert^{2}/n=1, and this norm identity turns the coordinate term into nCS⁡(γ)n\operatorname{CS}(\gamma), where γ\gamma is the law of the mean overlap of the new coordinates at the branching level, and the overlap estimate makes γ\gamma close to ζ\zeta. Hence, as N→∞N\to\infty, the increment per added coordinate is at least

PU(γ)−O(n−1/2)≥P∗(U)−O(n−1/2).\mathcal{P}_{U}(\gamma)-O(n^{-1/2})\geq\mathcal{P}_{*}(U)-O(n^{-1/2}).

Rotation invariance thus forces the cavity term to equal the Crisanti–Sommers entropy, and no minimizer or critical point of a cavity functional has to be identified. The main technical issue is the unbounded Gaussian integral over ε\varepsilon. We compute it on balls, pass to the limit N→∞N\to\infty there, evaluate the limit on finite cascades, and remove the cutoff on these cascades, where the precision identity gives uniform control. This is carried out in Chapter 4. Chapter 5 combines the two bounds. There a truncation argument, which moves the mass of an order parameter above the level rUr_{U} down to rUr_{U} and so decreases PU\mathcal{P}_{U}, gives the attainment of the infimum and the bound Qγ≤rUQ_{\gamma}\leq r_{U} for minimizers.

Step 3: The variational problem at hard margins. For the hard wall we prove an exact value identity, valid for every order parameter:

Pκ(γ)=12log⁡δ+αEh(Z)−δ2D(Q)−12∫0Qγ(q)D(q) dq.\mathcal{P}_{\kappa}(\gamma)=\frac{1}{2}\log\delta+\alpha\mathbb{E}h(Z)-\frac{\delta}{2}D(Q)-\frac{1}{2}\int_{0}^{Q}\gamma(q)D(q)\,\mathrm{d}q.

Here Q=QγQ=Q_{\gamma}, δ=1−Q\delta=1-Q, Z=(XQ−κ)/δZ=(X_{Q}-\kappa)/\sqrt{\delta} is the normalized top position of the Parisi diffusion, h(z)=log⁡Φ(z)+12R(z)2<0h(z)=\log\Phi(z)+\frac{1}{2}R(z)^{2}<0, and D(q)=αE ∂xuγ(q,Xq;κ)2−∫0qλγ−2D(q)=\alpha\mathbb{E}\,\partial_{x}u_{\gamma}(q,X_{q};\kappa)^{2}-\int_{0}^{q}\lambda_{\gamma}^{-2} is twice the density of the first variation of Pκ\mathcal{P}_{\kappa}. El Alaoui and Sellke computed this first variation at a minimizer whose support is an interval, in directions supported on that interval [ref-54] (Proposition 16) (see the opening of Chapter 2). Here it is used at every order parameter and in every direction that keeps the top of the support below a fixed level; Chapter 6 recalls the first-order conditions known for mixed pp-spin models. At a minimizer over the order parameters whose top is at most a fixed level r<1r<1, the last two terms are nonpositive, so Pκ(γ)<12log⁡δ\mathcal{P}_{\kappa}(\gamma)<\frac{1}{2}\log\delta. A finite value therefore keeps the top of every minimizer away from one, which gives existence and compactness of minimizers. The second tool is the universal slope bound

0<λγ(0) ∂xuγ(0,0;κ)<Cκ,0<\lambda_{\gamma}(0)\,\partial_{x}u_{\gamma}(0,0;\kappa)<C_{\kappa},

valid for every order parameter and proved by a maximum principle for an explicit combination of ∂xuγ\partial_x u_\gamma and ∂xxuγ\partial_{xx}u_\gamma. The map κ↦Pκ(γ)\kappa\mapsto\mathcal{P}_\kappa(\gamma) is concave, by Prékopa’s theorem [ref-108], with derivative −α ∂xuγ(0,0;κ)-\alpha\,\partial_xu_\gamma(0,0;\kappa), and λγ(0)≥δ\lambda_\gamma(0)\geq\delta, so the two tools transport values between margins. Take a constrained minimizer γ\gamma with a prescribed value ww at a margin k>κck>\kappa_c, where the unconstrained value is −∞-\infty. The value identity gives 1−Qγ>e2w1-Q_\gamma>e^{2w}, so its slope in the margin is at most αCke−2w\alpha C_k e^{-2w}. Concavity and the limit k↓κck\downarrow\kappa_c give P∗(κc−η)≤w+αηCκce−2w\mathcal{P}_*(\kappa_c-\eta)\leq w+\alpha\eta C_{\kappa_c}e^{-2w}, and the choice w=12log⁡(2αCκcη)w=\frac{1}{2}\log(2\alpha C_{\kappa_c}\eta) gives Theorem 1.3(iii).

For a minimizer over all order parameters, the value identity holds without the last two terms. Since h(z)h(z) is bounded below by −log⁡(1+z−)-\log(1+z_-) up to an additive constant, and the slope bound controls the negative part of ZZ, Jensen’s inequality gives the pinning bound of Theorem 1.3(iv). Combined with (iii) at the margin κc−η\kappa_c-\eta, it forces 1−Qγ≤Cη1/(1+α)1-Q_\gamma\leq C\eta^{1/(1+\alpha)}. Finally, weak top stability, a lower bound on the quantity S(Q)=αδ2E ∂xxuγ(Q,XQ;κ)2S(Q)=\alpha\delta^2\mathbb{E}\,\partial_{xx}u_\gamma(Q,X_Q;\kappa)^2 for constrained minimizers with more than one support point, gives λγ(0)≤C(1−Qγ)1/2\lambda_\gamma(0)\leq C(1-Q_\gamma)^{1/2}. At very negative margins the same lower bound cannot hold, so every minimizer is replica symmetric; this gives κc>−∞\kappa_c>-\infty and Theorem 1.3(v). No zero-temperature functional is needed. This is carried out in Chapter 6.

Step 4: The critical margin near α=2\alpha=2. To go below Gardner’s prediction we evaluate Pκ\mathcal{P}_\kappa at an explicit order parameter with two support points near one,

γδ=a1[p,Q)+1[Q,1],Q=1−δ,p=Q(1−ℓ),a=δQℓ,\gamma_\delta=a\mathbf{1}_{[p,Q)}+\mathbf{1}_{[Q,1]},\qquad Q=1-\delta,\qquad p=Q(1-\ell),\qquad a=\frac{\delta}{Q\ell},

which splits the atom of a replica-symmetric order parameter and places a thin layer of relative width ℓ\ell below the top. The homogeneity and the monotonicity of the Cole–Hopf step and a Gaussian tail bound give δPκ(γδ)≤QJℓ(κ/Q)+12δlog⁡δ\delta\mathcal{P}_\kappa(\gamma_\delta)\leq QJ_\ell(\kappa/\sqrt{Q})+\frac{1}{2}\delta\log\delta for an explicit Gaussian integral JℓJ_\ell, and

Jℓ(k)=1−αE(k−G)+24−ℓ2(log⁡2−12)(αΦ(k)−1)+O(ℓ3/2).J_\ell(k)=\frac{1-\alpha\mathbb{E}(k-G)_+^2}{4}-\frac{\ell}{2}\left(\log2-\frac{1}{2}\right)(\alpha\Phi(k)-1)+O(\ell^{3/2}).

If Jℓ(k)<0J_\ell(k)<0, then Pκ(γδ)→−∞\mathcal{P}_\kappa(\gamma_\delta)\to-\infty as δ↓0\delta\downarrow0 for every κ>k\kappa>k, so that κc≤k\kappa_c\leq k. At k=κRSk=\kappa_{\mathrm{RS}} the first term vanishes, while αΦ(κRS)−1\alpha\Phi(\kappa_{\mathrm{RS}})-1 is of exact order e=α−2e=\alpha-2. The layer therefore gains an amount of order eℓe\ell at a cost of order ℓ3/2\ell^{3/2}, and the choice ℓ≍e2\ell\asymp e^2 makes Jℓ(κRS)J_\ell(\kappa_{\mathrm{RS}}) negative, of order e3e^3. Lowering the margin by a small multiple of e3e^3 keeps it negative. The argument uses no minimizer and no regularity theory beyond the recursion for this single order parameter. Unlike the bounds of Stojnic and of Montanari, Zhong, and Zhou recalled at the beginning of this chapter, the bound comes from the Parisi functional at an explicit order parameter, and the thin layer of width ℓ≍e2\ell\asymp e^2 gives the order e3e^3 near α=2\alpha=2. An explicit trial order parameter was used in the same way by Toninelli to show that the Sherrington–Kirkpatrick model is not replica symmetric beyond the de Almeida–Thouless line [ref-129]. This is carried out in Chapter 7.

Step 5: From soft free energy to hard feasible volume. The upper bounds on the volume and on κN\kappa_N come from Step 1, applied to smooth approximations of the hard wall from above, together with a lemma showing that a single feasible configuration forces an exponentially nonnegligible feasible volume after a fixed decrease of the margin. For the lower bound, the Parisi formula for a concave penalty comparable to −β(ℓ−x)+2-\beta(\ell-x)_{+}^{2} bounds from below the volume of the configurations whose total squared violation ∑a(k0−ba⋅σ)+2\sum_a(k_0-b_a\cdot\sigma)_{+}^{2}, at a level k0k_0 slightly below ℓ\ell, is at most tNtN, for small t>0t>0. Two geometric facts turn this into feasible volume. First, a small violation can be repaired by a small displacement. Using the smallest singular value of the pattern matrix over sparse sets of rows (see [ref-130]), a separation argument moves such a configuration by at most 4tN4\sqrt{tN} to a point that satisfies every constraint at a slightly smaller margin. The opening of Chapter 8 compares this repair with that of El Alaoui and Sellke [ref-54]. Second, the repair map may collapse volume, so we do not use its image. Instead we add to the configuration a Gaussian vector tangent to the sphere, with variance ρ2\rho^{2} in each direction, and project back to SNS_N. This preserves μN\mu_N. Translating the Gaussian vector by a displacement dd lowers the probability of a symmetric set by at most the factor e−∥d∥2/(2ρ2)e^{-\lVert d\rVert^{2}/(2\rho^{2})}, and Šidák’s inequality [ref-116] shows that, with probability at least e−Nϵe^{-N\epsilon} for a small ϵ>0\epsilon>0, the Gaussian vector moves no constraint by more than a small amount. Hence the perturbed point is feasible at a slightly smaller margin with probability at least e−Nϵ′e^{-N\epsilon'}, where ϵ′\epsilon' is small when tt is small compared with ρ2\rho^{2}. The upper bounds are proved in §§3.8 and 3.9 and the lower bound in Chapter 8.

Step 6: Gap statistics. Gap averages enter through the free energy of the feasible set with a source term β∑aψ(ba⋅σ−k)\beta\sum_a\psi(b_a\cdot\sigma-k), for a bounded Lipschitz test function ψ\psi. Adding the source to a smooth approximation of the hard wall destroys monotonicity, and this is where the generality of class (A) is used. We prove a quadratic source estimate: the Parisi functional at γ\gamma with the source exceeds its linearization in β\beta, whose slope is αEψ(X1γ,k−k)\alpha\mathbb{E}\psi(X_1^{\gamma,k}-k), by at most 12αLip⁡(ψ)2β2λγ(0)\frac{1}{2}\alpha\operatorname{Lip}(\psi)^{2}\beta^{2}\lambda_\gamma(0). For a step order parameter the effect of the source is an iterated logarithmic moment along the Parisi diffusion observed at the breakpoints of γ\gamma. The transitions of this Markov chain are Gaussian laws tilted by log-concave weights. Like Gaussian laws, they have sub-Gaussian linear statistics and map Lipschitz functions to Lipschitz functions with the same constant (for log-concave tilts of Gaussian laws see [ref-25, ref-27]), and a backward induction with a second-order expansion at each step gives the estimate. Combined with Jensen’s inequality and the upper bound of Step 1, it yields one deviation inequality. With probability tending to one, simultaneously for every set T⊆SN(k)T\subseteq S_N(k) of positive volume, the average over TT of N−1∑aψ(ba⋅σ−k)N^{-1}\sum_a\psi(b_a\cdot\sigma-k) differs from αEψ(X1γ,k−k)\alpha\mathbb{E}\psi(X_1^{\gamma,k}-k) by at most

Pk(γ)−N−1log⁡μN(T)β+αLip⁡(ψ)2β2λγ(0)+ϵ.\frac{\mathcal{P}_k(\gamma)-N^{-1}\log\mu_N(T)}{\beta} +\frac{\alpha\operatorname{Lip}(\psi)^{2}\beta}{2}\lambda_\gamma(0)+\epsilon.

For Theorem 1.9 we apply it at a fixed margin κ<κc\kappa< \kappa_{\mathrm{c}} to the set of feasible configurations whose gap average deviates by more than ε\varepsilon, and Theorem 1.2(i) shows that this set is an exponentially small fraction of SN(κ)S_N(\kappa). For Theorem 1.8, a single configuration of nearly maximal margin is surrounded by a set of configurations that are feasible at a margin smaller by η\eta, have almost the same gaps, and have volume at least e−N(log⁡(1/η)+C)\mathrm{e}^{-N(\log(1/\eta)+C)}. Optimizing over β\beta leaves an error of order η+{λγ(0)(1+log⁡(1/η))}1/2\eta+\{\lambda\gamma(0)(1+\log(1/\eta))\}^{1/2}, which tends to zero by the collapse of minimizers in Theorem 1.3(iv).

The exact count of contacts does not use the variational problem. By Wendel’s theorem [ref-131], κN<0\kappa_N<0 with probability exponentially close to one when α>2\alpha>2. The maximizers of mNm_N then correspond to the points of the polytope {y∈RN:ga⋅y≥−1 for all a}\{y\in\mathbb{R}^N:g_a\cdot y\geq-1\text{ for all }a\} farthest from the origin. These are vertices, and Gaussian general position gives exactly NN contacts. A union bound over sets of NN rows gives a finite-NN form of (13) (Proposition 9.2). This is carried out in Chapter 9.

Chapter 2 collects the analytic properties of the Parisi functional used throughout: the Cole–Hopf recursion for step order parameters, Lipschitz continuity in the order parameter, and, for the hard wall, a priori bounds, signs of derivatives, and the first variation. Theorem 1.1 is proved in Chapter 5, Theorem 1.3 in Chapter 6, Theorem 1.4 in Chapter 7, Theorem 1.2 in Chapter 8, and Theorems 1.8 and 1.9 in Chapter 9. The corollaries follow from these theorems and, for Corollaries 1.6 and 1.7, from Part II, as explained next to their statements. The lemmas of §2.3, parts (i)–(iv) of Lemma 6.1, one Gaussian bound in the proof of Lemma 6.6 and the expansions of Chapter 7 are derived from Part II, §§2.1–2.5, Lemma 3.1 and Chapter A, which do not use Part I. Through them, Theorems 1.2–1.4, 1.8, and 1.9 also rest on Part II; Theorem 1.1 does not.

The Parisi functional

This chapter collects the deterministic facts about the Parisi functional that are used in the rest of this part. §2.1 introduces the Cole–Hopf step Tm,sT_{m,s} and the backward recursion that solves the Parisi equation (3) when the order parameter is a step function. It records the monotonicity of the recursion in its terminal datum, which is used in Chapters 3, 5, 8, and 9, and the homogeneity of the Cole–Hopf step, which is used in Chapter 7. It also shows that neither the Crisanti–Sommers entropy nor the hard-wall value depends on the cutoff used to define it. §2.2 proves that, for activations in class (A) or class (B), the value uγ(0,0;U)u_{\gamma}(0,0;U) is Lipschitz in γ\gamma for the L1L^{1} distance, with an explicit constant (Lemmas 2.3 and 2.4). This defines uγu_{\gamma} for general order parameters, and it lets the upper bound (Chapter 3), the lower bound (Chapter 4), the truncation argument of Chapter 5 and the passage to hard constraints in Chapter 8 move between step order parameters and general ones. These continuity estimates rest on one identity, (30) in Lemma 2.2, which writes the difference of two solutions with the same terminal datum as an integral along a diffusion. Its only input is a formula for the derivatives of one Cole–Hopf step (Lemma 2.1). §2.3 treats the hard-wall functional Pκ\mathcal{P}_{\kappa} on order parameters that equal one above a fixed level r<1r<1. There we state a priori bounds on the solution and on the associated diffusion (Lemma 2.6), Lipschitz continuity in the order parameter (Lemma 2.7), signs and a martingale (Lemma 2.8), and the first variation (Lemma 2.9). These are the inputs for the analysis of the variational problem in Chapter 6 and for the gap statistics in Chapter 9. Most of them are derived from Part II, as explained at the beginning of §2.3.

The hard-wall functional Pκ\mathcal{P}_{\kappa} was studied by El Alaoui and Sellke in [ref-54], Section 4. For step order parameters they solved the Parisi equation with the terminal datum (5) by the Cole–Hopf recursion [ref-54], (4.16) and proved a priori bounds on the solution [ref-54], Proposition 12. They proved Lipschitz continuity in the order parameter [ref-54], Proposition 13, extended the solution to general order parameters [ref-54], (4.12), and proved that it has one-sided derivatives in time [ref-54], Lemma 14. They also computed the first variation of Pκ\mathcal{P}_{\kappa} in terms of the Parisi diffusion [ref-54], (2.8), (4.13), and (4.15), at a minimizer whose support is an interval, in directions supported on that interval, and deduced the first-order conditions D=0D=0 and S=1S=1 on that interval [ref-54], Proposition 16, with DD and SS as in (44). Their arguments build on work on the Sherrington–Kirkpatrick and mixed pp-spin models: Guerra’s Lipschitz bound in the order parameter [ref-67], the Feynman–Kac proof of it by Jagannath and Tobasco [ref-76] (Lemma 14 and Remark 15), the diffusion and the regularity estimates of Auffinger and W.-K. Chen [ref-7, ref-8], and the first-variation formulas of W.-K. Chen [ref-38] and of El Alaoui, Montanari, and Sellke [ref-55] (Proposition 6.8). This chapter and Part II, Chapter 2 follow the same route, with the following differences. Concavity of the solution is obtained from Prékopa’s theorem (Lemma 2.1) rather than from an explicit computation along the recursion, which uses that the order parameter is nondecreasing [ref-54] (proof of (4.6)). The bound −∂xxuγ≤(1−qˉ)−1-\partial_{xx}u_{\gamma}\le(1-\bar q)^{-1} of [ref-54] ((4.6)) is sharpened to the strict bound −∂xxuγ<1/λγ-\partial_{xx}u_{\gamma}<1/\lambda_{\gamma} (Lemma 2.8). The constants in §2.3 are uniform for κ\kappa in a compact interval. The first variation is computed at every order parameter in the class Ur\mathcal{U}_{r} of order parameters equal to one on [r,1][r,1], in every direction that stays in Ur\mathcal{U}_{r} (Lemma 2.9), and not only at a minimizer. Finally, the Lipschitz estimates of §2.2 cover activations in class (B), which are concave and may tend to −∞-\infty quadratically.

Throughout this chapter, GG denotes a standard Gaussian variable and G′G' an independent copy of it, and ϕ\phi and Φ\Phi denote the standard Gaussian density and distribution function. All functions are Borel measurable, and ∥⋅∥L1\lVert\cdot\rVert_{L^{1}} is the norm of L1([0,1])L^{1}([0,1]).

Order parameters and the Cole–Hopf recursion

Recall from §1.1 that an order parameter is a nondecreasing, right-continuous function γ:[0,1]→[0,1]\gamma:[0,1]\to[0,1] that is identically one on [qˉ,1][\bar q,1] for some qˉ<1\bar q<1, and that U\mathcal{U} is the set of order parameters. For γ∈U\gamma\in\mathcal{U} we write

Qγ=inf⁡{q∈[0,1]:γ(q)=1}<1,λγ(q)=∫q1γ(t) dt.Q_{\gamma}=\inf\{q\in[0,1]:\gamma(q)=1\}<1,\qquad \lambda_{\gamma}(q)=\int_{q}^{1}\gamma(t)\,\mathrm{d}t.

Thus γ\gamma is the distribution function of a probability measure on [0,Qγ][0,Q_{\gamma}]. More generally, we identify every probability measure on [0,1][0,1] with its distribution function, a nondecreasing right-continuous function γ:[0,1]→[0,1]\gamma:[0,1]\to[0,1] with γ(1)=1\gamma(1)=1.

The Cole–Hopf step. For m≥0m\ge0, s≥0s\ge0, and a function f:R→[−∞,∞)f:\mathbb{R}\to[-\infty,\infty) that is bounded above, define

Tm,sf(x)={1mlog⁡Eemf(x+sG),m>0,Ef(x+sG),m=0,(16)T_{m,s}f(x)= \begin{cases} \frac{1}{m}\log\mathbb{E}e^{m f(x+\sqrt{s}G)}, & m>0,\\ \mathbb{E}f(x+\sqrt{s}G), & m=0, \end{cases} \tag*{(16)}

with the conventions e−∞=0e^{-\infty}=0 and log⁡0=−∞\log0=-\infty. The expectations exist because ff is bounded above, Tm,sfT_{m,s}f takes values in [−∞,sup⁡f][-\infty,\sup f], and Tm,0f=fT_{m,0}f=f. For functions f,gf,g that are bounded above, a∈Ra\in\mathbb{R}, m≥0m\ge0, and s,s′≥0s,s'\ge0, the following hold.

(T1) If f≤g+af\le g+a on R\mathbb{R}, then Tm,sf≤Tm,sg+aT_{m,s}f\le T_{m,s}g+a on R\mathbb{R}.

(T2) Tm,sTm,s′f=Tm,s+s′fT_{m,s}T_{m,s'}f=T_{m,s+s'}f.

(T3) t Tm,sf=Tm/t,s(tf)t\,T_{m,s}f=T_{m/t,s}(tf) for every t>0t>0.

(T4) Tm,sf≥T0,sf\mathcal{T}_{m,s}f \ge\mathcal{T}_{0,s}f.

If f≡−∞f \equiv-\infty, then Tm,sf≡−∞\mathcal{T}_{m,s}f \equiv-\infty for all m,sm,s, and (T2)–(T4) for ff are immediate; below we assume f≢−∞f \not\equiv-\infty, so that sup⁡f\sup f is finite. Property (T1) holds because y↦emyy \mapsto e^{my} is nondecreasing and Tm,s(g+a)=Tm,sg+a\mathcal{T}_{m,s}(g+a)=\mathcal{T}_{m,s}g+a. For (T2), write x+s+s′Gx+\sqrt{s+s'}G as x+sG+s′G′x+\sqrt{s}G+\sqrt{s'}G'. If m>0m>0, Tonelli’s theorem applied to emf≥0e^{mf}\ge0 gives

Eexp⁡{mTm,s′f(x+sG)}=Eemf(x+sG+s′G′)=Eemf(x+s+s′G),\mathbb{E}\exp\left\{m\mathcal{T}_{m,s'}f(x+\sqrt{s}G)\right\} = \mathbb{E}e^{mf(x+\sqrt{s}G+\sqrt{s'}G')} = \mathbb{E}e^{mf(x+\sqrt{s+s'}G)},

and if m=0m=0 the same argument applies to sup⁡f−f≥0\sup f-f\ge0. Property (T3) follows from (16), and (T4) is Jensen’s inequality Eemf≥emEf\mathbb{E}e^{mf}\ge e^{m\mathbb{E}f} for m>0m>0 and an equality for m=0m=0.

The operator Tm,s\mathcal{T}_{m,s} solves the Parisi equation over an interval on which the order parameter equals a constant mm; this is the Cole–Hopf transformation. Indeed, if uu solves (3) with γ=m>0\gamma=m>0, then h=emuh=e^{mu} satisfies

∂qh+12∂xxh=mh(∂qu+12∂xxu+m2(∂xu)2)=0,\partial_q h+\frac{1}{2}\partial_{xx}h = mh\left(\partial_q u+\frac{1}{2}\partial_{xx}u+\frac{m}{2}(\partial_xu)^2\right) = 0,

the backward heat equation, whose solutions are Gaussian averages of their later values. When m=0m=0 the equation is itself the backward heat equation. The recursion below turns this computation into a definition.

The step recursion. Let T∈[0,1]T\in[0,1]. A step function on [0,T)[0,T) is a function γ=∑j=0nmj1[tj,tj+1)\gamma=\sum_{j=0}^{n}m_j\mathbf{1}_{[t_j,t_{j+1})} with 0=t0<t1<⋯<tn+1=T0=t_0<t_1<\cdots<t_{n+1}=T and mj≥0m_j\ge0. For such γ\gamma and a function F:R→[−∞,∞)F:\mathbb{R}\to[-\infty,\infty) that is bounded above, define uγ(⋅,⋅;F)u_\gamma(\mathord{\cdot},\mathord{\cdot};F) on [0,T]×R[0,T]\times\mathbb{R} by the backward recursion

uγ(T,⋅;F)=F,uγ(q,⋅;F)=Tmj,tj+1−quγ(tj+1,⋅;F)(q∈[tj,tj+1)).(17)u_\gamma(T,\mathord{\cdot};F)=F,\qquad u_\gamma(q,\mathord{\cdot};F) = \mathcal{T}_{m_j,t_{j+1}-q}u_\gamma(t_{j+1},\mathord{\cdot};F) \qquad (q\in[t_j,t_{j+1})). \tag*{(17)}

We call FF the terminal datum and TT the terminal time; the terminal time is always clear from the context. By (T2), two consecutive steps with the same value compose into one step, so inserting a breakpoint does not change uγu_\gamma. Since two representations of γ\gamma have a common refinement, uγ(⋅,⋅;F)u_\gamma(\mathord{\cdot},\mathord{\cdot};F) depends only on γ\gamma and FF. For T=0T=0 there are no steps and uγ(0,⋅;F)=Fu_\gamma(0,\mathord{\cdot};F)=F. The recursion goes back to Parisi [ref-105]; see also [ref-68, ref-128], Chapter 14, and [ref-98], Chapter 3.

A step order parameter is an order parameter of the form

γ(q)=mjfor q∈[tj,tj+1), 0≤j≤n,0=t0<t1<⋯<tn<tn+1=1,(18)\gamma(q)=m_j \quad\text{for }q\in[t_j,t_{j+1}),\ 0\le j\le n, \qquad 0=t_0<t_1<\cdots<t_n<t_{n+1}=1, \tag*{(18)}

with 0≤m0≤⋯≤mn=10\le m_0\le\cdots\le m_n=1 and γ(1)=1\gamma(1)=1. For such γ\gamma and an activation UU that is bounded above, uγ(⋅,⋅;U)u_\gamma(\mathord{\cdot},\mathord{\cdot};U) is given by (17) with T=1T=1 and F=UF=U. The same definition applies to every step function γ\gamma that is the distribution function of a probability measure on [0,1][0,1]; its value mnm_n on the last interval [tn,1)[t_n,1) need not be one.

The properties of Tm,s\mathcal{T}_{m,s} pass to the recursion. By (T1), for functions F,F′F,F' bounded above and a∈Ra \in\mathbb{R},

F≤F′+a on R⟹uγ(q,x;F)≤uγ(q,x;F′)+a on [0,T]×R.(19)F \le F' + a \ \text{on } \mathbb{R} \quad\Longrightarrow\quad u_{\gamma}(q,x;F) \le u_{\gamma}(q,x;F') + a \ \text{on } [0,T] \times\mathbb{R}. \tag*{(19)}

In particular F↦uγ(q,x;F)F \mapsto u_{\gamma}(q,x;F) is nondecreasing and commutes with the addition of constants. For real-valued F,F′F,F' whose recursions are finite on [0,T]×R[0,T] \times\mathbb{R} (in particular, when both terminal data are bounded below by a quadratic),

sup⁡x∈R∣uγ(q,x;F)−uγ(q,x;F′)∣≤sup⁡x∈R∣F−F′∣.(20)\sup_{x \in\mathbb{R}} \left|u_{\gamma}(q,x;F) - u_{\gamma}(q,x;F')\right| \le \sup_{x \in\mathbb{R}} |F-F'|. \tag*{(20)}

By (T4), (T1) and the case m=0m=0 of (T2), uγ(q,x;F)≥EF(x+T−q G)u_{\gamma}(q,x;F) \ge\mathbb{E}F(x+\sqrt{T-q}\,G). Hence, if F(y)≥−C(1+y2)F(y) \ge-C(1+y^{2}) for all yy, then

uγ(q,x;F)≥EF(x+T−q G)≥−C(1+x2+T−q)≥−C(2+x2).(21)u_{\gamma}(q,x;F) \ge\mathbb{E}F(x+\sqrt{T-q}\,G) \ge-C(1+x^{2}+T-q) \ge-C(2+x^{2}). \tag*{(21)}

If γ=1\gamma=1 on [qˉ,T)[\bar{q},T) for some qˉ∈[0,T)\bar{q} \in[0,T), then (T2) gives

uγ(qˉ,x;F)=T1,T−qˉF(x)=log⁡EeF(x+T−qˉ G).(22)u_{\gamma}(\bar{q},x;F) = \mathcal{T}_{1,T-\bar{q}}F(x) = \log\mathbb{E}e^{F(x+\sqrt{T-\bar{q}}\,G)}. \tag*{(22)}

We next record the regularity of the recursion. Suppose that FF is real-valued and continuous, and that −C(1+y2)≤F(y)≤C-C(1+y^{2}) \le F(y) \le C for a constant CC. Then u=uγ(⋅,⋅;F)u=u_{\gamma}(\mathord{\cdot},\mathord{\cdot};F) is real-valued and continuous on [0,T]×R[0,T] \times\mathbb{R}, and on each open interval (tj,tj+1)(t_{j},t_{j+1}) it is C∞C^{\infty} and solves

∂qu+12∂xxu+mj2(∂xu)2=0.(23)\partial_{q}u+\frac{1}{2}\partial_{xx}u+\frac{m_{j}}{2}(\partial_{x}u)^{2}=0. \tag*{(23)}

By (21) and induction over the steps, it suffices to prove the following for a continuous function ff with −C′(1+y2)≤f(y)≤C′-C'(1+y^{2}) \le f(y) \le C', a number m≥0m \ge0, and g(s,x)=Tm,sf(x)g(s,x)=\mathcal{T}_{m,s}f(x): the function gg is C∞C^{\infty} on (0,∞)×R(0,\infty) \times\mathbb{R} with ∂sg=12∂xxg+m2(∂xg)2\partial_{s}g=\frac{1}{2}\partial_{xx}g+\frac{m}{2}(\partial_{x}g)^{2}, and g(s,x)→f(x0)g(s,x) \to f(x_{0}) as (s,x)→(0,x0)(s,x) \to(0,x_{0}). For m>0m>0 and s>0s>0,

emg(s,x)=∫Re−(y−x)2/(2s)2πsemf(y) dy,0<emf≤emC′.e^{mg(s,x)} = \int_{\mathbb{R}} \frac{e^{-(y-x)^{2}/(2s)}}{\sqrt{2\pi s}} e^{mf(y)}\,\mathrm{d}y, \qquad 0<e^{mf}\le e^{mC'}.

Differentiation under the integral shows that the right side is positive, C∞C^{\infty} in (s,x)(s,x), and solves ∂sh=12∂xxh\partial_{s}h=\frac{1}{2}\partial_{xx}h; reading the Cole–Hopf computation backwards gives the equation for gg. For m=0m=0 the same argument applies to gg itself, since ∣f(y)∣≤C′(1+y2)|f(y)| \le C'(1+y^{2}). For the limit, fix ρ>0\rho>0. For m>0m>0,

∣Eemf(x+s G)−emf(x)∣≤sup⁡∣y−x∣≤ρ∣emf(y)−emf(x)∣+2emC′P{s∣G∣>ρ},\left|\mathbb{E}e^{mf(x+\sqrt{s}\,G)}-e^{mf(x)}\right| \le \sup_{|y-x|\le\rho} \left|e^{mf(y)}-e^{mf(x)}\right| + 2e^{mC'}\mathbb{P}\{\sqrt{s}|G|>\rho\},

and for m=0m=0 the last term is replaced by E[∣f(x+s G)−f(x)∣;s∣G∣>ρ]\mathbb{E}[|f(x+\sqrt{s}\,G)-f(x)|;\sqrt{s}|G|>\rho], which is at most C′′(1+x2)P{s∣G∣>ρ}1/2C''(1+x^{2})\mathbb{P}\{\sqrt{s}|G|>\rho\}^{1/2} for s≤1s \le1, with C′′C'' depending only on C′C', by the Cauchy–Schwarz inequality. Letting s↓0s \downarrow0 and then ρ↓0\rho\downarrow0, the continuity of ff shows that Tm,sf→f\mathcal{T}_{m,s}f \to f locally uniformly as s↓0s \downarrow0 (for m>0m>0 because emfe^{mf} is continuous and positive), and hence g(s,x)→f(x0)g(s,x) \to f(x_{0}).

General order parameters. For γ∈U\gamma\in\mathcal{U} choose qˉ∈[Qγ,1)\bar{q} \in[Q_{\gamma},1). There are step order parameters γn\gamma_{n} with γn=1\gamma_{n}=1 on [qˉ,1][\bar{q},1] and ∥γn−γ∥L1≤1/n\lVert\gamma_{n}-\gamma\rVert_{L^{1}}\le1/n. Indeed, let 0=t0<⋯<tℓ=qˉ0=t_{0}<\cdots<t_{\ell}=\bar{q} have mesh at most 1/n1/n, and put γn(q)=γ(tj)\gamma_{n}(q)=\gamma(t_{j}) on [tj,tj+1)[t_{j},t_{j+1}) for j<ℓj<\ell and γn=1\gamma_{n}=1 on [qˉ,1][\bar{q},1]. Since γ\gamma is nondecreasing,

∫01∣γ−γn∣ dq≤∑j<ℓ(tj+1−tj)(γ(tj+1)−γ(tj))≤1n.\int_{0}^{1}\lvert\gamma-\gamma_{n}\rvert\,\mathrm{d}q \le \sum_{j<\ell}(t_{j+1}-t_{j})(\gamma(t_{j+1})-\gamma(t_{j})) \le \frac{1}{n}.

Further breakpoints may be added to the partition without affecting this bound. The same construction with qˉ=1\bar{q}=1 (and γn(1)=1\gamma_{n}(1)=1) approximates the distribution function of any probability measure on [0,1][0,1] by nondecreasing step functions with values in [0,1][0,1]. For general γ\gamma, and for the terminal data treated below (activations in class (A) or class (B), and the hard wall), the solution uγu_{\gamma} is defined as the limit of uγnu_{\gamma_{n}}. That the limit exists and does not depend on the approximating sequence follows from the Lipschitz estimates in γ\gamma: Lemma 2.3 for class (A), Lemma 2.4 for class (B), and Lemma 2.7 for the hard wall. The properties (19), (20) and (21) pass to these limits. For mixed pp-spin models, Auffinger and W.-K. Chen define the solution for general order parameters in the same way [ref-8].

The entropy and the cutoff. The definitions (4) and (5) involve a number qˉ∈[Qγ,1)\bar{q}\in[Q_{\gamma},1), and neither CS(γ)\mathrm{CS}(\gamma) nor uγ(⋅,⋅;κ)u_{\gamma}(\cdot,\cdot;\kappa) depends on this choice. We prove this here for the entropy and for step order parameters; the extension to general γ∈U\gamma\in\mathcal{U} is given after Lemma 2.7. For CS\mathrm{CS}, let Qγ≤qˉ<qˉ′<1Q_{\gamma}\le\bar{q}<\bar{q}'<1. Since λγ(q)=1−q\lambda_{\gamma}(q)=1-q on [Qγ,1][Q_{\gamma},1],

12∫qˉqˉ′dqλγ(q)+12log⁡(1−qˉ′)=12log⁡1−qˉ1−qˉ′+12log⁡(1−qˉ′)=12log⁡(1−qˉ).\frac{1}{2}\int_{\bar{q}}^{\bar{q}'}\frac{\mathrm{d}q}{\lambda_{\gamma}(q)} +\frac{1}{2}\log(1-\bar{q}') = \frac{1}{2}\log\frac{1-\bar{q}}{1-\bar{q}'} +\frac{1}{2}\log(1-\bar{q}') = \frac{1}{2}\log(1-\bar{q}).

Equivalently, since the integrand below vanishes on [Qγ,1][Q_{\gamma},1],

CS(γ)=12∫01(1λγ(q)−11−q)dq.(24)\mathrm{CS}(\gamma) = \frac{1}{2}\int_{0}^{1} \left( \frac{1}{\lambda_{\gamma}(q)} - \frac{1}{1-q} \right)\mathrm{d}q. \tag*{(24)}

For the hard wall, put

fqˉ(x)=log⁡Φ(x−κ1−qˉ)=log⁡P{x+1−qˉG≥κ},f_{\bar{q}}(x) = \log\Phi\left(\frac{x-\kappa}{\sqrt{1-\bar{q}}}\right) = \log\mathbb{P}\{x+\sqrt{1-\bar{q}}G\ge\kappa\},

the terminal datum (5) at time qˉ\bar{q}; its dependence on κ\kappa is suppressed. Let γ\gamma be a step order parameter and Qγ≤qˉ<qˉ′<1Q_{\gamma}\le\bar{q}<\bar{q}'<1. Since γ=1\gamma=1 on [qˉ,qˉ′][\bar{q},\bar{q}'], (22) with terminal time qˉ′\bar{q}' gives

T1,qˉ′−qˉfqˉ′(x)=log⁡P{x+qˉ′−qˉG′+1−qˉ′G≥κ}=fqˉ(x).\mathcal{T}_{1,\bar{q}'-\bar{q}}f_{\bar{q}'}(x) = \log\mathbb{P}\{x+\sqrt{\bar{q}'-\bar{q}}G' +\sqrt{1-\bar{q}'}G\ge\kappa\} = f_{\bar{q}}(x).

Thus the recursion from fqˉ′f_{\bar{q}'} at time qˉ′\bar{q}' produces fqˉf_{\bar{q}} at time qˉ\bar{q}, and the two definitions of uγ(q,x;κ)u_{\gamma}(q,x;\kappa) agree on [0,qˉ][0,\bar{q}]. The same computation with the hard constraint Hκ=log⁡1[κ,∞)H_{\kappa}=\log1_{[\kappa,\infty)} of §1.1 as terminal datum at time one gives T1,1−qˉHκ=fqˉT_{1,1-\bar q}H_{\kappa}=f_{\bar q}, so for every step order parameter γ\gamma and qˉ∈[Qγ,1)\bar q\in[Q_{\gamma},1),

uγ(q,x;Hκ)=uγ(q,x;κ)(q∈[0,qˉ]),hencePκ(γ)=αuγ(0,0;Hκ)+CS⁡(γ).(25)u_{\gamma}(q,x;H_{\kappa})=u_{\gamma}(q,x;\kappa)\qquad(q\in[0,\bar q]),\qquad\text{hence}\qquad\mathcal{P}_{\kappa}(\gamma)=\alpha u_{\gamma}(0,0;H_{\kappa})+\operatorname{CS}(\gamma). \tag*{(25)}

Taking qˉ=q\bar q=q in the definition, we obtain for every step order parameter

uγ(q,x;κ)=log⁡Φ(x−κ1−q),Qγ≤q<1.(26)u_{\gamma}(q,x;\kappa)=\log\Phi\left(\frac{x-\kappa}{\sqrt{1-q}}\right),\qquad Q_{\gamma}\leq q<1. \tag*{(26)}

Classes with a common cutoff. For 0≤r<10\leq r<1 let

Ur={γ∈U:γ=1 on [r,1]}={γ∈U:Qγ≤r}.\mathcal{U}_{r}=\{\gamma\in\mathcal{U}:\gamma=1\text{ on }[r,1]\}=\{\gamma\in\mathcal{U}:Q_{\gamma}\leq r\}.

The entropy is Lipschitz on each Ur\mathcal{U}_{r}. For γ1,γ2∈Ur\gamma_{1},\gamma_{2}\in\mathcal{U}_{r} use the common cutoff qˉ=r\bar q=r. On [0,r][0,r] we have λγi≥λγi(r)=1−r\lambda_{\gamma_i}\geq\lambda_{\gamma_i}(r)=1-r, and ∣λγ1(q)−λγ2(q)∣≤∥γ1−γ2∥L1|\lambda_{\gamma_1}(q)-\lambda_{\gamma_2}(q)|\leq\|\gamma_1-\gamma_2\|_{L^1}. Hence

∣CS⁡(γ1)−CS⁡(γ2)∣≤12∫0r∣λγ1−λγ2∣λγ1λγ2 dq≤r2(1−r)2∥γ1−γ2∥L1(γ1,γ2∈Ur).(27)|\operatorname{CS}(\gamma_1)-\operatorname{CS}(\gamma_2)|\leq\frac{1}{2}\int_{0}^{r}\frac{|\lambda_{\gamma_1}-\lambda_{\gamma_2}|}{\lambda_{\gamma_1}\lambda_{\gamma_2}}\,\mathrm{d}q\leq\frac{r}{2(1-r)^2}\|\gamma_1-\gamma_2\|_{L^1}\qquad(\gamma_1,\gamma_2\in\mathcal{U}_{r}). \tag*{(27)}

Each Ur\mathcal{U}_{r} is a compact metric space for the L1L^1 distance. The L1L^1 distance is a metric on Ur\mathcal{U}_{r}, because two right-continuous functions that agree almost everywhere agree on [0,1)[0,1). By Helly’s selection theorem, a sequence in Ur\mathcal{U}_{r} has a subsequence converging at every point of [0,1][0,1] to a nondecreasing function gg with values in [0,1][0,1] and g=1g=1 on [r,1][r,1]. The right-continuous version of gg lies in Ur\mathcal{U}_{r} and differs from gg at countably many points only, so the subsequence converges to it in L1L^1 by dominated convergence.

Continuity in the order parameter for smooth activations

We now show that γ↦uγ(0,0;U)\gamma\mapsto u_{\gamma}(0,0;U) is Lipschitz for the L1L^1 distance when UU belongs to class (A) or class (B). Three arguments need this. The upper bound of Chapter 3 is proved for step order parameters and is transferred to all of U\mathcal{U} within a class Ur\mathcal{U}_{r}. The lower bound of Chapter 4 compares two overlap laws in the Wasserstein distance, and these laws need not equal one near q=1q=1. The truncation argument of Chapter 5 needs an explicit Lipschitz constant. We therefore allow distribution functions γ,ν\gamma,\nu of arbitrary probability measures on [0,1][0,1], for which

W1(γ,ν)=∫01∣γ(q)−ν(q)∣ dq.W_{1}(\gamma,\nu)=\int_{0}^{1}|\gamma(q)-\nu(q)|\,\mathrm{d}q.

We use one pathwise bound for diffusions whose drift has linear growth. Let BB be a standard Brownian motion on [0,1][0,1] and B∗=sup⁡t≤1∣Bt∣B^{*}=\sup_{t\leq1}|B_t|, so that E(B∗)2≤4\mathbb{E}(B^{*})^2\leq4 by Doob’s inequality. If 0≤q≤T≤10\leq q\leq T\leq1 and dYs=β(s,Ys) ds+dBs\mathrm{d}Y_s=\beta(s,Y_s)\,\mathrm{d}s+\mathrm{d}B_s on [q,T][q,T] with Yq=xY_q=x and ∣β(s,y)∣≤a+K′∣y∣|\beta(s,y)| \le a + K'|y|, then ∣Ys∣≤∣x∣+2B∗+a+K′∫qs∣Yt∣ dt|Y_s| \le|x| + 2B^* + a + K'\int_q^s |Y_t|\,\mathrm{d}t, since T−q≤1T-q \le1, and Gronwall’s inequality gives

sup⁡s∈[q,T]∣Ys∣≤(∣x∣+2B∗+a)eK′.(28)\sup_{s \in[q,T]} |Y_s| \le(|x| + 2B^* + a)e^{K'}. \tag*{(28)}

The next lemma computes the first two derivatives of one Cole–Hopf step. It is the only place where the form of Tm,s\mathcal{T}_{m,s} enters the continuity estimates.

Lemma 2.1. Let f∈C2(R)f \in\mathrm{C}^2(\mathbb{R}) be bounded above with f′′f'' bounded, let m≥0m \ge0 and s>0s > 0, and put g=Tm,sfg = \mathcal{T}_{m,s}f. For x∈Rx \in\mathbb{R} let PxP_x be the law of Y=x+sGY = x + \sqrt{s}G tilted by emf(Y)e^{mf(Y)}, that is, dPx/dLaw⁡(Y)=emf(Y)/Eemf(Y)\mathrm{d}P_x/\mathrm{d}\operatorname{Law}(Y) = e^{mf(Y)}/\mathbb{E}e^{mf(Y)}. Then g∈C2(R)g \in\mathrm{C}^2(\mathbb{R}), g≤sup⁡fg \le\sup f, and

g′(x)=EPxf′(Y),g′′(x)=EPxf′′(Y)+mVar⁡Px(f′(Y)).(29)g'(x) = \mathbb{E}_{P_x}f'(Y), \qquad g''(x) = \mathbb{E}_{P_x}f''(Y) + m\operatorname{Var}_{P_x}\left(f'(Y)\right). \tag*{(29)}

Consequently the following hold.

(a) If ∣f′∣≤L|f'| \le L on R\mathbb{R}, then ∣g′∣≤L|g'| \le L and ∣g′′∣≤sup⁡∣f′′∣+mL2|g''| \le\sup|f''| + mL^2.

(b) If ff is concave and f′′≥−Kf'' \ge-K on R\mathbb{R}, then gg is concave and g′′≥−Kg'' \ge-K.

Proof. Since f′f' has at most linear growth, f′′f'' is bounded, and emf≤emsup⁡fe^{mf} \le e^{m\sup f}, the function emfe^{mf} and its first two derivatives mf′emfmf'e^{mf} and (mf′′+m2f′2)emf(mf'' + m^2f'^2)e^{mf}, evaluated at x+sGx + \sqrt{s}G, are dominated by C(1+∣G∣)2C(1+|G|)^2 for xx in a compact set, where CC depends on the compact set. Hence we may differentiate Eemf(x+sG)\mathbb{E}e^{mf(x+\sqrt{s}G)} twice under the expectation, and dividing by Eemf(Y)\mathbb{E}e^{mf(Y)} gives (29) for m>0m > 0. For m=0m = 0 we have Px=Law⁡(Y)P_x = \operatorname{Law}(Y), and the same argument applied to Ef(x+sG)\mathbb{E}f(x+\sqrt{s}G) gives g′=Ef′(Y)g' = \mathbb{E}f'(Y) and g′′=Ef′′(Y)g'' = \mathbb{E}f''(Y). Part (a) follows from (29) and Var⁡Pxf′(Y)≤L2\operatorname{Var}_{P_x}f'(Y) \le L^2. In part (b), g′′≥−Kg'' \ge-K by (29). If m=0m = 0, then gg is an average of concave functions. If m>0m > 0, put h(x,y)=exp⁡{mf(y)−(y−x)2/(2s)}h(x,y) = \exp\{mf(y) - (y-x)^2/(2s)\}. Its logarithm is concave in (x,y)(x,y), and ∫Rh(x,y) dy=2πs emg(x)\int_{\mathbb{R}} h(x,y)\,\mathrm{d}y = \sqrt{2\pi s}\,e^{mg(x)}. By Prékopa’s theorem [ref-108] (Theorem 6), the marginal of a log-concave function on R2\mathbb{R}^2 is log-concave, so mgmg is concave. □\square

For the hard-wall terminal datum, the formula for g′′g'' in (29) appears in [ref-54] (4.17). Part II, Lemma A.1 proves the same formulas for concave nondecreasing ff; here part (a) also covers data that are neither concave nor monotone, as in class (A). As in Part II, Lemma A.1, part (b) gives concavity one step at a time. Hence Lemma 2.2(i) holds for all step functions with values in [0,1][0,1], whereas Part II, Lemma A.2 and [ref-54] proof of (4.6) assume the coefficient nondecreasing.

The next lemma compares the solutions for two step coefficients with the same terminal datum, at a general terminal time TT. Part II, Lemma A.2 proves the identity (2.15) for a class of concave nondecreasing terminal data that contains the hard-wall datum.

Lemma 2.2. Let T∈(0,1]T \in(0,1] and let F∈C2(R)F \in C^{2}(\mathbb{R}) be bounded above with F′′F'' bounded. Let γ,ν:[0,T)→[0,1]\gamma,\nu:[0,T)\to[0,1] be step functions, let uγ=uγ(⋅,⋅;F)u_{\gamma}=u_{\gamma}(\cdot,\cdot;F) and uν=uν(⋅,⋅;F)u_{\nu}=u_{\nu}(\cdot,\cdot;F) be given by (17) with terminal datum FF at time TT, and put w=uγ−uνw=u_{\gamma}-u_{\nu}. Assume one of the following.

(A) ∣F′∣≤L|F'|\leq L on R\mathbb{R}.

(B) FF is concave and F′′≥−KF''\geq-K on R\mathbb{R} for a constant K≥0K\geq0. In this case put A=sup⁡F−F(0)+∣F′(0)∣+KA=\sup F-F(0)+|F'(0)|+K.

Then the following hold for all (q,x)∈[0,T]×R(q,x)\in[0,T]\times\mathbb{R}.

(i) In case (A), uγ(q,⋅)∈C2(R)u_{\gamma}(q,\cdot)\in C^{2}(\mathbb{R}) and ∣∂xuγ∣≤L|\partial_{x}u_{\gamma}|\leq L. In case (B), uγ(q,⋅)∈C2(R)u_{\gamma}(q,\cdot)\in C^{2}(\mathbb{R}) and

−K≤∂xxuγ≤0,F(0)−∣F′(0)∣∣x∣−K2(1+x2)≤uγ(q,x)≤sup⁡F,∣∂xuγ(q,x)∣≤A+K∣x∣.\begin{aligned} -K &\leq\partial_{xx}u_{\gamma}\leq0,\qquad F(0)-|F'(0)||x|-\frac{K}{2}(1+x^{2})\leq u_{\gamma}(q,x)\leq\sup F,\\ |\partial_{x}u_{\gamma}(q,x)|&\leq A+K|x|. \end{aligned}

The same holds for uνu_{\nu}.

(ii) Let BB be a standard Brownian motion and let YY solve dYs=β(s,Ys) ds+dBs\mathrm{d}Y_{s}=\beta(s,Y_{s})\,\mathrm{d}s+\mathrm{d}B_{s} on [q,T][q,T] with Yq=xY_{q}=x, where β=γ2(∂xuγ+∂xuν)\beta=\frac{\gamma}{2}(\partial_{x}u_{\gamma}+\partial_{x}u_{\nu}). This equation has a unique strong solution, and

w(q,x)=12∫qT(γ−ν)(s) E[(∂xuν)2(s,Ys)] ds.(30)w(q,x)=\frac{1}{2}\int_{q}^{T}(\gamma-\nu)(s)\,\mathbb{E}\left[(\partial_{x}u_{\nu})^{2}(s,Y_{s})\right]\,\mathrm{d}s. \tag*{(30)}

(iii) In case (A), ∣w(q,x)∣≤L22∫qT∣γ−ν∣ ds|w(q,x)|\leq\frac{L^{2}}{2}\int_{q}^{T}|\gamma-\nu|\,\mathrm{d}s. In case (B),

∣w(q,x)∣≤CA,K(1+x2)∫qT∣γ−ν∣ ds,CA,K=A2+3K2e2K(16+A2).|w(q,x)|\leq C_{A,K}(1+x^{2})\int_{q}^{T}|\gamma-\nu|\,\mathrm{d}s,\qquad C_{A,K}=A^{2}+3K^{2}e^{2K}(16+A^{2}).

Proof. Refining the partitions, we may assume that γ\gamma and ν\nu are constant on the intervals [tj,tj+1)[t_{j},t_{j+1}) of a common partition 0=t0<⋯<tn+1=T0=t_{0}<\cdots<t_{n+1}=T. We divide the proof into four steps.

Step 1. Proof of (i). For q∈[tj,tj+1)q\in[t_{j},t_{j+1}), uγ(q,⋅)u_{\gamma}(q,\cdot) is Tγ(tj),tj+1−q\mathcal{T}_{\gamma(t_{j}),t_{j+1}-q} applied to uγ(tj+1,⋅)u_{\gamma}(t_{j+1},\cdot). Induction over the intervals, starting from the last, and Lemma 2.1 give the following: in case (A), uγ(q,⋅)∈C2u_{\gamma}(q,\cdot)\in C^{2}, ∣∂xuγ∣≤L|\partial_{x}u_{\gamma}|\leq L, and ∣∂xxuγ∣≤sup⁡∣F′′∣+(n+1)L2|\partial_{xx}u_{\gamma}|\leq\sup|F''|+(n+1)L^{2} on [0,T]×R[0,T]\times\mathbb{R}; in case (B), uγ(q,⋅)∈C2u_{\gamma}(q,\cdot)\in C^{2} is concave with ∂xxuγ≥−K\partial_{xx}u_{\gamma}\geq-K. In case (B), uγ≤sup⁡Fu_{\gamma}\leq\sup F, and (21) together with F(y)≥F(0)+F′(0)y−K2y2F(y)\geq F(0)+F'(0)y-\frac{K}{2}y^{2} gives

uγ(q,x)≥EF(x+T−q G)≥F(0)+F′(0)x−K2(x2+T−q),u_{\gamma}(q,x)\geq\mathbb{E}F\left(x+\sqrt{T-q}\,G\right)\geq F(0)+F'(0)x-\frac{K}{2}(x^{2}+T-q),

which is the asserted lower bound, since T−q≤1T-q\leq1. For x∈{−1,0,1}x\in\{-1,0,1\} the right side is at least F(0)−∣F′(0)∣−KF(0)-|F'(0)|-K. By concavity,

uγ(q,1)−uγ(q,0)≤∂xuγ(q,0)≤uγ(q,0)−uγ(q,−1),u_{\gamma}(q,1)-u_{\gamma}(q,0)\leq\partial_{x}u_{\gamma}(q,0)\leq u_{\gamma}(q,0)-u_{\gamma}(q,-1),

so ∣∂xuγ(q,0)∣≤sup⁡F−F(0)+∣F′(0)∣+K=A|\partial_x u_\gamma(q,0)| \le\sup F-F(0)+|F'(0)|+K=A, and ∣∂xxuγ∣≤K|\partial_{xx}u_\gamma| \le K gives ∣∂xuγ(q,x)∣≤A+K∣x∣|\partial_xu_\gamma(q,x)| \le A+K|x|.

Step 2. The equation for ww. By §2.1, uγu_\gamma and uνu_\nu are continuous on [0,T]×R[0,T]\times\mathbb{R}, and on each open interval (tj,tj+1)(t_j,t_{j+1}) they are smooth and solve (2.8) with the constant coefficients γ\gamma and ν\nu. Subtracting the two equations and using

γ2(∂xuγ)2−ν2(∂xuν)2=β ∂xw+γ−ν2(∂xuν)2,\frac{\gamma}{2}(\partial_xu_\gamma)^2-\frac{\nu}{2}(\partial_xu_\nu)^2 =\beta\,\partial_xw+\frac{\gamma-\nu}{2}(\partial_xu_\nu)^2,

we obtain on each (tj,tj+1)×R(t_j,t_{j+1})\times\mathbb{R}

∂qw+12∂xxw+β ∂xw=−γ−ν2(∂xuν)2,w(T,⋅)=0.(31)\partial_qw+\frac{1}{2}\partial_{xx}w+\beta\,\partial_xw =-\frac{\gamma-\nu}{2}(\partial_xu_\nu)^2,\qquad w(T,\cdot)=0. \tag*{(31)}

By Step 1, β\beta is Lipschitz in yy uniformly in ss, and ∣β(s,y)∣≤a+K′∣y∣|\beta(s,y)|\le a+K'|y| with (a,K′)=(L,0)(a,K')=(L,0) in case (A) and (a,K′)=(A,K)(a,K')=(A,K) in case (B). It is Borel, since ∂xuγ\partial_xu_\gamma and ∂xuν\partial_xu_\nu are pointwise limits of difference quotients of continuous functions. Hence the equation for YY has a unique strong solution [ref-79], and (2.13) gives Esup⁡s∈[q,T]Ys2<∞\mathbb{E}\sup_{s\in[q,T]}Y_s^2<\infty. Moreover ∣∂xw(s,y)∣≤2(a+K′∣y∣)|\partial_xw(s,y)|\le2(a+K'|y|) by (i), and ∣w(s,y)∣≤C(1+y2)|w(s,y)|\le C(1+y^2) for a constant CC, because uγu_\gamma and uνu_\nu lie between the lower bound in (2.6) and sup⁡F\sup F (in case (A), F(y)≥F(0)−L∣y∣F(y)\ge F(0)-L|y|).

Step 3. Itô’s formula and the passage through the breakpoints. Fix jj with tj+1>qt_{j+1}>q, put tj′=max⁡{q,tj}t'_j=\max\{q,t_j\}, and let tj′<s1<s2<tj+1t'_j<s_1<s_2<t_{j+1}. On [s1,s2]×R[s_1,s_2]\times\mathbb{R} the function ww is smooth, so Itô’s formula applies to w(s,Ys)w(s,Y_s) on [s1,s2][s_1,s_2]. The stochastic integral ∫s1s2∂xw(s,Ys) dBs\int_{s_1}^{s_2}\partial_xw(s,Y_s)\,\mathrm{d}B_s has mean zero, because E∫s1s2(∂xw)2(s,Ys) ds≤4E(a+K′sup⁡s∣Ys∣)2<∞\mathbb{E}\int_{s_1}^{s_2}(\partial_xw)^2(s,Y_s)\,\mathrm{d}s\le4\mathbb{E}(a+K'\sup_s|Y_s|)^2<\infty. Taking expectations and using (31),

Ew(s2,Ys2)−Ew(s1,Ys1)=−12∫s1s2(γ−ν)(s) E(∂xuν)2(s,Ys) ds.\mathbb{E}w(s_2,Y_{s_2})-\mathbb{E}w(s_1,Y_{s_1}) =-\frac{1}{2}\int_{s_1}^{s_2}(\gamma-\nu)(s)\,\mathbb{E}(\partial_xu_\nu)^2(s,Y_s)\,\mathrm{d}s.

Now let s1↓tj′s_1\downarrow t'_j and s2↑tj+1s_2\uparrow t_{j+1}. The function ww is continuous on [0,T]×R[0,T]\times\mathbb{R} and YY has continuous paths. Since ∣w∣≤C(1+y2)|w|\le C(1+y^2) and (∂xuν)2(s,y)≤(a+K′∣y∣)2(\partial_xu_\nu)^2(s,y)\le(a+K'|y|)^2 are dominated along YY by C′(1+sup⁡sYs2)C'(1+\sup_sY_s^2), which is integrable, dominated convergence gives the same identity with s1=tj′s_1=t'_j and s2=tj+1s_2=t_{j+1}. Summing over jj, and using Yq=xY_q=x and w(T,⋅)=0w(T,\cdot)=0, gives (2.15).

Step 4. Proof of (iii). In case (A), (∂xuν)2≤L2(\partial_xu_\nu)^2\le L^2 in (2.15). In case (B), (∂xuν)2(s,y)≤2A2+2K2y2(\partial_xu_\nu)^2(s,y)\le2A^2+2K^2y^2 by (i). By (2.13) with (a,K′)=(A,K)(a,K')=(A,K) and E(B∗)2≤4\mathbb{E}(B^*)^2\le4,

EYs2≤e2KE(∣x∣+2B∗+A)2≤3e2K(x2+16+A2),\mathbb{E}Y_s^2\le e^{2K}\mathbb{E}(|x|+2B^*+A)^2 \le3e^{2K}(x^2+16+A^2),

so

E(∂xuν)2(s,Ys)≤2A2+6K2e2K(x2+16+A2)≤2CA,K(1+x2),\mathbb{E}(\partial_xu_\nu)^2(s,Y_s) \le2A^2+6K^2e^{2K}(x^2+16+A^2) \le2C_{A,K}(1+x^2),

and (2.15) gives the bound.

The identity (30) is the Feynman–Kac formula of Jagannath and Tobasco [ref-76], Lemma 14, and part (iii) is the analogue of Guerra’s bound (see the beginning of this chapter). Case (B) covers concave terminal data whose derivative grows linearly, such as the hard-wall datum and the activations in class (B).

For activations in class (A), the Lipschitz estimate holds in the uniform norm and at every time; this is the form used in the upper bound.

Lemma 2.3. Let UU belong to class (A) and put L=∥U′∥∞L=\lVert U'\rVert_{\infty}. Then for all q∈[0,1]q\in[0,1] and all distribution functions γ1,γ2\gamma_{1},\gamma_{2} of probability measures on [0,1][0,1], in particular for all γ1,γ2∈U\gamma_{1},\gamma_{2}\in\mathcal{U},

sup⁡x∈R∣uγ1(q,x;U)−uγ2(q,x;U)∣≤L22∫q1∣γ1−γ2∣ ds.(32)\sup_{x\in\mathbb{R}}\left|u_{\gamma_{1}}(q,x;U)-u_{\gamma_{2}}(q,x;U)\right| \le\frac{L^{2}}{2}\int_{q}^{1}\left|\gamma_{1}-\gamma_{2}\right|\,\mathrm{d}s. \tag*{(32)}

If moreover γ1,γ2∈Uqˉ\gamma_{1},\gamma_{2}\in\mathcal{U}_{\bar q} for some qˉ<1\bar q<1, then

∣PU(γ1)−PU(γ2)∣≤(αL22+qˉ2(1−qˉ)2)∥γ1−γ2∥L1.(33)\left|\mathcal{P}_{U}(\gamma_{1})-\mathcal{P}_{U}(\gamma_{2})\right| \le \left(\frac{\alpha L^{2}}{2}+\frac{\bar q}{2(1-\bar q)^{2}}\right) \lVert\gamma_{1}-\gamma_{2}\rVert_{L^{1}}. \tag*{(33)}

Proof. For step functions, (32) is Lemma 2.2(iii) in case (A), with T=1T=1 and F=UF=U. For general γ1,γ2\gamma_{1},\gamma_{2}, let γi,n\gamma_{i,n} be step approximations as in §2.1. By (32) for step functions, (uγi,n(q,⋅;U))n(u_{\gamma_{i,n}}(q,\cdot;U))_{n} is Cauchy in the uniform norm, and two approximating sequences of γi\gamma_{i} have the same limit, since they can be interleaved. This proves that uγiu_{\gamma_{i}} is well defined, and (32) passes to the limit. Finally, (33) is α\alpha times (32) at (q,x)=(0,0)(q,x)=(0,0) plus (27) with r=qˉr=\bar q. □\square

The next lemma gives the Lipschitz estimate for both classes at the origin, with an explicit constant.

Lemma 2.4. Let UU belong to class (A) or class (B). If UU belongs to class (A), put CUA=L2/2C_{U}^{\mathrm{A}}=L^{2}/2 with L=∥U′∥∞L=\lVert U'\rVert_{\infty}. If UU belongs to class (B), put CUB=CA,KC_{U}^{\mathrm{B}}=C_{A,K}, the constant of Lemma 2.2(iii) with

K=∥U′′∥∞,A=sup⁡U−U(0)+∣U′(0)∣+K.K=\lVert U''\rVert_{\infty},\qquad A=\sup U-U(0)+\left|U'(0)\right|+K.

Let CUC_{U} be the smallest of the numbers CUA,CUBC_{U}^{\mathrm{A}},C_{U}^{\mathrm{B}} that are defined. Then, for any two probability measures γ,ν\gamma,\nu on [0,1][0,1],

∣uγ(0,0;U)−uν(0,0;U)∣≤CUW1(γ,ν).(34)\left|u_{\gamma}(0,0;U)-u_{\nu}(0,0;U)\right| \le C_{U}W_{1}(\gamma,\nu). \tag*{(34)}

Consequently, for r∈[0,1)r\in[0,1) and γ,ν∈Ur\gamma,\nu\in\mathcal{U}_{r},

∣PU(γ)−PU(ν)∣≤(αCU+r2(1−r)2)∥γ−ν∥L1.(35)\left|\mathcal{P}_{U}(\gamma)-\mathcal{P}_{U}(\nu)\right| \le \left(\alpha C_{U}+\frac{r}{2(1-r)^{2}}\right) \lVert\gamma-\nu\rVert_{L^{1}}. \tag*{(35)}

Proof. If UU belongs to class (A), (34) with the constant CUAC_U^{\mathrm{A}} is (32) at (q,x)=(0,0)(q,x)=(0,0). Let UU belong to class (B). Then F=UF=U satisfies case (B) of Lemma 2.2 with T=1T=1 and the constants KK and AA above. For step functions γ,ν\gamma,\nu with values in [0,1][0,1], Lemma 2.2(iii) gives

∣uγ(q,x;U)−uν(q,x;U)∣≤CA,K(1+x2)∫q1∣γ−ν∣ dsfor all (q,x)∈[0,1]×R.\left|u_{\gamma}(q,x;U)-u_{\nu}(q,x;U)\right| \le C_{A,K}(1+x^{2})\int_{q}^{1}|\gamma-\nu|\,\mathrm{d}s \qquad\text{for all }(q,x)\in[0,1]\times\mathbb{R}.

For general γ\gamma, approximate it by step functions as in §2.1. By this estimate the values at (q,x)(q,x) form a Cauchy sequence, locally uniformly in (q,x)(q,x), and the limit does not depend on the approximating sequence. This defines uγ(q,x;U)u_{\gamma}(q,x;U) (for UU in both classes it agrees with the limit in Lemma 2.3), and the estimate passes to the limit. At (q,x)=(0,0)(q,x)=(0,0) it gives (34) with the constant CUBC_U^{\mathrm{B}}. Finally, (35) follows from (34), the identity W1(γ,ν)=∥γ−ν∥L1W_1(\gamma,\nu)=\|\gamma-\nu\|_{L^1} and (27). □\square

The hard-wall functional at a finite cutoff

We now study Pκ(γ)\mathcal{P}_{\kappa}(\gamma) as a deterministic functional of the order parameter. The top QγQ_{\gamma} of the support is not continuous for the L1L^1 distance, since adding a small mass near one moves it. We therefore work on the classes Ur\mathcal{U}_r, 0≤r<10\le r<1, and evaluate every γ∈Ur\gamma\in\mathcal{U}_r with the common terminal time rr. All constants below depend on rr and on a compact interval II containing κ\kappa, but not on the number of steps of a step order parameter. None of the results concerns the limit r→1r\to1. Part II, §§2.1–2.5 and Appendix A, which do not use Part I, prove the results of this section other than the martingale identity (43), in a stronger form with constants that are uniform for ∣κ∣≤K0|\kappa|\le K_0, and we derive them from there. Part II defines the hard-wall solution for step order parameters by the same recursion, with the terminal datum frf_r defined below, and for general order parameters by the same approximation within Ur\mathcal{U}_r (Part II, Definition 2.2(iii)), so the two definitions agree.

Gaussian quantities. We use the functions

R(z)=ϕ(z)Φ(z),Rˉ(z)=z+R(z),V(z)=R(z)Rˉ(z)(z∈R).R(z)=\frac{\phi(z)}{\Phi(z)},\qquad \bar{R}(z)=z+R(z),\qquad V(z)=R(z)\bar{R}(z) \qquad(z\in\mathbb{R}).

Lemma 2.5. The following statements hold.

(i) For t>0t>0,

tϕ(t)1+t2<Φ(−t)<ϕ(t)t,equivalentlyt<R(−t)<t+1t.\frac{t\phi(t)}{1+t^{2}}<\Phi(-t)<\frac{\phi(t)}{t}, \qquad\textit{equivalently}\qquad t<R(-t)<t+\frac{1}{t}.

(ii) R′=−VR'=-V and Rˉ′=1−V\bar{R}'=1-V.

(iii) For z∈Rz\in\mathbb{R} let TzT_z have the law of z+Gz+G conditioned on the event {z+G≥0}\{z+G\ge0\}, that is, the law on [0,∞)[0,\infty) with density proportional to ezt−t2/2e^{zt-t^{2}/2}. Then ETz=Rˉ(z)\mathbb{E}T_z=\bar{R}(z) and Var⁡(Tz)=1−V(z)\operatorname{Var}(T_z)=1-V(z). Consequently Rˉ>0\bar{R}>0 and 0<V<10<V<1.

Proof. The functions RR, Rˉ\bar R, VV, and the variable TZT_Z are those of Part II, §2.1. Part (i) is Part II, Lemma 2.1(b), and parts (ii) and (iii) are contained in Part II, Lemma 2.1(a). □\square

Part (i) consists of Gordon’s bounds for the Mills ratio [ref-65], and the bound V<1V<1 in part (iii) is due to Sampford [ref-113]. El Alaoui and Sellke collect these facts, in terms of the inverse Mills ratio, in [ref-54].

The terminal datum. Fix r∈[0,1)r\in[0,1). For γ∈Ur\gamma\in\mathcal U_r, the hard-wall solution uγ(⋅,⋅;κ)u_\gamma(\cdot,\cdot;\kappa) on [0,r]×R[0,r]\times\mathbb R is given by the recursion (17) with terminal time rr and terminal datum

fr(x)=log⁡Φ(x−κ1−r),f_r(x)=\log\Phi\left(\frac{x-\kappa}{\sqrt{1-r}}\right),

by (5) with qˉ=r\bar q=r; for step order parameters this is the definition, and for general γ∈Ur\gamma\in\mathcal U_r it is completed in Lemma 2.7. Using the same terminal time for all of Ur\mathcal U_r avoids the discontinuity of γ↦Qγ\gamma\mapsto Q_\gamma. With z=(x−κ)/1−rz=(x-\kappa)/\sqrt{1-r}, Lemma 2.5 gives

fr(x)≤0,fr′(x)=R(z)1−r>0,0<−fr′′(x)=V(z)1−r<11−r.(36)f_r(x)\le0,\qquad f_r'(x)=\frac{R(z)}{\sqrt{1-r}}>0,\qquad0<-f_r''(x)=\frac{V(z)}{1-r}<\frac{1}{1-r}. \tag*{(36)}

Every derivative of frf_r has polynomial growth, uniformly for κ\kappa in a compact interval; this gives the case r=0r=0 of Lemma 2.6(v). Indeed, R′=−zR−R2R'=-zR-R^2, so every derivative of RR is a polynomial in zz and RR, and 0<R(z)≤R(0)+∣z∣0<R(z)\le R(0)+|z| because −1<R′<0-1<R'<0 by Lemma 2.5(ii),(iii). Throughout the rest of this section, II is a compact interval, κ∈I\kappa\in I, K=(1−r)−1K=(1-r)^{-1}, and constants depend only on rr and II unless stated otherwise.

Notation. The following abbreviations are used in this section and in Chapter 6. We fix κ∈R\kappa\in\mathbb R and γ∈U\gamma\in\mathcal U, and write

Q=Qγ,δ=1−Q,λ=λγ,u=uγ(⋅,⋅;κ),b=∂xu,c=−∂xxu,v=λc.\begin{aligned} Q&=Q_\gamma,\qquad\delta=1-Q,\qquad\lambda=\lambda_\gamma,\qquad u=u_\gamma(\cdot,\cdot;\kappa),\\ b&=\partial_xu,\qquad c=-\partial_{xx}u,\qquad v=\lambda c. \end{aligned}

Thus λ(q)≥λ(Q)=δ>0\lambda(q)\ge\lambda(Q)=\delta>0 for q∈[0,Q]q\in[0,Q], and uu is defined on [0,1)×R[0,1)\times\mathbb R by (26). We write

L=∂q+12∂xx+γb ∂x\mathcal L=\partial_q+\frac{1}{2}\partial_{xx}+\gamma b\,\partial_x

for the generator of the diffusion

dXq=γ(q)b(q,Xq) dq+dBq,X0=0,(37)\mathrm dX_q=\gamma(q)b(q,X_q)\,\mathrm dq+\mathrm dB_q,\qquad X_0=0, \tag*{(37)}

where BB is a standard Brownian motion. For γ∈Ur\gamma\in\mathcal U_r we consider XX on [0,r][0,r], and we write Xs,xX^{s,x} for the solution of the equation in (37) on [s,r][s,r] with Xss,x=xX_s^{s,x}=x; it exists and is unique by Lemma 2.6(iii) for step order parameters and by Lemma 2.7 in general. On [0,Q][0,Q], XX is the Parisi diffusion (8) of §1.1, and the processes defined with two admissible values of rr agree on their common interval. It is also the diffusion of Part II, (2.14). When two order parameters γ,γ~\gamma,\widetilde{\gamma} are compared, we indicate the order parameter as a subscript, as in uγ~,bγu_{\widetilde{\gamma}},b_{\gamma}, or Xγ~X^{\widetilde{\gamma}}. With this notation, the hard-wall functional (1.6) with cutoff qˉ=Q\bar q=Q reads

Pκ(γ)=αu(0,0)+12∫0Qdqλ(q)+12log⁡δ=αu(0,0)+12∫01(1λ(q)−11−q)dq,(38)\mathcal{P}_{\kappa}(\gamma)=\alpha u(0,0)+\frac{1}{2}\int_{0}^{Q}\frac{\mathrm{d}q}{\lambda(q)}+\frac{1}{2}\log\delta =\alpha u(0,0)+\frac{1}{2}\int_{0}^{1}\left(\frac{1}{\lambda(q)}-\frac{1}{1-q}\right)\mathrm{d}q, \tag*{(38)}

the second form by (2.9), and P∗(κ)=inf⁡γ∈UPκ(γ)\mathcal{P}_{*}(\kappa)=\inf_{\gamma\in\mathcal{U}}\mathcal{P}_{\kappa}(\gamma).

A priori bounds and continuity. The next lemma collects the bounds on the solution and the diffusion that justify Itô’s formula and the Feynman–Kac formula in the rest of this part.

Lemma 2.6. Fix r∈[0,1)r\in[0,1) and a compact interval II. There is a constant CC, and for each p≥1p\geq1 and k≥1k\geq1 there are constants CpC_p and Ck′C'_k, depending only on rr, II, and pp or kk, such that for every κ∈I\kappa\in I and every step order parameter γ∈Ur\gamma\in\mathcal{U}_r the following hold on [0,r]×R[0,r]\times\mathbb{R}.

(i) −C(1+x2)≤u(q,x)≤0-C(1+x^2)\leq u(q,x)\leq0.

(ii) On each open interval where γ\gamma is constant, uu is smooth and

Lb=0,Lc=γc2.\mathcal{L}b=0,\qquad\mathcal{L}c=\gamma c^2.

(iii) 0≤c≤(1−r)−10\leq c\leq(1-r)^{-1} and ∣b(q,x)∣≤C(1+∣x∣)|b(q,x)|\leq C(1+|x|). In particular the drift γb\gamma b in (2.22) is Lipschitz in xx with constant (1−r)−1(1-r)^{-1} and has linear growth, so the equation in (2.22) has a unique strong solution Xs,xX^{s,x} from every starting point (s,x)∈[0,r]×R(s,x)\in[0,r]\times\mathbb{R}.

(iv) Esup⁡t∈[s,r]∣Xts,x∣p≤Cp(1+∣x∣)p\mathbb{E}\sup_{t\in[s,r]}|X_t^{s,x}|^p\leq C_p(1+|x|)^p; in particular Esup⁡q≤r∣Xq∣p≤Cp\mathbb{E}\sup_{q\leq r}|X_q|^p\leq C_p.

(v) ∣∂xku(q,x)∣≤Ck′(1+∣x∣)Ck′|\partial_x^k u(q,x)|\leq C'_k(1+|x|)^{C'_k}, and ∂xku\partial_x^k u is continuous on [0,r]×R[0,r]\times\mathbb{R}.

Proof. For r=0r=0 the statements concern u=f0u=f_0 only and follow from (2.21) and the polynomial growth of the derivatives of f0f_0 noted after it; let r>0r>0. Every statement is contained in Part II, Chapter 2, applied with K0=max⁡κ∈I∣κ∣K_0=\max_{\kappa\in I}|\kappa|; the constants there depend only on r,K0r,K_0, and pp or kk. Part (i) and the bound on bb in (iii) are Part II, Lemma 2.4(b). By Part II, Lemma 2.4(c), 0<c<1/λ≤(1−r)−10<c<1/\lambda\leq(1-r)^{-1} on [0,r]×R[0,r]\times\mathbb{R}, since λ≥λ(r)=1−r\lambda\geq\lambda(r)=1-r there. Hence ∣∂x(γb)∣=γc≤(1−r)−1|\partial_x(\gamma b)|=\gamma c\leq(1-r)^{-1}, and the drift γb\gamma b has linear growth. Strong existence and uniqueness for (2.22) are stated in Part II after (2.14). Part (ii) is Part II, Lemma 2.4(a) and Lemma 2.9(a). Part (iv) is Part II, Lemma 2.6(a), since 1+∣x∣p≤(1+∣x∣)p1+|x|^p\leq(1+|x|)^p for p≥1p\geq1. Part (v) is Part II, Lemma 2.4(a),(b): the bound there is Ck(1+∣x∣pk)≤2Ck(1+∣x∣)pkC_k(1+|x|^{p_k})\leq2C_k(1+|x|)^{p_k}, so one may take Ck′=max⁡{2Ck,pk}C'_k=\max\{2C_k,p_k\}.

Parts (i) and (iii) correspond to the bounds [ref-54], (4.4)–(4.6), and part (v) is a weaker form of [ref-54], (4.7), with polynomial growth in place of boundedness.

The next lemma is the hard-wall counterpart of Lemma 2.4. It extends the solution, its derivatives, and the diffusion to general order parameters, and it is the source of the compactness and continuity used in Chapter 6. For mixed pp-spin models, the convergence of all spatial derivatives of the solution as the order parameter converges is [ref-8], Proposition 1.

Lemma 2.7. Fix r∈[0,1)r \in[0,1) and a compact interval II. There is a constant CC depending only on rr and II, and for each p≥1p \ge1 a constant CpC_p depending only on rr, II, and pp, such that for all κ∈I\kappa\in I and all step order parameters γ,γ~∈Ur\gamma,\widetilde{\gamma} \in\mathcal{U}_r, with ε=∥γ~−γ∥L1\varepsilon=\lVert\widetilde{\gamma}-\gamma\rVert_{L^1};

∣uγ~(q,x)−uγ(q,x)∣≤Cε(1+x2),∣bγ~(q,x)−bγ(q,x)∣≤Cε(1+∣x∣)(39)\left|u_{\widetilde{\gamma}}(q,x)-u_{\gamma}(q,x)\right| \le C\varepsilon(1+x^2), \qquad\left|b_{\widetilde{\gamma}}(q,x)-b_{\gamma}(q,x)\right| \le C\varepsilon(1+|x|) \tag*{(39)}

on [0,r]×R[0,r]\times\mathbb{R}. For every k≥2k \ge2 and A0>0A_0>0,

sup⁡{∣∂xkuγ~(q,x)−∂xkuγ(q,x)∣:q∈[0,r], ∣x∣≤A0}⟶0as ε→0,\sup\left\{\left|\partial_x^k u_{\widetilde{\gamma}}(q,x)-\partial_x^k u_{\gamma}(q,x)\right|:q\in[0,r],\,|x|\le A_0\right\}\longrightarrow0 \qquad\text{as }\varepsilon\to0,

uniformly in κ∈I\kappa\in I and in the step order parameters. When Xγ~X^{\widetilde{\gamma}} and XγX^\gamma start from the same point (s,x)∈[0,r]×R(s,x)\in[0,r]\times\mathbb{R} and are driven by the same Brownian motion,

Esup⁡t∈[s,r]∣Xtγ~−Xtγ∣p≤Cpεp(1+∣x∣)p.(40)\mathbb{E}\sup_{t\in[s,r]}\left|X_t^{\widetilde{\gamma}}-X_t^\gamma\right|^p \le C_p\varepsilon^p(1+|x|)^p. \tag*{(40)}

Consequently the solution uu, its spatial derivatives, and the diffusion XX extend from step order parameters to every γ∈Ur\gamma\in\mathcal{U}_r by L1L^1 approximation within Ur\mathcal{U}_r, and the bounds of Lemma 2.6(i), (iii)–(v) and of this lemma hold for all γ,γ~∈Ur\gamma,\widetilde{\gamma}\in\mathcal{U}_r. If γ∈Ur\gamma\in\mathcal{U}_r equals a constant mm on an interval [a,a′)⊂[0,r][a,a')\subset[0,r], then u(q,⋅)=Tm,a′−qu(a′,⋅)u(q,\cdot)=T_{m,a'-q}u(a',\cdot) for q∈[a,a′)q\in[a,a'), and uu is smooth and satisfies Lb=0\mathcal{L}b=0 and Lc=γc2\mathcal{L}c=\gamma c^2 on (a,a′)×R(a,a')\times\mathbb{R}. Finally,

∣Pκ(γ~)−Pκ(γ)∣≤(αC+r2(1−r)2)∥γ~−γ∥L1(γ,γ~∈Ur, κ∈I).(41)\left|\mathcal{P}_\kappa(\widetilde{\gamma})-\mathcal{P}_\kappa(\gamma)\right| \le \left(\alpha C+\frac{r}{2(1-r)^2}\right) \lVert\widetilde{\gamma}-\gamma\rVert_{L^1} \qquad (\gamma,\widetilde{\gamma}\in\mathcal{U}_r,\ \kappa\in I). \tag*{(41)}

In particular Pκ\mathcal{P}_\kappa is continuous on the compact set Ur\mathcal{U}_r, uniformly for κ∈I\kappa\in I, and attains its minimum over Ur\mathcal{U}_r.

Proof. For r=0r=0 the class U0\mathcal{U}_0 has one element; let r>0r>0. We apply the results of Part II, Chapter 2 with K0=max⁡κ∈I∣κ∣K_0=\max_{\kappa\in I}|\kappa|; they hold for all γ,γ~∈Ur\gamma,\widetilde{\gamma}\in\mathcal{U}_r, and their constants depend only on rr, K0K_0, and pp or kk. The solution for general γ∈Ur\gamma\in\mathcal{U}_r is the limit of the solutions for step approximations in Ur\mathcal{U}_r (Part II, Definition 2.2(iii)). By Part II, Lemma 2.4(f), for every k≥0k\ge0,

∣∂xkuγ~(q,x)−∂xkuγ(q,x)∣≤Ckε(1+∣x∣pk)on [0,r]×R,\left|\partial_x^k u_{\widetilde{\gamma}}(q,x)-\partial_x^k u_\gamma(q,x)\right| \le C_k\varepsilon(1+|x|^{p_k}) \qquad\text{on }[0,r]\times\mathbb{R},

with p0=2p_0=2 and p1=1p_1=1. The cases k=0,1k=0,1 give (39), the cases k≥2k\ge2 give the convergence of the higher derivatives, and the derivatives of uγu_\gamma are the limits of those of its step approximations. The bound (40) is Part II, Lemma 2.6(b); in particular the diffusions of step approximations converge to XγX^\gamma. Part II, Lemma 2.4(a)–(c) and Lemma 2.6(a) hold for general γ∈Ur\gamma\in\mathcal{U}_r, so the proof of Lemma 2.6 gives its parts (i) and (iii)–(v) for all γ∈Ur\gamma\in\mathcal{U}_{r}. If γ=m\gamma=m on [a,a′)[a,a'), then u(q,⋅)=Tm,a′−qu(a′,⋅)u(q,\cdot)=\mathcal{T}_{m,a'-q}u(a',\cdot) for q∈[a,a′)q\in[a,a') by Part II, Lemma 2.5(a), and uu is smooth and satisfies Lb=0\mathcal{L}b=0 and Lc=γc2\mathcal{L}c=\gamma c^{2} on (a,a′)×R(a,a')\times\mathbb{R} by Part II, Lemma 2.4(a) and Lemma 2.9(a). Finally, the first bound in (39) at (0,0)(0,0) controls the energy term of Pκ\mathcal{P}_{\kappa}, and (27) controls the entropy term; this proves (41). Compactness of Ur\mathcal{U}_{r} was shown in §2.1. □\square

We record a consequence of Lemma 2.7. If γ∈Ur⊂Ur′\gamma\in\mathcal{U}_{r}\subset\mathcal{U}_{r'} with r<r′<1r<r'<1, then the step approximations of γ\gamma in Ur\mathcal{U}_{r} also lie in Ur′\mathcal{U}_{r'}, and for them the solutions computed with the cutoffs rr and r′r' agree on [0,r][0,r] by the discussion before (25). Passing to the limit, uγ(⋅,⋅ ;κ)u_{\gamma}(\cdot,\cdot\,;\kappa) and Pκ(γ)\mathcal{P}_{\kappa}(\gamma) do not depend on the cutoff for any γ∈U\gamma\in\mathcal{U}, and (26) holds for every γ∈U\gamma\in\mathcal{U}. This is also Part II, Lemma 2.3.

We next record the signs of bb, cc, and 1−v1-v, which sharpen Lemma 2.6(iii) to strict inequalities, and the martingale structure along the diffusion. They are the starting point of the analysis in Chapter 6.

Lemma 2.8. Let r∈[0,1)r\in[0,1), κ∈R\kappa\in\mathbb{R}, and γ∈Ur\gamma\in\mathcal{U}_{r}. Then, on [0,r]×R[0,r]\times\mathbb{R},

b>0,0<v<1,so that0<c<1λ≤11−r.(42)b>0,\qquad0<v<1,\qquad\text{so that}\qquad0<c<\frac{1}{\lambda}\leq\frac{1}{1-r}. \tag*{(42)}

Moreover, for q∈[0,r]q\in[0,r],

b(q,Xq)=b(0,0)−∫0qc(s,Xs) dBs;(43)b(q,X_{q})=b(0,0)-\int_{0}^{q}c(s,X_{s})\,\mathrm{d}B_{s}; \tag*{(43)}

in particular b(q,Xq)b(q,X_{q}) is a square-integrable martingale on [0,r][0,r].

Proof. The signs (42) are Part II, Lemma 2.4(b),(c), applied with K0=∣κ∣K_{0}=|\kappa|; the last inequality holds because λ≥λ(r)=1−r\lambda\geq\lambda(r)=1-r on [0,r][0,r]. It remains to prove (43). For r=0r=0 there is nothing to prove; let r>0r>0 and take I={κ}I=\{\kappa\}.

Step order parameters. Let γ\gamma be a step order parameter and fix (s,x)∈[0,r]×R(s,x)\in[0,r]\times\mathbb{R}. By Itô’s formula on each open interval of constancy and Lb=0\mathcal{L}b=0, db(t,Xts,x)=∂xb(t,Xts,x) dBt=−c(t,Xts,x) dBt\mathrm{d}b(t,X_{t}^{s,x})=\partial_{x}b(t,X_{t}^{s,x})\,\mathrm{d}B_{t}=-c(t,X_{t}^{s,x})\,\mathrm{d}B_{t} there. Since bb is continuous on [0,r]×R[0,r]\times\mathbb{R} with linear growth and 0≤c≤K0\leq c\leq K by Lemma 2.6(iii),(v), the identity extends to the endpoints of the intervals as in Step 3 of the proof of Lemma 2.2, and

b(t,Xts,x)=b(s,x)−∫stc(t′,Xt′s,x) dBt′(t∈[s,r]).b(t,X_{t}^{s,x})=b(s,x)-\int_{s}^{t}c(t',X_{t'}^{s,x})\,\mathrm{d}B_{t'}\qquad(t\in[s,r]).

With (s,x)=(0,0)(s,x)=(0,0) this is (43), and since 0≤c≤K0\leq c\leq K the stochastic integral is a square-integrable martingale on [0,r][0,r].

General order parameters. Let γ∈Ur\gamma\in\mathcal{U}_{r}, let γn∈Ur\gamma_{n}\in\mathcal{U}_{r} be step approximations, and write Xn=XγnX^{n}=X^{\gamma_{n}}, bn=bγnb_{n}=b_{\gamma_{n}}, and cn=cγnc_{n}=c_{\gamma_{n}}. By Lemma 2.7, cn→cc_{n}\to c uniformly on compact subsets of [0,r]×R[0,r] \times\mathbb{R}, and Xn→XX^{n} \to X in every moment, uniformly on [0,r][0,r]. By (39) and the Lipschitz bound ∣∂xb∣≤K|\partial_{x}b| \le K,

∣bn(q,Xqn)−b(q,Xq)∣≤Cεn(1+∣Xqn∣)+K∣Xqn−Xq∣,|b_{n}(q,X_{q}^{n})-b(q,X_{q})| \le C\varepsilon_{n}(1+|X_{q}^{n}|)+K|X_{q}^{n}-X_{q}|,

with εn=∥γn−γ∥L1\varepsilon_{n}=\|\gamma_{n}-\gamma\|_{L^{1}}, so bn(q,Xqn)→b(q,Xq)b_{n}(q,X_{q}^{n}) \to b(q,X_{q}) in L2L^{2}. Moreover cn(s,Xsn)→c(s,Xs)c_{n}(s,X_{s}^{n}) \to c(s,X_{s}) in probability for each ss, by the locally uniform convergence of cnc_{n}, the continuity of cc, and the convergence of XnX^{n}; since 0≤cn,c≤K0 \le c_{n},c \le K, dominated convergence and Itô’s isometry give ∫0qcn(s,Xsn) dBs→∫0qc(s,Xs) dBs\int_{0}^{q}c_{n}(s,X_{s}^{n})\,\mathrm{d}B_{s} \to\int_{0}^{q}c(s,X_{s})\,\mathrm{d}B_{s} in L2L^{2}. The identity (43) for γn\gamma_{n} therefore passes to the limit, almost surely for each fixed q∈[0,r]q \in[0,r]. Both sides of (43) have continuous paths, since bb and XX are continuous and the stochastic integral has a continuous version, so the identity holds simultaneously for all rational qq on one event of probability one, and then for all q∈[0,r]q \in[0,r] by continuity. □\square

The martingale identity (43) is the hard-wall form of [ref-7], Lemma 2. Part II, Lemma 2.9(b) states the martingale property of b(q,Xq)b(q,X_{q}) without the stochastic integral.

The first variation. For γ∈Ur\gamma\in\mathcal{U}_{r} and q∈[0,r]q \in[0,r] set

D(q)=αEb(q,Xq)2−∫0qλ(t)−2 dt,S(q)=αλ(q)2Ec(q,Xq)2.(44)D(q)=\alpha\mathbb{E}b(q,X_{q})^{2}-\int_{0}^{q}\lambda(t)^{-2}\,\mathrm{d}t, \qquad S(q)=\alpha\lambda(q)^{2}\mathbb{E}c(q,X_{q})^{2}. \tag*{(44)}

The first term of DD comes from the energy term of Pκ\mathcal{P}_{\kappa} and the second from the entropy term, as the proof of Part II, Proposition 2.11(a) shows. Since v=λcv=\lambda c, these are the functions DD and SS of Part II, (37). The function DD and the first variation below are the basic tools of Chapter 6.

Lemma 2.9. Let r∈[0,1)r \in[0,1), κ∈R\kappa\in\mathbb{R}, and γ,γ~∈Ur\gamma,\widetilde{\gamma} \in\mathcal{U}_{r}. The function DD is continuously differentiable on [0,r][0,r] with

D′=S−1λ2.D'=\frac{S-1}{\lambda^{2}}.

Moreover, with γθ=γ+θ(γ~−γ)∈Ur\gamma_{\theta}=\gamma+\theta(\widetilde{\gamma}-\gamma) \in\mathcal{U}_{r} for θ∈[0,1]\theta\in[0,1],

ddθPκ(γθ)∣θ=0+=12∫0r(γ~−γ)(q)D(q) dq.(45)\left.\frac{\mathrm{d}}{\mathrm{d}\theta}\mathcal{P}_{\kappa}(\gamma_{\theta})\right|_{\theta=0+} = \frac{1}{2}\int_{0}^{r}(\widetilde{\gamma}-\gamma)(q)D(q)\,\mathrm{d}q. \tag*{(45)}

Proof. The formula for D′D' is among the generator identities of Part II, Lemma 2.9(c)(i), and (45) is the first equality in Part II, (43) (Proposition 2.11(a)); both hold for all γ,γ~∈Ur\gamma,\widetilde{\gamma} \in\mathcal{U}_{r} and κ∈R\kappa\in\mathbb{R}. □\square

Formula (45) is the hard-wall analogue of the directional derivative of the Parisi functional of mixed pp-spin models computed by W.-K. Chen [ref-38], Theorem 2; see also [ref-55], Proposition 6.8. The first variation of El Alaoui and Sellke [ref-54], (4.15) involves the same function DD; the beginning of this chapter compares the two. The formula for D′D' is the analogue of [ref-8] (Proposition 3), where Auffinger and W.-K. Chen differentiate the corresponding function for mixed pp-spin models. Talagrand computed the derivative of the Parisi functional when the order parameter is transported by a map close to the identity [ref-126].

The upper bound

This chapter proves that PU(γ)\mathcal{P}_{U}(\gamma) bounds the free energy from above for every order parameter γ\gamma and every activation UU in class (A) (Theorem 3.1), and deduces upper bounds for the feasible volume and for the maximal margin. Lemma 3.2 is a change of measure on a Ruelle probability cascade with marks, which turns Gibbs averages over the leaves into expectations along a Markov chain, and Corollary 3.3 deduces that conditional overlaps are nondecreasing in the branching level. Lemma 3.4 computes a Gaussian cascade exactly. Lemma 3.5 gives the derivatives of the interpolating free energy. Lemma 3.6 proves concentration of the free energy and computes its derivatives in the perturbation parameters, and Lemmas 3.7 and 3.8 use it to control the conditional covariance of the two overlaps. Lemma 3.9 bounds the spherical term by the Crisanti–Sommers entropy. Proposition 3.10 combines these facts through a maximum principle in the interpolation time. Corollary 3.11 passes to the hard wall, with a bounded source term, and Lemma 3.13 and Corollary 3.14, with Lemma 3.12, bound the maximal margin.

Step 1 of §1.3 outlines the argument and introduces the notation used below. The proof follows that of [ref-92], Theorem 4.1, a step in Mourrat’s proof that the Hamilton–Jacobi limit bounds the free energy of bipartite models from above [ref-92], Theorem 1.1. As there, the atoms of the order parameter have equal mass, the derivative in tt is a sum of products of conditional overlap means and a conditional covariance (see [ref-92], (2.23)), and Mourrat’s perturbation enforces the Ghirlanda–Guerra identities only at contact points (§§3.3 and 3.4). Four things differ. The row side carries a non-Gaussian activation UU, with no sign condition on U′U', in place of Mourrat’s bilinear Gaussian energy. Only the spin side has free cascade levels, so the viscosity argument becomes a maximum principle in tt for the maximum over pp (§3.7). The monotonicity of the response overlaps ζl\zeta_{l} follows from a change of measure on a marked cascade; its Gaussian analogue is the monotonicity of the normalized derivatives of the enriched free energy in the cascade parameters [ref-92], Lemma 2.4 (see also [ref-128], Proposition 14.3.2). The bound closes on the Parisi functional PU(γh)\mathcal{P}_{U}(\gamma_{h}) instead of the solution of a Hamilton–Jacobi equation.

Throughout this chapter, ba=ga/Nb_{a}=g_{a}/\sqrt{N} for the rows gag_{a} of G\mathbf{G}, αN=M/N→α\alpha_{N}=M/N\to\alpha, and α′=sup⁡NαN<∞\alpha'=\sup_{N}\alpha_{N}<\infty. For UU in class (A) we write Umax⁡=sup⁡UU_{\max}=\sup U and L=∥U′∥∞L=\lVert U'\rVert_{\infty}.

Theorem 3.1. Let UU belong to class (A). For every γ∈U\gamma\in\mathcal{U},

lim sup⁡N→∞FN(U)≤PU(γ).(46)\limsup_{N\to\infty}F_{N}(U)\leq\mathcal{P}_{U}(\gamma). \tag*{(46)}

Moreover, P{fN(U)>PU(γ)+ε}→0\mathbb{P}\{f_N(U)>\mathcal{P}_U(\gamma)+\varepsilon\}\to0 for every ε>0\varepsilon>0.

Chapter 5 combines Theorem 3.1 with the lower bound of Chapter 4.

A change of measure on a marked cascade

The interpolation places Gaussian fields on the nodes of a Ruelle probability cascade. Its free energy is an iterated logarithmic moment over the levels of the tree, and to differentiate it we need to express averages over the cascade weights as expectations along a Markov chain on the marks. The following change-of-measure identity does this. It also gives the law of the branching level of two leaves sampled from the weights, and it shows that conditional overlaps are nondecreasing in the level.

The cascade. Fix k≥0k\ge0 and 0<θ1<⋯<θk<10<\theta_1<\cdots<\theta_k<1, and set θ0=0\theta_0=0, θk+1=1\theta_{k+1}=1, and wl=θl+1−θlw_l=\theta_{l+1}-\theta_l for 0≤l≤k0\le l\le k. The nodes of depth ll are the strings β∈N∗l\beta\in\mathbb{N}_*^l, with the empty string as the root, and the leaves are the nodes of depth kk. For a leaf β\beta we write β∣l\beta|l for its ancestor at depth ll, and for two leaves we put β∧β′=max⁡{l:β∣l=β′∣l}\beta\wedge\beta'=\max\{l:\beta|l=\beta'|l\}. At every node of depth l−1l-1, independently, place a Poisson point process on (0,∞)(0,\infty) with intensity measure θlx−1−θl dx\theta_lx^{-1-\theta_l}\,\mathrm{d}x. Its points, in decreasing order, are attached to the edges from that node to its children; πβ\pi_\beta denotes the point attached to the edge into the node β\beta. For a leaf β\beta put

Πβ=∏l=1kπβ∣l,Π=∑βΠβ,vβ=ΠβΠ.\Pi_\beta=\prod_{l=1}^{k}\pi_{\beta|l},\qquad \Pi=\sum_{\beta}\Pi_\beta,\qquad v_\beta=\frac{\Pi_\beta}{\Pi}.

The weights (vβ)(v_\beta) are the Ruelle probability cascade [ref-112]. For k=0k=0 the root is the only leaf, and Π=v=1\Pi=v=1.

Independently of the cascade, attach to each node β\beta of depth l≥1l\ge1 a random mark eβe_\beta with law PlP_l on a standard Borel space; marks at different nodes are independent. The root mark e0e_0 is fixed for now; below it contains the disorder shared by all leaves. For a leaf β\beta write e[β]=(e0,eβ∣1,…,eβ∣k)e_{[\beta]}=(e_0,e_{\beta|1},\ldots,e_{\beta|k}) for the marks along its path. Let Xk(e0,…,ek)X_k(e_0,\ldots,e_k) be a measurable function of a path of marks, and define backwards, for l=k,…,1l=k,\ldots,1,

Xl−1(e0,…,el−1)=1θllog⁡∫eθlXl dPl(el),Kl(del∣e0,…,el−1)=eθl(Xl−Xl−1)Pl(del).(47)X_{l-1}(e_0,\ldots,e_{l-1}) =\frac{1}{\theta_l}\log\int e^{\theta_lX_l}\,\mathrm{d}P_l(e_l), \qquad K_l(\mathrm{d}e_l\mid e_0,\ldots,e_{l-1}) =e^{\theta_l(X_l-X_{l-1})}P_l(\mathrm{d}e_l). \tag*{(47)}

We assume that every XlX_l is finite on every path. Then each KlK_l is a probability kernel. We write EK\mathbb{E}_K for expectation over the Markov chain (e0,e1,…,ek)(e_0,e_1,\ldots,e_k) started from e0e_0 with the kernels K1,…,KkK_1,\ldots,K_k, and, for 0≤l≤k0\le l\le k, EK,l\mathbb{E}_{K,l} for expectation over two paths (e,e′)(e,e') that coincide up to depth ll, follow the chain up to depth ll, and then continue independently with the kernels Kl+1,…,KkK_{l+1},\ldots,K_k.

Lemma 3.2. The quantities Π\Pi and W=∑βvβeXk(e[β])W=\sum_{\beta}v_{\beta}e^{X_{k}(e_{[\beta]})} are positive and finite almost surely. Let ωβ=vβeXk(e[β])/W\omega_{\beta}=v_{\beta}e^{X_{k}(e_{[\beta]})}/W be the associated Gibbs weights on the leaves. Then, for every measurable φ\varphi that is nonnegative, or integrable under the law on the right side,

E∑βωβφ(e[β])=EKφ(e),(48)\mathbb{E}\sum_{\beta}\omega_{\beta}\varphi(e_{[\beta]})=\mathbb{E}_{K}\varphi(e), \tag*{(48)}
E∑β,β′ωβωβ′1{β∧β′=l}φ(e[β],e[β′])=wlEK,lφ(e,e′),0≤l≤k.(49)\mathbb{E}\sum_{\beta,\beta'}\omega_{\beta}\omega_{\beta'}\mathbf{1}_{\{\beta\wedge\beta'=l\}}\varphi(e_{[\beta]},e_{[\beta']})=w_{l}\mathbb{E}_{K,l}\varphi(e,e'),\qquad0\le l\le k. \tag*{(49)}

Moreover, conditionally on every value of the root mark e0e_{0},

log⁡W=X0(e0)+log⁡Π′−log⁡Π,Π′=dΠ,E∣log⁡Π∣2<∞.(50)\log W=X_{0}(e_{0})+\log\Pi'-\log\Pi,\qquad\Pi'\stackrel{d}{=}\Pi,\qquad\mathbb{E}\lvert\log\Pi\rvert^{2}<\infty. \tag*{(50)}

In words, (48) says that the path of marks of a leaf sampled from the Gibbs weights is distributed as the chain with the kernels KlK_{l}, and (49) says that two independently sampled leaves branch at depth ll with probability wlw_{l}, whatever the function XkX_{k}, and that their paths then share the first ll marks and continue independently.

These are known properties of Ruelle probability cascades, and they rest on the invariance property of Bolthausen and Sznitman [ref-19]; see [ref-98], Chapter 2. With marks and the kernels KlK_{l}, (48) and (49) correspond to [ref-103], Theorems 6 and 7, and (50) to [ref-103], (3.12)–(3.13); Panchenko and Talagrand use them to write Guerra’s interpolation on a cascade. See also [ref-4]. Steps 1 and 3 of the proof follow [ref-103], Lemmas 2 and 3, and Step 4 follows Mourrat’s derivation of the law of β∧β′\beta\wedge\beta' by Gaussian integration by parts [ref-93], which adapts the derivation of [ref-98], (2.82).

Proof. For k=0k=0 we have W=eX0(e0)W=e^{X_{0}(e_{0})}, ω=1\omega=1, and w0=1w_{0}=1, and all assertions hold with Π=Π′=1\Pi=\Pi'=1. Let k≥1k\ge1. We divide the proof into four steps.

Step 1. A Poisson computation. Let (xn)(x_{n}) be a Poisson process with intensity θx−1−θ dx\theta x^{-1-\theta}\,\mathrm{d}x, let (en)(e_{n}) be independent marks with law PP, and let ff be measurable with c=∫eθf dP<∞c=\int e^{\theta f}\,\mathrm{d}P<\infty. By the mapping and marking theorems for Poisson processes [ref-81], Chapters 2 and 5, the pairs (xnef(en),en)(x_{n}e^{f(e_{n})},e_{n}) form a Poisson process with intensity

c θy−1−θ dy⊗c−1eθf(e)P(de).c\,\theta y^{-1-\theta}\,\mathrm{d}y\otimes c^{-1}e^{\theta f(e)}P(\mathrm{d}e).

Indeed, for a test function χ\chi the substitution y=xef(e)y=xe^{f(e)} gives

∬χ(xef(e),e) θx−1−θ dx P(de)=∬χ(y,e) θy−1−θeθf(e) dy P(de).\iint\chi(xe^{f(e)},e)\,\theta x^{-1-\theta}\,\mathrm{d}x\,P(\mathrm{d}e) = \iint\chi(y,e)\,\theta y^{-1-\theta}e^{\theta f(e)}\,\mathrm{d}y\,P(\mathrm{d}e).

After the first coordinate is multiplied by c−1/θc^{-1/\theta}, the first factor becomes θy−1−θ dy\theta y^{-1-\theta}\,\mathrm{d}y again. Since the intensity is a product, after ranking the first coordinates the marks are independent with law c−1eθfPc^{-1}e^{\theta f}P and independent of the ranked points. If each mark also carries an independent subtree and ff depends only on the first component of the mark, the conditional law of the subtree given the first component is unchanged.

Step 2. The totals. The sum SθS_{\theta} of the points of one process satisfies Ee−uSθ=exp⁡{−Γ(1−θ)uθ}\mathbb{E}e^{-uS_{\theta}}=\exp\{-\Gamma(1-\theta)u^{\theta}\} for u≥0u\geq0, so it is positive and finite almost surely. The identities

x−a=1Γ(a)∫0∞ua−1e−ux du(a>0),x^{-a}=\frac{1}{\Gamma(a)}\int_{0}^{\infty}u^{a-1}e^{-ux}\,\mathrm{d}u \qquad(a>0),
xa=aΓ(1−a)∫0∞(1−e−ux)u−a−1 du(0<a<1),x^{a}=\frac{a}{\Gamma(1-a)}\int_{0}^{\infty}(1-e^{-ux})u^{-a-1}\,\mathrm{d}u \qquad(0<a<1),

applied to x=Sθx=S_{\theta} and integrated, show that SθS_{\theta} has negative moments of all orders and positive moments of all orders a<θa<\theta; the second integral converges at u=0u=0 because 1−e−Γ(1−θ)uθ1-e^{-\Gamma(1-\theta)u^{\theta}} is of order uθu^{\theta}. In particular E∣log⁡Sθ∣2<∞\mathbb{E}|\log S_{\theta}|^{2}<\infty. Next, Π\Pi is a deterministic multiple of a copy of Sθ1S_{\theta_{1}}. If SnS_{n} denotes the total of the subtree below the nn-th child of the root, then Π=∑nxnSn\Pi=\sum_{n}x_{n}S_{n}, and by Step 1 with f=log⁡Snf=\log S_{n} the points xnSnx_{n}S_{n} form a Poisson process with intensity E[Sθ1] θ1x−1−θ1 dx\mathbb{E}[S^{\theta_{1}}]\,\theta_{1}x^{-1-\theta_{1}}\,\mathrm{d}x. By induction on kk, the subtree total SS is a multiple of a copy of Sθ2S_{\theta_{2}}, so ESθ1<∞\mathbb{E}S^{\theta_{1}}<\infty because θ1<θ2\theta_{1}<\theta_{2}. This proves the assertions about Π\Pi.

Step 3. The change of measure. For a node β\beta of depth l≥1l\geq1 with parent β−\beta^{-}, replace the edge weight by

π~β=πβeXl(e[β])−Xl−1(e[β−]),\widetilde{\pi}_{\beta}=\pi_{\beta}e^{X_{l}(e_{[\beta]})-X_{l-1}(e_{[\beta^{-}]})},

where e[β]e_{[\beta]} denotes the marks along the path from the root to β\beta. Given the ancestors of β\beta, the multiplier depends only on the mark eβe_{\beta}, and ∫eθl(Xl−Xl−1) dPl=1\int e^{\theta_{l}(X_{l}-X_{l-1})}\,\mathrm{d}P_{l}=1 by (47). Apply Step 1 at the root with c=1c=1, carrying the independent subtree of each child as part of its mark, and then apply it recursively below each child. By induction on the depth, after ranking at each node the modified weights π~\widetilde{\pi} form a cascade with the same law as the original one and independent of the marks, while the mark of a child of depth ll has conditional law KlK_{l} given its ancestors. The relabeling by ranks preserves the ancestry. The multipliers telescope along a path:

πβeXk(e[β])=eX0(e0)∏l=1kπ~β∣l.\pi_{\beta}e^{X_{k}(e_{[\beta]})}=e^{X_{0}(e_{0})}\prod_{l=1}^{k}\widetilde{\pi}_{\beta|l}.

Summing over β\beta gives W=eX0(e0)Π′/ΠW=e^{X_{0}(e_{0})}\Pi'/\Pi with Π′=∑β∏lπ~β∣l\Pi'=\sum_{\beta}\prod_{l}\widetilde{\pi}_{\beta|l}, which has the law of Π\Pi. This is (50), including the positivity and finiteness of WW. Moreover, ωβ=∏lπ~β∣l/Π′\omega_{\beta}=\prod_{l}\widetilde{\pi}_{\beta|l}/\Pi' are the weights of the relabeled cascade, which is independent of the marks. This proves (48). It also proves (49) with the coefficient wˉl=E∑β∧β′=lvβvβ′\bar{w}_{l}=\mathbb{E}\sum_{\beta\wedge\beta'=l}v_{\beta}v_{\beta'} in place of wlw_{l}, because, given the relabeled cascade, the marks along the paths to two leaves with β∧β′=l\beta\wedge\beta'=l share the first ll marks and then continue independently.

Step 4. The branching probabilities. The coefficient wˉl\bar{w}_{l} does not depend on the marks, so we may compute it with a convenient choice. Fix 1≤l≤k1 \le l \le k and s>0s>0, let the mark of each node ν\nu of depth ll be a standard Gaussian variable GνG_{\nu} (the other marks play no role), and take Xk(e[β])=sGβ∣lX_{k}(e_{[\beta]})=sG_{\beta|_{l}}. Then Xm=sGβ∣lX_{m}=sG_{\beta|_{l}} for m≥lm \ge l and Xl−1=θls2/2X_{l-1}=\theta_{l}s^{2}/2, so KlK_{l} is the law N(θls,1)N(\theta_{l}s,1), and (3.3) gives

E∑∣ν∣=lωνGν=θls,ων=∑β∣l=νωβ.\mathbb{E}\sum_{|\nu|=l}\omega_{\nu}G_{\nu}=\theta_{l}s, \qquad \omega_{\nu}=\sum_{\beta|_{l}=\nu}\omega_{\beta}.

The sum over nodes is absolutely integrable, by (3.3) applied to the path function ∣Gβ∣l∣|G_{\beta|_{l}}|. Conditionally on all other variables, ων=AesGν/(AesGν+B)\omega_{\nu}=Ae^{sG_{\nu}}/(Ae^{sG_{\nu}}+B) with A,BA,B independent of GνG_{\nu}, so Gaussian integration by parts gives E[Gνων]=s E[ων(1−ων)]\mathbb{E}[G_{\nu}\omega_{\nu}]=s\,\mathbb{E}[\omega_{\nu}(1-\omega_{\nu})]. Summing over ν\nu and using ∑νων=1\sum_{\nu}\omega_{\nu}=1 gives E∑νων2=1−θl\mathbb{E}\sum_{\nu}\omega_{\nu}^{2}=1-\theta_{l}. By (3.4) with the coefficients wˉm\bar{w}_{m} and φ=1\varphi=1, the left side equals ∑m≥lwˉm\sum_{m\ge l}\bar{w}_{m}. Hence ∑m≥lwˉm=1−θl\sum_{m\ge l}\bar{w}_{m}=1-\theta_{l} for 1≤l≤k1\le l\le k, and also for l=0l=0 since ∑mwˉm=1\sum_{m}\bar{w}_{m}=1. Taking differences gives wˉl=θl+1−θl=wl\bar{w}_{l}=\theta_{l+1}-\theta_{l}=w_{l}. □\square

We write ⟨⋅⟩\langle\cdot\rangle for averages under the Gibbs weights and their independent replicas, that is, leaves β1,β2,…\beta^{1},\beta^{2},\ldots sampled independently from (ωβ)(\omega_{\beta}). Conditional averages given β1∧β2=l\beta^{1}\wedge\beta^{2}=l are taken with respect to the probability E⟨⋅ 1{β1∧β2=l}⟩/wl\mathbb{E}\langle\cdot\,\mathbf{1}_{\{\beta^{1}\wedge\beta^{2}=l\}}\rangle/w_{l}.

Corollary 3.3. Let ff be a bounded function of a path of marks with values in a Euclidean space, and let Fl=EK[f∣e0,…,el]F_{l}=\mathbb{E}_{K}[f\mid e_{0},\ldots,e_{l}]. Then

E⟨f(e[β1])⋅f(e[β2])∣β1∧β2=l⟩=EK∥Fl∥2,\mathbb{E}\langle f(e_{[\beta^{1}]})\cdot f(e_{[\beta^{2}]})\mid\beta^{1}\wedge\beta^{2}=l\rangle = \mathbb{E}_{K}\lVert F_{l}\rVert^{2},

and the right side is nonnegative and nondecreasing in ll. Moreover,

E[log⁡W∣e0]=X0(e0),Var⁡(log⁡W∣e0)≤4Var⁡(log⁡Π).(51)\mathbb{E}[\log W\mid e_{0}]=X_{0}(e_{0}), \qquad \operatorname{Var}(\log W\mid e_{0})\le4\operatorname{Var}(\log\Pi). \tag*{(51)}

Proof. In (3.4), condition on the common part of the two paths. Their continuations are independent, so the conditional expectation of f(e)⋅f(e′)f(e)\cdot f(e') is ∥Fl∥2\lVert F_{l}\rVert^{2}. The sequence (Fl)(F_{l}) is a martingale for the chain, so EK∥Fl∥2\mathbb{E}_{K}\lVert F_{l}\rVert^{2} is nondecreasing by conditional Jensen. For (51), take conditional expectations in (3.5), using Π′=dΠ\Pi'\stackrel{d}{=}\Pi given e0e_{0}. For the variance, center both logarithms at Elog⁡Π\mathbb{E}\log\Pi and use (x−y)2≤2x2+2y2(x-y)^{2}\le2x^{2}+2y^{2}; independence of Π\Pi and Π′\Pi' is not needed. □\square

In Lemma 3.5(iv) we apply the monotonicity in ll to the response overlap of the activation.

Gaussian cascades

Both bounds on the free energy use one explicit computation: a Gaussian field on a cascade, integrated against a Gaussian weight at each leaf. In this chapter it bounds the spherical term (Lemma 3.9); in Chapter 4 it evaluates the cavity term. The useful observation is that the field along a sampled path, divided by its precision, is a Gaussian martingale. Its variances give every overlap of the added coordinates, and the normalization of their squared norm identifies the value with a Crisanti–Sommers entropy.

Let K={p∈Rk+1:0≤p0≤⋯≤pk}\mathcal{K}=\{p\in\mathbb{R}^{k+1}:0\leq p_{0}\leq\cdots\leq p_{k}\}, and put Δp0=p0\Delta p_{0}=p_{0} and Δpl=pl−pl−1\Delta p_{l}=p_{l}-p_{l-1} for 1≤l≤k1\leq l\leq k. For x∈Kx\in\mathcal{K} with xk<1x_{k}<1 let

γx(q)=∑l=0kwl1{q≥xl},λγx(s)=∫s1γx(q) dq=1−∑l=0kwlmax⁡(xl,s);\gamma_{x}(q)=\sum_{l=0}^{k}w_{l}\mathbf{1}_{\{q\geq x_{l}\}},\qquad \lambda_{\gamma_{x}}(s)=\int_{s}^{1}\gamma_{x}(q)\,\mathrm{d}q =1-\sum_{l=0}^{k}w_{l}\max(x_{l},s);

γx∈U\gamma_{x}\in\mathcal{U} is the order parameter with atoms of mass wlw_{l} at xlx_{l}.

Fix n≥1n\geq1 and p∈Kp\in\mathcal{K}. Let (yν)(y_{\nu}) be independent standard Gaussian vectors in Rn\mathbb{R}^{n} indexed by all nodes ν\nu of the tree, including the root, independent of the cascade, and put

Yβ=∑l=0kΔplyβ∣l,Yβ,l=∑j=0lΔpjyβ∣j(0≤l≤k),Y_{\beta}=\sum_{l=0}^{k}\sqrt{\Delta p_{l}}y_{\beta|l},\qquad Y_{\beta,l}=\sum_{j=0}^{l}\sqrt{\Delta p_{j}}y_{\beta|j} \qquad(0\leq l\leq k),

for a leaf β\beta. The coordinates of YY are independent, and each has covariance pβ∧β′p_{\beta\wedge\beta'} between the leaves β\beta and β′\beta'. We call YY the Gaussian cascade field with levels pp. Summation by parts gives

∑l=1kθlΔpl=pk−∑l=0kwlpl.(52)\sum_{l=1}^{k}\theta_{l}\Delta p_{l} =p_{k}-\sum_{l=0}^{k}w_{l}p_{l}. \tag*{(52)}

For b>∑l=1kθlΔplb>\sum_{l=1}^{k}\theta_{l}\Delta p_{l} define, for 0≤l≤k0\leq l\leq k,

dl=b−∑j=l+1kθjΔpj;ml=p0d02+∑j=1lΔpjdj−1dj;G(b,p)=p02d0+12∑j=1k1θjlog⁡djdj−1,(53)d_{l}=b-\sum_{j=l+1}^{k}\theta_{j}\Delta p_{j};\qquad m_{l}=\frac{p_{0}}{d_{0}^{2}}+\sum_{j=1}^{l}\frac{\Delta p_{j}}{d_{j-1}d_{j}};\qquad \mathcal{G}(b,p)=\frac{p_{0}}{2d_{0}}+\frac{1}{2}\sum_{j=1}^{k}\frac{1}{\theta_{j}}\log\frac{d_{j}}{d_{j-1}}, \tag*{(53)}

and put ϱ=mk+1/b\varrho=m_{k}+1/b. Since dl−dl−1=θlΔpl≥0d_{l}-d_{l-1}=\theta_{l}\Delta p_{l}\geq0 and d0=b−pk+∑lwlpl>0d_{0}=b-p_{k}+\sum_{l}w_{l}p_{l}>0, we have 0<d0≤⋯≤dk=b0<d_{0}\leq\cdots\leq d_{k}=b and 0≤m0≤⋯≤mk0\leq m_{0}\leq\cdots\leq m_{k}. We call dld_{l} the precision at depth ll. The top precision bb is a scalar and is unrelated to the rows bab_{a}. Let νn\nu_{n} be the standard Gaussian measure on Rn\mathbb{R}^{n}, and let M\mathcal{M} be the random probability measure on pairs (β,ε)(\beta,\varepsilon) of a leaf and a vector of Rn\mathbb{R}^{n} given by

M(β,dε)∝vβνn(dε)exp⁡{ε⋅Yβ−b−12∥ε∥2}.\mathcal{M}(\beta,\mathrm{d}\varepsilon)\propto v_{\beta}\nu_{n}(\mathrm{d}\varepsilon)\exp\left\{\varepsilon\cdot Y_{\beta}-\frac{b-1}{2}\lVert\varepsilon\rVert^{2}\right\}.

We write EM\mathbb{E}_{\mathcal{M}} and PM\mathbb{P}_{\mathcal{M}} for expectation and probability over the cascade, the field YY, and independent samples (β1,ε1),(β2,ε2),…(\beta^{1},\varepsilon^{1}),(\beta^{2},\varepsilon^{2}),\ldots from M\mathcal{M}, and we put J=β1∧β2J=\beta^{1}\wedge\beta^{2}. Conditional expectations given J=lJ=l are taken with respect to EM[⋅ 1{J=l}]/PM(J=l)\mathbb{E}_{\mathcal{M}}[\cdot\,\mathbf{1}_{\{J=l\}}]/\mathbb{P}_{\mathcal{M}}(J=l).

Lemma 3.4. In this setting the following hold. (i) The normalization of M\mathcal{M} is finite, and

1nElog⁡∑βvβe∥Yβ∥2/(2b)=G(b,p),(54)\frac{1}{n}\mathbb{E}\log\sum_{\beta}v_{\beta}e^{\lVert Y_{\beta}\rVert^{2}/(2b)}=\mathcal{G}(b,p), \tag*{(54)}
1nElog⁡∑βvβ∫νn(dε) eε⋅Yβ−(b−1)∥ε∥2/2=G(b,p)−12log⁡b.(55)\frac{1}{n}\mathbb{E}\log\sum_{\beta}v_{\beta}\int\nu_{n}(\mathrm{d}\varepsilon)\,e^{\varepsilon\cdot Y_{\beta}-(b-1)\lVert\varepsilon\rVert^{2}/2}=\mathcal{G}(b,p)-\frac{1}{2}\log b. \tag*{(55)}

(ii) Under PM\mathbb{P}_{\mathcal{M}}, the process Ml=Yβ1,l/dlM_{l}=Y_{\beta^{1},l}/d_{l}, 0≤l≤k0\leq l\leq k, is a centered Gaussian martingale with independent increments and

M0∼N(0,m0In),Ml−Ml−1∼N(0,(ml−ml−1)In)(1≤l≤k),(56)M_{0}\sim N(0,m_{0}I_{n}),\qquad M_{l}-M_{l-1}\sim N(0,(m_{l}-m_{l-1})I_{n})\qquad(1\leq l\leq k), \tag*{(56)}

and conditionally on β1\beta^{1} and YY the vector ε1\varepsilon^{1} has the law N(Mk,b−1In)N(M_{k},b^{-1}I_{n}). In particular,

Law⁡M(ε1)=N(0,ϱIn),(57)\operatorname{Law}_{\mathcal{M}}(\varepsilon^{1})=N(0,\varrho I_{n}), \tag*{(57)}
PM(J=l)=wl,EM[ε1⋅ε2n∣J=l]=ml(0≤l≤k).(58)\mathbb{P}_{\mathcal{M}}(J=l)=w_{l},\qquad\mathbb{E}_{\mathcal{M}}\left[\left.\frac{\varepsilon^{1}\cdot\varepsilon^{2}}{n}\right|J=l\right]=m_{l}\qquad(0\leq l\leq k). \tag*{(58)}

(iii) The levels satisfy the duality identity

ϱpk−∑l=0kwlmlpl=bϱ−1.(59)\varrho p_{k}-\sum_{l=0}^{k}w_{l}m_{l}p_{l}=b\varrho-1. \tag*{(59)}

(iv) The order parameter γm/ϱ\gamma_{m/\varrho}, with atoms of mass wlw_{l} at ml/ϱm_{l}/\varrho, belongs to U\mathcal{U}, its largest atom is mk/ϱ=1−1/(ϱb)m_{k}/\varrho=1-1/(\varrho b), and

G(b,p)−12log⁡b=CS(γm/ϱ)+12log⁡ϱ.(60)\mathcal{G}(b,p)-\frac{1}{2}\log b=\mathrm{CS}(\gamma_{m/\varrho})+\frac{1}{2}\log\varrho. \tag*{(60)}

If ϱ=1\varrho=1, then moreover, for 0≤l≤k0\leq l\leq k,

λγm(ml)=1dl,pl=∫0mldsλγm(s)2,CS(γm)=G(b,p)−12log⁡b,(61)\lambda_{\gamma_{m}}(m_{l})=\frac{1}{d_{l}},\qquad p_{l}=\int_{0}^{m_{l}}\frac{\mathrm{d}s}{\lambda_{\gamma_{m}}(s)^{2}},\qquad\mathrm{CS}(\gamma_{m})=\mathcal{G}(b,p)-\frac{1}{2}\log b, \tag*{(61)}
pk−∑l=0kwlmlpl=b−1,(62)p_{k}-\sum_{l=0}^{k}w_{l}m_{l}p_{l}=b-1, \tag*{(62)}

and dl≥1d_{l}\geq1, 1≤b≤1+pk1\leq b\leq1+p_{k}, and mk≤1−(1+pk)−1m_{k}\leq1-(1+p_{k})^{-1}.

(v) If YcY^{c} is a real Gaussian cascade field with levels c∈Kc\in\mathcal{K} (the case n=1n=1), then

Elog⁡∑βvβeYβc=12(ck−∑l=0kwlcl).(63)\mathbb{E}\log\sum_{\beta}v_{\beta}e^{Y_{\beta}^{c}}=\frac{1}{2}\left(c_{k}-\sum_{l=0}^{k}w_{l}c_{l}\right). \tag*{(63)}

By (49), the branching probabilities wlw_l do not depend on the function integrated at the leaves. In particular, they remain wlw_l when the ε\varepsilon-integral in the definition of M\mathcal{M} is restricted to a ball. The restricted measure, used in §4.5, is defined for every bb, including the case d0≤0d_0 \le0 in which M\mathcal{M} is not defined. In (60), ϱ\varrho is the mean of ∥ε1∥2/n\lVert\varepsilon^1 \rVert^2/n under PM\mathbb{P}_{\mathcal{M}}, by (57); when it equals one, the value (55) is exactly the Crisanti–Sommers entropy of the law of the conditional overlaps mJm_J.

The value G(b,p)\mathcal{G}(b,p), together with the term (b−1−log⁡b)/2(b-1-\log b)/2, is the spherical part of the expression minimized over bb in the Parisi formula for spherical models of Talagrand [ref-125] and W.-K. Chen [ref-37]. Parts (iii) and (iv) identify it with a Crisanti–Sommers entropy, in line with Talagrand’s theorem that the spherical Parisi formula agrees with the Crisanti–Sommers formula [ref-44, ref-125]. Part (v) is the linear case computed by Arguin [ref-4]; see also [ref-100] (22).

Proof. We divide the proof into five steps. Steps 1, 2, and 5 apply Lemma 3.2 to the Gaussian increments; Steps 3 and 4 are algebraic.

Step 1. The recursion. We use the marks e0=p0 y∅e_0=\sqrt{p_0}\,y_{\varnothing} at the root and eν=Δpl yνe_\nu=\sqrt{\Delta p_l}\,y_\nu at a node ν\nu of depth l≥1l\ge1, with laws Pl=N(0,ΔplIn)P_l=N(0,\Delta p_l I_n), and write xl=e0+e1+⋯+elx_l=e_0+e_1+\cdots+e_l for the partial sums along a path, so that Yβ,l=xlY_{\beta,l}=x_l along the path of β\beta. Since ∫νn(dε)eε⋅x−(b−1)∥ε∥2/2=b−n/2e∥x∥2/(2b)\int\nu_n(\mathrm{d}\varepsilon)e^{\varepsilon\cdot x-(b-1)\lVert\varepsilon\rVert^2/2}=b^{-n/2}e^{\lVert x\rVert^2/(2b)} for b>0b>0, the leaf function of (55) is Xk=∥xk∥2/(2b)−n2log⁡bX_k=\lVert x_k\rVert^2/(2b)-\frac{n}{2}\log b, and that of (54) is ∥xk∥2/(2b)\lVert x_k\rVert^2/(2b). For one coordinate, θ∈(0,1)\theta\in(0,1), Δ≥0\Delta\ge0, d>θΔd>\theta\Delta, and a standard Gaussian variable GG, completing the square gives

1θlog⁡Eexp⁡(θ(x+ΔG)22d)=x22(d−θΔ)+12θlog⁡dd−θΔ.\frac{1}{\theta}\log\mathbb{E}\exp\left(\frac{\theta(x+\sqrt{\Delta}G)^2}{2d}\right) = \frac{x^2}{2(d-\theta\Delta)} + \frac{1}{2\theta}\log\frac{d}{d-\theta\Delta}.

Since dl−θlΔpl=dl−1>0d_l-\theta_l\Delta p_l=d_{l-1}>0, backward induction in (47) shows that Xl(e0,…,el)=∥xl∥2/(2dl)+alX_l(e_0,\ldots,e_l)=\lVert x_l\rVert^2/(2d_l)+a_l with al−1=al+n2θllog⁡(dl/dl−1)a_{l-1}=a_l+\frac{n}{2\theta_l}\log(d_l/d_{l-1}), where ak=−n2log⁡ba_k=-\frac{n}{2}\log b for (55) and ak=0a_k=0 for (54). All values are finite, so Lemma 3.2 applies. By (50), nn times the left side of (55) equals EX0(e0)=np0/(2d0)+a0=n[G(b,p)−12log⁡b]\mathbb{E}X_0(e_0)=np_0/(2d_0)+a_0=n[\mathcal{G}(b,p)-\frac{1}{2}\log b], and the same computation with ak=0a_k=0 gives (54). This proves (i).

Step 2. The law of a sampled path. Under M\mathcal{M} the leaf β1\beta^1 is sampled from the Gibbs weights of Lemma 3.2 with the leaf function XkX_k of (55), and, given the leaf and YY, the vector ε1\varepsilon^1 has density proportional to eε⋅Yβ1−b∥ε∥2/2e^{\varepsilon\cdot Y_{\beta^1}-b\lVert\varepsilon\rVert^2/2} with respect to Lebesgue measure, that is, the law N(Yβ1/b,b−1In)N(Y_{\beta^1}/b,b^{-1}I_n). By (47), the kernel Kl(del∣e0,…,el−1)K_l(\mathrm{d}e_l\mid e_0,\ldots,e_{l-1}) has density proportional to exp⁡{θl∥xl−1+el∥2/(2dl)−∥el∥2/(2Δpl)}\exp\{\theta_l\lVert x_{l-1}+e_l\rVert^2/(2d_l)-\lVert e_l\rVert^2/(2\Delta p_l)\}. The coefficient of −∥el∥2/2-\lVert e_l\rVert^2/2 in the exponent is 1/Δpl−θl/dl=dl−1/(dlΔpl)1/\Delta p_l-\theta_l/d_l=d_{l-1}/(d_l\Delta p_l), so KlK_l is the Gaussian law with mean θlΔplxl−1/dl−1\theta_l\Delta p_l x_{l-1}/d_{l-1} and covariance (Δpldl/dl−1)In(\Delta p_l d_l/d_{l-1})I_n (the point mass at 00 if Δpl=0\Delta p_l=0). Since dl−1+θlΔpl=dld_{l-1}+\theta_l\Delta p_l=d_l, under the chain

xl=dldl−1xl−1+(Δpldldl−1)1/2ζl,hencexldl=xl−1dl−1+(Δpldl−1dl)1/2ζl,x_l=\frac{d_l}{d_{l-1}}x_{l-1} +\left(\frac{\Delta p_l d_l}{d_{l-1}}\right)^{1/2}\zeta_l, \qquad \text{hence}\qquad \frac{x_l}{d_l} = \frac{x_{l-1}}{d_{l-1}} +\left(\frac{\Delta p_l}{d_{l-1}d_l}\right)^{1/2}\zeta_l,

with independent standard Gaussian vectors ξl\xi_l that are independent of the root mark. The root mark keeps its law N(0,p0In)N(0,p_0I_n), so x0/d0∼N(0,(p0/d02)In)x_0/d_0\sim N(0,(p_0/d_0^2)I_n). By (48), averaged over the root mark, the process Ml=xl/dlM_l=x_l/d_l along the path of β1\beta^1 has this law, which is (56) by the definition of mlm_l. Since Yβ1/b=MkY_{\beta^1}/b=M_k, the vector ε1\varepsilon^1 is Mk+b−1/2ξ′M_k+b^{-1/2}\xi' with an independent standard Gaussian vector ξ′\xi', which gives (57). By (49), J=lJ=l with probability wlw_l, and the two paths then share x0,…,xlx_0,\ldots,x_l and continue independently with the same kernels. Given xlx_l, the vectors Mk1M_k^1 and Mk2M_k^2 of the two paths are therefore independent, each with conditional mean Ml=xl/dlM_l=x_l/d_l, and EM[ε1⋅ε2∣J=l]=E∥Ml∥2=nml\mathbb{E}_M[\varepsilon^1\cdot\varepsilon^2\mid J=l]=\mathbb{E}\lVert M_l\rVert^2=nm_l. All functions involved are polynomials in Gaussian vectors, so the integrability required in Lemma 3.2 holds. This proves (ii).

Step 3. The duality identity. Let λϱ(s)=ϱ−∑lwlmax⁡(ml,s)\lambda_\varrho(s)=\varrho-\sum_l w_l\max(m_l,s) for s∈[0,mk]s\in[0,m_k]. This function is nonincreasing, with λϱ′(s)=−∑l:ml<swl\lambda_\varrho'(s)=-\sum_{l:m_l<s}w_l for s∉{ml}s\notin\{m_l\}; it is constant on [0,m0][0,m_0] and affine with slope −θj-\theta_j on [mj−1,mj][m_{j-1},m_j]. Moreover λϱ(mk)=ϱ−mk=1/b=1/dk\lambda_\varrho(m_k)=\varrho-m_k=1/b=1/d_k, and

λϱ(mj−1)−λϱ(mj)=θj(mj−mj−1)=θjΔpjdj−1dj=1dj−1−1dj,\lambda_\varrho(m_{j-1})-\lambda_\varrho(m_j) =\theta_j(m_j-m_{j-1}) =\frac{\theta_j\Delta p_j}{d_{j-1}d_j} =\frac{1}{d_{j-1}}-\frac{1}{d_j},

since dj−dj−1=θjΔpjd_j-d_{j-1}=\theta_j\Delta p_j. By downward induction, λϱ(mj)=1/dj\lambda_\varrho(m_j)=1/d_j for all jj; in particular λϱ≥1/b>0\lambda_\varrho\ge1/b>0 on [0,mk][0,m_k]. On [0,m0][0,m_0] we have λϱ=1/d0\lambda_\varrho=1/d_0, and on [mj−1,mj][m_{j-1},m_j] we have (1/λϱ)′=θj/λϱ2(1/\lambda_\varrho)'=\theta_j/\lambda_\varrho^2 and (log⁡λϱ)′=−θj/λϱ(\log\lambda_\varrho)'=-\theta_j/\lambda_\varrho. Hence the constant piece contributes m0d02=p0m_0d_0^2=p_0 and m0d0=p0/d0m_0d_0=p_0/d_0, and the jj-th affine piece contributes (dj−dj−1)/θj=Δpj(d_j-d_{j-1})/\theta_j=\Delta p_j and θj−1log⁡(dj/dj−1)\theta_j^{-1}\log(d_j/d_{j-1}), to the two integrals

∫0mldsλϱ(s)2=pl(0≤l≤k),∫0mkdsλϱ(s)=2G(b,p).(64)\int_0^{m_l}\frac{\mathrm{d}s}{\lambda_\varrho(s)^2}=p_l\quad(0\le l\le k),\qquad \int_0^{m_k}\frac{\mathrm{d}s}{\lambda_\varrho(s)}=2\mathcal{G}(b,p). \tag*{(64)}

(If mj−1=mjm_{j-1}=m_j, then Δpj=0\Delta p_j=0 and dj=dj−1d_j=d_{j-1}, and both contributions vanish.) Insert the first identity of (64) into the left side of (59) and use Fubini’s theorem:

ϱpk−∑lwlmlpl=∫0mkϱ−∑l:ml>swlmlλϱ(s)2 ds.\varrho p_k-\sum_l w_lm_lp_l =\int_0^{m_k}\frac{\varrho-\sum_{l:m_l>s}w_lm_l}{\lambda_\varrho(s)^2}\,\mathrm{d}s.

The numerator equals λϱ(s)+s∑l:ml≤swl=λϱ(s)−sλϱ′(s)\lambda_\varrho(s)+s\sum_{l:m_l\le s}w_l=\lambda_\varrho(s)-s\lambda_\varrho'(s) for almost every ss, so the integrand is (s/λϱ(s))′(s/\lambda_\varrho(s))', and the integral is mk/λϱ(mk)=bmk=bϱ−1m_k/\lambda_\varrho(m_k)=bm_k=b\varrho-1. This proves (59) and (iii).

Step 4. The entropy. Suppose first that ϱ=1\varrho=1. Then mk=1−1/b<1m_k=1-1/b<1, γm∈U\gamma_m\in\mathcal{U}, and λϱ=λγm\lambda_\varrho=\lambda_{\gamma_m} on [0,mk][0,m_k]. Step 3 gives the first two identities in (61), and (59) becomes (62). By (64) and the definition of CS⁡\operatorname{CS} with the cutoff mk∈[Qγm,1)m_k\in[Q_{\gamma_m},1),

CS⁡(γm)=12∫0mkdsλγm(s)+12log⁡(1−mk)=G(b,p)−12log⁡b.\operatorname{CS}(\gamma_m) =\frac{1}{2}\int_0^{m_k}\frac{\mathrm{d}s}{\lambda_{\gamma_m}(s)} +\frac{1}{2}\log(1-m_k) =\mathcal{G}(b,p)-\frac{1}{2}\log b.

Moreover 1/dl=λγm(ml)≤λγm(0)≤11/d_l=\lambda_{\gamma_m}(m_l)\leq\lambda_{\gamma_m}(0)\leq1, so b=dk≥dl≥1b=d_k\geq d_l\geq1; since ml,pl≥0m_l,p_l\geq0, (62) gives b−1≤pkb-1\leq p_k, and then mk=1−1/b≤1−(1+pk)−1m_k=1-1/b\leq1-(1+p_k)^{-1}. For general ϱ>0\varrho>0, multiply the levels and the top precision by ϱ\varrho. The parameters (ϱb,ϱp)(\varrho b,\varrho p) satisfy the standing condition, and by (53) they have precisions ϱdl\varrho d_l, conditional levels ml/ϱm_l/\varrho, and the same value G(ϱb,ϱp)=G(b,p)\mathcal{G}(\varrho b,\varrho p)=\mathcal{G}(b,p), because p0/d0p_0/d_0 and the ratios dj/dj−1d_j/d_{j-1} do not change. Their top level satisfies mk/ϱ+1/(ϱb)=1m_k/\varrho+1/(\varrho b)=1. The case ϱ=1\varrho=1, applied to them, shows that γm/ϱ∈U\gamma_m/\varrho\in\mathcal{U} with largest atom mk/ϱ=1−1/(ϱb)m_k/\varrho=1-1/(\varrho b), and

CS(γm/ϱ)=G(ϱb,ϱp)−12log⁡(ϱb)=G(b,p)−12log⁡b−12log⁡ϱ.\mathrm{CS}(\gamma_m/\varrho)=\mathcal{G}(\varrho b,\varrho p)-\frac{1}{2}\log(\varrho b)=\mathcal{G}(b,p)-\frac{1}{2}\log b-\frac{1}{2}\log\varrho.

This proves (iv).

Step 5. The linear field. Apply Lemma 3.2 with the marks of Step 1 for the levels cc (with n=1n=1) and the leaf function Xk=xkX_k=x_k. For θ∈(0,1)\theta\in(0,1) and Δ≥0\Delta\geq0, θ−1log⁡Eeθ(x+ΔG)=x+θΔ/2\theta^{-1}\log\mathbb{E}e^{\theta(x+\sqrt{\Delta}G)}=x+\theta\Delta/2, so Xl=xl+12∑j>lθjΔcjX_l=x_l+\frac{1}{2}\sum_{j>l}\theta_j\Delta c_j and EX0(e0)=12∑j=1kθjΔcj\mathbb{E}X_0(e_0)=\frac{1}{2}\sum_{j=1}^{k}\theta_j\Delta c_j. By (50) and (52) for cc, this is (63). For k=0k=0 both sides of (63) vanish. □\square

The interpolation and its derivatives

We now define the interpolating free energy and compute its derivatives. Throughout this section the activation UU is in class (A). From now on we fix an integer K≥2K\geq2 and use the cascade of §3.1 with

k=K−1,θl=lK(0≤l≤K),so that wl=1K(0≤l≤k).k=K-1,\qquad\theta_l=\frac{l}{K}\quad(0\leq l\leq K),\qquad\text{so that }w_l=\frac{1}{K}\quad(0\leq l\leq k).

We keep writing wlw_l, which makes the role of each factor visible. We also fix the levels of the row field,

h=(h0,…,hk)∈K,hk<1,h=(h_0,\ldots,h_k)\in\mathcal{K},\qquad h_k<1,

so that γh\gamma_h is the order parameter with atoms of mass wlw_l at hlh_l. The spin field has levels p∈Kp\in\mathcal{K}. At every node β\beta of the cascade, including the root, take independent standard Gaussian vectors yβ∈RNy_\beta\in\mathbb{R}^N and zβ∈RMz_\beta\in\mathbb{R}^M, independent of GG, and set, for a leaf β\beta,

Yβ=∑l=0kΔpl yβ∣l,Ξβ=∑l=0kΔhl zβ∣l,Y_\beta=\sum_{l=0}^{k}\sqrt{\Delta p_l}\,y_{\beta|l},\qquad \Xi_\beta=\sum_{l=0}^{k}\sqrt{\Delta h_l}\,z_{\beta|l},

where Δh0=h0\Delta h_0=h_0 and Δhl=hl−hl−1\Delta h_l=h_l-h_{l-1}. Thus YY and Ξ\Xi are the Gaussian cascade fields with levels pp and hh in dimensions NN and MM. The remaining variance 1−hk1-h_k of the row field is carried by an auxiliary standard Gaussian vector ξ∈RM\xi\in\mathbb{R}^M inside each leaf, so that each coordinate of Ξβ+1−hk ξ\Xi_\beta+\sqrt{1-h_k}\,\xi has variance one, as does ba⋅σb_a\cdot\sigma for fixed σ∈SN\sigma\in S_N. This equality of variances is what makes the diagonal terms cancel when the activation is differentiated in tt. Without the cascade, interpolations of this form were used for the perceptron by Talagrand [ref-127].

The perturbation. To obtain approximate Ghirlanda–Guerra identities we add a small Gaussian field whose covariance is a polynomial in the spin overlap and in the branching level. Fix an enumeration (ϑn)n≥1(\vartheta_n)_{n \ge1} of Q∩[0,1]\mathbb{Q} \cap[0,1] and an integer j+≥1j_{+} \ge1, and let I={1,…,j+}3\mathcal{I} = \{1,\ldots,j_{+}\}^{3}. For j=(j1,j2,j3)∈Ij=(j_1,j_2,j_3) \in\mathcal{I} put

cj(R,q)=(ϑj1R+ϑj2q)j3,so that∣cj(R,q)∣≤2j3(∣R∣≤1, 0≤q≤1).c_j(R,q) = (\vartheta_{j_1}R+\vartheta_{j_2}q)^{j_3}, \qquad\text{so that} \qquad|c_j(R,q)| \le2^{j_3} \quad(|R| \le1,\ 0 \le q \le1).

We use centered Gaussian fields Hj(σ,β)H^j(\sigma,\beta), j∈Ij \in\mathcal{I}, indexed by σ∈SN\sigma\in S_N and leaves β\beta, independent of each other and of all previous variables, with covariance

EHj(σ,β)Hj(σ′,β′)=Ncj(R,J/k),R=σ⋅σ′N,J=β∧β′.(65)\mathbb{E}H^j(\sigma,\beta)H^j(\sigma',\beta') = Nc_j(R,J/k), \qquad R=\frac{\sigma\cdot\sigma'}{N}, \quad J=\beta\wedge\beta'. \tag*{(65)}

This perturbation, together with the prefactor sN=N−1/16s_N=N^{-1/16} used below, is the one of Mourrat [ref-92]. It has the form of Panchenko’s perturbation for multi-species models [ref-100] (see [ref-98] for one species), with the branching level J/kJ/k in place of the overlap of a second species. Panchenko averages over random coefficients in the perturbation; here, as in [ref-92], the perturbation coefficients are chosen at a contact point (§3.4). Mourrat proves that such a field exists in [ref-92]; the construction below places its Gaussian tensors on the nodes of the cascade, so that they become part of the marks. For integrability we use the following explicit construction. For m≥1m \ge1 let am,0=0a_{m,0}=0 and am,l=(l/k)m−((l−1)/k)ma_{m,l}=(l/k)^m-((l-1)/k)^m for 1≤l≤k1 \le l \le k, and let a0,0=1a_{0,0}=1 and a0,l=0a_{0,l}=0 for l≥1l \ge1. Thus am,l≥0a_{m,l} \ge0 and ∑l≤Jam,l=(J/k)m\sum_{l \le J}a_{m,l}=(J/k)^m for all m≥0m \ge0 and 0≤J≤k0 \le J \le k. Attach to each node β\beta independent standard Gaussian tensors gβj,i∈(RN)⊗ig_{\beta}^{j,i} \in(\mathbb{R}^N)^{\otimes i}, 0≤i≤j30 \le i \le j_3, and set

Hj(σ,β)=∑i=0j3∑l=0kcj,i,lN−i/2⟨gβ∣lj,i,σ⊗i⟩,cj,i,l2=N(j3i)ϑj1iϑj2j3−iaj3−i,l.H^j(\sigma,\beta) = \sum_{i=0}^{j_3}\sum_{l=0}^{k} c_{j,i,l}N^{-i/2} \left\langle g_{\beta|l}^{j,i},\sigma^{\otimes i}\right\rangle, \qquad c_{j,i,l}^{2} = N\binom{j_3}{i}\vartheta_{j_1}^{i}\vartheta_{j_2}^{j_3-i}a_{j_3-i,l}.

Only the nodes β∣l\beta|l with l≤β∧β′l \le\beta\wedge\beta' are shared by two leaves, so the covariance is

∑i=0j3(j3i)ϑj1iϑj2j3−iN1−i(σ⋅σ′)i∑l≤Jaj3−i,l=N∑i=0j3(j3i)(ϑj1R)i(ϑj2J/k)j3−i,\sum_{i=0}^{j_3}\binom{j_3}{i}\vartheta_{j_1}^{i}\vartheta_{j_2}^{j_3-i}N^{1-i}(\sigma\cdot\sigma')^i \sum_{l \le J}a_{j_3-i,l} = N\sum_{i=0}^{j_3}\binom{j_3}{i}(\vartheta_{j_1}R)^i(\vartheta_{j_2}J/k)^{j_3-i},

which is (65) by the binomial theorem. Since am,0=0a_{m,0}=0 unless m=0m=0, only the tensor of order i=j3i=j_3 has a root component, with cj,j3,02=Nϑj1j3≤Nc_{j,j_3,0}^{2}=N\vartheta_{j_1}^{j_3} \le N.

The interpolating free energy. Let sN=N−1/16s_N=N^{-1/16} and x=(xj)j∈I∈RIx=(x_j)_{j \in\mathcal{I}} \in\mathbb{R}^{\mathcal{I}}. A configuration is a triple (σ,ξ,β)(\sigma,\xi,\beta) with σ∈SN\sigma\in S_N, ξ∈RM\xi\in\mathbb{R}^M, and β\beta a leaf, with prior μN(dσ) N(0,IM)(dξ) vβ\mu_N(\mathrm{d}\sigma)\,\mathrm{N}(0,I_M)(\mathrm{d}\xi)\,v_\beta. For t∈[0,1]t \in[0,1], p∈Kp \in K, and x∈RIx \in\mathbb{R}^{\mathcal{I}} the Hamiltonian is

H(σ,ξ,β)=∑a=1MU(Sa)+Yβ⋅σ+sN∑j∈IxjHj(σ,β),Sa=t ba⋅σ+1−t(Ξβ,a+1−hk ξa).\begin{aligned} \mathscr{H}(\sigma,\xi,\beta) &= \sum_{a=1}^{M}U(S_a)+Y_\beta\cdot\sigma+s_N\sum_{j \in\mathcal{I}}x_jH^j(\sigma,\beta), \\ S_a &= \sqrt{t}\,b_a \cdot\sigma+\sqrt{1-t}\left(\Xi_{\beta,a}+\sqrt{1-h_k}\,\xi_a\right). \end{aligned}

Let W=∑βvβ∬eH(σ,ξ,β) μN(dσ) N(0,IM)(dξ)W=\sum_{\beta}v_{\beta}\iint e^{\mathscr{H}(\sigma,\xi,\beta)}\,\mu_{N}(\mathrm{d}\sigma)\,\mathrm{N}(0,I_{M})(\mathrm{d}\xi) be the partition function, and put

uN(t,p,x)=1NElog⁡W−pk2,ψσ,N(p)=1NElog⁡∑βvβ∫eYβ⋅σμN(dσ)−pk2.u_{N}(t,p,x)=\frac{1}{N}\mathbb{E}\log W-\frac{p_{k}}{2}, \qquad \psi_{\sigma,N}(p)=\frac{1}{N}\mathbb{E}\log\sum_{\beta}v_{\beta}\int e^{Y_{\beta}\cdot\sigma}\mu_{N}(\mathrm{d}\sigma)-\frac{p_{k}}{2}.

The function ψσ,N\psi_{\sigma,N} is the spherical part of uNu_{N} at t=0t=0 (see (70)), and the subtraction of pk/2p_{k}/2 normalizes it: ψσ,N(0)=0\psi_{\sigma,N}(0)=0. This is a marked cascade in the sense of §3.1. The root mark consists of GG, y∅y_{\varnothing}, z∅z_{\varnothing}, and the root tensors; the mark of a node β\beta of depth l≥1l\geq1 consists of yβy_{\beta}, zβz_{\beta}, and the tensors gβj,ig_{\beta}^{j,i}; and Xk=log⁡∬eH dμN dN(0,IM)X_{k}=\log\iint e^{\mathscr{H}}\,\mathrm{d}\mu_{N}\,\mathrm{d}\mathrm{N}(0,I_{M}) is the logarithm of the leaf integral. From now on ⟨⋅⟩\langle\cdot\rangle denotes the Gibbs measure on configurations, proportional to vβeHv_{\beta}e^{\mathscr{H}} times the product prior, together with its independent replicas (σl,ξl,βl)(\sigma^{l},\xi^{l},\beta^{l}). We write

R12=σ1⋅σ2N,J=J12=β1∧β2,X12=1M∑a=1MU′(Sa1)U′(Sa2),R_{12}=\frac{\sigma^{1}\cdot\sigma^{2}}{N}, \qquad J=J_{12}=\beta^{1}\wedge\beta^{2}, \qquad X_{12}=\frac{1}{M}\sum_{a=1}^{M}U'(S_{a}^{1})U'(S_{a}^{2}),

where SalS_{a}^{l} is SaS_{a} evaluated at the ll-th replica, and

rl=E⟨R12∣J=l⟩,ζl=αNE⟨X12∣J=l⟩.r_{l}=\mathbb{E}\langle R_{12}\mid J=l\rangle, \qquad \zeta_{l}=\alpha_{N}\mathbb{E}\langle X_{12}\mid J=l\rangle.

By Lemma 3.2, E⟨1{J=l}⟩=wl\mathbb{E}\langle\mathbf{1}_{\{J=l\}}\rangle=w_{l} whatever the parameters, and conditional averages given J=lJ=l are as in §3.1. For random variables A,BA,B of two replicas we write Cov⁡(A,B∣J=l)=E⟨AB∣J=l⟩−E⟨A∣J=l⟩E⟨B∣J=l⟩\operatorname{Cov}(A,B\mid J=l)=\mathbb{E}\langle AB\mid J=l\rangle-\mathbb{E}\langle A\mid J=l\rangle\mathbb{E}\langle B\mid J=l\rangle.

The next lemma collects what the comparison uses: continuity, Guerra’s formula for the derivative in tt, the right derivatives in pp in the directions 1≥m\mathbf{1}_{\geq m}, and the monotonicity of ζl\zeta_{l}. For 0≤m≤k0\leq m\leq k let 1≥m∈K\mathbf{1}_{\geq m}\in\mathcal{K} be the vector with coordinates 1{l≥m}\mathbf{1}_{\{l\geq m\}}; moving pp in this direction increases Δpm\Delta p_{m} and leaves the other increments unchanged.

Lemma 3.5. The following hold.

(i) The function uNu_{N} is continuous on [0,1]×K×RI[0,1]\times\mathcal{K}\times\mathbb{R}^{I}, and so is E⟨f(R12,J12)⟩\mathbb{E}\langle f(R_{12},J_{12})\rangle for every bounded measurable function ff on [−1,1]×{0,…,k}[-1,1]\times\{0,\ldots,k\}.

(ii) For 0<t<10<t<1, uNu_{N} is differentiable in tt, and

∂tuN=−αN2E⟨X12(R12−hJ)⟩=−12∑l=0kwl[(rl−hl)ζl+Cov⁡(R12,αNX12∣J=l)].(66)\partial_{t}u_{N} =-\frac{\alpha_{N}}{2}\mathbb{E}\langle X_{12}(R_{12}-h_{J})\rangle =-\frac{1}{2}\sum_{l=0}^{k}w_{l}\left[(r_{l}-h_{l})\zeta_{l}+\operatorname{Cov}\left(R_{12},\alpha_{N}X_{12}\mid J=l\right)\right]. \tag*{(66)}

(iii) For every (t,p,x)(t,p,x) and 0≤m≤k0\leq m\leq k, the function ϵ↦uN(t,p+ϵ1≥m,x)\epsilon\mapsto u_{N}(t,p+\epsilon\mathbf{1}_{\geq m},x) on [0,∞)[0,\infty) has right derivative at ϵ=0\epsilon=0 equal to

−12E⟨R121{J≥m}⟩=−12∑l=mkwlrl.(67)-\frac{1}{2}\mathbb{E}\langle R_{12}\mathbf{1}_{\{J\geq m\}}\rangle =-\frac{1}{2}\sum_{l=m}^{k}w_{l}r_{l}. \tag*{(67)}

(iv) The conditional response overlaps are ordered:

0≤ζ0≤ζ1≤⋯≤ζk≤αNL2.(68)0 \le\zeta_{0} \le\zeta_{1} \le\cdots\le\zeta_{k} \le\alpha_{N}L^{2}. \tag*{(68)}

Proof. Given its leaf β\beta, a replica (σ,ξ)(\sigma,\xi) has the leaf measure, proportional to eH(σ,ξ,β)\mathrm{e}^{\mathcal{H}(\sigma,\xi,\beta)} times the prior μN⊗N(0,IM)\mu_{N}\otimes N(0,I_{M}). It depends only on the path of marks e[β]e_{[\beta]}, and we write ⟨⋅⟩β\langle\cdot\rangle_{\beta} for averages under it. Given the leaves, replicas are independent. We divide the proof into six steps.

Step 1. Integrability. For a node β\beta of depth mm let Θm(eβ)\Theta_{m}(e_{\beta}) be the sum of the Euclidean norms of yβy_{\beta}, zβz_{\beta}, and all tensors gβj,ig_{\beta}^{j,i}; at the root add the Frobenius norm ∥G∥F\lVert\mathbf{G}\rVert_{F}. On any compact set of parameters there is CC, depending also on NN, such that

∣Xl(e0,…,el)∣≤C(1+∑m=0lΘm(em)),0≤l≤k.(69)\left|X_{l}(e_{0},\ldots,e_{l})\right| \le C\left(1+\sum_{m=0}^{l}\Theta_{m}(e_{m})\right), \qquad0 \le l \le k. \tag*{(69)}

For l=kl=k, the upper bound follows from U≤Umax⁡U\le U_{\max}, ∣Yβ⋅σ∣≤N∣Yβ∣|Y_{\beta}\cdot\sigma|\le\sqrt{N}|Y_{\beta}|, and ∣⟨g,σ⊗i⟩∣≤Ni/2∥g∥|\langle g,\sigma^{\otimes i}\rangle|\le N^{i/2}\lVert g\rVert, which hold uniformly in (σ,ξ)(\sigma,\xi). For the lower bound, apply Jensen’s inequality to the leaf integral and use U(y)≥U(0)−L∣y∣U(y)\ge U(0)-L|y|, ∫∣ba⋅σ∣ dμN≤∣ba∣\int|b_{a}\cdot\sigma|\,\mathrm{d}\mu_{N}\le|b_{a}| (since ∫σσT dμN=IN\int\sigma\sigma^{\mathsf{T}}\,\mathrm{d}\mu_{N}=I_{N}), ∫∣ξa∣ dN(0,IM)≤1\int|\xi_{a}|\,\mathrm{d}N(0,I_{M})\le1, ∫Yβ⋅σ dμN=0\int Y_{\beta}\cdot\sigma\,\mathrm{d}\mu_{N}=0, and the bound on the tensors. Backward induction in (47) gives (69) at every level: Jensen’s inequality gives the lower bound, and the upper bound uses that each Θm\Theta_{m} has finite exponential moments of all orders. In particular the hypotheses of Lemma 3.2 hold, and each kernel density eθl(Xl−Xl−1)\mathrm{e}^{\theta_{l}(X_{l}-X_{l-1})} is at most exp⁡{C(1+∑m≤lΘm)}\exp\{C(1+\sum_{m\le l}\Theta_{m})\}.

Step 2. Continuity. Fix the marks. The leaf integral is continuous in (t,p,x)(t,p,x) by dominated convergence, since its integrand is continuous in the parameters and bounded, uniformly in (σ,ξ)(\sigma,\xi) and on compact parameter sets, by exp⁡{C(1+∑mΘm)}\exp\{C(1+\sum_{m}\Theta_{m})\}. Backward induction in (47), with dominated convergence justified by (69), shows that every XlX_{l} and every kernel density is continuous in the parameters. Since Elog⁡W=EX0(e0)\mathbb{E}\log W=\mathbb{E}X_{0}(e_{0}) by (50) and ∣X0∣≤C(1+Θ0)|X_{0}|\le C(1+\Theta_{0}), uNu_{N} is continuous. For the second claim in (i), let φ(e,e′)\varphi(e,e') be the average of f(σ⋅σ′/N,l)f(\sigma\cdot\sigma'/N,l) under the product of the leaf measures of two paths e,e′e,e'. It is bounded by sup⁡∣f∣\sup|f| and continuous in the parameters. By (49), applied for a fixed root mark, E⟨f(R12,J)1{J=l}⟩=wlEEK,lφ\mathbb{E}\langle f(R_{12},J)\mathbf{1}_{\{J=l\}}\rangle=w_{l}\mathbb{E}\mathbb{E}_{K,l}\varphi, and EK,l\mathbb{E}_{K,l} integrates φ\varphi against a product of kernel densities, each continuous and bounded by exp⁡{C(1+∑mΘm)}\exp\{C(1+\sum_{m}\Theta_{m})\}. Dominated convergence proves (i).

Step 3. Differentiation under the recursion. Let ∂\partial denote the derivative in one of the parameters t∈(0,1)t\in(0,1) and Δpm\Delta p_{m}, the latter in the range Δpm>0\Delta p_{m}>0. Locally uniformly in the parameters, ∣∂H∣≤C(1+∣ξ∣+∑mΘm)|\partial\mathcal{H}|\le C(1+|\xi|+\sum_{m}\Theta_{m}). Hence ∂Xk=⟨∂H⟩β\partial X_{k}=\langle\partial\mathcal{H}\rangle_{\beta}, and, bounding ⟨∣ξ∣⟩β≤esup⁡H−Xk∫∣ξ∣ dN(0,IM)\langle|\xi|\rangle_{\beta}\le\mathrm{e}^{\sup\mathcal{H}-X_{k}}\int|\xi|\,\mathrm{d}N(0,I_{M}) with (69), we get ∣∂Xk∣≤exp⁡{C(1+∑mΘm)}|\partial X_{k}|\le\exp\{C(1+\sum_{m}\Theta_{m})\}. Differentiating (47) gives ∂Xl−1=∫∂Xl dKl\partial X_{l-1}=\int\partial X_{l}\,\mathrm{d}K_{l}, and the same kind of bound propagates to every level, which justifies each differentiation under the integral. Hence ∂X0=EK∂Xk\partial X_{0}=\mathbb{E}_{K}\partial X_{k}, and by (50) and (48)

∂Elog⁡W=EEK∂Xk=E⟨∂H⟩.\partial\mathbb{E}\log W=\mathbb{E}\mathbb{E}_{K}\partial X_{k}=\mathbb{E}\langle\partial\mathcal{H}\rangle.

Step 4. Gaussian integration by parts. Each Gaussian coordinate at a node of depth mm enters the leaf integrals of all leaves below that node, and only those. Fix one such coordinate and let it vary in a compact set. Since ∣U′∣≤L|U'|\leq L and ∣σ∣=N|\sigma|=\sqrt{N}, the derivative of each descendant Hamiltonian in this coordinate is bounded, uniformly over the spin and auxiliary variables. Thus the perturbed descendant leaf integrals are bounded above and below by common finite multiples of their original values, their first derivatives satisfy the same domination, and so do their second derivatives, by the bound on U′′U''. The series over leaves defining the Gibbs averages therefore converge locally uniformly together with their derivatives, and by (48) and (69) the sum over nodes of the expected absolute contributions is finite. This justifies Gaussian integration by parts node by node. Summing over nodes and coordinates gives, for every row aa and every 0≤m≤k0\leq m\leq k,

E⟨yβ∣m⋅σ⟩=ΔpmN(1−E⟨R121{J≥m}⟩),E⟨zβ∣m,aU′(Sa)⟩=(1−t)Δhm(E⟨U′′(Sa)+U′(Sa)2⟩−E⟨U′(Sa1)U′(Sa2)1{J≥m}⟩),E⟨ζaU′(Sa)⟩=(1−t)(1−hk) E⟨U′′(Sa)+U′(Sa)2⟩,∑i=1NE⟨gaiU′(Sa)σi⟩=tN(E⟨U′′(Sa)+U′(Sa)2⟩−E⟨U′(Sa1)U′(Sa2)R12⟩).\begin{aligned} \mathbb{E}\langle y_{\beta|m}\cdot\sigma\rangle &=\sqrt{\Delta p_{m}}N\left(1-\mathbb{E}\langle R_{12}\mathbf{1}_{\{J\geq m\}}\rangle\right),\\ \mathbb{E}\langle z_{\beta|m,a}U'(S_{a})\rangle &=\sqrt{(1-t)\Delta h_{m}}\left(\mathbb{E}\langle U''(S_{a})+U'(S_{a})^{2}\rangle-\mathbb{E}\langle U'(S_{a}^{1})U'(S_{a}^{2})\mathbf{1}_{\{J\geq m\}}\rangle\right),\\ \mathbb{E}\langle\zeta_{a}U'(S_{a})\rangle &=\sqrt{(1-t)(1-h_{k})}\,\mathbb{E}\langle U''(S_{a})+U'(S_{a})^{2}\rangle,\\ \sum_{i=1}^{N}\mathbb{E}\langle g_{ai}U'(S_{a})\sigma_{i}\rangle &=\sqrt{tN}\left(\mathbb{E}\langle U''(S_{a})+U'(S_{a})^{2}\rangle-\mathbb{E}\langle U'(S_{a}^{1})U'(S_{a}^{2})R_{12}\rangle\right). \end{aligned}

Here β\beta is the leaf of the sampled configuration. In the first line, the derivative of ⟨σi1{β∣m=ν}⟩\langle\sigma_{i}\mathbf{1}_{\{\beta|m=\nu\}}\rangle in the coordinate yν,iy_{\nu,i} produces the one-replica term Δpm⟨σi21{β∣m=ν}⟩\sqrt{\Delta p_{m}}\langle\sigma_{i}^{2}\mathbf{1}_{\{\beta|m=\nu\}}\rangle and the two-replica term −Δpm⟨σi1σi21{β1∣m=β2∣m=ν}⟩-\sqrt{\Delta p_{m}}\langle\sigma_{i}^{1}\sigma_{i}^{2}\mathbf{1}_{\{\beta^{1}|m=\beta^{2}|m=\nu\}}\rangle; summing over ii and over the nodes ν\nu of depth mm gives the formula, because ∣σ∣2=N|\sigma|^{2}=N and 1{β1∣m=β2∣m}=1{J≥m}\mathbf{1}_{\{\beta^{1}|m=\beta^{2}|m\}}=\mathbf{1}_{\{J\geq m\}}. The second and fourth lines are obtained in the same way; there the factor U′′+(U′)2U''+(U')^{2} comes from differentiating U′(Sa)eU(Sa)U'(S_{a})e^{U(S_{a})}. The variable ζ\zeta is integrated inside the leaf measure, so in the third line only the one-replica term appears.

Step 5. The derivative in tt. For 0<t<10<t<1, ∂tH=∑aU′(Sa)∂tSa\partial_{t}\mathcal{H}=\sum_{a}U'(S_{a})\partial_{t}S_{a} with

∂tSa=ba⋅σ2t−Ξβ,a+1−hk ζa21−t,ba⋅σ=N−1/2∑igaiσi,Ξβ,a=∑mΔhm zβ∣m,a.\partial_{t}S_{a} =\frac{b_{a}\cdot\sigma}{2\sqrt{t}} -\frac{\Xi_{\beta,a}+\sqrt{1-h_{k}}\,\zeta_{a}}{2\sqrt{1-t}}, \qquad b_{a}\cdot\sigma=N^{-1/2}\sum_{i}g_{ai}\sigma_{i}, \qquad \Xi_{\beta,a}=\sum_{m}\sqrt{\Delta h_{m}}\,z_{\beta|m,a}.

The last three identities of Step 4, together with ∑mΔhm=hk\sum_{m}\Delta h_{m}=h_{k} and ∑mΔhm1{J≥m}=hJ\sum_{m}\Delta h_{m}\mathbf{1}_{\{J\geq m\}}=h_{J}, give for each row aa

E⟨U′(Sa)∂tSa⟩=12(E⟨U′′+(U′)2⟩−E⟨U′(Sa1)U′(Sa2)R12⟩)−12(E⟨U′′+(U′)2⟩−E⟨U′(Sa1)U′(Sa2)hJ⟩),\begin{aligned} \mathbb{E}\langle U'(S_{a})\partial_{t}S_{a}\rangle ={}&\frac{1}{2}\left(\mathbb{E}\langle U''+(U')^{2}\rangle-\mathbb{E}\langle U'(S_{a}^{1})U'(S_{a}^{2})R_{12}\rangle\right)\\ &-\frac{1}{2}\left(\mathbb{E}\langle U''+(U')^{2}\rangle-\mathbb{E}\langle U'(S_{a}^{1})U'(S_{a}^{2})h_{J}\rangle\right), \end{aligned}

with U′′+(U′)2U''+(U')^{2} evaluated at SaS_{a}. The one-replica terms cancel because ba⋅σb_{a}\cdot\sigma and Ξβ,a+1−hk ξa\Xi_{\beta,a}+\sqrt{1-h_{k}}\,\xi_{a} both have variance one. Summing over aa and dividing by NN gives the first identity in (66), by Step 3; the term −pk/2-p_{k}/2 of uNu_{N} does not depend on tt. For the second, split the average according to the value of JJ and use E⟨R12αNX12∣J=l⟩=rlζl+Cov⁡(R12,αNX12∣J=l)\mathbb{E}\langle R_{12}\alpha_{N}X_{12}\mid J=l\rangle=r_{l}\zeta_{l}+\operatorname{Cov}(R_{12},\alpha_{N}X_{12}\mid J=l). This proves (ii).

Step 6. The derivatives in pp and the monotonicity. Moving pp along 1≥m\mathbf{1}_{\ge m} increases Δpm\Delta p_{m} and pkp_{k} at unit rate and leaves the other increments fixed. If Δpm>0\Delta p_{m}>0, then ∂ΔpmH=yβ∣m⋅σ/(2Δpm)\partial_{\Delta p_{m}}\mathcal{H}=y_{\beta\mid m}\cdot\sigma/(2\sqrt{\Delta p_{m}}), and the first identity of Step 4 and Step 3 show that 1NElog⁡W\frac{1}{N}\mathbb{E}\log W has derivative 12(1−E⟨R121{J≥m}⟩)\frac{1}{2}(1-\mathbb{E}\langle R_{12}\mathbf{1}_{\{J\ge m\}}\rangle); subtracting 12\frac{1}{2} for the term −pk/2-p_{k}/2 gives (67). If Δpm=0\Delta p_{m}=0, this formula holds at p+ϵ1≥mp+\epsilon\mathbf{1}_{\ge m} for every ϵ>0\epsilon>0, and by (i), applied with f(R,l)=R1{l≥m}f(R,l)=R\mathbf{1}_{\{l\ge m\}}, both uNu_{N} and the right side of (67) are continuous as ϵ↓0\epsilon\downarrow0. The mean value theorem then gives the right derivative at ϵ=0\epsilon=0. Finally, E⟨R121{J≥m}⟩=∑l≥mwlrl\mathbb{E}\langle R_{12}\mathbf{1}_{\{J\ge m\}}\rangle=\sum_{l\ge m}w_{l}r_{l} by the definition of rlr_{l}. This proves (iii).

For (iv), fix the root mark and apply Corollary 3.3 to the path function with coordinates ⟨U′(Sa)⟩β/M\langle U'(S_{a})\rangle_{\beta}/\sqrt{M}, a≤Ma\le M, whose norm is at most LL. Since replicas are independent given their leaves, ⟨U′(Sa1)U′(Sa2)⟩\langle U'(S_{a}^{1})U'(S_{a}^{2})\rangle given the two leaves equals ⟨U′(Sa)⟩β1⟨U′(Sa)⟩β2\langle U'(S_{a})\rangle_{\beta^{1}}\langle U'(S_{a})\rangle_{\beta^{2}}, so for fixed root mark E⟨X12∣J=l⟩\mathbb{E}\langle X_{12}\mid J=l\rangle is the quantity EK∥Fl∥2\mathbb{E}_{\mathsf{K}}\lVert F_{l}\rVert^{2} of that corollary. It is nonnegative, nondecreasing in ll, and at most L2L^{2}. These properties survive averaging over the root mark, because wlw_{l} does not depend on it. This proves (68). □\square

The endpoints. At t=1t=1, p=0p=0, and x=0x=0 the Hamiltonian is ∑aU(ba⋅σ)\sum_{a}U(b_{a}\cdot\sigma), which depends neither on ξ~\widetilde{\xi} nor on the leaf, so W=ZN(U)W=Z_{N}(U). At t=0t=0 and x=0x=0 the leaf integral is the product of ∫eYβ⋅σμN(dσ)\int e^{Y_{\beta}\cdot\sigma}\mu_{N}(\mathrm{d}\sigma), which depends only on the vectors yy, and of

∏a=1M∫exp⁡{U(Ξβ,a+1−hk ξa)}N(0,1)(dξa),\prod_{a=1}^{M}\int\exp\left\{U\left(\Xi_{\beta,a}+\sqrt{1-h_{k}}\,\xi_{a}\right)\right\}N(0,1)(\mathrm{d}\xi_{a}),

which depends only on the vectors zz; these are independent components of the marks. For independent F,F′F,F' we have θ−1log⁡Eeθ(F+F′)=θ−1log⁡EeθF+θ−1log⁡EeθF′\theta^{-1}\log\mathbb{E}e^{\theta(F+F')}=\theta^{-1}\log\mathbb{E}e^{\theta F}+\theta^{-1}\log\mathbb{E}e^{\theta F'}, so the recursion (47) factorizes into a spin part and a row part, and by (50) so does Elog⁡W\mathbb{E}\log W. The spin part is Nψσ,N(p)+Npk/2N\psi_{\sigma,N}(p)+Np_{k}/2. The row part factorizes in the same way over the rows, whose coordinates z⋅,az_{\cdot,a} of the node vectors are independent. For one row it is the recursion (17) for the order parameter γh\gamma_{h}, evaluated at (0,0)(0,0): the auxiliary variable gives the step T1,1−hk\mathcal{T}_{1,1-h_{k}} over [hk,1][h_{k},1]; the level l≥1l\ge1 gives the step Tθl,Δhl\mathcal{T}_{\theta_{l},\Delta h_{l}} over [hl−1,hl)[h_{l-1},h_{l}), where γh=θl\gamma_{h}=\theta_{l}; and the root gives the heat step T0,h0\mathcal{T}_{0,h_{0}} over [0,h0)[0,h_{0}), where γh=0\gamma_{h}=0. Levels with Δhl=0\Delta h_{l}=0 give empty intervals and the identity step Tθl,0\mathcal{T}_{\theta_{l},0}, so they can be omitted. Hence

uN(0,p,0)=ψσ,N(p)+αNψγh(0,0;U),uN(1,0,0)=FN(U).(70)u_{N}(0,p,0)=\psi_{\sigma,N}(p)+\alpha_{N}\psi_{\gamma_{h}}(0,0;U),\qquad u_{N}(1,0,0)=F_{N}(U). \tag*{(70)}

The identification of the row part with the recursion (17) is the representation of the Parisi functional through Ruelle probability cascades (see [ref-98] Chapter 2 and [ref-4]).

Concentration and identities at a contact

The comparison argument evaluates the interpolation only at maximum points of a deterministic function of the parameters. At such points we need the Ghirlanda–Guerra identities for the perturbed Gibbs measure to hold approximately. They follow from concentration of the derivatives of log⁡W\log W in the perturbation parameters xx. Since the points are deterministic, concentration at each fixed parameter suffices.

This section follows Steps 1–3 of the proof of [ref-92], Theorem 4.1. There a quadratic penalty around the point (1,…,1)(1,\ldots,1) [ref-92], (4.18)–(4.19) keeps the perturbation coefficients at the contact point close to one; the bound on the Hessian in xx at a contact point and the concentration of the free energy [ref-92], Proposition 4.2 control the fluctuations of the perturbation Hamiltonians; and Gaussian integration by parts gives approximate Ghirlanda–Guerra identities [ref-92], (4.28). Lemmas 3.6 and 3.7 carry out these steps for the present interpolation, and the penalty appears in the proof of Proposition 3.10.

Lemma 3.6. Fix KK, hh, and I\mathcal{I}, and a compact subset of [0,1]×K×RI[0,1]\times K\times\mathbb{R}^{\mathcal{I}}. There is a constant CC such that

Var⁡(1Nlog⁡W)≤CN\operatorname{Var}\left(\frac{1}{N}\log W\right)\leq\frac{C}{N}

on this set. Moreover, almost surely x↦log⁡Wx\mapsto\log W is convex and C2C^{2} on RI\mathbb{R}^{\mathcal{I}}, with

∂xjlog⁡W=sN⟨Hj⟩,(71)\partial_{x_j}\log W=s_N\langle H^j\rangle, \tag*{(71)}
∂xj2log⁡W=sN2Var⁡⟨⋅⟩(Hj).(72)\partial_{x_j}^{2}\log W=s_N^{2}\operatorname{Var}_{\langle\cdot\rangle}(H^j). \tag*{(72)}

The function uNu_N is convex and differentiable in xx, with N∂xjuN=sNE⟨Hj⟩N\partial_{x_j}u_N=s_N\mathbb{E}\langle H^j\rangle, ∣∂xjuN∣≤2j++1N−1/8∣xj∣|\partial_{x_j}u_N|\leq2^{j_++1}N^{-1/8}|x_j|, and, with Cj+=2j+C_{j_+}=2^{j_+},

∣uN(t,p,x)−uN(t,p,0)∣≤Cj+N−1/8∣x∣2((t,p,x)∈[0,1]×K×RI).(73)|u_N(t,p,x)-u_N(t,p,0)|\leq C_{j_+}N^{-1/8}|x|^{2} \qquad\left((t,p,x)\in[0,1]\times K\times\mathbb{R}^{\mathcal{I}}\right). \tag*{(73)}

Proof. We divide the proof into three steps.

Step 1. The variance. By the law of total variance, Var⁡(log⁡W)=EVar⁡(log⁡W∣e0)+Var⁡(E[log⁡W∣e0])\operatorname{Var}(\log W)=\mathbb{E}\operatorname{Var}(\log W\mid e_0)+\operatorname{Var}(\mathbb{E}[\log W\mid e_0]). By (51) the first term is at most 4Var⁡(log⁡Π)4\operatorname{Var}(\log\Pi), a constant depending only on the cascade, and E[log⁡W∣e0]=X0(e0)\mathbb{E}[\log W\mid e_0]=X_0(e_0) is a function of the root Gaussian variables G\boldsymbol{G}, y∅y_{\varnothing}, z∅z_{\varnothing}, and the root tensors. As in Steps 2 and 3 of the proof of Lemma 3.5, applied to these variables (the derivatives of H\mathcal{H} in them are bounded), X0X_0 is continuously differentiable in them and its gradient is the EK\mathbb{E}_K-average of the gradient of the leaf logarithm, so by Jensen’s inequality its squared norm is at most the EK\mathbb{E}_K-average of the squared norm of the leaf gradient. The derivative of the leaf logarithm in gaig_{ai} is t/N⟨U′(Sa)σi⟩β\sqrt{t/N}\langle U'(S_a)\sigma_i\rangle_\beta; by Jensen’s inequality and ∣σ∣2=N|\sigma|^2=N the sum of its squares over a,ia,i is at most (t/N)MNL2=NtαNL2(t/N)MNL^2=Nt\alpha_NL^2. The derivative in y∅y_{\varnothing} is p0⟨σ⟩β\sqrt{p_0}\langle\sigma\rangle_\beta, with squared norm at most Np0Np_0. The derivative in z∅,az_{\varnothing,a} is (1−t)h0⟨U′(Sa)⟩β\sqrt{(1-t)h_0}\langle U'(S_a)\rangle_\beta, with squared norm summed over aa at most NαNh0L2N\alpha_{N}h_{0}L^{2}. The root tensor of HjH^{j} has coefficient cj,j3,02≤Nc_{j,j_{3},0}^{2}\leq N, and the derivative of the leaf logarithm in it has squared norm at most sN2xj2cj,j3,02N−j3∣σ∣2j3≤sN2Nxj2s_{N}^{2}x_{j}^{2}c_{j,j_{3},0}^{2}N^{-j_{3}}|\sigma|^{2j_{3}}\leq s_{N}^{2}Nx_{j}^{2}. Altogether

∣∇X0∣2≤N(tαNL2+p0+αNh0L2+N−1/8∣x∣2),|\nabla X_{0}|^{2}\leq N\left(t\alpha_{N}L^{2}+p_{0}+\alpha_{N}h_{0}L^{2}+N^{-1/8}|x|^{2}\right),

and the Gaussian Poincaré inequality [ref-21] gives Var⁡(X0)≤E∣∇X0∣2≤CN\operatorname{Var}(X_{0})\leq\mathbb{E}|\nabla X_{0}|^{2}\leq CN on the compact set. Hence Var⁡(N−1log⁡W)≤N−2(4Var⁡log⁡Π+CN)≤C/N\operatorname{Var}(N^{-1}\log W)\leq N^{-2}(4\operatorname{Var}\log\Pi+CN)\leq C/N.

Step 2. The derivatives in xx. Let Θˉβ=∑j∈I∑i,l∣cj,i,l∣∥gβ∣lj,i∥\bar{\Theta}_{\beta}=\sum_{j\in\mathcal{I}}\sum_{i,l}|c_{j,i,l}|\|g_{\beta|l}^{j,i}\|, so that ∑jsup⁡σ∣Hj(σ,β)∣≤Θˉβ\sum_{j}\sup_{\sigma}|H^{j}(\sigma,\beta)|\leq\bar{\Theta}_{\beta}, and let Aβ(x)A_{\beta}(x) be the leaf integral. For each x0∈QIx_{0}\in\mathbb{Q}^{\mathcal{I}} and each integer n≥1n\geq1,

∑βvβAβ(x0)(1+Θˉβ+Θˉβ2)esNnΘˉβ<∞almost surely.\sum_{\beta}v_{\beta}A_{\beta}(x_{0})(1+\bar{\Theta}_{\beta}+\bar{\Theta}_{\beta}^{2})e^{s_{N}n\bar{\Theta}_{\beta}}<\infty \qquad\text{almost surely}.

Indeed, the sum equals W(x0)W(x_{0}) times a Gibbs average whose expectation, by (48), is an EK\mathbb{E}_{K}-expectation of (1+Θˉ+Θˉ2)esNnΘˉ(1+\bar{\Theta}+\bar{\Theta}^{2})e^{s_{N}n\bar{\Theta}}, which is finite by the kernel density bound of Step 1 in the proof of Lemma 3.5. On the intersection of these countably many events, since Aβ(x)≤Aβ(x0)esN∣x−x0∣1ΘˉβA_{\beta}(x)\leq A_{\beta}(x_{0})e^{s_{N}|x-x_{0}|_{1}\bar{\Theta}_{\beta}}, the series defining W(x)W(x) and the series of its first two derivatives converge locally uniformly in xx. Termwise differentiation gives (71) and (72). More generally, the Hessian of log⁡W\log W in xx is sN2s_{N}^{2} times the covariance matrix of (Hj)j∈I(H^{j})_{j\in\mathcal{I}} under ⟨⋅⟩\langle\cdot\rangle, which is positive semidefinite, so log⁡W\log W is convex in xx. For a convex function the difference quotients (log⁡W(x)−log⁡W(x−ϵej))/ϵ(\log W(x)-\log W(x-\epsilon\mathbf{e}_{j}))/\epsilon and (log⁡W(x+ϵej)−log⁡W(x))/ϵ(\log W(x+\epsilon\mathbf{e}_{j})-\log W(x))/\epsilon, where ej\mathbf{e}_{j} is the jj-th coordinate vector of RI\mathbb{R}^{\mathcal{I}}, bound ∂xjlog⁡W\partial_{x_{j}}\log W from below and above, are integrable, and converge monotonically as ϵ↓0\epsilon\downarrow0. Hence ∂xjElog⁡W=E∂xjlog⁡W\partial_{x_{j}}\mathbb{E}\log W=\mathbb{E}\partial_{x_{j}}\log W, and uNu_{N} is convex in xx. A convex function on RI\mathbb{R}^{\mathcal{I}} whose partial derivatives exist is differentiable, so uNu_{N} is differentiable in xx.

Step 3. Replica integration by parts. For a bounded measurable function ff of the overlap arrays (Rll′,Jll′)l,l′≤n(R_{ll'},J_{ll'})_{l,l'\leq n} of nn replicas, Gaussian integration by parts in the tensors gives

E⟨fHj(1)⟩=sNxjN(∑l=1nE⟨f cj(R1l,J1l/k)⟩−nE⟨f cj(R1,n+1,J1,n+1/k)⟩),(74)\mathbb{E}\langle fH^{j}(1)\rangle = s_{N}x_{j}N\left( \sum_{l=1}^{n}\mathbb{E}\langle f\,c_{j}(R_{1l},J_{1l}/k)\rangle - n\mathbb{E}\langle f\,c_{j}(R_{1,n+1},J_{1,n+1}/k)\rangle \right), \tag*{(74)}

where Hj(1)=Hj(σ1,β1)H^{j}(1)=H^{j}(\sigma^{1},\beta^{1}). To justify the exchange of the sum over nodes with the expectation, differentiate with respect to one tensor gνj,ig_{\nu}^{j,i} at one node ν\nu of depth mm. Each replica ll contributes the indicator that βl\beta^{l} descends from ν\nu, and contracting the two tensors gives NiR1liN^{i}R_{1l}^{i}. Summing the indicators over the nodes of depth mm gives 1{J1l≥m}\mathbf{1}_{\{J_{1l}\geq m\}}, and summing the squared coefficients over ii and mm gives the covariance (65). The expected absolute sum over nodes is bounded by a constant times E⟨∥gβ∣mj,i∥⟩\mathbb{E}\langle\|g_{\beta|m}^{j,i}\|\rangle, which is finite by (48) and (69). Taking n=1n=1 and f=1f=1 in (74), and using ∣cj∣≤2j3|c_{j}|\leq2^{j_{3}}, gives ∣E⟨Hj⟩∣≤2j3+1sN∣xj∣N|\mathbb{E}\langle H^{j}\rangle|\leq2^{j_{3}+1}s_{N}|x_{j}|N, hence ∣∂xjuN∣=sNN−1∣E⟨Hj⟩∣≤2j3+1N−1/8∣xj∣|\partial_{x_{j}}u_{N}|=s_{N}N^{-1}|\mathbb{E}\langle H^{j}\rangle|\leq2^{j_{3}+1}N^{-1/8}|x_{j}|. Integrating along the segment from 00 to xx gives (73), since j3≤j+j_{3}\leq j_{+}. □\square

For two families of overlap arrays All′,Bll′A^{ll'}, B^{ll'}, a test function ff of the arrays of the first nn replicas, and a function cc of two real variables, the Ghirlanda–Guerra discrepancy is

Dn(f,c)=E⟨fc(A1,n+1,B1,n+1)⟩−1nE⟨f⟩E⟨c(A12,B12)⟩−1n∑l=2nE⟨fc(A1l,B1l)⟩.\mathcal{D}_{n}(f,c)=\mathbb{E}\langle f c(A^{1,n+1},B^{1,n+1})\rangle-\frac{1}{n}\mathbb{E}\langle f\rangle\mathbb{E}\langle c(A^{12},B^{12})\rangle-\frac{1}{n}\sum_{l=2}^{n}\mathbb{E}\langle f c(A^{1l},B^{1l})\rangle.

The Ghirlanda–Guerra identities [ref-64] state that Dn(f,c)=0\mathcal{D}_{n}(f,c)=0. For c=cjc=c_{j} this is the form used in [ref-92], (4.28) and (5.6), and it is the analogue, for two overlap arrays, of the multi-species discrepancy in [ref-100], (31).

Lemma 3.7. Fix a compact set of parameters (t,p)∈[0,1]×K(t,p)\in[0,1]\times K. There is a constant CC with the following property. Suppose that (t,p)(t,p) lies in this set, that ∣x−1∣≤1/2|x-\mathbf{1}|\leq1/2, where 1∈RI\mathbf{1}\in\mathbb{R}^{\mathcal{I}} has all coordinates equal to one, and that for every y∈RIy\in\mathbb{R}^{\mathcal{I}}

0≤uN(x+y)−uN(x)−∇xuN(x)⋅y≤∣y∣2,(75)0\leq u_{N}(x+y)-u_{N}(x)-\nabla_{x}u_{N}(x)\cdot y\leq|y|^{2}, \tag*{(75)}

where uN(x)u_{N}(x) abbreviates uN(t,p,x)u_{N}(t,p,x). Then for A=RA=R and B=J/kB=J/k, every j∈Ij\in\mathcal{I}, every n≥1n\geq1, and every measurable function ff of the overlap arrays of nn replicas with ∣f∣≤1|f|\leq1,

∣Dn(f,cj)∣≤CN−1/8.|\mathcal{D}_{n}(f,c_{j})|\leq CN^{-1/8}.

Proof. We show that HjH^{j} has small fluctuations, both under the Gibbs measure and over the disorder, at the scale relevant to (3.29). We divide the proof into three steps.

Step 1. Thermal fluctuation. Adding (75) at y=±ϵejy=\pm\epsilon e_{j} gives uN(x+ϵej)+uN(x−ϵej)−2uN(x)≤2ϵ2u_{N}(x+\epsilon e_{j})+u_{N}(x-\epsilon e_{j})-2u_{N}(x)\leq2\epsilon^{2}, and the expectation of the corresponding second difference of log⁡W\log W is NN times the left side. This second difference is nonnegative by convexity, and, divided by ϵ2\epsilon^{2}, it converges to sN2Var⁡⟨⋅⟩(Hj)s_{N}^{2}\operatorname{Var}_{\langle\cdot\rangle}(H^{j}) by (3.27). Fatou’s lemma gives EVar⁡⟨⋅⟩(Hj)≤2NsN−2=2N9/8\mathbb{E}\operatorname{Var}_{\langle\cdot\rangle}(H^{j})\leq2Ns_{N}^{-2}=2N^{9/8}.

Step 2. Disorder fluctuation. Put u^N=N−1log⁡W−pk/2\widehat{u}_{N}=N^{-1}\log W-pk/2, so that uN=Eu^Nu_{N}=\mathbb{E}\widehat{u}_{N}, and fix ϵ∈(0,1]\epsilon\in(0,1]. Let Δ=max⁡∣u^N(z)−uN(z)∣\Delta=\max|\widehat{u}_{N}(z)-u_{N}(z)| over the three points z∈{x,x+ϵej,x−ϵej}z\in\{x,x+\epsilon e_{j},x-\epsilon e_{j}\}. By convexity of u^N\widehat{u}_{N} in xjx_{j} and the upper bound in (75),

∂ju^N(x)≤u^N(x+ϵej)−u^N(x)ϵ≤uN(x+ϵej)−uN(x)+2Δϵ≤∂juN(x)+ϵ+2Δϵ,\partial_{j}\widehat{u}_{N}(x)\leq\frac{\widehat{u}_{N}(x+\epsilon e_{j})-\widehat{u}_{N}(x)}{\epsilon}\leq\frac{u_{N}(x+\epsilon e_{j})-u_{N}(x)+2\Delta}{\epsilon}\leq\partial_{j}u_{N}(x)+\epsilon+\frac{2\Delta}{\epsilon},

and in the same way, with the secant through x−ϵejx-\epsilon e_{j}, ∂ju^N(x)≥∂juN(x)−ϵ−2Δ/ϵ\partial_{j}\widehat{u}_{N}(x)\geq\partial_{j}u_{N}(x)-\epsilon-2\Delta/\epsilon. The three points are deterministic and lie in a fixed compact set, so Lemma 3.6 gives EΔ2≤C/N\mathbb{E}\Delta^{2}\leq C/N. With ϵ=N−1/4\epsilon=N^{-1/4},

E(∂ju^N(x)−∂juN(x))2≤2ϵ2+8EΔ2ϵ2≤CN−1/2.\mathbb{E}\left(\partial_{j}\widehat{u}_{N}(x)-\partial_{j}u_{N}(x)\right)^{2}\leq2\epsilon^{2}+\frac{8\mathbb{E}\Delta^{2}}{\epsilon^{2}}\leq CN^{-1/2}.

By (3.26), ∂ju^N=N−1sN⟨Hj⟩\partial_{j}\widehat{u}_{N}=N^{-1}s_{N}\langle H^{j}\rangle and ∂juN=N−1sNE⟨Hj⟩\partial_{j}u_{N}=N^{-1}s_{N}\mathbb{E}\langle H^{j}\rangle, so E(⟨Hj⟩−E⟨Hj⟩)2≤N2sN−2CN−1/2=CN13/8\mathbb{E}(\langle H^{j}\rangle-\mathbb{E}\langle H^{j}\rangle)^{2}\leq N^{2}s_{N}^{-2}CN^{-1/2}=CN^{13/8}. Adding the thermal fluctuation, E⟨(Hj−E⟨Hj⟩)2⟩≤CN13/8\mathbb{E}\langle(H^{j}-\mathbb{E}\langle H^{j}\rangle)^{2}\rangle\leq CN^{13/8}.

Step 3. Conclusion. Subtract from (3.29) for nn replicas the same identity with n=1n=1 and f=1f=1, multiplied by E⟨f⟩\mathbb{E}\langle f\rangle. The terms with l=1l=1 cancel, because cj(R11,J11/k)=cj(1,1)c_j(R_{11},J_{11}/k)=c_j(1,1) is a constant, and what remains is

E⟨f(Hj(1)−E⟨Hj⟩)⟩=−sNxjNn Dn(f,cj).\mathbb{E}\left\langle f\left(H^j(1)-\mathbb{E}\langle H^j\rangle\right)\right\rangle=-s_Nx_jNn\,\mathcal{D}_n(f,c_j).

By the Cauchy–Schwarz inequality the left side is at most CN13/16CN^{13/16}. Dividing by nsNxjN≥12N15/16ns_Nx_jN\ge\frac{1}{2}N^{15/16}, which uses xj≥1/2x_j\ge1/2, gives ∣Dn(f,cj)∣≤CN−1/8\lvert\mathcal{D}_n(f,c_j)\rvert\le CN^{-1/8}. □

Synchronization

We use Panchenko’s synchronization of overlaps, in the quantitative form proved by Mourrat. Panchenko proved that, under the Ghirlanda–Guerra identities, the overlap of each species in a multi-species model is a nondecreasing Lipschitz function of the total overlap [ref-100]. The proof rests on his theorem that the Ghirlanda–Guerra identities imply ultrametricity [ref-97]. He proved the analogue for vector spins in [ref-101]. Mourrat’s synchronization theorem [ref-92] is a variant for two overlap arrays under approximate identities. The following form of it bounds the conditional variance of one overlap given the other, once approximate Ghirlanda–Guerra identities hold for a sufficiently rich finite family of test functions. Chapter 4 uses the same synchronization in the lower bound.

Lemma 3.8. Fix an enumeration (ϑn)(\vartheta_n) of Q∩[0,1]\mathbb{Q}\cap[0,1]. For every ε>0\varepsilon>0 there is an integer nε≥1n_\varepsilon\ge1 with the following property. Let a random probability measure be supported on the product of the unit balls of two Hilbert spaces, and let All′A^{ll'} and Bll′B^{ll'} be the two inner-product arrays of independent samples from it. Suppose that

∣Dn(f,cj)∣≤1nεfor all n≤nε and j∈{1,…,nε}3,\left|\mathcal{D}_n(f,c_j)\right|\le\frac{1}{n_\varepsilon} \qquad\text{for all }n\le n_\varepsilon\text{ and }j\in\{1,\ldots,n_\varepsilon\}^3,

and every continuous test function ff of the arrays of nn samples with ∣f∣≤1\lvert f\rvert\le1, where cj(u,v)=(ϑj1u+ϑj2v)j3c_j(u,v)=(\vartheta_{j_1}u+\vartheta_{j_2}v)^{j_3}. If the averaged law of B12B^{12} is K−1∑l=1KδqlK^{-1}\sum_{l=1}^{K}\delta_{q_l}, where −1=q0<q1<⋯<qK≤1-1=q_0<q_1<\cdots<q_K\le1, then

E⟨(A12−E[A12∣B12])2⟩≤12K+εK2max⁡0≤l<K(ql+1−ql)−1.(76)\mathbb{E}\left\langle\left(A^{12}-\mathbb{E}\left[A^{12}\mid B^{12}\right]\right)^2\right\rangle \le\frac{12}{K}+\varepsilon K^2\max_{0\le l<K}(q_{l+1}-q_l)^{-1}. \tag*{(76)}

Proof. This is [ref-92], stated there with a number δ>0\delta>0 in place of 1/nε1/n_\varepsilon and with the conditions required for n,j1,j2,j3≤⌊δ−1⌋n,j_1,j_2,j_3\le\lfloor\delta^{-1}\rfloor; its hypothesis on the test functions and on the law of B12B^{12} is the one above, and E⟨A12∣B12⟩\mathbb{E}\langle A^{12}\mid B^{12}\rangle is the conditional expectation under the averaged measure. Our form follows by taking nε=⌈δ−1⌉n_\varepsilon=\lceil\delta^{-1}\rceil, because then 1/nε≤δ1/n_\varepsilon\le\delta and ⌊δ−1⌋≤nε\lfloor\delta^{-1}\rfloor\le n_\varepsilon. □

We apply Lemma 3.8 with the following choices. Recall that K=k+1K=k+1 and wl=1/Kw_l=1/K for 0≤l≤k0\le l\le k. The maps σ↦σ/N\sigma\mapsto\sigma/\sqrt{N} and β↦k−1/2∑l=1kuβ∣l\beta\mapsto k^{-1/2}\sum_{l=1}^{k}\mathbf{u}_{\beta|l}, where (uν)(\mathbf{u}_\nu) is an orthonormal family indexed by the nodes of depth at least one, represent RR and J/kJ/k as inner products of unit vectors, since k−1#{1≤l≤k:β∣l=β′∣l}=(β∧β′)/kk^{-1}\#\{1\le l\le k:\beta|l=\beta'|l\}=(\beta\wedge\beta')/k. Up to the factor k−1/2k^{-1/2}, this is the embedding of [ref-92]. The image of the Gibbs measure under the product of these maps is a random probability measure on the product of two unit balls. By Lemma 3.2, J/kJ/k has averaged law exactly uniform on {0,1/k,…,1}\{0,1/k,\ldots,1\}, for every value of the parameters, including the perturbation. In the notation of Lemma 3.8 this is ql=(l−1)/kq_l=(l-1)/k for 1≤l≤K1\le l\le K, so the gaps ql+1−qlq_{l+1}-q_l are 11 (for l=0l=0) and 1/k1/k, and the right side of (76) is at most 12/K+εK312/K+\varepsilon K^3.

We choose the parameters in the following order. Given KK, let ε=K−4\varepsilon=K^{-4}, let nK=nεn_K=n_\varepsilon be given by Lemma 3.8, and take j+=nKj_+=n_K, so that the functions cjc_j, j∈Ij\in\mathcal{I}, are exactly those of Lemma 3.8. At any point satisfying the hypotheses of Lemma 3.7, and for NN so large that CN−1/8≤1/nKCN^{-1/8}\le1/n_K, the hypotheses of Lemma 3.8 hold. The left side of (76) equals ∑lwlVar⁡(R12∣J=l)\sum_l w_l\operatorname{Var}(R_{12}\mid J=l), since E⟨R12∣J⟩=rJ\mathbb{E}\langle R_{12}\mid J\rangle=r_J, and 12/K+K−4K3=13/K12/K+K^{-4}K^3=13/K. Hence, with

EN=12∑l=0kwlCov⁡(R12,αNX12∣J=l)andηN=αNL2213K,(77)\mathcal{E}_N=\frac{1}{2}\sum_{l=0}^{k}w_l\operatorname{Cov}\left(R_{12},\alpha_NX_{12}\mid J=l\right) \qquad\text{and}\qquad \eta_N=\frac{\alpha_NL^2}{2}\sqrt{\frac{13}{K}}, \tag*{(77)}

we have ∑lwlVar⁡(R12∣J=l)≤13/K\sum_l w_l\operatorname{Var}(R_{12}\mid J=l)\le13/K and ∣EN∣≤ηN|\mathcal{E}_N|\le\eta_N at such a point. The second bound follows from the first: since ∣X12∣≤L2|X_{12}|\le L^2, the conditional Cauchy–Schwarz inequality gives ∣Cov⁡(R12,αNX12∣J=l)∣≤αNL2Var⁡(R12∣J=l)1/2|\operatorname{Cov}(R_{12},\alpha_NX_{12}\mid J=l)|\le\alpha_NL^2\operatorname{Var}(R_{12}\mid J=l)^{1/2}, and the Cauchy–Schwarz inequality over ll, with ∑lwl=1\sum_l w_l=1, gives ∑lwlVar⁡(R12∣J=l)1/2≤(13/K)1/2\sum_l w_l\operatorname{Var}(R_{12}\mid J=l)^{1/2}\le(13/K)^{1/2}. By (66),

∂tuN=−12∑l=0kwl(rl−hl)ζl−EN.(78)\partial_tu_N=-\frac{1}{2}\sum_{l=0}^{k}w_l(r_l-h_l)\zeta_l-\mathcal{E}_N. \tag*{(78)}

The tolerance ε=K−4\varepsilon=K^{-4} and the bound 13/K13/K are those of Mourrat [ref-92]. There the conditional variances of both overlaps enter. Here only the spin overlap is synchronized with JJ, and the response overlap enters through the bound ∣X12∣≤L2|X_{12}|\le L^2, which gives the square root in ηN\eta_N.

The spherical term

At t=0t=0 the spins interact only with the cascade field YY, through the term ψσ,N(p)\psi_{\sigma,N}(p) of (70). This is the only geometric term in the comparison. It enters through the maximum over pp of ψσ,N(p)+12∑lwlhlpl\psi_{\sigma,N}(p)+\frac{1}{2}\sum_l w_lh_lp_l, and we show that this maximum is at most CS(γh)+oN(1)\mathrm{CS}(\gamma_h)+o_N(1). The idea is to replace the uniform measure on the sphere by a Gaussian measure with a well-chosen variance 1/b1/b, for which the cascade recursion is computed exactly by Lemma 3.4. This replacement, at the cost (b−1−log⁡b)/2(b-1-\log b)/2, is the large deviation argument behind the Parisi formula for spherical models [ref-125]; see also [ref-37, ref-93]. Here only an upper bound is needed, and it holds for each NN with an error of order log⁡N/N\log N/N.

Lemma 3.9. The following hold.

(i) For every NN and every p∈Kp \in K, ψσ,N(p)≤pk−12∑lwlpl\psi_{\sigma,N}(p) \le\sqrt{p_k}-\frac{1}{2}\sum_l w_l p_l.

(ii) For h∈Kh \in K with hk<1h_k<1,

ψˉN≔sup⁡p∈K{ψσ,N(p)+12∑l=0kwlhlpl}≤CS(γh)+oN(1),(79)\bar{\psi}_N \coloneqq\sup_{p\in K}\left\{\psi_{\sigma,N}(p)+\frac{1}{2}\sum_{l=0}^{k}w_lh_lp_l\right\}\le\mathrm{CS}(\gamma_h)+o_N(1), \tag*{(79)}

where oN(1)o_N(1) depends only on KK and hh.

Proof. Part (ii) rests on part (i).

Proof of (i). The leaf logarithm log⁡∫eYβ⋅σ dμN\log\int e^{Y_\beta\cdot\sigma}\,\mathrm{d}\mu_N is NΔpl\sqrt{N\Delta p_l}-Lipschitz in the Gaussian vector yβ∣ly_{\beta|l}, since its gradient is Δpl⟨σ⟩β\sqrt{\Delta p_l}\langle\sigma\rangle_\beta. Differentiating (47) shows that each Xl−1X_{l-1} keeps these Lipschitz bounds in the earlier vectors, since its gradient is a KlK_l-average. Gaussian concentration in the form Eeθ(F−EF)≤eθ2ℓ2/2\mathbb{E}e^{\theta(F-\mathbb{E}F)}\le e^{\theta^2\ell^2/2} for an ℓ\ell-Lipschitz function FF of a standard Gaussian vector [ref-21], Theorem 5.5, applied at each level l=k,…,1l=k,\ldots,1, shows that Xl−1X_{l-1} exceeds the plain average of XlX_l over yβ∣ly_{\beta|l} by at most NθlΔpl/2N\theta_l\Delta p_l/2. The root is averaged without a logarithm. Therefore, by (50),

1NElog⁡∑βvβ∫eYβ⋅σ dμN≤1NElog⁡∫eY⋅σ dμN+12∑l=1kθlΔpl≤pk+12∑l=1kθlΔpl,\frac{1}{N}\mathbb{E}\log\sum_{\beta}v_\beta\int e^{Y_\beta\cdot\sigma}\,\mathrm{d}\mu_N \le\frac{1}{N}\mathbb{E}\log\int e^{Y\cdot\sigma}\,\mathrm{d}\mu_N+\frac{1}{2}\sum_{l=1}^{k}\theta_l\Delta p_l \le\sqrt{p_k}+\frac{1}{2}\sum_{l=1}^{k}\theta_l\Delta p_l,

where Y∼N(0,pkIN)Y\sim N(0,p_kI_N), and we used Y⋅σ≤N∣Y∣Y\cdot\sigma\le\sqrt{N}|Y| and E∣Y∣≤Npk\mathbb{E}|Y|\le\sqrt{Np_k}. By (52) this proves (i).

Proof of (ii). We divide the proof into three steps.

Step 1. Localization. By (i), since hl≤hkh_l\le h_k and pl≥0p_l\ge0,

ψσ,N(p)+12∑lwlhlpl≤pk−12∑lwl(1−hl)pl≤pk−12wk(1−hk)pk.\psi_{\sigma,N}(p)+\frac{1}{2}\sum_lw_lh_lp_l \le\sqrt{p_k}-\frac{1}{2}\sum_lw_l(1-h_l)p_l \le\sqrt{p_k}-\frac{1}{2}w_k(1-h_k)p_k.

The value at p=0p=0 is zero, and the right side is negative when pk>pˉ≔4/[wk(1−hk)]2p_k>\bar p\coloneqq4/[w_k(1-h_k)]^2. So the supremum may be restricted to the set {p∈K:pk≤pˉ}\{p\in K:p_k\le\bar p\}, and we prove the bound uniformly over this set. Fix such a pp.

Step 2. Choice of the variance. In the notation of (53) with n=Nn=N, as bb increases every dld_l increases, so mk+1/bm_k+1/b decreases continuously. It tends to zero as b→∞b\to\infty, and to infinity as b↓∑j≥1θjΔpjb\downarrow\sum_{j\ge1}\theta_j\Delta p_j: if p0>0p_0>0, the term p0/d02p_0/d_0^2 diverges; if p0=0p_0=0 and p≠0p\ne0, then for the first index j∗j_* with Δpj∗>0\Delta p_{j_*}>0 we have dj∗−1=b−∑j≥1θjΔpjd_{j_*-1}=b-\sum_{j\ge1}\theta_j\Delta p_j, and the j∗j_*-th term of mkm_k diverges; if p=0p=0, the term 1/b1/b diverges. We choose bb as the unique solution of mk+1/b=1m_k+1/b=1, so that Lemma 3.4(iv) with ϱ=1\varrho=1 applies. In particular b≥1b\geq1. Also b<2(pˉ+1)b<2(\bar p+1): since dl≥d0≥b−pk≥b−pˉd_l\geq d_0\geq b-p_k\geq b-\bar p, we have mk≤pk/(b−pˉ)2≤pˉ/(b−pˉ)2m_k\leq p_k/(b-\bar p)^2\leq\bar p/(b-\bar p)^2 for b>pˉb>\bar p, and at b=2(pˉ+1)b=2(\bar p+1) this plus 1/b1/b is at most 18+12<1\frac{1}{8}+\frac{1}{2}<1, because (pˉ+2)2≥8pˉ(\bar p+2)^2\geq8\bar p.

Step 3. Comparison with the sphere. For fixed YY, the average of eρY⋅σe^{\rho Y\cdot\sigma} over σ∈SN\sigma\in S_N is nondecreasing in ρ≥0\rho\geq0: by the symmetry σ↦−σ\sigma\mapsto-\sigma it equals the average of cosh⁡(ρY⋅σ)\cosh(\rho Y\cdot\sigma), whose derivative in ρ\rho is nonnegative. If ε∼N(0,b−1IN)\varepsilon\sim N(0,b^{-1}I_N), then ε=∣ε∣N−1/2σ\varepsilon=|\varepsilon|N^{-1/2}\sigma with σ\sigma uniform on SNS_N and independent of ∣ε∣|\varepsilon|, and b∣ε∣2b|\varepsilon|^2 has the chi-square law χN2\chi_N^2 with NN degrees of freedom. Restricting EeY⋅ε=e∣Y∣2/(2b)\mathbb{E}e^{Y\cdot\varepsilon}=e^{|Y|^2/(2b)} to {∣ε∣2≥N}\{|\varepsilon|^2\geq N\} and using the monotonicity gives, pathwise,

log⁡∑βvβ∫eYβ⋅σ dμN≤log⁡∑βvβe∣Yβ∣2/(2b)−log⁡P{χN2≥bN}.\log\sum_{\beta}v_{\beta}\int e^{Y_{\beta}\cdot\sigma}\,\mathrm{d}\mu_N \leq \log\sum_{\beta}v_{\beta}e^{|Y_{\beta}|^2/(2b)} -\log\mathbb{P}\{\chi_N^2\geq bN\}.

For N≥2N\geq2 and b≥1b\geq1,

−1Nlog⁡P{χN2≥bN}≤b−1−log⁡b2+12log⁡(πN)+log⁡b+1+16NN.(80)-\frac{1}{N}\log\mathbb{P}\{\chi_N^2\geq bN\} \leq \frac{b-1-\log b}{2} + \frac{\frac{1}{2}\log(\pi N)+\log b+1+\frac{1}{6N}}{N}. \tag*{(80)}

Indeed, the chi-square density xN/2−1e−x/2/(2N/2Γ(N/2))x^{N/2-1}e^{-x/2}/(2^{N/2}\Gamma(N/2)) is nonincreasing on [N−2,∞)[N-2,\infty), so P{χN2≥bN}\mathbb{P}\{\chi_N^2\geq bN\} is at least twice its value at bN+2bN+2. Then (80) follows from (bN+2)N/2−1≥(bN)N/2−1(bN+2)^{N/2-1}\geq(bN)^{N/2-1} and Stirling’s bound Γ(z)≤2πzz−1/2e−z+1/(12z)\Gamma(z)\leq\sqrt{2\pi}z^{z-1/2}e^{-z+1/(12z)} with z=N/2z=N/2, after collecting terms. Take expectations, divide by NN, and use (54) (with n=Nn=N), (61) and (62) in the form b−12−pk2=−12∑lwlmlpl\frac{b-1}{2}-\frac{p_k}{2}=-\frac{1}{2}\sum_l w_lm_lp_l. Since 1≤b<2(pˉ+1)1\leq b<2(\bar p+1), this gives

ψσ,N(p)+12∑lwlhlpl≤CS⁡(γm)+12∑lwlpl(hl−ml)+Cpˉlog⁡NN.\psi_{\sigma,N}(p)+\frac{1}{2}\sum_l w_lh_lp_l \leq \operatorname{CS}(\gamma_m) +\frac{1}{2}\sum_l w_lp_l(h_l-m_l) +C_{\bar p}\frac{\log N}{N}.

It remains to show that CS⁡(γh)≥CS⁡(γm)+12∑lwlpl(hl−ml)\operatorname{CS}(\gamma_h)\geq\operatorname{CS}(\gamma_m)+\frac{1}{2}\sum_l w_lp_l(h_l-m_l). The map x↦CS⁡(γx)x\mapsto\operatorname{CS}(\gamma_x) is convex on the convex set {x∈K:xk<1}\{x\in K:x_k<1\}. Indeed, compute the entropies at two points x,x′x,x' with a common cutoff q′∈[max⁡(xk,xk′),1)q'\in[\max(x_k,x'_k),1): for each s<q′s<q', λγx(s)=1−∑lwlmax⁡(xl,s)\lambda_{\gamma_x}(s)=1-\sum_lw_l\max(x_l,s) is positive and concave in xx, so its reciprocal is convex. Hence f(ϵ)=CS⁡(γm+ϵ(h−m))f(\epsilon)=\operatorname{CS}(\gamma_{m+\epsilon(h-m)}) is convex on [0,1][0,1], and CS⁡(γh)≥CS⁡(γm)+f′(0+)\operatorname{CS}(\gamma_h)\geq\operatorname{CS}(\gamma_m)+f'(0+); this one-sided derivative is all we need, also when mm lies on the boundary of KK. With the common cutoff q′=max⁡(hk,mk)q'=\max(h_k,m_k), all the functions λ\lambda involved are at least 1−q′>01-q'>0 on [0,q′][0,q'], and ∣max⁡(ml+ϵ(hl−ml),s)−max⁡(ml,s)∣≤ϵ∣hl−ml∣|\max(m_l+\epsilon(h_l-m_l),s)-\max(m_l,s)|\leq\epsilon|h_l-m_l|. For almost every ss, the right derivative of max⁡(ml+ϵ(hl−ml),s)\max(m_l+\epsilon(h_l-m_l),s) at ϵ=0\epsilon=0 is (hl−ml)1{ml>s}(h_l-m_l)\mathbf{1}_{\{m_l>s\}}. Dominated convergence and (61) give

f′(0+)=∑lwl2(hl−ml)∫0mldsλγm(s)2=12∑lwlpl(hl−ml).f'(0+) = \sum_l\frac{w_l}{2}(h_l-m_l)\int_0^{m_l}\frac{\mathrm{d}s}{\lambda_{\gamma_m}(s)^2} = \frac{1}{2}\sum_lw_lp_l(h_l-m_l).

This bounds the left side of the previous display by CS⁡(γh)+Cpˉlog⁡N/N\operatorname{CS}(\gamma_h)+C_{\bar p}\log N/N, uniformly over p∈Kp\in K with pk≤pˉp_k\leq\bar p, and proves (ii). □\square

The comparison

We now show that, up to the synchronization error, the free energy at t=1t=1 is at most the maximum over pp of the interpolation at t=0t=0. The reason is the following. At a point where pp maximizes uN(t,⋅,x)+12∑lwlhlplu_N(t,\cdot,x)+\frac{1}{2}\sum_l w_lh_lp_l, the product term in (78) has a sign, so ∂tuN\partial_tu_N is at most the synchronization error. A maximum principle in tt turns this into a bound on the growth of the maximum over pp. Maximizing also over xx, with a quadratic penalty, selects a perturbation at which Lemma 3.7 applies.

This is a finite-NN form of Step 5 of the proof of [ref-92] (Theorem 4.1), where the first-order conditions [ref-92] ((4.39)–(4.40)) at a contact point and the monotonicity of [ref-92] (Lemma 2.4) give the sign [ref-92] ((4.32)) needed for the supersolution property. Here the maximum principle in tt for the maximum over pp takes the place of viscosity solutions and of the comparison principle [ref-92] (Proposition 3.2). The use of first-order conditions in the levels to control the derivative in tt also parallels the stationarization of Stojnic’s fully lifted interpolation for bilinearly indexed random processes [ref-120, ref-121]. There the levels on both sides are chosen at stationary points along the interpolation path, and the derivative in tt vanishes when the overlaps concentrate.

Proposition 3.10. Let UU belong to class (A), let K≥2K\geq2 and h∈Kh\in\mathcal K with hk<1h_k<1, let nKn_K and j+=nKj_+=n_K be as in §3.5, and let ηN\eta_N and ψˉN\bar\psi_N be given by (77) and (79). There is N0N_0, depending only on UU, KK, hh, and α′\alpha', such that for N≥N0N\geq N_0

FN(U)≤ψˉN+αNuγh(0,0;U)+ηN+3Cj+∣I∣N−1/8.(81)F_N(U)\leq\bar\psi_N+\alpha_Nu_{\gamma_h}(0,0;U)+\eta_N+3C_{j_+}|\mathcal I|N^{-1/8}. \tag*{(81)}

Consequently,

lim sup⁡N→∞FN(U)≤PU(γh)+αL2213K.(82)\limsup_{N\to\infty}F_N(U)\leq\mathcal P_U(\gamma_h)+\frac{\alpha L^2}{2}\sqrt{\frac{13}{K}}. \tag*{(82)}

Proof. Fix t0∈(0,1)t_0\in(0,1) and ς∈(0,1]\varsigma\in(0,1], and consider on [0,t0]×K×RI[0,t_0]\times\mathcal K\times\mathbb R^{\mathcal I} the function

Ψ(t,p,x)=uN(t,p,x)+12∑lwlhlpl−∣x−1∣2−(ηN+ς)t.\Psi(t,p,x)=u_N(t,p,x)+\frac{1}{2}\sum_l w_lh_lp_l-|x-\mathbf1|^2-(\eta_N+\varsigma)t.

We divide the proof into four steps.

Step 1. Localization. We show that Ψ\Psi attains its maximum and that all maximum points lie in a compact set that does not depend on NN, t0t_0, or ς\varsigma. We take N≥N0N\geq N_0, with N0N_0 fixed at the end of this step; in particular

(a)2Cj+N−1/8≤min⁡{12,∣I∣−1},\text{(a)}\quad2C_{j_+}N^{-1/8}\leq\min\left\{\frac{1}{2},|\mathcal I|^{-1}\right\},

where Cj+C_{j_+} is the constant of (73). Then Cj+N−1/8∣x∣2≤2Cj+N−1/8(∣x−1∣2+∣I∣)≤12∣x−1∣2+1C_{j_+}N^{-1/8}|x|^2\leq2C_{j_+}N^{-1/8}(|x-\mathbf1|^2+|\mathcal I|)\leq\frac{1}{2}|x-\mathbf1|^2+1. Since U≤Umax⁡U\leq U_{\max}, the leaf integral is at most eMUmax⁡∫eYβ⋅σdμNe^{MU_{\max}}\int e^{Y_\beta\cdot\sigma}\mathrm d\mu_N when x=0x=0, so uN(t,p,0)≤αNUmax⁡+ψσ,N(p)u_N(t,p,0)\leq\alpha_NU_{\max}+\psi_{\sigma,N}(p). With (73), Lemma 3.9(i) and hl≤hkh_l\leq h_k,

Ψ(t,p,x)≤α′Umax⁡++1+pk−12wk(1−hk)pk−12∣x−1∣2.\Psi(t,p,x)\leq\alpha'U_{\max}^{+}+1+\sqrt{p_k}-\frac{1}{2}w_k(1-h_k)p_k-\frac{1}{2}|x-\mathbf1|^2.

On the other hand, by (73), (70), ψσ,N(0)=0\psi_{\sigma,N}(0)=0, and the bound uγh(0,0;U)≥EU(G)≥U(0)−Lu_{\gamma h}(0,0;U)\geq\mathbb{E}U(G)\geq U(0)-L from (2.6),

Ψ(0,0,1)≥uN(0,0,0)−Cj+N−1/8∣I∣≥−α′(∣U(0)∣+L)−1.\Psi(0,0,1)\geq u_N(0,0,0)-C_{j_+}N^{-1/8}|\mathcal{I}|\geq-\alpha'(|U(0)|+L)-1.

Since wk(1−hk)>0w_k(1-h_k)>0 and 0≤pl≤pk0\leq p_l\leq p_k on KK, we have Ψ<Ψ(0,0,1)\Psi<\Psi(0,0,1) outside [0,t0]×K0[0,t_0]\times K_0, for a compact set K0⊂K×RIK_0\subset K\times\mathbb{R}^{\mathcal{I}} that depends only on U,K,h,IU,K,h,\mathcal{I}, and α′\alpha'. As Ψ\Psi is continuous by Lemma 3.5(i), it attains its maximum, and every maximum point lies in [0,t0]×K0[0,t_0]\times K_0. Let N0N_0 be the least integer such that, for all N≥N0N\geq N_0, condition (a) holds and also

(b) 2j+N−1/8∣x∣≤122^{j_+}N^{-1/8}|x|\leq\frac{1}{2} for all (p,x)∈K0(p,x)\in K_0, and

(c) CN−1/8≤1/nKCN^{-1/8}\leq1/n_K, where CC is the constant of Lemma 3.7 for the compact set [0,1]×{p:(p,x)∈K0 for some x}[0,1]\times\{p:(p,x)\in K_0\text{ for some }x\}.

Since K0,IK_0,\mathcal{I}, and nKn_K are determined by U,K,hU,K,h, and α′\alpha', so is N0N_0; it does not depend on t0t_0 or ς\varsigma.

Step 2. Conditions at a maximum point. Let (t∗,p∗,x∗)(t_*,p_*,x_*) be a maximum point. First, x∗x_* maximizes x↦uN(t∗,p∗,x)−∣x−1∣2x\mapsto u_N(t_*,p_*,x)-|x-1|^2, so ∇xuN=2(x∗−1)\nabla_xu_N=2(x_*-1) there. The derivative bound of Lemma 3.6 gives ∣x∗−1∣≤2j+N−1/8∣x∗∣|x_*-1|\leq2^{j_+}N^{-1/8}|x_*|, so ∣x∗−1∣≤12|x_*-1|\leq\frac{1}{2} by (b). Maximality in xx gives uN(x∗+y)−uN(x∗)≤∣x∗+y−1∣2−∣x∗−1∣2=∇xuN(x∗)⋅y+∣y∣2u_N(x_*+y)-u_N(x_*)\leq|x_*+y-1|^2-|x_*-1|^2=\nabla_xu_N(x_*)\cdot y+|y|^2 for all y∈RIy\in\mathbb{R}^{\mathcal{I}}, and convexity of uNu_N in xx gives the lower bound in (75). Thus Lemma 3.7 applies, and by (c) and §3.5, ∣EN∣≤ηN|\mathcal{E}_N|\leq\eta_N at (t∗,p∗,x∗)(t_*,p_*,x_*).

Second, p∗p_* maximizes p↦uN(t∗,p,x∗)+12∑lwlhlplp\mapsto u_N(t_*,p,x_*)+\frac{1}{2}\sum_l w_lh_lp_l over KK. Since p∗+ϵ1≥m∈Kp_*+\epsilon\mathbf{1}_{\geq m}\in K for ϵ≥0\epsilon\geq0, the right derivative in this direction is nonpositive, and by (67)

Am≔12∑l=mkwl(hl−rl)≤0(0≤m≤k).(83)A_m\coloneqq\frac{1}{2}\sum_{l=m}^{k}w_l(h_l-r_l)\leq0\qquad(0\leq m\leq k). \tag*{(83)}

Third, t∗t_* maximizes t↦Ψ(t,p∗,x∗)t\mapsto\Psi(t,p_*,x_*) over [0,t0][0,t_0].

Step 3. Every maximum point has t∗=0t_*=0. Suppose that t∗>0t_*>0. Since t∗≤t0<1t_*\leq t_0<1, uNu_N is differentiable in tt at (t∗,p∗,x∗)(t_*,p_*,x_*) by Lemma 3.5(ii), and maximality in tt gives ∂tuN(t∗,p∗,x∗)≥ηN+ς\partial_tu_N(t_*,p_*,x_*)\geq\eta_N+\varsigma. We now show the opposite inequality. Put ζ−1=0\zeta_{-1}=0, so that ζl=∑m≤l(ζm−ζm−1)\zeta_l=\sum_{m\leq l}(\zeta_m-\zeta_{m-1}). Summation by parts, (68) and (83) give

∑lwl(rl−hl)ζl=∑m=0k(ζm−ζm−1)∑l=mkwl(rl−hl)=−2∑m=0k(ζm−ζm−1)Am≥0.\sum_l w_l(r_l-h_l)\zeta_l = \sum_{m=0}^{k}(\zeta_m-\zeta_{m-1})\sum_{l=m}^{k}w_l(r_l-h_l) = -2\sum_{m=0}^{k}(\zeta_m-\zeta_{m-1})A_m \geq0.

By (78) and Step 2, ∂tuN=−12∑lwl(rl−hl)ζl−EN≤∣EN∣≤ηN\partial_tu_N=-\frac{1}{2}\sum_lw_l(r_l-h_l)\zeta_l-\mathcal{E}_N\leq|\mathcal{E}_N|\leq\eta_N, a contradiction. Hence t∗=0t_*=0.

Step 4. Conclusion. Compare Ψ\Psi at (t0,0,1)(t_{0},0,1) with its value at a maximum point, which has t∗=0t_{*}=0. By (73) and (70),

uN(t0,0,1)−(ηN+ς)t0≤Ψ(0,p∗,x∗)≤ψσ,N(p∗)+12∑lwlhlp∗,l+αNuγh(0,0;U)+Cj+N−1/8∣x∗∣2−∣x∗−1∣2.\begin{aligned} u_{N}(t_{0},0,1)-(\eta_{N}+\varsigma)t_{0} &\leq\Psi(0,p_{*},x_{*}) \\ &\leq\psi_{\sigma,N}(p_{*})+\frac{1}{2}\sum_{l}w_{l}h_{l}p_{*,l} +\alpha_{N}u_{\gamma_{h}}(0,0;U) \\ &\qquad+C_{j_{+}}N^{-1/8}|x_{*}|^{2}-|x_{*}-1|^{2}. \end{aligned}

By (a), the last two terms are at most 2Cj+N−1/8∣I∣2C_{j_{+}}N^{-1/8}|\mathcal{I}|, and the first two are at most ψˉN\bar{\psi}_{N}. Letting t0↑1t_{0}\uparrow1, by continuity of uNu_{N}, and then ς↓0\varsigma\downarrow0,

uN(1,0,1)≤ψˉN+αNuγh(0,0;U)+ηN+2Cj+N−1/8∣I∣.u_{N}(1,0,1)\leq\bar{\psi}_{N}+\alpha_{N}u_{\gamma_{h}}(0,0;U)+\eta_{N}+2C_{j_{+}}N^{-1/8}|\mathcal{I}|.

Finally, FN(U)=uN(1,0,0)≤uN(1,0,1)+Cj+N−1/8∣I∣F_{N}(U)=u_{N}(1,0,0)\leq u_{N}(1,0,1)+C_{j_{+}}N^{-1/8}|\mathcal{I}| by (70) and (73). This proves (81). By Lemma 3.9(ii), ψˉN≤CS(γh)+oN(1)\bar{\psi}_{N}\leq\mathrm{CS}(\gamma_{h})+o_{N}(1), and αN→α\alpha_{N}\to\alpha, so (81) gives (82). □\square

Proof of Theorem 3.1. Let γ∈U\gamma\in\mathcal{U} and K≥2K\geq2. Approximate γ\gamma by the order parameter γK=⌊Kγ⌋/K\gamma_{K}=\lfloor K\gamma\rfloor/K, which has KK atoms of mass 1/K1/K at the quantiles hl=inf⁡{q:γ(q)≥(l+1)/K}h_{l}=\inf\{q:\gamma(q)\geq(l+1)/K\}, 0≤l≤k=K−10\leq l\leq k=K-1. Indeed, by right-continuity hl≤qh_{l}\leq q if and only if γ(q)≥(l+1)/K\gamma(q)\geq(l+1)/K, so γh(q)=K−1#{l:hl≤q}=⌊Kγ(q)⌋/K\gamma_{h}(q)=K^{-1}\#\{l:h_{l}\leq q\}=\lfloor K\gamma(q)\rfloor/K. Thus h∈Kh\in\mathcal{K}, hk=Qγ<1h_{k}=Q_{\gamma}<1, γ−1/K<γK≤γ\gamma-1/K<\gamma_{K}\leq\gamma, so that ∥γK−γ∥L1≤1/K\|\gamma_{K}-\gamma\|_{L^{1}}\leq1/K, and γK=γ=1\gamma_{K}=\gamma=1 on [Qγ,1][Q_{\gamma},1]. By (82) for this hh, and by (33) with qˉ=Qγ\bar{q}=Q_{\gamma}, since γ,γK∈UQγ\gamma,\gamma_{K}\in\mathcal{U}_{Q_{\gamma}},

lim sup⁡N→∞FN(U)≤PU(γ)+1K(αL22+Qγ2(1−Qγ)2)+αL2213K.\limsup_{N\to\infty}F_{N}(U) \leq\mathcal{P}_{U}(\gamma) +\frac{1}{K}\left(\frac{\alpha L^{2}}{2}+\frac{Q_{\gamma}}{2(1-Q_{\gamma})^{2}}\right) +\frac{\alpha L^{2}}{2}\sqrt{\frac{13}{K}}.

Letting K→∞K\to\infty proves (46).

If L=0L=0, the activation is constant and fN(U)=FN(U)f_{N}(U)=F_{N}(U), so the probability bound follows from (46). Let L>0L>0. The map G↦fN(U)G\mapsto f_{N}(U) is Lipschitz: its derivative in gaig_{ai} is N−3/2⟨U′(ba⋅σ)σi⟩N^{-3/2}\langle U'(b_{a}\cdot\sigma)\sigma_{i}\rangle, where now ⟨⋅⟩\langle\cdot\rangle is the Gibbs measure of ZN(U)Z_{N}(U), and by Jensen’s inequality and ∣σ∣2=N|\sigma|^{2}=N the sum of the squares is at most αNL2/N\alpha_{N}L^{2}/N. Gaussian concentration [ref-21] (Theorem 5.6) gives P{fN(U)−FN(U)>ε/2}≤exp⁡(−ε2N/(8αNL2))\mathbb{P}\{f_{N}(U)-F_{N}(U)>\varepsilon/2\}\leq\exp(-\varepsilon^{2}N/(8\alpha_{N}L^{2})), and FN(U)≤PU(γ)+ε/2F_{N}(U)\leq\mathcal{P}_{U}(\gamma)+\varepsilon/2 for large NN by (46).

The parameters and limits are chosen, with UU and γ\gamma fixed, in the following order: first KK (and with it hh), then the synchronization tolerance ε=K−4\varepsilon=K^{-4}, the integer nKn_{K}, and the perturbation family I\mathcal{I}; then, for each N≥N0N\geq N_{0}, the auxiliary t0↑1t_{0}\uparrow1 and ς↓0\varsigma\downarrow0 inside the proof of Proposition 3.10; then N→∞N\to\infty; and finally K→∞K\to\infty. □\square

Hard constraints

We now pass to the hard wall. Recall that VN(κ)=ZN(Hκ)V_{N}(\kappa)=Z_{N}(H_{\kappa}) with Hκ=log⁡1[κ,∞)H_{\kappa}=\log\mathbf{1}_{[\kappa,\infty)} and log⁡0=−∞\log0=-\infty. Chapter 9 also needs the feasible volume weighted by a bounded function of the fields, so we allow a source term from the start.

Terminal data with a hard wall. Let γ\gamma be a step order parameter as in (18), and let φ:R→R\varphi:\mathbb{R}\to\mathbb{R} be bounded and measurable. The terminal datum Hκ+φH_{\kappa}+\varphi is bounded above by ∥φ∥∞\lVert\varphi\rVert_{\infty}, so the recursion (17) with T=1T=1 defines uγ(⋅,⋅;Hκ+φ)u_{\gamma}(\cdot,\cdot;H_{\kappa}+\varphi), independently of the representation of γ\gamma, and we put

PHκ+φ(γ)=αuγ(0,0;Hκ+φ)+CS(γ).\mathcal{P}_{H_{\kappa}+\varphi}(\gamma)=\alpha u_{\gamma}(0,0;H_{\kappa}+\varphi)+\mathrm{CS}(\gamma).

The function uγ(q,x;Hκ)u_{\gamma}(q,x;H_{\kappa}) is finite for q<1q<1: for q≤tnq\leq t_{n}, the last breakpoint of γ\gamma, it equals uγ(q,x;κ)u_{\gamma}(q,x;\kappa) by (25), and for tn≤q<1t_{n}\leq q<1 it equals log⁡P{x+1−q G≥κ}\log\mathbb{P}\{x+\sqrt{1-q}\,G\geq\kappa\} by (22). By (19), applied in both directions, for bounded measurable φ,φ′\varphi,\varphi' the functions uγ(⋅,⋅;Hκ+φ)u_{\gamma}(\cdot,\cdot;H_{\kappa}+\varphi) are finite on [0,1)×R[0,1)\times\mathbb{R}, and

∣uγ(q,x;Hκ+φ)−uγ(q,x;Hκ+φ′)∣≤sup⁡y∈R∣φ(y)−φ′(y)∣,(q,x)∈[0,1)×R.(84)\left|u_{\gamma}(q,x;H_{\kappa}+\varphi)-u_{\gamma}(q,x;H_{\kappa}+\varphi')\right| \leq\sup_{y\in\mathbb{R}}\left|\varphi(y)-\varphi'(y)\right|, \qquad(q,x)\in[0,1)\times\mathbb{R}. \tag*{(84)}

By (25), PHκ(γ)=Pκ(γ)\mathcal{P}_{H_{\kappa}}(\gamma)=\mathcal{P}_{\kappa}(\gamma).

The wall approximation. We approximate the hard wall from above by activations in class (A) that decrease to it. Let

Ω(y)=y+−arctan⁡(y+),Uℓ(y;κ)=−ℓΩ(κ−y)(ℓ≥1).(85)\Omega(y)=y^{+}-\arctan(y^{+}),\qquad U_{\ell}(y;\kappa)=-\ell\Omega(\kappa-y)\qquad(\ell\geq1). \tag*{(85)}

The function Ω\Omega vanishes on (−∞,0](-\infty,0], is positive on (0,∞)(0,\infty), and is C2\mathrm{C}^{2} with Ω′(y)=(y+)2/(1+(y+)2)∈[0,1]\Omega'(y)=(y^{+})^{2}/(1+(y^{+})^{2})\in[0,1] and Ω′′(y)=2y+/(1+(y+)2)2∈[0,1]\Omega''(y)=2y^{+}/(1+(y^{+})^{2})^{2}\in[0,1]. Hence Uℓ(⋅;κ)U_{\ell}(\cdot;\kappa) belongs to class (A), with sup⁡Uℓ=0\sup U_{\ell}=0 and ∣Uℓ′∣,∣Uℓ′′∣≤ℓ|U'_{\ell}|,|U''_{\ell}|\leq\ell. Moreover, eUℓ=1e^{U_{\ell}}=1 on [κ,∞)[\kappa,\infty), and eUℓ(y)e^{U_{\ell}(y)} decreases to zero as ℓ→∞\ell\to\infty for y<κy<\kappa. Thus eUℓ↓1[κ,∞)e^{U_{\ell}}\downarrow\mathbf{1}_{[\kappa,\infty)} pointwise.

Corollary 3.11. Let κ∈R\kappa\in\mathbb{R} and ε>0\varepsilon>0.

(i) For every γ∈U\gamma\in\mathcal{U},

P{1Nlog⁡VN(κ)>Pκ(γ)+ε}⟶0.\mathbb{P}\left\{\frac{1}{N}\log V_{N}(\kappa)>\mathcal{P}_{\kappa}(\gamma)+\varepsilon\right\}\longrightarrow0.

(ii) Let γ\gamma be a step order parameter and let φ:R→R\varphi:\mathbb{R}\to\mathbb{R} be bounded and Lipschitz. Then

P{1Nlog⁡∫SN(κ)exp⁡{∑a=1Mφ(ba⋅σ)}μN(dσ)>PHκ+φ(γ)+ε}⟶0.\mathbb{P}\left\{ \frac{1}{N}\log\int_{S_{N}(\kappa)} \exp\left\{\sum_{a=1}^{M}\varphi(b_{a}\cdot\sigma)\right\} \mu_{N}(\mathrm{d}\sigma) > \mathcal{P}_{H_{\kappa}+\varphi}(\gamma)+\varepsilon \right\}\longrightarrow0.

Part (ii) with φ=0\varphi=0 is part (i) for step order parameters. Chapter 9 uses part (ii) with φ\varphi a real multiple of a bounded Lipschitz test function translated by κ\kappa, and Chapter 8 uses part (i).

Proof. We prove (ii) first, since (i) is deduced from it.

Proof of (ii). Let φ\varphi be LφL_{\varphi}-Lipschitz. The activation must be C2\mathrm{C}^{2}, so we first smooth φ\varphi. Let χ\chi be a smooth probability density supported in [−1,1][-1,1], and for τ∈(0,1]\tau\in(0,1] let φτ=φ∗χτ\varphi_{\tau}=\varphi\ast\chi_{\tau} with χτ(y)=τ−1χ(y/τ)\chi_{\tau}(y)=\tau^{-1}\chi(y/\tau). Then φτ\varphi_{\tau} is smooth, ∥φτ∥∞≤∥φ∥∞\lVert\varphi_{\tau}\rVert_{\infty}\leq\lVert\varphi\rVert_{\infty}, and ∣φτ−φ∣≤Lφτ\lvert\varphi_{\tau}-\varphi\rvert\leq L_{\varphi}\tau. Moreover φτ′=φ′∗χτ\varphi_{\tau}'=\varphi'*\chi_{\tau} and φτ′′=φ′∗χτ′\varphi_{\tau}''=\varphi'*\chi_{\tau}', where φ′\varphi' is the almost everywhere derivative, so ∣φτ′∣≤Lφ\lvert\varphi_{\tau}'\rvert\leq L_{\varphi} and ∣φτ′′∣≤Lφ∥χ′∥L1/τ\lvert\varphi_{\tau}''\rvert\leq L_{\varphi}\lVert\chi'\rVert_{L^{1}}/\tau. Hence U~ℓ=Uℓ(⋅;κ)+φτ\widetilde U_{\ell}=U_{\ell}(\mathord{\cdot};\kappa)+\varphi_{\tau}, with UℓU_{\ell} from (85), belongs to class (A) for every ℓ\ell.

Since Uℓ(⋅;κ)=0U_{\ell}(\mathord{\cdot};\kappa)=0 on [κ,∞)[\kappa,\infty) and the integrand of ZN(U~ℓ)Z_{N}(\widetilde U_{\ell}) is positive,

∫SN(κ)e∑aφ(ba⋅σ) μN(dσ)≤eMLφτ∫SN(κ)e∑aφτ(ba⋅σ) μN(dσ)≤eMLφτZN(U~ℓ).\int_{S_{N}(\kappa)} \mathrm{e}^{\sum_{a}\varphi(b_{a}\cdot\sigma)}\,\mu_{N}(\mathrm{d}\sigma) \leq \mathrm{e}^{M L_{\varphi}\tau} \int_{S_{N}(\kappa)} \mathrm{e}^{\sum_{a}\varphi_{\tau}(b_{a}\cdot\sigma)}\,\mu_{N}(\mathrm{d}\sigma) \leq \mathrm{e}^{M L_{\varphi}\tau}Z_{N}(\widetilde U_{\ell}).

So the normalized logarithm in (ii) is at most fN(U~ℓ)+α′Lφτf_{N}(\widetilde U_{\ell})+\alpha' L_{\varphi}\tau.

We show that uγ(0,0;U~ℓ)u_{\gamma}(0,0;\widetilde U_{\ell}) decreases to uγ(0,0;Hκ+φτ)u_{\gamma}(0,0;H_{\kappa}+\varphi_{\tau}) as ℓ→∞\ell\to\infty. The terminal data satisfy eU~ℓ↓1[κ,∞)eφτ=eHκ+φτ\mathrm{e}^{\widetilde U_{\ell}}\downarrow\mathbf{1}_{[\kappa,\infty)}\mathrm{e}^{\varphi_{\tau}}=\mathrm{e}^{H_{\kappa}+\varphi_{\tau}} pointwise, with values in (0,e∥φ∥∞](0,\mathrm{e}^{\lVert\varphi\rVert_{\infty}}]. The recursion (2.2) applies the same operators Tm,sT_{m,s} to both terminal data. Each is nondecreasing, so by induction over the steps the functions obtained from U~ℓ\widetilde U_{\ell} decrease in ℓ\ell; they are bounded above by ∥φ∥∞\lVert\varphi\rVert_{\infty}, and their limits are the functions obtained from Hκ+φτH_{\kappa}+\varphi_{\tau}, which are finite by (84). Indeed, at a step with m>0m>0 the functions emfℓ∈(0,em∥φ∥∞]\mathrm{e}^{m f_{\ell}}\in(0,\mathrm{e}^{m\lVert\varphi\rVert_{\infty}}] decrease to emf\mathrm{e}^{m f}, and at a step with m=0m=0 the functions ∥φ∥∞−fℓ≥0\lVert\varphi\rVert_{\infty}-f_{\ell}\geq0 increase to ∥φ∥∞−f\lVert\varphi\rVert_{\infty}-f; bounded and monotone convergence apply. Hence PU~ℓ(γ)↓PHκ+φτ(γ)\mathcal{P}_{\widetilde U_{\ell}}(\gamma)\downarrow\mathcal{P}_{H_{\kappa}+\varphi_{\tau}}(\gamma), and PHκ+φτ(γ)≤PHκ+φ(γ)+αLφτ\mathcal{P}_{H_{\kappa}+\varphi_{\tau}}(\gamma)\leq\mathcal{P}_{H_{\kappa}+\varphi}(\gamma)+\alpha L_{\varphi}\tau by (84).

Now choose τ\tau with (α′+α)Lφτ≤ε/3(\alpha'+\alpha)L_{\varphi}\tau\leq\varepsilon/3, and then ℓ\ell with PU~ℓ(γ)≤PHκ+φτ(γ)+ε/3\mathcal{P}_{\widetilde U_{\ell}}(\gamma)\leq\mathcal{P}_{H_{\kappa}+\varphi_{\tau}}(\gamma)+\varepsilon/3. Then PU~ℓ(γ)+α′Lφτ≤PHκ+φ(γ)+2ε/3\mathcal{P}_{\widetilde U_{\ell}}(\gamma)+\alpha' L_{\varphi}\tau\leq\mathcal{P}_{H_{\kappa}+\varphi}(\gamma)+2\varepsilon/3, and Theorem 3.1, applied to U~ℓ\widetilde U_{\ell} and γ\gamma, gives

P{1Nlog⁡∫SN(κ)e∑aφ(ba⋅σ) μN(dσ)>PHκ+φ(γ)+ε}≤P{fN(U~ℓ)>PU~ℓ(γ)+ε3}⟶0.\mathbb{P}\left\{ \frac{1}{N}\log\int_{S_{N}(\kappa)} \mathrm{e}^{\sum_{a}\varphi(b_{a}\cdot\sigma)}\,\mu_{N}(\mathrm{d}\sigma) > \mathcal{P}_{H_{\kappa}+\varphi}(\gamma)+\varepsilon \right\} \leq \mathbb{P}\left\{ f_{N}(\widetilde U_{\ell}) > \mathcal{P}_{\widetilde U_{\ell}}(\gamma)+\frac{\varepsilon}{3} \right\} \longrightarrow0.

Proof of (i). Let γ∈U\gamma\in\mathcal{U} and r=Qγ<1r=Q_{\gamma}<1. The step order parameters γn∈Ur\gamma_{n}\in\mathcal{U}_{r} of §2.1 converge to γ\gamma in L1([0,1])L^{1}([0,1]), so uγn(0,0;κ)→uγ(0,0;κ)u_{\gamma_{n}}(0,0;\kappa)\to u_{\gamma}(0,0;\kappa) by the definition of uγu_{\gamma} through Lemma 2.7, and CS(γn)→CS(γ)\mathrm{CS}(\gamma_{n})\to\mathrm{CS}(\gamma) by (2.12). Hence there is a step order parameter γ′∈Ur\gamma'\in\mathcal{U}_{r} with Pκ(γ′)≤Pκ(γ)+ε/2\mathcal{P}_{\kappa}(\gamma')\leq\mathcal{P}_{\kappa}(\gamma)+\varepsilon/2. Part (ii) with φ=0\varphi=0, the order parameter γ′\gamma', and ε/2\varepsilon/2 gives the claim, since PHκ(γ′)=Pκ(γ′)\mathcal{P}_{H_{\kappa}}(\gamma')=\mathcal{P}_{\kappa}(\gamma'). □\square

Volume from one feasible point

A bound on the volume does not by itself bound the largest feasible margin, because a feasible set may have zero spherical volume. Lemma 3.13 shows that one feasible point at margin κ\kappa produces exponentially small, but not superexponentially small, volume at margin κ−η\kappa-\eta, on a set where the fields ba⋅σb_{a}\cdot\sigma change little. The set it constructs is used again in Chapter 9. The proof uses Šidák’s inequality for a Gaussian vector whose covariance may be singular; Chapter 8 uses the same extension.

Lemma 3.12. Let (ξ1,…,ξM)(\xi_{1},\ldots,\xi_{M}) be a centered Gaussian vector with an arbitrary, possibly singular, covariance matrix, and let c1,…,cM≥0c_{1},\ldots,c_{M}\ge0. Then

P{∣ξa∣≤ca for all a≤M}≥∏a=1MP{∣ξa∣≤ca}.\mathbb{P}\{|\xi_{a}|\le c_{a}\ \text{for all }a\le M\}\ge\prod_{a=1}^{M}\mathbb{P}\{|\xi_{a}|\le c_{a}\}.

Proof. For a nonsingular covariance matrix this is Šidák’s inequality [ref-116], Corollary 1, proved independently by Khatri [ref-80]. If ca=0c_{a}=0 and Var⁡ξa>0\operatorname{Var}\xi_{a}>0 for some aa, the right side vanishes, and a coordinate with ca=0c_{a}=0 and ξa=0\xi_{a}=0 contributes the factor one to both sides and may be removed. So we may assume that all ca>0c_{a}>0. Let ξ′∼N(0,IM)\xi'\sim N(0,I_{M}) be independent of ξ\xi, and for ϵ>0\epsilon>0 put ξϵ=ξ+ϵξ′\xi^{\epsilon}=\xi+\epsilon\xi', whose covariance matrix is nonsingular. Then P{∣ξaϵ∣≤ca ∀a}≥∏aP{∣ξaϵ∣≤ca}\mathbb{P}\{|\xi_{a}^{\epsilon}|\le c_{a}\ \forall a\}\ge\prod_{a}\mathbb{P}\{|\xi_{a}^{\epsilon}|\le c_{a}\}. Let ϵ↓0\epsilon\downarrow0 along a sequence. Each factor converges to P{∣ξa∣≤ca}\mathbb{P}\{|\xi_{a}|\le c_{a}\}: if Var⁡ξa>0\operatorname{Var}\xi_{a}>0, because ξaϵ\xi_{a}^{\epsilon} and ξa\xi_{a} are centered Gaussian with variances converging and a continuous limiting law; if ξa=0\xi_{a}=0, because P{∣ϵξa′∣≤ca}→1\mathbb{P}\{|\epsilon\xi_{a}'|\le c_{a}\}\to1. Since ξϵ→ξ\xi^{\epsilon}\to\xi pointwise and the intervals are closed, the indicator of {∣ξaϵ∣≤ca ∀a}\{|\xi_{a}^{\epsilon}|\le c_{a}\ \forall a\} has lim sup⁡\limsup at most the indicator of {∣ξa∣≤ca ∀a}\{|\xi_{a}|\le c_{a}\ \forall a\}. Fatou’s lemma, applied to one minus the indicators, proves the lemma. □

Lemma 3.13. Fix α′<∞\alpha'<\infty. There are constants Kcap≥1K_{\mathrm{cap}}\ge1, Ccap<∞C_{\mathrm{cap}}<\infty, and N0N_{0}, depending only on α′\alpha', with the following property. Let N≥N0N\ge N_{0}, M≤α′NM\le\alpha'N, and let b1,…,bM∈RNb_{1},\ldots,b_{M}\in\mathbb{R}^{N} be deterministic vectors with ∣ba∣2≤2|b_{a}|^{2}\le2. Let κ∈R\kappa\in\mathbb{R} and η∈(0,1]\eta\in(0,1], suppose that some σ0∈SN\sigma_{0}\in S_{N} satisfies ba⋅σ0≥κb_{a}\cdot\sigma_{0}\ge\kappa for all aa, and put

s0=min⁡{η2Kcap,η2∣κ∣+1}.s_{0}=\min\left\{\frac{\eta}{2K_{\mathrm{cap}}},\sqrt{\frac{\eta}{2|\kappa|+1}}\right\}.

There is a Borel set T⊂{σ∈SN:ba⋅σ≥κ−η for all a}T\subset\{\sigma\in S_{N}:b_{a}\cdot\sigma\ge\kappa-\eta\ \text{for all }a\} with

μN(T)≥exp⁡{−N(log⁡(2/s0)+Ccap)},∣ba⋅σ−ba⋅σ0∣≤s02∣ba⋅σ0∣+s0Kcap(σ∈T, a≤M).(86)\begin{aligned} \mu_{N}(T)&\ge\exp\left\{-N\left(\log(2/s_{0})+C_{\mathrm{cap}}\right)\right\},\\ |b_{a}\cdot\sigma-b_{a}\cdot\sigma_{0}|&\le s_{0}^{2}|b_{a}\cdot\sigma_{0}|+s_{0}K_{\mathrm{cap}}\qquad(\sigma\in T,\ a\le M). \tag*{(86)} \end{aligned}

If κ\kappa ranges over a compact interval II, then log⁡(2/s0)≤log⁡(1/η)+CI\log(2/s_{0})\le\log(1/\eta)+C_{I} with CIC_{I} depending only on α′\alpha' and II.

Proof. We divide the proof into four steps.

Step 1. Constants. Let c0=12(log⁡2−12)>0c_{0}=\frac{1}{2}(\log2-\frac{1}{2})>0. For X∼χd2X\sim\chi_{d}^{2} and 0<y<10<y<1, Markov’s inequality applied to e−uXe^{-uX} with u=(1−y)/(2y)u=(1-y)/(2y) gives the Chernoff bound P{X≤yd}≤euyd(1+2u)−d/2=exp⁡{−d2(y−1−log⁡y)}\mathbb{P}\{X\le yd\}\le e^{uyd}(1+2u)^{-d/2}=\exp\{-\frac{d}{2}(y-1-\log y)\}; with y=12y=\frac{1}{2} this is P{X≤d/2}≤e−c0d\mathbb{P}\{X\le d/2\}\le e^{-c_{0}d}. Let π(z)=P{∣2G∣≤z/2}\pi(z)=\mathbb{P}\{|\sqrt{2}G|\le z/2\}, and choose Kcap≥1K_{\mathrm{cap}}\ge1 with α′log⁡(1/π(Kcap))≤c0/4\alpha'\log(1/\pi(K_{\mathrm{cap}}))\le c_{0}/4, which is possible because π(z)→1\pi(z)\to1 as z→∞z\to\infty. Put Ccap=c0/4+log⁡2C_{\mathrm{cap}}=c_{0}/4+\log2, and let N0≥3N_{0}\ge3 be such that e−c0(N−1)≤12e−c0N/4e^{-c_{0}(N-1)}\le\frac{1}{2}e^{-c_{0}N/4} for N≥N0N\ge N_{0}.

Step 2. Coordinates. Let PP be the orthogonal projection onto σ0⊥\sigma_{0}^{\perp}. Every σ∈SN\sigma\in S_{N} with ∣σ⋅σ0∣<N|\sigma\cdot\sigma_{0}| < N can be written uniquely as

σ=ρσ0+sN v,ρ=σ⋅σ0N∈(−1,1),s=1−ρ2,v∈σ0⊥, ∣v∣=1.\sigma= \rho\sigma_{0} + s\sqrt{N}\,v,\qquad\rho= \frac{\sigma\cdot\sigma_{0}}{N} \in(-1,1),\qquad s = \sqrt{1-\rho^{2}},\qquad v \in\sigma_{0}^{\perp},\ |v|=1.

Under μN\mu_{N}, ρ\rho and vv are independent, vv is uniform on the unit sphere of σ0⊥\sigma_{0}^{\perp}, and ρ\rho has density cN′(1−ρ2)(N−3)/2c_{N}'(1-\rho^{2})^{(N-3)/2} on (−1,1)(-1,1), with cN′=Γ(N/2)/(πΓ((N−1)/2))c_{N}'=\Gamma(N/2)/(\sqrt{\pi}\Gamma((N-1)/2)). This constant is nondecreasing in NN by log-convexity of Γ\Gamma, and equals 1/21/2 at N=3N=3. Let

T={σ:ρ∈[1−s02,1−s02/4], N∣ba⋅v∣≤Kcap for all a≤M}.T=\left\{\sigma:\rho\in\left[\sqrt{1-s_{0}^{2}},\sqrt{1-s_{0}^{2}/4}\right],\ \sqrt{N}|b_{a}\cdot v|\le K_{\mathrm{cap}}\ \text{for all }a\le M\right\}.

It is a Borel set, and on it s∈[s0/2,s0]s\in[s_{0}/2,s_{0}] and ρ≥0\rho\ge0. Also s0≤η/(2Kcap)≤1/2s_{0}\le\eta/(2K_{\mathrm{cap}})\le1/2.

Step 3. Fields. For σ∈T\sigma\in T, ba⋅σ−ba⋅σ0=(ρ−1)ba⋅σ0+sN ba⋅vb_{a}\cdot\sigma-b_{a}\cdot\sigma_{0}=(\rho-1)b_{a}\cdot\sigma_{0}+s\sqrt{N}\,b_{a}\cdot v, and 0≤1−ρ≤s2≤s020\le1-\rho\le s^{2}\le s_{0}^{2} because ρ=1−s2≥1−s2\rho=\sqrt{1-s^{2}}\ge1-s^{2}. This gives the second bound in (86). For feasibility, ba⋅σ0≥κb_{a}\cdot\sigma_{0}\ge\kappa and ρ∈[1−s02,1]\rho\in[1-s_{0}^{2},1] imply ρba⋅σ0≥κ−s02∣κ∣\rho b_{a}\cdot\sigma_{0}\ge\kappa-s_{0}^{2}|\kappa|: if κ>0\kappa>0, then ρba⋅σ0≥(1−s02)κ\rho b_{a}\cdot\sigma_{0}\ge(1-s_{0}^{2})\kappa; if ba⋅σ0≥0≥κb_{a}\cdot\sigma_{0}\ge0\ge\kappa, the left side is nonnegative; and if κ≤ba⋅σ0<0\kappa\le b_{a}\cdot\sigma_{0}<0, then ρba⋅σ0≥ba⋅σ0\rho b_{a}\cdot\sigma_{0}\ge b_{a}\cdot\sigma_{0}. Therefore

ba⋅σ≥κ−s02∣κ∣−s0Kcap≥κ−η2−η2,b_{a}\cdot\sigma\ge\kappa-s_{0}^{2}|\kappa|-s_{0}K_{\mathrm{cap}}\ge\kappa-\frac{\eta}{2}-\frac{\eta}{2},

using s02≤η/(2∣κ∣+1)s_{0}^{2}\le\eta/(2|\kappa|+1) and s0≤η/(2Kcap)s_{0}\le\eta/(2K_{\mathrm{cap}}). Thus every σ∈T\sigma\in T is feasible at margin κ−η\kappa-\eta.

Step 4. Volume. On the interval of ρ\rho defining TT we have 1−ρ2≥s02/41-\rho^{2}\ge s_{0}^{2}/4, and the interval has length 34s02/(1−s02/4+1−s02)≥38s02\frac{3}{4}s_{0}^{2}/(\sqrt{1-s_{0}^{2}/4}+\sqrt{1-s_{0}^{2}})\ge\frac{3}{8}s_{0}^{2}. Hence its probability is at least 12(s0/2)N−3⋅38s02=32s0(s0/2)N≥(s0/2)N\frac{1}{2}(s_{0}/2)^{N-3}\cdot\frac{3}{8}s_{0}^{2}=\frac{3}{2s_{0}}(s_{0}/2)^{N}\ge(s_{0}/2)^{N}. For the direction, represent v=Pζ~/∣Pζ~∣v=P\widetilde{\zeta}/|P\widetilde{\zeta}| with ζ~∼N(0,IN)\widetilde{\zeta}\sim N(0,I_{N}), so that ba⋅v=(Pba)⋅ζ~/∣Pζ~∣b_{a}\cdot v=(Pb_{a})\cdot\widetilde{\zeta}/|P\widetilde{\zeta}|. On the event {∣Pζ~∣2≥N/4}∩{∣(Pba)⋅ζ~∣≤Kcap/2 for all a}\{|P\widetilde{\zeta}|^{2}\ge N/4\}\cap\{|(Pb_{a})\cdot\widetilde{\zeta}|\le K_{\mathrm{cap}}/2\text{ for all }a\} we have N∣ba⋅v∣≤Kcap\sqrt{N}|b_{a}\cdot v|\le K_{\mathrm{cap}} for all aa. The variables (Pba)⋅ζ~(Pb_{a})\cdot\widetilde{\zeta} are centered and jointly Gaussian with variances ∣Pba∣2≤2|Pb_{a}|^{2}\le2, and P{∣N(0,v)∣≤Kcap/2}\mathbb{P}\{|N(0,v)|\le K_{\mathrm{cap}}/2\} is nonincreasing in vv. By Lemma 3.12, the choice of KcapK_{\mathrm{cap}}, and M≤α′NM\le\alpha' N,

P{∣(Pba)⋅ζ~∣≤Kcap/2 for all a}≥π(Kcap)M≥e−c0N/4.\mathbb{P}\left\{|(Pb_{a})\cdot\widetilde{\zeta}|\le K_{\mathrm{cap}}/2\ \text{for all }a\right\}\ge\pi(K_{\mathrm{cap}})^{M}\ge e^{-c_{0}N/4}.

Since ∣Pζ~∣2|P\widetilde{\zeta}|^{2} has the law χN−12\chi_{N-1}^{2} and N/4≤(N−1)/2N/4\le(N-1)/2, Step 1 gives P{∣Pζ~∣2<N/4}≤e−c0(N−1)≤12e−c0N/4\mathbb{P}\{|P\widetilde{\zeta}|^{2}<N/4\}\le e^{-c_{0}(N-1)}\le\frac{1}{2}e^{-c_{0}N/4}. The directional event therefore has probability at least 12e−c0N/4\frac{1}{2}e^{-c_{0}N/4}. By independence of ρ\rho and vv,

μN(T)≥(s02)N12e−c0N/4≥exp⁡{−N(log⁡(2/s0)+Ccap)}.\mu_{N}(T)\ge\left(\frac{s_{0}}{2}\right)^{N}\frac{1}{2}e^{-c_{0}N/4}\ge\exp\left\{-N\left(\log(2/s_{0})+C_{\mathrm{cap}}\right)\right\}.

Finally, for κ∈I\kappa\in I and η≤1\eta\le1 we have η/(2∣κ∣+1)≥η/2∣κ∣+1\sqrt{\eta/(2|\kappa|+1)}\ge\eta/\sqrt{2|\kappa|+1}, so s0≥cIηs_{0}\ge c_{I}\eta with cI=min⁡{1/(2Kcap),(2sup⁡κ∈I∣κ∣+1)−1/2}c_{I}=\min\{1/(2K_{\mathrm{cap}}),(2\sup_{\kappa\in I}|\kappa|+1)^{-1/2}\}, and log⁡(2/s0)≤log⁡(1/η)+log⁡(2/cI)\log(2/s_{0})\le\log(1/\eta)+\log(2/c_{I}). □\square

Corollary 3.14. Let α>0\alpha>0. For every real κ>κc(α)\kappa>\kappa_{\mathrm{c}}(\alpha), P{κN≥κ}→0\mathbb{P}\{\kappa_{N}\ge\kappa\}\to0.

Proof. Choose η∈(0,1]\eta\in(0,1] with κ−η>κc\kappa-\eta>\kappa_{c}, and let N0N_{0} and CV=log⁡(2/s0)+CcapC_{V}=\log(2/s_{0})+C_{\mathrm{cap}} be given by Lemma 3.13 for α′=sup⁡NαN\alpha'=\sup_{N}\alpha_{N}, the margin κ\kappa, and this η\eta. Since κ−η>κc\kappa-\eta>\kappa_{c}, the definition of κc\kappa_{c} gives P∗(κ−η)=−∞\mathcal{P}_{*}(\kappa-\eta)=-\infty, so there is γ∈U\gamma\in\mathcal{U} with Pκ−η(γ)<−CV−1\mathcal{P}_{\kappa-\eta}(\gamma)<-C_{V}-1. By Corollary 3.11(i) with ε=1\varepsilon=1,

P{1Nlog⁡VN(κ−η)≥−CV}≤P{1Nlog⁡VN(κ−η)>Pκ−η(γ)+1}⟶0.\mathbb{P}\left\{\frac{1}{N}\log V_{N}(\kappa-\eta)\geq-C_{V}\right\} \leq \mathbb{P}\left\{\frac{1}{N}\log V_{N}(\kappa-\eta)>\mathcal{P}_{\kappa-\eta}(\gamma)+1\right\} \longrightarrow0.

On the other hand, suppose that N≥N0N\geq N_{0}, κN≥κ\kappa_{N}\geq\kappa, and max⁡a∣ba∣2≤2\max_{a}|b_{a}|^{2}\leq2. A maximizer σ0\sigma_{0} of min⁡aba⋅σ\min_{a}b_{a}\cdot\sigma over the compact set SNS_{N} satisfies ba⋅σ0≥κb_{a}\cdot\sigma_{0}\geq\kappa for all aa, so Lemma 3.13 gives VN(κ−η)≥μN(T)≥e−NCVV_{N}(\kappa-\eta)\geq\mu_{N}(T)\geq e^{-NC_{V}}. Finally, N∣ba∣2N|b_{a}|^{2} has the law χN2\chi_{N}^{2}, and the Chernoff bound P{χN2≥2N}≤e−N(1−log⁡2)/2\mathbb{P}\{\chi_{N}^{2}\geq2N\}\leq e^{-N(1-\log2)/2} (the upper-tail analogue of Step 1 in the proof of Lemma 3.13) gives

P{max⁡a∣ba∣2>2}≤α′Ne−N(1−log⁡2)/2⟶0.\mathbb{P}\left\{\max_{a}|b_{a}|^{2}>2\right\} \leq \alpha'Ne^{-N(1-\log2)/2} \longrightarrow0.

Hence P{κN≥κ}≤P{N−1log⁡VN(κ−η)≥−CV}+P{max⁡a∣ba∣2>2}→0\mathbb{P}\{\kappa_{N}\geq\kappa\}\leq\mathbb{P}\{N^{-1}\log V_{N}(\kappa-\eta)\geq-C_{V}\}+\mathbb{P}\{\max_{a}|b_{a}|^{2}>2\}\to0. □\square

The lower bound

In this chapter we prove the lower bound lim inf⁡NFN(U)≥P∗(U)\liminf_{N} F_{N}(U) \ge\mathcal{P}_{*}(U) for smooth activations UU (Theorem 4.1). Together with the upper bound of Chapter 3, it gives the Parisi formula in Chapter 5. The argument is outlined in Step 2 of §1.3.

§4.1 introduces the spherical coordinates, the Poisson row count, and the three systems that are compared, and §4.2 proves the overlap representation theorem (Theorem 4.6) from the external results behind it. §4.3 defines the perturbation and states the perturbation, continuity, and precision lemmas, which are proved in §4.8. §4.4 reduces the increment to the row term and the coordinate term. §4.5 evaluates them on cascades with Lemma 3.4 and removes the cutoff on ∥ε∥\lVert\varepsilon\rVert there (Lemma 4.10). §4.6 proves the rotation estimate, and §4.7 proves Theorem 4.1.

The two cavity steps, adding constraints and adding coordinates, follow the cavity method for the perceptron in the replica-symmetric regime [ref-115], [ref-127], Chapters 2 and 3, [ref-128], Chapter 8. As in [ref-37, ref-99], the trial order parameter comes from the limiting overlap law of the system. Compared with W.-K. Chen’s spherical form [ref-37] of the Aizenman–Sims–Starr scheme and with its multi-species version by Bates and Sohn [ref-16], two points differ. First, in both works the cavity functional is identified with the spherical Parisi functional, as the number of cavity coordinates grows, through Talagrand’s computation [ref-125]; here the coordinate term is an explicit Gaussian integral on cascades, and the rotation invariance of the patterns identifies it with the Crisanti–Sommers entropy. Second, the other species in the Ghirlanda–Guerra identities is not a group of coordinates, as in [ref-16, ref-100], but the response vector.

Activation class. Throughout this chapter, UU satisfies, for some Umax⁡<∞U_{\max} < \infty and L<∞L < \infty (enlarged, if necessary, so that L≥1+∣U(0)∣L \ge1 + |U(0)|),

U∈C4(R),U≤Umax⁡,∣U′(x)∣≤L(1+∣x∣),∣U(j)(x)∣≤L(2≤j≤4),(87)U \in C^{4}(\mathbb{R}), \qquad U \le U_{\max}, \qquad|U'(x)| \le L(1+|x|), \qquad|U^{(j)}(x)| \le L \quad(2 \le j \le4), \tag*{(87)}

and in addition either U′U' is bounded or UU is concave. In the first case UU belongs to class (A), and in the second to class (B). The derivatives of orders three and four are used only in the perturbation of §4.3, in the error estimates of the cavity interpolation, and in Lemma 4.9. They control the first and second derivatives of xU′(x)xU'(x) and U′(x)2U'(x)^{2}. In Chapter 5 this extra smoothness is removed by mollification. By (87) and Taylor’s formula,

−C0(1+x2)≤U(x)≤Umax⁡,C0=2L.(88)-C_{0}(1+x^{2}) \le U(x) \le U_{\max}, \qquad C_{0}=2L. \tag*{(88)}

Conventions. Throughout this chapter, CC and cc denote positive constants that may depend on α\alpha, Umax⁡U_{\max}, and LL, but not on NN; their values may change from one occurrence to the next. A dependence on the cavity dimension nn or on the cutoff Λ\Lambda is indicated by a subscript, as in Cn,ΛC_{n,\Lambda} or On,Λ(⋅)O_{n,\Lambda}(\cdot). The cavity dimension nn is fixed until the last step of the proof. We write νn\nu_{n} for the standard Gaussian measure on Rn\mathbb{R}^{n} and GG for a standard Gaussian variable. The letter aa is a row index, and G,G′,G′′\mathbf{G},\mathbf{G}',\mathbf{G}'' denote Gaussian matrices with rows ga,ga′,ga′′g_{a},g'_{a},g''_{a}. The letter bb, with or without subscripts, denotes a precision parameter of a Gaussian integral, as in Lemma 3.4. The rows ba=ga/Nb_{a}=g_{a}/\sqrt{N} of Chapter 1 and the derivative b=∂xub=\partial_{x}u of §2.3 are not used in this chapter.

The goal of the chapter is the following theorem.

Theorem 4.1. For every α>0\alpha>0 and every activation UU satisfying (87), with either bounded U′U' or concave UU,

lim inf⁡N→∞FN(U)≥P∗(U)=inf⁡γ∈U{αuγ(0,0;U)+CS(γ)}.\liminf_{N\to\infty}F_{N}(U)\ge P_{*}(U) =\inf_{\gamma\in\mathcal{U}}\left\{\alpha u_{\gamma}(0,0;U)+\mathrm{CS}(\gamma)\right\}.

Spherical coordinates, the row count, and the three systems

We first separate the added coordinates without changing the surface measure. Throughout this chapter

N′=N+n.N'=N+n.

For σ∈SN\sigma\in S_{N} and ε∈Rn\varepsilon\in\mathbb{R}^{n} with ∥ε∥2<N′\lVert\varepsilon\rVert^{2}<N', define

χ(ε)=N′−∥ε∥2N,ρ(σ,ε)=(χ(ε)σ,ε)∈SN′.\chi(\varepsilon)=\sqrt{\frac{N'-\lVert\varepsilon\rVert^{2}}{N}}, \qquad \rho(\sigma,\varepsilon)=(\chi(\varepsilon)\sigma,\varepsilon)\in S_{N'}.

The uniform probability measure μN′\mu_{N'} on SN′S_{N'} is the image of μN(dσ)⊗fN,n(ε) dε\mu_{N}(\mathrm{d}\sigma)\otimes f_{N,n}(\varepsilon)\,\mathrm{d}\varepsilon under (σ,ε)↦ρ(σ,ε)(\sigma,\varepsilon)\mapsto\rho(\sigma,\varepsilon), where

fN,n(ε)=Γ(N′/2)Γ(N/2)(N′/2)n/21(2π)n/2(1−∥ε∥2N′)N/2−11{∥ε∥2<N′}.(89)f_{N,n}(\varepsilon) = \frac{\Gamma(N'/2)} {\Gamma(N/2)(N'/2)^{n/2}} \frac{1}{(2\pi)^{n/2}} \left(1-\frac{\lVert\varepsilon\rVert^{2}}{N'}\right)^{N/2-1} \mathbf{1}_{\{\lVert\varepsilon\rVert^{2}<N'\}}. \tag*{(89)}

Indeed, write a uniform point of SN′S_{N'} as N′ η/∥η∥\sqrt{N'}\,\eta/\lVert\eta\rVert for a standard Gaussian vector η=(η1,η2)∈RN×Rn\eta=(\eta_{1},\eta_{2})\in\mathbb{R}^{N}\times\mathbb{R}^{n}. Its last nn coordinates are ε=N′η2/∥η∥\varepsilon=\sqrt{N'}\eta_{2}/\lVert\eta\rVert, and its first NN coordinates are χ(ε)σ\chi(\varepsilon)\sigma with σ=Nη1/∥η1∥\sigma=\sqrt{N}\eta_{1}/\lVert\eta_{1}\rVert. The vector σ\sigma is uniform on SNS_{N} and independent of (∥η1∥,η2)(\lVert\eta_{1}\rVert,\eta_{2}), hence of ε\varepsilon. The variable ∥ε∥2/N′=∥η2∥2/∥η∥2\lVert\varepsilon\rVert^{2}/N'=\lVert\eta_{2}\rVert^{2}/\lVert\eta\rVert^{2} has the beta distribution with parameters n/2n/2 and N/2N/2, and the direction of ε\varepsilon is uniform and independent of its norm. Writing the beta density in polar coordinates on Rn\mathbb{R}^{n} gives (89). For Λ≥1\Lambda\ge1 let

BΛ={ε∈Rn:∥ε∥2≤Λn}.\mathbb{B}_{\Lambda} = \left\{\varepsilon\in\mathbb{R}^{n}:\lVert\varepsilon\rVert^{2}\le\Lambda n\right\}.

Stirling’s formula, Γ(N′/2)/Γ(N/2)=(N′/2)n/2(1+On(N−1))\Gamma(N'/2)/\Gamma(N/2)=(N'/2)^{n/2}(1+O_n(N^{-1})), and the expansion (N/2−1)log⁡(1−∥ε∥2/N′)=−∥ε∥2/2+On,Λ(N−1)(N/2-1)\log(1-\lVert\varepsilon\rVert^2/N')=-\lVert\varepsilon\rVert^2/2+O_{n,\Lambda}(N^{-1}) give, uniformly on BΛ\mathbb{B}_{\Lambda},

fN,n(ε)=e−∥ε∥2/2(2π)n/2(1+On,Λ(N−1)),χ(ε)=1+n−∥ε∥22N+On,Λ(N−2).(90)f_{N,n}(\varepsilon)=\frac{e^{-\lVert\varepsilon\rVert^2/2}}{(2\pi)^{n/2}}\left(1+O_{n,\Lambda}(N^{-1})\right),\qquad \chi(\varepsilon)=1+\frac{n-\lVert\varepsilon\rVert^2}{2N}+O_{n,\Lambda}(N^{-2}). \tag*{(90)}

[ref-37] separates the cavity coordinates of spherical models by a decomposition of the uniform measure of the same kind. The expansion of fN,nf_{N,n} reflects the classical fact that a fixed number of coordinates of a uniform point on a high-dimensional sphere are asymptotically independent standard Gaussian variables; see [ref-47].

Gaussian matrices. We use the following standard bound on the operator norm of a Gaussian matrix [ref-130], Corollary 5.35, here and in Chapters 5 and 8.

Lemma 4.2. Let AA be an M×NM\times N matrix with independent standard Gaussian entries. For every t≥0t\ge0,

P{∥A∥op>M+N+t}≤2e−t2/2.\mathbb{P}\left\{\lVert A\rVert_{\mathrm{op}}>\sqrt{M}+\sqrt{N}+t\right\}\le2e^{-t^2/2}.

Consequently E∥A∥op2≤4(M+N)+8\mathbb{E}\lVert A\rVert_{\mathrm{op}}^2\le4(M+N)+8, and Eexp⁡(c∥A∥op2/N)≤3exp⁡(4c(M+N)/N)\mathbb{E}\exp(c\lVert A\rVert_{\mathrm{op}}^2/N)\le3\exp(4c(M+N)/N) for 0<c≤N/80<c\le N/8.

Proof. The tail bound is the upper half of [ref-130], Corollary 5.35. Put T=(∥A∥op−M−N)+T=(\lVert A\rVert_{\mathrm{op}}-\sqrt{M}-\sqrt{N})_+, so that P(T>t)≤2e−t2/2\mathbb{P}(T>t)\le2e^{-t^2/2} and ∥A∥op2≤4(M+N)+2T2\lVert A\rVert_{\mathrm{op}}^2\le4(M+N)+2T^2. Then ET2=∫0∞2tP(T>t) dt≤4\mathbb{E}T^2=\int_0^\infty2t\mathbb{P}(T>t)\,\mathrm{d}t\le4, which gives the second bound. For 0<s≤1/40<s\le1/4,

EesT2=1+∫0∞2stest2P(T>t) dt≤1+2s1/2−s≤3,\mathbb{E}e^{sT^2} =1+\int_0^\infty2st e^{st^2}\mathbb{P}(T>t)\,\mathrm{d}t \le1+\frac{2s}{1/2-s}\le3,

and the choice s=2c/Ns=2c/N gives the third bound. □\square

Poissonization. We replace the number MNM_N of rows by a Poisson variable M∼Poi⁡(αN)M\sim\operatorname{Poi}(\alpha N), and when passing to dimension N′N' we add independently π∼Poi⁡(αn)\pi\sim\operatorname{Poi}(\alpha n) rows, so that the N′N'-dimensional system has M+π∼Poi⁡(αN′)M+\pi\sim\operatorname{Poi}(\alpha N') rows; from now on the letter π\pi denotes this count. This changes the free energy by o(1)o(1). To check this, let hN(j)h_N(j) be the expected log partition function in dimension NN with jj rows. Adding a fresh row gg changes hNh_N by Elog⁡⟨eU(g⋅σ/N)⟩\mathbb{E}\log\langle e^{U(g\cdot\sigma/\sqrt{N})}\rangle, where ⟨⋅⟩\langle\cdot\rangle is the Gibbs measure before the row is added. By Jensen’s inequality and U≤Umax⁡U\le U_{\max}, this lies between EU(G)\mathbb{E}U(G) and Umax⁡U_{\max}, because g⋅σ/Ng\cdot\sigma/\sqrt{N} is standard Gaussian for every fixed σ\sigma and gg is independent of the Gibbs measure. By (88), ∣EU(G)∣<∞\lvert\mathbb{E}U(G)\rvert<\infty. Hence hNh_N is Lipschitz with a constant independent of NN, and

∣EhN(M)−hN(MN)∣≤CE∣M−MN∣≤C(αN+∣αN−MN∣)=o(N).\left\lvert\mathbb{E}h_N(M)-h_N(M_N)\right\rvert \le C\mathbb{E}\lvert M-M_N\rvert \le C\left(\sqrt{\alpha N}+\lvert\alpha N-M_N\rvert\right) =o(N).

Poisson numbers of constraints are standard in the study of diluted spin glasses [ref-57, ref-102].

Conditioning on the number of rows. The concentration estimates below are proved at a fixed number of rows, so we carry out the cavity computation conditionally on the value of MM, for values in the window

WN={j∈N:∣j−αN∣≤N2/3}.W_N=\{j\in\mathbb{N}:|j-\alpha N|\leq N^{2/3}\}.

The Poisson tail bound P(∣M−αN∣>t)≤2exp⁡(−t2/(2(αN+t)))\mathbb{P}(|M-\alpha N|>t)\leq2\exp(-t^2/(2(\alpha N+t))) gives P(M∉WN)≤2e−cN1/3\mathbb{P}(M\notin W_N)\leq2e^{-cN^{1/3}}, and for large NN every j∈WNj\in W_N satisfies j≤2αNj\leq2\alpha N. Step 2 of §4.7 reduces the proof to one value of MM in WNW_N. From §4.3 on, MM denotes a fixed integer in WNW_N, expectations are conditional on this value, and all bounds and all terms o(1)o(1) are uniform in M∈WNM\in W_N; the exceptions are Lemma 4.7(i) and the quantities ΞN\Xi_N of Step 1 of §4.7, which concern the Poissonized systems. The number π∼Poi⁡(αn)\pi\sim\operatorname{Poi}(\alpha n) of rows added in dimension N′N' remains Poisson throughout.

The three systems. We compare three systems. Each has MM rows, and the N′N'-dimensional system has π\pi further rows.

(a) The NN-dimensional system has configuration space SNS_N, rows ga(N)∈RNg_a^{(N)}\in\mathbb{R}^N for a≤Ma\leq M, and Hamiltonian ∑a≤MU(ga(N)⋅σ/N)\sum_{a\leq M}U(g_a^{(N)}\cdot\sigma/\sqrt{N}).

(b) The N′N'-dimensional system has configuration space SN′S_{N'} and rows (ga,ga′)∈RN×Rn(g_a,g'_a)\in\mathbb{R}^N\times\mathbb{R}^n for a≤Ma\leq M, together with π\pi further rows g~i∈RN′\widetilde{g}_i\in\mathbb{R}^{N'}. We write ⟨⋅⟩N′\langle\cdot\rangle_{N'} for its Gibbs measure with the first MM rows only.

(c) The reference system has configuration space SNS_N, the rows ga∈RNg_a\in\mathbb{R}^N for a≤Ma\leq M (the first NN coordinates of the rows of the N′N'-dimensional system), and the fields

Sa(σ)=ga⋅σN′.S_a(\sigma)=\frac{g_a\cdot\sigma}{\sqrt{N'}}.

All rows are independent standard Gaussian vectors. We write G,G′\mathbf{G},\mathbf{G}' for the M×NM\times N and M×nM\times n matrices with rows gag_a and ga′g'_a. In the coordinates ρ=ρ(σ,ε)\rho=\rho(\sigma,\varepsilon) the fields of the N′N'-dimensional system are

(ga,ga′)⋅ρN′=χ(ε)Sa(σ)+wa(ε),wa(ε)=ga′⋅εN′.(91)\frac{(g_a,g'_a)\cdot\rho}{\sqrt{N'}}=\chi(\varepsilon)S_a(\sigma)+w_a(\varepsilon),\qquad w_a(\varepsilon)=\frac{g'_a\cdot\varepsilon}{\sqrt{N'}}. \tag*{(91)}

The fields of the NN-dimensional system can be realized from the same rows. Let ga′′∈RNg''_a\in\mathbb{R}^N be independent standard Gaussian vectors, forming the matrix G′′\mathbf{G}'', and put

ηa(σ)=cNga′′⋅σ,cN2=1N−1N′=nNN′.(92)\eta_a(\sigma)=c_N g''_a\cdot\sigma,\qquad c_N^2=\frac{1}{N}-\frac{1}{N'}=\frac{n}{NN'}. \tag*{(92)}

As processes indexed by σ∈SN\sigma\in S_N, the centered Gaussian families (Sa+ηa)a≤M(S_a+\eta_a)_{a\leq M} and (ga(N)⋅σ/N)a≤M(g_a^{(N)}\cdot\sigma/\sqrt{N})_{a\leq M} have the same covariance σ⋅σ′/N\sigma\cdot\sigma'/N, hence the same law. Both the N′N'-dimensional and the NN-dimensional partition functions are therefore integrals against the reference system with modified fields.

Observables of the reference system. The following quantities determine the coordinate term:

AN(σ)=1N∑a≤MSa(σ)U′(Sa(σ)),BN(σ)=1N′∑a≤MU′′(Sa(σ)),XN(σ,σ′)=1N′∑a≤MU′(Sa(σ))U′(Sa(σ′)),RN(σ,σ′)=σ⋅σ′N,bN(σ)=1+AN(σ)−BN(σ).(93)\begin{aligned} A_N(\sigma) &= \frac{1}{N}\sum_{a \le M} S_a(\sigma)U'(S_a(\sigma)), \qquad B_N(\sigma) = \frac{1}{N'}\sum_{a \le M} U''(S_a(\sigma)), \\ X_N(\sigma,\sigma') &= \frac{1}{N'}\sum_{a \le M} U'(S_a(\sigma))U'(S_a(\sigma')), \qquad R_N(\sigma,\sigma') = \frac{\sigma\cdot\sigma'}{N}, \\ b_N(\sigma) &= 1 + A_N(\sigma) - B_N(\sigma). \tag*{(93)} \end{aligned}

The spin overlap is RNR_N. We call XNX_N the response overlap; it is a Gram matrix, hence positive semidefinite, and its diagonal is XN(σ,σ)=N′−1∑aU′(Sa(σ))2X_N(\sigma,\sigma)=N'^{-1}\sum_a U'(S_a(\sigma))^2. The terms ANA_N and BNB_N arise below from the radial factor χ(ε)\chi(\varepsilon) and from the second-order term of a Gaussian integration by parts, respectively, and bN−1b_N-1 will be the coefficient of −(∥ε∥2−n)/2-(\lVert\varepsilon\rVert^2-n)/2 in the exponent of the Gaussian integral over ε\varepsilon.

Let ΩN\Omega_N be the event on which the operator norms of G\mathbf{G}, G′′\mathbf{G}'', and G′\mathbf{G}' are at most C1NC_1\sqrt{N}. Since M≤2αNM \le2\alpha N, Lemma 4.2 with t=Nt=\sqrt{N} shows that P(ΩNc)≤6e−N/2\mathbb{P}(\Omega_N^c)\le6e^{-N/2} for N≥nN \ge n, with the constant C1=2α+4C_1=\sqrt{2\alpha}+4. Since ∑aSa(σ)2=∥Gσ∥2/N′≤∥G∥op2\sum_a S_a(\sigma)^2=\lVert\mathbf{G}\sigma\rVert^2/N' \le\lVert\mathbf{G}\rVert_{\mathrm{op}}^2, the growth bounds in (4.1) give, for every σ∈SN\sigma\in S_N,

∑a≤MSa(σ)2≤∥G∥op2,∣bN(σ)∣+XN(σ,σ)+∣BN(σ)∣≤C(1+∥G∥op2N).(94)\sum_{a \le M} S_a(\sigma)^2 \le\lVert\mathbf{G}\rVert_{\mathrm{op}}^2,\qquad |b_N(\sigma)|+X_N(\sigma,\sigma)+|B_N(\sigma)| \le C\left(1+\frac{\lVert\mathbf{G}\rVert_{\mathrm{op}}^2}{N}\right). \tag*{(94)}

On ΩN\Omega_N the right sides are at most CNCN and CC. By Lemma 4.2, for every c>0c>0 and N≥N0(c)N \ge N_0(c),

Eexp⁡(csup⁡σ∈SN(∣bN(σ)∣+XN(σ,σ)))≤Cc,(95)\mathbb{E}\exp\left(c\sup_{\sigma\in S_N}\left(|b_N(\sigma)|+X_N(\sigma,\sigma)\right)\right)\le C_c, \tag*{(95)}

uniformly in M≤2αNM \le2\alpha N. In particular, all polynomial moments of these quantities and of ∥G∥op2/N\lVert\mathbf{G}\rVert_{\mathrm{op}}^2/N are bounded.

The overlap representation theorem

The cavity computation requires a description of the limiting joint law of the spin overlap RNR_N and the response overlap XNX_N. We use a general result about random probability measures on Hilbert spaces, which does not involve the perceptron Hamiltonian. In this section we state it, recall the three external results on which it rests, and derive it from them.

Setting. For each N≥1N \ge1, let GN\mathfrak{G}_N be a random probability measure on the product of the unit balls of two separable Hilbert spaces, which may depend on NN. For independent samples (xNℓ,yNℓ)ℓ≥1(x_N^\ell,y_N^\ell)_{\ell\ge1} from GN\mathfrak{G}_N put

Rℓℓ′1,N=xNℓ⋅xNℓ′,Rℓℓ′2,N=yNℓ⋅yNℓ′(ℓ,ℓ′≥1).R_{\ell\ell'}^{1,N}=x_N^\ell\cdot x_N^{\ell'},\qquad R_{\ell\ell'}^{2,N}=y_N^\ell\cdot y_N^{\ell'} \qquad(\ell,\ell' \ge1).

We write E⟨⋅⟩N\mathbb{E}\langle\cdot\rangle_{N} for expectation over the random measure and its samples. The arrays take values in [−1,1][-1,1], they are positive semidefinite, and they are invariant in law under finite permutations of the sample labels. Suppose that the pair of arrays converges in finite-dimensional distribution to a pair (R1,R2)(R^{1},R^{2}), and write E⟨⋅⟩\mathbb{E}\langle\cdot\rangle for expectations of the limiting arrays. The limiting arrays have the same three properties. For r≥1r\geq1, rational v1,v2≥0v_{1},v_{2}\geq0 and an integer m≥1m\geq1 put Q(x,y)=(v1x+v2y)mQ(x,y)=(v_{1}x+v_{2}y)^{m}. For a bounded measurable function ff of the arrays of the first rr samples, the Ghirlanda–Guerra discrepancy of §3.4, written with rr replicas since nn is the cavity dimension here, is

DrN(f,Q)=E⟨fQ(R1,r+11,N,R1,r+12,N)⟩N−1rE⟨f⟩NE⟨Q(R121,N,R122,N)⟩N−1r∑ℓ=2rE⟨fQ(R1ℓ1,N,R1ℓ2,N)⟩N.(96)\begin{aligned} \mathscr{D}_{r}^{N}(f,Q) &=\mathbb{E}\left\langle fQ\left(R_{1,r+1}^{1,N},R_{1,r+1}^{2,N}\right)\right\rangle_{N} -\frac{1}{r}\mathbb{E}\langle f\rangle_{N}\mathbb{E}\left\langle Q\left(R_{12}^{1,N},R_{12}^{2,N}\right)\right\rangle_{N} \\ &\quad-\frac{1}{r}\sum_{\ell=2}^{r}\mathbb{E}\left\langle fQ\left(R_{1\ell}^{1,N},R_{1\ell}^{2,N}\right)\right\rangle_{N}. \tag*{(96)} \end{aligned}

Here ff may depend on the diagonal entries of the arrays. For a single random measure G\mathfrak{G} we drop the index NN and write Dr(f,Q)\mathscr{D}_{r}(f,Q).

Ruelle probability cascade arrays with paired levels. We use the Ruelle probability cascade of §3.1: a depth k≥0k\geq0, parameters 0=θ0<θ1<⋯<θk<θk+1=10=\theta_{0}<\theta_{1}<\cdots<\theta_{k}<\theta_{k+1}=1, masses wl=θl+1−θlw_{l}=\theta_{l+1}-\theta_{l} for 0≤l≤k0\leq l\leq k, weights (vβ)(v_{\beta}) indexed by the leaves β\beta of the infinitely branching tree of depth kk, and the depth β∧β′\beta\wedge\beta' of the last common ancestor of two leaves. If β1\beta^{1} and β2\beta^{2} are sampled independently from the weights (vβ)(v_{\beta}), then the averaged probability that β1∧β2=l\beta^{1}\wedge\beta^{2}=l is wlw_{l}, by (3.4) with the leaf function zero. Given levels 0≤q0≤⋯≤qk0\leq q_{0}\leq\cdots\leq q_{k} and 0≤p0≤⋯≤pk0\leq p_{0}\leq\cdots\leq p_{k}, the Ruelle probability cascade array with paired levels (ql,pl)(q_{l},p_{l}) is the pair of arrays with off-diagonal entries (qβℓ∧βℓ′,pβℓ∧βℓ′)(q_{\beta^{\ell}\wedge\beta^{\ell'}},p_{\beta^{\ell}\wedge\beta^{\ell'}}), ℓ≠ℓ′\ell\neq\ell', where β1,β2,…\beta^{1},\beta^{2},\ldots are independent samples from (vβ)(v_{\beta}). The law of its pair of (1,2)(1,2) entries is ∑l=0kwlδ(ql,pl)\sum_{l=0}^{k}w_{l}\delta_{(q_{l},p_{l})}. A single sequence of levels 0≤s0<⋯<sk0\leq s_{0}<\cdots<s_{k} gives the Ruelle probability cascade array with levels (sl)(s_{l}), whose overlap law is ∑lwlδsl\sum_{l}w_{l}\delta_{s_{l}}; every probability measure with finite support in [0,∞)[0,\infty) is of this form.

External results. We use three results from the literature. The first is the Dovbysh–Sudakov representation [ref-51] of positive semidefinite exchangeable arrays, in the form of [ref-96]. A random array (Rℓℓ′)ℓ,ℓ′≥1(R_{\ell\ell'})_{\ell,\ell'\geq1} is called a Gram–de Finetti array if it is symmetric, every finite principal submatrix is positive semidefinite, and its law is invariant under finite permutations of the indices.

Lemma 4.3. Let (Rℓℓ′)(R_{\ell\ell'}) be a Gram–de Finetti array with Rℓℓ≤1R_{\ell\ell}\leq1 almost surely. There is a random probability measure G\mathfrak{G} on the unit ball of a separable Hilbert space such that, for independent samples σ1,σ2,…\sigma^{1},\sigma^{2},\ldots from G\mathfrak{G}, the off-diagonal array (Rℓℓ′)ℓ≠ℓ′(R_{\ell\ell'})_{\ell\neq\ell'} has the law of (σℓ⋅σℓ′)ℓ≠ℓ′(\sigma^{\ell}\cdot\sigma^{\ell'})_{\ell\neq\ell'}.

Proof. By [ref-96], the law of the array is a mixture over a random parameter ω\omega of the laws of (hℓ⋅hℓ′+aℓδℓℓ′)ℓ,ℓ′(h_{\ell}\cdot h_{\ell'}+a_{\ell}\delta_{\ell\ell'})_{\ell,\ell'}, where (hℓ,aℓ)(h_{\ell},a_{\ell}) are independent samples from a probability measure ηω\eta_{\omega} on ℓ2×[0,∞)\ell^{2} \times[0,\infty). Since ∥hℓ∥2≤∥hℓ∥2+aℓ=Rℓℓ≤1\lVert h_{\ell}\rVert^{2} \le\lVert h_{\ell}\rVert^{2}+a_{\ell}=R_{\ell\ell}\le1 almost surely, ηω\eta_{\omega} is concentrated on pairs with ∥h∥≤1\lVert h\rVert\le1 for almost every ω\omega. The law of the first coordinate under ηω\eta_{\omega} is the required random measure. □\square

The second result, from [ref-49], Sections 5.6–5.7, characterizes the scalar arrays that satisfy the Ghirlanda–Guerra identities. Let G\mathcal{G} be a random probability measure on the unit ball of a separable Hilbert space, with overlap array R=(σℓ⋅σℓ′)ℓ≠ℓ′R=(\sigma^{\ell}\cdot\sigma^{\ell'})_{\ell\ne\ell'} and overlap law ζ(A)=E⟨1{R12∈A}⟩\zeta(A)=\mathbb{E}\langle\mathbf{1}_{\{R_{12}\in A\}}\rangle. We say that G\mathcal{G} satisfies the Ghirlanda–Guerra identities [ref-64] if for every r≥1r\ge1, every bounded measurable function ff of (Rℓℓ′)ℓ≠ℓ′≤r(R_{\ell\ell'})_{\ell\ne\ell'\le r}, and every bounded measurable ψ:R→R\psi:\mathbb{R}\to\mathbb{R},

E⟨fψ(R1,r+1)⟩=1rE⟨f⟩E⟨ψ(R12)⟩+1r∑ℓ=2rE⟨fψ(R1ℓ)⟩.(97)\mathbb{E}\langle f\psi(R_{1,r+1})\rangle = \frac{1}{r}\mathbb{E}\langle f\rangle\mathbb{E}\langle\psi(R_{12})\rangle + \frac{1}{r}\sum_{\ell=2}^{r}\mathbb{E}\langle f\psi(R_{1\ell})\rangle. \tag*{(97)}

Part (a) of the following lemma is Panchenko’s characterization of such measures, which rests on his proof of ultrametricity [ref-97] (see also [ref-98], Theorem 2.14); the positivity of the overlap is Talagrand’s positivity principle (see [ref-98], Theorem 2.16). The identities for Ruelle probability cascades used in part (b), [ref-49], Theorem 5.28, are due to Bovier and Kurkova [ref-24].

Lemma 4.4. Let G\mathcal{G} satisfy the Ghirlanda–Guerra identities, with overlap array RR and overlap law ζ\zeta.

(a) The measure ζ\zeta is supported in [0,1][0,1], and the law of RR is determined by ζ\zeta.

(b) If ζh\zeta_{h}, h≥1h\ge1, are probability measures with finite support in [0,1][0,1] that converge weakly to ζ\zeta, then the Ruelle probability cascade arrays whose overlap laws are ζh\zeta_{h} converge in finite-dimensional distribution to RR.

Proof. Part (a) is [ref-49], Theorem 5.29, together with ∣R12∣≤1|R_{12}|\le1. For (b), [ref-49], Corollary 5.32 shows that the cascade arrays converge in finite-dimensional distribution (the Poisson–Dirichlet cascades of [ref-49] are the Ruelle probability cascades of §3.1). The limit has overlap law ζ\zeta, and it satisfies the identities (97): each cascade array satisfies them by [ref-49], Theorem 5.28, so the limit satisfies them for continuous ff and ψ\psi, and hence for bounded measurable ff and ψ\psi, because for fixed ψ\psi (respectively ff) both sides are integrals against finite signed measures. The cascade array with levels s0<⋯<sks_{0}<\cdots<s_{k} in [0,1][0,1] is the off-diagonal part of the Gram array of independent samples from the random measure ∑βvβδhβ\sum_{\beta}v_{\beta}\delta_{h_{\beta}}, where hβ=∑l=0k(sl−sl−1)1/2eβ∣lh_{\beta}=\sum_{l=0}^{k}(s_{l}-s_{l-1})^{1/2}e_{\beta|l} with s−1=0s_{-1}=0 and an orthonormal family (eν)(e_{\nu}) indexed by the nodes; its diagonal entries equal sk≤1s_{k}\le1. Along a subsequence these Gram arrays converge in finite-dimensional distribution, and the limit is a Gram–de Finetti array with diagonal entries at most one whose off-diagonal part is the limit of the cascade arrays. By Lemma 4.3 this off-diagonal part is the overlap array of a random measure on a unit ball, and this measure satisfies the Ghirlanda–Guerra identities, since these are statements about the law of the off-diagonal array. By (a), the limit has the law of RR. □\square

The third result is Mourrat’s quantitative synchronization theorem [ref-92], Theorem 5.3, a finitary variant of Panchenko’s synchronization theorem; see §3.5. For a real random variable YY let FY−1(v)=inf⁡{s∈R:P(Y≤s)≥v}F_Y^{-1}(v)=\inf\{s\in\mathbb{R}:\mathbb{P}(Y\le s)\ge v\} for v∈[0,1]v\in[0,1], and let VV be uniform on [0,1][0,1]. The pair (FY1−1(V),FY2−1(V))(F_{Y_1}^{-1}(V),F_{Y_2}^{-1}(V)) is the monotone coupling of the laws of Y1Y_1 and Y2Y_2.

Lemma 4.5. Fix an enumeration (ϑi)i≥1(\vartheta_i)_{i\ge1} of Q∩[0,1]\mathbb{Q}\cap[0,1], and put cj(x,y)=(ϑj1x+ϑj2y)j3c_j(x,y)=(\vartheta_{j_1}x+\vartheta_{j_2}y)^{j_3} for j=(j1,j2,j3)j=(j_1,j_2,j_3), as in Lemma 3.8.

(a) For every ϵ>0\epsilon>0 there is δ>0\delta>0 with the following property. Let G\mathfrak{G} be a random probability measure on the product of the unit balls of two Hilbert spaces, with arrays R1,R2R^1,R^2 as above. Suppose that ∣Dr(f,cj)∣≤δ|\mathcal{D}_r(f,c_j)|\le\delta for all r,j1,j2,j3≤⌊δ−1⌋r,j_1,j_2,j_3\le\lfloor\delta^{-1}\rfloor and all continuous functions ff of the arrays of rr samples with ∥f∥∞≤1\lVert f\rVert_\infty\le1, where Dr\mathcal{D}_r is (4.10) for G\mathfrak{G}. Then, for every φ∈C∞(R2)\varphi\in C^\infty(\mathbb{R}^2),

∣E⟨φ(R121,R122)⟩−Eφ(F1−1(V),F2−1(V))∣≤ϵ(∥φ∥∞+∥∇φ∥∞),\left|\mathbb{E}\left\langle\varphi(R_{12}^1,R_{12}^2)\right\rangle-\mathbb{E}\varphi\left(F_1^{-1}(V),F_2^{-1}(V)\right)\right| \le\epsilon\left(\lVert\varphi\rVert_\infty+\lVert\nabla\varphi\rVert_\infty\right),

where Fa−1F_a^{-1} is the function FY−1F_Y^{-1} for the law of R12aR_{12}^a under E⟨⋅⟩\mathbb{E}\langle\cdot\rangle.

(b) If Y1,N→Y1Y_{1,N}\to Y_1 and Y2,N→Y2Y_{2,N}\to Y_2 in law, then the monotone couplings converge in law:

(FY1,N−1(V),FY2,N−1(V))→(FY1−1(V),FY2−1(V)).(F_{Y_{1,N}}^{-1}(V),F_{Y_{2,N}}^{-1}(V))\to(F_{Y_1}^{-1}(V),F_{Y_2}^{-1}(V)).

Proof. Part (a) is [ref-92], Theorem 5.3, whose hypothesis is formulated with the same defects, for continuous test functions of the arrays of rr samples including their diagonal entries. Part (b) is [ref-92], Lemma 5.4. □

Theorem 4.6. In the setting above, assume that for every r≥1r\ge1, every integer m≥1m\ge1, and all rational v1,v2≥0v_1,v_2\ge0,

sup⁡∥f∥∞≤1∣DrN(f,Q)∣⟶0(N→∞),(98)\sup_{\lVert f\rVert_\infty\le1}\left|\mathcal{D}_r^N(f,Q)\right|\longrightarrow0 \qquad(N\to\infty), \tag*{(98)}

where the supremum is over bounded measurable functions of the arrays of the first rr samples. Then the limiting pair satisfies the Ghirlanda–Guerra identities

E⟨fQ(R1,r+11,R1,r+12)⟩=1rE⟨f⟩ E⟨Q(R121,R122)⟩+1r∑ℓ=2rE⟨fQ(R1ℓ1,R1ℓ2)⟩(99)\mathbb{E}\left\langle fQ(R_{1,r+1}^1,R_{1,r+1}^2)\right\rangle = \frac{1}{r}\mathbb{E}\langle f\rangle\,\mathbb{E}\left\langle Q(R_{12}^1,R_{12}^2)\right\rangle + \frac{1}{r}\sum_{\ell=2}^{r}\mathbb{E}\left\langle fQ(R_{1\ell}^1,R_{1\ell}^2)\right\rangle \tag*{(99)}

for every bounded measurable ff of the arrays of its first rr samples. Moreover:

(i) The off-diagonal entries of R1R^1 and R2R^2 are nonnegative almost surely.

(ii) The off-diagonal pair of arrays is a limit in finite-dimensional distribution of Ruelle probability cascade arrays with paired nondecreasing levels (ql,pl)∈[0,1]2(q_l,p_l)\in[0,1]^2.

Part (ii) is the form, for a pair of Gram arrays, of Panchenko’s synchronization combined with his approximation by Ruelle probability cascades [ref-100]. Step 1 of the proof follows the proof of [ref-100], Theorem 3 in Section 3 there, Step 2 the positivity argument of [ref-100], Section 4, and Step 4 the approximation by cascades in [ref-100], Section 5. In Step 3 we use Mourrat’s quantitative form, Lemma 4.5(a). It applies to the Gram pairs at finite NN, before passing to the limit, so no representing measure of the limiting arrays needs to preserve their diagonal values.

Proof. We divide the proof into four steps.

Step 1. The identities in the limit. For continuous ff, finite-dimensional convergence and (98) give (99), since all arrays take values in [−1,1][-1,1]. Both sides of (99) are integrals of ff against finite signed measures on the arrays of rr samples, and they agree on continuous functions, so they agree for bounded measurable ff. Choosing (v1,v2)=(1,0)(v_{1},v_{2})=(1,0), (0,1)(0,1) and (1/2,1/2)(1/2,1/2) gives (97) with ψ(x)=xm\psi(x)=x^{m}, m≥1m\geq1, for the scalar arrays R1R^{1}, R2R^{2}, and S/2S/2, where S=R1+R2S=R^{1}+R^{2}, and for test functions ff of their off-diagonal entries. It also holds for constant ψ\psi, since the sum of its three coefficients is one. By Weierstrass approximation it holds for continuous ψ\psi on [−1,1][-1,1], and, again by equality of finite measures, for bounded measurable ψ\psi.

Step 2. Positivity. Each of the arrays R1R^{1}, R2R^{2}, and S/2S/2 is a Gram–de Finetti array with diagonal entries at most one, as a limit of such arrays. By Lemma 4.3, its off-diagonal part is the overlap array of a random measure on a unit ball, and this measure satisfies the Ghirlanda–Guerra identities by Step 1. Lemma 4.4(a) gives (i), and it shows that the law of the off-diagonal array S/2S/2 is determined by the law of S12/2S_{12}/2.

Step 3. Monotone coupling. Let ϵ>0\epsilon>0, and let δ>0\delta>0 be given by Lemma 4.5(a). The hypothesis of Lemma 4.5(a) involves finitely many tuples (r,j1,j2,j3)(r,j_{1},j_{2},j_{3}), and by (98) it holds for GN\mathfrak{G}_{N} for all sufficiently large NN. Hence, for φ∈C∞(R2)\varphi\in C^{\infty}(\mathbb{R}^{2}),

∣E⟨φ(R121,N,R122,N)⟩N−Eφ(F1,N−1(V),F2,N−1(V))∣≤ϵ(∥φ∥∞+∥∇φ∥∞),\left|\mathbb{E}\left\langle\varphi\left(R_{12}^{1,N},R_{12}^{2,N}\right)\right\rangle_{N}-\mathbb{E}\varphi\left(F_{1,N}^{-1}(V),F_{2,N}^{-1}(V)\right)\right| \leq\epsilon\left(\lVert\varphi\rVert_{\infty}+\lVert\nabla\varphi\rVert_{\infty}\right),

where Fa,N−1F_{a,N}^{-1} is defined from the law of R12a,NR_{12}^{a,N}. As N→∞N\to\infty, the first term converges to E⟨φ(R121,R122)⟩\mathbb{E}\langle\varphi(R_{12}^{1},R_{12}^{2})\rangle, and by Lemma 4.5(b) the second converges to Eφ(F1−1(V),F2−1(V))\mathbb{E}\varphi(F_{1}^{-1}(V),F_{2}^{-1}(V)). Since ϵ\epsilon is arbitrary, the law of (R121,R122)(R_{12}^{1},R_{12}^{2}) is the monotone coupling of its marginals. The maps F1−1F_{1}^{-1} and F2−1F_{2}^{-1} are nondecreasing, so the image of (0,1)(0,1) under v↦(F1−1(v),F2−1(v))v\mapsto(F_{1}^{-1}(v),F_{2}^{-1}(v)) is totally ordered for the coordinatewise order, and so is its closure. Thus the support CC of the law of (R121,R122)(R_{12}^{1},R_{12}^{2}), a compact subset of [0,1]2[0,1]^{2} by (i), is totally ordered. Two points (x,y)(x,y) and (x′,y′)(x',y') of CC with x+y≤x′+y′x+y\leq x'+y' therefore satisfy

0≤x′−x≤(x′+y′)−(x+y),0≤y′−y≤(x′+y′)−(x+y).0\leq x'-x\leq(x'+y')-(x+y),\qquad0\leq y'-y\leq(x'+y')-(x+y).

In particular the sum determines both coordinates: on the compact set {x+y:(x,y)∈C}\{x+y:(x,y)\in C\}, which is the support of the law of S12S_{12}, there are nondecreasing 1-Lipschitz functions Q1,Q2Q_{1},Q_{2} with (x,y)=(Q1(x+y),Q2(x+y))(x,y)=(Q_{1}(x+y),Q_{2}(x+y)) on CC. We extend them to nondecreasing 1-Lipschitz functions on R\mathbb{R}, by linear interpolation across the gaps of this compact set and by constants outside its hull. Almost surely R121=Q1(S12)R_{12}^{1}=Q_{1}(S_{12}) and R122=Q2(S12)R_{12}^{2}=Q_{2}(S_{12}). By exchangeability and a countable intersection, the same holds for every off-diagonal pair of indices.

Step 4. Approximation by cascades. Let F−1F^{-1} be the function Fγ−1F_{\gamma}^{-1} for the law ζS\zeta_{S} of S12/2S_{12}/2, and let ζh\zeta_{h} be the law of F−1(⌈hV⌉/h)F^{-1}(\lceil hV\rceil/h). It has finite support contained in the support of ζS\zeta_{S}, and it converges weakly to ζS\zeta_{S} as h→∞h\to\infty, since 0≤⌈hV⌉/h−V<1/h0\leq\lceil hV\rceil/h-V<1/h almost surely and F−1F^{-1} is continuous at almost every point of (0,1)(0,1). By Steps 1 and 2 and Lemma 4.4(b), the cascade arrays with overlap laws ζh\zeta_{h}, with levels s0h<⋯<skhs_{0}^{h}<\cdots<s_{k}^{h} in [0,1][0,1], converge in finite-dimensional distribution to (Sℓℓ′/2)ℓ≠ℓ′(S_{\ell\ell'}/2)_{\ell\ne\ell'}. Apply the maps s↦(Q1(2s),Q2(2s))s\mapsto(Q_{1}(2s),Q_{2}(2s)) to their entries. The resulting arrays are the cascade arrays with the paired levels (ql,pl)=(Q1(2slh),Q2(2slh))(q_{l},p_{l})=(Q_{1}(2s_{l}^{h}),Q_{2}(2s_{l}^{h})), which are nondecreasing in ll and lie in [0,1]2[0,1]^{2}, because 2slh2s_{l}^{h} lies in the support of the law of S12S_{12}. By continuity of Q1Q_{1} and Q2Q_{2} they converge in finite-dimensional distribution to (Q1(Sℓℓ′),Q2(Sℓℓ′))ℓ≠ℓ′=(Rℓℓ′1,Rℓℓ′2)ℓ≠ℓ′(Q_{1}(S_{\ell\ell'}),Q_{2}(S_{\ell\ell'}))_{\ell\ne\ell'}=(R_{\ell\ell'}^{1},R_{\ell\ell'}^{2})_{\ell\ne\ell'}. This proves (ii). □\square

Perturbations, continuity, and the precision identity

The cavity computation needs three properties of the Gibbs measures, which we obtain by adding a small perturbation to each Hamiltonian. First, the Ghirlanda–Guerra identities (99) for the spin and response overlaps must hold in the limit, so that Theorem 4.6 applies. Second, the per-configuration observables bNb_{N} and XN(σ,σ)X_{N}(\sigma,\sigma) must concentrate, so that they can be replaced by constants in the limit. Third, the perturbations must be compatible with adding coordinates and rows. In this section we define the perturbation and state its properties (Lemma 4.7), a continuity lemma for the functionals that appear in the cavity computation (Lemma 4.8), and an identity for the mean precision (Lemma 4.9). We then explain how these lemmas are combined with Theorem 4.6. The three lemmas are proved in §4.8. Every assertion in this section concerns a fixed number nn of added coordinates and a fixed cutoff Λ\Lambda for their squared norm; removing the cutoff is a separate step.

The perturbation. Consider a dd-dimensional system, d∈{N,N′}d\in\{N,N'\}, with configurations ρ∈Sd\rho\in S_{d} and rows ga∈Rdg_{a}\in\mathbb{R}^{d} indexed by a finite set A\mathcal{A} (in the N′N'-dimensional system, A={1,…,M}\mathcal{A}=\{1,\ldots,M\} for the Gibbs measure ⟨⋅⟩N′\langle\cdot\rangle_{N'}, and A\mathcal{A} also contains the π\pi added rows when these are present), and write xa(ρ)=ga⋅ρ/dx_{a}(\rho)=g_{a}\cdot\rho/\sqrt{d}. The perturbation has two parts.

Scalar part. Put

τ1(x)=xU′(x)−U′′(x),τ2(x)=U′(x)2,ζd=d−1/4,\tau_{1}(x)=xU'(x)-U''(x),\qquad\tau_{2}(x)=U'(x)^{2},\qquad\zeta_{d}=d^{-1/4},

and add ζd∑a∈A(u1τ1+u2τ2)(xa(ρ))\zeta_{d}\sum_{a\in\mathcal{A}}(u_{1}\tau_{1}+u_{2}\tau_{2})(x_{a}(\rho)) to the Hamiltonian, where u1,u2∈[1,2]u_{1},u_{2}\in[1,2] are parameters. Equivalently, the activation UU is replaced by

Uξ~=U+ξ~(u1τ1+u2τ2),ξ~=ζd.(100)U^{\widetilde{\xi}}=U+\widetilde{\xi}(u_{1}\tau_{1}+u_{2}\tau_{2}),\qquad\widetilde{\xi}=\zeta_{d}. \tag*{(100)}

By (87), ∣τi(x)∣≤Cτ(1+x2)|\tau_i(x)| \le C_\tau(1+x^2) for a constant CτC_\tau depending only on LL, and the first and second derivatives of τi\tau_i grow at most linearly. The derivative of the Hamiltonian in uiu_i is ξ~d∑aτi(xa(ρ))\widetilde{\xi}_d \sum_a \tau_i(x_a(\rho)). For the reference fields these sums are exactly

∑a≤Mτ2(Sa)=N′XN(σ,σ),∑a≤Mτ1(Sa)=NAN−N′BN=N(bN−1)−nBN,(101)\sum_{a \le M}\tau_2(S_a)=N'X_N(\sigma,\sigma), \qquad\sum_{a \le M}\tau_1(S_a)=NA_N-N'B_N=N(b_N-1)-nB_N, \tag*{(101)}

so concentration of these derivatives gives concentration of XN(σ,σ)X_N(\sigma,\sigma) and bNb_N.

Gaussian part. We first bound the response vectors without changing them on a typical event. Fix a smooth function ψ0:[0,∞)→(0,1]\psi_0:[0,\infty)\to(0,1] with ψ0=1\psi_0=1 on [0,3/4][0,3/4], ψ0(t)≤1/t\psi_0(t)\le1/t for t≥1t\ge1, and ψ0(t)=1/t\psi_0(t)=1/t for t≥2t\ge2, and define ϖ(y)=ψ0(∥y∥)y\varpi(y)=\psi_0(\lVert y\rVert)y on every Euclidean space. Then ϖ\varpi equals the identity on the ball of radius 3/43/4, takes values in the closed unit ball, and its first and second derivatives are bounded by a constant that does not depend on the dimension. Put

Ψd(ρ)=(U′(xa(ρ)))a∈ACXd,\Psi_d(\rho)=\frac{\left(U'(x_a(\rho))\right)_{a\in\mathcal A}}{\sqrt{C_Xd}},

where CXC_X is a constant depending only on α\alpha and LL. We choose CXC_X so large that, on the event ΩN\Omega_N, the vector Ψd(ρ)\Psi_d(\rho) with A={1,…,M}\mathcal A=\{1,\ldots,M\} has norm at most 1/21/\sqrt{2} for every ρ∈Sd\rho\in S_d in each of the three systems. We also require the same bound in the N′N'-dimensional system when A\mathcal A contains the π\pi added rows as well, on the event

ΩN′=ΩN∩{π≤N,∥G~∥op≤C1N},\Omega'_N=\Omega_N\cap\{\pi\le N,\lVert\widetilde{\mathbf G}\rVert_{\mathrm{op}}\le C_1\sqrt{N}\},

where G~\widetilde{\mathbf G} is the π×N′\pi\times N' matrix with rows g~i\widetilde{g}_i. This is possible because ∥Ψd(ρ)∥2≤2L2(#A+∑axa(ρ)2)/(CXd)\lVert\Psi_d(\rho)\rVert^2\le2L^2(\#\mathcal A+\sum_a x_a(\rho)^2)/(C_Xd) by (87), on ΩN\Omega_N the fields of the three systems ((a)–(c) of §4.1, with the fields Sa+ηaS_a+\eta_a for the NN-dimensional system) satisfy ∑a≤Mxa(ρ)2≤CN\sum_{a\le M}x_a(\rho)^2\le CN, and the fields of the added rows satisfy ∑i≤π(g~i⋅ρ)2/N′≤∥G~∥op2\sum_{i\le\pi}(\widetilde{g}_i\cdot\rho)^2/N'\le\lVert\widetilde{\mathbf G}\rVert_{\mathrm{op}}^2. By Lemma 4.2 and the Poisson tail of π\pi, P(ΩN′c)≤Cne−N/2\mathbb P(\Omega_N'{}^{c})\le C_ne^{-N/2}. Set

Ψ^d=ϖ∘Ψd,X^d(ρ,ρ′)=Ψ^d(ρ)⋅Ψ^d(ρ′),R(ρ,ρ′)=ρ⋅ρ′d.\widehat{\Psi}_d=\varpi\circ\Psi_d,\qquad \widehat{X}_d(\rho,\rho')=\widehat{\Psi}_d(\rho)\cdot\widehat{\Psi}_d(\rho'),\qquad R(\rho,\rho')=\frac{\rho\cdot\rho'}{d}.

Since 1/2<3/41/\sqrt{2}<3/4, we have Ψ^d=Ψd\widehat{\Psi}_d=\Psi_d on ΩN\Omega_N when A={1,…,M}\mathcal A=\{1,\ldots,M\}. Hence, on ΩN\Omega_N, X^d\widehat{X}_d is the response overlap d−1∑aU′(xa(ρ))U′(xa(ρ′))d^{-1}\sum_a U'(x_a(\rho))U'(x_a(\rho')) of the dd-dimensional system divided by CXC_X.

Enumerate the pairs v=(v1,v2)v=(v_1,v_2) of nonnegative rationals with v1+v2≤1v_1+v_2\le1 as v(1),v(2),…v^{(1)},v^{(2)},\ldots, write j(v)j(v) for the index of vv, and put cv,m=2−j(v)−mc_{v,m}=2^{-j(v)-m} for integers m≥1m\ge1. Let hv,mh_{v,m} be independent (conditionally on the rows) centered Gaussian processes on SdS_d with covariance

Qv,m(ρ,ρ′)=(v1R(ρ,ρ′)+v2X^d(ρ,ρ′))m.Q_{v,m}(\rho,\rho')=\left(v_1R(\rho,\rho')+v_2\widehat{X}_d(\rho,\rho')\right)^m.

These processes exist. Indeed, by the binomial theorem

Qv,m=∑i=0m(mi)v1iv2m−iRiX^dm−i,Q_{v,m}=\sum_{i=0}^{m}\binom{m}{i}v_{1}^{i}v_{2}^{m-i}R^{i}\widehat{X}_{d}^{m-i},

and R(ρ,ρ′)iX^d(ρ,ρ′)m−iR(\rho,\rho')^{i}\widehat{X}_{d}(\rho,\rho')^{m-i} is the inner product of the tensors (ρ/d)⊗i⊗Ψ^d(ρ)⊗(m−i)(\rho/\sqrt{d})^{\otimes i}\otimes\widehat{\Psi}_{d}(\rho)^{\otimes(m-i)} and (ρ′/d)⊗i⊗Ψ^d(ρ′)⊗(m−i)(\rho'/\sqrt{d})^{\otimes i}\otimes\widehat{\Psi}_{d}(\rho')^{\otimes(m-i)}. Hence

hv,m(ρ)=⟨Tv,m,Fv,m(ρ)⟩,Fv,m(ρ)=((mi)1/2v1i/2v2(m−i)/2(ρ/d)⊗i⊗Ψ^d(ρ)⊗(m−i))0≤i≤m,(102)\begin{aligned} h_{v,m}(\rho)&=\langle T_{v,m},F_{v,m}(\rho)\rangle,\\ F_{v,m}(\rho)&=\left(\binom{m}{i}^{1/2}v_{1}^{i/2}v_{2}^{(m-i)/2}(\rho/\sqrt{d})^{\otimes i}\otimes\widehat{\Psi}_{d}(\rho)^{\otimes(m-i)}\right)_{0\le i\le m}, \tag*{(102)} \end{aligned}

where Tv,m=(Tv,m,i)0≤i≤mT_{v,m}=(T_{v,m,i})_{0\le i\le m} is a family of independent standard Gaussian tensors of the matching shapes, has covariance Fv,m(ρ)⋅Fv,m(ρ′)=Qv,m(ρ,ρ′)F_{v,m}(\rho)\cdot F_{v,m}(\rho')=Q_{v,m}(\rho,\rho'). Every diagonal covariance satisfies ∥Fv,m(ρ)∥2=Qv,m(ρ,ρ)≤1\lVert F_{v,m}(\rho)\rVert^{2}=Q_{v,m}(\rho,\rho)\le1, since R(ρ,ρ)=1R(\rho,\rho)=1, X^d(ρ,ρ)≤1\widehat{X}_{d}(\rho,\rho)\le1 and v1+v2≤1v_{1}+v_{2}\le1. With sd=d1/3s_{d}=d^{1/3} and parameters uv,m∈[1,2]u_{v,m}\in[1,2] we add

sd∑v,mcv,muv,mhv,m(ρ)s_{d}\sum_{v,m}c_{v,m}u_{v,m}h_{v,m}(\rho)

to the Hamiltonian. The series converges almost surely, uniformly in ρ\rho (see eq:4.41 below). The Gaussian part is modeled on Panchenko’s perturbation for the multi-species Sherrington–Kirkpatrick model [ref-100], which has weights of the same form 2−j−m2^{-j-m}, parameters uniform on [1,2][1,2], and a strength NγN^{\gamma} with 1/4<γ<1/21/4<\gamma<1/2 (here sd=d1/3s_{d}=d^{1/3}). In [ref-100] the covariance is a power of a weighted sum of the overlaps within the species. Here it is a power of a weighted sum of the spin overlap and the overlap of the response vectors, which the cutoff ϖ\varpi keeps in the unit ball.

Parameters and the three systems. The parameters u1,u2u_{1},u_{2} and uv,mu_{v,m} are independent and uniform on [1,2][1,2], and they take the same values in the NN-dimensional, the N′N'-dimensional, and the reference system. The first two systems use the construction above with d=Nd=N and d=N′d=N'. The reference system uses the scalar part ξ~N′∑a≤M(u1τ1+u2τ2)(Sa(σ))\widetilde{\xi}_{N'}\sum_{a\le M}(u_{1}\tau_{1}+u_{2}\tau_{2})(S_{a}(\sigma)), and as its Gaussian part it uses, by definition, the Gaussian part of the N′N'-dimensional system with the rows a≤Ma\le M, evaluated at ρ(σ,0)=(N′/N σ,0)\rho(\sigma,0)=(\sqrt{N'/N}\,\sigma,0). This common restriction is what makes the comparison in Lemma 4.7(iv) possible. We have R(ρ(σ,0),ρ(σ′,0))=RN(σ,σ′)R(\rho(\sigma,0),\rho(\sigma',0))=R_{N}(\sigma,\sigma') and xa(ρ(σ,0))=N′/N Sa(σ)x_{a}(\rho(\sigma,0))=\sqrt{N'/N}\,S_{a}(\sigma). Thus the response vector of the reference perturbation is built from U′(N′/N Sa)U'(\sqrt{N'/N}\,S_{a}) rather than U′(Sa)U'(S_{a}). Since ∣U′(N′/N S)−U′(S)∣≤L∣S∣n/N\lvert U'(\sqrt{N'/N}\,S)-U'(S)\rvert\le L\lvert S\rvert n/N, the Cauchy–Schwarz inequality shows that on ΩN\Omega_{N}

∣X^N′(ρ(σ,0),ρ(σ′,0))−XN(σ,σ′)CX∣≤CnN(σ,σ′∈SN).(103)\left\lvert\widehat{X}_{N'}(\rho(\sigma,0),\rho(\sigma',0))-\frac{X_{N}(\sigma,\sigma')}{C_{X}}\right\rvert\le\frac{Cn}{N}\qquad(\sigma,\sigma'\in S_{N}). \tag*{(103)}

Finally, the reference system is tilted by exp⁡(nBN(σ)/2)\exp(nB_{N}(\sigma)/2). We write ⟨⋅⟩ref\langle\cdot\rangle_{\mathrm{ref}} for the Gibbs measure of the perturbed and tilted reference system. The tilt will cancel between the two cavity terms in §4.4.

Invariance. The Gaussian tensors are invariant in law under orthogonal transformations of their spin slots, and Ψd(Oρ)\Psi_d(O\rho) computed with the rows gag_a equals Ψd(ρ)\Psi_d(\rho) computed with the rows OTgaO^{\mathsf T}g_a. Since the rows are standard Gaussian vectors, the perturbed N′N'-dimensional system is invariant in law under a simultaneous orthogonal transformation of the configurations and of the rows. In particular, for every orthogonal OO of RN′\mathbb{R}^{N'} and every bounded measurable FF,

E⟨F(Oρ1,…,Oρr)⟩N′=E⟨F(ρ1,…,ρr)⟩N′.(104)\mathbb{E}\langle F(O\rho^1,\ldots,O\rho^r)\rangle_{N'} = \mathbb{E}\langle F(\rho^1,\ldots,\rho^r)\rangle_{N'}. \tag*{(104)}

Properties of the perturbation. For a system with Gibbs measure ⟨⋅⟩\langle\cdot\rangle and a Gaussian component with covariance Q=Qv,mQ=Q_{v,m}, write Qℓℓ′=Q(ρℓ,ρℓ′)Q_{\ell\ell'}=Q(\rho^\ell,\rho^{\ell'}). For a bounded measurable function ff of the arrays (R(ρℓ,ρℓ′),X^d(ρℓ,ρℓ′))ℓ,ℓ′≤r(R(\rho^\ell,\rho^{\ell'}),\widehat{X}_d(\rho^\ell,\rho^{\ell'}))_{\ell,\ell'\le r} of rr replicas, the defect in the identity (99) is ∣D^r(f,Q)∣|\widehat{\mathscr{D}}_r(f,Q)|, where D^r\widehat{\mathscr{D}}_r is the discrepancy (96) of the pair (R,X^d)(R,\widehat{X}_d) with (v1,v2)=v(v_1,v_2)=v; its terms Q(Rℓℓ′1,Rℓℓ′2)Q(R_{\ell\ell'}^1,R_{\ell\ell'}^2) are then the values Qℓℓ′Q_{\ell\ell'}. For the reference system the arrays are those of the points ρ(σℓ,0)\rho(\sigma^\ell,0).

Lemma 4.7. Fix nn and Λ\Lambda. In the following statements o(1)o(1) denotes a quantity that tends to zero as N→∞N\to\infty, uniformly in the perturbation parameters and in M∈WNM\in W_N, except where a statement specifies otherwise: in (ii) the defects tend to zero after averaging over the parameters, and in (iii) uniformly on the sets GN(M)\mathcal{G}_N(M).

(i) For the NN-dimensional system with M∼Poi⁡(αN)M\sim\operatorname{Poi}(\alpha N) rows, the perturbation changes the expected log partition function by o(N)o(N).

(ii) For the reference system let Ai(σ)=N′−1∑a≤Mτi(Sa(σ))A_i(\sigma)=N'^{-1}\sum_{a\le M}\tau_i(S_a(\sigma)), i=1,2i=1,2. After averaging over the perturbation parameters, E⟨∣Ai−E⟨Ai⟩ref∣⟩ref→0\mathbb{E}\langle|A_i-\mathbb{E}\langle A_i\rangle_{\mathrm{ref}}|\rangle_{\mathrm{ref}}\to0. By (101), the observables bNb_N and XN(σ,σ)X_N(\sigma,\sigma) therefore concentrate under ⟨⋅⟩ref\langle\cdot\rangle_{\mathrm{ref}} about their expectations, and so does the diagonal entry X^N′(ρ(σ,0),ρ(σ,0))\widehat{X}_{N'}(\rho(\sigma,0),\rho(\sigma,0)) of its perturbation.

(iii) There are sets GN=GN(M)\mathcal{G}_N=\mathcal{G}_N(M) of perturbation parameters with inf⁡M∈WNP(GN(M))→1\inf_{M\in W_N}\mathbb{P}(\mathcal{G}_N(M))\to1 such that

sup⁡M∈WN sup⁡GN(M) sup⁡∥f∥∞≤1∣D^r(f,Qv,m)∣→0\sup_{M\in W_N}\ \sup_{\mathcal{G}_N(M)}\ \sup_{\lVert f\rVert_\infty\le1} \left|\widehat{\mathscr{D}}_r(f,Q_{v,m})\right|\to0

for the reference system and every fixed rr and (v,m)(v,m), and for the N′N'-dimensional system with the Gibbs measure ⟨⋅⟩N′\langle\cdot\rangle_{N'} and every fixed rr and (v,m)(v,m) with v2=0v_2=0, and such that the concentration defects of (ii), without the average over the parameters, also tend to zero uniformly on GN(M)\mathcal{G}_N(M) and in M∈WNM\in W_N. The sets GN(M)\mathcal{G}_N(M) do not depend on Λ\Lambda.

(iv) Comparison with the reference perturbation. (a) In the partition function of the N′N'-dimensional system with MM rows, restricted to ε∈BΛ\varepsilon\in\mathbb{B}_\Lambda, replacing the Gaussian part of the perturbation at ρ(σ,ε)\rho(\sigma,\varepsilon) by its value at ρ(σ,0)\rho(\sigma,0) changes the expected logarithm by o(1)o(1), and changes the Gibbs expectations of bounded measurable functions of finitely many spin overlaps and cavity coordinates by o(1)o(1) times their supremum norm. (b) In the representation of the NN-dimensional system through the fields Sa+ηaS_a+\eta_a of (92), replacing its Gaussian perturbation by that of the reference system, and the strengths sN,ζNs_N,\zeta_N by sN′,ζN′s_{N'},\zeta_{N'}, changes the expected log partition function by o(1)o(1).

(v) Rows. Adding the π\pi further rows to the N′N'-dimensional system, including the corresponding change of its perturbation, increases the expected log partition function by at least

Elog⁡⟨exp⁡{∑i≤πU(yi(ρ))}⟩N′−o(1),\mathbb{E}\log\left\langle\exp\left\{\sum_{i\leq\pi}U(y_i(\rho))\right\}\right\rangle_{N'}-o(1),

where yi(ρ)=g~i⋅ρ/N′y_i(\rho)=\widetilde{g}_i\cdot\rho/\sqrt{N'} are the fields of the added rows. Moreover, for N≥N0(n)N\geq N_0(n) the difference between the expected log partition functions of the N′N'-dimensional system with M+πM+\pi rows and of the NN-dimensional system with MM rows is at least −Cn-C_n, uniformly in the perturbation parameters and in M∈WNM\in W_N, and for every NN and every number MM of rows it is at most Cn(N+M)C_n(N+M) in absolute value.

The second lemma provides continuity of the quantities in the cavity computation under convergence of overlap arrays. We apply it both to the Gibbs measures of §4.4 and to finite Ruelle probability cascades, so we state it for a random probability measure G\mathfrak{G} on a measurable space E\mathrm{E}. We write ⟨⋅⟩\langle\cdot\rangle for the average over independent samples ρ1,ρ2,…\rho^1,\rho^2,\ldots from G\mathfrak{G} (replicas) and E⟨⋅⟩\mathbb{E}\langle\cdot\rangle for the expectation over G\mathfrak{G} and the replicas. Let RR and XX be random positive semidefinite kernels on E\mathrm{E} with R(ρ,ρ)=1R(\rho,\rho)=1, let bb be a random function on E\mathrm{E}, and let z:E→Rnz:\mathrm{E}\to\mathbb{R}^n and y,y1,y2,…:E→Ry,y_1,y_2,\ldots:\mathrm{E}\to\mathbb{R} be random fields, all defined on one probability space with G\mathfrak{G} and jointly measurable. For rr replicas write Ar=(R(ρℓ,ρℓ′),X(ρℓ,ρℓ′),b(ρℓ))ℓ,ℓ′≤rA_r=(R(\rho^\ell,\rho^{\ell'}),X(\rho^\ell,\rho^{\ell'}),b(\rho^\ell))_{\ell,\ell'\leq r} for their arrays. For such an array aa let Na\mathrm{N}_a be the centered Gaussian law of vectors (zℓ,yℓ,y1ℓ,y2ℓ,…)ℓ≤r(z^\ell,y^\ell,y_1^\ell,y_2^\ell,\ldots)_{\ell\leq r} under which the families z,y,y1,y2,…z,y,y_1,y_2,\ldots are independent and

Eziℓzjℓ′=δijXℓℓ′,Eyℓyℓ′=nRℓℓ′Xℓℓ′,Eyiℓyiℓ′=Rℓℓ′;\mathbb{E}z_i^\ell z_j^{\ell'}=\delta_{ij}X_{\ell\ell'},\qquad \mathbb{E}y^\ell y^{\ell'}=nR_{\ell\ell'}X_{\ell\ell'},\qquad \mathbb{E}y_i^\ell y_i^{\ell'}=R_{\ell\ell'};

these matrices are positive semidefinite by the Schur product theorem. We assume:

(G) For every rr and every bounded measurable function Φ\Phi of ArA_r and of the values at ρ1,…,ρr\rho^1,\ldots,\rho^r of z,yz,y, and finitely many yiy_i, E⟨Φ⟩=E⟨Φ‾(Ar)⟩\mathbb{E}\langle\Phi\rangle=\mathbb{E}\langle\overline{\Phi}(A_r)\rangle, where Φ‾(a)\overline{\Phi}(a) is the integral of Φ(a,⋅)\Phi(a,\cdot) against Na\mathrm{N}_a.

This holds, for instance, if E=Sd\mathrm{E}=S_d and, conditionally on (G,R,X,b)(\mathfrak{G},R,X,b), the fields are independent centered Gaussian processes with covariances δijX\delta_{ij}X, nRXnRX, and RR, and are independent of the replicas. This is the situation in §4.4. The numerator, or cavity partition function on BΛ\mathbb{B}_\Lambda, and the denominator are

NΛ=⟨∫BΛνn(dε)eε⋅z(ρ)−(∥ε∥2−n)(b(ρ)−1)/2⟩,D=⟨ey(ρ)⟩,\mathcal{N}_\Lambda = \left\langle \int_{\mathbb{B}_\Lambda} \nu_n(\mathrm{d}\varepsilon) e^{\varepsilon\cdot z(\rho)-(\lVert\varepsilon\rVert^2-n)(b(\rho)-1)/2} \right\rangle, \qquad D=\left\langle e^{y(\rho)}\right\rangle,

and the cavity measure is the random probability measure on E×BΛ\mathrm{E}\times\mathbb{B}_\Lambda with density proportional to the integrand of NΛ\mathcal{N}_\Lambda with respect to G⊗νn\mathfrak{G}\otimes\nu_n.

Lemma 4.8. Consider a sequence of such random measures satisfying (G) for which, for every c>0c>0, the quantities

E⟨exp⁡(cX(ρ,ρ)+c∣b(ρ)∣)⟩\mathbb{E}\left\langle\exp\left(cX(\rho,\rho)+c|b(\rho)|\right)\right\rangle

are bounded along the sequence from some index on, and suppose that the finite-dimensional distributions of the arrays converge. Then the following quantities converge, and their limits depend only on the limiting law of the arrays:

  1. Elog⁡NΛ\mathbb{E}\log\mathcal{N}_{\Lambda} and Elog⁡D\mathbb{E}\log\mathcal{D};

  2. expectations under the cavity measure of bounded continuous functions of the arrays and the cavity coordinates of finitely many replicas;

  3. the expected logarithmic increment Elog⁡⟨exp⁡{∑i≤πU(yi(ρ))}⟩\mathbb{E}\log\left\langle\exp\left\{\sum_{i\leq\pi}U(y_i(\rho))\right\}\right\rangle, for π∼Poi⁡(αn)\pi\sim\operatorname{Poi}(\alpha n) independent of the other variables, whenever UU is continuous and satisfies (88).

In particular, if E⟨∣b−bˉN∣⟩→0\mathbb{E}\langle|b-\bar b_N|\rangle\to0 for constants bˉN→b∞\bar b_N\to b_\infty, then bb may be replaced by the constant b∞b_\infty in these limits.

Lemma 4.8 is the analogue for the perceptron of the continuity of the cavity functionals in the law of the overlap array, proved for mixed pp-spin models in [ref-99] and for spherical models in [ref-37]. The difference is that the functionals here also depend on the response overlap and on the precision bb, and that part (iii) treats the row term, whose activation may be unbounded below.

The third lemma is an identity for the mean of bNb_N under the reference system, obtained by Gaussian integration by parts in the rows gag_a; it turns SaU′(Sa)S_aU'(S_a) into U′′(Sa)+U′(Sa)2U''(S_a)+U'(S_a)^2 minus a two-replica term. In the limit it shows that the smallest precision of the Gaussian integral over the added coordinates is at least one (Step 3 of §4.7).

Lemma 4.9. Fix nn. For two replicas σ1,σ2\sigma^1,\sigma^2 of the reference system, uniformly in the perturbation parameters and in M∈WNM\in W_N,

E⟨bN⟩ref=1+E⟨XN(σ,σ)⟩ref−E⟨RN(σ1,σ2)XN(σ1,σ2)⟩ref+o(1).\mathbb{E}\langle b_N\rangle_{\mathrm{ref}} = 1+\mathbb{E}\langle X_N(\sigma,\sigma)\rangle_{\mathrm{ref}} -\mathbb{E}\left\langle R_N(\sigma^1,\sigma^2)X_N(\sigma^1,\sigma^2)\right\rangle_{\mathrm{ref}} +o(1).

Finite cascades as random measures. Lemma 4.8 is also applied to finite Ruelle probability cascades carrying prescribed diagonal values. Take a cascade as in §4.2, with depth kk, weights (vβ)(v_\beta), and paired levels 0≤q0≤⋯≤qk≤10\leq q_0\leq\cdots\leq q_k\leq1 and 0≤p0≤⋯≤pk0\leq p_0\leq\cdots\leq p_k, and let pˉ≥pk\bar p\geq p_k and bˉ∈R\bar b\in\mathbb{R}. Independently of (vβ)(v_\beta), let YpY^p, YcY^c, and Yq,1,Yq,2,…Y^{q,1},Y^{q,2},\ldots be independent Gaussian cascade fields (§3.2), with values in Rn\mathbb{R}^n, R\mathbb{R}, and R\mathbb{R}, and with levels pp, c=(nqlpl)l≤kc=(nq_lp_l)_{l\leq k}, and qq, respectively; the levels clc_l are nonnegative and nondecreasing. The cascade realization with diagonal values (1,pˉ,bˉ)(1,\bar p,\bar b) is the random measure G=∑βvβδβ⊗νn⊗ν1⊗ν1⊗N\mathfrak{G}=\sum_\beta v_\beta\delta_\beta\otimes\nu_n\otimes\nu_1\otimes\nu_1^{\otimes\mathbb{N}} on E=N∗k×Rn×R×RNE=\mathbb{N}_*^k\times\mathbb{R}^n\times\mathbb{R}\times\mathbb{R}^{\mathbb{N}}, with b≡bˉb\equiv\bar b, with kernels RR and XX equal to qβ∧β′q_{\beta\wedge\beta'} and pβ∧β′p_{\beta\wedge\beta'} at distinct points with leaves β,β′\beta,\beta' and to 11 and pˉ\bar p at equal points, and with the fields

z(ρ)=Yβp+pˉ−pk x,y(ρ)=Yβc+n(pˉ−qkpk) x′,yi(ρ)=Yβq,i+1−qk xi′,z(\rho)=Y_\beta^p+\sqrt{\bar p-p_k}\,x,\qquad y(\rho)=Y_\beta^c+\sqrt{n(\bar p-q_kp_k)}\,x',\qquad y_i(\rho)=Y_\beta^{q,i}+\sqrt{1-q_k}\,x_i',

at ρ=(β,x,x′,x′′)\rho=(\beta,x,x',x''). The within-state coordinates x,x′,x′′x,x',x'' carry the nonnegative differences between the diagonal values and the top levels. Two independent samples from G\mathfrak{G} are distinct points almost surely, so the arrays of the replicas are cascade arrays with paired levels (ql,pl)(q_l,p_l) off the diagonal, with diagonal values 11, pˉ\bar{p}, and bˉ\bar{b}. Given the cascade weights and the leaves of the replicas, the values of the fields at the replicas are centered Gaussian with the covariances in (G), because the cascade fields are independent of the weights and the within-state coordinates are independent standard Gaussian. This gives (G), and the moment condition of Lemma 4.8 holds because X(ρ,ρ)=pˉX(\rho,\rho)=\bar{p} and b=bˉ\mathsf{b}=\bar{b} are constant.

How the lemmas are combined with the overlap theorem. Fix M∈WNM\in W_N and parameters in the set GN(M)\mathcal{G}_N(M) of Lemma 4.7(iii) (in §4.7, M=MNM=M_N depends on NN), and pass to a subsequence along which the finite-dimensional distributions of the arrays (RN,XN,bN)(R_N,X_N,b_N) under E⟨⋅⟩ref\mathbb{E}\langle\cdot\rangle_{\mathrm{ref}} converge; they are tight by (94) and (95). The reference system satisfies (G) with G=⟨⋅⟩ref\mathfrak{G}=\langle\cdot\rangle_{\mathrm{ref}}, R=RNR=R_N, X=XNX=X_N, and b=bN\mathsf{b}=b_N, for fields that are, conditionally on the disorder, independent Gaussian processes with the covariances in (G); for instance

zi(σ)=∑a≤MU′(Sa(σ))N′gai,y(σ)=n∑a≤M∑i≤NσiU′(Sa(σ))NN′gai′,yj(σ)=∑i≤NσiNgji′′,\begin{aligned} z_i(\sigma)&=\sum_{a\leq M}\frac{U'(S_a(\sigma))}{\sqrt{N'}}g_{ai}, \qquad y(\sigma)=\sqrt{n}\sum_{a\leq M}\sum_{i\leq N}\frac{\sigma_iU'(S_a(\sigma))}{\sqrt{NN'}}g'_{ai},\\ y_j(\sigma)&=\sum_{i\leq N}\frac{\sigma_i}{\sqrt{N}}g''_{ji}, \end{aligned}

with independent standard Gaussian variables gai,gai′,gji′′g_{ai},g'_{ai},g''_{ji}, independent of the reference system. By (95), it satisfies the moment condition of Lemma 4.8. We apply Theorem 4.6 to the Gram pairs generated by σ/N\sigma/\sqrt{N} and Ψ^N′(ρ(σ,0))\widehat{\Psi}_{N'}(\rho(\sigma,0)). Their limiting arrays have R1=RR^1=R and R2=X/CXR^2=X/C_X, where (R,X)(R,X) is the limit of (RN,XN)(R_N,X_N), for three reasons.

(a) Both feature vectors lie in unit balls, and on ΩN\Omega_N the response Gram array differs from XN/CXX_N/C_X by O(n/N)O(n/N), by (103); moreover P(ΩNc)→0\mathbb{P}(\Omega_N^c)\to0.

(b) The diagonals concentrate by Lemma 4.7(ii)–(iii), so the limiting diagonal values are 11 and p∞/CXp_\infty/C_X, where

p∞=lim⁡E⟨XN(σ,σ)⟩refp_\infty=\lim\mathbb{E}\langle X_N(\sigma,\sigma)\rangle_{\mathrm{ref}}

along a further subsequence.

(c) Lemma 4.7(iii) gives vanishing discrepancies DNN(f,Q)\mathfrak{D}_N^N(f,Q) of (96), uniformly over bounded tests, for each fixed number of replicas and each Q=Qv,mQ=Q_{v,m}, that is, for v1+v2≤1v_1+v_2\leq1. Homogeneity of degree mm extends the conclusion to every nonnegative rational pair (v1,v2)(v_1,v_2), so (98) holds.

By Theorem 4.6(ii), after multiplying the second levels of its approximating cascades by CXC_X, the off-diagonal limit of (RN,XN)(R_N,X_N) is a limit in finite-dimensional distribution of Ruelle probability cascade arrays with paired nondecreasing levels (ql,pl)(q_l,p_l). The observable bNb_N concentrates about E⟨bN⟩ref\mathbb{E}\langle b_{N}\rangle_{\mathrm{ref}}, which converges along a further subsequence to a constant b∞b_{\infty}; by the last assertion of Lemma 4.8, bNb_{N} may be replaced by b∞b_{\infty} in all limits.

In the limit, 0≤R12≤10 \le R_{12} \le1 and 0≤X12≤p∞0 \le X_{12} \le p_{\infty} almost surely. Nonnegativity is Theorem 4.6(i), and RN≤1R_{N} \le1. For the last bound, the Cauchy–Schwarz inequality gives XN(σ1,σ2)≤(XN(σ1,σ1)+XN(σ2,σ2))/2X_{N}(\sigma^{1},\sigma^{2}) \le(X_{N}(\sigma^{1},\sigma^{1})+X_{N}(\sigma^{2},\sigma^{2}))/2, so E⟨(XN(σ1,σ2)−p∞)+⟩ref≤E⟨∣XN(σ,σ)−p∞∣⟩ref→0\mathbb{E}\langle(X_{N}(\sigma^{1},\sigma^{2})-p_{\infty})_{+}\rangle_{\mathrm{ref}} \le\mathbb{E}\langle|X_{N}(\sigma,\sigma)-p_{\infty}|\rangle_{\mathrm{ref}} \to0, and the left side converges to the corresponding expectation in the limit by uniform integrability. Projecting the levels of the approximating cascades onto [0,1]×[0,p∞][0,1]\times[0,p_{\infty}], which is monotone, continuous, and the identity on the values of the limiting arrays, we may assume 0≤ql≤10 \le q_{l} \le1 and 0≤pl≤p∞0 \le p_{l} \le p_{\infty}. The cascade realizations of these cascades with diagonal values (1,p∞,b∞)(1,p_{\infty},b_{\infty}) then have arrays converging to the limiting arrays of (RN,XN,bN)(R_{N},X_{N},b_{N}). We apply Lemma 4.8 to the sequence that alternates the reference systems along our subsequence (with bNb_{N} replaced by b∞b_{\infty}) and these cascade realizations, whose arrays converge to the same limit. We conclude that the limits along our subsequence of the quantities in Lemma 4.8(i)–(ii) for the reference system are the limits of their values on the cascade realizations.

For the spin overlaps of the N′N'-dimensional system we apply Theorem 4.6 to the Gram pairs generated by ρ/N′\rho/\sqrt{N'} and the zero vector. Then every QQ in its hypothesis is (v1x)m=v1mxm(v_{1}x)^{m}=v_{1}^{m}x^{m}, so the hypothesis reduces to the defects with v=(1,0)v=(1,0) of Lemma 4.7(iii). By parts (i)–(ii) of the theorem, the limiting spin overlap array is a limit in finite-dimensional distribution of Ruelle probability cascade arrays with nondecreasing levels in [0,1][0,1], and the overlap laws of these cascades converge to the limiting overlap law ζ~\widetilde{\zeta} of the N′N'-dimensional system.

The two cavity terms

We now compute the increment of the expected log partition function from dimension NN to dimension N′N' at fixed perturbation parameters and a fixed number M∈WNM\in W_{N} of rows. Write log⁡ZN′M+π\log Z_{N'}^{M+\pi}, log⁡ZN′M\log Z_{N'}^{M}, and log⁡ZNM\log Z_{N}^{M} for the log partition functions of the perturbed N′N'-dimensional system with all M+πM+\pi rows, of the same system with its first MM rows only, and of the perturbed NN-dimensional system. The increment splits as

Elog⁡ZN′M+π−Elog⁡ZNM=(Elog⁡ZN′M+π−Elog⁡ZN′M)⏟row term+(Elog⁡ZN′M−Elog⁡ZNM)⏟coordinate term.(105)\mathbb{E}\log Z_{N'}^{M+\pi}-\mathbb{E}\log Z_{N}^{M} = \underbrace{\left(\mathbb{E}\log Z_{N'}^{M+\pi}-\mathbb{E}\log Z_{N'}^{M}\right)}_{\text{row term}} + \underbrace{\left(\mathbb{E}\log Z_{N'}^{M}-\mathbb{E}\log Z_{N}^{M}\right)}_{\text{coordinate term}}. \tag*{(105)}

The row term. By Lemma 4.7(v), the row term is at least Elog⁡⟨exp⁡{∑i≤πU(yi(ρ))}⟩N′−o(1)\mathbb{E}\log\langle\exp\{\sum_{i\le\pi}U(y_{i}(\rho))\}\rangle_{N'}-o(1), where the fields yiy_{i} of the added rows are centered Gaussian with covariance ρ⋅ρ′/N′\rho\cdot\rho'/N' and are independent of ⟨⋅⟩N′\langle\cdot\rangle_{N'}. Thus ⟨⋅⟩N′\langle\cdot\rangle_{N'}, with R(ρ,ρ′)=ρ⋅ρ′/N′R(\rho,\rho')=\rho\cdot\rho'/N', X≡0X\equiv0, b≡0b\equiv0, and the fields yiy_{i} of the added rows, satisfies (G) and the moment condition of Lemma 4.8. By Lemma 4.8(iii), which applies by (88), the limit of this quantity along a subsequence along which the spin overlap array converges depends only on the limiting law of that array. For parameters in GN(M)\mathcal{G}_{N}(M) this array is a limit of Ruelle probability cascade arrays with nondecreasing levels in [0,1][0,1] whose overlap laws ζ~h\widetilde{\zeta}_{h} converge to the limiting overlap law ζ~\widetilde{\zeta} (§4.3). Consider the hh-th of these cascades, with levels 0≤q0≤⋯≤qk≤10 \le q_{0} \le\cdots\le q_{k} \le1 (we drop the index hh from its levels and masses), and its cascade realization with diagonal values (1,0,0)(1,0,0) and p≡0p \equiv0. Given π\pi, the fields yi(ρ)=Yβq,i+1−qkxi′′y_{i}(\rho)=Y_{\beta}^{q,i}+\sqrt{1-q_{k}}x_{i}^{\prime\prime} of different rows are independent, so

⟨exp⁡{∑i≤πU(yi(ρ))}⟩=∑βvβ∏i≤πE[eU(Yβq,i+1−qkG)∣Yq,i].\left\langle\exp\left\{\sum_{i\le\pi}U(y_i(\rho))\right\}\right\rangle = \sum_{\beta}v_{\beta}\prod_{i\le\pi} \mathbb{E}\left[\left.e^{U\left(Y_{\beta}^{q,i}+\sqrt{1-q_k}G\right)}\right|Y^{q,i}\right].

Apply Lemma 3.2 with the marks (eν(i))i≤π(e_{\nu}^{(i)})_{i\le\pi}, where eν(i)e_{\nu}^{(i)} are the Gaussian increments of Yq,iY^{q,i} at the node ν\nu (with variance q0q_{0} at the root and ql−ql−1q_{l}-q_{l-1} at depth ll), and with the leaf function ∑i≤πT1,1−qkU(Yβq,i)\sum_{i\le\pi}\mathcal{T}_{1,1-q_k}U(Y_{\beta}^{q,i}). The leaf function is bounded above and at least −C∑i(1+(Yβq,i)2)-C\sum_i(1+(Y_{\beta}^{q,i})^2) by (4.2) and (2.6), so all quantities in (3.2) are finite. Since the marks of different rows are independent and the leaf function is a sum over the rows, the recursion (3.2) factorizes over the rows. For one row, its steps are Tθl,ql−ql−1\mathcal{T}_{\theta_l,q_l-q_{l-1}} at depth ll, and the average over the root mark is T0,q0\mathcal{T}_{0,q_0}. By (3.5) and the recursion (2.2), the expected logarithm of the display equals

π[T0,q0Tθ1,q1−q0⋯Tθk,qk−qk−1T1,1−qkU](0)=πuζh(0,0;U),\pi\left[\mathcal{T}_{0,q_0}\mathcal{T}_{\theta_1,q_1-q_0}\cdots\mathcal{T}_{\theta_k,q_k-q_{k-1}}\mathcal{T}_{1,1-q_k}U\right](0) = \pi u_{\zeta_h}(0,0;U),

where ζh=∑lwlδql\zeta_h=\sum_l w_l\delta_{q_l} is the overlap law of the cascade, a probability measure on [0,1][0,1] whose distribution function equals 00 on [0,q0)[0,q_0) and θl+1\theta_{l+1} on [ql,ql+1)[q_l,q_{l+1}) for 0≤l≤k0\le l\le k, with qk+1=1q_{k+1}=1 (some of these intervals may be empty); its steps are exactly the operators in the display. Averaging over π∼Poi⁡(αn)\pi\sim\operatorname{Poi}(\alpha n) gives αn uζh(0,0;U)\alpha n\,u_{\zeta_h}(0,0;U). As h→∞h\to\infty, these values converge, by Lemma 4.8(iii) applied to the sequence that alternates the N′N'-dimensional systems and the cascade realizations, to the limit of the lower bound for the row term. Moreover uζh(0,0;U)→uζ~(0,0;U)u_{\zeta_h}(0,0;U)\to u_{\widetilde{\zeta}}(0,0;U) by Lemma 2.4, since the laws ζh\zeta_h converge weakly on [0,1][0,1] and hence in the Wasserstein distance. Thus, along the subsequences used below,

lim inf⁡N(row term)≥αn uζ~(0,0;U).(106)\liminf_{N}(\text{row term})\ge\alpha n\,u_{\widetilde{\zeta}}(0,0;U). \tag*{(106)}

In §4.7 we show that ζ~\widetilde{\zeta} equals the limiting overlap law ζ\zeta of the reference system.

Reduction of the coordinate term to the reference system. By (4.3) and (4.5), ZN′MZ_{N'}^{M} is an integral over (σ,ε)(\sigma,\varepsilon) against μN(dσ)fN,n(ε) dε\mu_N(\mathrm{d}\sigma)f_{N,n}(\varepsilon)\,\mathrm{d}\varepsilon. Restricting ε\varepsilon to the ball BΛ\mathbb{B}_{\Lambda} decreases it, so the coordinate term is bounded below by the same expression with ZN′MZ_{N'}^{M} replaced by its restriction to BΛ\mathbb{B}_{\Lambda}. By Lemma 4.7(iv)(a) we may replace the Gaussian part of the perturbation at ρ(σ,ε)\rho(\sigma,\varepsilon) by its value at ρ(σ,0)\rho(\sigma,0), at a cost o(1)o(1). After this replacement the Gaussian part depends only on σ\sigma. We add and subtract the activation terms ∑aUζ~(Sa)\sum_a U^{\widetilde{\zeta}}(S_a) at the reference fields, with ζ~=ζ~N′\widetilde{\zeta}=\widetilde{\zeta}_{N'}, which include the scalar part, and factor out the partition function of the tilted reference system. Similarly, by Lemma 4.7(iv)(b) the NN-dimensional system may be written, at a cost o(1)o(1), with the fields Sa+ηaS_a+\eta_a of (4.6), the reference perturbation, and the strength ξ~N′\tilde{\xi}_{N'}. Since the reference partition function is common to both terms, it cancels, and the coordinate term is at least

Elog⁡⟨∫BΛfN,n(ε)e∑a[Uξ~(χ(ε)Sa+wa(ε))−Uξ~(Sa)] dε e−nBN/2⟩ref−Elog⁡⟨e∑a[Uξ~(Sa+ηa)−Uξ~(Sa)]e−nBN/2⟩ref−o(1).(107)\begin{aligned} &\mathbb{E}\log\left\langle\int_{\mathbb{B}_{\Lambda}} f_{N,n}(\varepsilon)e^{\sum_a[U^{\tilde{\xi}}(\chi(\varepsilon)S_a+w_a(\varepsilon))-U^{\tilde{\xi}}(S_a)]}\,\mathrm{d}\varepsilon\,e^{-nB_N/2}\right\rangle_{\mathrm{ref}} \\ &\qquad-\mathbb{E}\log\left\langle e^{\sum_a[U^{\tilde{\xi}}(S_a+\eta_a)-U^{\tilde{\xi}}(S_a)]}e^{-nB_N/2}\right\rangle_{\mathrm{ref}}-o(1). \tag*{(107)} \end{aligned}

The factors e−nBN/2e^{-nB_N/2} undo the tilt of ⟨⋅⟩ref\langle\cdot\rangle_{\mathrm{ref}}; they appear in both terms. Here Sa=Sa(σ)S_a=S_a(\sigma), BN=BN(σ)B_N=B_N(\sigma), and ⟨⋅⟩ref\langle\cdot\rangle_{\mathrm{ref}} acts on σ\sigma.

Gaussian interpolation for the numerator. We replace the activation increment in the first term of (107) by a Gaussian field. Let z(σ)=(z1(σ),…,zn(σ))z(\sigma)=(z_1(\sigma),\ldots,z_n(\sigma)) be, conditionally on the disorder, a centered Gaussian process with Ezi(σ)zj(σ′)=δijXN(σ,σ′)\mathbb{E}z_i(\sigma)z_j(\sigma')=\delta_{ij}X_N(\sigma,\sigma'), independent of everything else (for instance the process zz of §4.3). For 0≤t≤10\le t\le1 put Sat=χ(ε)Sa+t wa(ε)S_a^t=\chi(\varepsilon)S_a+\sqrt{t}\,w_a(\varepsilon) and

Δt=∑a[Uξ~(Sat)−Uξ~(Sa)]+1−t ε⋅z(σ)+(1−t)∥ε∥22BN(σ),\Delta_t=\sum_a[U^{\tilde{\xi}}(S_a^t)-U^{\tilde{\xi}}(S_a)]+\sqrt{1-t}\,\varepsilon\cdot z(\sigma)+(1-t)\frac{\lVert\varepsilon\rVert^2}{2}B_N(\sigma),

and let φ(t)\varphi(t) be the first term of (107) with the exponent replaced by Δt\Delta_t; thus φ(1)\varphi(1) is that term. Write ⟨⋅⟩t\langle\cdot\rangle_t for the associated Gibbs measure on SN×BΛS_N\times\mathbb{B}_{\Lambda}. We first discuss the terms produced by the unperturbed activation UU and then the terms that carry a factor ξ~\tilde{\xi} or ξ~2\tilde{\xi}^2.

The field wa(ε)w_a(\varepsilon) is centered Gaussian with Ewa(ε)wa(ε′)=ε⋅ε′/N′\mathbb{E}w_a(\varepsilon)w_a(\varepsilon')=\varepsilon\cdot\varepsilon'/N', independently over aa, and it is independent of the reference system. Gaussian integration by parts in the fresh columns ga′g'_a produces the diagonal terms ∥ε∥2(U′′+U′2)(Sat)\lVert\varepsilon\rVert^2(U''+U'^2)(S_a^t) (the U′′U'' part from differentiating U′U', the U′2U'^2 part from differentiating the Gibbs density of the same replica) and a two-replica term. Integration by parts in zz produces the U′2U'^2 part and the two-replica term at the reference fields SaS_a, with the opposite sign, and the derivative of the explicit term (1−t)∥ε∥2BN/2(1-t)\lVert\varepsilon\rVert^2B_N/2 supplies the U′′U'' part at the reference fields, also with the opposite sign. The result, for the activation UU, is

φ′(t)=12N′∑aE⟨∥ε∥2[(U′′+U′2)(Sat)−(U′′+U′2)(Sa)]⟩t−12N′∑aE⟨(ε1⋅ε2)[U′(Sat,1)U′(Sat,2)−U′(Sa1)U′(Sa2)]⟩t′,(108)\begin{aligned} \varphi'(t) ={}&\frac{1}{2N'}\sum_a\mathbb{E}\left\langle \lVert\varepsilon\rVert^2\left[(U''+U'^2)(S_a^t)-(U''+U'^2)(S_a)\right] \right\rangle_t \\ &-\frac{1}{2N'}\sum_a\mathbb{E}\left\langle (\varepsilon^1\cdot\varepsilon^2)\left[U'(S_a^{t,1})U'(S_a^{t,2})-U'(S_a^1)U'(S_a^2)\right] \right\rangle_{t'}, \tag*{(108)} \end{aligned}

where the superscripts 1,21,2 refer to two replicas (σ1,ε1)(\sigma^1,\varepsilon^1), (σ2,ε2)(\sigma^2,\varepsilon^2). The first line collects the diagonal contributions and the second the two-replica covariance contributions.

We bound (108) on ΩN\Omega_N. Uniformly on SN×BΛS_N\times\mathbb{B}_{\Lambda},

∑aSa2≤CN,∑a∣Sat−Sa∣2≤Cn,Λ.\sum_a S_a^2\le CN,\qquad\sum_a\lvert S_a^t-S_a\rvert^2\le C_{n,\Lambda}.

The second bound follows from Sat−Sa=(χ(ε)−1)Sa+t wa(ε)S_a^t-S_a=(\chi(\varepsilon)-1)S_a+\sqrt{t}\,w_a(\varepsilon), from χ(ε)−1=On,Λ(N−1)\chi(\varepsilon)-1=O_{n,\Lambda}(N^{-1}), and from ∑awa(ε)2≤∥G′∥op2∥ε∥2/N′≤C∥ε∥2\sum_a w_a(\varepsilon)^2\leq\lVert\mathbf{G}'\rVert_{\mathrm{op}}^2\lVert\varepsilon\rVert^2/N'\leq C\lVert\varepsilon\rVert^2. Since ∣U′∣\lvert U'\rvert grows at most linearly and U′′,U′′′U'',U''' are bounded, the function ωU=U′′+U′2\omega_U=U''+U'^2 satisfies ∣ωU(x)−ωU(x′)∣≤C(1+∣x∣+∣x−x′∣)∣x−x′∣\lvert\omega_U(x)-\omega_U(x')\rvert\leq C(1+\lvert x\rvert+\lvert x-x'\rvert)\lvert x-x'\rvert, and the Cauchy–Schwarz inequality gives

1N′∑a∣ωU(Sat)−ωU(Sa)∣≤CN′(∑a(1+∣Sa∣+∣Sat−Sa∣)2)1/2(∑a∣Sat−Sa∣2)1/2≤Cn,ΛN−1/2.\begin{aligned} \frac{1}{N'}\sum_a\left\lvert\omega_U(S_a^t)-\omega_U(S_a)\right\rvert &\leq\frac{C}{N'}\left(\sum_a(1+\lvert S_a\rvert+\lvert S_a^t-S_a\rvert)^2\right)^{1/2} \left(\sum_a\lvert S_a^t-S_a\rvert^2\right)^{1/2}\\ &\leq C_{n,\Lambda}N^{-1/2}. \end{aligned}

The product difference in the second line obeys the same estimate, after writing it as [U′(Sat,1)−U′(Sa1)]U′(Sat,2)+U′(Sa1)[U′(Sat,2)−U′(Sa2)][U'(S_a^{t,1})-U'(S_a^1)]U'(S_a^{t,2})+U'(S_a^1)[U'(S_a^{t,2})-U'(S_a^2)]. Since ∥ε∥2≤Λn\lVert\varepsilon\rVert^2\leq\Lambda n and ∣ε1⋅ε2∣≤Λn\lvert\varepsilon^1\cdot\varepsilon^2\rvert\leq\Lambda n on BΛ\mathrm{B}_\Lambda, the right side of (108) is On,Λ(N−1/2)O_{n,\Lambda}(N^{-1/2}) on ΩN\Omega_N. On ΩNc\Omega_N^c the integrand is bounded by a polynomial in the operator norms, whose moments are bounded, so its contribution is at most Cn,ΛP(ΩNc)1/2C_{n,\Lambda}\mathbb{P}(\Omega_N^c)^{1/2}.

The terms carrying ξ\xi or ξ2\xi^2 arise because the interpolation uses the activation UξU^\xi of (100) while BNB_N and XNX_N are defined with UU. In (108) they replace ωU(Sat)\omega_U(S_a^t) by ωUξ(Sat)\omega_{U^\xi}(S_a^t) and U′(Sat,1)U′(Sat,2)U'(S_a^{t,1})U'(S_a^{t,2}) by (Uξ)′(Sat,1)(Uξ)′(Sat,2)(U^\xi)'(S_a^{t,1})(U^\xi)'(S_a^{t,2}). By (87) the first and second derivatives of τ1\tau_1 and τ2\tau_2 grow at most linearly, and the first derivatives of UU as well; hence ∣ωUξ(x)−ωU(x)∣≤C(ξ+ξ2)(1+x2)\lvert\omega_{U^\xi}(x)-\omega_U(x)\rvert\leq C(\xi+\xi^2)(1+x^2) and ∣(Uξ)′(x)(Uξ)′(x′)−U′(x)U′(x′)∣≤C(ξ+ξ2)(1+x2+x′2)\lvert(U^\xi)'(x)(U^\xi)'(x')-U'(x)U'(x')\rvert\leq C(\xi+\xi^2)(1+x^2+x'^2). Since N′−1∑a(1+(Sat)2)≤CN'^{-1}\sum_a(1+(S_a^t)^2)\leq C on ΩN\Omega_N, these terms contribute On,Λ(ξN′+ξN′2)=o(1)O_{n,\Lambda}(\xi_{N'}+\xi_{N'}^2)=o(1) on ΩN\Omega_N; this is sufficient, although it is weaker than the rate N−1/2N^{-1/2} available for the activation UU.

We conclude that φ(1)−φ(0)=o(1)\varphi(1)-\varphi(0)=o(1) for fixed nn and Λ\Lambda. The same holds for Gibbs expectations: differentiating E⟨F⟩t\mathbb{E}\langle F\rangle_t for a bounded measurable function FF of the spin overlaps and cavity coordinates of rr replicas, which does not depend on the fresh Gaussian variables, inserts at most r+2r+2 replicas and the same covariance differences, so ∣∂tE⟨F⟩t∣≤Cr,n,Λ∥F∥∞(N−1/2+ξN′+ξN′2)\lvert\partial_t\mathbb{E}\langle F\rangle_t\rvert\leq C_{r,n,\Lambda}\lVert F\rVert_\infty(N^{-1/2}+\xi_{N'}+\xi_{N'}^2) on ΩN\Omega_N, up to a contribution Cr,n,Λ∥F∥∞P(ΩNc)1/2C_{r,n,\Lambda}\lVert F\rVert_\infty\mathbb{P}(\Omega_N^c)^{1/2}.

The numerator at t=0t=0. By Taylor’s formula and (90), uniformly on BΛ\mathrm{B}_\Lambda and on ΩN\Omega_N,

∑a[U(χ(ε)Sa)−U(Sa)]=(χ(ε)−1)∑aSaU′(Sa)+O((χ(ε)−1)2∑aSa2)=n−∥ε∥22AN+On,Λ(N−1),\begin{aligned} \sum_a[U(\chi(\varepsilon)S_a)-U(S_a)] &=(\chi(\varepsilon)-1)\sum_aS_aU'(S_a) +O\left((\chi(\varepsilon)-1)^2\sum_aS_a^2\right)\\ &=\frac{n-\lVert\varepsilon\rVert^2}{2}A_N+O_{n,\Lambda}(N^{-1}), \end{aligned}

and the corresponding terms carrying ξ\xi are On,Λ(ξN′)O_{n,\Lambda}(\xi_{N'}). Adding ∥ε∥2BN/2\lVert\varepsilon\rVert^2B_N/2 and using bN−1=AN−BNb_N-1=A_N-B_N, the exponent Δ0\Delta_0 equals

ε⋅z(σ)−∥ε∥2−n2(bN−1)+n2BN+o(1).\varepsilon\cdot z(\sigma)-\frac{\lVert\varepsilon\rVert^2-n}{2}(b_N-1)+\frac{n}{2}B_N+o(1).

The term nBN/2nB_N/2 cancels the factor e−nBN/2\mathrm{e}^{-nB_N/2} in (107), and fN,n(ε) dεf_{N,n}(\varepsilon)\,\mathrm{d}\varepsilon may be replaced by νn(dε)\nu_n(\mathrm{d}\varepsilon) by (90). This explains the definitions of ANA_N, BNB_N, and bNb_N.

Gaussian interpolation for the denominator. The second term of (107) is treated by the same interpolation. The field ηa(σ)=cNgN′′a⋅σ\eta_a(\sigma)=c_N g_N^{\prime\prime a}\cdot\sigma has covariance cN2σ⋅σ′=(n/N′)RN(σ,σ′)c_N^2\sigma\cdot\sigma'=(n/N')R_N(\sigma,\sigma'), and on ΩN\Omega_N we have ∑aηa2≤cN2∥G′′∥op2N≤Cn\sum_a\eta_a^2\leq c_N^2\lVert G^{\prime\prime}\rVert_{\mathrm{op}}^2N\leq Cn. Let y(σ)y(\sigma) be, conditionally on the disorder, a centered Gaussian process with covariance nRN(σ,σ′)XN(σ,σ′)nR_N(\sigma,\sigma')X_N(\sigma,\sigma'), and interpolate with

∑a[Uξˉ(Sa+tηa)−Uξˉ(Sa)]+1−t y(σ)+(1−t)n2BN(σ).\sum_a\left[U^{\bar{\xi}}(S_a+\sqrt{t}\eta_a)-U^{\bar{\xi}}(S_a)\right]+\sqrt{1-t}\,y(\sigma)+(1-t)\frac{n}{2}B_N(\sigma).

Integration by parts gives (108) with ∥ε∥2\lVert\varepsilon\rVert^2 replaced by nn and ε1⋅ε2\varepsilon^1\cdot\varepsilon^2 replaced by nRN(σ1,σ2)nR_N(\sigma^1,\sigma^2), and the same estimates show that the two endpoints differ by o(1)o(1). At t=0t=0 the activation increment is replaced by y(σ)+nBN(σ)/2y(\sigma)+nB_N(\sigma)/2, and nBN/2nB_N/2 again cancels the factor e−nBN/2e^{-nB_N/2}.

The coordinate term. Combining the preceding paragraphs, for fixed nn and Λ\Lambda the coordinate term is at least

Elog⁡⟨∫BΛνn(dε) eε⋅z(σ)−(∥ε∥2−n)(bN(σ)−1)/2⟩ref−Elog⁡⟨ey(σ)⟩ref−o(1).(109)\mathbb{E}\log\left\langle\int_{\mathbb{B}_{\Lambda}}\nu_n(\mathrm{d}\varepsilon)\,e^{\varepsilon\cdot z(\sigma)-(\lVert\varepsilon\rVert^2-n)(b_N(\sigma)-1)/2}\right\rangle_{\mathrm{ref}} -\mathbb{E}\log\left\langle e^{y(\sigma)}\right\rangle_{\mathrm{ref}}-o(1). \tag*{(109)}

The same comparison holds for Gibbs expectations of bounded continuous functions of finitely many replicas, including their cavity coordinates: under the Gibbs measure of the N′N'-dimensional system with MM rows, conditioned on ε∈BΛ\varepsilon\in\mathbb{B}_{\Lambda} for each replica, such expectations differ by o(1)o(1) from those under the cavity measure of the first term of (109), where the spin overlaps are RN(σℓ,σℓ′)R_N(\sigma^\ell,\sigma^{\ell'}). Indeed, the chain of comparisons above consists of Lemma 4.7(iv)(a), the interpolation, whose effect on Gibbs expectations was bounded above, and the replacement of an exponent and a density by quantities that differ from them by o(1)o(1) uniformly on ΩN×BΛ\Omega_N\times\mathbb{B}_{\Lambda}. Moreover, on BΛ\mathbb{B}_{\Lambda} the spin overlap ρℓ⋅ρℓ′/N′=[χ(εℓ)χ(εℓ′)NRN(σℓ,σℓ′)+εℓ⋅εℓ′]/N′\rho^\ell\cdot\rho^{\ell'}/N'=[\chi(\varepsilon^\ell)\chi(\varepsilon^{\ell'})NR_N(\sigma^\ell,\sigma^{\ell'})+\varepsilon^\ell\cdot\varepsilon^{\ell'}]/N' of the N′N'-dimensional system differs from RN(σℓ,σℓ′)R_N(\sigma^\ell,\sigma^{\ell'}) by On,Λ(N−1)O_{n,\Lambda}(N^{-1}), and a continuous function is uniformly continuous on the compact set of overlaps in [−1,1][-1,1] and cavity coordinates in BΛ\mathbb{B}_{\Lambda}.

For parameters in GN(M)\mathcal{G}_N(M) and along the subsequences of §4.3, Lemma 4.8 allows us to replace bNb_N by its limit b∞b_{\infty}. Thus, for fixed nn and Λ\Lambda, the coordinate term is bounded below, up to o(1)o(1), by

Elog⁡⟨∫BΛνn(dε) eε⋅z(σ)−(∥ε∥2−n)(b∞−1)/2⟩ref−Elog⁡⟨ey(σ)⟩ref,(110)\mathbb{E}\log\left\langle\int_{\mathbb{B}_{\Lambda}}\nu_n(\mathrm{d}\varepsilon)\,e^{\varepsilon\cdot z(\sigma)-(\lVert\varepsilon\rVert^2-n)(b_{\infty}-1)/2}\right\rangle_{\mathrm{ref}} -\mathbb{E}\log\left\langle e^{y(\sigma)}\right\rangle_{\mathrm{ref}}, \tag*{(110)}

where, conditionally on the reference system, zz and yy are centered Gaussian with

Ezi(σ)zj(σ′)=δijXN(σ,σ′),Ey(σ)y(σ′)=nRN(σ,σ′)XN(σ,σ′).\mathbb{E}z_i(\sigma)z_j(\sigma')=\delta_{ij}X_N(\sigma,\sigma'),\qquad \mathbb{E}y(\sigma)y(\sigma')=nR_N(\sigma,\sigma')X_N(\sigma,\sigma').

The first term of (110) is the expected logarithm of the numerator NΛ\mathcal{N}_{\Lambda} of Lemma 4.8 for the reference system with b=b∞b=b_{\infty}, and the second is minus that of the denominator D\mathcal{D}; we call the two expected logarithms the numerator and the denominator. By Lemma 4.8(i), both converge along the subsequence, and by §4.3 their limits are the limits of their values on cascade realizations. We compute these values next.

The cavity quantities on cascades

We evaluate the numerator and the denominator of (110) on a cascade realization, using Lemma 3.4, and we show that the cutoff Λ\Lambda can be removed when the smallest precision is bounded below. Throughout this section we use the notation of §3.2: a Ruelle probability cascade of depth kk with parameters θl\theta_l and masses wlw_l, the branching level J=β1∧β2J=\beta^1\wedge\beta^2 of two replicas, and, for levels p∈Kp\in\mathcal{K} and a precision bb, the precisions dld_l, the mean levels mlm_l, the function G(b,p)G(b,p), the mean squared norm ϱ=mk+1/b\varrho=m_k+1/b, and the tilted measure M\mathcal{M} defined in (53) and below it.

Fix a cascade with paired levels 0≤q0≤⋯≤qk≤10\le q_0\le\cdots\le q_k\le1 and 0≤p0≤⋯≤pk0\le p_0\le\cdots\le p_k, and its cascade realization with diagonal values (1,pˉ,bˉ)(1,\bar p,\bar b) (§4.3). The overlap law of the cascade is ∑lwlδql\sum_l w_l\delta_{q_l}, and the law of the response overlap is ∑lwlδpl\sum_l w_l\delta_{p_l}. On the realization, the response field zz is the cascade field YβpY_\beta^p plus the within-state variable pˉ−pk x\sqrt{\bar p-p_k}\,x. Averaging xx inside the Gibbs average multiplies the integrand of the numerator by e(pˉ−pk)∥ε∥2/2e^{(\bar p-p_k)\lVert\varepsilon\rVert^2/2}, which changes the coefficient of −∥ε∥2/2-\lVert\varepsilon\rVert^2/2 from bˉ−1\bar b-1 to b−1b-1, where

b=bˉ−(pˉ−pk).(111)b=\bar b-(\bar p-p_k). \tag*{(111)}

By (52), the smallest precision d0=b−∑l≥1θlΔpld_0=b-\sum_{l\ge1}\theta_l\Delta p_l of (53) is expressed through the mean of the response overlap:

d0=b−pk+∑l=0kwlpl.(112)d_0=b-p_k+\sum_{l=0}^{k}w_lp_l. \tag*{(112)}

When d0>0d_0>0, all precisions are positive and Lemma 3.4 applies.

The cavity quantities on a cascade realization. Averaging the within-state coordinates inside the Gibbs average gives the following values. The numerator of (110), with b∞b_\infty replaced by bˉ\bar b, equals ΘΛ+n(bˉ−1)/2\Theta_\Lambda+n(\bar b-1)/2, where

ΘΛ=Elog⁡∑βvβ∫BΛνn(dε)eε⋅Yβp−(b−1)∥ε∥2/2,\Theta_\Lambda=\mathbb{E}\log\sum_\beta v_\beta\int_{\mathbb{B}_\Lambda}\nu_n(\mathrm{d}\varepsilon)e^{\varepsilon\cdot Y_\beta^p-(b-1)\lVert\varepsilon\rVert^2/2},

and we write Θ\Theta for the same quantity with ε\varepsilon integrated over Rn\mathbb{R}^n. When d0>0d_0>0, (55) gives Θ=n[G(b,p)−12log⁡b]\Theta=n[G(b,p)-\frac12\log b]. The denominator equals

Elog⁡∑βvβeYβc+n(pˉ−qkpk)/2=n2(pˉ−∑l=0kwlqlpl),\mathbb{E}\log\sum_\beta v_\beta e^{Y_\beta^c+n(\bar p-q_kp_k)/2} =\frac{n}{2}\left(\bar p-\sum_{l=0}^{k}w_lq_lp_l\right),

by (63) with the levels cl=nqlplc_l=nq_lp_l; here pˉ≥pk≥qkpk\bar p\ge p_k\ge q_kp_k. Hence, when d0>0d_0>0, the unrestricted numerator minus the denominator, divided by nn, is

G(b,p)−12log⁡b+12(bˉ−1−pˉ+∑l=0kwlqlpl).(113)G(b,p)-\frac12\log b+\frac12\left(\bar b-1-\bar p+\sum_{l=0}^{k}w_lq_lp_l\right). \tag*{(113)}

The last term vanishes in the limit by Lemma 4.9 (Step 3 of §4.7).

The cutoff is removed with the following lemma. It concerns the unrestricted tilted measure M\mathcal{M} of §3.2 with the field Y=YρY=Y^{\rho}, defined when d0>0d_{0}>0, and the random probability measure MΛ\mathcal{M}_{\Lambda} on pairs (β,ε)(\beta,\varepsilon) with density proportional to vβνn(dε)eε⋅Yβ−(b−1)∥ε∥2/21BΛ(ε)v_{\beta}\nu_{n}(\mathrm{d}\varepsilon)e^{\varepsilon\cdot Y_{\beta}-(b-1)\lVert\varepsilon\rVert^{2}/2}\mathbf{1}_{B_{\Lambda}}(\varepsilon), which is defined for every bb; when d0>0d_{0}>0, MΛ\mathcal{M}_{\Lambda} is M\mathcal{M} conditioned on {ε∈BΛ}\{\varepsilon\in B_{\Lambda}\}.

Lemma 4.10. Fix the cavity dimension nn and constants P0,B0,c0>0P_{0},B_{0},c_{0}>0, and consider finite cascades with

0≤pk≤P0,d0≥c0,b≤B0.0\le p_{k}\le P_{0},\qquad d_{0}\ge c_{0},\qquad b\le B_{0}.

There are constants C,c′>0C,c^{\prime}>0, depending only on n,P0,B0,c0n,P_{0},B_{0},c_{0}, such that for all Λ≥1\Lambda\ge1

0≤Θ−ΘΛ≤Ce−c′Λ,PM(ε1∉BΛ)≤Ce−c′Λ.(114)0\le\Theta-\Theta_{\Lambda}\le Ce^{-c^{\prime}\Lambda},\qquad \mathbb{P}_{\mathcal{M}}(\varepsilon^{1}\notin B_{\Lambda})\le Ce^{-c^{\prime}\Lambda}. \tag*{(114)}

Moreover ϱ=mk+1/b≤P0/c02+1/c0\varrho=m_{k}+1/b\le P_{0}/c_{0}^{2}+1/c_{0}.

Proof. Every denominator dl−1dld_{l-1}d_{l} and d02d_{0}^{2} in (53) is at least c02c_{0}^{2}, and the increments of pp sum to pk≤P0p_{k}\le P_{0}, so mk≤P0/c02m_{k}\le P_{0}/c_{0}^{2}; also 1/b≤1/d0≤1/c01/b\le1/d_{0}\le1/c_{0}. By (57), one replica ε1\varepsilon^{1} has the law N(0,ϱIn)N(0,\varrho I_{n}) under PM\mathbb{P}_{\mathcal{M}}, and the Gaussian tail bound for ∥ε1∥2\lVert\varepsilon^{1}\rVert^{2} gives the second estimate in (114), for Λ\Lambda larger than a constant; enlarging CC covers the remaining Λ≥1\Lambda\ge1.

A small averaged probability of {ε∉BΛ}\{\varepsilon\notin B_{\Lambda}\} alone would not bound Θ−ΘΛ\Theta-\Theta_{\Lambda}, so we use convexity. Define the leaf functions

f(y)=log⁡∫νn(dε)eε⋅y−(b−1)∥ε∥2/2,fΛ(y)=log⁡∫BΛνn(dε)eε⋅y−(b−1)∥ε∥2/2(y∈Rn).\begin{aligned} f(y)&=\log\int\nu_{n}(\mathrm{d}\varepsilon)e^{\varepsilon\cdot y-(b-1)\lVert\varepsilon\rVert^{2}/2},\\ f_{\Lambda}(y)&=\log\int_{B_{\Lambda}}\nu_{n}(\mathrm{d}\varepsilon)e^{\varepsilon\cdot y-(b-1)\lVert\varepsilon\rVert^{2}/2}\qquad(y\in\mathbb{R}^{n}). \end{aligned}

Then fΛ≤ff_{\Lambda}\le f, so ΘΛ≤Θ\Theta_{\Lambda}\le\Theta. The map φ↦Elog⁡∑βvβeφ(Yβ)\varphi\mapsto\mathbb{E}\log\sum_{\beta}v_{\beta}e^{\varphi(Y_{\beta})} is convex on leaf functions, and for a bounded direction ψ\psi its derivative at ff is EMψ(Yβ1)\mathbb{E}_{\mathcal{M}}\psi(Y_{\beta_{1}}), by dominated convergence. Applying convexity with the bounded direction ψt=max⁡{fΛ,f−t}−f∈[−t,0]\psi_{t}=\max\{f_{\Lambda},f-t\}-f\in[-t,0] gives

Θ−Elog⁡∑βvβemax⁡{fΛ,f−t}(Yβ)≤EM[f(Yβ1)−fΛ(Yβ1)],\Theta-\mathbb{E}\log\sum_{\beta}v_{\beta}e^{\max\{f_{\Lambda},f-t\}(Y_{\beta})} \le\mathbb{E}_{\mathcal{M}}\left[f(Y_{\beta_{1}})-f_{\Lambda}(Y_{\beta_{1}})\right],

and letting t→∞t\to\infty, by monotone convergence, Θ−ΘΛ≤EM[f(Yβ1)−fΛ(Yβ1)]\Theta-\Theta_{\Lambda}\le\mathbb{E}_{\mathcal{M}}[f(Y_{\beta_{1}})-f_{\Lambda}(Y_{\beta_{1}})]. Now exp⁡(fΛ−f)(y)\exp(f_{\Lambda}-f)(y) is the probability that a vector with law N(y/b,b−1In)N(y/b,b^{-1}I_{n}) lies in BΛB_{\Lambda}, and by Lemma 3.4(ii), Yβ1/b=Mk∼N(0,mkIn)Y_{\beta_{1}}/b=M_{k}\sim N(0,m_{k}I_{n}) under PM\mathbb{P}_{\mathcal{M}}. Hence

Θ−ΘΛ≤E[−log⁡P{Mk+b−1/2G′∈BΛ∣Mk}],\Theta-\Theta_{\Lambda}\le \mathbb{E}\left[-\log\mathbb{P}\left\{M_{k}+b^{-1/2}G^{\prime}\in B_{\Lambda}\mid M_{k}\right\}\right],

where G′G^{\prime} is a standard Gaussian vector in Rn\mathbb{R}^{n} independent of Mk∼N(0,mkIn)M_{k}\sim N(0,m_{k}I_{n}). We split the last expectation according to the size of MkM_k. On {∥Mk∥≤Λn/2}\{\lVert M_k\rVert\le\sqrt{\Lambda n}/2\}, the conditional probability of leaving BΛ\mathbb{B}_{\Lambda} is at most P(∥G′∥2>bΛn/4)≤P(∥G′∥2>c0Λn/4)≤Ce−c′Λ\mathbb{P}(\lVert G'\rVert^{2}>b\Lambda n/4)\le\mathbb{P}(\lVert G'\rVert^{2}>c_{0}\Lambda n/4)\le Ce^{-c'\Lambda}, and for Λ\Lambda larger than a constant the negative logarithm is at most twice this bound, since −log⁡(1−x)≤2x-\log(1-x)\le2x for 0≤x≤1/20\le x\le1/2. For all MkM_k and Λ≥1\Lambda\ge1, integrating the Gaussian density of Mk+b−1/2G′M_k+b^{-1/2}G' over the fixed ball B1\mathbb{B}_{1} and using ∥x−Mk∥2≤2∥x∥2+2∥Mk∥2\lVert x-M_k\rVert^{2}\le2\lVert x\rVert^{2}+2\lVert M_k\rVert^{2} gives

P{Mk+b−1/2G′∈BΛ∣Mk}≥c2exp⁡[−C2(1+∥Mk∥2)],\mathbb{P}\{M_k+b^{-1/2}G'\in\mathbb{B}_{\Lambda}\mid M_k\}\ge c_{2}\exp\left[-C_{2}(1+\lVert M_k\rVert^{2})\right],

with c2,C2c_{2},C_{2} depending only on n,c0,B0n,c_{0},B_{0}, since c0≤b≤B0c_{0}\le b\le B_{0}. The contribution of {∥Mk∥>Λn/2}\{\lVert M_k\rVert>\sqrt{\Lambda n}/2\} is therefore bounded by a Gaussian tail expectation of log⁡(1/c2)+C2(1+∥Mk∥2)\log(1/c_{2})+C_{2}(1+\lVert M_k\rVert^{2}), which is also at most Ce−c′ΛCe^{-c'\Lambda} uniformly for mk≤P0/c02m_k\le P_{0}/c_{0}^{2}. Enlarging CC covers the remaining Λ≥1\Lambda\ge1. □\square

The measure MΛ\mathcal{M}_{\Lambda} is the law of the leaves and cavity coordinates under the cavity measure of the cascade realization on BΛ\mathbb{B}_{\Lambda}, after the within-state coordinates are integrated out. For rr replicas, MΛ\mathcal{M}_{\Lambda} is M\mathcal{M} conditioned on the event that εℓ∈BΛ\varepsilon^{\ell}\in\mathbb{B}_{\Lambda} for ℓ≤r\ell\le r, realization by realization. If a probability measure gives mass 1−δ1-\delta to an event, then conditioning each of rr replicas on this event changes the expectation of a bounded function FF of the rr replicas by at most 2r∥F∥∞δ2r\lVert F\rVert_{\infty}\delta, because the conditioning event for the rr replicas has mass at least 1−rδ1-r\delta. Since this bound is linear in δ\delta, it survives averaging over the disorder:

∣EMΛF−EMF∣≤2r∥F∥∞PM(ε1∉BΛ).(115)\left|\mathbb{E}_{\mathcal{M}_{\Lambda}}F-\mathbb{E}_{\mathcal{M}}F\right|\le2r\lVert F\rVert_{\infty}\mathbb{P}_{\mathcal{M}}(\varepsilon^{1}\notin\mathbb{B}_{\Lambda}). \tag*{(115)}

Rotation invariance

The two identities that determine the entropy come from the following lemma. It uses only the invariance (104) of the N′N'-dimensional system under simultaneous rotations of the configurations and the rows. In the coordinates ρ=ρ(σ,ε)\rho=\rho(\sigma,\varepsilon), the quantity RcavR^{\mathrm{cav}} below is the overlap ε1⋅ε2/n\varepsilon^{1}\cdot\varepsilon^{2}/n of the added coordinates.

Lemma 4.11. Suppose that the disorder-averaged law of a random Gibbs measure on SN′S_{N'} is invariant under simultaneous orthogonal transformations of its replicas. For two replicas ρ1,ρ2\rho^{1},\rho^{2} define

R=ρ1⋅ρ2N′,Rcav=1n∑i=N+1N′ρi1ρi2.R=\frac{\rho^{1}\cdot\rho^{2}}{N'},\qquad R^{\mathrm{cav}}=\frac{1}{n}\sum_{i=N+1}^{N'}\rho_{i}^{1}\rho_{i}^{2}.

Then

E⟨(Rcav−R)2⟩≤3N′n(N′−1).(116)\mathbb{E}\left\langle(R^{\mathrm{cav}}-R)^{2}\right\rangle\le\frac{3N'}{n(N'-1)}. \tag*{(116)}

Moreover, the averaged one-replica law is uniform on SN′S_{N'}, and the law of its last nn coordinates converges to νn\nu_n as N→∞N\to\infty.

Proof. By the invariance, we may replace two given replicas by their images under an independent Haar-distributed rotation without changing the averaged law. Condition on two unit vectors ρ1/N′\rho^{1}/\sqrt{N'} and ρ2/N′\rho^{2}/\sqrt{N'} with inner product RR. Their images can be written as xx and Rx+1−R2x⊥Rx+\sqrt{1-R^{2}}x^{\perp}, where xx is uniform on the unit sphere of RN′\mathbb{R}^{N'} and, given xx, x⊥x^{\perp} is uniform on the unit sphere of the orthogonal complement of xx. For the nn coordinate indices i∈{N+1,…,N′}i\in\{N+1,\ldots,N'\} put A=∑ixi2A=\sum_{i}x_{i}^{2} and B=∑ixixi⊥B=\sum_{i}x_{i}x_{i}^{\perp}, so that Rcav=(N′/n)(RA+1−R2B)R^{\mathrm{cav}}=(N'/n)(RA+\sqrt{1-R^{2}}B). The variable AA has the beta distribution with parameters n/2n/2 and N/2N/2, so

EA=nN′,Var⁡(A)=2nNN′2(N′+2)≤2nN′(N′−1).\mathbb{E}A=\frac{n}{N'},\qquad\operatorname{Var}(A)=\frac{2nN}{N'^{2}(N'+2)}\leq\frac{2n}{N'(N'-1)}.

Given xx, E[xi⊥xj⊥∣x]=(δij−xixj)/(N′−1)\mathbb{E}[x_{i}^{\perp}x_{j}^{\perp}\mid x]=(\delta_{ij}-x_{i}x_{j})/(N'-1), hence E[B∣x]=0\mathbb{E}[B\mid x]=0 and E[B2∣x]=(A−A2)/(N′−1)\mathbb{E}[B^{2}\mid x]=(A-A^{2})/(N'-1). Therefore

EB2≤nN′(N′−1),E[(A−n/N′)B]=0.\mathbb{E}B^{2}\leq\frac{n}{N'(N'-1)},\qquad\mathbb{E}[(A-n/N')B]=0.

Consequently

E[(RA+1−R2B−Rn/N′)2]=R2Var⁡(A)+(1−R2)EB2≤3nN′(N′−1),\mathbb{E}\left[(RA+\sqrt{1-R^{2}}B-Rn/N')^{2}\right] =R^{2}\operatorname{Var}(A)+(1-R^{2})\mathbb{E}B^{2} \leq\frac{3n}{N'(N'-1)},

and multiplying by (N′/n)2(N'/n)^{2} gives (116). The averaged one-replica law is a rotation-invariant probability measure on SN′S_{N'}, hence uniform. By (89) its last nn coordinates have the density fN,nf_{N,n}, which converges pointwise to the standard Gaussian density; by Scheffé’s lemma the laws converge to νn\nu_{n} in total variation. □\square

Proof of the lower bound

The order of the limits matters, because the Gaussian integral over all of Rn\mathbb{R}^{n} in the numerator is not a bounded function of the overlaps. For fixed nn we first let N→∞N\to\infty along a subsequence at a fixed cutoff Λ\Lambda, then let the index hh of the approximating cascades tend to infinity, then let Λ→∞\Lambda\to\infty; the cavity dimension nn is sent to infinity last.

Proof of Theorem 4.1. We divide the proof into eight steps.

Step 1. Increments. Let ΞN\Xi_{N} be the expected log partition function of the perturbed NN-dimensional system with M∼Poi⁡(αN)M\sim\operatorname{Poi}(\alpha N) rows, averaged also over the perturbation parameters. By the Poissonization estimate of §4.1 and Lemma 4.7(i),

ΞNN=FN(U)+o(1).(117)\frac{\Xi_{N}}{N}=F_{N}(U)+o(1). \tag*{(117)}

For every fixed nn,

lim inf⁡N→∞ΞNN≥1nlim inf⁡N→∞(ΞN+n−ΞN).(118)\liminf_{N \to\infty}\frac{\Xi_N}{N} \ge\frac{1}{n}\liminf_{N \to\infty}(\Xi_{N+n}-\Xi_N). \tag*{(118)}

To verify this, let ℓ′\ell' be smaller than nn times the right side. Then ΞN+n−ΞN≥ℓ′\Xi_{N+n}-\Xi_N \ge\ell' for all N≥N0N \ge N_0. Every N≥N0+nN \ge N_0+n can be written as N=N1+jnN=N_1+jn with N0≤N1<N0+nN_0 \le N_1<N_0+n and j≥1j \ge1, and telescoping along the residue class of N1N_1 gives ΞN≥min⁡N0≤N1<N0+nΞN1+jℓ′\Xi_N \ge\min_{N_0 \le N_1<N_0+n}\Xi_{N_1}+j\ell'. Since j/N→1/nj/N \to1/n, the left side of (118) is at least ℓ′/n\ell'/n. This is the inequality on which the Aizenman–Sims–Starr scheme rests; see [ref-100].

Let ℓn=lim inf⁡N(ΞN+n−ΞN)\ell_n=\liminf_N(\Xi_{N+n}-\Xi_N). Then ℓn<∞\ell_n<\infty, since ΞN/N\Xi_N/N is bounded above by αUmax⁡+o(1)\alpha U_{\max}+o(1) and (118) holds; and ℓn≥−Cn\ell_n \ge-C_n by Step 2.

Step 2. Choice of parameters and row count. For fixed perturbation parameters and a number MM of rows, let ΔN(M)\Delta_N(M) be the increment on the left of (4.19), computed conditionally on the number of rows; thus ΞN+n−ΞN\Xi_{N+n}-\Xi_N is the average of ΔN(M)\Delta_N(M) over the parameters and over M∼Poi⁡(αN)M \sim\operatorname{Poi}(\alpha N), because M+π∼Poi⁡(αN′)M+\pi\sim\operatorname{Poi}(\alpha N'). By Lemma 4.7(v), ΔN(M)≥−Cn\Delta_N(M) \ge-C_n uniformly in the parameters and in M∈WNM \in W_N, and ∣ΔN(M)∣≤Cn(N+M)|\Delta_N(M)| \le C_n(N+M) for every MM. Since P(M∉WN)≤2e−cN1/3\mathbb{P}(M \notin W_N) \le2e^{-cN^{1/3}} and E(N+M)2≤CN2\mathbb{E}(N+M)^2 \le CN^2, the Cauchy–Schwarz inequality shows that the event {M∉WN}\{M \notin W_N\} contributes o(1)o(1) to the average; in particular ℓn≥−Cn\ell_n \ge-C_n. Choose a subsequence along which ΞN+n−ΞN→ℓn\Xi_{N+n}-\Xi_N \to\ell_n. Let GN′\mathcal{G}'_N be the event that M∈WNM \in W_N and that the parameters lie in the set GN(M)\mathcal{G}_N(M) of Lemma 4.7(iii). Its probability is at least P(M∈WN)inf⁡M∈WNP(GN(M))→1\mathbb{P}(M \in W_N)\inf_{M \in W_N}\mathbb{P}(\mathcal{G}_N(M)) \to1. Since ΔN≥−Cn\Delta_N \ge-C_n on {M∈WN}\{M \in W_N\}, the average of ΔN(M)\Delta_N(M) over GN′\mathcal{G}'_N is at most

ΞN+n−ΞN+CnP((GN′)c)+o(1)P(GN′)=ℓn+o(1)\frac{\Xi_{N+n}-\Xi_N+C_n\mathbb{P}((\mathcal{G}'_N)^c)+o(1)}{\mathbb{P}(\mathcal{G}'_N)} =\ell_n+o(1)

along the subsequence. Hence we can choose MN∈WNM_N \in W_N and parameters in GN(MN)\mathcal{G}_N(M_N) whose increment satisfies ΔN(MN)≤ℓn+o(1)\Delta_N(M_N) \le\ell_n+o(1). We fix these parameters and this row count from now on.

Step 3. One limit of the overlap arrays, and positive precision. Along a further subsequence, the finite-dimensional distributions of the arrays (RN,XN,bN)(R_N,X_N,b_N) under E⟨⋅⟩ref\mathbb{E}\langle\cdot\rangle_{\mathrm{ref}} converge, E⟨bN⟩ref→b∞\mathbb{E}\langle b_N\rangle_{\mathrm{ref}} \to b_\infty, E⟨XN(σ,σ)⟩ref→p∞\mathbb{E}\langle X_N(\sigma,\sigma)\rangle_{\mathrm{ref}} \to p_\infty, and the spin overlap array of the N′N'-dimensional system under E⟨⋅⟩N′\mathbb{E}\langle\cdot\rangle_{N'} converges. Write L\mathrm{L} for the limiting law of the arrays of the reference system, with diagonal values (1,p∞,b∞)(1,p_\infty,b_\infty), and EL\mathbb{E}_{\mathrm{L}} for expectations under it; let ζ\zeta be the law of R12R_{12} under L\mathrm{L}, and ζ~\widetilde{\zeta} the limiting overlap law of the N′N'-dimensional system. These limits do not depend on Λ\Lambda. The error terms o(1)o(1) of §4.4 depend on Λ\Lambda, but for every fixed Λ\Lambda they tend to zero along the whole sequence.

By §4.3, there are finite cascades, indexed by h≥1h \ge1, with depths khk_h and paired levels 0≤qlh≤10 \le q_l^h \le1 and 0≤plh≤p∞0 \le p_l^h \le p_\infty, whose cascade realizations with diagonal values (1,p∞,b∞)(1,p_\infty,b_\infty) have arrays converging to L\mathrm{L}. We attach the index hh to the quantities

References

No references were extracted for this paper.

Paper details

Contents