Introduction

For an integer n≥2n \ge2, let P+(n)P^{+}(n) denote its largest prime divisor, and put P+(1)=1P^{+}(1)=1. The Dickman–de Bruijn function ρ\rho is the continuous function on [0,∞)[0,\infty) determined by

ρ(u)=1(0≤u≤1),uρ′(u)=−ρ(u−1)(u>1).\rho(u)=1 \quad(0\le u\le1), \qquad u\rho'(u)=-\rho(u-1) \quad(u>1).

The classical theory of smooth numbers, beginning with Dickman and developed by Ramaswami and de Bruijn [4, 6, 22], gives

1X#{n≤X:P+(n)≤Xa}⟶ρ(1/a)(0<a≤1).\frac{1}{X}\#\{n\le X:P^{+}(n)\le X^{a}\}\longrightarrow\rho(1/a) \qquad(0<a\le1).

Our result identifies the joint law at two consecutive integers.

Theorem 1.1 (Joint Dickman law). For every fixed a,b∈(0,1)a,b\in(0,1),

lim⁡X→∞1X#{2≤n≤X:P+(n)≤na, P+(n+1)≤nb}=ρ(1/a)ρ(1/b).\lim_{X\to\infty}\frac{1}{X}\#\{2\le n\le X:P^{+}(n)\le n^{a},\ P^{+}(n+1)\le n^{b}\}=\rho(1/a)\rho(1/b).

Here the limit is taken over all real XX, with ordinary, unweighted counting.

Thus log⁡P+(n)/log⁡n\log P^{+}(n)/\log n and log⁡P+(n+1)/log⁡n\log P^{+}(n+1)/\log n have, in natural density, independent limiting distributions with distribution function a↦ρ(1/a)a\mapsto\rho(1/a), continuously extended to [0,1][0,1]. Section 11 proves the corresponding law with fixed thresholds XaX^{a}, XbX^{b}, derives the displayed moving-threshold statement, and identifies this continuous limiting law. Inclusion–exclusion using the fixed-threshold law and its marginals gives the upper-tail independence conjecture of Erdős and Pomerance [7], p. 311.

The comparison conjecture usually attributed to Erdős and Turán is a consequence, rather than a separate main theorem.

Corollary 1.2. One has

lim⁡X→∞1X#{2≤n≤X:P+(n)<P+(n+1)}=12.\lim_{X\to\infty}\frac{1}{X}\#\{2\le n\le X:P^{+}(n)<P^{+}(n+1)\}=\frac{1}{2}.

The reverse ordering has the same natural density.

The limiting marginal is continuous, so its product measure gives zero mass to the diagonal; symmetry then gives the corollary. In particular, no quantitative separation estimate is needed for this deduction.

Previous work

Erdős and Pomerance [7] formulated the joint independence problem and proved that each ordering has positive lower natural density. Their explicit lower bound was 0.0099. They also proved that for every ε>0\varepsilon>0, some δ>0\delta>0 makes the number of integers n<Xn<X satisfying X−δ<P+(n)/P+(n+1)<XδX^{-\delta}<P^{+}(n)/P^{+}(n+1)<X^{\delta} less than εX\varepsilon X for all sufficiently large XX [7], Theorem 1. The lower-density bound was increased to 0.05544 by de la Bretèche, Pomerance and Tenenbaum [5], Section 3, and to 0.05866 by an observation of Fouvry recorded in the same section. Wang [27, 28] obtained 0.1063 and 0.1356, and Lü and Wang [15] obtained 0.2017. Yang’s recent preprint [30], Theorem 1.4 gives 0.280 for both orderings. These are bounds on lower natural densities; they do not assert that a natural density exists.

For the joint law itself, Teräväinen [26] proved the predicted product in logarithmic density. That is, for fixed 0<a,b<10 < a,b < 1, the indicator in eq:1, averaged with weight 1/n1/n and normalization 1/log⁡X1/\log X, tends to ρ(1/a)ρ(1/b)\rho(1/a)\rho(1/b). His Theorem 1.16 gives logarithmic density 1/21/2 for the ordering, while Theorem 1.19 proves positive lower natural density for every nondegenerate rectangle of normalized largest-prime-factor values inside (0,1)2(0,1)^2. Tao and Teräväinen [24] then obtained the joint product law for ordinary averages outside an exceptional set of scales of logarithmic density zero; their Corollary 1.16 gives the corresponding ordering result. Wang [29] proved the ordinary joint law under the Elliott–Halberstam conjecture for friable integers. Jiang, Lü and Wang [13] established averaged-over-shift forms of the conjectures, a different conclusion from a result at the fixed shift one.

The more recent work of Tao and Teräväinen [25] gives a quantitative joint law outside an exceptional set of scales, with uniformity in the smoothness parameters and a power saving in log⁡X\log X. Theorem 1.1 concerns fixed parameters and has no exceptional scales; it does not assert their quantitative error term or their uniformity for growing parameters. Relative to the results just described, the distinction is therefore the unconditional ordinary limit at every sufficiently large scale, not a new marginal distribution or a logarithmically averaged independence statement.

Method and organization

We first replace the largest prime factor by finitely many prime counts. For a fixed integer J≥2J \ge2, divide the primes in (x1/J,x](x^{1/J},x] into bins Bk,x=(xk/J,x(k+1)/J]\mathcal{B}_{k,x} = (x^{k/J},x^{(k+1)/J}], 1≤k<J1 \le k < J. If ΩE(n)\Omega_E(n) counts prime factors in EE with multiplicity, a choice of unit complex numbers ζk\zeta_k defines the bin character

fx(n)=∏k=1J−1ζkΩBk,x(n).f_x(n) = \prod_{k=1}^{J-1} \zeta_k^{\Omega_{\mathcal{B}_{k,x}}(n)}.

Let gxg_x be defined by a second arbitrary phase vector, and let Fx=fx−μF_x = f_x - \mu, where μ\mu is the limiting mean of fxf_x. Finite Fourier inversion reduces the joint law to the mixed decorrelation

1x∑n<xgx(n)‾Fx(n+1)⟶0.\frac{1}{x}\sum_{n<x}\overline{g_x(n)}F_x(n+1) \longrightarrow0.

Here xx runs through integer scales. The two phase vectors may differ. This freedom gives independence of the two count vectors.

Two properties of these functions connect the beginning and end of the proof. First, every fixed positive integer multiplier has all its prime factors below the bins once xx is large. Thus Fx(un)=Fx(n)F_x(un)=F_x(n) and gx(un)=gx(n)g_x(un)=g_x(n) for each fixed uu. Second, FxF_x has cancellation in averages over growing intervals that remain short relative to the original scale. Let BB be the auxiliary parameter and let T=T(B)→∞T=T(B)\to\infty be the growing shift scale. If PB\mathcal{P}_B is a finite set of primes with min⁡PB→∞\min\mathcal{P}_B\to\infty, define

GB,ν(n)=∏p∈PBp∣n1+p−ν/B2,ν≥0.G_{B,\nu}(n)=\prod_{\substack{p\in\mathcal{P}_B\\p\mid n}}\frac{1+p^{-\nu/B}}{2},\qquad\nu\ge0.

For any interval IBI_B of LB→∞L_B\to\infty consecutive integer shifts, and fixed 0<s0<s10<s_0<s_1, positive integer DD, residue r(modD)r\pmod D, and ν≥0\nu\ge0, Section 2 proves

lim⁡B→∞lim sup⁡x→∞1Tx∑s0Tx≤n≤s1Tx∣1LB∑i∈IBi≡r(modD)Fx(n+i)GB,ν(n+i)∣2=0.\lim_{B\to\infty}\limsup_{x\to\infty}\frac{1}{Tx}\sum_{s_0Tx\le n\le s_1Tx}\left|\frac{1}{L_B}\sum_{\substack{i\in I_B\\i\equiv r\;(\mathrm{mod}\;D)}}F_x(n+i)G_{B,\nu}(n+i)\right|^2=0.

At fixed BB the shift interval is short relative to xx, while the base points have size TxTx. This is the scale produced by the divisor substitution below.

The centered function FxF_x is not multiplicative. On the relevant ranges n=OB(x)n = O_B(x), the bin counts take values in a fixed finite set. This permits interpolation by a fixed finite family of real, nonnegative multiplicative functions. After resolving the residue condition by Dirichlet characters, the real short-interval theorem of Matomäki and Radziwiłł [16] compares the principal components with long means whose centered combination vanishes. The complex theorem of Matomäki, Radziwiłł and Tao [17] controls the nonprincipal twists. The rest of the proof must reduce the mixed correlation to this particular cancellation statement.

Suppose, then, that the mixed correlation stays nonzero along a sequence of scales. The profinite integers Z^\widehat{\mathbb{Z}} encode compatible residue classes modulo all positive integers and carry Haar probability measure. Passing to a subsequence gives a bounded profile W(t,w)W(t,w) on (0,∞)×Z^(0,\infty) \times\widehat{\mathbb{Z}}. Its integral against each fixed compactly supported continuous test Φ(t,w)\Phi(t,w) is the subsequential limit of the mixed averages weighted by Φ(n/x,n)\Phi(n/x,n), with nn in the second argument viewed through its residues. A fixed nonnegative compactly supported smooth cutoff ϕ\phi can be chosen so that

∫ϕ(t)W(t,w) dt dw≠0,\int\phi(t)W(t,w)\,dt\,dw \ne0,

where dtdt is Lebesgue measure and dwdw is Haar measure.

For an integer nn, the amplifier defines DB(n)D_B(n) as a nonnegative weighted count of factorizations n=amn=am and n+1=clan+1=cla, restricted by eB<c<e2Be^B<c<e^{2B} and Tc<a<2TcTc<a<2Tc. At fixed BB its coefficient list is finite, and the same formula defines a function on Z^\widehat{\mathbb{Z}} depending only on residues at primes above a cutoff tending to infinity. The construction gives DBD_B Haar mean at least d0>0d_0>0 for large BB and bounded L2L^2 norm. To prove these bounds, the argument passes to a comparison model in which each site prime is assigned by an independent fair coin to the coefficient divisor or the remaining factor. The first and second moments are analyzed through one split and two conditionally independent splits of the same site prime sets. For each fixed BB, Φ(t,w)=ϕ(t)DB(w)\Phi(t,w)=\phi(t)D_B(w) is an allowed profile test, so along the selected subsequence

lim⁡x→∞1x∑n≥1ϕ(n/x)gx(n)Fx(n+1)DB(n)=∫ϕ(t)W(t,w)DB(w) dt dw.\lim_{x\to\infty}\frac{1}{x}\sum_{n\ge1}\phi(n/x)g_x(n)F_x(n+1)D_B(n)=\int\phi(t)W(t,w)D_B(w)\,dt\,dw.

Independence from every fixed residue coordinate and L2L^2 approximation of ϕW\phi W show that the right-hand side has absolute value bounded away from zero as B→∞B\to\infty.

For fixed BB and large xx, writing n=amn=am in this weighted average replaces gx(n)g_x(n) by gx(m)g_x(m). For each fixed pair (c,m)(c,m), the remaining terms form a sum over aa satisfying c∣am+1c\mid am+1, with nonnegative divisor and smooth weights multiplying Fx(am+1)F_x(am+1). Cauchy–Schwarz in (c,m)(c,m) removes the unit-modulus factor gx(m)g_x(m) and squares this inner sum. The resulting normalized quadratic energy has a positive lower bound, while its equal-coefficient diagonal tends to zero in the same iterated limit. A positive contribution therefore remains from distinct indices a,ba,b. Section 4 proves these moment and energy claims.

For such a pair, cc divides both am+1am+1 and bm+1bm+1. Thus mm is a unit modulo cc and a−b=jca-b=jc, with 0<∣j∣<T0<|j|<T by the coefficient windows. Writing la=(am+1)/cl_a=(am+1)/c and lb=(bm+1)/cl_b=(bm+1)/c, the exact substitution

n′=bla,n′+j=albn'=bl_a,\qquad n'+j=al_b

turns the label product into Fx(n′)‾Fx(n′+j)\overline{F_x(n')}F_x(n'+j) by invariance under the fixed multipliers a,b,ca,b,c. The new endpoints have size TxTx. The off-diagonal energy is therefore a weighted graph of ordinary additive shifts. Its edge weights depend on endpoint prime sets and on additional residue data from each divisor representation.

Group the new endpoints into blocks N+iN+i, 1≤i≤M1\le i\le M, with M=⌊C6T⌋M=\lfloor C_6T\rfloor for a fixed large C6C_6, and put s=N/(Tx)s=N/(Tx). Section 5 couples the actual auxiliary-prime divisibility sets at these positions to independent model sets SiS_i, each formed by including every auxiliary prime pp independently with probability 1/p1/p, and averages the extra residue weight of each representation. This compares the arithmetic graph with an endpoint kernel Lik(Si,Sk;s)L_{ik}(S_i,S_k;s) in expected cut norm. Here cut norm is the largest absolute normalized pairing with two real vectors bounded by one, chosen after the matrix is known. Section 6 controls representation multiplicity and provides the averaged second-moment bounds needed for sampling.

The aim is now to replace this endpoint kernel by products of functions of its two endpoints. For an endpoint type SS, let bb be the product obtained by retaining each prime independently with probability 1/21/2. The features used in the comparison are

Vl(S)=1∣Il∣Psplit(log⁡bB∈Il∣S),V_l(S)=\frac{1}{\lvert I_l\rvert}\mathbb{P}_{\mathrm{split}}\left(\frac{\log b}{B}\in I_l\mid S\right),

where the IlI_l form a fixed coarse partition of a compact logarithmic interval, chosen for a desired accuracy. The channel estimates in Section 7 show that fair splitting suppresses residue dependence and makes logarithmic outputs uniformly approximable on this partition. Fourier detection of a−b=jca-b=jc then lets Section 8 apply arithmetic estimates to the coefficient cc and these channel estimates to the endpoints. The resulting comparison has the form

∑lcl,B(j,s)Vl(Si)Vl(Sk),j=k−i,\sum_l c_{l,B}(j,s)V_l(S_i)V_l(S_k), \qquad j=k-i,

where i,ki,k are the endpoint positions. The scalar coefficients separate an arithmetic factor in ∣j∣\lvert j\rvert from a smooth factor in the normalized lag ∣j∣/T\lvert j\rvert/T and the common block origin ss. The number of features is fixed, while their values and the coefficients may depend on BB.

The feature comparison first holds after integration against separate functions of the two independent endpoint sets. The actual labels need not have this single-site form, and an optimizing test can depend on the whole configuration. Section 9 upgrades the comparison to expected cut norm by approximating an optimizing cut with a small sample of columns. Finally, polynomial approximation of each log-window indicator expresses VlV_l, up to a small mean-square error, as a fixed finite combination of

Esplit[e−vlog⁡b/B∣S]=∏p∈S1+p−v/B2.\mathbb{E}_{\mathrm{split}}\left[e^{-v\log b/B}\mid S\right]=\prod_{p\in S}\frac{1+p^{-v/B}}{2}.

At an integer endpoint these are exactly GB,vG_{B,v}. The arithmetic lag factor is approximated in mean over lags by a fixed periodic function. For a fixed small δ>0\delta>0 chosen for the desired accuracy, subdivide each position block into a fixed number of intervals whose lengths grow with TT and are at most δT\delta T. Within each interval pair, the normalized lag varies by at most 2δ2\delta. Approximating the smooth lag factor there and resolving the periodic factor into fixed residue classes turns the feature energy, up to a controlled error, into products of the weighted short averages already controlled in Section 2. Their vanishing contradicts the positive energy. Sections 10 and 11 complete this argument and the passage from bin events to the joint law.

The divisor amplification and graph comparison have close antecedents in Tao’s logarithmic correlation argument [23], the prime-divisibility graphs of Helfgott and Radziwiłł [12], Pilatte’s product-of-primes amplification [21], and the mixed-function decoupling of Tao and Teräväinen [25], Section 3.1. The present rough-divisor graph and its endpoint comparison are proved locally. The sampling argument likewise builds on the cut-based methods of Frieze and Kannan [10], Alon, Fernandez de la Vega, Kannan and Karpinski [1], Lemma 3, and Borgs, Chayes, Lovász, Sós and Vesztergombi [2] (Theorem 4.6). Because the kernels here vary with BB and can be large, the proof establishes the moment bounds needed before applying bounded-differences concentration [18].

The arithmetic estimates are developed in Section 3 using Selberg–Delange theory, character estimates and upper-bound sieves in the forms recorded by Granville and Koukoulopoulos [11], Koukoulopoulos [14], and Ford [8, 9]. The minor-arc input in Section 8 is the exponential-sum estimate of Montgomery and Vaughan [20]. Throughout the proof, the original scale xx tends to infinity before the auxiliary scale BB. Bin data, feature degrees, residue moduli and divisor multipliers are fixed in that inner limit. The contradiction applies to a subsequence of every possible bad sequence of original scales, and hence yields the full ordinary limit.

Large-prime labels and their short averages

We encode the large prime factors by two independently chosen multiplicative labels. Their one-variable distributions have limits independent of fixed residue conditions. We establish those limits and the weighted short-average estimate needed to prove that the two labels decorrelate at consecutive integers. In the short-average estimate the original counting scale xx tends to infinity before the auxiliary parameter BB does. The scale x>1x > 1 is real unless a statement explicitly restricts it to integers.

Labels and their one-variable distributions

Fix an integer J≥2J \ge2. For 1≤k<J1 \le k < J put

Bk,x=(xk/J,x(k+1)/J].(1)\mathcal{B}_{k,x} = \left(x^{k/J},x^{(k+1)/J}\right]. \tag*{(1)}

If EE is a set of primes, ΩE(n)\Omega_E(n) denotes the number of prime factors of nn belonging to EE, counted with multiplicity. Fix complex numbers ζ1,…,ζJ−1\zeta_1,\ldots,\zeta_{J-1} and ξ1,…,ξJ−1\xi_1,\ldots,\xi_{J-1} of absolute value one, independent of xx, with no relation required between the two vectors. Define the completely multiplicative function fxf_x by

fx(p)={ζk,p∈Bk,x,1≤k<J,1,p∉⋃1≤k<JBk,x,fx(n)=∏k=1J−1ζkΩBk,x(n).(2)\begin{aligned} f_x(p) &= \begin{cases} \zeta_k, & p \in\mathcal{B}_{k,x},\quad1 \le k < J,\\ 1, & p \notin\bigcup_{1 \le k < J}\mathcal{B}_{k,x}, \end{cases} &\qquad f_x(n) &= \prod_{k=1}^{J-1}\zeta_k^{\Omega_{\mathcal{B}_{k,x}}(n)}. \tag*{(2)} \end{aligned}

Define the completely multiplicative label associated with (ξk)(\xi_k) by

gx(n)=∏k=1J−1ξkΩBk,x(n).(3)g_x(n)=\prod_{k=1}^{J-1}\xi_k^{\Omega_{\mathcal{B}_{k,x}}(n)}. \tag*{(3)}

Both labels have absolute value one. For every fixed C>0C > 0, every integer n≤Cxn \le Cx has at most JJ prime factors in the union of the bins, once xx is sufficiently large in terms of C,JC,J. Indeed, J+1J+1 such factors would have product exceeding x1+1/J>Cxx^{1+1/J} > Cx.

Lemma 2.1 (One-variable distribution). For every fixed 0<α<β<∞0 < \alpha< \beta< \infty, positive integer qq, and residue class a(modq)a \pmod q, the distribution of

(ΩBk,x(n))1≤k<J\left(\Omega_{\mathcal{B}_{k,x}}(n)\right)_{1 \le k < J}

among the integers αx<n≤βx\alpha x < n \le\beta x, n≡a(modq)n \equiv a \pmod q, normalized to have mass one, converges as x→∞x \to\infty. Its limit depends only on JJ, and not on $\alpha,\beta,q,a.

Consequently there is a number μ=μ(J,ζ1,…,ζJ−1)\mu=\mu(J,\zeta_1,\ldots,\zeta_{J-1}) with ∣μ∣≤1|\mu|\le1 such that

lim⁡x→∞q(β−α)x∑αx<n≤βxn≡a(modq)fx(n)=μ.(4)\lim_{x\to\infty}\frac{q}{(\beta-\alpha)x}\sum_{\substack{\alpha x<n\le\beta x\\ n\equiv a\pmod q}}f_x(n)=\mu. \tag*{(4)}

The same assertion holds for gxg_x, with a mean μg\mu_g depending only on JJ and the vector (ξk)k=1J−1(\xi_k)_{k=1}^{J-1}. Write

Fx(n)=fx(n)−μ.F_x(n)=f_x(n)-\mu.

Then ∣Fx(n)∣≤2|F_x(n)|\le2, and for every fixed positive integer uu,

Fx(un)=Fx(n),gx(un)=gx(n)(n≥1)F_x(un)=F_x(n),\qquad g_x(un)=g_x(n)\qquad(n\ge1)

for all sufficiently large xx in terms of u,Ju,J.

Proof. We prove convergence by computing all mixed factorial moments of the bin counts. We first record the elementary prime estimates used in this calculation. Put ϑ(z)=∑p≤zlog⁡p\vartheta(z)=\sum_{p\le z}\log p. For z≥2z\ge2 we have

ϑ(z)≪z,π(z)≪zlog⁡z,∑p≤zlog⁡pp=log⁡z+O(1).(5)\vartheta(z)\ll z,\qquad\pi(z)\ll\frac{z}{\log z},\qquad\sum_{p\le z}\frac{\log p}{p}=\log z+O(1). \tag*{(5)}

Indeed, the primes in (m,2m](m,2m] divide (2mm)\binom{2m}{m}, giving ϑ(2m)−ϑ(m)≪m\vartheta(2m)-\vartheta(m)\ll m. Dyadic summation proves the first estimate, and separating the primes at z\sqrt{z} proves the second. Moreover,

∑pν≤zlog⁡p=∑1≤ν≤log⁡2zϑ(z1/ν)≪z,\sum_{p^\nu\le z}\log p=\sum_{1\le\nu\le\log_2 z}\vartheta(z^{1/\nu})\ll z,

because the terms with ν≥2\nu\ge2 are O(zlog⁡z)O(\sqrt{z}\log z). Expanding log⁡n\log n over prime powers now gives

∑n≤zlog⁡n=∑pν≤zlog⁡p⌊zpν⌋=z∑p≤zlog⁡pp+O(z).\sum_{n\le z}\log n=\sum_{p^\nu\le z}\log p\left\lfloor\frac{z}{p^\nu}\right\rfloor =z\sum_{p\le z}\frac{\log p}{p}+O(z).

The floor errors use the preceding prime-power bound, and the terms with ν≥2\nu\ge2 use ∑p∑ν≥2(log⁡p)/pν<∞\sum_p\sum_{\nu\ge2}(\log p)/p^\nu<\infty. Integral comparison on the left proves the third estimate in (2.7). Partial summation yields

∑xs<p≤xt1p⟶log⁡(t/s)(0<s<t fixed),(6)\sum_{x^s<p\le x^t}\frac{1}{p}\longrightarrow\log(t/s)\qquad(0<s<t\text{ fixed}), \tag*{(6)}

as well as ∑p≤z1/p=log⁡log⁡z+O(1)\sum_{p\le z}1/p=\log\log z+O(1). The number of integers n≤βxn\le\beta x divisible by p2p^2 for some prime p>x1/Jp>x^{1/J} is at most

βx∑p>x1/J1p2+Oβ(x)=o(x).\beta x\sum_{p>x^{1/J}}\frac{1}{p^2}+O_\beta(\sqrt{x})=o(x).

This remains negligible relative to the size of any one fixed progression in the lemma. We may therefore replace multiplicity counts by counts of distinct primes.

For the mixed factorial moments of these distinct-prime counts, write (b)r=b(b−1)⋯(b−r+1)(b)_r=b(b-1)\cdots(b-r+1) and (b)0=1(b)_0=1. Fix nonnegative integers r1,…,rJ−1r_1,\ldots,r_{J-1}, and let r=r1+⋯+rJ−1r=r_1+\cdots+r_{J-1}. The case r=0r=0 is immediate. For r>0r>0, a term in the factorial-moment expansion selects, in order, rkr_k distinct primes from bin kk, for each kk. Denote their product by dd; only d≤βxd \le\beta x can contribute. The number of such selections is OJ,β,r(x/log⁡x)O_{J,\beta,r}(x / \log x). To see this, fix every selected prime except one, and write d1d_1 for their product. If the last prime has any admissible choice, its upper bound βx/d1\beta x/d_1 exceeds x1/Jx^{1/J}. The prime-counting upper bound therefore gives at most

OJ,β(xd1log⁡x)O_{J,\beta}\left(\frac{x}{d_1 \log x}\right)

choices. Summing this bound over the other primes costs a bounded product of reciprocal prime sums over the bins. We may drop distinctness and product restrictions when taking this upper bound.

For sufficiently large xx, every selected prime is coprime to qq. The Chinese remainder theorem then gives

#{αx<n≤βx:n≡a(modq), d∣n}=(β−α)xqd+O(1).\#\{\alpha x<n\le\beta x:n\equiv a\pmod q,\ d\mid n\}=\frac{(\beta-\alpha)x}{qd}+O(1).

The sum of the errors over all selections is o(x)o(x). After normalization, the factorial moment is consequently the reciprocal sum over the selections with d≤βxd\le\beta x, up to o(1)o(1).

In the variable v=log⁡p/log⁡xv=\log p/\log x, eq:2 says that the reciprocal prime measure on bin kk converges to dv/vdv/v on (k/J,(k+1)/J](k/J,(k+1)/J]. The product restriction is

∑selected primesv≤1+log⁡βlog⁡x.\sum_{\text{selected primes}}v\le1+\frac{\log\beta}{\log x}.

The limiting product measure is absolutely continuous, so the boundary ∑v=1\sum v=1 has measure zero. Repeated selections do not change the limit: their reciprocal contribution is bounded by a constant times ∑p>x1/Jp−2=o(1)\sum_{p>x^{1/J}}p^{-2}=o(1). Returning from distinct-prime counts to multiplicity counts also costs o(1)o(1) in the normalized moment, because the exceptional set has size o(x)o(x) and the counts are bounded.

List the selected bin indices as k1,…,krk_1,\ldots,k_r, with rkr_k occurrences of kk. We have proved the explicit limit

lim⁡x→∞q(β−α)x∑αx<n≤βxn≡a(modq)∏k=1J−1(ΩBk,x(n))rk=∫vi∈(ki/J,(ki+1)/J](1≤i≤r)v1+⋯+vr≤1∏i=1rdvivi.(7)\begin{aligned} \lim_{x\to\infty}\frac{q}{(\beta-\alpha)x} \sum_{\substack{\alpha x<n\le\beta x\\ n\equiv a\pmod q}} \prod_{k=1}^{J-1}\bigl(\Omega_{B_{k,x}}(n)\bigr)_{r_k} &= \int_{\substack{v_i\in(k_i/J,(k_i+1)/J]\\(1\le i\le r)\\v_1+\cdots+v_r\le1}} \prod_{i=1}^{r}\frac{dv_i}{v_i}. \tag*{(7)} \end{aligned}

For r=0r=0, both sides are one. The right-hand side depends only on JJ and the multi-index (rk)(r_k). All count vectors lie in a fixed finite set, and mixed factorial polynomials span the functions on that set. The asserted distributional convergence follows.

Taking the expectation of the function (bk)↦∏kζkbk(b_k)\mapsto\prod_k\zeta_k^{b_k} gives eq:2; using the vector (ζk)(\zeta_k) gives the mean μg\mu_g. Finally, if every prime factor of uu is at most x1/Jx^{1/J}, then fx(u)=gx(u)=1f_x(u)=g_x(u)=1. Complete multiplicativity of fxf_x and gxg_x proves eq:2.

The centered function FxF_x is not asserted to be multiplicative. Its eventual invariance under each fixed multiplier is the exact property in (2.6) that will be used below.

The same distributional limit holds for averages over 1≤n≤x1\le n\le x. Indeed, apply the lemma to any fixed function of the count vector on αx<n≤x\alpha x<n\le x, then let α↓0\alpha\downarrow0; that function is bounded on the finite set of possible vectors, so the omitted interval has normalized contribution O(α)O(\alpha). In particular, the limit in (7) is unchanged when its normalized progression sum is replaced by x−1∑1≤n≤x∏k(ΩBk,x(n))rkx^{-1}\sum_{1\le n\le x}\prod_k(\Omega_{B_k,x}(n))^{r_k}. This gives (4) on the initial segment as well. A fixed shift of the integer argument changes only a bounded number of terms after nonpositive arguments are omitted, so the label means have the same limits after such a shift.

The mixed correlation and its short-average input

The joint law will follow by finite Fourier inversion from the following statement for the two independently chosen phase vectors. The factor Fx(n+1)F_x(n+1) is centered, while gx(n)g_x(n) has absolute value one.

Proposition 2.2 (Mixed decorrelation). For every fixed integer J≥2J\ge2 and every pair of fixed phase vectors (ζk)k=1J−1(\zeta_k)_{k=1}^{J-1} and (ξk)k=1J−1(\xi_k)_{k=1}^{J-1} of absolute value one, the corresponding labels satisfy

lim⁡x→∞1x∑1≤n<xgx(n)Fx(n+1)=0,(8)\lim_{x\to\infty}\frac{1}{x}\sum_{1\le n<x}g_x(n)F_x(n+1)=0, \tag*{(8)}

where Fx=fx−μF_x=f_x-\mu and the limit is through all integer scales.

After establishing the short-average input below, we will suppose that the average in (8) stays a fixed positive distance from zero along a sequence and derive a contradiction. The input is used in Section 10 at the final step of the proof.

For each auxiliary parameter BB, let PBP_B be a finite set of primes, and suppose that min⁡PB→∞\min P_B\to\infty as B→∞B\to\infty. An empty set may be allowed by interpreting its minimum as +∞+\infty and its product as one. For a fixed real v≥0v\ge0, define

GB,v(m)=∏p∣mp∈PB(12+12p−v/B).(9)G_{B,v}(m)=\prod_{\substack{p\mid m\\p\in P_B}}\left(\frac{1}{2}+\frac{1}{2}p^{-v/B}\right). \tag*{(9)}

This is a real, nonnegative, 1-bounded multiplicative function. For fixed BB it is periodic, with period dividing ∏p∈PBp\prod_{p\in P_B}p. To see its fair-split meaning, put SB(m)={p∈PB:p∣m}S_B(m)=\{p\in P_B:p\mid m\} and form a random product bb by including each prime of SB(m)S_B(m) independently with probability 1/21/2. The empty product is one. Independence gives

GB,v(m)=Esplit[exp⁡(−vlog⁡bB)].G_{B,v}(m)=\mathbb{E}_{\mathrm{split}}\left[\exp\left(-\frac{v\log b}{B}\right)\right].

These are the fair-split averages used to approximate the endpoint features in Section 10.

Lemma 2.3 (Weighted short averages of the centered labels). Let T=T(B)→∞T=T(B)\to\infty, let LBL_B be positive integers tending to infinity, and let aBa_B be any integer for each BB. Fix 0<s0<s1<∞0<s_0<s_1<\infty, a positive integer DD, a residue r(modD)r\pmod D, and v≥0v\ge0. Then

lim⁡B→∞lim sup⁡x→∞1Tx∑s0Tx≤n≤s1Tx∣1LB∑i=aB+1i≡r(modD)aB+LBFx(n+i)GB,v(n+i)∣2=0.(10)\lim_{B\to\infty}\limsup_{x\to\infty}\frac{1}{Tx}\sum_{s_0Tx\le n\le s_1Tx}\left|\frac{1}{L_B}\sum_{\substack{i=a_B+1\\i\equiv r\pmod D}}^{a_B+L_B}F_x(n+i)G_{B,v}(n+i)\right|^2=0. \tag*{(10)}

*For each fixed BB all the arguments of the arithmetic functions are positive once xx is sufficiently large. The same conclusion holds if the inner upper limit is taken along any sequence x→∞x\to\infty.、】【

The proof interpolates the centered function of the finitely many bin counts by real, nonnegative multiplicative functions. Once the residue of the origin nn is fixed, the congruence on ii becomes a fixed residue condition on the argument m=n+im=n+i. Resolving that condition by Dirichlet characters produces two different tasks. For the principal character, short averages must be compared with long means whose centered linear combination vanishes. For a nonprincipal character, the short averages themselves must be small. We record the two published inputs for these tasks and prove the required distance estimate before returning to the weighted lemma.

The short-interval inputs

For a bounded arithmetic function gg and H>0H>0, set

AHg(z)=1H∑z<n≤z+Hg(n),LXg=1X∑X<n≤2Xg(n).A_Hg(z)=\frac{1}{H}\sum_{z<n\le z+H}g(n),\qquad L_Xg=\frac{1}{X}\sum_{X<n\le2X}g(n).

For a 1-bounded multiplicative function gg, define

D(g,nit;X)2=∑p≤X1−Re⁡(g(p)p−it)p,M(g;X)=inf⁡∣t∣≤XD(g,nit;X)2.(11)\mathrm{D}(g,n^{it};X)^2=\sum_{p\le X}\frac{1-\operatorname{Re}(g(p)p^{-it})}{p},\qquad M(g;X)=\inf_{\lvert t\rvert\le X}\mathrm{D}(g,n^{it};X)^2. \tag*{(11)}

We use the following two forms of the short-interval theorems. Their uniformity in the multiplicative function is part of the statements.

Theorem 2.4 (Real short-interval comparison). There is a function r(H)r(H) tending to zero as H→∞H\to\infty such that, for every real multiplicative function g:N→[−1,1]g:\mathbb{N}\to[-1,1] and 2≤H≤X2\le H\le X,

1X∫X2X∣AHg(z)−LXg∣2 dz≤r(H).(12)\frac{1}{X}\int_X^{2X}\lvert A_Hg(z)-L_Xg\rvert^2\,dz\le r(H). \tag*{(12)}

This is a consequence of the uniform exceptional-set estimate in [16], using boundedness to pass to mean square. The terms in that estimate depending on XX can be absorbed into r(H)r(H) because X≥HX\ge H.

Theorem 2.5 (Complex short averages). For every multiplicative function g:N→Cg:\mathbb{N}\to\mathbb{C} with ∣g∣≤1\lvert g\rvert\le1 and X≥H≥10X\ge H\ge10,

1X∫X2X∣AHg(z)∣2 dz≪(1+M(g;X))e−M(g;X)+(log⁡log⁡Hlog⁡H)2+(log⁡X)−1/50,(13)\frac{1}{X}\int_X^{2X}\lvert A_Hg(z)\rvert^2\,dz\ll(1+M(g;X))e^{-M(g;X)}+\left(\frac{\log\log H}{\log H}\right)^2+(\log X)^{-1/50}, \tag*{(13)}

with an absolute implied constant.

This is [17], in the revised version with the corrected proof of Proposition A.3. That proof permits the factor e−Me^{-M}; we use the weaker (1+M)e−M(1+M)e^{-M}, which also covers M<1M<1. Thus divergence of the minimum over ∣t∣≤X\lvert t\rvert\le X suffices for our application.

The preceding statements also hold with integer origins, at the cost of an error tending to zero with HH. Indeed, for ∣g∣≤1\lvert g\rvert\le1,

∣AHg(z)−AHg(⌊z⌋)∣≤2H.\lvert A_Hg(z)-A_Hg(\lfloor z\rfloor)\rvert\le\frac{2}{H}.

Integrating over unit intervals proves the claim, with O(1/X)O(1/X) errors at the endpoints. Changing a short interval endpoint by a bounded amount similarly costs O(1/H)O(1/H).

A uniform estimate for the character twists

Lemma 2.6 (Fixed nonprincipal characters). Fix J≥2J \ge2, positive constants c,Cc,C, a nonprincipal Dirichlet character χ\chi of modulus qχq_\chi, and a finite set EE of primes. Suppose that, for each xx, hxh_x is a 1-bounded multiplicative function satisfying

hx(p)=χ(p)(p≤x1/J, p∉E).h_x(p)=\chi(p)\qquad(p\le x^{1/J},\ p\notin E).

With X=cxX=cx, one has

inf⁡∣t∣≤CXD(hx,nit;X)2⟶∞(x→∞).(14)\inf_{|t|\le CX}\mathbb{D}(h_x,n^{it};X)^2\longrightarrow\infty\qquad(x\to\infty). \tag*{(14)}

The assertion is uniform over all the functions hxh_x with this prime agreement. The threshold for xx may depend on J,c,C,χ,EJ,c,C,\chi,E.

Proof. Put y=x1/Jy=x^{1/J} and σ=1+1/log⁡x\sigma=1+1/\log x. For large xx we have y<Xy<X. Since the summands defining the distance are nonnegative, we can restrict to p≤yp\le y and delete EE and the primes dividing the modulus of χ\chi. Deleting those finitely many primes costs only a constant in the lower bounds below.

We record a prime-sum comparison, uniform in any factors of absolute value at most one. Replacing the weight 1/p1/p on p≤yp\le y by p−σp^{-\sigma} on all primes has absolute error OJ(1)O_J(1). Below yy, the error is at most

1log⁡x∑p≤ylog⁡pp=OJ(1).\frac{1}{\log x}\sum_{p\le y}\frac{\log p}{p}=O_J(1).

For the tail, partial summation and the prime-counting bound in (5) give

∑p>yp−σ≪∫y∞u−1−1/log⁡xlog⁡u du=∫1/J∞e−vv dv=OJ(1).\sum_{p>y}p^{-\sigma}\ll\int_y^\infty\frac{u^{-1-1/\log x}}{\log u}\,du=\int_{1/J}^\infty\frac{e^{-v}}{v}\,dv=O_J(1).

The prime-power terms in the logarithm of either a zeta or a Dirichlet LL Euler product are also uniformly O(1)O(1) for σ>1\sigma>1.

For ∣t∣≤1|t|\le1 it follows that

Re⁡∑p≤yχ(p)p−itp=log⁡∣L(σ+it,χ)∣+OJ(1)≤CJ,χ.\operatorname{Re}\sum_{p\le y}\frac{\chi(p)p^{-it}}{p}=\log|L(\sigma+it,\chi)|+O_J(1)\le C_{J,\chi}.

Here we used only an upper bound for the logarithm: a nonprincipal Dirichlet LL-function has no pole at 1 and is bounded on the compact region under consideration. Since (6) and the accompanying partial-summation estimate give ∑p≤y1/p=log⁡log⁡x+OJ(1)\sum_{p\le y}1/p=\log\log x+O_J(1), this proves a lower bound log⁡log⁡x−OJ,χ,E(1)\log\log x-O_{J,\chi,E}(1) for this frequency range.

For ∣t∣>1|t|>1, let mm be the order of χ\chi on the units. The elementary inequality ∣1−zm∣≤m∣1−z∣|1-z^m|\le m|1-z|, for ∣z∣=1|z|=1, gives

1−Re⁡(χ(p)p−it)≥1m2(1−Re⁡(p−imt))(p∤qχ).1-\operatorname{Re}(\chi(p)p^{-it})\ge\frac{1}{m^2}(1-\operatorname{Re}(p^{-imt}))\qquad(p\nmid q_\chi).

The Vinogradov–Korobov bound, with any fixed logarithmic exponent strictly between 2/32/3 and 1, gives

∣ζ(σ+iu)∣≪(log⁡(∣u∣+3))3/4(σ≥1, ∣u∣≥1).(15)|\zeta(\sigma+iu)|\ll(\log(|u|+3))^{3/4}\qquad(\sigma\ge1,\ |u|\ge1). \tag*{(15)}

The bound on the line 1 follows from [8]. Its extension to the right can be seen by applying the Phragmén–Lindelöf principle in 1≤σ≤21\le\sigma\le2 to

s−1s+1ζ(s)(log⁡(s+3))3/4.\frac{s-1}{s+1}\frac{\zeta(s)}{(\log(s+3))^{3/4}}.

The pole at 1 is removed, the logarithmic power has an analytic branch in this strip, and both vertical boundaries are bounded. For ∣Im⁡s∣≥1|\operatorname{Im}s| \ge1, the removed rational factor is bounded away from zero, which gives (2.17). For σ≥2\sigma\ge2 the same estimate follows from absolute convergence. Applying the prime-sum comparison just proved, we obtain uniformly for 1<∣t∣≤CX1 < |t| \le CX,

∑p≤y1−Re⁡(p−imt)p=log⁡log⁡x−log⁡∣ζ(σ+imt)∣+OJ(1)≥14log⁡log⁡x−OJ,c,C,χ(1).\sum_{p \le y} \frac{1-\operatorname{Re}(p^{-\mathrm{i}mt})}{p}=\log\log x-\log|\zeta(\sigma+\mathrm{i}mt)|+O_J(1) \ge\frac{1}{4}\log\log x-O_{J,c,C,\chi}(1).

Combining the two frequency ranges proves

D(hx,nit;X)2≥14m2log⁡log⁡x−OJ,c,C,χ,E(1)(∣t∣≤CX),\mathbb{D}(h_x,n^{\mathrm{i}t};X)^2 \ge\frac{1}{4m^2}\log\log x-O_{J,c,C,\chi,E}(1)\qquad(|t|\le CX),

which implies (2.16).

Proof of the weighted short-average lemma

Proof of Lemma 2.3. Teräväinen [26] uses real multiplicative generating functions for large-prime counts and recovers event coefficients from them. Here the finite count range gives a pointwise tensor interpolation of the centered function of all bin counts by completely multiplicative functions. We prove this interpolation with coefficients independent of B,xB,x. Choose J+1J+1 distinct numbers in (0,1](0,1]. The Vandermonde matrix whose entries are their powers of orders 0,…,J0,\ldots,J is invertible. Taking tensor products over the J−1J-1 count coordinates shows that there are finitely many complex constants cℓc_\ell and vectors zℓ=(zℓ,1,…,zℓ,J−1)∈(0,1]J−1z_\ell=(z_{\ell,1},\ldots,z_{\ell,J-1})\in(0,1]^{J-1} such that

∏k=1J−1zkbk−μ=∑ℓcℓ∏k=1J−1zℓ,kbk(0≤bk≤J).(16)\prod_{k=1}^{J-1}z_k^{b_k}-\mu=\sum_\ell c_\ell\prod_{k=1}^{J-1}z_{\ell,k}^{b_k}\qquad(0\le b_k\le J). \tag*{(16)}

The coefficients and vectors depend only on the fixed labels and JJ. Define

hℓ,x(m)=∏k=1J−1zℓ,kΩBk,x(m).h_{\ell,x}(m)=\prod_{k=1}^{J-1}z_{\ell,k}^{\Omega_{B_k,x}(m)}.

Each hℓ,xh_{\ell,x} is real, nonnegative, 1-bounded, and completely multiplicative. On every interval m≤CBxm\le C_Bx with CBC_B fixed for fixed BB, the identity Fx(m)=∑ℓcℓhℓ,x(m)F_x(m)=\sum_\ell c_\ell h_{\ell,x}(m) holds for all sufficiently large xx.

We next prove the needed estimate for each fixed residue class of the integer argument. For a class r′r' (mod DD), put d=gcd⁡(r′,D)d=\gcd(r',D) and q=D/dq=D/d, and write m=dkm=dk. Then m≡r′m\equiv r' (mod DD) is equivalent to k≡r′/dk\equiv r'/d (mod qq), a unit class modulo qq. For all sufficiently large BB, no prime factor of dd belongs to PBP_B, so GB,v(dk)=GB,v(k)G_{B,v}(dk)=G_{B,v}(k). For fixed BB and sufficiently large xx, (2.6) also gives Fx(dk)=Fx(k)F_x(dk)=F_x(k), and the same equality holds for every interpolant hℓ,xh_{\ell,x}.

By character orthogonality, the restricted kk-sum is a fixed linear combination of sums of

hℓ,x(k)GB,v(k)χ(k),χ(modq).h_{\ell,x}(k)G_{B,v}(k)\chi(k),\qquad\chi\pmod q.

All these functions are 1-bounded and multiplicative. The sum length after division by dd is HB=LB/dH_B=L_B/d. Normalization by LBL_B is d−1d^{-1} times normalization by HBH_B. Changing integer endpoints introduces an error OD(LB−1)O_D(L_B^{-1}).

Consider first the principal character χ0\chi_0 (mod qq). The functions in (2.19) are then real, so Theorem 2.4 applies to each interpolant. Although their individual long means need not vanish, their centered linear combination does. More precisely, on any dyadic interval [X,2X][X,2X] with X=cBxX=c_Bx and cB>0c_B>0 fixed for fixed BB,

∑ℓcℓLX(hℓ,xGB,vχ0)=1X∑X<k≤2XFx(k)GB,v(k)χ0(k)⟶0.\sum_{\ell}c_\ell\mathcal{L}_X(h_{\ell,x}G_{B,v}\chi_0)=\frac{1}{X}\sum_{X<k\le2X}F_x(k)G_{B,v}(k)\chi_0(k)\longrightarrow0.

To justify the last limit, split the sum into residue classes modulo the fixed common period of GB,vG_{B,v} and χ0\chi_0, and apply Lemma 2.1 to every class. The number and the sizes of these residue classes may depend on BB; they are fixed in this xx-limit. The mean square of the centered short average thus has inner upper limit at most a fixed multiple of r(HB)r(H_B), where the multiple depends on the interpolation coefficients and DD, but not on BB.

For a nonprincipal χ\chi, the function in (2.19) equals χ(p)\chi(p) at all primes p≤x1/Jp\le x^{1/J} outside the finite set PBP_B. Lemma 2.6 therefore applies for each fixed BB, with X=cBxX=c_Bx and the frequency range ∣t∣≤X|t|\le X. For all sufficiently large BB we have HB≥10H_B\ge10, so Theorem 2.5 applies once xx is sufficiently large. Taking the inner upper limit in (13) leaves at most a constant times (log⁡log⁡HB/log⁡HB)2(\log\log H_B/\log H_B)^2. This tends to zero as B→∞B\to\infty. There are only finitely many interpolants and characters, so the same holds after their linear combination.

Finally, for a given origin nn, the condition i≡r(modD)i\equiv r\pmod D in (10) is the condition m=n+i≡n+r(modD)m=n+i\equiv n+r\pmod D. There are only DD possible argument residues r′r', and it suffices to sum the bounds just proved for these possibilities. After m=dkm=dk, the short-interval origin is (n+aB)/d(n+a_B)/d. For fixed BB and all sufficiently large xx,

s0T2dx≤n+aBd≤2s1Tdx.\frac{s_0T}{2d}x\le\frac{n+a_B}{d}\le\frac{2s_1T}{d}x.

Choose a dyadic cover of the fixed scaled interval [s0T/(2d),2s1T/d][s_0T/(2d),2s_1T/d]. It gives intervals [cB,jx,2cB,jx][c_{B,j}x,2c_{B,j}x] with every cB,j>0c_{B,j}>0 fixed in the inner xx-limit, and with a number of intervals bounded in terms of s0,s1,Ds_0,s_1,D independently of BB. Their scales are comparable to Tx/dTx/d. Mapping the origins to their integer parts has multiplicity at most d+1d+1, and replacing an origin by its integer part costs O(HB−1)O(H_B^{-1}). The integer-origin version of the preceding estimates consequently applies. Normalizing by TxTx instead of the length of each dyadic range changes only fixed factors. Since HB=LB/d→∞H_B=L_B/d\to\infty for every d∣Dd\mid D, this proves (10).

All assertions used an unrestricted upper limit as x→∞x\to\infty; restricting that upper limit to a subsequence preserves them.

A profile of a nonzero mixed correlation

We begin the proof by supposing that Proposition 2.2 fails. Fix an offending JJ and pair of phase vectors. Boundedness permits passing to a subsequence on which the mixed correlation has a nonzero limit:

lim⁡x→∞1x∑1≤n<xgx(n)‾Fx(n+1)≠0.(17)\lim_{x\to\infty}\frac{1}{x}\sum_{1\le n<x}\overline{g_x(n)}F_x(n+1)\ne0. \tag*{(17)}

The arithmetic amplification in the following sections will contradict this limit; Section 10 completes the argument.

Let Z^\widehat{\mathbb{Z}} be the profinite integers, with Haar probability measure dwdw, and identify every integer with its natural image in Z^\widehat{\mathbb{Z}}. On (0,∞)×Z^(0,\infty)\times\widehat{\mathbb{Z}} define the locally finite measures

λx=1x∑n≥1δ(n/x,n),νx=1x∑n≥1gx(n)‾Fx(n+1)δ(n/x,n).\lambda_x=\frac{1}{x}\sum_{n\ge1}\delta_{(n/x,n)},\qquad \nu_x=\frac{1}{x}\sum_{n\ge1}\overline{g_x(n)}F_x(n+1)\delta_{(n/x,n)}.

Unweighted scale counting gives λx→dt dw\lambda_x \to dt\,dw against compactly supported continuous tests. This follows first for a continuous test in tt times a residue-class indicator in ww by progression counting, and then for every compactly supported continuous test by uniform approximation. Also ∣νx∣≤2λx|\nu_x| \le2\lambda_x. The resulting uniform variation bounds on compact sets, weak compactness on a countable exhaustion, and a diagonal extraction give a further subsequence along which νx\nu_x converges against these tests to a locally finite complex measure ν\nu.

For every compactly supported continuous Φ\Phi, the weak convergences give

∣∫Φ dν∣=lim⁡x→∞∣∫Φ dνx∣≤2lim⁡x→∞∫∣Φ∣ dλx=2∫0∞∫Z^∣Φ(t,w)∣ dw dt.\left|\int\Phi\,d\nu\right|=\lim_{x\to\infty}\left|\int\Phi\,d\nu_x\right|\le2\lim_{x\to\infty}\int|\Phi|\,d\lambda_x=2\int_0^\infty\int_{\widehat{\mathbb{Z}}}|\Phi(t,w)|\,dw\,dt.

The dual characterization of variation implies ∣ν∣≤2 dt dw|\nu|\le2\,dt\,dw. The Radon–Nikodym theorem therefore gives a measurable function

W:(0,∞)×Z^⟶C,∣W(t,w)∣≤2almost everywhere,(18)W:(0,\infty)\times\widehat{\mathbb{Z}}\longrightarrow\mathbb{C},\qquad|W(t,w)|\le2\quad\text{almost everywhere}, \tag*{(18)}

such that

lim⁡x→∞1x∑n≥1gx(n)‾Fx(n+1)Φ(n/x,n)=∫0∞∫Z^Φ(t,w)W(t,w) dw dt(19)\lim_{x\to\infty}\frac{1}{x}\sum_{n\ge1}\overline{g_x(n)}F_x(n+1)\Phi(n/x,n)=\int_0^\infty\int_{\widehat{\mathbb{Z}}}\Phi(t,w)W(t,w)\,dw\,dt \tag*{(19)}

for every compactly supported continuous Φ\Phi on the product. Here and henceforth limits in xx may use the fixed subsequence.

Boundedness also permits replacing the compactly supported continuous tests in (19) by the indicator of 0<t<10<t<1: truncate near 00, approximate the remaining interval at its endpoints, and use the uniform bound on both weighted counting measures and WW. Equation (17) therefore implies

∫01∫Z^W(t,w) dw dt≠0.\int_0^1\int_{\widehat{\mathbb{Z}}}W(t,w)\,dw\,dt\ne0.

Approximating this interval indicator by nonnegative smooth functions supported in (0,1)(0,1), we can fix a real nonnegative ϕ∈Cc∞((0,∞))\phi\in C^\infty_c((0,\infty)) with

β∗=∣∫0∞∫Z^ϕ(t)W(t,w) dw dt∣>0.(20)\beta_*=\left|\int_0^\infty\int_{\widehat{\mathbb{Z}}}\phi(t)W(t,w)\,dw\,dt\right|>0. \tag*{(20)}

The rest of the proof derives a contradiction from this fixed profile and bump. Every auxiliary choice will be made after J,fx,gx,Fx,W,ϕJ,f_x,g_x,F_x,W,\phi and the subsequence have been fixed.

Arithmetic preliminaries at an auxiliary log scale

This section supplies local asymptotic formulae and upper bounds for the coefficient and residue weights used in the amplification. All limits in this section are as B→∞B\to\infty. Constants may depend on explicitly fixed compact scale ranges, on the finitely many smooth norms indicated below, and later on the fixed regularity grid and its tolerance. They are uniform in the integer variables and moduli in the stated ranges. The parameter BB will always be held fixed when an inner limit in xx is taken elsewhere in the proof. We use the notation e(t)=exp⁡(2πit)e(t)=\exp(2\pi it).

Take BB through sufficiently large positive integers and put

T=⌊B0.32⌋,P0=B1000,P={p:P0<p≤exp⁡(4B)},R=Blog⁡P0,ℓ=log⁡R.T=\lfloor B^{0.32}\rfloor,\qquad P_0=B^{1000},\qquad\mathcal{P}=\{p:P_0<p\le\exp(4B)\},\qquad R=\frac{B}{\log P_0},\qquad\ell=\log R.

In particular R→∞R \to\infty, ℓ/log⁡B→1\ell/\log B \to1, and T<P0T < P_0. An integer is rough if it has no prime divisor at most P0P_0. Write μMob\mu_{\mathrm{Mob}} for the Möbius function, ω\omega for the number of distinct prime divisors, and ωE\omega_E for this count restricted to a set of primes EE.

The factors 2−ω2^{-\omega} turn divisor sums into averages over fair splits of prime sets. Set

A0(a)=(log⁡P0)R/2μMob2(a)2−ω(a)1a rough,K0(v)=R1/22−ωP(v).A_0(a)=(\log P_0)^{R/2}\mu_{\mathrm{Mob}}^2(a)2^{-\omega(a)}1_{a\ \mathrm{rough}}, \qquad K_0(v)=R^{1/2}2^{-\omega_{\mathcal{P}}(v)}.

The second definition also applies to v∈Z^v \in\widehat{\mathbb{Z}}, and to a subset S⊂PS \subset\mathcal{P} by taking ωP(S)=∣S∣\omega_{\mathcal{P}}(S)=|S|. For U⊂S⊂PU \subset S \subset\mathcal{P}, let aU=∏p∈Upa_U=\prod_{p\in U}p, with a∅=1a_{\varnothing}=1. The definitions give

A0(aU)K0(S∖U)=(log⁡P0)R/22−∣S∣=B2−∣S∣.(21)A_0(a_U)K_0(S\setminus U)=(\log P_0)^{R/2}2^{-|S|}=B2^{-|S|}. \tag*{(21)}

A fair split selects each prime of SS independently with probability 1/21/2. Thus summing the left side against a function of UU gives BB times its fair-split expectation.

The coefficient supports used below are contained in [1,exp⁡(bB)][1,\exp(bB)] for some fixed b<4b<4 once BB is large. Thus every prime divisor of a rough coefficient on these supports belongs to P\mathcal{P}.

Uniform local asymptotics

Our local estimates have two outputs. For A0A_0, they give a smooth mean on multiplicative windows at log scale BB, including additive phases with moduli and real frequencies bounded by fixed powers of BB. For a product aa formed by selecting each p∈Pp\in\mathcal{P} independently with probability z/pz/p, where z∈{1/4,1/2}z\in\{1/4,1/2\}, they give a smooth density for log⁡a/B\log a/B after restriction to a unit residue class, with an absolute error useful even on intervals of width O(1/B)O(1/B) in that coordinate. Both outputs follow by partial summation from cumulative estimates with a saving of any prescribed power of BB. We first recall the unrestricted Selberg–Delange estimates, then exclude the primes at most P0P_0 uniformly as that cutoff moves. For these two values of zz, define

gz(n)=μMob2(n)zω(n),gz,P0(n)=gz(n)1n rough.g_z(n)=\mu_{\mathrm{Mob}}^2(n)z^{\omega(n)}, \qquad g_{z,P_0}(n)=g_z(n)1_{n\ \mathrm{rough}}.

Lemma 3.1 (Classical estimates used in roughness removal). For each fixed nonnegative integer HH and either of the above values of zz, there are real constants cj(z)c_j(z) such that, for Y≥3Y\ge3,

∑n≤Ygz(n)=Y∑j=0Hcj(z)(log⁡Y)z−1−j+OH(Y(log⁡Y)z−2−H),\sum_{n\le Y}g_z(n)=Y\sum_{j=0}^{H}c_j(z)(\log Y)^{z-1-j}+O_H\bigl(Y(\log Y)^{z-2-H}\bigr),

where

c0(z)=1Γ(z)∏p(1+z/p)(1−1/p)z>0.c_0(z)=\frac{1}{\Gamma(z)}\prod_p(1+z/p)(1-1/p)^z>0.

For each fixed C,D>0C,D>0, uniformly over nonprincipal Dirichlet characters χ\chi of modulus d≤(log⁡Y)Cd\le(\log Y)^C, one has

∑n≤Ygz(n)χ(n)≪C,DY(log⁡Y)−D.\sum_{n\le Y}g_z(n)\chi(n)\ll_{C,D}Y(\log Y)^{-D}.

The constants in the second estimate, and the threshold after which it holds, need not be effective.

Proof. The first statement is the classical fixed-order Selberg–Delange expansion; see [11] and [14]. The normalization and the moving roughness cutoff required here are recorded explicitly below. In ℜs>1\Re s>1,

∑ngz(n)ns=ζ(s)zGz(s),Gz(s)=∏p(1+zp−s)(1−p−s)z.\sum_n\frac{g_z(n)}{n^s}=\zeta(s)^zG_z(s), \qquad G_z(s)=\prod_p(1+zp^{-s})(1-p^{-s})^z.

Define the local powers by their power-series logarithms. The logarithm of each local factor of GzG_z is O(p−2ℜs)O(p^{-2\Re s}). Consequently GzG_z is analytic, nonzero, and bounded on ℜs≥0.9\Re s \ge0.9, with all derivatives bounded on fixed smaller compact sets. Near s=1s=1, put u=s−1u=s-1. The analytic factor

(uζ(1+u))zGz(1+u)1+u\frac{(u\zeta(1+u))^zG_z(1+u)}{1+u}

has a convergent Taylor series and has value Gz(1)G_z(1) at u=0u=0. Truncated Perron inversion, followed by a contour around the cut to the left of 11, integrates its jjth Taylor term against Yeulog⁡Yu−j−zY^e u \log Y u^{-j-z}. Hankel’s reciprocal-gamma formula gives the factor (log⁡Y)z−j−1/Γ(z−j)(\log Y)^{z-j-1}/\Gamma(z-j). Taylor’s remainder, integrated on that contour, is OH(Y(log⁡Y)z−H−2)O_H(Y(\log Y)^{z-H-2}). For completeness, the contour may have right edge 1+1/log⁡Y1+1/\log Y, height exp⁡(clog⁡Y)\exp(c\sqrt{\log Y}), left edge 1−c′/log⁡Y1-c'/\sqrt{\log Y}, and small circular part of radius 1/log⁡Y1/\log Y. The classical zero-free region for ζ\zeta permits fixed sufficiently small c,c′>0c,c'>0; the other contour pieces and the Perron truncation error are OH(Yexp⁡(−c′log⁡Y)(log⁡Y)OH(1))O_H(Y\exp(-c'\sqrt{\log Y})(\log Y)^{O_H(1)}) and are absorbed in the displayed remainder. The local expansion just described also identifies c0(z)=Gz(1)/Γ(z)c_0(z)=G_z(1)/\Gamma(z).

We spell out the character uniformity. The twisted series is

L(s,χ)zGz,χ(s),Gz,χ(s)=∏p(1+zχ(p)p−s)(1−χ(p)p−s)z.L(s,\chi)^zG_{z,\chi}(s), \qquad G_{z,\chi}(s)=\prod_p(1+z\chi(p)p^{-s})(1-\chi(p)p^{-s})^z.

The same cancellation of the linear local term bounds Gz,χG_{z,\chi} uniformly in χ\chi on ℜs≥0.9\Re s\ge0.9. Write LY=log⁡YL_Y=\log Y and Θ=exp⁡(cLY)\Theta=\exp(c\sqrt{L_Y}), and use the rectangle

1−c′LY≤ℜs≤1+1LY,∣ℑs∣≤Θ.1-\frac{c'}{\sqrt{L_Y}}\le\Re s\le1+\frac{1}{L_Y}, \qquad|\Im s|\le\Theta.

The classical Dirichlet zero-free region excludes zeros in this rectangle except possibly a real zero of a real primitive character inducing χ\chi. Siegel’s bound states that for every fixed ϵ>0\epsilon>0 such a zero obeys 1−β≫ϵd−ϵ1-\beta\gg_\epsilon d^{-\epsilon}; see [14] for the zero-free region and exceptional-zero bound. Choose ϵ<1/(2C)\epsilon<1/(2C). Since d≤LYCd\le L_Y^C, this distance is eventually larger than c′/LYc'/\sqrt{L_Y}, uniformly in the allowed moduli. The Euler factors removed for imprimitive characters have no zeros in ℜs>0\Re s>0. Thus an analytic logarithm of L(s,χ)L(s,\chi) exists throughout the rectangle, agreeing with its Euler-product logarithm on its intersection with ℜs>1\Re s>1.

There is also a uniform polynomial bound in LYL_Y on this rectangle. For s=σ+its=\sigma+it truncate the Dirichlet series at V=⌈d(2+∣t∣)⌉V=\lceil d(2+|t|)\rceil. The periodic character sums have modulus at most dd, so partial summation bounds the tail by O(d(1+∣s∣)V−σ)O(d(1+|s|)V^{-\sigma}). The initial segment is at most O((1+log⁡V)Vmax⁡(1−σ,0))O((1+\log V)V^{\max(1-\sigma,0)}). Since log⁡V≪CLY\log V\ll_C\sqrt{L_Y} and (1−σ)log⁡V=OC(1)(1-\sigma)\log V=O_C(1) in the rectangle, these estimates bound L(s,χ)L(s,\chi) by a fixed power of LYL_Y. Because zz is real, the modulus of the chosen analytic power is ∣L(s,χ)z∣=∣L(s,χ)∣z|L(s,\chi)^z|=|L(s,\chi)|^z.

The coefficients have modulus at most one. At an endpoint Y∈Z+12Y\in\mathbb Z+\frac12, truncated Perron therefore has error O(Y(1+log⁡Y)2exp⁡(−cLY))O(Y(1+\log Y)^2\exp(-c\sqrt{L_Y})); this follows also by summing its usual min⁡(1,(Θ∣log⁡(Y/n)∣)−1)\min(1,(\Theta|\log(Y/n)|)^{-1}) error. Shift the Perron contour to the left edge. There is no pole or branch cut in the twisted case. The new vertical integral is bounded by Yexp⁡(−c′LY)Y\exp(-c'\sqrt{L_Y}) times a fixed power of LYL_Y, and the horizontal integrals have the additional factor Θ−1\Theta^{-1}. This proves an exponential square-root-log saving and hence the claimed saving of any fixed logarithmic power. For arbitrary Y≥3Y\ge3, let Y′Y' be the least element of Z+12\mathbb Z+\frac12 with Y′≥YY'\ge Y. Then 0≤Y′−Y<10\le Y'-Y<1 and d≤(log⁡Y)C≤(log⁡Y′)Cd\le(\log Y)^C\le(\log Y')^C. Replacing YY by Y′Y' changes the sum by at most one term and any smooth main term by an amount absorbed in the stated errors. This gives both estimates for arbitrary real YY.

Lemma 3.2 (The moving roughness cutoff). Fix D>0D > 0. There is a fixed integer H=H(D)H = H(D) and real coefficients Cj,z(P0)C_{j,z}(P_0), 0≤j≤H0 \le j \le H, such that uniformly for B0.89≤log⁡Y≤4BB^{0.89} \le\log Y \le4B,

∑n≤Ygz,P0(n)=Y∑j=0HCj,z(P0)(log⁡Y)z−1−j+OD(YB−D).\sum_{n \le Y} g_{z,P_0}(n) = Y \sum_{j=0}^{H} C_{j,z}(P_0)(\log Y)^{z-1-j} + O_D(YB^{-D}).

There is an absolute constant C0C_0 such that for every fixed jj,

∣Cj,z(P0)∣≪j(log⁡P0)C0+j,C0,z(P0)=c0(z)∏p≤P0(1+z/p)−1∼e−γEzΓ(z)(log⁡P0)−z,\lvert C_{j,z}(P_0)\rvert\ll_j (\log P_0)^{C_0+j}, \qquad C_{0,z}(P_0) = c_0(z) \prod_{p \le P_0}(1+z/p)^{-1} \sim\frac{e^{-\gamma_E z}}{\Gamma(z)}(\log P_0)^{-z},

where γE\gamma_E is Euler’s constant. Uniformly for nonprincipal χ\chi of modulus d≤B15d \le B^{15} in the same size range,

∑n≤Ygz,P0(n)χ(n)≪DYB−D.\sum_{n \le Y} g_{z,P_0}(n)\chi(n) \ll_D YB^{-D}.

Proof. Define the multiplicative function hzh_z by hz(pe)=(−z)eh_z(p^e)=(-z)^e for p≤P0p \le P_0 and e≥1e \ge1, and by hz(pe)=0h_z(p^e)=0 for p>P0p>P_0 and e≥1e \ge1. Its Euler factors give the exact identity gz,P0=gz∗hzg_{z,P_0}=g_z * h_z. If ϵ0=1/log⁡P0\epsilon_0=1/\log P_0, then

∑v∣hz(v)∣v1−ϵ0=∏p≤P0(1−zp−1+ϵ0)−1≪(log⁡P0)C0.\sum_v \frac{\lvert h_z(v)\rvert}{v^{1-\epsilon_0}} = \prod_{p \le P_0}(1-zp^{-1+\epsilon_0})^{-1} \ll(\log P_0)^{C_0}.

Indeed pϵ0≤ep^{\epsilon_0} \le e in the product, its logarithm is bounded by a constant times ∑p≤P0p−1+ϵ0\sum_{p \le P_0}p^{-1+\epsilon_0} plus a bounded sum of square terms, and Mertens’ estimate bounds that first sum by O(log⁡log⁡P0)O(\log\log P_0). Increasing the absolute constant C0C_0 if necessary gives the displayed assertion for both values of zz. Since (log⁡v)k≤k!ϵ0−kvϵ0(\log v)^k \le k!\epsilon_0^{-k}v^{\epsilon_0}, it also gives

∑v∣hz(v)∣v(log⁡v)k≪k(log⁡P0)C0+k.\sum_v \frac{\lvert h_z(v)\rvert}{v}(\log v)^k \ll_k (\log P_0)^{C_0+k}.

Put LY=log⁡YL_Y=\log Y. In the convolution sum the terms v>Yv>\sqrt{Y} contribute, after division by YY, at most

∑v>Y∣hz(v)∣v≪(log⁡P0)C0exp⁡(−LY2log⁡P0),\sum_{v>\sqrt{Y}}\frac{\lvert h_z(v)\rvert}{v} \ll(\log P_0)^{C_0}\exp\left(-\frac{L_Y}{2\log P_0}\right),

because the inner untwisted or twisted sum is at most Y/vY/v in absolute value. For v≤Yv \le\sqrt{Y}, apply Lemma 3.1 at Y/vY/v. For the untwisted sum expand each factor by Taylor’s formula:

(LY−log⁡v)z−1−j=∑k=0H−j(z−1−jk)(−log⁡v)kLYz−1−j−k+OH(LYz−2−H(log⁡v)H−j+1).(L_Y-\log v)^{z-1-j} = \sum_{k=0}^{H-j}\binom{z-1-j}{k}(-\log v)^kL_Y^{z-1-j-k} + O_H\left(L_Y^{z-2-H}(\log v)^{H-j+1}\right).

This remainder is uniform for 0≤log⁡v≤LY/20 \le\log v \le L_Y/2. After summing with hz(v)/vh_z(v)/v, the total Taylor and Selberg–Delange remainders are at most

OH(YLYz−2−H(log⁡P0)C0+H+1).O_H\left(YL_Y^{z-2-H}(\log P_0)^{C_0+H+1}\right).

The moment sums in the polynomial coefficients may be extended to all vv. To see that the resulting error is negligible, combine (log⁡v)k≪kϵ0−kvϵ0/2(\log v)^k \ll_k \epsilon_0^{-k}v^{\epsilon_0/2} with the preceding Rankin bound: the tail of each such moment is Ok((log⁡P0)C0+kexp⁡(−LY/(4log⁡P0)))O_k((\log P_0)^{C_0+k}\exp(-L_Y/(4\log P_0))). Thus explicitly

Cν,z(P0)=∑j+k=νcj(z)(z−1−jk)(−1)k∑vhz(v)v(log⁡v)k.C_{\nu,z}(P_0)=\sum_{j+k=\nu}c_j(z)\binom{z-1-j}{k}(-1)^k\sum_v\frac{h_z(v)}{v}(\log v)^k.

The moment bounds prove the claimed coefficient bounds. Since log⁡P0=1000log⁡B\log P_0=1000\log B and LY≥B0.89L_Y\ge B^{0.89}, choosing the fixed integer HH sufficiently large makes all these errors OD(YB−D)O_D(YB^{-D}). The zeroth moment is ∏p≤P0(1+z/p)−1\prod_{p\le P_0}(1+z/p)^{-1}. Combining it with the Euler product for c0(z)c_0(z) gives

C0,z(P0)=1Γ(z)∏p≤P0(1−1/p)z∏p>P0(1+z/p)(1−1/p)z.C_{0,z}(P_0)=\frac{1}{\Gamma(z)}\prod_{p\le P_0}(1-1/p)^z\prod_{p>P_0}(1+z/p)(1-1/p)^z.

The second product tends to one, and Mertens’ product formula proves its stated asymptotic and positivity.

For a nonprincipal character, use the same convolution with hz(v)χ(v)h_z(v)\chi(v) and the second estimate of Lemma 3.1. On v≤Yv\le\sqrt{Y} we have log⁡(Y/v)≥LY/2\log(Y/v)\ge L_Y/2 and, for all large BB, d≤B15≤(log⁡(Y/v))18d\le B^{15}\le(\log(Y/v))^{18}. Take an arbitrarily large fixed saving exponent in that lemma. The factor ∑v∣hz(v)∣/v≪(log⁡P0)c0\sum_v|h_z(v)|/v\ll(\log P_0)^{c_0} is absorbed by increasing that exponent. The terms v>Yv>\sqrt{Y} have already been bounded independently of χ\chi. This proves the required uniform character estimate.

We now turn these counting estimates into local densities for the coefficient weights and probability laws for randomly selected prime products. Choose the expansion order once, sufficiently large to use Lemma 3.2 with D=200D=200 for both values of zz. For L>0L>0 write

Pz,B(L)=∑j=0HCj,z(P0)Lz−1−j,Dz,B(L)=Pz,B(L)+Pz,B′(L).P_{z,B}(L)=\sum_{j=0}^{H}C_{j,z}(P_0)L^{z-1-j},\qquad D_{z,B}(L)=P_{z,B}(L)+P'_{z,B}(L).

Thus Dz,B(log⁡Y)D_{z,B}(\log Y) is exactly the derivative with respect to YY of YPz,B(log⁡Y)YP_{z,B}(\log Y). Define

mB(s)=(log⁡P0)R1/2D1/2,B(Bs),Qz,B=∏p∈P(1−z/p),fz,B(s)=BQz,BDz,B(Bs).m_B(s)=(\log P_0)R^{1/2}D_{1/2,B}(Bs),\qquad Q_{z,B}=\prod_{p\in\mathcal{P}}(1-z/p),\qquad f_{z,B}(s)=BQ_{z,B}D_{z,B}(Bs).

These are finite combinations of smooth powers on s>0s>0. The coefficient bounds and log⁡P0=Bo(1)\log P_0=B^{o(1)} imply, for every fixed nonnegative derivative order, convergence of these functions and their derivatives uniformly on compact subintervals of the indicated domains:

mB(s)⟶m(s)=cAs−1/2(0<s<4),fz,B(s)⟶czsz−1(0<s<3.3),m_B(s)\longrightarrow m(s)=c_A s^{-1/2}\quad(0<s<4),\qquad f_{z,B}(s)\longrightarrow c_zs^{z-1}\quad(0<s<3.3),

where

cA=e−γE/2Γ(1/2),cz=e−γEzzΓ(z).c_A=\frac{e^{-\gamma_E/2}}{\Gamma(1/2)},\qquad c_z=\frac{e^{-\gamma_E z}}{z\Gamma(z)}.

Here Qz,B∼(4R)−zQ_{z,B}\sim(4R)^{-z} by Mertens’ estimate. To check the derivative assertion directly, the j=0j=0 term has the stated limit after normalization. Every term with j≥1j\ge1, as well as the Pz,B′P'_{z,B} part, gains at least one power of B−1B^{-1} relative to that term on a fixed compact interval, with only a fixed power of log⁡P0\log P_0 lost. The same reasoning applies after any fixed number of ss-derivatives. Moreover the leading term dominates uniformly for Bs≥B0.89Bs\ge B^{0.89}: each ratio of a lower-order term to it is bounded by a fixed power of log⁡P0\log P_0 divided by a positive power of BsBs. Consequently fz,B(s)≪sz−1f_{z,B}(s)\ll s^{z-1} for B−1/10≤s≤3.2B^{-1/10}\le s\le3.2, and all these densities are positive on that range for sufficiently large BB.

Proposition 3.3 (Local coefficient and product laws). Let w3w_3 be smooth and supported in a fixed compact subinterval of (0,∞)(0,\infty). Uniformly for 0.9B≤log⁡X≤2.2B0.9B\le\log X\le2.2B, q≤B15q\le B^{15}, (u,q)=1(u,q)=1, and real ζ\zeta with ∣ζ∣≤B14|\zeta|\le B^{14},

∑cA0(c)w3(c/X)e(cu/q+ζc/X)=XμMob(q)φ(q)∫w3(t)mB(log⁡(Xt)/B)e(ζt) dt+O(XB−50).(22)\sum_c A_0(c)w_3(c/X)e(cu/q+\zeta c/X) = X\frac{\mu_{\mathrm{Mob}}(q)}{\varphi(q)} \int w_3(t)m_B(\log(Xt)/B)e(\zeta t)\,dt+O(XB^{-50}). \tag*{(22)}

The convention q=1q = 1 is included. The constant in the error is controlled by a fixed constant times finitely many low-order smooth norms of w3w_3; the expansion order defining mBm_B is independent of w3w_3.

Let νz\nu_z be the law of the product of primes of P\mathcal{P} selected independently with probabilities z/pz/p, and write s=log⁡a/Bs = \log a/B for its log coordinate. For d≤B15d \le B^{15}, every such product is a unit modulo dd. Uniformly for all intervals I′⊂[B−1/10,3.2]I' \subset[B^{-1/10}, 3.2] and unit residues r(modd)r \pmod d,

Pνz(s∈I′, a≡r(modd))=1φ(d)∫I′fz,B(s) ds+O(B−80).(23)\mathbb{P}_{\nu_z}(s \in I',\ a \equiv r \pmod d) = \frac{1}{\varphi(d)} \int_{I'} f_{z,B}(s)\,ds + O(B^{-80}). \tag*{(23)}

Uniformly for 0≤δ≤10 \le\delta\le1,

Pνz(s≤δ)≪(δ+1/R)z.(24)\mathbb{P}_{\nu_z}(s \le\delta) \ll(\delta+ 1/R)^z. \tag*{(24)}

For each positive integer n≤exp⁡(3.2B)n \le\exp(3.2B),

ν1/2(n)≪A0(n)Bn,\nu_{1/2}(n) \ll\frac{A_0(n)}{Bn},

where both sides are zero unless nn is a squarefree rough product, apart from the allowed empty product n=1n = 1.

Proof. Since P0>B15P_0 > B^{15}, all rough integers are units for every modulus in the statement. Character orthogonality and Lemma 3.2 give, for each unit class,

∑n≤Yn≡r (d)gz,P0(n)=γφ(d)Pz,B(log⁡Y)+O(YB−200).\sum_{\substack{n \le Y\\ n \equiv r\ (d)}} g_{z,P_0}(n) = \frac{\gamma}{\varphi(d)} P_{z,B}(\log Y) + O(YB^{-200}).

There is no loss of a factor φ(d)\varphi(d) here: it is cancelled by the normalization in character orthogonality. Partial summation, multiplication by (log⁡P0)R1/2(\log P_0)^{R^{1/2}} when z=1/2z = 1/2, and summation over the unit residues against e(ur/q)e(ur/q) now prove (22). The sum of the latter phases is the Ramanujan sum μMob(q)\mu_{\mathrm{Mob}}(q) because (u,q)=1(u,q) = 1. For explicit error accounting, summation over residues costs at most qq, the normalizing factor is at most BB eventually, and differentiating w3(t)e(ξt)w_3(t)e(\xi t) costs at most a constant times (1+∣ξ∣)(∥w3∥∞+∥w3′∥1+∥w3∥1)(1 + |\xi|)(\|w_3\|_\infty+ \|w_3'\|_1 + \|w_3\|_1) on the fixed support. Thus the initial B−200B^{-200} saving exceeds all these losses by much more than the claimed B−50B^{-50}. On that support log⁡(Xt)\log(Xt) is in the range of the preceding lemma once BB is sufficiently large.

For a squarefree product nn of the allowed primes, independence gives the exact mass

νz(n)=Qz,Bzω(n)n∏p∣n(1−z/p)−1.\nu_z(n) = Q_{z,B}\frac{z^{\omega(n)}}{n}\prod_{p\mid n}(1-z/p)^{-1}.

On n≤exp⁡(3.3B)n \le\exp(3.3B) the last product is 1+O(B/P0)1 + O(B/P_0), uniformly: its logarithm is O(ω(n)/P0)=O(B/P0)O(\omega(n)/P_0) = O(B/P_0). In this range every rough prime factor is below the upper cutoff exp⁡(4B)\exp(4B). Hence the mass without that last product is exactly Qz,Bgz,P0(n)/nQ_{z,B}g_{z,P_0}(n)/n. Partial summation of the unit-class formula between any two endpoints eBa,eBbe^{B a},e^{B b}, with B−1/10≤a≤b≤3.2B^{-1/10} \le a \le b \le3.2, gives

Qz,B∑eBa<n≤eBbn≡r (d)gz,P0(n)n=1φ(d)∫abBQz,BDz,B(Bs) ds+O(B−199).Q_{z,B}\sum_{\substack{e^{Ba}<n\le e^{Bb}\\ n\equiv r\ (d)}}\frac{g_{z,P_0}(n)}{n} = \frac{1}{\varphi(d)}\int_a^b BQ_{z,B}D_{z,B}(Bs)\,ds + O(B^{-199}).

Indeed each endpoint error is O(B−200)O(B^{-200}), and integrating the error O(tB−200)O(tB^{-200}) against dt/t2dt/t^2 over a log interval of length O(B)O(B) costs O(B−199)O(B^{-199}); also Qz,B≤1Q_{z,B} \le1. The O(B/P0)O(B/P_0) relative mass correction contributes at most O(B/P0)O(B/P_0) in total, since the uncorrected masses are bounded above by the true probabilities. Endpoint atoms have exponentially small mass because n≥exp⁡(B0.9)n \ge\exp(B^{0.9}). This proves (3.3) for arbitrary endpoint conventions and arbitrarily short or partial intervals. No relative error for a short interval is asserted or needed.

If δ<1/R\delta< 1/R, the event s≤δs \le\delta forces the product to be empty, so its probability is QZ,B≪R−zQ_{Z,B} \ll R^{-z}. If 1/R≤δ≤11/R \le\delta\le1, the event forces all primes with log⁡p>δB\log p > \delta B to be absent. Mertens’ product estimate therefore bounds its probability by

∏eδB<p≤e4B(1−z/p)≪δz.\prod_{\mathrm{e}^{\delta B}<p\le\mathrm{e}^{4B}}(1-z/p)\ll\delta^{z}.

These estimates prove (3.4), including δ=0\delta=0. Finally, on a squarefree rough nn in the range of (3.5),

ν1/2(n)A0(n)/(Bn)=Q1/2,BR1/2∏p∣n(1−1/(2p))−1=12+o(1)\frac{\nu_{1/2}(n)}{A_0(n)/(Bn)}=Q_{1/2,B}R^{1/2}\prod_{p\mid n}(1-1/(2p))^{-1}=\frac{1}{2}+o(1)

uniformly, and also for n=1n=1. This proves (3.5); outside the squarefree rough support both masses vanish.

Upper sieve bounds with reducing local weights

We state precisely the classical upper-sieve input, in its event-space form. For a fixed positive integer kk, choose once a bound Pk>2kP_k>2k depending only on kk. For a sieve bound Z≥2Z\ge2, let RZ⊂{p:Pk<p≤Z}\mathcal{R}_Z\subset\{p:P_k<p\le Z\} be the set of retained primes. Suppose their forbidden local densities obey 0≤θp≤k/p0\le\theta_p\le k/p. Mertens’ theorem then gives the upper dimension condition

∏p∈RZy<p≤z(1−θp)−1≤(log⁡zlog⁡y)k(1+Ok(1/log⁡y))(2≤y≤z).\prod_{\substack{p\in\mathcal{R}_Z\\y<p\le z}}(1-\theta_p)^{-1}\le\left(\frac{\log z}{\log y}\right)^k(1+O_k(1/\log y))\qquad(2\le y\le z).

Let V\mathcal{V} be the ambient mass. For each squarefree dd whose prime factors all lie in RZ\mathcal{R}_Z, let Vd\mathcal{V}_d be the mass satisfying the forbidden condition at every prime dividing dd, and put

rd=Vd−V∏p∣dθp.r_d=\mathcal{V}_d-\mathcal{V}\prod_{p\mid d}\theta_p.

Extend rdr_d by zero to all other positive integers. The classical fundamental lemma of the upper sieve supplies sk>0s_k>0, depending only on kk, and upper weights of absolute value at most one, supported on squarefree d≤Dd\le D with prime factors in RZ\mathcal{R}_Z, such that, when Z≤D1/skZ\le D^{1/s_k}, the mass avoiding the retained forbidden conditions is at most

CkV∏p∈RZ(1−θp)+∑d≤D∣rd∣.C_k\mathcal{V}\prod_{p\in\mathcal{R}_Z}(1-\theta_p)+\sum_{d\le D}|r_d|.

The constants are uniform over the forbidden sets satisfying the displayed dimension condition. This is the bounded-dimension upper-sieve statement of [9], Theorems 2.4 and 3.6.

Lemma 3.4 (Interval, rectangle, and random-root sieves). There is a sufficiently small ck>0c_k>0, depending only on fixed kk, with the following properties for Z≥2Z\ge2.

In an interval of length N≥1N\ge1, at most kk forbidden residues at each prime p≤Zp\le Z give an upper bound

CkN∏p≤Z(1−θp)+Ok(N0.8)(Z≤Nck).C_kN\prod_{p\le Z}(1-\theta_p)+O_k(N^{0.8})\qquad(Z\le N^{c_k}).

In a rectangle of side lengths N1,N2≥1N_1, N_2 \ge1, suppose that at each prime the forbidden set is a union of at most kk proper affine lines. With N∗=min⁡(N1,N2)N_* = \min(N_1,N_2), there is the corresponding bound

CkN1N2∏p≤Z(1−θp)+Ok(N1N2N∗−0.4)(Z≤N∗ck).C_kN_1N_2 \prod_{p\le Z}(1-\theta_p) + O_k(N_1N_2N_*^{-0.4}) \qquad(Z \le N_*^{c_k}).

An application may omit specified primes from the restrictions, provided it also omits their factors from the products.

Both assertions also apply to reducing local weights. At each prime, attach a factor in [0,1][0,1] to each of at most kk specified residue classes modulo pp in the interval case, or at most kk proper affine lines in the rectangle case. The factor is applied on its class or line and equals one off it. Replace 1−θp1-\theta_p by the residue average of the product of these factors.

Proof. At a squarefree modulus dd, the one-dimensional forbidden intersection has at most kω(d)k^{\omega(d)} residue classes, each counted with error O(1)O(1). Its remainder is therefore O(kω(d))O(k^{\omega(d)}). In the rectangle there are at most kω(d)dk^{\omega(d)}d residue pairs: at each prime there are at most kpkp pairs, and the Chinese remainder theorem multiplies these bounds. Each residue pair has count

N1N2d2+O(N1+N2d+1).\frac{N_1N_2}{d^2}+O\left(\frac{N_1+N_2}{d}+1\right).

Thus its remainder is O(kω(d)(N1+N2+d))O(k^{\omega(d)}(N_1+N_2+d)). For fixed kk, ∑d≤Dkω(d)≪kD(1+log⁡D)Ok(1)\sum_{d\le D} k^{\omega(d)} \ll_k D(1+\log D)^{O_k(1)}; one may obtain this by bounding kω(d)k^{\omega(d)} by a fixed divisor function and applying the elementary hyperbola bound to its sum. Use the preceding upper sieve with D=N11/2D=N_1^{1/2} in an interval, or D=N∗1/2D=N_*^{1/2} in a rectangle, and choose ck<1/(2sk)c_k<1/(2s_k). The interval remainder is N11/2+o(1)N_1^{1/2+o(1)}. The rectangle remainder is at most

(N1+N2)N∗1/2+o(1)+N∗1+o(1)≪N1N2N∗−0.4(N_1+N_2)N_*^{1/2+o(1)}+N_*^{1+o(1)}\ll N_1N_2N_*^{-0.4}

for sufficiently large N∗N_*. Enlarge the constants to cover smaller lengths.

The preceding sieve application temporarily omitted the primes p≤Pkp\le P_k. For each unweighted choice of forbidden conditions, restore the factors of those primes whose restrictions the application retains. If one such prime forbids the whole residue space, the fully sifted count is zero. Otherwise its surviving fraction 1−θp1-\theta_p is at least 1/p1/p in an interval and at least 1/p21/p^2 in a rectangle. The product of the reciprocals of these fractions over the bounded initial set is therefore bounded in terms of kk. Enlarging CkC_k restores all their factors in the displayed main term, without changing the remainder. This argument is uniform when conditions coincide and when ZZ includes only part of the initial set. Primes that an application chooses to omit remain absent from both its restrictions and its product.

For the weighted assertion, at every prime independently choose whether to forbid each specified condition, using probability one minus its local weight. Make these choices independently between conditions too, even if some coincide. For a fixed integer point, its probability of avoiding all the randomly chosen conditions is exactly the product of its local weights. Apply the unweighted bound for each choice and take expectations. The constants and remainders are uniform because every choice has at most kk residue conditions or proper lines. Independence between primes makes the expectation of the local-density product the product of the expected local densities. The latter is precisely the residue average stated in the lemma. The initial-prime factors were restored before this averaging, so no lower bound for an averaged local factor is needed. This proves the weighted version, including coincident conditions. Squarefreeness conditions and all factors in [0,1][0,1] above ZZ may be dropped when applying an upper bound. ∎

In the remaining estimates, n≍Xn \asymp X and analogous notation mean membership in an interval between fixed positive constant multiples of the indicated scale. Restricting to additional size or positivity conditions only decreases the upper bounds. They are uniform when (log⁡X)/B(\log X)/B lies in any fixed compact subinterval of (0,4)(0,4), with constants allowed to depend on that subinterval. In particular, they apply throughout 0.9B≤log⁡X≤2.2B0.9B \le\log X \le2.2B, the range used later. In applications involving TXTX, this convention still leaves all coefficient logarithms below 4B4B.

Lemma 3.5 (First and second coefficient moments). On these intervals,

X−1∑n≍XA0(n)≪1,X−1∑n≍XA0(n)2≪B1/4+o(1).X^{-1}\sum_{n\asymp X} A_0(n) \ll1,\qquad X^{-1}\sum_{n\asymp X} A_0(n)^2 \ll B^{1/4+o(1)}.

Proof. Choose a sufficiently small fixed c>0c>0 and sieve up to Z=exp⁡(cB)Z=\exp(cB), as permitted by Lemma 3.4 throughout the fixed log-size range. For a single form, local weight zero at p≤P0p\le P_0 and local weight t∈{1/2,1/4}t\in\{1/2,1/4\} above P0P_0 give residue averages 1−1/p1-1/p and 1−(1−t)/p1-(1-t)/p, respectively. Their product is

≪(log⁡P0)−1(log⁡Zlog⁡P0)t−1≪c(log⁡P0)−1Rt−1.\ll(\log P_0)^{-1}\left(\frac{\log Z}{\log P_0}\right)^{t-1}\ll_c (\log P_0)^{-1}R^{t-1}.

Dropping squarefreeness only increases the sum. Put HB=(log⁡P0)R1/2H_B=(\log P_0)R^{1/2}. Multiplication by HBH_B for the first moment gives one. Multiplication by HB2H_B^2 for the second moment gives

HB2(log⁡P0)−1R−3/4=(log⁡P0)R1/4=B1/4+o(1).H_B^2(\log P_0)^{-1}R^{-3/4}=(\log P_0)R^{1/4}=B^{1/4+o(1)}.

The sieve remainders remain negligible after these polynomial factors. The omitted primes above ZZ up to any of the fixed log-size endpoints have bounded reciprocal sum; in particular using this fixed small cc does not alter any power of RR in the bound.

Fix a sufficiently large absolute constant C2C_2 and define, for nonzero integers jj,

Σ(j)=∏p∣j(1+C2/p).\Sigma(j)=\prod_{p\mid j}(1+C_2/p).

For every fixed r>0r>0 this function satisfies ∑1≤j≤UΣ(j)r≪rU\sum_{1\le j\le U}\Sigma(j)^r\ll_r U for U≥1U\ge1. Indeed expand Σ(j)r=∑d∣jbr(d)\Sigma(j)^r=\sum_{d\mid j}b_r(d), with brb_r nonnegative, supported on squarefree integers, and br(p)=(1+C2/p)r−1=Or(1/p)b_r(p)=(1+C_2/p)^r-1=O_r(1/p). Then ∑dbr(d)/d=∏p(1+br(p)/p)<∞\sum_d b_r(d)/d=\prod_p(1+b_r(p)/p)<\infty, proving the assertion by summing the divisor expansion.

Lemma 3.6 (Two- and three-form coefficient bounds). Uniformly for 0<∣j∣≤T0<|j|\le T and squarefree rough b≍TXb\asymp TX,

∑c≍Xb+jc≍TX, b+jc>0A0(c)A0(b+jc)≪XΣ(j).\sum_{\substack{c\asymp X\\ b+jc\asymp TX,\ b+jc>0}} A_0(c)A_0(b+jc)\ll X\Sigma(j).

On the same fixed log-size ranges,

∑b≍TX, c≍Xa=b+jc≍TX, a>0A0(b)A0(a)A0(c)≪TX2Σ(j).\sum_{\substack{b\asymp TX,\ c\asymp X\\ a=b+jc\asymp TX,\ a>0}} A_0(b)A_0(a)A_0(c)\ll TX^2\Sigma(j).

The implied constants do not depend on bb, jj, XX, or BB.િ

Proof. Use the same sieve limit Z=exp⁡(cB)Z=\exp(cB) with cc sufficiently small. For (3.7), at p∤jbp\nmid jb the roots of cc and b+jcb+jc are distinct. At a prime at most P0P_0 the local density of avoiding them is 1−2/p1-2/p; at a larger prime the residue average of the two weights 1/21/2 is 1−1/p1-1/p. Apart from a factor 1+O(p−2)1+O(p^{-2}), these are the squares of the single-form factors in the preceding proof. Bounded small primes may be omitted. If p∣jp\mid j, then p<P0p<P_0 and bb is a unit at pp. Only the root of cc is forbidden. Its loss compared with two ordinary roots is at most 1+O(1/p)1+O(1/p) outside the omitted bounded set. The product of these losses is covered by Σ(j)\Sigma(j) after fixing C2C_2 sufficiently large. If p∣bp\mid b, then p>P0p>P_0 and we may drop both local conditions; their reciprocal sum is

∑p∣b1p≤ω(b)P0≪BP0.\sum_{p\mid b}\frac{1}{p}\leq\frac{\omega(b)}{P_0}\ll\frac{B}{P_0}.

Thus these exceptional primes have bounded total cost. Lemma 3.4 now gives density at most CΣ(j)(log⁡P0)−2R−1C\Sigma(j)(\log P_0)^{-2}R^{-1} in an interval of length comparable to XX. Its normalization HB2H_B^2 proves (3.7).

For (3.8), use the rectangle in (b,c)(b,c) of area comparable to TX2TX^2. At p∤jp\nmid j, the equations b=0b=0, c=0c=0, and b+jc=0b+jc=0 are three distinct lines. Their pairwise and triple intersections have density p−2p^{-2}. Inclusion–exclusion, also with the local reducing weights, therefore gives the product of the three single-form averages up to 1+O(p−2)1+O(p^{-2}). At p∣jp\mid j the lines for aa and bb coincide; since p<P0p<P_0, the loss from three ordinary rough exclusions to two is at most 1+O(1/p)1+O(1/p). Again their product is covered by Σ(j)\Sigma(j). The rectangular sieve gives density at most CΣ(j)(log⁡P0)−3R−3/2C\Sigma(j)(\log P_0)^{-3}R^{-3/2}, which is cancelled by HB3H_B^3. All sieve remainders are negligible because XX is exponential in BB, whereas TT, the normalizations, and the possible singular losses are polynomial in BB. Extending a support to a rectangle or interval for the sieve is harmless: the local congruence restrictions continue to hold on the original positive support, and the extension is used only for an upper bound.

Regularity cutoffs and their loss

Fix an integer L≥1L\geq1 and let Δ=1/L\Delta=1/L. The grid consists of g=i/Lg=i/L, 1≤i≤L1\leq i\leq L. Fix a small τ>0\tau>0 and, later, a sufficiently large C∗>0C_*>0. The choices of LL and τ\tau remain available for the additional fixed requirements in Section 6: first take LL sufficiently large, then take τ\tau sufficiently small. None of the estimates of the present section requires an upper bound for LL or a positive lower bound for τ\tau.

The prefix bounds will control competing divisor representations in Section 6; the tail lower bounds will make the two-split second moment in Section 4 summable.

Definition 3.7 (Regular prime sets). For S⊂PS\subset\mathcal{P} put

Ng(S)={∣{p∈S:log⁡p≤Bg}∣,g<1,∣S∣,g=1.N_g(S)= \begin{cases} \left|\{p\in S:\log p\leq B^g\}\right|, & g<1,\\ |S|, & g=1. \end{cases}

Let Yi=2ilog⁡P0Y_i=2^i\log P_0, 0≤i≤imax⁡0\leq i\leq i_{\max}, where imax⁡i_{\max} is the first index with Yimax⁡≥4BY_{i_{\max}}\geq4B. The set SS is regular if both of the following conditions hold:

(gℓ/2−τℓ)≤Ng(S)≤(gℓ/2+τℓ)for every grid point g,(g\ell/2-\tau\ell)\leq N_g(S)\leq(g\ell/2+\tau\ell)\quad\text{for every grid point }g,

and

∣{p∈S:log⁡p>Yi}∣≥0.4log⁡(B/Yi)−C∗(0≤i≤imax⁡).|\{p\in S:\log p>Y_i\}|\geq0.4\log(B/Y_i)-C_*\quad(0\leq i\leq i_{\max}).

An integer or profinite integer is regular when its set of dividing primes from P\mathcal{P} is regular. Multiplicities are not counted in these cutoffs.

Put A=A01regA=A_{0\mathrm{1reg}} and K=K01regK=K_{0\mathrm{1reg}}, with the analogous definition for K(S)K(S). The lower total-count bound in the definition gives

A(a), K(v), K(S)≤Bh0+τlog⁡2+o(1),h0=1−log⁡22.A(a),\ K(v),\ K(S) \le B^{h_0+\tau\log2+o(1)},\qquad h_0=\frac{1-\log2}{2}.

For AA, any prime factors outside P\mathcal{P} only decrease the weight; the additional factor log⁡P0\log P_0 is Bo(1)B^{o(1)}. Choose τ\tau sufficiently small that

2(h0+τlog⁡2)<0.32,1−6(h0+τlog⁡2)>0.07,1−5(h0+τlog⁡2)>0.229.2(h_0+\tau\log2)<0.32,\qquad1-6(h_0+\tau\log2)>0.07,\qquad1-5(h_0+\tau\log2)>0.229.

These are compatible strict conditions, since

2h0=0.3068528…,1−6h0=0.0794415…,1−5h0=0.2328679….2h_0=0.3068528\ldots,\qquad1-6h_0=0.0794415\ldots,\qquad1-5h_0=0.2328679\ldots.

Any further decrease of τ\tau preserves them.

Lemma 3.8 (Loss from imposing regularity). There is an absolute c3>0c_3>0 such that, for each fixed choice of the grid, tolerance, and compact log-size ranges above, there are C>0C>0 and a function ϵB→0\epsilon_B\to0 such that the part of the untruncated sum in (3.8) where at least one coefficient is not regular is at most

CTX2Σ(j)(ϵB+e−c3C∗).CTX^2\Sigma(j)\left(\epsilon_B+e^{-c_3C_*}\right).

The constant CC may depend on the fixed grid, tolerance, and scale ranges, but not on C∗,j,X,BC_*,j,X,B. The function ϵB\epsilon_B may depend on the same fixed data but is independent of C∗C_*. The analogous bound with scale XX applies to the first-moment sum of A0A_0 in (3.6). For the product of two independent coefficient weights in a box, it applies with the corresponding area in place of TX2Σ(j)TX^2\Sigma(j).

The same upper probability O(ϵB+e−c3C∗)O(\epsilon_B+e^{-c_3C_*}) holds for failure of regularity in a set of independent prime indicators on P\mathcal{P} with parameters 1/(2p)+O(p−2)1/(2p)+O(p^{-2}), uniformly when the constant in the OO term is fixed. It also holds after omitting a deterministic subset EB⊂PE_B\subset\mathcal{P} satisfying

∑p∈EB1p≤CEBP0\sum_{p\in E_B}\frac{1}{p}\le C_E\frac{B}{P_0}

for a fixed CE>0C_E>0. The subset may depend on BB, and the bound is uniform over all such subsets for each fixed CEC_E.

Proof. We give direct upper bounds relative to the scales in the statement; no division by the actual mass of a weighted sum is used. Sieve with the fixed Z=exp⁡(cB)Z=\exp(cB) used in Lemmas 3.5 and 3.6. For a coefficient whose count is being tested, replace its local weight 1/21/2 on a subset Q⊂PQ\subset\mathcal{P} by (1/2)es(1/2)e^s or (1/2)e−s(1/2)e^{-s}, where s>0s>0 is fixed and small enough that (1/2)es<1(1/2)e^s<1. The random-root sieve still applies. At ordinary primes the logarithm of its product, relative to the unmodified single-form factor, changes by

12(e±s−1)∑p∈Qp≤Z1p+O(1).\frac{1}{2}(e^{\pm s}-1)\sum_{\substack{p\in Q\\p\le Z}}\frac{1}{p}+O(1).

The O(1)O(1) is uniform for small fixed ss. In the three-form rectangle, line intersections alter it by a convergent sum of O(p−2)O(p^{-2}); the factors at primes dividing jj are still bounded by Σ(j)\Sigma(j), independently of such ss. The one-form case and independent two-variable case have the same conclusion with their respective scale bounds. All modified local weights lie in [0,1][0,1], so dropping those above ZZ preserves the direction of the bound, even for the positive tilt.

For a prefix g<1g<1, take Q={p∈P:log⁡p≤Bg}Q=\{p\in\mathcal{P}:\log p\leq B^{g}\}; for g=1g=1 take all of P\mathcal{P}. Mertens’ estimate gives

∑p∈Qp≤Z1p=gℓ+o(ℓ).\sum_{\substack{p\in Q\\p\leq Z}}\frac{1}{p}=g\ell+o(\ell).

For example, for g<1g<1 the left side is glog⁡B−log⁡log⁡P0+O(1)g\log B-\log\log P_0+O(1) once BB is large, and its difference from gℓg\ell is O(log⁡log⁡B)=o(ℓ)O(\log\log B)=o(\ell). There are only fixed finitely many grid points. The exponential Markov factor for a count exceeding (g/2+τ)ℓ(g/2+\tau)\ell is e−s(g/2+τ)ℓe^{-s(g/2+\tau)\ell}; the tilted sieve therefore bounds its weighted contribution by the appropriate base scale times

O(exp⁡([−sτ+g2(es−1−s)]ℓ+o(ℓ))).O\left(\exp\left(\left[-s\tau+\frac{g}{2}(e^s-1-s)\right]\ell+o(\ell)\right)\right).

For a count below (g/2−τ)ℓ(g/2-\tau)\ell the corresponding bound is

O(exp⁡([−sτ+g2(e−s−1+s)]ℓ+o(ℓ))).O\left(\exp\left(\left[-s\tau+\frac{g}{2}(e^{-s}-1+s)\right]\ell+o(\ell)\right)\right).

If the latter threshold is negative the event is empty. Otherwise choose ss sufficiently small in terms of the fixed τ\tau; since both exponential remainders in brackets are O(s2)O(s^2) and g≤1g\leq1, both brackets are at most −sτ/2-s\tau/2. Thus every prefix violation saves a fixed positive power of BB. A union bound over the grid and the at most three coefficients contributes o(1)o(1) times the scale bound.

For a tail endpoint Y=Yi<BY=Y_i<B, use the negative tilt on Q={p∈P:log⁡p>Y}Q=\{p\in\mathcal{P}:\log p>Y\}. Uniformly in these endpoints,

∑p∈Qp≤Z1p=log⁡(B/Y)+Oc(1).\sum_{\substack{p\in Q\\p\leq Z}}\frac{1}{p}=\log(B/Y)+O_c(1).

When Y≤cBY\leq cB this is Mertens’ estimate with upper endpoint ecBe^{cB}. When cB<Y<BcB<Y<B the sum is zero and log⁡(B/Y)=Oc(1)\log(B/Y)=O_c(1), which gives the same formula. Exponential Markov at the threshold 0.4log⁡(B/Y)−C∗0.4\log(B/Y)-C_* gives, relative to the appropriate base scale,

Oc(exp⁡(−sC∗+[0.4s+12(e−s−1)]log⁡(B/Y))).O_c\left(\exp\left(-sC_*+\left[0.4s+\frac{1}{2}(e^{-s}-1)\right]\log(B/Y)\right)\right).

Now choose a small absolute s>0s>0, independently of the grid. The bracket is −0.1s+O(s2)-0.1s+O(s^2) and is negative. With as=12(1−e−s)−0.4s>0a_s=\frac{1}{2}(1-e^{-s})-0.4s>0, the sum of these bounds over the dyadic endpoints Y<BY<B is

Oc(e−sC∗)∑Yi<B(Yi/B)as≪ce−sC∗.O_c(e^{-sC_*})\sum_{Y_i<B}(Y_i/B)^{a_s}\ll_c e^{-sC_*}.

For Y≥BY\geq B the required lower threshold is nonpositive, so there is no failure. Taking c3=sc_3=s proves the asserted tail loss.

The sieve remainders remain negligible in each application: their savings are powers of intervals exponential in BB, whereas normalization and Markov factors are only fixed powers of BB, and there are O(log⁡B)O(\log B) tail endpoints. For C∗≥0C_*\geq0, the Markov factors involving C∗C_* are at most their values at zero. Thus the residual function ϵB\epsilon_B can be chosen independently of C∗C_*. The prefix factors depend only on the fixed grid and τ\tau. This proves (3.11) and its one- and two-weight variants.

Finally let XpX_p be the independent indicators in the probability statement, with parameters λp=1/(2p)+O(p−2)\lambda_p=1/(2p)+O(p^{-2}). For either sign and any subset QQ,

Eexp⁡(±s∑p∈QXp)=∏p∈Q(1+λp(e±s−1)).\mathbb{E}\exp\left(\pm s\sum_{p\in Q}X_p\right)=\prod_{p\in Q}(1+\lambda_p(e^{\pm s}-1)).

Its logarithm is (e±s−1)∑p∈Qλp+O(∑pλp2)(e^{\pm s}-1)\sum_{p\in Q}\lambda_p+O(\sum_p\lambda_p^2). The latter error is bounded uniformly, while the parameter sums are gℓ/2+o(ℓ)g\ell/2+o(\ell) for a prefix and 12log⁡(B/Y)+O(1)\frac{1}{2}\log(B/Y)+O(1) for a tail with Y<BY<B. Omitting EBE_B changes any parameter sum by at most OCE(B/P0)=o(1)O_{C_E}(B/P_0)=o(1), uniformly for fixed CEC_E. Apply exactly the two Markov calculations above, now to this product moment-generating function and with no sieve remainder. This proves all the probability assertions. □\square

Amplifying a mixed correlation

Starting from the fixed nonzero correlation profile in Section 2, we construct a nonnegative divisor weight DBD_B that preserves it. We then expand the weighted correlation so that Cauchy–Schwarz removes gxg_x and leaves a positive quadratic energy in FxF_x. The divisor weight depends only on primes tending to infinity with BB; its positive mean and bounded second moment are the properties that will preserve the profile.

Retain WW, ϕ\phi, and

β∗=∣∫(0,∞)×Z^ϕ(t)W(t,w) dt dw∣>0\beta_*=\left|\int_{(0,\infty)\times\widehat{\mathbb Z}}\phi(t)W(t,w)\,dt\,dw\right|>0

from Section 2, and all parameters and weights from Section 3. In particular, ϕ\phi is real and nonnegative. Fix real, nonnegative, nonzero functions ρ0,ψ∈Cc∞((1,2))\rho_0,\psi\in C_c^\infty((1,2)). The regularity grid and its tolerance τ\tau are fixed throughout. A constant C∗C_* is also fixed whenever BB tends to infinity. Every limit in xx below is taken first, along the bad subsequence already selected in Section 2. Constants may depend on these fixed functions, the labels, and the grid and tolerance; dependence on C∗C_* will be indicated when it matters.

The divisor weight and the correlation it preserves

For an integer n≥1n\ge1, consider the two factorizations

n=am,n+1=cl.n=am,\qquad n+1=cl.

We select cc with log⁡c/B\log c/B in the support of ρ0\rho_0 and aa with a/(Tc)a/(Tc) in the support of ψ\psi. The weights A(a),A(c)A(a),A(c) apply to the chosen divisors and K(m),K(l)K(m),K(l) to the two quotients. Define their normalized divisor sum on the profinite integers by

DB(w)=1B∑a∣wc∣w+1A(a)K(w/a)A(c)K((w+1)/c)ρ0(log⁡cB)ψ(aTc),w∈Z^.(25)D_B(w)=\frac{1}{B}\sum_{\substack{a\mid w\\c\mid w+1}}A(a)K(w/a)A(c)K((w+1)/c)\rho_0\left(\frac{\log c}{B}\right)\psi\left(\frac{a}{Tc}\right),\qquad w\in\widehat{\mathbb Z}. \tag*{(25)}

For integer w=nw=n, this is exactly the sum over the two factorizations above. In Z^\widehat{\mathbb Z}, divisibility means membership in the image of multiplication by the divisor. Multiplication by a positive integer is injective on Z^\widehat{\mathbb Z}, so each quotient in the sum is well defined.

For fixed BB, the coefficient support is finite: eB<c<e2Be^B<c<e^{2B} and Tc<a<2TcTc<a<2Tc. Every coefficient is less than e4Be^{4B} for sufficiently large BB, so every prime of a nonzero coefficient weight belongs to PP. The modulus ∏p∈Pp2\prod_{p\in P}p^2 suffices to determine DBD_B: a coefficient is squarefree, and testing whether pp divides its quotient requires at most the residue modulo p2p^2. Thus DBD_B is a nonnegative function of finitely many high-prime residues.

Define I1,xI_{1,x} by the weighted correlation

I1,xB=1x∑n≥1ϕ(n/x)gx(n)‾Fx(n+1)DB(n).(26)\frac{I_{1,x}}{B}=\frac{1}{x}\sum_{n\ge1}\phi(n/x)\overline{g_x(n)}F_x(n+1)D_B(n). \tag*{(26)}

For each fixed BB, the function ϕ(t)DB(w)\phi(t)D_B(w) is a compactly supported continuous profile test. The convergence in (19) therefore gives

lim⁡x→∞I1,xB=∫(0,∞)×Z^ϕ(t)W(t,w)DB(w) dt dw.(27)\lim_{x\to\infty}\frac{I_{1,x}}{B}=\int_{(0,\infty)\times\widehat{\mathbb{Z}}}\phi(t)W(t,w)D_B(w)\,dt\,dw. \tag*{(27)}

The factor 1/B1/B in (25) is explained by the fair-split identity (21). Each finite prime set contributes a factor BB when its divisors are summed as a fair split. For two independent products with law ν1/2\nu_{1/2}, the first-moment calculation below will show that the smooth size-and-ratio expectation is a positive constant divided by BB, up to a smaller error. The two factors BB and the external division by BB then give a bounded positive mean in that model. The next reduction compares this model with the actual divisor weight and bounds the exceptional cases.

The moment target and the independent split model

Lemma 4.1 (Moments of the divisor weight). There exist C0>0C_0>0 and d0>0d_0>0 such that, for each fixed C∗≥C0C_* \ge C_0 and all sufficiently large BB,

∥DB∥L2(Z^)≪C∗1,∫Z^DB(w) dw≥d0.(28)\lVert D_B\rVert_{L^2(\widehat{\mathbb{Z}})}\ll_{C_*}1,\qquad\int_{\widehat{\mathbb{Z}}}D_B(w)\,dw\ge d_0. \tag*{(28)}

*The constant d0d_0 is independent of C∗C_* and BB. In addition, ∫DB≪1\int D_B\ll1 with a constant independent of C∗C_*. *

We prove this lemma after reducing the actual divisor weight to an independent model and establishing the concentration estimate needed for its second moment. A single split supplies the positive mean. Two splits of the same prime sets govern the second moment.

Let S1,S2S_1,S_2 be independent random subsets of P\mathcal{P}, each formed by including every prime pp independently with probability 1/p1/p. Conditional on these sets, split each SνS_\nu by independent fair coins into a coefficient subset CνC_\nu and a remaining subset RνR_\nu. Write

a=∏p∈C1p,c=∏p∈C2p,a=\prod_{p\in C_1}p,\qquad c=\prod_{p\in C_2}p,

with empty products equal to one, and define

D~B(S1,S2)=BEsplit[ρ0(log⁡cB)ψ(aTc)1all four subsets are regular∣S1,S2].(29)\widetilde{D}_B(S_1,S_2)=B\mathbb{E}_{\mathrm{split}}\left[\rho_0\left(\frac{\log c}{B}\right)\psi\left(\frac{a}{T_c}\right)\mathbf{1}_{\text{all four subsets are regular}}\mathrel{\Big|}S_1,S_2\right]. \tag*{(29)}

In particular, 0≤D~B≪B0\le\widetilde{D}_B\ll B pointwise.

Lemma 4.2 (Reduction to independent fair splits). For each fixed C∗≥0C_* \ge0, as B→∞B\to\infty,

∫Z^DB(w)r dw=ED~Br+o(1)(r=1,2).\int_{\widehat{\mathbb{Z}}}D_B(w)^r\,dw=\mathbb{E}\widetilde{D}_B^r+o(1)\qquad(r=1,2).

Proof. For Haar distributed ww, let

S1(w)={p∈P:p∣w},S2(w)={p∈P:p∣w+1}.S_1(w)=\{p\in\mathcal{P}:p\mid w\},\qquad S_2(w)=\{p\in\mathcal{P}:p\mid w+1\}.

At each prime these are disjoint hits, each of probability 1/p1/p. Their joint law differs by O(1/p2)O(1/p^2) in total variation from two independent Bernoulli(1/p)(1/p) indicators. Independence across primes therefore couples them to the independent sets S1,S2S_1,S_2 above, with total failure probability

O(∑p>P0p−2)=O(1/P0).O\left(\sum_{p>P_0}p^{-2}\right)=O(1/P_0).

The probability that any p∈Pp \in\mathcal{P} has p2∣wp^2 \mid w or p2∣w+1p^2 \mid w+1 is also O(1/P0)O(1/P_0).

On the complement of these exceptional events, coefficient divisors are subset products of the corresponding site sets, and the prime set of each quotient is exactly the complementary subset. By (21), a divisor and its complement then have untruncated weight B2−∣Sν∣B2^{-|S_\nu|}. The truncated weight inserts exactly the regularity indicators for those two subsets. Summing the two fair splits shows that the value of (25) is precisely (29) on this successful coupling.

We can discard the exceptional events in both moments. Any nonzero representation in (25) forces each entire site to have at most (1+2τ)ℓ(1+2\tau)\ell distinct primes, by the total regularity cutoffs for the coefficient and quotient. This remains true in the presence of squares: their two prime sets still cover the site’s distinct primes, although they need not be disjoint. Hence there are at most 22(1+2τ)ℓ2^{2(1+2\tau)\ell} coefficient choices across the two sites. Combining this with (3.9) gives, for example, the uniform bound DB≪B4D_B \ll B^4. Together with D~B≪B\widetilde D_B \ll B, the coupling and site-square errors change its first two moments by at most O(B8/P0)=o(1)O(B^8/P_0)=o(1). This proves the comparison.

The two-split question and an addition-product estimate

The second moment of (29) involves two conditionally independent fair splits of the same two site sets. Denote their coefficient products by (a(h),c(h))(a^{(h)},c^{(h)}), h=1,2h=1,2, and their coefficient and remaining prime sets by Cν(h),Rν(h)C_\nu^{(h)},R_\nu^{(h)}, ν=1,2\nu=1,2. Choose fixed closed intervals Iρ0⊂(1,2)I_{\rho_0}\subset(1,2) and Iψ⊂(1,2)I_\psi\subset(1,2) containing the supports of ρ0\rho_0 and ψ\psi, respectively. Let Hh\mathcal{H}_h be the size-and-ratio event

log⁡c(h)B∈Iρ0,a(h)Tc(h)∈Iψ.\frac{\log c^{(h)}}{B}\in I_{\rho_0}, \qquad\frac{a^{(h)}}{Tc^{(h)}}\in I_\psi.

and let Eh\mathcal{E}_h be Hh\mathcal{H}_h together with regularity of its four subsets. Since the smooth weights are bounded and nonnegative,

ED~B2≪B2P(E1∩E2)(30)\mathbb{E}\widetilde D_B^2 \ll B^2\mathbb{P}(\mathcal{E}_1\cap\mathcal{E}_2) \tag*{(30)}

It therefore suffices to prove the global probability bound P(E1∩E2)≪C∗B−2\mathbb{P}(\mathcal{E}_1\cap\mathcal{E}_2)\ll_{C_*}B^{-2}.

To prove this target, we will group the changes between the two splits by their largest logarithmic scale YY. A further unconditional coupling will replace the first coefficient and remaining classes by independent prime processes, with its error paid before conditioning. In that comparison model, the additions from a remaining set at primes with log⁡p≤Y\log p\leq Y are independent selections with probabilities 1/(4p)1/(4p), independent also of the data above YY. After fixing the retained coefficient primes and the additions at the other site, the second ratio window confines the logarithm of the recipient’s addition product to an interval of fixed length. A new prime with logarithm above Y/2Y/2 also forces that product logarithm above Y/2Y/2. The next estimate gives the resulting O(1/Y)O(1/Y) probability uniformly in the interval’s location. The second-moment proof will combine it with the retained first-remainder tail restrictions and the high assignment coins. The product itself need not be bounded by the largest allowed prime.

Lemma 4.3 (Addition-product concentration). Fix H>0H>0. Write L=log⁡P0L=\log P_0. For sufficiently large BB, let L≤Y≤8BL\leq Y\leq8B, put U=min⁡(Y,4B)U=\min(Y,4B), and form a random product QYQ_Y by including each prime P0<p≤eUP_0<p\leq e^U independently with probability 1/(4p)1/(4p). Uniformly over all intervals J⊂RJ\subset\mathbb{R} of length at most HH,

P(log⁡QY∈J∩[Y/2,∞))≪H1Y.\mathbb{P}(\log Q_Y\in J\cap[Y/2,\infty))\ll_H \frac{1}{Y}.

The implicit constant is independent of BB, YY, and JJ.

Proof. If the intersection is empty there is nothing to prove. Otherwise enclose it in an interval [v,v+H][v,v+H] with v≥Y/2v \ge Y/2, and set X=eYX=e^Y. The probability of a squarefree product nn of the permitted primes is

qYwY(n)n,qY=∏P0<p≤eU(1−14p),wY(n)=∏p∣n1/41−1/(4p).q_Y\frac{w_Y(n)}{n}, \qquad q_Y=\prod_{P_0<p\le e^U}\left(1-\frac{1}{4p}\right), \qquad w_Y(n)=\prod_{p\mid n}\frac{1/4}{1-1/(4p)}.

Set wY(n)=0w_Y(n)=0 for other integers. Since Y/2≤U≤YY/2\le U\le Y and U≥LU\ge L, Mertens’ estimate gives

qY≪(L/U)1/4≪(L/Y)1/4,qY≤1.q_Y\ll(L/U)^{1/4}\ll(L/Y)^{1/4}, \qquad q_Y\le1.

Let cs>0c_s>0 be an admissible exponent for the one-dimensional upper sieve of Section 3, with a remainder O(N.8)O(N^{.8}) on intervals of length NN. Choose a fixed c>0c>0 sufficiently small that c≤1/4c\le1/4 and 2c<cs2c<c_s, and put Z=ecYZ=e^{cY}. Then Z≤eUZ\le e^U and, since X≥eY/2X\ge e^{Y/2}, Z≤XcsZ\le X^{c_s}. Dropping squarefreeness, the restriction on prime factors greater than ZZ, and the reducing weights above ZZ majorizes wY(n)w_Y(n) by

w(n)=∏p≤min⁡(P0,Z)1p∤n∏P0<p≤Zap1p∣n,ap=14−1/p<1.w(n)=\prod_{p\le\min(P_0,Z)}1_{p\nmid n}\prod_{P_0<p\le Z}a_p^{1_{p\mid n}}, \qquad a_p=\frac{1}{4-1/p}<1.

Apply the random-weight form of that sieve on an interval containing [X,eHX][X,e^H X] and of length comparable to XX, with constants depending only on HH. It yields

∑X≤n≤eHXwY(n)≪HXD+X.8,\sum_{X\le n\le e^H X}w_Y(n)\ll_H XD+X^{.8},

where

D=∏p≤min⁡(P0,Z)(1−1/p)∏P0<p≤Z(1−1−app).D=\prod_{p\le\min(P_0,Z)}(1-1/p)\prod_{P_0<p\le Z}\left(1-\frac{1-a_p}{p}\right).

If Z≥P0Z\ge P_0, the relation 1−ap=3/4+O(1/p)1-a_p=3/4+O(1/p) and Mertens’ estimate imply

D≪c1L(LY)3/4.D\ll_c \frac{1}{L}\left(\frac{L}{Y}\right)^{3/4}.

Consequently qYD≪c1/Yq_YD\ll_c 1/Y. If Z<P0Z<P_0, ordinary rough exclusion instead gives D≪1/log⁡Z=1/(cY)D\ll1/\log Z=1/(cY), and the same conclusion follows from qY≤1q_Y\le1. Finally,

qYX−.2≤e−Y/10≪1Y.q_YX^{-.2}\le e^{-Y/10}\ll\frac{1}{Y}.

Using 1/n≤1/X1/n\le1/X on the summation interval proves

P(X≤QY≤eHX)≤qYX∑X≤n≤eHXwY(n)≪H1Y,\mathbb{P}(X\le Q_Y\le e^H X)\le\frac{q_Y}{X}\sum_{X\le n\le e^H X}w_Y(n)\ll_H \frac{1}{Y},

as required.

The first and second moments

Proof of Lemma 4.1. By Lemma 4.2, it suffices to prove the asserted mean and second-moment bounds for D‾B\overline{D}_B. The first moment. Without the regularity indicator in (29), the products a,ca,c are independent and each has law ν1/2\nu_{1/2}. Let fB=f1/2,Bf_B=f_{1/2,B} and f(s)=c1/2s−1/2f(s)=c_{1/2}s^{-1/2} on the compact positive log ranges under consideration. For fixed s=log⁡c/Bs=\log c/B on the support of ρ0\rho_0, (23) and partial summation give, uniformly there,

BEaψ(aTc)=∫0∞ψ(r)fB(s+log⁡T+log⁡rB)drr+o(1).B\mathbb{E}_a\psi\left(\frac{a}{Tc}\right)=\int_0^\infty\psi(r)f_B\left(s+\frac{\log T+\log r}{B}\right)\frac{dr}{r}+o(1).

To justify the use of (23), the relevant interval in log⁡a/B\log a/B has length O(1/B)O(1/B); its absolute distribution-function error is O(B−80)O(B^{-80}), and partial summation against the smooth window, whose total variation is bounded, leaves an error o(1/B)o(1/B) before the displayed multiplication by BB. Since log⁡T/B→0\log T/B\to0 and fBf_B converges uniformly with derivatives on these compact intervals, another application of (23) to cc shows that the untruncated first moment tends to

d∗=(∫0∞ρ0(s)f(s)2 ds)(∫0∞ψ(r)drr)>0.d_*=\left(\int_0^\infty\rho_0(s)f(s)^2\,ds\right)\left(\int_0^\infty\psi(r)\frac{dr}{r}\right)>0.

The same calculation, or its upper-bound version for fixed intervals containing the supports, shows that the size-and-ratio event H1\mathcal{H}_1 has probability O(1/B)O(1/B) under the independent coefficient-product law.

We next control the regularity losses in this normalization. Cover the coefficient support by O(B)O(B) dyadic-size boxes c≍Xc\asymp X, a≍TXa\asymp TX. All their logarithmic scales lie in a fixed compact subinterval of the ranges for the arithmetic estimates. From (3.5), the joint product probability at (a,c)(a,c) is bounded by

CA0(a)A0(c)B2ac.C\frac{A_0(a)A_0(c)}{B^2ac}.

Within one such box ac≍TX2ac\asymp TX^2. The independent-coefficient version of (3.11) bounds the untruncated weighted sum where either coefficient is nonregular by

O(TX2{o(1)+e−c3C∗}).O\left(TX^2\{o(1)+e^{-c_3C_*}\}\right).

After the harmonic denominators, summing the O(B)O(B) boxes, and multiplying by the BB in (29), the coefficient-regularity loss is therefore o(1)+O(e−c3C∗)o(1)+O(e^{-c_3C_*}).

Conditional on a selected coefficient subset at a site, the remaining indicators at primes not selected for that coefficient are independent, with probabilities

1/(2p)1−1/(2p)=12p−1.\frac{1/(2p)}{1-1/(2p)}=\frac{1}{2p-1}.

The selected coefficient primes have reciprocal sum O(B/P0)O(B/P_0), uniformly on the coefficient support, because the logarithm of their product is O(B)O(B). Thus the independent-indicator regularity estimate of Section 3 applies uniformly after these primes are omitted. The conditional chance of a remaining subset being nonregular is o(1)+O(e−c3C∗)o(1)+O(e^{-c_3C_*}). Multiplying by the bounded smooth weights and the coefficient-event probability O(1/B)O(1/B) shows that the remaining-subset loss in (29) is again o(1)+O(e−c3C∗)o(1)+O(e^{-c_3C_*}) after its outside factor BB.

It follows that

ED~B=d∗+O(e−c3C∗)+o(1).\mathbb{E}\widetilde{D}_B=d_*+O(e^{-c_3C_*})+o(1).

where the error constant in the exponential term does not depend on C∗C_*. This proves the asserted positive lower bound after choosing C0C_0 large, for instance with a fixed d0<d∗/2d_0<d_*/2 and then allowing the negligible coupling error. The upper bound ∫DB≪1\int D_B\ll1 follows by omitting all regularity indicators in the independent model. Both bounds can use constants independent of C∗C_*. For the lower bound one may also note that increasing C∗C_* only relaxes the tail cutoffs.

The second moment. We use the two-split notation introduced before Lemma 4.3. By (30), the required bound is the global probability estimate stated there.

Compare the two assignments of each site prime. If there is a change, its largest logarithm lies in one of the disjoint intervals (γ/2,γ](\gamma/2,\gamma], where

γ=2ilog⁡P0,1≤i≤imax⁡.\gamma= 2^i \log P_0,\qquad1 \le i \le i_{\max}.

Here 4B≤2imax⁡log⁡P0<8B4B \le2^{i_{\max}}\log P_0 < 8B for large BB. Write AγA_\gamma for the event that this is the largest changed interval. Every change in it is either a move from remaining to coefficient or a move in the reverse direction. The joint law and E1∩E2\mathcal{E}_1 \cap\mathcal{E}_2 are invariant under exchanging the splits. Hence

P(E1∩E2∩Aγ)≤2∑v=12P(E1∩E2∩Aγ∩Uγ,v).\mathbb{P}(\mathcal{E}_1 \cap\mathcal{E}_2 \cap A_\gamma) \le2 \sum_{v=1}^{2} \mathbb{P}(\mathcal{E}_1 \cap\mathcal{E}_2 \cap A_\gamma\cap U_{\gamma,v}).

where Uγ,vU_{\gamma,v} requires at least one prime at site vv with logarithm in (γ/2,γ](\gamma/2,\gamma] to move from the first remaining set to the second coefficient set. This orientation inequality has been taken in the original symmetric two-split law.

For the upper bound on its right side, couple the four first-split sets Cv(1),Rv(1)C_v^{(1)}, R_v^{(1)} to four mutually independent prime processes, each with inclusion probabilities 1/(2p)1/(2p). At a prime the two classes at one site were mutually exclusive, so the cost of this coupling is O(1/p2)O(1/p^2), and the total cost is O(1/P0)O(1/P_0). Assign independent fair second-split coins to each occurrence in these processes. On the successful, collision-free coupling this constructs exactly the original second split. On the exceptional event there may be two occurrences of a prime, for which we still assign independent coins and define coefficient products with multiplicity. This provides a convenient coupled model for upper bounds. Even if the coupling error is counted separately in every one of the O(log⁡B)O(\log B) intervals, its contribution after multiplication by B2B^2 is O(B2log⁡B/P0)=o(1)O(B^2\log B/P_0)=o(1). No uniform coupling assertion conditioned on a rare coefficient value is needed.

Work now in that independent model. Its first coefficient products have independent ν1/2\nu_{1/2} laws, so the first-moment calculation gives P(H1)=O(1/B)\mathbb{P}(\mathcal{H}_1)=O(1/B). Fix first coefficient sets satisfying H1\mathcal{H}_1 and the first coefficient regularity conditions. For each term on the oriented right side, we will retain the needed first-remainder restrictions while bounding its second ratio condition. The sum of these conditional bounds over the changed scales, together with the no-change contribution, must be OC∗(1/B)O_{C_*}(1/B) uniformly in the fixed coefficient sets. The orientation factor and coupling error remain outside this conditional calculation.

For a prime set SS let NS(>γ)=#{p∈S:log⁡p>γ}N_S(> \gamma)=\#\{p\in S:\log p>\gamma\}, and put

q(γ)=.4log⁡(B/γ)−C∗.q(\gamma)=.4\log(B/\gamma)-C_*.

The first coefficient tail cutoffs give NCv(1)(>γ)≥q(γ)N_{C_v^{(1)}}(>\gamma)\ge q(\gamma). Expose the first remaining sets above γ\gamma, and retain the two necessary tail restrictions NRv(1)(>γ)≥q(γ)N_{R_v^{(1)}}(>\gamma)\ge q(\gamma). Conditional on all these first high-prime sets, the chance that none of their second assignments changes is exactly

2−∑v=12{NCv(1)(>γ)+NRv(1)(>γ)}.2^{-\sum_{v=1}^{2}\{N_{C_v^{(1)}}(>\gamma)+N_{R_v^{(1)}}(>\gamma)\}}.

On the retained restrictions this is at most

2−4q(γ)=24C∗(γ/B)α0,α0=1.6log⁡2>1.2^{-4q(\gamma)}=2^{4C_*}(\gamma/B)^{\alpha_0},\qquad\alpha_0=1.6\log2>1.

If q(γ)<0q(\gamma)<0 this upper bound exceeds one, which remains a valid bound. Averaging over the first remaining high-prime sets does not increase it.

All first remaining restrictions below YY may now be omitted. For each site, a prime below this threshold is newly added to the second coefficient exactly when it lies in the independent first remaining process and its second coin selects the coefficient. These new-addition indicators are independent and have probabilities 1/(4p)1/(4p). The new additions at the two sites are independent of one another, of the first coefficient sets and their second coins, and of all the exposed variables above YY.

Fix a recipient site ν\nu. Condition on the retained first-coefficient primes at both sites and on the additions at the other site. The second ratio condition in H2\mathcal{H}_2 then restricts the logarithm of the addition product at site ν\nu to an interval of length at most

Hψ=log⁡(sup⁡Iψinf⁡Iψ).H_{\psi}=\log\left(\frac{\sup I_{\psi}}{\inf I_{\psi}}\right).

For example, for recipient site 1 the second coefficient products have the form a(2)=aretbnewa^{(2)}=a_{\mathrm{ret}}b_{\mathrm{new}} and c(2)=cretcnewc^{(2)}=c_{\mathrm{ret}}c_{\mathrm{new}}, and solving a(2)/(Tc(2))∈Iψa^{(2)}/(T c^{(2)})\in I_{\psi} gives precisely such an interval for log⁡bnew\log b_{\mathrm{new}}. The same calculation with the inequality inverted applies at site 2. On UY,ν\mathcal{U}_{Y,\nu} there is a new prime with logarithm greater than Y/2Y/2, so this addition product also has logarithm at least Y/2Y/2. There are no new additions above YY on AYA_Y.

We can drop the second coefficient size restriction and all second regularity restrictions for this upper bound. Lemma 4.3 then bounds the conditional chance of the remaining addition-product conditions by O(1/Y)O(1/Y), uniformly in every value on which we have conditioned. To make the conditioning order explicit, write C\mathcal{C} for the fixed first coefficient sets, GYG_Y for the retained first remaining tail restrictions, and HYH_Y for the event of no change above YY. Let F\mathcal{F} include C\mathcal{C} and reveal all second coins of the first coefficient sets, all first remaining data and second coins above YY, and the additions at the other site. The recipient addition product QνQ_\nu below YY is independent of this information, and the ratio condition specifies an F\mathcal{F}-measurable interval J(F)J(\mathcal{F}). For the oriented event under consideration, the tower property gives

P(oriented event∣C)≤E[1GY1HYP(log⁡Qν∈J(F)∩[Y/2,∞)∣F)∣C]≪1YE[1GY1HY∣C]≤2−4q(Y)Y.\begin{aligned} \mathbb{P}(\text{oriented event}\mid\mathcal{C}) &\leq\mathbb{E}\left[1_{G_Y}1_{H_Y}\mathbb{P}\left(\log Q_\nu\in J(\mathcal{F})\cap[Y/2,\infty)\mid\mathcal{F}\right)\mid\mathcal{C}\right]\\ &\ll\frac{1}{Y}\mathbb{E}\left[1_{G_Y}1_{H_Y}\mid\mathcal{C}\right]\leq\frac{2^{-4q(Y)}}{Y}. \end{aligned}

In the last step the high second coins are averaged after the uniform small-ball bound is applied; no conditional bound on P(HY∣F)\mathbb{P}(H_Y\mid\mathcal{F}) is asserted. Thus, conditional on any good first coefficient sets, the contribution for this orientation at scale YY is at most

OC∗((Y/B)α0Y).O_{C_*}\left(\frac{(Y/B)^{\alpha_0}}{Y}\right).

The constants here may depend on the fixed C∗C_*, but not on the first coefficient sets, BB, or YY.

If there is no change at all, use the tail cutoff at Y=log⁡P0Y=\log P_0. All four first subsets then have at least .4log⁡P0−C∗.4\log P_0-C_* primes. The same calculation bounds the conditional probability, retaining the first remaining tail restrictions, by

OC∗(R−α0).O_{C_*}(R^{-\alpha_0}).

The dyadic scales satisfy

∑Y(Y/B)α0Y=B−α0∑YYα0−1≪1B.\sum_Y \frac{(Y/B)^{\alpha_0}}{Y} = B^{-\alpha_0}\sum_Y Y^{\alpha_0-1} \ll\frac{1}{B}.

because α0>1\alpha_0 > 1 and the largest YY is less than 8B8B. Moreover,

BR−α0=B1−α0(1000log⁡B)α0⟶0.BR^{-\alpha_0}=B^{1-\alpha_0}(1000\log B)^{\alpha_0}\longrightarrow0.

Requiring first coefficient regularity only decreases the O(1/B)O(1/B) mass of H1\mathcal{H}_1 established above. Combining the preceding bounds, including the orientation factor and the coupling errors, gives

P(E1∩E2)≪C∗1B(1B+R−α0)+O(log⁡B/P0)≪C∗1B2+O(log⁡B/P0).\mathbb{P}(\mathcal{E}_1\cap\mathcal{E}_2)\ll_{C_*}\frac{1}{B}\left(\frac{1}{B}+R^{-\alpha_0}\right)+O(\log B/P_0)\ll_{C_*}\frac{1}{B^2}+O(\log B/P_0).

Multiplication by B2B^2 proves ED~B2≪C∗1\widetilde{E D}_B^2\ll_{C_*}1. Lemma 4.2 transfers these bounds to the Haar divisor weight and completes the proof of (28).

The correlation survives the divisor weights

Lemma 4.4 (Preservation of the profinite profile). For each fixed C∗≥C0C_* \ge C_0, put H(t,w)=ϕ(t)W(t,w)H(t,w)=\phi(t)W(t,w). Then

∫H(t,w)DB(w) dt dw−(∫H dt dw)(∫DB dw)⟶0(B→∞).\int H(t,w)D_B(w)\,dt\,dw-\left(\int H\,dt\,dw\right)\left(\int D_B\,dw\right)\longrightarrow0\qquad(B\to\infty).

Proof. The function HH is bounded and supported on a fixed compact interval in tt, so it is square-integrable. Conditional expectations with respect to tt and a finite residue coordinate w mod qw\bmod q approximate HH in L2L^2 as these residue coordinates increase. Equivalently, for every ε>0\varepsilon>0 there is a function Hq(t,w)H_q(t,w) depending only on tt and w mod qw\bmod q such that

∥H−Hq∥2<ε,∫Hq dt dw=∫H dt dw.\lVert H-H_q\rVert_2<\varepsilon,\qquad\int H_q\,dt\,dw=\int H\,dt\,dw.

This follows, for example, by first approximating in the dense space of finite sums of products of L2L^2 functions of tt and locally constant functions of ww, and then taking conditional expectation.

For all sufficiently large BB, no prime dividing qq belongs to P\mathcal{P}. The finite-residue function DBD_B is therefore independent of w mod qw\bmod q under Haar measure. Consequently

∫Hq(t,w)DB(w) dt dw=(∫H dt dw)(∫DB dw).\int H_q(t,w)D_B(w)\,dt\,dw=\left(\int H\,dt\,dw\right)\left(\int D_B\,dw\right).

By Cauchy–Schwarz and (28), the error on replacing HqH_q by HH is OC∗(ε)O_{C_*}(\varepsilon), uniformly in large BB; the length of the fixed tt-support is absorbed in the constant. Since ε\varepsilon is arbitrary, this proves (4.7). □

The factor ∫DB dw\int D_B\,dw in (4.7) is real, nonnegative, and at least d0d_0 eventually. Taking complex absolute values in (27) and using (4.7) therefore gives

lim inf⁡B→∞lim⁡x→∞∣I1,x∣B≥d0β∗.(31)\liminf_{B\to\infty}\lim_{x\to\infty}\frac{|I_{1,x}|}{B}\ge d_0\beta_*. \tag*{(31)}

The possible dependence of the L2L^2 constant on C∗C_* is harmless: this approximation is performed with C∗C_* fixed, and the positive lower constant in (31) is independent of it.

A positive quadratic energy and its diagonal

The preserved correlation now has an expansion suited to Cauchy–Schwarz. For c∣am+1c \mid am+1, put la=(am+1)/cl_a=(am+1)/c. At fixed BB, the coefficient list in (25) is finite. Thus gx(am)=gx(m)g_x(am)=g_x(m) holds simultaneously for every coefficient aa on that list once xx is sufficiently large, by (2.6). Expanding (26), writing n=amn=am, and using this fixed-multiplier invariance gives the exact identity

I1,x=1x∑c≥1ρ0(log⁡cB)A(c)∑m≥1K(m)gx(m)‾∑a≥1c∣am+1A(a)K(la)ψ(aTc)ϕ(amx)Fx(am+1).(32)I_{1,x}=\frac{1}{x}\sum_{c\ge1}\rho_0\left(\frac{\log c}{B}\right)A(c)\sum_{m\ge1}K(m)\overline{g_x(m)}\sum_{\substack{a\ge1\\c\mid am+1}}A(a)K(l_a)\psi\left(\frac{a}{Tc}\right)\phi\left(\frac{am}{x}\right)F_x(am+1). \tag*{(32)}

The values of FxF_x are unchanged in this reindexing. The expanded form places gx(m)g_x(m) outside the inner sum, where ∣gx(m)∣=1|g_x(m)|=1 supplies the unit factor in Cauchy–Schwarz.

Proposition 4.5. There is c5>0c_5>0, independent of all sufficiently large fixed C∗C_\ast, for which the nonnegative energy below satisfies the displayed lower bound:

I2,x=1x∑c≥1ρ0(log⁡c/B)A(c)∑m≥1K(m)∣∑a≥1c∣am+1A(a)K(la)ψ(aTc)ϕ(amx)Fx(am+1)∣2,(33)I_{2,x}=\frac{1}{x}\sum_{c\ge1}\rho_0(\log c/B)A(c)\sum_{m\ge1}K(m)\left|\sum_{\substack{a\ge1\\c\mid am+1}}A(a)K(l_a)\psi\left(\frac{a}{Tc}\right)\phi\left(\frac{am}{x}\right)F_x(am+1)\right|^2, \tag*{(33)}
lim inf⁡B→∞lim inf⁡x→∞I2,xBT≥c5.\liminf_{B\to\infty}\liminf_{x\to\infty}\frac{I_{2,x}}{BT}\ge c_5.

Its diagonal contribution is negligible:

lim⁡B→∞lim sup⁡x→∞I2,xdiagBT=0.\lim_{B\to\infty}\limsup_{x\to\infty}\frac{I_{2,x}^{\mathrm{diag}}}{BT}=0.

Here I2,xdiagI_{2,x}^{\mathrm{diag}} denotes the terms with equal coefficient indices on expanding the square, and la=(am+1)/cl_a=(am+1)/c as in (32).

Proof. Choose 0<αϕ<βϕ<∞0<\alpha_\phi<\beta_\phi<\infty with supp⁡ϕ⊂[αϕ,βϕ]\operatorname{supp}\phi\subset[\alpha_\phi,\beta_\phi], and write u−=inf⁡Iψu_-=\inf I_\psi, u+=sup⁡Iψu_+=\sup I_\psi. If the inner sum in (32) is nonzero, then

αϕxu++Tc≤m≤βϕxu−−Tc.\frac{\alpha_\phi x}{u_++Tc}\le m\le\frac{\beta_\phi x}{u_--Tc}.

Let Ic,x\mathcal{I}_{c,x} denote this interval. Use the nonnegative weights x−1ρ0(log⁡c/B)A(c)K(m)x^{-1}\rho_0(\log c/B)A(c)K(m) on this interval. Cauchy–Schwarz gives

∣I1,x∣2≤VB,xI2,x,VB,x=1x∑cρ0(log⁡cB)A(c)∑m∈Ic,xK(m).|I_{1,x}|^2\le V_{B,x}I_{2,x},\qquad V_{B,x}=\frac{1}{x}\sum_c\rho_0\left(\frac{\log c}{B}\right)A(c)\sum_{m\in\mathcal{I}_{c,x}}K(m).

where ∣gx(m)∣=1|g_x(m)|=1 supplies the unit factor. No multiplicativity of the centered function FxF_x is needed.

For fixed BB, the weight K0K_0 is periodic and has mean

∫ZK0(w) dw=R1/2∏p∈P(1−12p)≪1.\int_{\mathbb{Z}}K_0(w)\,dw=R^{1/2}\prod_{p\in\mathcal{P}}\left(1-\frac{1}{2p}\right)\ll1.

Progression counting, K≤K0K \le K_0, and the length of Ic,x\mathcal{I}_{c,x} therefore imply

lim sup⁡x→∞VB,x≪1T∑cρ0(log⁡cB)A0(c)c≪BT.\limsup_{x\to\infty} V_{B,x} \ll\frac{1}{T}\sum_c \rho_0\left(\frac{\log c}{B}\right)\frac{A_0(c)}{c} \ll\frac{B}{T}.

For the last bound, partition the compact log support of ρ0\rho_0 into O(B)O(B) dyadic-size boxes and apply (3.6) in each box. Thus we may fix C4>0C_4>0, independently of C∗C_\ast, such that

lim sup⁡x→∞VB,x≤C4B/T\limsup_{x\to\infty} V_{B,x} \le C_4B/T

for all sufficiently large BB. Combining this with (31) proves (33), with any fixed

0<c5<(d0β∗)2C4.0<c_5<\frac{(d_0\beta_\ast)^2}{C_4}.

This makes the uniformity of the positive lower constant explicit.

For the diagonal, set

QB=sup⁡A(a)K(l),Q_B=\sup A(a)K(l),

where aa ranges over the coefficient support and ll over positive integers for which the truncated weights are nonzero. The total-count cutoff and (3.9) give

QB≪B2(h0+τlog⁡2)+o(1)=o(T),Q_B\ll B^{2(h_0+\tau\log2)+o(1)}=o(T),

using the strict inequality 2(h0+τlog⁡2)<.322(h_0+\tau\log2)<.32. Since all smooth weights here are nonnegative,

ψ(a/(Tc))2ϕ(am/x)2≤∥ψ∥∞∥ϕ∥∞ψ(a/(Tc))ϕ(am/x).\psi(a/(Tc))^2\phi(am/x)^2\le\|\psi\|_\infty\|\phi\|_\infty\psi(a/(Tc))\phi(am/x).

Bounding one of the two factors A(a)K(la)A(a)K(l_a) by QBQ_B and ∣Fx(am+1)∣2|F_x(am+1)|^2 by 4 yields

I2,xdiag≪QBJB,x,I^{\mathrm{diag}}_{2,x}\ll Q_B\mathcal{J}_{B,x},

where

JB,x=1x∑cρ0(log⁡cB)A(c)∑m≥1K(m)∑a≥1c∣am+1A(a)K(la)ψ(aTc)ϕ(amx).\mathcal{J}_{B,x}=\frac{1}{x}\sum_c\rho_0\left(\frac{\log c}{B}\right)A(c)\sum_{m\ge1}K(m)\sum_{\substack{a\ge1\\c\mid am+1}}A(a)K(l_a)\psi\left(\frac{a}{Tc}\right)\phi\left(\frac{am}{x}\right).

The same divisor reindexing as before gives

JB,x=Bx∑n≥1ϕ(n/x)DB(n).\mathcal{J}_{B,x}=\frac{B}{x}\sum_{n\ge1}\phi(n/x)D_B(n).

Here no label is present, so finite progression counting directly gives

lim⁡x→∞JB,x=B(∫0∞ϕ(t) dt)(∫ZDB(w) dw)≪B\lim_{x\to\infty}\mathcal{J}_{B,x} =B\left(\int_0^\infty\phi(t)\,dt\right)\left(\int_{\mathbb{Z}}D_B(w)\,dw\right)\ll B

by the first-moment upper bound in Lemma 4.1. It follows that

lim sup⁡x→∞I2,xdiagBT≪QBT⟶0,\limsup_{x\to\infty}\frac{I^{\mathrm{diag}}_{2,x}}{BT}\ll\frac{Q_B}{T}\longrightarrow0,

which completes the proof.

From mixed amplification to an independent-site kernel

The mixed amplification has eliminated the unit-modulus label gxg_x by Cauchy–Schwarz. Its positive square I2,xI_{2,x} contains only the centered label FxF_x. We first turn its off-diagonal terms into edges between integers at small additive distance. We then separate the residue data at the endpoints from the additional residue data in each representation of an edge. Throughout this section, all parameters other than BB and xx are fixed; xx tends to infinity first. Constants are uniform when the smooth scale variable belongs to the compact range specified below.

The change of variables

Consider distinct coefficients a,ba,b in the inner square defining I2,xI_{2,x}, for fixed m,cm,c, and put la=(am+1)/cl_a=(am+1)/c and lb=(bm+1)/cl_b=(bm+1)/c. The congruences imply that mm is a unit modulo cc and that

a−b=jc,0<∣j∣≤T.a-b=jc,\qquad0<|j|\le T.

They also imply (a,c)=(b,c)=1(a,c)=(b,c)=1. Since a,ba,b are rough and ∣j∣≤T<P0|j|\le T<P_0, we have (a,j)=(b,j)=1(a,j)=(b,j)=1, and hence (a,b)=1(a,b)=1. For the ordered term in which the aa-factor is conjugated, set

n′=bla,n′+j=alb.n'=bl_a,\qquad n'+j=al_b.

For these fixed coefficients, this is a bijective reindexing by the conditions b∣n′b\mid n' and a∣n′+ja\mid n'+j. Indeed, write n′=bzn'=bz. The latter congruence, together with b=a−jcb=a-jc, gives j(1−cz)≡0(moda)j(1-cz)\equiv0\pmod a. Thus

m=cz−1am=\frac{cz-1}{a}

is an integer, and substitution recovers both original congruences and both displayed identities. Positive compact scale supports ensure all arguments are positive for large xx.

Figure 1 records the two fixed multiplications that produce the additive shift. Only fixed-multiplier invariance is used; the centered function FxF_x need not be multiplicative.

The exact reindexing of two divisor representations

Figure 1. The exact reindexing of two divisor representations. The coefficient identity makes the new endpoints differ by jj, while fixed-multiplier invariance preserves their centered labels.

For fixed BB, the coefficient supports are finite, so (2.6) applies simultaneously to the multipliers c,b,ac,b,a once xx is sufficiently large. Removing cc and then inserting b,ab,a gives

Fx(am+1)‾Fx(bm+1)‾=Fx(la)‾Fx(lb)‾=Fx(n′)‾Fx(n′+j)‾.\overline{F_x(am+1)}\overline{F_x(bm+1)}=\overline{F_x(l_a)}\overline{F_x(l_b)}=\overline{F_x(n')}\overline{F_x(n'+j)}.

Define

Ψs(u1,u2)=ψ(u1)ψ(u2)ϕ(s/u1)ϕ(s/u2).\Psi_s(u_1,u_2)=\psi(u_1)\psi(u_2)\phi(s/u_1)\phi(s/u_2).

If s=n′/(Tx)s=n'/(Tx), u1=a/(Tc)u_1=a/(Tc) and u2=b/(Tc)u_2=b/(Tc), then

s/u1=bm/x+b/(ax),s/u2=am/x+1/x.s/u_1=bm/x+b/(ax),\qquad s/u_2=am/x+1/x.

The two original ϕ\phi-arguments are therefore interchanged, with errors OB(1/x)O_B(1/x). Since they occur as a product, their replacement by Ψs\Psi_s has vanishing error for fixed BB. There is a fixed compact interval S⊂(0,∞)S \subset(0,\infty) outside which Ψs\Psi_s vanishes on the coefficient supports. It depends only on ϕ,ψ\phi,\psi.

Consequently, the off-diagonal part of I2,x/(BT)I_{2,x}/(BT) is, up to ox(1)o_x(1),

1Tx∑n′∑0<∣j∣≤TFx(n′)‾Fx(n′+j)Kj(n′,n′/(Tx)),(34)\frac{1}{Tx}\sum_{n'}\sum_{0<|j|\leq T}\overline{F_x(n')}F_x(n'+j)\mathcal{K}_j(n',n'/(Tx)), \tag*{(34)}

where, for u∈Z^u \in\widehat{\mathbb{Z}},

Kj(u,s)=1B∑a−b=jc, (a,b)=(a,c)=(b,c)=1b∣u, a∣u+jA(a)A(b)A(c)K(u/b)K((u+j)/a)(c(u/b)−1a)ρ0(log⁡c/B)Ψs(a/(Tc),b/(Tc)).(35)\mathcal{K}_j(u,s)=\frac{1}{B}\sum_{\substack{a-b=jc,\ (a,b)=(a,c)=(b,c)=1\\ b\mid u,\ a\mid u+j}}A(a)A(b)A(c)K(u/b)K((u+j)/a) \left(\frac{c(u/b)-1}{a}\right)\rho_0(\log c/B)\Psi_s(a/(Tc),b/(Tc)). \tag*{(35)}

All coefficient variables are positive integers. Divisibility in Z^\widehat{\mathbb{Z}} means membership in the image of multiplication by the divisor; that multiplication is injective. The preceding congruence calculation also proves that every quotient in (35) is defined on its summation domain. This kernel is nonnegative, depends on finitely many prime-power residues, and satisfies

Kj(u,s)=K−j(u+j,s).\mathcal{K}_j(u,s)=\mathcal{K}_{-j}(u+j,s).

For this last identity, exchange a,ba,b: the last quotient is unchanged because c(u+j)−a=cu−bc(u+j)-a=cu-b.

Lemma 5.1 (Mean edge mass). Uniformly for 0<∣j∣≤T0<|j|\leq T and s∈Ss\in S,

Eu∈Z^Kj(u,s)≪Σ(j)T.(36)\mathbb{E}_{u\in\widehat{\mathbb{Z}}}\mathcal{K}_j(u,s)\ll\frac{\Sigma(j)}{T}. \tag*{(36)}

The implied constant can be independent of the regularity parameter C∗C_*.

Proof. Replace all truncated weights by their untruncated majorants. For fixed pairwise coprime a,b,ca,b,c, the two endpoint congruences have Haar probability 1/(ab)1/(ab). At a prime p∈Pp\in\mathcal{P} not dividing abcabc, the three residue weights have distinct zero classes 0,−j,b/c0,-j,b/c for uu. Here p>∣j∣p>|j|, and b/c≢−j(modp)b/c\not\equiv-j\pmod p follows from a=b+jca=b+jc. Their local mean is 1−3/(2p)1-3/(2p). At primes dividing abcabc discard the local reducing factors; this changes the Euler-product bound by at most exp⁡(O(∑p∣abc1/p))=O(1)\exp(O(\sum_{p\mid abc}1/p))=O(1), since that sum is O(B/P0)O(B/P_0). Independence over primes and Mertens’ formula give

R3/2∏p∈P(1−3/(2p))≪1.R^{3/2}\prod_{p\in\mathcal{P}}(1-3/(2p))\ll1.

On a dyadic box c≍Xc\asymp X, the denominator is ab≪T2X2ab\ll T^2X^2. The three-form estimate (3.8) bounds its contribution before the factor 1/B1/B by O(Σ(j)/T)O(\Sigma(j)/T). There are O(B)O(B) boxes. This proves the assertion using only untruncated upper bounds.

Blocks and the norm used for comparison

Fix 0<η<10<\eta<1 and C6>1C_6>1, and set M=⌈C6T⌉M=\lceil C_6T\rceil. Removing the lags ∣j∣<ηT|j|<\eta T in (34) costs O(η)+ox(1)O(\eta)+o_x(1) in absolute upper limit, by Lemma 5.1 and the bounded ordinary averages of Σ\Sigma. Average the origins over i=1,…,Mi=1,\ldots,M, writing n′=n+in'=n+i. For fixed BB, replacing (n+i)/(Tx)(n+i)/(Tx) by n/(Tx)n/(Tx) has vanishing error. At lag jj, the fraction of ii for which k=i+j∉[1,M]k=i+j\notin[1,M] is at most ∣j∣/M|j|/M. The total cost of these endpoints is O(T/M)=O(1/C6)O(T/M)=O(1/C_6). The remaining expression is

QB,xarith:=1Tx∑n1M∑1≤i,k≤MηT≤∣k−i∣≤TFx(n+i)‾Fx(n+k)Kk−i(n+i,n/(Tx)).(37)\mathcal{Q}^{\mathrm{arith}}_{B,x}:=\frac{1}{Tx}\sum_n\frac{1}{M}\sum_{\substack{1\leq i,k\leq M\\ \eta T\leq|k-i|\leq T}}\overline{F_x(n+i)}F_x(n+k)K_{k-i}(n+i,n/(Tx)). \tag*{(37)}

Thus (33) and the diagonal estimate give a positive lower bound for the iterated lower limit of (37), once η\eta is small enough and C6C_6 is large enough. These choices are made after C∗C_\ast and before letting BB grow.

Call a pair (i,k)(i,k) allowed if 1≤i,k≤M1\leq i,k\leq M and ηT≤∣k−i∣≤T\eta T\leq|k-i|\leq T. All matrices below are zero on other pairs. For a real matrix HH, define

∥H∥□=1Mmax⁡ϵ,δ∈{−1,1}M∣∑i,kHikϵiδk∣.(38)\lVert H\rVert_{\Box}=\frac{1}{M}\max_{\epsilon,\delta\in\{-1,1\}^{M}}\left|\sum_{i,k}H_{ik}\epsilon_i\delta_k\right|. \tag*{(38)}

The same maximum results from allowing real test coordinates in [−1,1][-1,1]. Splitting real and imaginary parts shows that for complex ziz_i with ∣zi∣≤2|z_i|\leq2,

∣1M∑i,kHikzizk‾∣≤16∥H∥□.(39)\left|\frac{1}{M}\sum_{i,k}H_{ik}z_i\overline{z_k}\right|\leq16\lVert H\rVert_{\Box}. \tag*{(39)}

For fixed BB, a matrix whose entries are finite-residue functions continuous in ss has the same property after taking this finite maximum. Its ordinary scale averages converge to its Haar-product integral. Therefore bounds on the Haar expectation of the norm, uniformly for s∈Ss\in\mathcal{S}, control errors in (37). This comparison is deterministic in the labels; it requires no independence between the labels and the residue data.

Endpoint types and candidate representations

The block energy is now expressed in the norm that will control its approximation. We next separate the prime-divisibility data at the endpoints from the additional residue data belonging to each edge representation.

Put Si(u)={p∈P:p∣u+i}S_i(u)=\{p\in\mathcal{P}:p\mid u+i\}. We compare them with independent sets SiS_i, each formed by independently including every p∈Pp\in\mathcal{P} with probability 1/p1/p. At a fixed p>P0>Mp>P_0>M, the actual law has disjoint hits of probability 1/p1/p at each site. Its total variation distance from independent hits is O(M2/p2)O(M^2/p^2). Consequently the joint site laws can be coupled with failure probability O(M2/P0)O(M^2/P_0). Also

P{p2∣u+i for some p∈P, i≤M}≪M/P0.(40)\mathbb{P}\{p^2\mid u+i\text{ for some }p\in\mathcal{P},\ i\leq M\}\ll M/P_0. \tag*{(40)}

For an allowed pair i,k=i+ji,k=i+j with j>0j>0, a candidate consists of a subset product bb of SiS_i, a subset product aa of SkS_k, and the positive integer c=(a−b)/jc=(a-b)/j. Require pairwise coprimality, rough squarefreeness, all three coefficient regularity conditions, and regularity of the two remaining sets Si∖bS_i\setminus b and Sk∖aS_k\setminus a. Here S∖bS\setminus b means removal of the primes dividing bb. It is harmless to use the closed support windows

B≤log⁡c≤2B,Tc≤a,b≤2Tc;B\leq\log c\leq2B,\qquad Tc\leq a,b\leq2Tc;

the smooth factors still impose the original supports. A reversed ordered pair refers to the same candidate.

A contributing endpoint has at most (1+2τ)ℓ(1+2\tau)\ell primes. There are at most 2(1+2τ)ℓ2^{(1+2\tau)\ell} subset choices at a pair, so the number of candidates in the entire block is O(Bℓ)O(B^\ell). This rough bound also applies to representations that contribute to (35) in the presence of site squares: the union of the coefficient and quotient prime sets covers the site’s distinct primes, and their total-count restrictions give the same bound. Together with (3.9), this gives, for example, a deterministic O(B10)O(B^{10}) bound for all total matrix masses used below. This bound allows small-probability errors to be discarded before sharper mean estimates are available.

Set kB=ESK(S)k_B=\mathbb{E}_S K(S) under the independent 1/p1/p law. The untruncated Euler product gives kB≪1k_B\ll1. Define the latent kernel

Lik(Si,Sk;s)=kBB∑candidatesA(a)A(b)A(c)K(Si∖b)K(Sk∖a)ρ0(log⁡c/B)Ψs(a/(Tc),b/(Tc))(41)\mathcal{L}_{ik}(S_i,S_k;s)=\frac{k_B}{B}\sum_{\text{candidates}} A(a)A(b)A(c)K(S_i\setminus b)K(S_k\setminus a)\rho_0(\log c/B)\Psi_s\left(a/(Tc),b/(Tc)\right) \tag*{(41)}

and extend it by symmetry and by zero on nonpairs. For actual endpoint sets without site squares, the terms allowed by the smooth factors have the same coefficient representations in (35) and (41), and the two endpoint quotient prime sets are the complementary sets displayed in (41). The latter kernel keeps these terms and replaces the remaining factor K(m)K(m) by its independent mean kBk_B.

Lemma 5.2 (Uniform conditional means). For every possible endpoint type SiS_i, not merely almost every typical type, and every allowed lag j=k−ij=k-i,

ESkLik(Si,Sk;s)≪Σ(j)T.(42)\mathbb{E}_{S_k}\mathcal{L}_{ik}(S_i,S_k;s)\ll\frac{\Sigma(j)}{T}. \tag*{(42)}

In particular, every conditional expected degree is O(1)O(1), and E∑i,kLik=O(M)\mathbb{E}\sum_{i,k}\mathcal{L}_{ik}=O(M).

Proof. For an admissible fair split bb at the first endpoint, A(b)K(Si∖b)=B2−∣Si∣A(b)K(S_i\setminus b)=B2^{-|S_i|}. At the other endpoint, the analogous identity and then averaging over its site set give the split-product law ν1/2\nu_{1/2}. By (3.5), its mass at aa is at most CA0(a)/(Ba)CA_0(a)/(Ba). Dropping all remaining restrictions for an upper bound yields

ESkLik(Si,Sk;s)≪Eb∣Si,split⁡1{b in the possible range}∑c≍b/T, a=b+jc≍bB≤log⁡c≤2BA0(c)A0(a)a≪Σ(j)TEb∣Si,split⁡1{b in the possible range}≤CΣ(j)T,\begin{aligned} \mathbb{E}_{S_k}\mathcal{L}_{ik}(S_i,S_k;s)\ll\mathbb{E}_{b\mid S_i,\operatorname{split}}\mathbf{1}_{\{b\text{ in the possible range}\}}\sum_{\substack{c\asymp b/T,\ a=b+jc\asymp b\\ B\le\log c\le2B}}\frac{A_0(c)A_0(a)}{a} \\ &\ll\frac{\Sigma(j)}{T}\mathbb{E}_{b\mid S_i,\operatorname{split}}\mathbf{1}_{\{b\text{ in the possible range}\}}\le C\frac{\Sigma(j)}{T}, \end{aligned}

using (3.7). This is a restricted expectation of mass at most one, with no division by its probability. If the fixed type has no admissible split, its kernel is zero. The same estimate with signed lag reversed handles either endpoint. Summing the bounded averages of Σ\Sigma proves the degree assertions. □\square

Separating the auxiliary roots

The following comparison justifies this averaging in expected cut norm, uniformly against a proposed endpoint kernel. It is stronger than an estimate against any one predetermined pair of tests.

Proposition 5.3 (Independent-root comparison). Let Mik(Si,Sk;s)M_{ik}(S_i,S_k;s) be a symmetric real kernel, zero on nonpairs, with total absolute mass O(B10)O(B^{10}) uniformly in its arguments. Then

Eu∥(Kk−i(u+i,s)−Mik(Si(u),Sk(u);s))i,k∥□≤o(1)+E(Si) indep∥(Lik(Si,Sk;s)−Mik(Si,Sk;s))i,k∥□.(43)\mathbb{E}_u\left\|\left(K_{k-i}(u+i,s)-M_{ik}(S_i(u),S_k(u);s)\right)_{i,k}\right\|_{\square} \le o(1)+\mathbb{E}_{(S_i)\ \mathrm{indep}}\left\|\left(\mathcal{L}_{ik}(S_i,S_k;s)-M_{ik}(S_i,S_k;s)\right)_{i,k}\right\|_{\square}. \tag*{(43)}

The error is uniform for s∈Ss \in\mathcal{S}.

Proof. In the absence of site squares, quotient prime sets at the endpoints equal the complementary sets used in (41). For a candidate at ii, k=i+jk=i+j, the remaining argument in (35) is

m=c(u+i)−bab=cab(u−r),r=b/c−i=a/c−k.m=\frac{c(u+i)-b}{ab}=\frac{c}{ab}(u-r),\qquad r=b/c-i=a/c-k.

We show that, at a negligible total error, the set of primes dividing this argument can be replaced for each unordered candidate by an independent 1/p1/p set, independent of the site sets and of the other candidates’ extra sets. The following argument first bounds two obstructions determined by the endpoint sets: coincident rational roots and a root forced to hit a third occupied site. It then fixes a surviving actual site configuration and couples the remaining prime tests, without conditioning on absence of site squares. The site-square event is paid for in the matrix weights, and the final step controls the independent extra-weight fluctuations in cut norm.

Coincident rational roots. Since b/cb/c is reduced, equal rr for two candidates forces the same denominator cc. Given r,cr,c, the coefficient at any position s′s' is fixed, namely b+(s′−i)cb+(s'-i)c. On the same pair this determines the candidate uniquely. Any other unordered pair with this root uses at least a third position with a prescribed positive coefficient at least TcTc. Conditional on the first two independent site sets, that coefficient is available at the third site with probability at most 1/(Tc)1/(Tc): it has probability zero unless it is a squarefree product of P\mathcal{P} primes, and otherwise the probability is its reciprocal. The polynomial candidate and position bounds, and c≥eBc \ge e^B, make the union of these events negligible. The site coupling transfers the conclusion to the actual model.

Forced hits at a third endpoint. Also discard the event that, for some candidate and s′≠i,ks' \ne i,k, an occupied prime pp at site ss satisfies p∤abcp \nmid abc and r≡−s′r \equiv-s' (mod pp). In the independent-site model, condition on the candidate’s two endpoint sets. Such a prime divides b+(s′−i)cb+(s'-i)c, a nonzero integer of logarithmic size O(B)O(B). It is nonzero because c>1c>1 and (b,c)=1(b,c)=1. For any nonzero integer of logarithmic size O(B)O(B), the sum of reciprocals of its prime divisors above P0P_0 is O(B/P0)O(B/P_0). Thus the conditional union probability is O(B/P0)O(B/P_0) per candidate and third position. A polynomial union bound and the site coupling again suffice.

Coefficient primes and occupied primes. The two structural exclusions above depend only on the endpoint sets, and their actual-law probabilities have been bounded through the site coupling. Fix an actual site configuration avoiding them; its hits at each prime are disjoint. We still do not condition on absence of site squares. At an unoccupied prime, uu mod pp is uniform outside the MM forbidden classes −1,…,−M-1,\ldots,-M. At an occupied prime, its residue is fixed and the higher pp-adic digits remain uniform. These conditional laws are independent over primes.

If p∣abp \mid ab, then the coefficients are squarefree and pairwise coprime, and the test p∣mp \mid m in (5.11) has conditional probability 1/p1/p, determined by the next digit. If p∣cp \mid c, it is impossible. Set these exceptional tests to zero temporarily, in both the actual model and the independent comparison model. The total error per candidate is O(B/P0)O(B/P_0). This use of uniform higher digits is made before removing site-square events from the weights; we do not condition those digits on absence of squares.

At p∤abcp \nmid abc, the root test is u≡r(modp)u \equiv r \pmod p. If the prime is occupied at an endpoint, such coincidence would force p∣bp \mid b or p∣ap \mid a. At any other occupied site it is excluded by the preceding discarded event. The actual root test is therefore zero at the remaining occupied primes. The cost of setting the independent test to zero there is bounded in expectation by the deterministic candidate bound times

E∑p occupied1p≤M∑p>P01p2.\mathbb{E}\sum_{p\ \mathrm{occupied}}\frac{1}{p}\le M\sum_{p>P_0}\frac{1}{p^2}.

This domination does not require candidates to be independent of occupied primes.

Unoccupied primes. The roots are now distinct rational numbers. A collision between two of their residues modulo pp, after excluding coefficient primes, requires pp to divide their nonzero cross-multiplied difference. That integer has logarithmic size O(B)O(B). A collision with a forbidden site class has the same bound, or concerns a coefficient prime already handled. For each candidate pair or candidate--site pair, these defective primes have reciprocal sum O(B/P0)O(B/P_0). At all such primes, discard all involved tests. Conditional on the sites, each defined test has probability at most 1/(p−M)≤2/p1/(p-M)\le2/p, so another polynomial loss is harmless.

At a remaining unoccupied prime the root classes are distinct and avoid all forbidden classes. Their joint law is a disjoint-choice law with individual hit probabilities 1/(p−M)1/(p-M). If there are Ncand=O(B4)N_{\mathrm{cand}}=O(B^4) tests, its total variation distance from independent Bernoulli(1/p)(1/p) tests is

O((M+Ncand)2p2).O\left(\frac{(M+N_{\mathrm{cand}})^2}{p^2}\right).

For example, compare first to independent tests of probability 1/(p−M)1/(p-M), at cost O(Ncand2/p2)O(N_{\mathrm{cand}}^2/p^2) from multiple hits, and then change each probability to 1/p1/p, at cost O(NcandM/p2)O(N_{\mathrm{cand}}M/p^2). Summing over primes constructs the claimed conditional coupling. The preceding union bounds give total failure probability O(BC/P0)+O(BCe−B)O(B^C/P_0)+O(B^C e^{-B}) for a fixed crude exponent CC. On these failures, use the deterministic polynomial matrix-mass bounds. After including that loss and enlarging the crude exponent, their contribution to the expected cut norm is at most O(B50/P0)+O(B50e−B)=o(1)O(B^{50}/P_0)+O(B^{50}e^{-B})=o(1). Equation (40) is handled in the resulting weights, as just explained.

Fluctuations of the extra weights. The preceding coupling has separated the extra roots from the endpoint types. It remains to show that their independent fluctuations are small even after maximizing the testing signs. The comparison model has independent endpoint types and an additional independent set SrootS_{\mathrm{root}} for each unordered candidate. In its weight, replace kBk_B in (41) by K(Sroot)K(S_{\mathrm{root}}). Conditional on the endpoints, the candidate weights are independent, and their means give precisely L\mathcal{L}. By (3.9), each full candidate weight, including its six truncated factors and the factor 1/B1/B, is at most

CB−1+6(h0+τlog⁡2)+o(1)≤B−0.07CB^{-1+6(h_0+\tau\log2)+o(1)}\le B^{-0.07}

for large BB, with harmless enlargement of constants. For fixed sign vectors in (38), an unordered candidate has sign coefficient of absolute value at most two. Positivity bounds the conditional variance sum by a constant times B−0.07∑i,kLikB^{-0.07}\sum_{i,k}\mathcal{L}_{ik}.

Bernstein’s inequality, union over at most 22M2^{2M} sign choices, and integration of the tail show that the conditional expected cut norm of the centered fluctuation is at most

C(B−0.071M∑i,kLik+B−0.07).(44)C\left(\sqrt{B^{-0.07}\frac{1}{M}\sum_{i,k}\mathcal{L}_{ik}}+B^{-0.07}\right). \tag*{(44)}

Indeed, before division by MM, a threshold

C(B−0.07(∑i,kLik)(M+y)+B−0.07(M+y))C\left(\sqrt{B^{-0.07}\left(\sum_{i,k}\mathcal{L}_{ik}\right)(M+y)+B^{-0.07}(M+y)}\right)

has tail O(e−y)O(e^{-y}) for y≥0y \ge0, with CC absorbing the sign enumeration. Lemma 5.2 and Jensen’s inequality make the expectation of (67) tend to zero. The triangle inequality now proves eq:10; the deterministic polynomial mass bounds control the coupling failures for both kernels. All estimates used only uniform bounds for the smooth factors on S\mathcal{S}.

A second moment for the latent rows

The independent-root comparison leaves the latent kernel L\mathcal{L} of (3.9). Its conditional first moments are bounded uniformly over every site type by (3.11). For the sampling argument we also need an integrated second-moment estimate. A small bound on each individual representation does not suffice, because an edge may have several representations. Expanding the square and weighting one representation will identify the measure under which that multiplicity must be controlled.

The grid size LL and tolerance τ\tau will be chosen once below, consistently with Section 3. In every limit B→∞B \to\infty in this section, those choices, η>0\eta> 0, the block constant C6C_6, C∗C_*, and the smooth weights are held fixed. All estimates are uniform in the positions and in the smooth parameter ss on the fixed compact range used in Section 5. Constants may depend on the fixed parameters but not on the particular first representation.

Proposition 6.1 (Integrated row moments). For a sufficiently fine fixed regularity grid and then a sufficiently small fixed tolerance, the independent-site latent kernel satisfies

∑k≠iE Lik2≪B−.21(1≤i≤M).(45)\sum_{k \ne i} \mathbb{E}\,\mathcal{L}_{ik}^{2} \ll B^{-.21} \qquad(1 \le i \le M). \tag*{(45)}

Moreover, its normalized total mass has bounded second moment:

E(1M∑i,kLik)2≪1.\mathbb{E}\left(\frac{1}{M}\sum_{i,k}\mathcal{L}_{ik}\right)^2 \ll1.

For every set I⊆{1,…,M}I \subseteq\{1,\ldots,M\}, the same estimates hold after replacing Lik\mathcal{L}_{ik} by 1{i,k∈I}Lik\mathbf{1}_{\{i,k\in I\}}\mathcal{L}_{ik}; the normalization remains MM.

Write hτ=h0+τlog⁡2h_\tau=h_0+\tau\log2. There are five truncated weights in an individual summand of (3.9). Since kB≪1k_B\ll1 and the smooth factors are bounded, (3.11) gives the uniform bound

one contribution to Lik≤B−κ1+5τlog⁡2+o(1),κ1=1−5h0=0.232867951…>.2328.(46)\text{one contribution to }\mathcal{L}_{ik}\le B^{-\kappa_1+5\tau\log2+o(1)}, \qquad \kappa_1=1-5h_0=0.232867951\ldots>.2328. \tag*{(46)}

The measure obtained from a first representation

Fix an allowed pair i,k=i+ji,k=i+j. Let Cj\mathcal{C}_j be the finite set of positive triples (a,b,c)(a,b,c) satisfying

a−b=jc,B≤log⁡c≤2B,Tc≤a,b≤2Tc,a-b=jc,\qquad B\le\log c\le2B,\qquad T c\le a,b\le2T c,

whose entries are pairwise coprime, rough, squarefree, and regular. These are precisely the coefficient conditions for a candidate; the two remainder regularity conditions still depend on the site types. We initially take j>0j>0; the negative case follows by exchanging the endpoints. Fix (a,b,c)∈Cj(a,b,c)\in\mathcal{C}_j. For an integer vv whose prime factors belong to P\mathcal{P}, write P(v)\mathcal{P}(v) for its set of prime factors.

In the independent-site model, the event P(b)⊂Si\mathcal{P}(b) \subset S_i, P(a)⊂Sk\mathcal{P}(a) \subset S_k has probability 1/(ab)1/(ab). Conditional on that event, put

Tb=Si∖P(b),Ta=Sk∖P(a).\mathcal{T}_b = S_i \setminus\mathcal{P}(b), \qquad\mathcal{T}_a = S_k \setminus\mathcal{P}(a).

These sets are independent, and at each prime not excluded by the corresponding coefficient the inclusion probability is 1/p1/p. Weighting one such inclusion by its factor 1/21/2 in K0K_0 changes this probability to

(1/p)/21−1/(2p)=12p−1.\frac{(1/p)/2}{1-1/(2p)}=\frac{1}{2p-1}.

Thus weighting by K0(Tb)K0(Ta)K_0(\mathcal{T}_b)K_0(\mathcal{T}_a) gives a product probability measure, denoted by Pa,btilt\mathbb{P}^{\mathrm{tilt}}_{a,b}, under which the two remainder sets are independent and their available prime indicators have probabilities 1/(2p−1)1/(2p-1).

For a coefficient vv the normalizing factor is

Ξ(v)=R1/2∏p∈P∖P(v)(1−12p).\Xi(v)=R^{1/2}\prod_{p\in\mathcal{P}\setminus\mathcal{P}(v)}\left(1-\frac{1}{2p}\right).

The full product with no exclusions is bounded by Mertens’ estimate. Furthermore log⁡v=O(B)\log v=O(B) and every prime of vv exceeds P0P_0, so

∑p∣v1p≪BP0.\sum_{p\mid v}\frac{1}{p}\ll\frac{B}{P_0}.

Removing these factors changes the product by 1+o(1)1+o(1). Consequently Ξ(v)≪1\Xi(v)\ll1 uniformly for the coefficients under consideration. For every nonnegative function HH of the two remainder sets we therefore have the exact change-of-measure identity

E[1P(b)⊂Si,P(a)⊂SkK0(Tb)K0(Ta)H(Tb,Ta)]=Ξ(b)Ξ(a)abEa,btiltH(Tb,Ta).\mathbb{E}\left[1_{\mathcal{P}(b)\subset S_i,\mathcal{P}(a)\subset S_k}K_0(\mathcal{T}_b)K_0(\mathcal{T}_a)H(\mathcal{T}_b,\mathcal{T}_a)\right] =\frac{\Xi(b)\Xi(a)}{ab}\mathbb{E}^{\mathrm{tilt}}_{a,b}H(\mathcal{T}_b,\mathcal{T}_a).

The indicators that the two first remainders are regular will stay inside HH. In particular, we never divide by their probability.

Given these remainder sets, reconstruct Si=P(b)∪TbS_i=\mathcal{P}(b)\cup\mathcal{T}_b and Sk=P(a)∪TaS_k=\mathcal{P}(a)\cup\mathcal{T}_a, and let Na,b,c(Tb,Ta)N_{a,b,c}(\mathcal{T}_b,\mathcal{T}_a) be the number of valid candidates at this pair for the reconstructed types, including the first triple whenever it is valid.

We can now see exactly how this measure enters the square. Expand the square of the sum of nonnegative candidate contributions, designate one candidate as the first representation, and bound every alternative contribution by (6.2). The first contribution retains the two factors K0(Tb)K0(Ta)K_0(\mathcal{T}_b)K_0(\mathcal{T}_a) and their regularity indicators. The change-of-measure identity gives

E Lik2≪B−κ1+5τlog⁡2+o(1)1B∑(a,b,c)∈CjA(a)A(b)A(c)abΞ(a)Ξ(b)Ea,btilt[1Tb regular1Ta regularNa,b,c(Tb,Ta)].(47)\mathbb{E}\,L_{ik}^{2}\ll B^{-\kappa_1+5\tau\log2+o(1)}\frac{1}{B}\sum_{(a,b,c)\in\mathcal{C}_j}\frac{A(a)A(b)A(c)}{ab}\Xi(a)\Xi(b) \mathbb{E}^{\mathrm{tilt}}_{a,b}\left[1_{\mathcal{T}_b\ \mathrm{regular}}1_{\mathcal{T}_a\ \mathrm{regular}}N_{a,b,c}(\mathcal{T}_b,\mathcal{T}_a)\right]. \tag*{(47)}

The implied constant absorbs kBk_B and the bounded smooth factors. Thus the remaining multiplicity question is the following uniform estimate; the coefficient sum will then be bounded by (3.8).

Lemma 6.2 (Weighted representation multiplicity). The grid and tolerance can be chosen so that, uniformly for (a,b,c)∈Cj(a,b,c)\in\mathcal{C}_j and ηT≤∣j∣≤T\eta T\leq|j|\leq T,

Ea,btilt[1Tb regular1Ta regularNa,b,c(Tb,Ta)]≪B.009.\mathbb{E}^{\mathrm{tilt}}_{a,b}\left[1_{\mathcal{T}_b\ \mathrm{regular}}1_{\mathcal{T}_a\ \mathrm{regular}}N_{a,b,c}(\mathcal{T}_b,\mathcal{T}_a)\right]\ll B^{.009}.

Proof. Every alternative candidate has a unique decomposition

b′=d1y,d1∣b,P(y)⊂Tba′=e1z,e1∣a,P(z)⊂Taa′−b′=jc′.b' = d_1y,\quad d_1 \mid b,\quad\mathcal{P}(y) \subset\mathcal{T}_b \qquad a' = e_1z,\quad e_1 \mid a,\quad\mathcal{P}(z) \subset\mathcal{T}_a \qquad a' - b' = jc'.

Here all the products are squarefree, and yy is coprime to bb while zz is coprime to aa. The coefficient c′c' is determined once the four products are chosen. We may discard any of the alternative candidate conditions for upper bounds; however, we retain the coefficient regularity conditions that are used below. Since a′,b′≤2Texp⁡(2B)a',b' \le2T\exp(2B), we have log⁡y,log⁡z≤3B\log y,\log z \le3B for all sufficiently large BB.

Set Δ=1/L\Delta= 1/L, δ0=3Δ\delta_0 = 3\Delta, and U=max⁡(log⁡y,log⁡z)U = \max(\log y,\log z). We count the alternatives with U≤Bδ0U \le B^{\delta_0} directly from prefix regularity. When U>Bδ0U > B^{\delta_0}, fixing the omissions and the new product at one endpoint places the other new product in one residue class modulo ∣j∣\lvert j\rvert, through e1z−d1y=jc′e_1z-d_1y=jc'. The resulting progression density is the gain that will pay for the remaining subset choices.

For a grid point g<1g<1, define

Qg={p∈P:log⁡p≤Bg},Q1=P,Q_g = \{p \in\mathcal{P} : \log p \le B^g\},\qquad Q_1=\mathcal{P},

and write ωg(v)=∣P(v)∩Qg∣\omega_g(v)=\lvert\mathcal{P}(v)\cap Q_g\rvert, with the same notation for a prime set. Regularity gives

(g/2−τ)ℓ≤ωg(v)≤(g/2+τ)ℓ(g/2-\tau)\ell\le\omega_g(v) \le(g/2+\tau)\ell

for each of the first and alternative coefficients, and for each regular first remainder. Mertens’ estimate, on this fixed grid, gives

∑p∈Qg1p=gℓ+o(ℓ).\sum_{p\in Q_g}\frac{1}{p}=g\ell+o(\ell).

The o(ℓ)o(\ell) is uniform over the finitely many grid points.

Omissions above an addition prefix. Suppose that all primes of zz and yy lie in QgQ_g. Let Da=P(a)∖P(e1)D_a=\mathcal{P}(a)\setminus\mathcal{P}(e_1) be the primes omitted from aa. Regularity of the total counts of a,a′a,a' gives

∣Da∣−ω(z)≤2τℓ.\left\lvert D_a\right\rvert-\omega(z)\le2\tau\ell.

Regularity at gg gives the analogous inequality

∣Da∩Qg∣−ω(z)≤2τℓ.\left\lvert D_a\cap Q_g\right\rvert-\omega(z)\le2\tau\ell.

Subtracting shows that at most 4τℓ4\tau\ell primes above the prefix can be omitted. The same conclusion holds for the omissions from bb. When g=1g=1 there are no primes above the prefix.

Write

H(u)=−ulog⁡u−(1−u)log⁡(1−u),0≤u≤1,H(u)=-u\log u-(1-u)\log(1-u),\qquad0\le u\le1,

with the endpoint values defined by continuity. For completeness, if m≤8τℓm\le8\tau\ell, the number of such subsets is at most 2m≤28τℓ2^m\le2^{8\tau\ell}. For 8τℓ<m≤(1/2+τ)ℓ8\tau\ell<m\le(1/2+\tau)\ell, the binomial bound

∑q≤4τℓ(mq)≤(m+1)exp⁡(mH(4τℓ/m))≤(m+1)exp⁡((1/2+τ)ℓH(4τ1/2+τ))\sum_{q\le4\tau\ell}\binom{m}{q}\le(m+1)\exp\left(mH(4\tau\ell/m)\right)\le(m+1)\exp\left((1/2+\tau)\ell H\left(\frac{4\tau}{1/2+\tau}\right)\right)

applies: HH is increasing up to 1/21/2, and mH(k/m)mH(k/m) is increasing in mm for fixed 0<k<m0<k<m. Thus the choices at both coefficients together cost Bεhigh(τ)+o(1)B^{\varepsilon_{\mathrm{high}}(\tau)+o(1)}, where εhigh(τ)→0\varepsilon_{\mathrm{high}}(\tau)\to0 as τ→0\tau\to0. The estimate is uniform also when the relevant set is empty.

Small additions. If

U≤Bg0,U \le B^{g_0},

all additions lie in Qg0Q_{g_0}. On the two first-remainder regularity events, the four relevant prefix sets—the prime sets of a,ba,b and the two first remainders—have at most (2g0+4τ)ℓ(2g_0+4\tau)\ell primes altogether. Allowing arbitrary subset choices there, and then the higher omissions just bounded, gives the pointwise estimate

Nsmall≤2(2g0+4τ)ℓBεhigh(τ)+o(1)=B(2g0+4τ)log⁡2+εhigh(τ)+o(1).N_{\mathrm{small}} \le2^{(2g_0+4\tau)\ell}B^{\varepsilon_{\mathrm{high}}(\tau)+o(1)} = B^{(2g_0+4\tau)\log2+\varepsilon_{\mathrm{high}}(\tau)+o(1)}.

This also bounds its expectation with the first-remainder regularity indicators. By taking LL large and then τ\tau small, its exponent apart from o(1)o(1) can be made strictly smaller than .008.008.

Large-addition classes and combinatorial choices. For U>Bg0U>B^{g_0} use the classes

Bg−Δ<U≤Bg(g0<g<1),B1−Δ<U≤3B(g=1),B^{g-\Delta}<U\le B^g \qquad(g_0<g<1),\qquad B^{1-\Delta}<U\le3B \qquad(g=1),

where gg ranges over grid points. All addition primes again belong to QgQ_g. It suffices to treat the portion of a class in which log⁡z>Bg−Δ\log z>B^{g-\Delta}. The portion with log⁡y>Bg−Δ\log y>B^{g-\Delta} is estimated by the same argument with the endpoints exchanged; the numeric sieve below permits either sign for its second slope. If both inequalities hold, counting the alternative twice only increases the upper bound.

Define the exact normalized counts and their clipped values by

rz∗=ω(z)gℓ,ry∗=ω(y)gℓ,rz=min⁡(rz∗,1/2),ry=min⁡(ry∗,1/2).r_z^*=\frac{\omega(z)}{g\ell},\qquad r_y^*=\frac{\omega(y)}{g\ell},\qquad r_z=\min(r_z^*,1/2),\qquad r_y=\min(r_y^*,1/2).

On the first-remainder regularity events, 0≤rz∗,ry∗≤1/2+τ/g0\le r_z^*,r_y^*\le1/2+\tau/g. There are O(ℓ2)O(\ell^2) possible pairs of integer counts. We estimate one such pair at a time.

The number of omitted primes in the aa prefix is rz∗gℓ+O(τℓ)r_z^*g\ell+O(\tau\ell), among gℓ/2+O(τℓ)g\ell/2+O(\tau\ell) available primes. Applying (mq)≤exp⁡(mH(q/m))\binom{m}{q}\le\exp(mH(q/m)), and summing over the O(ℓ)O(\ell) permissible values of qq, gives, with the higher-omission factor included,

#{e1}≤B12gH(2rz)+εL(τ)+o(1).(48)\#\{e_1\}\le B^{\frac{1}{2}gH(2r_z)+\varepsilon_L(\tau)+o(1)}. \tag*{(48)}

The function εL(τ)\varepsilon_L(\tau) can be chosen to tend to zero for fixed LL. To justify this uniformly in the counts, divide m,qm,q by ℓ\ell and use uniform continuity of HH on [0,1][0,1]; the perturbations are O(τ/g)O(\tau/g), and g≥g0>0g\ge g_0>0 is fixed. The clipping from rz∗r_z^* to rzr_z changes these arguments by the same amount. Also ℓ/log⁡B→1\ell/\log B\to1, and all polynomial factors in ℓ\ell contribute Bo(1)B^{o(1)}. We use εL(τ)\varepsilon_L(\tau) below for a possibly enlarged function with this same limiting property.

The choices of d1d_1 have the corresponding bound with ryr_y. Conditional on a regular first remainder TbT_b, the number of its subsets yy of the prescribed cardinality is at most

B12gH(2ry)+εL(τ)+o(1).B^{\frac{1}{2}gH(2r_y)+\varepsilon_L(\tau)+o(1)}.

This bound is uniform in that remainder set. We may condition on it and sum over these yy while the remainder at the other endpoint still has its independent tilted law.

Availability of a fixed numeric addition. The entropy bounds count the possible omitted prime sets. To control large additions, we also need the probability that each numerically specified product is present in the other remainder set. The next bound keeps the first remainder’s regularity requirement in this probability. For a fixed squarefree zz in this class, with P(z)⊂Qg∖P(a)\mathcal{P}(z) \subset\mathcal{Q}_g \setminus\mathcal{P}(a), its inclusion probability in Ta\mathcal{T}_a is

∏p∣z12p−1=2−ω(z)z∏p∣z(1−12p)−1≪2−ω(z)z.\prod_{p\mid z} \frac{1}{2p-1}=\frac{2^{-\omega(z)}}{z}\prod_{p\mid z}\left(1-\frac{1}{2p}\right)^{-1}\ll\frac{2^{-\omega(z)}}{z}.

The last bound is uniform because log⁡z≤3B\log z\le3B and all its primes exceed P0P_0. Conditional on those inclusions, the remaining available indicators in Qg\mathcal{Q}_g are independent with parameter sum

Λ=∑p∈Qg∖(P(a)∪P(z))12p−1=12gℓ+o(ℓ).\Lambda=\sum_{p\in\mathcal{Q}_g\setminus(\mathcal{P}(a)\cup\mathcal{P}(z))}\frac{1}{2p-1}=\frac{1}{2}g\ell+o(\ell).

The excluded primes have reciprocal sum O(B/P0)O(B/P_0), so this estimate is uniform in the fixed coefficient and numeric addition. First prefix regularity requires the number of these other inclusions to be at most (1/2−rz∗)gℓ+τℓ(1/2-r_z^*)g\ell+\tau\ell.

If XX is a sum of independent Bernoulli variables with parameter sum Λ\Lambda, then for t≥0t\ge0,

P(X≤k)≤exp⁡(tk+(e−t−1)Λ).\mathbb{P}(X\le k)\le\exp\bigl(tk+(e^{-t}-1)\Lambda\bigr).

For 0≤k≤Λ0\le k\le\Lambda, minimizing gives exponent −Λ+k−klog⁡(k/Λ)-\Lambda+k-k\log(k/\Lambda), with its continuous value at k=0k=0. For k≥Λk\ge\Lambda use the bound 11, and for k<0k<0 the event is empty. Uniform continuity of the resulting rate, including the endpoints, and the fixed-grid relation g≥g0g\ge g_0 therefore imply

Pa,btilt(P(z)⊂Ta, Ta regular at g)≪(1/2)ω(z)zB−gI1/2(1/2−rz)+εL(τ)+o(1),I1/2(u)=12−u+ulog⁡u1/2.(49)\mathbb{P}^{\mathrm{tilt}}_{a,b}\bigl(\mathcal{P}(z)\subset\mathcal{T}_a,\ \mathcal{T}_a\text{ regular at }g\bigr) \ll\frac{(1/2)^{\omega(z)}}{z}B^{-gI_{1/2}(1/2-r_z)+\varepsilon_L(\tau)+o(1)}, \qquad I_{1/2}(u)=\frac{1}{2}-u+u\log\frac{u}{1/2}. \tag*{(49)}

Here ulog⁡u=0u\log u=0 at u=0u=0. Retaining only prefix regularity on the left gives an upper bound for retaining all of first-remainder regularity, which is what the multiplicity expectation requires.

A harmonic sieve for the numeric addition. We now sum the availability bound over numeric additions satisfying the alternative coefficient relation. The factor 1/z1/z in that bound is the reason for the harmonic normalization below. Fix e1,d1,ye_1,d_1,y from the preceding choices. The alternative size conditions imply

z≍Z′:=d1ye1z\asymp Z':=\frac{d_1y}{e_1}

with absolute comparison constants; we may enlarge the interval to [Z′/2,2Z′][Z'/2,2Z'].

For these fixed data and the fixed cardinality class, let Z\mathcal{Z} consist of the positive integers zz with the following properties:

z is a squarefree product of primes in Qg∖P(a),ω(z)=rz∗gℓ,z\text{ is a squarefree product of primes in }\mathcal{Q}_g\setminus\mathcal{P}(a),\qquad\omega(z)=r_z^*g\ell,
Z′/2≤z≤2Z′,log⁡z∈{[Bg−Δ,Bg],g<1,[B1−Δ,3B],g=1,Z'/2\le z\le2Z',\qquad \log z\in \begin{cases} [B^{g-\Delta},B^g], & g<1,\\ [B^{1-\Delta},3B], & g=1, \end{cases}
e1z≡d1y(mod∣j∣),c′(z):=e1z−d1yj>0,(e1z,d1y,c′(z))∈Cj.e_1z\equiv d_1y\pmod{|j|},\qquad c'(z):=\frac{e_1z-d_1y}{j}>0,\qquad (e_1z,d_1y,c'(z))\in\mathcal{C}_j.

The congruence makes c′(z)c'(z) an integer. Every valid alternative in this designated class yields an element of Z\mathbb{Z}. The inclusion P(z)⊂Ta\mathcal{P}(z) \subset\mathcal{T}_a remains the event estimated in (49); the two alternative remainder regularity conditions are discarded. Equation (49) applies to every z∈Zz \in\mathcal{Z}, because each is a possible product in the tilted remainder at the aa endpoint. If Z\mathcal{Z} is empty there is nothing to prove; otherwise its interval and log conditions give Z′≥12exp⁡(Bg−Δ)Z' \ge\frac{1}{2}\exp(B^{g-\Delta}).

All of e1,d1,ye_1,d_1,y are units modulo ∣j∣|j|, since their prime factors exceed P0>TP_0>T. Thus the displayed congruence is one unit residue class. We claim

∑z∈Z(1/2)ω(z)z≪1∣j∣Bg[F0(rz)−1−h0]+3Δ+rmin⁡+τlog⁡2+o(1),F0(u)=u(1+log⁡1/2u),F0(0)=0,(50)\sum_{z\in\mathcal{Z}} \frac{(1/2)^{\omega(z)}}{z} \ll\frac{1}{|j|}B^{g[F_0(r_z)-1-h_0]+3\Delta+r_{\min}+\tau\log2+o(1)}, \qquad F_0(u)=u\left(1+\log\frac{1/2}{u}\right),\qquad F_0(0)=0, \tag*{(50)}

where rmin⁡∈(0,1/2)r_{\min}\in(0,1/2) is any fixed number. The excess exponent 3Δ+rmin⁡+τlog⁡2+o(1)3\Delta+r_{\min}+\tau\log2+o(1) can be made small by the later choices of grid, clamp, and tolerance.

To prove the deterministic estimate, put λ=max⁡(rmin⁡,rz)\lambda=\max(r_{\min},r_z), so rmin⁡≤λ≤1/2r_{\min}\le\lambda\le1/2. On Z\mathcal{Z} the fixed cardinality gives

(1/2)ω(z)=λω(z)(1/2λ)rz∗gℓ.(1/2)^{\omega(z)}=\lambda^{\omega(z)}\left(\frac{1/2}{\lambda}\right)^{r_z^*g\ell}.

The membership of (e1z,d1y,c′(z))(e_1z,d_1y,c'(z)) in Cj\mathcal{C}_j also gives the upper prefix cutoff for c′(z)c'(z), and hence

1≤2(g/2+τ)ℓ2−ωg(c′(z)).1\le2^{(g/2+\tau)\ell}2^{-\omega_g(c'(z))}.

The prefix factor is introduced to supply a second reducing root in the sieve below. At the leading gg-scale its saving −g/2-g/2 combines with the cutoff cost glog⁡2/2g\log2/2, leaving −gh0-gh_0 in the exponent. The calculation below verifies the root conditions and retains the grid and tolerance losses. Insert these two factors before enlarging the summation domain, and keep only the reducing prime factors up to

Zsieve=exp⁡(Bg−2Δ).Z_{\mathrm{sieve}}=\exp(B^{g-2\Delta}).

Deleting the other reducing factors increases each summand. Since z≥Z′/2z\ge Z'/2 on Z\mathcal{Z}, we obtain the deterministic majorant

∑z∈Z(1/2)ω(z)z≤2Z′(1/2λ)rz∗gℓ2(g/2+τ)ℓ∑z∈Z∩[Z′/2,2Z′]e1z≡d1y(mod∣j∣)∏P0<p≤Zsieveλ1{p∣z}2−1{p∣c′(z)}.\sum_{z\in\mathcal{Z}}\frac{(1/2)^{\omega(z)}}{z} \le\frac{2}{Z'}\left(\frac{1/2}{\lambda}\right)^{r_z^*g\ell}2^{(g/2+\tau)\ell} \sum_{\substack{z\in\mathbb{Z}\cap[Z'/2,2Z']\\ e_1z\equiv d_1y\pmod{|j|}}} \prod_{P_0<p\le Z_{\mathrm{sieve}}}\lambda^{1\{p\mid z\}}2^{-1\{p\mid c'(z)\}}.

The finite congruence product is defined for every integer value of c′(z)c'(z), including zero. The hard prefix cutoff, the fixed-cardinality condition, and all other candidate restrictions have been removed only in this deterministic sum. Equation (49) is applied only to z∈Zz\in\mathcal{Z}; the displayed enlargement bounds the weighted sum that results.

For explicit congruence accounting put J1=∣j∣J_1=|j| and σ=j/J1\sigma=j/J_1, and choose a unit z0z_0 modulo J1J_1 satisfying the progression condition. Write

z=z0+J1t,c′=c0+σe1t,c0=e1z0−d1yj∈Z.z=z_0+J_1t,\qquad c'=c_0+\sigma e_1t,\qquad c_0=\frac{e_1z_0-d_1y}{j}\in\mathbb{Z}.

The parameter tt runs through an interval of length N≍Z′/J1N\asymp Z'/J_1. Its logarithm is at least Bg−Δ−O(log⁡B)B^{g-\Delta}-O(\log B). Consequently

log⁡Zsievelog⁡N=O(B−Δ)⟶0.\frac{\log Z_{\mathrm{sieve}}}{\log N}=O(B^{-\Delta})\longrightarrow0.

so the interval upper sieve of Section 3, in fixed dimension two, applies uniformly.

The displayed product already omits every prime at most P0P_0, including every prime of jj. Among the remaining primes also drop those dividing e1d1ye_1d_1y. Their reciprocal sum is O(B/P0)O(B/P_0), uniformly, since log⁡(e1d1y)=O(B)\log(e_1d_1y)=O(B). At each retained prime the two slopes are invertible, and the two roots are distinct. Indeed their determinant is

J1c0−σe1z0=−σd1y,J_1c_0-\sigma e_1z_0=-\sigma d_1y,

which is nonzero modulo such a prime. The local weights of the roots are λ\lambda and 1/21/2, so the average local weight is exactly

1−3/2−λp.1-\frac{3/2-\lambda}{p}.

This also explains why no coefficient-height factor enters the sieve: only the number and distinctness of residue roots are used. If one keeps primes dividing jj instead, the first form is a fixed unit and the second has invertible slope; dropping them avoids any need for a separate local factor.

The random-weight interval sieve bounds the sum of the retained reducing weights by a constant times

N∏P0<p≤Zsievep∤e1d1y(1−3/2−λp)+O(N.8).N\prod_{\substack{P_0<p\le Z_{\mathrm{sieve}}\\p\nmid e_1d_1y}}\left(1-\frac{3/2-\lambda}{p}\right)+O(N^{.8}).

The first term is NB(g−2Δ)(λ−3/2)+o(1)NB^{(g-2\Delta)(\lambda-3/2)+o(1)}, by Mertens’ estimate. Indeed the reciprocal sum in that product is (g−2Δ)log⁡B−log⁡log⁡P0+O(1)=(g−2Δ)log⁡B+o(log⁡B)(g-2\Delta)\log B-\log\log P_0+O(1)=(g-2\Delta)\log B+o(\log B), and the sum of the squared reciprocals is negligible. Discarding the small-prime conditions costs at most powers of log⁡P0\log P_0 relative to their possible sieve factors, and those powers are Bo(1)B^{o(1)}; alternatively, the displayed large-prime product already proves the upper bound directly.

Using z≍Z′z\asymp Z' and restoring the threshold factors, the harmonic normalization of the main sieve bound is N/Z′≍1/J1N/Z'\asymp1/J_1. The remainder contributes at most

1J1BO(1)N−.2,\frac{1}{J_1}B^{O(1)}N^{-.2},

which is negligible uniformly in the class, since g≥4Δg\ge4\Delta and log⁡N≥Bg−Δ−O(log⁡B)\log N\ge B^{g-\Delta}-O(\log B). The total power of BB in the main term is

(g−2Δ)(λ−3/2)+grz∗log⁡1/2λ+(g/2+τ)log⁡2+o(1).(g-2\Delta)(\lambda-3/2)+gr_z^*\log\frac{1/2}{\lambda}+(g/2+\tau)\log2+o(1).

If rz∗≤1/2r_z^*\le1/2, then rz=rz∗r_z=r_z^*, and this power equals

g[F0(rz)−1−h0]+g(λ−rz+rzlog⁡rzλ)+2Δ(3/2−λ)+τlog⁡2+o(1).g[F_0(r_z)-1-h_0]+g\left(\lambda-r_z+r_z\log\frac{r_z}{\lambda}\right)+2\Delta(3/2-\lambda)+\tau\log2+o(1).

We used −3/2+(log⁡2)/2=−1−h0-3/2+(\log2)/2=-1-h_0. The clamp contribution in parentheses vanishes when rz≥rmin⁡r_z\ge r_{\min} and lies between 00 and rmin⁡r_{\min} when 0≤rz<rmin⁡0\le r_z<r_{\min}, including the value rmin⁡r_{\min} at rz=0r_z=0. The grid contribution is at most 3Δ3\Delta. If rz∗>1/2r_z^*>1/2, our choice is rz=λ=1/2r_z=\lambda=1/2; the Rankin factor is then exactly 1 and the same upper bound follows directly with the clipped rzr_z. These observations prove (50). If yy is the designated large numeric variable, the identical proof uses the slope −d1-d_1 for c′c' instead, and its determinant has the same nonvanishing property outside the already discarded coefficient primes.

Combining the bounds and choosing parameters. Condition on the first remainder at the bb endpoint. If it is not regular its contribution vanishes. If it is regular, enumerate the choices of d1,yd_1,y by the two entropy bounds with ryr_y, and enumerate e1e_1 by (48). For each such choice, sum the availability bound (49) using (50). Independence of the other first remainder justifies this conditioning, and every bound is uniform in the conditioned regular remainder. Thus we may average it out without any additional factor.

A direct expansion of HH gives the exact identity

F0(rz)−I1/2(1/2−rz)=12H(2rz)(0≤rz≤1/2).F_0(r_z)-I_{1/2}(1/2-r_z)=\frac{1}{2}H(2r_z) \qquad(0\le r_z\le1/2).

The three combinatorial factors have entropy exponents gH(2rz)/2gH(2r_z)/2, gH(2ry)/2gH(2r_y)/2, and gH(2ry)/2gH(2r_y)/2. Adding the probability and sieve exponents therefore bounds the contribution of one class, one orientation, and one count pair, with both first-remainder regularity indicators, by

≪1∣j∣Bg[H(2rz)+H(2ry)−1−h0]+3Δ+rmin⁡+εL(τ)+o(1)≤1∣j∣Bκ1+3Δ+rmin⁡+εL(τ)+o(1).(51)\begin{aligned} \ll\frac{1}{|j|}B^{g[H(2r_z)+H(2r_y)-1-h_0]+3\Delta+r_{\min}+\varepsilon_L(\tau)+o(1)}\\ \le\frac{1}{|j|}B^{\kappa_1+3\Delta+r_{\min}+\varepsilon_L(\tau)+o(1)}. \tag*{(51)} \end{aligned}

The tolerance error here includes the preceding τlog⁡2\tau\log2 and the finitely many entropy and rate perturbations. The last inequality uses H≤log⁡2H\le\log2 and 0<g≤10<g\le1. If the bracket in the first exponent is negative, multiplying it by gg still gives a nonpositive number; otherwise its maximum is at most 2log⁡2−1−h0=κ12\log2-1-h_0=\kappa_1.

Here is an explicit consistent order of choices. First choose LL large enough that 6Δlog⁡2<.0046\Delta\log2<.004 and 3Δ<.013\Delta<.01. Next choose a fixed rmin⁡<.01r_{\min}<.01. Finally make τ\tau small enough that the small-addition exponent is strictly below .008.008, that

3Δ+rmin⁡+εL(τ)<.04,3\Delta+r_{\min}+\varepsilon_L(\tau)<.04,

and that all the strict weight inequalities in (3.10) hold. These choices are possible because εL(τ)\varepsilon_L(\tau) and εhigh(τ)\varepsilon_{\mathrm{high}}(\tau) tend to zero for the fixed grid. They impose no condition on C∗C_\ast, since only the prefix and total regularity conditions were needed for this estimate.

There are finitely many grid classes and orientations and O(ℓ2)=Bo(1)O(\ell^2)=B^{o(1)} count pairs. Since ∣j∣≥ηT≍ηB.32|j|\ge\eta T\asymp_\eta B^{.32} and

κ1+.04−.32=−.047132048…<0,\kappa_1+.04-.32=-.047132048\ldots<0,

the sum of all the large-addition bounds is o(1)o(1). The small-addition bound, with a strict exponent below .008.008, is O(B.009)O(B^{.009}) after the o(1)o(1) in its exponent is absorbed. This proves the lemma, including the first candidate itself when valid.

Squaring the kernel and averaging the rows

The multiplicity estimate has controlled how many alternative representations survive after weighting a first one. Combining it with the small weight of each representation now gives the required integrated row-square bound.

Proof of Proposition 6.1. For an allowed edge i,k=i+ji,k=i+j, apply Lemma 6.2 to (47) and use the uniform bound on Ξ\Xi. This gives

ELik2≪B−κ1+5τlog⁡2+.009+o(1)1B∑(a,b,c)∈CjA(a)A(b)A(c)ab.\mathbb{E}L_{ik}^{2}\ll B^{-\kappa_1+5\tau\log2+.009+o(1)}\frac{1}{B}\sum_{(a,b,c)\in\mathcal{C}_j}\frac{A(a)A(b)A(c)}{ab}.

The remaining coefficient sum is bounded using (3.8). Cover the support of cc by O(B)O(B) dyadic boxes c≍Xc \asymp X, so a,b≍TXa,b \asymp TX in each box. On such a box, dropping coefficient cutoffs for an upper bound gives

1B∑a−b=jcA0(a)A0(b)A0(c)ab≪1B(TX)2TX2Σ(j)=Σ(j)BT.\frac{1}{B}\sum_{a-b=jc}\frac{A_0(a)A_0(b)A_0(c)}{ab}\ll\frac{1}{B(TX)^2}TX^2\Sigma(j)=\frac{\Sigma(j)}{BT}.

Summing the boxes yields O(Σ(j)/T)O(\Sigma(j)/T). The same proof works for negative jj by symmetry. Consequently

ELik2≪B−κ1+5τlog⁡2+.009+o(1)Σ(k−i)T.\mathbb{E}L_{ik}^2\ll B^{-\kappa_1+5\tau\log2+.009+o(1)}\frac{\Sigma(k-i)}{T}.

The choice in (3.10) gives κ1−5τlog⁡2=1−5hτ>.229\kappa_1-5\tau\log2=1-5h\tau>.229. Thus the exponent on the right is strictly below −.220-.220 before the o(1)o(1) term. Finally

1T∑0<∣j∣≤TΣ(j)≪1.\frac{1}{T}\sum_{0<|j|\leq T}\Sigma(j)\ll1.

so summing over a row proves (45), with room to absorb the o(1)o(1) in the exponent.

For the last assertion set Di=∑kLikD_i=\sum_k L_{ik}. Given SiS_i, the summands with different k≠ik\ne i are independent, because the other site types are independent and the roots have already been averaged out. Therefore

EVar⁡(Di∣Si)≤∑kELik2≪B−.21.\mathbb{E}\operatorname{Var}(D_i\mid S_i)\leq\sum_k\mathbb{E}L_{ik}^2\ll B^{-.21}.

The conditional mean is bounded for every SiS_i by (42). The variance identity consequently gives

EDi2=EVar⁡(Di∣Si)+E(E[Di∣Si])2≪1.\mathbb{E}D_i^2=\mathbb{E}\operatorname{Var}(D_i\mid S_i)+\mathbb{E}\bigl(\mathbb{E}[D_i\mid S_i]\bigr)^2\ll1.

Notice that only an integrated conditional variance bound has been used; a pointwise conditional second-moment estimate is unnecessary. Jensen’s inequality across the rows now gives

E(1M∑iDi)2≤1M∑iEDi2≪1.\mathbb{E}\left(\frac{1}{M}\sum_iD_i\right)^2\leq\frac{1}{M}\sum_i\mathbb{E}D_i^2\ll1.

For the kernel restricted to II as in the statement, nonnegativity can only reduce the conditional means, row-square sums, and total mass. Keeping the same normalization MM therefore gives the stated bounds.

Smoothing channels on logarithmic and residue space

The site model admits a useful smoothing operation: choose a fair subset of the primes at one site and record the logarithm and a residue of its product. We prove that this operation suppresses residue dependence in operator norm and that its logarithmic outputs are uniformly approximable on a fixed coarse partition. Operator norm, rather than convergence for each fixed test, is needed because the single-site tests used later may depend on the position and on an externally conditioned sample.

All limits in this section are as B→∞B\to\infty. The regularity parameters fixed in Section 3 remain fixed, although no regularity cutoff is imposed in the channels below. We use only the prime-product estimates (23)–(24) and the independent site law. An occurrence of xx inside a channel denotes a logarithmic coordinate, not the original counting scale.

The channel and its two-split kernel

Let ΩB\Omega_B be the finite set of subsets of P\mathcal{P}, with probability measure under which the inclusions p∈Sp \in S are independent and have probabilities 1/p1/p. Given SS, retain each of its primes independently with probability 1/21/2, and let bb be their product. Unconditionally bb has law ν1/2\nu_{1/2}. Write

I=[1/2,3],xb=log⁡bB.I=[1/2,3], \qquad x_b=\frac{\log b}{B}.

For a positive integer m1m_1, partition II into m1m_1 equal coarse intervals IlI_l. Choose the number nBn_B of fine intervals to be

nB=m1⌈∣I∣B11/10m1⌉,δB=∣I∣nB.n_B=m_1\left\lceil\frac{|I|B^{11/10}}{m_1}\right\rceil,\qquad\delta_B=\frac{|I|}{n_B}.

Thus the fine partition refines the coarse partition and δB≍B−11/10\delta_B\asymp B^{-11/10} for every fixed m1m_1. Use consistent half-open conventions, assigning the final endpoint to the last cell.

For an integer 1≤d≤B1\le d\le B, let Rd=(Z/dZ)×R_d=(\mathbb{Z}/d\mathbb{Z})^\times with uniform probability measure; for d=1d=1 this is the one-point group. All products in question are units modulo dd, since P0>BP_0>B for sufficiently large BB. Give I×RdI\times R_d the product measure

dλd(x,r)=dx dunif⁡Rd(r).d\lambda_d(x,r)=dx\,d\operatorname{unif}_{R_d}(r).

For g∈L2(ΩB)g\in L^2(\Omega_B), define a function constant on each fine logarithmic cell by

Udg(x,r)=φ(d)δBE[g(S)1{xb∈Ixfine, b≡r(modd)}],x∈Ixfine⊂I.(52)U_dg(x,r)=\frac{\varphi(d)}{\delta_B}\mathbb{E}\left[g(S)1_{\{x_b\in I_x^{\mathrm{fine}},\,b\equiv r\pmod d\}}\right],\qquad x\in I_x^{\mathrm{fine}}\subset I. \tag*{(52)}

The expectation in this definition includes both the site and its fair split. The domain norm of UdU_d is the site L2L^2 norm, and its range norm is that of L2(λd)L^2(\lambda_d).

Let PavP_{\mathrm{av}} average the residue coordinate. Then

ug:=PavUdg=U1gu_g:=P_{\mathrm{av}}U_dg=U_1g

is independent of dd. Let Pm1P_{m_1} average on the coarse logarithmic intervals and act identically on the residue coordinate when one is present. Let PfineP_{\mathrm{fine}} denote averaging on the fine logarithmic intervals. These averaging operators are orthogonal projections, and PfineP_{\mathrm{fine}} commutes with Pm1P_{m_1}.

Proposition 7.1 (Uniform channel bounds). For every fixed coarse partition, uniformly for 1≤d≤B1\le d\le B,

∥Ud∥≪1,∥(1−Pav)Ud∥≪B−c7,c7=1200.(53)\lVert U_d\rVert\ll1,\qquad\lVert(1-P_{\mathrm{av}})U_d\rVert\ll B^{-c_7},\qquad c_7=\frac{1}{200}. \tag*{(53)}

Moreover, for every ϵ1>0\epsilon_1>0, one can choose m1m_1 such that

lim sup⁡B→∞sup⁡∥g∥2≤1∥ug−Pm1ug∥2≤ϵ1.(54)\limsup_{B\to\infty}\sup_{\lVert g\rVert_2\le1}\lVert u_g-P_{m_1}u_g\rVert_2\le\epsilon_1. \tag*{(54)}

All constants are independent of the input gg. The sufficiently large threshold for BB may depend on the chosen fixed partition.

We first identify the kernel used to prove the proposition. A cell A=J×{r}A=J\times\{r\}, with JJ a fine logarithmic interval, has measure

λd(A)=δBφ(d)=:μd.\lambda_d(A)=\frac{\delta_B}{\varphi(d)}=:\mu_d.

Set

pA(S)=Psplit⁡(xb∈J, b≡r(modd)∣S).p_A(S)=\mathbb{P}_{\operatorname{split}}(x_b \in J,\ b \equiv r \pmod d \mid S).

For a step function FF in the range space, the adjoint of the channel is

Ud∗F(S)=∑ApA(S)F∣A.U_d^*F(S)=\sum_A p_A(S)F|_A.

Consequently UdUd∗U_dU_d^* has the step kernel

Kd(A,A′)=E[pA(S)pA′(S)]μd2.(55)K_d(A,A')=\frac{\mathbb{E}[p_A(S)p_{A'}(S)]}{\mu_d^2}. \tag*{(55)}

This is the density on fine cells of two conditionally independent fair splits of the same site, restricted to II in each logarithmic coordinate. The kernel is nonnegative and symmetric. Its row integral satisfies

∑A′Kd(A,A′)λd(A′)=E[pA(S)∑A′pA′(S)]μd≤Pν1/2(xb∈J, b≡r(modd))μd.\sum_{A'}K_d(A,A')\lambda_d(A')=\frac{\mathbb{E}\left[p_A(S)\sum_{A'}p_{A'}(S)\right]}{\mu_d}\leq\frac{\mathbb{P}_{\nu_{1/2}}(x_b\in J,\ b\equiv r\pmod d)}{\mu_d}.

By (23), the numerator is at most CδB/φ(d)+O(B−80)C\delta_B/\varphi(d)+O(B^{-80}). Indeed its logarithmic interval is contained in the fixed compact interval II, where the densities are uniformly bounded. Since μd−1≪B21/10\mu_d^{-1}\ll B^{21/10}, the row integrals are bounded uniformly. The column integrals have the same bound. Schur’s inequality gives ∥UdUd∗∥≪1\lVert U_dU_d^*\rVert\ll1, and therefore proves the first assertion of (53). This proof already applies to every L2L^2 input, including signed or complex inputs.

Lemma 7.2 (Independent common and exclusive products). Let b1,b2b_1,b_2 be two conditionally independent fair splits of the same site. Their joint law can be coupled, with failure probability O(P0−1)O(P_0^{-1}), to

(CE1,CE2),(CE_1,CE_2),

where C,E1,E2C,E_1,E_2 are independent products with law ν1/4\nu_{1/4}. Replacing the two-split step kernel by this independent-product step kernel changes its operator norm by O(B5/P0)O(B^5/P_0), uniformly for d≤Bd\leq B.

Proof. At a fixed prime pp, the categories common to both splits, exclusive to the first, and exclusive to the second have probabilities 1/(4p)1/(4p) each, and are mutually exclusive. The joint law of three independent Bernoulli indicators with these marginal probabilities differs from this categorical law in total variation by O(p−2)O(p^{-2}). For example, the probability of two or more independent successes is O(p−2)O(p^{-2}), and each remaining probability differs from its categorical value by O(p−2)O(p^{-2}). Couple these laws independently over the primes. A union bound gives total failure probability

O(∑p>P0p−2)=O(P0−1).O\left(\sum_{p>P_0}p^{-2}\right)=O(P_0^{-1}).

On successful coupling the products agree. Collisions that would put a squared prime in one of the independent products are included in the failure event.

If two joint probability laws differ in total variation by ε\varepsilon, each cell-pair probability differs by at most a constant times ε\varepsilon. Their step kernel difference is thus bounded pointwise by O(εμd−2)O(\varepsilon\mu_d^{-2}). The total measure of the output space is ∣I∣|I|, so Schur’s inequality bounds the operator difference by the same expression times ∣I∣|I|. Now

μd−2≪B42/10≪B5.\mu_d^{-2}\ll B^{42/10}\ll B^5.

Taking ε=O(P0−1)\varepsilon=O(P_0^{-1}) proves the assertion.

Removing small exclusive products

The two-split model reduces the channel operator to independent common and exclusive products. An exclusive product near 1 provides too little averaging to erase its residue, so we first bound the operator contribution of those products. Both row and column bounds are needed.

For a product vv, write ℓB(v)=(log⁡v)/B\ell_B(v) = (\log v)/B. Let KdindK_d^{\mathrm{ind}} be the independent-product kernel from Lemma 7.2. Its entries are

Kdind(A,A′)=P((ℓB(CE1),CE1 mod d)∈A,  (ℓB(CE2),CE2 mod d)∈A′)μd2.K_d^{\mathrm{ind}}(A,A') = \frac{\mathbb{P}\bigl((\ell_B(CE_1),CE_1 \bmod d)\in A,\;(\ell_B(CE_2),CE_2 \bmod d)\in A'\bigr)}{\mu_d^2}.

The following estimate concerns the operator obtained by retaining only the specified event in this probability.

Lemma 7.3 (Two-sided Schur estimate for small exclusives). For 0<θ≤1/100 < \theta\le1/10, the contribution to KdindK_d^{\mathrm{ind}} from ℓB(Ei)≤θ\ell_B(E_i) \le\theta, for either i=1i = 1 or i=2i = 2, has operator norm

≪(θ+1/R)1/4+B−70+B5/P0,\ll(\theta+ 1/R)^{1/4} + B^{-70} + B^5/P_0,

uniformly for d≤Bd \le B. The same bound, with a different absolute constant, holds for the union of the two events.

Proof. It suffices to consider i=2i = 2. Put

σθ=Pν1/4(ℓB(E2)≤θ)≪(θ+1/R)1/4,\sigma_\theta= \mathbb{P}_{\nu_{1/4}}(\ell_B(E_2) \le\theta) \ll(\theta+ 1/R)^{1/4},

where the last inequality is (24). The row integral at A=J×{r}A = J \times\{r\} of the restricted kernel is at most

P((ℓB(CE1),CE1 mod d)∈A,  ℓB(E2)≤θ)μd=σθP((ℓB(CE1),CE1 mod d)∈A)μd.\frac{\mathbb{P}\bigl((\ell_B(CE_1),CE_1 \bmod d)\in A,\;\ell_B(E_2)\le\theta\bigr)}{\mu_d} = \sigma_\theta\frac{\mathbb{P}\bigl((\ell_B(CE_1),CE_1 \bmod d)\in A\bigr)}{\mu_d}.

Here E2E_2 is independent of C,E1C,E_1. The marginal law of CE1CE_1 differs from ν1/2\nu_{1/2} by O(P0−1)O(P_0^{-1}) in total variation: the independent common and exclusive indicators can be coupled to their mutually exclusive counterparts prime by prime exactly as in the preceding lemma. Thus (23) and μd−1≪B21/10\mu_d^{-1} \ll B^{21/10} bound the last display by

Cσθ+O(B5/P0)+O(B−70).C\sigma_\theta+ O(B^5/P_0) + O(B^{-70}).

For the column integral at A′=J′×{r′}A' = J' \times\{r'\}, first omit the condition that CE1CE_1 has logarithm in II. Condition on E2=eE_2 = e with ℓB(e)≤θ\ell_B(e) \le\theta. The remaining condition on CC is

ℓB(C)∈J′−ℓB(e),C≡r′e−1(modd).\ell_B(C) \in J' - \ell_B(e), \qquad C \equiv r'e^{-1} \pmod d.

The shifted interval lies in [2/5,3][2/5,3], has length δB\delta_B, and the prescribed residue is a unit. Applying (23) with z=1/4z = 1/4, uniformly in ee, gives conditional probability at most

CδBφ(d)+O(B−80).C\frac{\delta_B}{\varphi(d)} + O(B^{-80}).

Integration over the event ℓB(E2)≤θ\ell_B(E_2) \le\theta and division by μd\mu_d therefore give a column bound Cσθ+O(B−70)C\sigma_\theta+ O(B^{-70}). Schur’s inequality proves the claimed operator bound. The i=1i = 1 case follows on transposing the kernel. The kernel of the union is nonnegative and is entrywise at most the sum of the two restricted kernels, so the row and column estimates also prove its bound. □

The separate column estimate is essential: an estimate for the total probability of a discarded event would not by itself control its operator norm.

Flattening all nonconstant residue modes

We prove the second assertion of (53). Set θ=B−1/10\theta= B^{-1/10} and retain the part of KdindK_d^{\mathrm{ind}} where ℓB(E1),ℓB(E2)>θ\ell_B(E_1),\ell_B(E_2)>\theta. For a fixed common product C=cC=c, the probability that one retained exclusive lands in the output cell A=J×{r}A=J\times\{r\} is

pA(θ)(c)=P(ℓB(E)∈(J−ℓB(c))∩(θ,∞), E≡rc−1(modd)).p_A^{(\theta)}(c)=\mathbb{P}\left(\ell_B(E)\in(J-\ell_B(c))\cap(\theta,\infty),\ E\equiv rc^{-1}\pmod d\right).

If the interval is nonempty it is contained in [B−1/10,3][B^{-1/10},3], since ℓB(c)≥0\ell_B(c)\ge0. The interval and its endpoint conventions, including a partial cell at θ\theta, are within the uniform statement of (23). Consequently

pA(θ)(c)=1φ(d)∫(J−ℓB(c))∩(θ,∞)f1/4,B(t) dt+O(B−80).(56)p_A^{(\theta)}(c)=\frac{1}{\varphi(d)}\int_{(J-\ell_B(c))\cap(\theta,\infty)}f_{1/4,B}(t)\,dt+O(B^{-80}). \tag*{(56)}

The main term depends on JJ and ℓB(c)\ell_B(c), but not on rr or on the residue of cc. All terms are uniformly bounded: the probability is at most one and its main term differs by O(B−80)O(B^{-80}). Alternatively the bound f1/4,B(t)≪t−3/4f_{1/4,B}(t)\ll t^{-3/4} gives the explicit cell bound O(δBB3/40/φ(d))O(\delta_BB^{3/40}/\varphi(d)).

Conditional on CC, the exclusives are independent. Multiplying (56) for two output cells and integrating over CC, the retained kernel differs from a kernel independent of both residues by at most

O(B−80μd−2)=O(B−75)O(B^{-80}\mu_d^{-2})=O(B^{-75})

pointwise, and hence by O(B−70)O(B^{-70}) in operator norm. A kernel independent of both residues is annihilated by projection onto 1−Pav1-P_{\mathrm{av}} on either side. Lemmas 7.2 and 7.3 therefore yield

∥(1−Pav)UdUd∗(1−Pav)∥≪(B−1/10+1/R)1/4+B−70+B5/P0≪B−1/40,\left\|(1-P_{\mathrm{av}})U_dU_d^*(1-P_{\mathrm{av}})\right\| \ll(B^{-1/10}+1/R)^{1/4}+B^{-70}+B^5/P_0 \ll B^{-1/40},

because 1/R=(log⁡P0)/B=O(log⁡B/B)1/R=(\log P_0)/B=O(\log B/B) and P0=B1000P_0=B^{1000}. Taking the square root of this operator identity gives the stronger bound

∥(1−Pav)Ud∥≪B−1/80.\left\|(1-P_{\mathrm{av}})U_d\right\|\ll B^{-1/80}.

In particular (53) holds with c7=1/200c_7=1/200. This reasoning did not fix an input function at any stage, so the estimate is uniform on the whole unit ball of L2(ΩB)L^2(\Omega_B).

Uniform coarse compactness

Residue dependence is now negligible uniformly in the input function. The remaining task is to approximate the logarithmic output uniformly on a fixed finite partition. This will give a finite family of endpoint features for the graph kernel.

We now prove (54). It suffices to use d=1d=1, since ug=U1gu_g=U_1g. Fix 0<θ<1/200<\theta<1/20 and choose a smooth function aθ:R→[0,1]a_\theta:\mathbb{R}\to[0,1] equal to zero on (−∞,θ](-\infty,\theta] and to one on [2θ,∞)[2\theta,\infty). It may be chosen nondecreasing. In the independent-product kernel insert the weight aθ(ℓB(E1))aθ(ℓB(E2))a_\theta(\ell_B(E_1))a_\theta(\ell_B(E_2)). The change is a nonnegative kernel supported where at least one exclusive logarithm is at most 2θ2\theta. Lemma 7.3 bounds its operator norm by

O((2θ+1/R)1/4)+O(B−70+B5/P0).O((2\theta+1/R)^{1/4})+O(B^{-70}+B^5/P_0).

Choose also a smooth upper cutoff equal to one on [0,3][0,3] and supported below 16/516/5. On the positive axis let f1/4,Bf_{1/4,B} be the product of this upper cutoff, aθa_\theta, and f1/4,Bf_{1/4,B}, extended by zero to the real line. For each fixed θ\theta, all derivatives of any fixed order of this function are bounded uniformly in BB. This follows from the derivative bounds for f1/4,Bf_{1/4,B} established before Proposition 3.3; the cutoff keeps the argument away from zero. The upper cutoff changes none of the probabilities relevant to the output interval II.

For completeness, the weighted version of the local law used here follows directly from its interval version. For fixed z=ℓB(c)z=\ell_B(c) and a fine interval JJ, apply (23) to subintervals of (J−z)∩[θ,3](J-z)\cap[\theta,3]. The difference between the exclusive-product measure and the density f1/4,B(t) dtf_{1/4,B}(t)\,dt has cumulative integral O(B−80)O(B^{-80}) there. Stieltjes integration by parts against aθa_\theta bounds its weighted integral by Oθ(B−80)O_\theta(B^{-80}): the endpoint values and the total variation of aθa_\theta are bounded. Thus, uniformly in c,Jc,J,

E[aθ(ℓB(E))1{ℓB(cE)∈J}]=∫Jf1/4,Bcut(x−ℓB(c)) dx+Oθ(B−80).\mathbb{E}\left[a_\theta\bigl(\ell_B(E)\bigr)\mathbf{1}_{\{\ell_B(cE)\in J\}}\right]=\int_J f^{\mathrm{cut}}_{1/4,B}\bigl(x-\ell_B(c)\bigr)\,dx+O_\theta(B^{-80}).

Conditional independence of the two exclusives now identifies the weighted step kernel, up to Oθ(B−70)O_\theta(B^{-70}) in operator norm, with the fine-cell compression of the continuous kernel

KB,θ(x,y)=EC[f1/4,Bcut(x−ℓB(C))f1/4,Bcut(y−ℓB(C))],(x,y)∈I2.(57)K_{B,\theta}(x,y)=\mathbb{E}_{C}\left[f^{\mathrm{cut}}_{1/4,B}\bigl(x-\ell_B(C)\bigr)f^{\mathrm{cut}}_{1/4,B}\bigl(y-\ell_B(C)\bigr)\right],\qquad(x,y)\in I^2. \tag*{(57)}

Indeed multiplying the two weighted cell formulas makes an Oθ(B−80)O_\theta(B^{-80}) error in each cell-pair probability; division by δ2\delta^2 still leaves an operator error smaller than Oθ(B−70)O_\theta(B^{-70}).

The kernel in (57) and its first derivatives are bounded uniformly in BB, with constants depending on θ\theta. Differentiation may be taken under the expectation because the cut densities and their derivatives have uniform bounds and the law of CC is a probability measure. In particular no limiting law for CC, and no derivative bound for that law, is needed. Let TB,θT_{B,\theta} be the integral operator with this kernel. If a=∣I∣/m1a=|I|/m_1 is the coarse mesh, the mean value theorem gives

sup⁡x,y∈I∣KB,θ(x,y)−(Pm1(x)Pm1(y)KB,θ)(x,y)∣≪θa.\sup_{x,y\in I}\left|K_{B,\theta}(x,y)-\left(P_{m_1}^{(x)}P_{m_1}^{(y)}K_{B,\theta}\right)(x,y)\right|\ll_\theta a.

Here the two superscripts denote averaging the indicated kernel coordinate. Schur’s inequality consequently gives

∥TB,θ−Pm1TB,θPm1∥≪θm1−1.(58)\left\|T_{B,\theta}-P_{m_1}T_{B,\theta}P_{m_1}\right\|\ll_\theta m_1^{-1}. \tag*{(58)}

Write Q=1−Pm1Q=1-P_{m_1}. Fine and coarse averaging commute, and fine averaging is a contraction. Therefore

∥QPfineTB,θPfineQ∥≤∥QTB,θQ∥≪θm1−1.\left\|QP_{\mathrm{fine}}T_{B,\theta}P_{\mathrm{fine}}Q\right\|\leq\left\|QT_{B,\theta}Q\right\|\ll_\theta m_1^{-1}.

Combining this with the independent-model comparison and the cutoff estimate proves, for every fixed θ,m1\theta,m_1,

∥QU1∥2=∥QU1U1∗Q∥≤C(2θ+1/R)1/4+Cθm1−1+oθ,m1(1).\|QU_1\|^2=\|QU_1U_1^*Q\| \leq C(2\theta+1/R)^{1/4}+C_\theta m_1^{-1}+o_{\theta,m_1}(1).

Given ϵ1>0\epsilon_1>0, first choose θ\theta so that C(2θ)1/4<ϵ12/4C(2\theta)^{1/4}<\epsilon_1^2/4, and then choose m1m_1 so that Cθ/m1<ϵ12/4C_\theta/m_1<\epsilon_1^2/4. Taking the upper limit in BB and the square root proves (54), and completes the proof of Proposition 7.1.

Corollary 7.4 (Coarse site features). For the chosen coarse partition, define

VI(S)=1∣I∣Psplit(xb∈I∣S).V_I(S)=\frac{1}{|I|}\mathbb{P}_{\mathrm{split}}(x_b\in I\mid S).

Then 0≤Vl≤∣Il∣−10 \le V_l \le|I_l|^{-1}, and for every g∈L2(ΩB)g \in L^2(\Omega_B) the coarse value of Pm1ugP_{m_1}u_g on IlI_l is exactly

E[g(S)Vl(S)].\mathbb{E}[g(S)V_l(S)].

In particular the bounds (53)–(54) and this identity hold uniformly when gg varies with a position, with BB, or with an external parameter.

Proof. Integrating (52) with d=1d=1 over IlI_l sums precisely the fine cells contained in IlI_l. The result is E[g(S)Psplit(xb∈Il∣S)]\mathbb{E}[g(S)\mathbb{P}_{\mathrm{split}}(x_b \in I_l \mid S)]. Divide by ∣Il∣|I_l|. The bounds on VlV_l follow from its definition; the final uniformity follows because the preceding results are operator bounds rather than fixed-input limits.

Integral approximation by endpoint features

We now approximate the latent matrix L\mathcal{L} from (41) after integrating over its endpoint types. For an allowed pair k=i+jk=i+j with j>0j>0, let Si,SkS_i,S_k be independent site types and let g,h:ΩB→[−1,1]g,h:\Omega_B \to[-1,1] be real measurable functions. The quantity to be approximated is

ESi,Sk[Lik(Si,Sk;s)g(Si)h(Sk)].\mathbb{E}_{S_i,S_k}\left[\mathcal{L}_{ik}(S_i,S_k;s)g(S_i)h(S_k)\right].

For a coarse partition supplied by (54) at a chosen accuracy ϵ1\epsilon_1, let VlV_l be the features from Corollary 7.4. Our target is a finite sum of scalar multiples of

E[g(S)Vl(S)]E[h(S)Vl(S)].\mathbb{E}[g(S)V_l(S)]\mathbb{E}[h(S)V_l(S)].

The coefficients may depend on B,j,s,lB,j,s,l, but not on the tests. The partition is fixed before BB tends to infinity. The comparison will be uniform in the tests, allowing a different function at every position. After summing over allowed pairs and dividing by MM, its error will be controlled by the loss e−c3C∗e^{-c_3C_*} from removing regularity restrictions, the chosen coarse accuracy ϵ1\epsilon_1, and a term tending to zero. The final proposition records the constants and their parameter dependence. Section 9 then converts this integrated comparison into control for tests chosen after the whole matrix is sampled.

Throughout this section the regularity grid and its tolerances are fixed, as are C∗C_*, η>0\eta>0, and C6C_6. The parameter BB tends to infinity, T=⌊B0.32⌋T=\lfloor B^{0.32}\rfloor, and ηT≤j≤T\eta T \le j \le T. All estimates involving the smooth parameter ss are uniform on the fixed compact interval containing its support. Constants with a subscript η\eta may depend on η\eta; constants without that subscript in the cutoff-removal estimate below do not. The parameter xx from the original counting problem does not occur in this section. The letters x,yx,y below denote log coordinates in I=[1/2,3]I=[1/2,3].

Removing the regularity restrictions

Let S∈ΩBS \in\Omega_B have the independent-site law: each prime in P\mathcal{P} belongs to SS independently with probability 1/p1/p. A fair split of SS selects each of its primes independently with probability 1/21/2; write bb for the product of the selected primes. The unconditional law of bb is ν1/2\nu_{1/2}. For a real measurable function gg on the site space, define the signed mass

γg(n)=E[g(S)1{b=n}].\gamma_g(n)=\mathbb{E}\left[g(S)\mathbf{1}_{\{b=n\}}\right].

In particular, ∣γg(n)∣≤ν1/2(n)|\gamma_g(n)|\le\nu_{1/2}(n) whenever ∣g∣≤1|g|\le1. All statements below allow gg to depend on BB. For a subset product bb of SS, the untruncated weights satisfy the exact identity

A0(b)K0(S∖b)=(log⁡P0)R2−∣S∣=B2−∣S∣.A_0(b)K_0(S\setminus b)=(\log P_0)R2^{-|S|}=B2^{-|S|}.

The factor 2−∣S∣2^{-|S|} is the conditional probability of each fair split. This identity connects the endpoint weights with the unrestricted splits defining the channels. The next lemma bounds the cost of removing the regularity indicators inside the integrated test; after that removal, the endpoint averages are the unrestricted ones defining the channels.

Lemma 8.1 (Uniform removal of cutoffs). Let k=i+jk=i+j be an allowed pair with j>0j>0, and let ∣g∣,∣h∣≤1|g|,|h|\le1. Replacing the coefficient and remaining-site regularity restrictions in E[Likg(Si)h(Sk)]\mathbb{E}[\mathcal{L}_{ik}g(S_i)h(S_k)] by no restrictions, while keeping the scalar kBk_B, incurs an absolute error at most

CΣ(j)T(rB+e−c3C∗),rB⟶0.(59)C\frac{\Sigma^{(j)}}{T}\left(r_B+e^{-c_3C_*}\right),\qquad r_B\longrightarrow0. \tag*{(59)}

Here rBr_B may depend on the fixed regularity parameters and on C∗C_*, but the constant CC is independent of C∗C_*, η\eta, and C6C_6. The pairwise gcd restrictions may then be removed at an additional error O(B−998)O(B^{-998}). The resulting integral is

kBB∑a−b=jcγg(b)γh(a)A0(c)ρ0(log⁡c/B)Ψs(a/(Tc),b/(Tc)).(60)k_BB\sum_{a-b=jc}\gamma_g(b)\gamma_h(a)A_0(c)\rho_0(\log c/B)\Psi_s\left(a/(Tc),b/(Tc)\right). \tag*{(60)}

The sum is over positive integers a,b,ca,b,c.

Proof. Applying the subset identity above at both endpoints in (41) turns their subset sums into fair-split expectations, still with regularity indicators attached to the selected and unselected sets, and changes the coefficient kB/Bk_B/B into kBBk_BB. The remaining coefficient is c=(a−b)/jc=(a-b)/j, and its weight is A0(c)A_0(c) times its regularity indicator.

All weights other than g,hg,h are nonnegative. We may therefore bound the cost of removing an indicator by putting ∣g∣,∣h∣≤1|g|,|h|\le1 and summing the corresponding positive mass. On the support under consideration, c≍Xc\asymp X, a,b≍TXa,b\asymp TX, and 0.9B≤log⁡X≤2.2B0.9B\le\log X\le2.2B. There are O(B)O(B) dyadic choices of XX. Formula (3.5) bounds the split-product probabilities by

ν1/2(a)ν1/2(b)≪A0(a)A0(b)B2ab.\nu_{1/2}(a)\nu_{1/2}(b)\ll\frac{A_0(a)A_0(b)}{B^2ab}.

Consequently the untruncated mass in one such box is at most

CB∑a−b=jcc≍X, a,b≍TXA0(a)A0(b)A0(c)ab≪Σ(j)BT\frac{C}{B} \sum_{\substack{a-b=jc\\ c\asymp X,\ a,b\asymp TX}} \frac{A_0(a)A_0(b)A_0(c)}{ab} \ll\frac{\Sigma^{(j)}}{BT}

by (3.8), since ab≍T2X2ab\asymp T^2X^2. If any one of the three coefficients is nonregular, (3.11) gives the same bound with the additional factor rB+e−c3C∗r_B+e^{-c_3C_*}. Summing over the boxes proves the required estimate for coefficient failures.

It remains to treat the two unselected endpoint sets. At a given prime pp, the probabilities of the three possibilities “selected”, “unselected”, and “absent” are respectively 1/(2p)1/(2p), 1/(2p)1/(2p), 1−1/p1-1/p. Conditional on the selected product being aa, the unselected indicators at primes not dividing aa are therefore independent with probabilities

1/(2p)1−1/(2p)=12p−1.\frac{1/(2p)}{1-1/(2p)}=\frac{1}{2p-1}.

At primes dividing aa they are zero. The same statement holds for bb, independently at the other site. Each selected coefficient on our support has log⁡\log size O(B)O(B), and all its prime factors exceed P0P_0. In particular, the reciprocal sum of its prime factors is O(B/P0)O(B/P_0). The uniform cutoff probability estimate established after (3.11) thus bounds either unselected-set failure by rB+O(e−c3C∗)r_B + O(e^{-c_3 C_*}), uniformly in the chosen coefficients. Multiplying by the preceding positive mass bound and summing the boxes proves (59). This argument uses (3.8) and (3.11) for 1≤j≤T1 \le j \le T; it has made no use of j≥ηTj \ge\eta T or of the block length MM. This proves the asserted independence of its constant.

Finally, if one of (a,b)(a,b), (a,c)(a,c), (b,c)(b,c) exceeds one, a prime dividing both aa and bb must occur. Indeed a−b=jca-b=jc, and every prime factor of a coefficient is greater than P0>jP_0>j. The split products at the two sites are independent, so the probability of a common selected prime is at most

∑p∈P14p2≪P0−1.\sum_{p\in\mathcal{P}}\frac{1}{4p^2}\ll P_0^{-1}.

For a fixed pair of selected products there is at most one value of cc, and the remaining positive integrand is at most CBsup⁡cA0(c)CB\sup_c A_0(c). Since sup⁡A0≤(log⁡P0)R1/2=B1/2+o(1)\sup A_0\le(\log P_0)R^{1/2}=B^{1/2+o(1)}, the resulting error is O(B3/2+o(1)/P0)O(B^{3/2+o(1)}/P_0). After removal of these restrictions, the fair splits and independence of the two sites give (60) exactly.

Fourier detection and the discarded arcs

The support of (60) already predicts the structure of its real-variable comparison. For a nonzero term put

x=log⁡aB,y=log⁡bB,u1=aTc,u2=bTc.x=\frac{\log a}{B},\qquad y=\frac{\log b}{B},\qquad u_1=\frac{a}{Tc},\qquad u_2=\frac{b}{Tc}.

The exact relation a−b=jca-b=jc gives

u1−u2=jT,B(x−y)=log⁡(u1/u2).u_1-u_2=\frac{j}{T},\qquad B(x-y)=\log(u_1/u_2).

Since u1,u2u_1,u_2 lie in a fixed compact subinterval of (1,2)(1,2), the two log coordinates satisfy 0<x−y≪B−10<x-y\ll B^{-1}. We will enforce the discrete relation by additive Fourier orthogonality. This places the cc sum in a factor to which the arithmetic estimates apply, while the two signed endpoint sums are controlled by their mass and L2L^2 bounds. For the surviving small denominators, the channel estimates allow us to average the residue dependence of the endpoint tests with a uniform error, leaving an explicit finite residue sum. Fourier inversion will then give an integral on the narrow logarithmic band, where the coarse channel approximation produces the feature products described above.

We first analyze (60) in a single smooth dyadic box. Choose a smooth compactly supported dyadic partition of unity on (0,∞)(0,\infty) for the variable cc, and write X=2mX=2^m, NX=TXN_X=TX. A box has c/Xc/X in a fixed compact subinterval of (0,∞)(0,\infty). On the support of Ψs\Psi_s, both a/NXa/N_X and b/NXb/N_X are also in fixed positive compact intervals. The smooth weight in the three variables

(a/NX,b/NX,c/X)(a/N_X,b/N_X,c/X)

has all fixed-order derivatives bounded uniformly in m,B,sm,B,s. This includes ρ0((log⁡X+log⁡(c/X))/B)\rho_0((\log X+\log(c/X))/B), whose differentiated log-scale factors only improve the bound.

Extend this weight smoothly inside a larger fixed cube, expand the extension in a periodic Fourier series, and multiply each coordinate factor by a fixed smooth compact cutoff equal to one on the original support. We obtain a sum of separated products

w1(a/NX)w2(b/NX)w3(c/X).w_1(a/N_X)w_2(b/N_X)w_3(c/X).

More precisely, if ν∈Z3\nu\in\mathbb{Z}^3 indexes the Fourier terms, their scalar coefficients are OA((1+∣ν∣)−A)O_A((1+|\nu|)^{-A}) for every fixed AA, uniformly in the boxes and in ss. The fixed-order smooth norms of the factors grow at most polynomially in ∣ν∣|\nu|. All estimates below have only finitely many such smooth-norm losses; choosing AA larger than those losses plus four makes every Fourier sum absolutely convergent. When a bound is summed over boxes, we use the uniform coefficient bound at each fixed ν\nu. It therefore suffices to write the calculation for one separated term.

For that term set

S1(α)=∑aγh(a)w1(a/NX)e(aα),S2(α)=∑bγg(b)w2(b/NX)e(bα),S3(α)=∑cA0(c)w3(c/X)e(cα).\begin{aligned} S_1(\alpha) &= \sum_a \gamma_h(a)w_1(a/N_X)e(a\alpha),\\ S_2(\alpha) &= \sum_b \gamma_g(b)w_2(b/N_X)e(b\alpha),\\ S_3(\alpha) &= \sum_c A_0(c)w_3(c/X)e(c\alpha). \end{aligned}

Additive orthogonality expresses its contribution as

kBB∫01S1(α)S2(−α)S3(−jα) dα.k_B B\int_0^1 S_1(\alpha)S_2(-\alpha)S_3(-j\alpha)\,d\alpha.

The elementary product formula and Mertens’ estimate give

0≤kB≤EK0(S)=R1/2∏p∈P(1−12p)≪1,0 \leq k_B \leq E K_0(S)=R^{1/2}\prod_{p\in\mathcal{P}}\left(1-\frac{1}{2p}\right)\ll1,

uniformly in the truncation constant. By (3.5), (3.6), and Parseval,

∫01∣Si(α)∣2 dα≪1B2NX2∑n≍NXA0(n)2≪B−2+1/4+o(1)NX,i=1,2.(61)\int_0^1 |S_i(\alpha)|^2\,d\alpha\ll\frac{1}{B^2N_X^2}\sum_{n\asymp N_X} A_0(n)^2 \ll\frac{B^{-2+1/4+o(1)}}{N_X},\qquad i=1,2. \tag*{(61)}

The first-moment part of (3.6) also gives

sup⁡α∣Si(α)∣≪B−1,i=1,2.\sup_\alpha|S_i(\alpha)|\ll B^{-1},\qquad i=1,2.

The implicit constants here and below include the specified smooth norms of the separated factors. On the circle of jαj\alpha (mod 1), take major arcs of radius B13/XB^{13}/X about the reduced fractions u/qu/q with q≤B12q\leq B^{12}. They are disjoint for large BB: the distance between distinct such fractions is at least B−24B^{-24}, whereas XX is exponential in BB. On the complement we have

∣S3(−jα)∣≪XB−1/2+o(1).|S_3(-j\alpha)|\ll XB^{-1/2+o(1)}.

Here are the details of the imported estimate and its application. The Montgomery–Vaughan bound [20] states that a multiplicative function ff satisfying ∣f(p)∣≤1|f(p)|\leq1 for every prime pp and ∑n≤Z∣f(n)∣2≤Z\sum_{n\leq Z}|f(n)|^2\leq Z for every Z≥1Z\geq1 obeys

∣∑n≤Yf(n)e(nθ)∣≪γlog⁡Y+γ(log⁡R′)3/2R′\left|\sum_{n\leq Y}f(n)e(n\theta)\right|\ll\frac{\gamma}{\log Y}+\frac{\gamma(\log R')^{3/2}}{\sqrt{R'}}

whenever ∣θ−u′/q′∣≤1/(q′)2|\theta-u'/q'|\leq1/(q')^2, (u′,q′)=1(u',q')=1, and 2≤R′≤q′≤Y/R′2\leq R'\leq q'\leq Y/R'. Apply Dirichlet approximation with denominator limit ⌊X/B12⌋\lfloor X/B^{12}\rfloor. If the resulting denominator is at most B12B^{12}, the approximation error is at most 2B12/X<B13/X2B^{12}/X < B^{13}/X, so the point is on a major arc. Off the major arcs we consequently have

B12<q′≤X/B12,∣θ−u′q′∣≤1/(q′)2.B^{12} < q' \le X/B^{12}, \qquad\left|\theta- \frac{u'}{q'}\right| \le1/(q')^2.

For every Y≍XY \asymp X arising in partial summation, these inequalities permit R′=B10R'=B^{10}. The function f(n)=μMob2(n)2−ω(n)1{n rough}f(n)=\mu_{\mathrm{Mob}}^2(n)2^{-\omega(n)}1_{\{n\ \mathrm{rough}\}} is multiplicative and 1-bounded, so it satisfies the theorem’s hypotheses. Smooth partial summation and restoration of the factor

(log⁡P0)R1/2=Blog⁡P0=B1/2+o(1)(\log P_0)R^{1/2}=\sqrt{B\log P_0}=B^{1/2+o(1)}

give (8.5).

For later reference, if a discarded region satisfies ∣S3∣≪XB−a+o(1)\lvert S_3\rvert\ll XB^{-a+o(1)}, its contribution in one box is at most

CBXB−a+o(1)B−7/4+o(1)TX=B−3/4−a+o(1)T,CBXB^{-a+o(1)}\frac{B^{-7/4+o(1)}}{TX} =\frac{B^{-3/4-a+o(1)}}{T},

by Cauchy–Schwarz and (61). Summing over the O(B)O(B) boxes gives a total minor-arc error of

B−1/4+o(1)T.\frac{B^{-1/4+o(1)}}{T}.

On a major arc, apply (22) to S3S_3. When q>B0.4q>B^{0.4}, its main term is at most X/φ(q)≪XB−0.4+o(1)X/\varphi(q)\ll XB^{-0.4+o(1)}, uniformly along the arc. Formula (8.6) shows that these arcs cost B−0.15+o(1)/TB^{-0.15+o(1)}/T in total. The error O(XB−50)O(XB^{-50}) in (22), on all the major arcs together, costs at most B−49.75+o(1)/TB^{-49.75+o(1)}/T. All three errors are uniform in the endpoint tests and are o(1/T)o(1/T).

Put Q=B0.4Q=B^{0.4}. The remaining arcs for α\alpha have the following exact parametrization:

d=jq,h′(modd),(h′,q)=1,α=h′d+ξjX,∣ξ∣≤B13.d=jq,\qquad h'\pmod d,\qquad(h',q)=1,\qquad\alpha=\frac{h'}{d}+\frac{\xi}{jX},\qquad|\xi|\le B^{13}.

For each reduced fraction on the jαj\alpha circle there are exactly jj lifts in this list. In particular the parametrization neither omits nor repeats an arc. Define

w~3,B,X(ξ)=∫w3(z)mB(log⁡(Xz)/B)e(−ξz) dz.\widetilde w_{3,B,X}(\xi)=\int w_3(z)m_B(\log(Xz)/B)e(-\xi z)\,dz.

These transforms are uniformly Schwartz, with bounds controlled by fixed smooth norms. The main term from (22) contains the factor XX, while dα=dξ/(jX)d\alpha=d\xi/(jX). The resulting box expression is therefore

kBBj∑q≤QμMob(q)φ(q)∑h′(modd)(h′,q)=1∫Rw~3,B,X(ξ)S1(h′d+ξjX)S2(−h′d−ξjX) dξ.(62)\frac{k_BB}{j}\sum_{q\le Q}\frac{\mu_{\mathrm{Mob}}(q)}{\varphi(q)} \sum_{\substack{h'\pmod d\\(h',q)=1}} \int_{\mathbb R}\widetilde w_{3,B,X}(\xi) S_1\left(\frac{h'}{d}+\frac{\xi}{jX}\right) S_2\left(-\frac{h'}{d}-\frac{\xi}{jX}\right)\,d\xi. \tag*{(62)}

In this formula the integration has been extended from [−B13,B13][-B^{13},B^{13}] to R\mathbb R. To justify it, use (8.4), the bound of jφ(q)j\varphi(q) for the number of lifts at denominator qq, and a Schwartz bound of any sufficiently large fixed order. The number of boxes and the total arc count are polynomial in BB, so the tails give o(1/T)o(1/T) after all sums. Notice also that d≤TQ≤B0.72<Bd\le TQ\le B^{0.72}<B, as required for the channel estimates.

Fine histograms and residue averaging

The discarded arcs already have negligible total contribution. The remaining major arcs carry a logarithmic coordinate and a unit residue at each endpoint. We next use the channel bounds to retain the former while averaging the latter, uniformly for the endpoint tests.

We replace the split-product measures in (62) by the fine log and residue histograms of (52). For instance, the replacement for its first sum is

1φ(d)∑r∈Rde(h′r/d)∫IUdh(x,r)w1(eBx/NX)e(ξeBx/(jX)) dx.(63)\frac{1}{\varphi(d)}\sum_{r\in\mathcal{R}_d}e(h'r/d)\int_I U_{dh}(x,r)w_1(e^{Bx}/N_X)e(\xi e^{Bx}/(jX))\,dx. \tag*{(63)}

The residue phases have not been approximated. We give the norm estimate which controls this step, its denominator sum, and the subsequent residue projection.

Let IX⊂II_X\subset I be an enlarged log window about log⁡NX/B\log N_X/B, of length O(1/B)O(1/B), which contains all fine cells meeting the support of the factors in a box. These windows have bounded overlap as XX ranges over its dyadic values, because successive centers are spaced by (log⁡2)/B(\log2)/B. They lie in the interior of II for large BB. If FF is supported in IXI_X and bounded, finite Fourier Parseval on Z/dZ\mathbb{Z}/d\mathbb{Z}, with the residue vector extended by zero off the units, gives

∑h′ mod d∣1φ(d)∑r∈Rde(h′r/d)∫IXV(x,r)F(x) dx∣2≤dφ(d)∣IX∣∥F∥∞2∥V∥L2(IX×Rd)2.(64)\sum_{h'\bmod d}\left\lvert\frac{1}{\varphi(d)}\sum_{r\in\mathcal{R}_d}e(h'r/d)\int_{I_X}V(x,r)F(x)\,dx\right\rvert^2 \le\frac{d}{\varphi(d)}\lvert I_X\rvert\lVert F\rVert_\infty^2\lVert V\rVert_{L^2(I_X\times\mathcal{R}_d)}^2. \tag*{(64)}

Indeed the exact finite Fourier identity, before applying Cauchy–Schwarz to the integrals, is

∑h′ mod d∣1φ(d)∑r∈Rde(h′r/d)vr∣2=dφ(d)1φ(d)∑r∈Rd∣vr∣2.\sum_{h'\bmod d}\left\lvert\frac{1}{\varphi(d)}\sum_{r\in\mathcal{R}_d}e(h'r/d)v_r\right\rvert^2=\frac{d}{\varphi(d)}\frac{1}{\varphi(d)}\sum_{r\in\mathcal{R}_d}\lvert v_r\rvert^2.

Here and throughout, the residue component of the L2L^2 norm uses uniform probability on Rd\mathcal{R}_d.

For clarity, the histogram replacement can be estimated even when the signed masses γg\gamma_g have no regularity whatsoever. For a fine cell JJ and a unit residue rr, the definition (52) says exactly that

1φ(d)∫JUdh(x,r) dx=E[h(S)1{log⁡b/B∈J, b≡r (d)}].\frac{1}{\varphi(d)}\int_J U_{dh}(x,r)\,dx=\mathbb{E}\left[h(S)1_{\{\log b/B\in J,\ b\equiv r\ (d)\}}\right].

For F(x)=w1(eBx/NX)e(ξeBx/(jX))F(x)=w_1(e^{Bx}/N_X)e(\xi e^{Bx}/(jX)), differentiation on the relevant enlarged window gives

∥F′∥∞≤Cw,η(1+∣ξ∣)B.\lVert F'\rVert_\infty\le C_{w,\eta}(1+\lvert\xi\rvert)B.

We used eBx≍TXe^{Bx}\asymp TX and j≥ηTj\ge\eta T. Thus the oscillation of FF on a fine cell is at most

aB(ξ):=Cw,η(1+∣ξ∣)BδB.a_B(\xi):=C_{w,\eta}(1+\lvert\xi\rvert)B\delta_B.

After taking the factor 1/φ(d)1/\varphi(d) outside as in (63), the residue vector of the replacement error has absolute value at residue rr at most

aB(ξ)∫IXUd1(x,r) dx.a_B(\xi)\int_{I_X}U_{d1}(x,r)\,dx.

This follows by subtracting the cell average of FF from its value at the actual split product and using ∣h∣≤1|h| \le1. Positivity of the measure for Ud1U_{d1} is the only pointwise information used. Formula (64) therefore bounds the squared sum of this error over h′h' by

Cηdφ(d)∣IX∣aB(ζ)2∥Ud1∥L2(IX×Rd)2.\frac{C_\eta d}{\varphi(d)} |I_X| a_B(\zeta)^2 \lVert U_{d1}\rVert_{L^2(I_X \times\mathbb{R}_d)}^2.

The actual, unreplaced sum has the same type of estimate without aBa_B, by bounding its residue vector with ∥F∥∞∫IXUd1\lVert F\rVert_\infty\int_{I_X} U_{d1}. The replaced sum has (64) directly with V=UdhV = U_{d h}.

We apply these estimates to the difference of the two products in (62), replacing one factor at a time. Cauchy–Schwarz over h′h' is valid even though the sum is restricted by (h′,q)=1(h',q)=1, since extending either squared sum to all residues increases it. Afterwards Cauchy–Schwarz over the boxes and bounded overlap of IXI_X bound the sum of products of local norms by the product of global norms. These global norms are bounded by [](#eq:7.2. Finally the Schwartz moments absorb the factors 1+∣ζ∣1+|\zeta|, as well as the polynomial smooth losses in the separated expansion. Since ∣IX∣≪1/B|I_X| \ll1/B, this proves that at a fixed qq the total histogram error, including all boxes, is at most

Cηj1φ(q)jqφ(jq)BδδB.(65)\frac{C_\eta}{j}\frac{1}{\varphi(q)}\frac{jq}{\varphi(jq)}B^\delta\delta_B. \tag*{(65)}

In particular, the number O(B)O(B) of boxes has not introduced an extra factor of BB: the local window length cancels the BB in (62), and the remaining local norms are summed by bounded overlap.

We record explicitly a bound for the denominator sum. The elementary inequality

jqφ(jq)≤jφ(j)qφ(q)\frac{jq}{\varphi(jq)} \le\frac{j}{\varphi(j)}\frac{q}{\varphi(q)}

and partial summation imply

∑q≤Q1φ(q)jqφ(jq)≪jφ(j)log⁡(2Q).(66)\sum_{q\le Q}\frac{1}{\varphi(q)}\frac{jq}{\varphi(jq)} \ll\frac{j}{\varphi(j)}\log(2Q). \tag*{(66)}

For completeness, to verify the required mean-value input write (n/φ(n))2=∑d∣nb(d)(n/\varphi(n))^2=\sum_{d\mid n}b(d), where bb is supported on squarefree integers and

b(p)=(1−1/p)−2−1=O(1/p).b(p)=(1-1/p)^{-2}-1=O(1/p).

All these coefficients are nonnegative, and

∑db(d)d=∏p(1+b(p)/p)<∞.\sum_d\frac{b(d)}{d}=\prod_p(1+b(p)/p)<\infty.

Hence ∑q≤Q(q/φ(q))2≪Q\sum_{q\le Q}(q/\varphi(q))^2\ll Q, which after partial summation gives (66). Since j/φ(j)=Bo(1)j/\varphi(j)=B^{o(1)} uniformly for j≤Tj\le T and BδδB≍B−1/10B^\delta\delta_B\asymp B^{-1/10}, summing (65) gives o(1/T)o(1/T), uniformly on the allowed lag range.

We now replace Udh,UdgU_{d h},U_{d g} in the histogram expression by their residue averages uh,ugu_h,u_g. Expand the difference of products one factor at a time and repeat (64). The only new input is

∥Udh−uh∥2+∥Udg−ug∥2≪B−c7\lVert U_{d h}-u_h\rVert_2+\lVert U_{d g}-u_g\rVert_2\ll B^{-c_7}

from [](#eq:7.2, uniformly in d≤Bd\le B and the bounded tests. It follows that the error is bounded by

Cηjjφ(j)log⁡(2Q)B−c7=o(1/T).\frac{C_\eta}{j}\frac{j}{\varphi(j)}\log(2Q)B^{-c_7}=o(1/T).

This establishes the residue projection for arbitrary bounded single-site tests, rather than just for the constant test.

The exact residue factor

After the channel projection, the endpoint tests no longer depend on a residue coordinate. The remaining finite residue sum can therefore be computed exactly; it is the arithmetic factor governing the allowed lags.

After residue averaging, the unit phase average in (3.9) is cd(h′)/φ(d)c_d(h')/\varphi(d), where

cd(h′)=∑r∈Rde(h′r/d)c_d(h')=\sum_{r\in\mathcal{R}_d}e(h'r/d)

is the Ramanujan sum. The second endpoint contributes its complex conjugate phase average. Since the coefficient of a nonsquarefree qq in (3.8) is zero, only squarefree qq matter. For such qq the exact identity is

∑h′modjq(h′,q)=1∣cjq(h′)∣2φ(jq)2={jφ(j)φ(q),(j,q)=1,0,(j,q)>1.(67)\sum_{\substack{h'\bmod jq\\(h',q)=1}}\frac{\lvert c_{jq}(h')\rvert^2}{\varphi(jq)^2} = \begin{cases} \dfrac{j}{\varphi(j)\varphi(q)}, & (j,q)=1,\\ 0, & (j,q)>1. \end{cases} \tag*{(67)}

Here is a direct proof that also covers nonsquarefree jj. If p∣(j,q)p\mid(j,q), the modulus jqjq is divisible by p2p^2, while (h′,q)=1(h',q)=1 forces p∤h′p\nmid h'. For every a≥2a\ge2,

cpa(h′)=∑r mod pae(h′r/pa)−∑r mod pa−1e(h′r/pa−1)=0(p∤h′).c_{p^a}(h')=\sum_{r\bmod p^a}e(h'r/p^a)-\sum_{r\bmod p^{a-1}}e(h'r/p^{a-1})=0\quad(p\nmid h').

Chinese remaindering therefore makes the whole Ramanujan sum zero. If (j,q)=1(j,q)=1, the Ramanujan sums split over the two coprime moduli. At the squarefree modulus qq, a unit frequency has cq(h′)=μMob(q)c_q(h')=\mu_{\mathrm{Mob}}(q), of absolute value one. At modulus jj, orthogonality gives

∑h′ mod j∣cj(h′)∣2=∑r,r′∈Rj∑h′ mod je(h′(r−r′)/j)=jφ(j).\sum_{h'\bmod j}\lvert c_j(h')\rvert^2 = \sum_{r,r'\in\mathcal{R}_j}\sum_{h'\bmod j}e(h'(r-r')/j) =j\varphi(j).

The numerator on the left of (67) is thus jφ(j)φ(q)j\varphi(j)\varphi(q), and division by φ(jq)2=φ(j)2φ(q)2\varphi(jq)^2=\varphi(j)^2\varphi(q)^2 proves the identity.

Combining (67) with the factor μMob(q)/φ(q)\mu_{\mathrm{Mob}}(q)/\varphi(q) in (3.8) yields the singular series

S(j)=jφ(j)∑q(q,j)=1μMob(q)φ(q)2=jφ(j)∏p∤j(1−1(p−1)2).(68)\mathfrak{S}(j)=\frac{j}{\varphi(j)}\sum_{\substack{q\\(q,j)=1}}\frac{\mu_{\mathrm{Mob}}(q)}{\varphi(q)^2} =\frac{j}{\varphi(j)}\prod_{p\nmid j}\left(1-\frac{1}{(p-1)^2}\right). \tag*{(68)}

The series is absolutely convergent: for every fixed ϵ>0\epsilon>0, 1/φ(q)≪ϵq−1+ϵ1/\varphi(q)\ll_\epsilon q^{-1+\epsilon}, so

∑q>Q1φ(q)2≪ϵQ−1+2ϵ(0<ϵ<1/2).\sum_{q>Q}\frac{1}{\varphi(q)^2}\ll_\epsilon Q^{-1+2\epsilon}\quad(0<\epsilon<1/2).

This tail bound is independent of jj. Also

0≤S(j)≤j/φ(j).0\le\mathfrak{S}(j)\le j/\varphi(j).

The series vanishes when jj is odd, consistently with the parity obstruction in a−b=jca-b=jc for rough coefficients.

After the residue projection, the archimedean integrals in (3.8) no longer depend on qq. The total over boxes of their absolute values, before the factor B/jB/j, is O(1/B)O(1/B): apply Cauchy–Schwarz on their log windows, then use bounded overlap and ∥ug∥2,∥uh∥2≪1\lVert u_g\rVert_2,\lVert u_h\rVert_2\ll1. The Schwartz integrations and smooth expansions have the same summability as before. It follows that replacing the sum q≤Qq \le Q by eq:8.13 costs at most

Cjjφ(j)∑q>Q1φ(q)2.(69)\frac{C}{j}\frac{j}{\varphi(j)}\sum_{q>Q}\frac{1}{\varphi(q)^2}. \tag*{(69)}

In particular this is a uniform vanishing multiple of (j/φ(j))/j(j/\varphi(j))/j.

Archimedean inversion and coarse features

For one separated box, the phase remaining in the product of archimedean integrals is

e(ζeBx−eByjX).e\left(\zeta\frac{e^{Bx}-e^{By}}{jX}\right).

Fourier inversion in ζ\zeta therefore evaluates w3(z′)mB(log⁡(Xz′)/B)w_3(z')m_B(\log(Xz')/B) at z′=(eBx−eBy)/(jX)z'=(e^{Bx}-e^{By})/(jX). The inversion introduces no further Jacobian: its integration variable is ζ\zeta. Recombining the separated factors and the dyadic partition gives

κBBjS(j)∫x>yx,y∈Iuh(x)ug(y)mB(z)ρ0(z)⋅Ψs(λeB(x−y)eB(x−y)−1,λeB(x−y)−1) dx dy,λ=jT,z=1Blog⁡(eBx−eByj).(70)\frac{\kappa_B B}{j}\mathfrak{S}(j)\int_{\substack{x>y\\x,y\in I}}u_h(x)u_g(y)m_B(z)\rho_0(z)\cdot\Psi_s\left(\frac{\lambda e^{B(x-y)}}{e^{B(x-y)}-1},\frac{\lambda}{e^{B(x-y)}-1}\right)\,dx\,dy, \qquad\lambda=\frac{j}{T},\qquad z=\frac{1}{B}\log\left(\frac{e^{Bx}-e^{By}}{j}\right). \tag*{(70)}

Weights are interpreted as zero off their positive domains and supports. The inversions and rearrangements are justified for ug,uh∈L2(I)u_g,u_h\in L^2(I) by the bounded log windows, the Schwartz transforms, and the absolutely summable smooth expansion, with the absolute-product bound just used for (69).

We spell out the support properties needed to approximate (70). There is a compact interval [a0,b0]⊂(1,2)[a_0,b_0]\subset(1,2) containing the support of ψ\psi. Write v=B(x−y)v=B(x-y). If the Ψs\Psi_s factor is nonzero, its two arguments u1,u2u_1,u_2 lie in [a0,b0][a_0,b_0], with

u1−u2=λ,ev=u1/u2.u_1-u_2=\lambda,\qquad e^v=u_1/u_2.

For η≤λ≤1\eta\le\lambda\le1, this implies

0<cη≤v≤C,cη=log⁡(1+η/b0),C=log⁡(b0/a0).0<c_\eta\le v\le C,\qquad c_\eta=\log(1+\eta/b_0),\qquad C=\log(b_0/a_0).

If these inequalities have no solution for a particular λ\lambda, the weight is simply zero. On this support,

z=y+log⁡(ev−1)−log⁡jB=y+Oη(log⁡T/B).z=y+\frac{\log(e^v-1)-\log j}{B}=y+O_\eta(\log T/B).

The functions mBρ0m_B\rho_0, extended by zero beyond a fixed compact interval inside (0,4)(0,4), converge uniformly to mρ0m\rho_0, with uniformly bounded derivatives. It follows that

mB(z)ρ0(z)=m(y)ρ0(y)+oη(1)m_B(z)\rho_0(z)=m(y)\rho_0(y)+o_\eta(1)

uniformly wherever the integrand can be nonzero.

To quantify its effect for arbitrary tests, consider the integral operator with nonnegative kernel

B1{0<x−y<C/B},x,y∈I.B\mathbf{1}_{\{0<x-y<C/B\}}, \qquad x,y \in I.

Its row and column integrals are at most CC, so Schur’s inequality gives a uniform L2L^2 operator bound. Multiplication by the bounded smooth weights preserves that bound. Since ∥ug∥2,∥uh∥2≪1\lVert u_g\rVert_2,\lVert u_h\rVert_2 \ll1, the replacement of mB(z)ρ0(z)m_B(z)\rho_0(z) in (70) costs oη(1)S(j)/jo_\eta(1)\mathcal{S}(j)/j.

Given ϵ1>0\epsilon_1>0, choose the fixed coarse partition in (54), and write ugc=Pm1ugu_g^c=P_{m_1}u_g, uhc=Pm1uhu_h^c=P_{m_1}u_h. The same operator bound and the contraction property of Pm1P_{m_1} show that replacing the two endpoint functions by ugc,uhcu_g^c,u_h^c costs at most

Cη(ϵ1+o(1))S(j)j.C_\eta(\epsilon_1+o(1))\frac{\mathcal{S}(j)}{j}.

This constant does not grow with the chosen coarse partition; the errors tending to zero may of course depend on that fixed partition. Its coarse functions satisfy

∥ugc∥∞,∥uhc∥∞≤max⁡l∣Il∣−1/2max⁡(∥ug∥2,∥uh∥2)≪m11.\lVert u_g^c\rVert_\infty,\lVert u_h^c\rVert_\infty \leq\max_l |I_l|^{-1/2}\max(\lVert u_g\rVert_2,\lVert u_h\rVert_2) \ll_{m_1} 1.

By (8.16), xx and yy belong to the same coarse cell except when yy is within C/BC/B of one of the finitely many cell boundaries. The area of these exceptional pairs is Om1(B−2)O_{m_1}(B^{-2}): the total possible length for yy is Om1(B−1)O_{m_1}(B^{-1}), and for each such yy the possible length for xx is O(B−1)O(B^{-1}). Thus replacing uhc(x)u_h^c(x) by uhc(y)u_h^c(y) costs Oη,m1(B−1)S(j)/jO_{\eta,m_1}(B^{-1})\mathcal{S}(j)/j. The support of m(y)ρ0(y)m(y)\rho_0(y) is a fixed positive distance from the boundary of II, so there is no further boundary contribution from the condition x∈Ix\in I for large BB.

Recall the coarse features from Corollary 7.4, and define their deterministic coefficients by

Vl(S)=1∣Il∣Psplit(log⁡b/B∈Il∣S),hl=∫Ilm(y)ρ0(y) dy.V_l(S)=\frac{1}{|I_l|}\mathbb{P}_{\mathrm{split}}(\log b/B\in I_l\mid S), \qquad h_l=\int_{I_l}m(y)\rho_0(y)\,dy.

These are nonnegative, Vl(S)≤∣Il∣−1V_l(S)\leq|I_l|^{-1}, and hl≥0h_l\geq0. Since the fine partition refines the coarse partition, averaging (52) over a coarse cell gives the exact identity

(Pm1ug)∣Il=E[g(S)Vl(S)].(P_{m_1}u_g)|_{I_l}=\mathbb{E}[g(S)V_l(S)].

After the preceding replacements in (70), substitute x=y+v/Bx=y+v/B. The factor dx=dv/Bdx=dv/B cancels its prefactor BB, and 1/j=(1/T)(1/λ)1/j=(1/T)(1/\lambda). The resulting integral tests are therefore precisely those of the following site matrix:

Mik(Si,Sk;s)=kBTS(∣j∣)w(s,∣j∣/T)∑lhlVl(Si)Vl(Sk),j=k−i,(71)\mathcal{M}_{ik}(S_i,S_k;s) =\frac{k_B}{T}\mathcal{S}(|j|)w(s,|j|/T)\sum_l h_lV_l(S_i)V_l(S_k), \qquad j=k-i, \tag*{(71)}
w(s,λ)=1λ∫v>0Ψs(λevev−1,λev−1) dv,η≤λ≤1.w(s,\lambda)=\frac{1}{\lambda}\int_{v>0}\Psi_s\left(\frac{\lambda e^v}{e^v-1},\frac{\lambda}{e^v-1}\right)\,dv, \qquad\eta\leq\lambda\leq1.

The matrix is zero off the allowed pairs. Formula (71) defines it for both signs of jj; it is symmetric because its displayed dependence on the endpoints is symmetric. For positive jj, independence of the endpoint types turns its integral test against g,hg,h into the product of the two expectations just displayed. The case of negative jj follows by exchanging the endpoints.

The function ww is nonnegative and smooth, with all fixed-order derivatives bounded on the relevant compact ss-range and η≤λ≤1\eta\leq\lambda\leq1. There is in fact an expression giving bounds independent of η\eta. In (71) put u=λ/(ev−1)u=\lambda/(e^v-1). Then dv=−λ du/(u(u+λ))dv=-\lambda\,du/(u(u+\lambda)), and hence

w(s,λ)=∫0∞Ψs(u+λ,u)u(u+λ) du.w(s,\lambda)=\int_0^\infty\frac{\Psi_s(u+\lambda,u)}{u(u+\lambda)}\,du.

The factors ψ(u)ψ(u+λ)\psi(u)\psi(u+\lambda) restrict the integration to a fixed compact subset of u>0u>0. Consequently this formula extends ww smoothly to 0≤λ≤10\leq\lambda\leq1, with uniformly bounded derivatives there, and gives zero when λ>b0−a0\lambda>b_0-a_0. Differentiation under the integral is valid because its denominators are bounded away from zero on that fixed support.

Uniform integral cut comparison and row bounds

The preceding calculation has produced the finite-feature matrix M\mathcal{M}. We now collect its approximation error in the same block normalization as the original energy and record the row bounds needed when the endpoint types are sampled.

Proposition 8.2 (Integral comparison). For the matrix M\mathcal{M} in (71), put Eik=Lik−Mik\mathcal{E}_{ik}=\mathcal{L}_{ik}-\mathcal{M}_{ik}, with all three matrices zero on nonpairs. Given ϵ1>0\epsilon_1>0, choose the coarse partition as in (54). Then

sup⁡gi,hk:ΩB→[−1,1]gi,hk real measurable∣1M∑i,kESi,Sk[Eik(Si,Sk;s)gi(Si)hk(Sk)]∣≤Ce−c3C∗+Cηϵ1+o(1).(72)\sup_{\substack{g_{i,h_k}:\Omega_B\to[-1,1]\\ g_{i,h_k}\ \mathrm{real\ measurable}}} \left|\frac{1}{M}\sum_{i,k}\mathbb{E}_{S_i,S_k}\left[\mathcal{E}_{ik}(S_i,S_k;s)g_i(S_i)h_k(S_k)\right]\right| \leq C e^{-c_3 C_*}+C_\eta\epsilon_1+o(1). \tag*{(72)}

The o(1)o(1) is uniform in ss and in all the indicated tests. The coefficient of e−c3C∗e^{-c_3C_*} is independent of η\eta, C6C_6, and C∗C_*. In addition, for every realization of the site types,

sup⁡i∑k∣Mik∣≪η,m11,sup⁡i∑k∣Mik∣2≪η,m1T−1.(73)\sup_i\sum_k|\mathcal{M}_{ik}|\ll_{\eta,m_1}1,\qquad \sup_i\sum_k|\mathcal{M}_{ik}|^2\ll_{\eta,m_1}T^{-1}. \tag*{(73)}

Proof. Every estimate above was uniform in a pair of tests with absolute value at most one. Apart from (59), the pre-projection arc errors are uniform o(1/T)o(1/T). The histogram and residue projection errors have the explicit common upper bound

Cηjjφ(j)log⁡(2Q)(BδB+B−c7).\frac{C_\eta}{j}\frac{j}{\varphi(j)}\log(2Q)(B\delta_B+B^{-c_7}).

The series tail is bounded by (69), and the archimedean and coarse replacements have total error

(Cηϵ1+oη,m1(1))S(j)j.(C_\eta\epsilon_1+o_{\eta,m_1}(1))\frac{\mathfrak{S}(j)}{j}.

All these bounds hold separately for every allowed edge. They thus remain valid if the endpoint tests differ from edge to edge, and in particular for the collection of tests in (72).

For a fixed positive lag there are at most MM ordered pairs with that lag; including the negative lag at most doubles this number. Division by MM therefore reduces their total error to at most twice the sum of the edge error over ηT≤j≤T\eta T\leq j\leq T. The first-moment bound for Σ\Sigma gives

1T∑1≤j≤TΣ(j)≪1.\frac{1}{T}\sum_{1\leq j\leq T}\Sigma(j)\ll1.

Consequently the sum of (59) is O(rB+e−c3C∗)O(r_B+e^{-c_3C_*}), with a constant independent of η\eta, C6C_6, C∗C_*. A uniform o(1/T)o(1/T) error sums to o(1)o(1). For the other errors use the bounded mean of j/φ(j)j/\varphi(j) and

∑ηT≤j≤T1jjφ(j)≤1ηT∑j≤Tjφ(j)≪η1.\sum_{\eta T\leq j\leq T}\frac{1}{j}\frac{j}{\varphi(j)} \leq\frac{1}{\eta T}\sum_{j\leq T}\frac{j}{\varphi(j)} \ll_\eta1.

Since log⁡(2Q)(BδB+B−c7)→0\log(2Q)(B\delta_B+B^{-c_7})\to0, the denominator sum contributes o(1)o(1). Absolute convergence makes the series tail tend to zero as well. Finally S(j)≤j/φ(j)\mathcal{S}(j)\le j/\varphi(j) controls the archimedean and coarse errors, giving the stated Cηϵ1+o(1)C_\eta\epsilon_1+o(1). There is no factor of C6C_6, because the count of pairs for each lag has already been divided by MM. This proves (72).

For the deterministic row estimates, boundedness of kBk_B, ww, and the finitely many features gives the pointwise bound

∣Mik∣≤Cη,m1T∣k−i∣φ(∣k−i∣)|\mathcal{M}_{ik}|\le\frac{C_{\eta,m_1}}{T}\frac{|k-i|}{\varphi(|k-i|)}

on allowed pairs. Each row has at most two entries for each positive lag. The first and second powers of j/φ(j)j/\varphi(j) have bounded means. For the second power this was proved before (66), and the first follows by Cauchy–Schwarz. Summing the displayed pointwise bound and its square over 1≤j≤T1\le j\le T proves (73), uniformly in every collection of site types and in the block length.

Sampling the integral cut comparison

The comparison in (72) allows arbitrary bounded functions at each single site. To use it against the labels in (37), we must also allow the testing signs to depend on the entire sampled matrix. We do this by approximating an optimizing row-sign vector using a small set of columns, and then applying concentration simultaneously to the resulting finite family of optimizations.

The small-column approximation is related to the sampling argument of Alon, Fernández de la Vega, Kannan, and Karpinski [1], Lemma 3 and to the proof of Borgs et al. [2], Theorem 4.6. Their arguments approximate optimizing cuts by a small sample and then control a finite family of tests. Here the kernels depend on BB and on position, and are controlled by degree bounds and integrated row-second-moment bounds. We therefore prove the required sampling transfer directly rather than invoking those results as a black box.

Throughout this section the regularity parameters, C∗C_*, η\eta, C6C_6, ϵ1\epsilon_1, and the coarse partition are fixed. In particular, T=⌊B0.32⌋T=\lfloor B^{0.32}\rfloor and M=⌈C6T⌉≍B0.32M=\lceil C_6T\rceil\asymp B^{0.32}. All limits and o(1)o(1) terms in the probabilistic argument refer to B→∞B\to\infty with these parameters fixed. We work at one value of ss in the fixed compact range of the smooth weights. Every bound below is uniform in that value of ss; no simultaneous event over all ss will be needed.

Let S=(S1,…,SM)S=(S_1,\ldots,S_M) have the independent site law from Section 5. Write

Eik(Si,Sk;s)=Lik(Si,Sk;s)−Mik(Si,Sk;s),E0=Ce−c3C∗+Cηϵ1.\mathcal{E}_{ik}(S_i,S_k;s)=\mathcal{L}_{ik}(S_i,S_k;s)-\mathcal{M}_{ik}(S_i,S_k;s),\qquad E_0=Ce^{-c_3C_*}+C_\eta\epsilon_1.

Here the fixed constants in E0E_0 are large enough for (72). As established there, the coefficient CC of e−c3C∗e^{-c_3C_*} can be chosen independently of η\eta and C6C_6. The matrices are real and symmetric, and are zero on the diagonal and on all nonallowed pairs. We retain the normalization

∥H∥□=1Mmax⁡ϵ,δ∈{−1,1}M∣∑i,kHikϵiδk∣.\lVert H\rVert_{\square}=\frac{1}{M}\max_{\epsilon,\delta\in\{-1,1\}^M}\left|\sum_{i,k}H_{ik}\epsilon_i\delta_k\right|.

Proposition 9.1 (Sampling the type-kernel comparison). For every fixed ϵ2>0\epsilon_2>0,

E∥(Eik(Si,Sk;s))∥□≤E0+ϵ2+o(1),(74)\mathbb{E}\left\lVert\bigl(\mathcal{E}_{ik}(S_i,S_k;s)\bigr)\right\rVert_{\square}\le E_0+\epsilon_2+o(1), \tag*{(74)}

uniformly for ss in the fixed compact range.

We first record the degree and integrability estimates required in the proof. Use the nonnegative symmetric envelope

Hik=Lik+∣Mik∣,∣Eik∣≤Hik.\mathcal{H}_{ik}=\mathcal{L}_{ik}+\lvert\mathcal{M}_{ik}\rvert,\qquad\lvert\mathcal{E}_{ik}\rvert\leq\mathcal{H}_{ik}.

Dependence on the endpoint types and on ss is suppressed when no confusion is possible. By the every-type bound (5.9) and the deterministic absolute row bounds for (8.17), there is a constant C0C_0 such that

E[∑kHik∣Si]≤C0for every row type Si.\mathbb{E}\left[\left.\sum_k\mathcal{H}_{ik}\right|S_i\right]\leq C_0\qquad\text{for every row type }S_i.

By (6.1) and the squared-row bounds following (8.19), there is a constant C1C_1 such that

∑kEHik2≤qB,∑kEEik2≤qB,qB=C1B−0.21,1≤i≤M.\sum_k\mathbb{E}\mathcal{H}_{ik}^{2}\leq q_B,\qquad\sum_k\mathbb{E}\mathcal{E}_{ik}^{2}\leq q_B,\qquad q_B=C_1B^{-0.21},\qquad1\leq i\leq M.

Indeed (Lik+∣Mik∣)2≤2Lik2+2Mik2(\mathcal{L}_{ik}+\lvert\mathcal{M}_{ik}\rvert)^2\leq2\mathcal{L}_{ik}^{2}+2\mathcal{M}_{ik}^{2}, and 1/T=O(B−0.32)1/T=O(B^{-0.32}). These squared-row bounds are integrated over the row type; no uniform conditional second-moment assertion is being made.

Lemma 9.2 (A common degree event and uniform integrability). There is a fixed K>0K>0 such that, on an event GB\mathcal{G}_B of probability 1−O(B−0.01)−O(B−0.03)1-O(B^{-0.01})-O(B^{-0.03}),

1M∑i,kEik2≤B−0.20,max⁡i∑k∣Eik∣≤KB0.07.(75)\frac{1}{M}\sum_{i,k}\mathcal{E}_{ik}^{2}\leq B^{-0.20},\qquad\max_i\sum_k\lvert\mathcal{E}_{ik}\rvert\leq KB^{0.07}. \tag*{(75)}

The event may also be required to satisfy max⁡i∑kHik≤KB0.07\max_i\sum_k\mathcal{H}_{ik}\leq KB^{0.07}. Moreover, the normalized total absolute mass

AB=1M∑i,k∣Eik∣A_B=\frac{1}{M}\sum_{i,k}\lvert\mathcal{E}_{ik}\rvert

has bounded second moment, uniformly in BB and ss.

Proof. Put Di=∑kHikD_i=\sum_k\mathcal{H}_{ik} and dB=KB0.07d_B=KB^{0.07}, where KK is any fixed sufficiently large constant. Conditional on SiS_i, the summands of DiD_i are independent, since each off-diagonal summand depends on one other independent site. Consequently

EVar⁡(Di∣Si)≤∑kEHik2≤qB,EDi2≤C02+qB.\mathbb{E}\operatorname{Var}(D_i\mid S_i)\leq\sum_k\mathbb{E}\mathcal{H}_{ik}^{2}\leq q_B,\qquad\mathbb{E}D_i^{2}\leq C_0^{2}+q_B.

For sufficiently large BB, dB≥2C0d_B\geq2C_0. Conditional Chebyshev followed by averaging the row type gives

P(Di>dB)≤qB(dB−C0)2≤4qBdB2.\mathbb{P}(D_i>d_B)\leq\frac{q_B}{(d_B-C_0)^2}\leq\frac{4q_B}{d_B^2}.

Thus the event

Ωdeg={max⁡iDi≤dB}\Omega_{\mathrm{deg}}=\{\max_i D_i\leq d_B\}

has complement of probability O(MqB/dB2)=O(B−0.03)O(Mq_B/d_B^2)=O(B^{-0.03}). Markov’s inequality gives

P(1M∑i,kEik2>B−0.20)≤qBB0.20=O(B−0.01).\mathbb{P}\left(\frac{1}{M}\sum_{i,k}\mathcal{E}_{ik}^{2}>B^{-0.20}\right)\leq q_BB^{0.20}=O(B^{-0.01}).

Take

GB=Ωdeg∩{1M∑i,kEik2≤B−0.20}.\mathcal{G}_B=\Omega_{\mathrm{deg}}\cap\left\{\frac{1}{M}\sum_{i,k}\mathcal{E}_{ik}^{2}\leq B^{-0.20}\right\}.

This implies (75). Finally, Jensen’s inequality across rows gives

EAB2≤E(1M∑iDi)2≤1M∑iEDi2≤C02+qB.\mathbb{E}A_B^2 \le\mathbb{E}\left(\frac{1}{M}\sum_i D_i\right)^2 \le\frac{1}{M}\sum_i \mathbb{E}D_i^2 \le C_0^2+q_B.

In particular, for any events BB\mathcal{B}_B with probabilities tending to zero uniformly in ss, E[AB1BB]=o(1)\mathbb{E}[A_B\mathbf{1}_{\mathcal{B}_B}]=o(1) by Cauchy–Schwarz.

The same degree estimate applies to the graph induced by any fixed subset of indices: deleting terms decreases both its conditional expected degrees and its integrated squared-row sums. In particular, the probability of a degree exceeding dBd_B is still O(B−0.03)O(B^{-0.03}) after a union over its at most MM rows. This remains true conditional on the types at indices outside the subset, since the remaining types retain their original independent laws.

Lemma 9.3 (Approximation by a small set of columns). Put m2=⌊MB−0.18⌋m_2=\lfloor MB^{-0.18}\rfloor. On GB\mathcal{G}_B there are a subset J′⊂{1,…,M}J'\subset\{1,\ldots,M\} of size m2m_2 and signs (δj)j∈J′(\delta_j)_{j\in J'} such that, with

g^i=sgn⁡(∑j∈J′Eijδj),R′={1,…,M}∖J′,\widehat{g}_i=\operatorname{sgn}\left(\sum_{j\in J'}\mathcal{E}_{ij}\delta_j\right),\qquad R'=\{1,\ldots,M\}\setminus J',

and assigning the sign +1+1 at zero for all sign choices in this section,

∥E∥□≤1M∑k∈R′∣∑i∈R′Eikg^i∣+O(B−0.01)+O(B−0.11).\|\mathcal{E}\|_{\square}\le\frac{1}{M}\sum_{k\in R'}\left|\sum_{i\in R'}\mathcal{E}_{ik}\widehat{g}_i\right|+O(B^{-0.01})+O(B^{-0.11}).

The constants are independent of the realized matrix on GB\mathcal{G}_B.

Proof. Fix the realized matrix. Choose column signs δk\delta_k attaining the cut norm and choose the best-response row signs, so the unnormalized maximum is the positive quantity ∑i∣ri∣\sum_i|r_i|, where ri=∑kEikδkr_i=\sum_k\mathcal{E}_{ik}\delta_k. For this deterministic matrix only, sample a uniform subset J′J' of m2m_2 columns and form the unbiased estimate

r^i=Mm2∑j∈J′Eijδj.\widehat{r}_i=\frac{M}{m_2}\sum_{j\in J'}\mathcal{E}_{ij}\delta_j.

The variance formula for sampling without replacement yields

EJ′∣r^i−ri∣2≤Mm2∑kEik2.\mathbb{E}_{J'}|\widehat{r}_i-r_i|^2\le\frac{M}{m_2}\sum_k\mathcal{E}_{ik}^2.

For real numbers rr and r^\widehat{r}, ∣r∣−rsgn⁡(r^)≤2∣r−r^∣|r|-r\operatorname{sgn}(\widehat{r})\le2|r-\widehat{r}|. Applying this inequality and then Cauchy–Schwarz across rows, the expected loss in the unnormalized bilinear sum is at most

2∑i(Mm2∑kEik2)1/2≤2M(Mm2B−0.20)1/2=O(MB−0.01).2\sum_i\left(\frac{M}{m_2}\sum_k\mathcal{E}_{ik}^2\right)^{1/2}\le2M\left(\frac{M}{m_2}B^{-0.20}\right)^{1/2}=O(MB^{-0.01}).

Here m2≍MB−0.18m_2\asymp MB^{-0.18} for large BB. Hence at least one subset has this loss bound, with sgn⁡(r^i)=g^i\operatorname{sgn}(\widehat{r}_i)=\widehat{g}_i. Its signs on J′J' are the restrictions of the chosen optimizing column signs.

Deleting every ordered pair with an endpoint in J′J' changes any such bilinear sum in absolute value by at most 2m2dB2m_2d_B, using symmetry and the degree event. After deletion, maximizing the column signs gives exactly ∑k∈R′∣∑i∈R′Eikg^i∣\sum_{k\in R'}|\sum_{i\in R'}\mathcal{E}_{ik}\widehat{g}_i|. Since m2dB/M=O(B−0.11)m_2d_B/M=O(B^{-0.11}), the claim follows.

Lemma 9.4 (Uniform mean after fixing a column set). Fix a deterministic subset J′J' of size m2m_2 and a deterministic assignment σ∈{−1,1}J′\sigma\in\{-1,1\}^{J'}. Condition on its site types SJ′=t\mathbf{S}_{J'}=\mathbf{t}. For i∈R′i \in R', define the fixed single-site function

gi(u)=sgn⁡(∑j∈J′Eij(u,tj;s)σj).g_i(u)=\operatorname{sgn}\left(\sum_{j\in J'}\mathcal{E}_{ij}(u,t_j;s)\sigma_j\right).

Then, writing

ZJ′,σ,t=1M∑k∈R′∣∑i∈R′Eik(Si,Sk;s)gi(Si)∣,Z_{J',\sigma,\mathbf{t}}=\frac{1}{M}\sum_{k\in R'}\left|\sum_{i\in R'}\mathcal{E}_{ik}(S_i,S_k;s)g_i(S_i)\right|,

we have E[ZJ′,σ,t∣SJ′=t]≤E0+o(1)\mathbb{E}[Z_{J',\sigma,\mathbf{t}}\mid\mathbf{S}_{J'}=\mathbf{t}]\le E_0+o(1), uniformly in J′J', σ\sigma, t\mathbf{t}, and ss.

Proof. Only the types in the deterministic set J′J' have been conditioned on. In particular, no event describing which signs an optimizer would choose has been imposed. The types in R′R' remain independent with their usual laws, and ∣gi∣≤1|g_i|\le1.

For k∈R′k\in R' and a possible value vv of SkS_k, put

ak(v)=∑i∈R′i≠kESi[Eik(Si,v;s)gi(Si)].a_k(v)=\sum_{\substack{i\in R'\\i\ne k}}\mathbb{E}_{S_i}\left[\mathcal{E}_{ik}(S_i,v;s)g_i(S_i)\right].

Choose hk(v)=sgn⁡(ak(v))h_k(v)=\operatorname{sgn}(a_k(v)), and set both families of tests equal to zero on removed indices. Equation (72), which is uniform over all bounded single-site tests, gives

1M∑k∈R′ESk∣ak(Sk)∣=1M∑i,k∈R′E[Eikgi(Si)hk(Sk)]≤E0+o(1).\frac{1}{M}\sum_{k\in R'}\mathbb{E}_{S_k}|a_k(S_k)|=\frac{1}{M}\sum_{i,k\in R'}\mathbb{E}\left[\mathcal{E}_{ik}g_i(S_i)h_k(S_k)\right]\le E_0+o(1).

This application is valid for every fixed value of t\mathbf{t}, however the resulting functions gig_i depend on that value.

Conditional on SkS_k, the centered terms

Eik(Si,Sk;s)gi(Si)−ESi[Eik(Si,Sk;s)gi(Si)],i∈R′∖{k},\mathcal{E}_{ik}(S_i,S_k;s)g_i(S_i)-\mathbb{E}_{S_i}\left[\mathcal{E}_{ik}(S_i,S_k;s)g_i(S_i)\right],\qquad i\in R'\setminus\{k\},

are independent. The diagonal contributes zero. Conditional variance followed by Cauchy–Schwarz therefore gives

E∣∑i∈R′Eikgi(Si)−ak(Sk)∣≤(∑i∈R′EEik2)1/2≤qB1/2.\mathbb{E}\left|\sum_{i\in R'}\mathcal{E}_{ik}g_i(S_i)-a_k(S_k)\right|\le\left(\sum_{i\in R'}\mathbb{E}\mathcal{E}_{ik}^{2}\right)^{1/2}\le q_B^{1/2}.

Symmetry supplies the squared-column bound from the squared-row bound. Averaging this estimate over kk costs at most qB1/2=O(B−0.105)q_B^{1/2}=O(B^{-0.105}). This proves the asserted mean bound, including its uniformity in the conditioned choice.

Proof of Proposition 9.1. The skeleton lemma reduces the adaptive cut norm to a finite family of optimizations, each with conditional mean at most E0+o(1)E_0+o(1). The remaining step is a concentration estimate strong enough to hold for the entire family. The common degree event will be excluded only once. We prove concentration for each of the deterministic subset/sign assignments in Lemma 9.4, and then take a union over those assignments. Fix such an assignment and condition on SJ′=t\mathbf{S}_{J'}=\mathbf{t}. Abbreviate its random quantity by ZZ. Let GR′\mathcal{G}_{R'} be the set of configurations of the remaining site types satisfying

max⁡i∈R′∑k∈R′Hik(Si,Sk;s)≤dB.\max_{i\in R'}\sum_{k\in R'}\mathcal{H}_{ik}(S_i,S_k;s)\le d_B.

The observation after Lemma 9.2 shows that

P(GR′c∣SJ′=t)=O(B−0.03),\mathbb{P}\left(\mathcal{G}_{R'}^{c}\mid S_{J'}=\mathbf{t}\right)=O\left(B^{-0.03}\right),

uniformly in the conditioning and assignment.

We next verify a Lipschitz bound between any two configurations u,v∈GR′\mathbf{u},\mathbf{v}\in\mathcal{G}_{R'}, rather than only between good configurations differing in one coordinate. Let A={i∈R′:ui≠vi}A=\{i\in R':u_i\ne v_i\}. The row functions gig_i have already been fixed by t\mathbf{t} and σ\boldsymbol{\sigma}. Thus a term Eik(ui,uk)gi(ui)\mathcal{E}_{ik}(u_i,u_k)g_i(u_i) can change only if ii or kk belongs to AA. The reverse triangle inequality for each column sum gives

∣Z(u)−Z(v)∣≤1M∑i,k∈R′i∈A or k∈A(Hik(ui,uk;s)+Hik(vi,vk;s))≤4dBM∣A∣.\lvert Z(\mathbf{u})-Z(\mathbf{v})\rvert\le\frac{1}{M}\sum_{\substack{i,k\in R'\\i\in A\ \text{or}\ k\in A}}\left(\mathcal{H}_{ik}(u_i,u_k;s)+\mathcal{H}_{ik}(v_i,v_k;s)\right)\le\frac{4d_B}{M}\lvert A\rvert.

The last inequality uses both endpoint degree bounds and symmetry. It also includes the change of gig_i when i∈Ai\in A. No connecting path of good configurations is required.

Choose a fixed L′>E0+2+ϵ2L'>E_0+2+\epsilon_2. On GR′\mathcal{G}_{R'} the function f=min⁡(Z,L′)f=\min(Z,L') is nonnegative and Lipschitz for the site Hamming distance dHd_H, with constant cB=4dB/Mc_B=4d_B/M. For sufficiently large BB the good set is nonempty. We use the McShane extension [19] in its infimum form [3], followed by clamping:

Z~(u)=min⁡{L′,inf⁡v∈GR′(f(v)+cBdH(u,v))}.\widetilde{Z}(\mathbf{u})=\min\left\{L',\inf_{\mathbf{v}\in\mathcal{G}_{R'}}\left(f(\mathbf{v})+c_Bd_H(\mathbf{u},\mathbf{v})\right)\right\}.

The site spaces are finite for every BB, so the formula is measurable. The infimum defines a cBc_B-Lipschitz function by the triangle inequality, and its restriction to the good set is ff: the Lipschitz inequality for ff gives the lower bound, and taking v=u\mathbf{v}=\mathbf{u} gives the upper bound. Clamping preserves the Lipschitz bound. Hence 0≤Z~≤L′0\le\widetilde{Z}\le L' everywhere and Z~=min⁡(Z,L′)\widetilde{Z}=\min(Z,L') on GR′\mathcal{G}_{R'}. By Lemma 9.4,

E[Z~∣SJ′=t]≤E[Z∣SJ′=t]+L′P(GR′c∣SJ′=t)≤E0+o(1),\mathbb{E}\left[\widetilde{Z}\mid S_{J'}=\mathbf{t}\right]\le\mathbb{E}\left[Z\mid S_{J'}=\mathbf{t}\right]+L'\mathbb{P}\left(\mathcal{G}_{R'}^{c}\mid S_{J'}=\mathbf{t}\right)\le E_0+o(1),

again uniformly in every fixed choice.

Changing any one of the independent remaining site types changes Z~\widetilde{Z} by at most cBc_B. McDiarmid’s bounded-differences inequality [18] therefore gives, for every a>0a>0,

P(Z~−E[Z~∣SJ′]>a∣SJ′=t)≤exp⁡(−2a2∣R′∣cB2)≤exp⁡(−a2M8dB2).\mathbb{P}\left(\widetilde{Z}-\mathbb{E}\left[\widetilde{Z}\mid S_{J'}\right]>a\mathrel{\big|}S_{J'}=\mathbf{t}\right)\le\exp\left(-\frac{2a^2}{\lvert R'\rvert c_B^2}\right)\le\exp\left(-\frac{a^2M}{8d_B^2}\right).

Taking a=ϵ2/4a=\epsilon_2/4 and using the uniform mean bound shows that, for large BB,

P(Z~>E0+ϵ2/2∣SJ′=t)≤exp⁡(−cϵ2M/B0.14).\mathbb{P}\left(\widetilde{Z}>E_0+\epsilon_2/2\mid S_{J'}=\mathbf{t}\right)\le\exp\left(-c_{\epsilon_2}M/B^{0.14}\right).

On GR′\mathcal{G}_{R'}, the event Z>E0+ϵ2/2Z>E_0+\epsilon_2/2 is equivalent to this event for Z~\widetilde{Z}, since the threshold is strictly below L′L'.

There are at most

NB=(Mm2)m2≤(2eMm2)m2N_B=\binom{M}{m_2}^{m_2}\le\left(\frac{2eM}{m_2}\right)^{m_2}

deterministic subset/sign assignments. For each assignment, integrate the last conditional tail bound over its subset types. We may then union-bound the extension tails, because

log⁡NB=O(m2log⁡M)=O(B0.14log⁡B)=o(M/B0.14),M/B0.14≍B0.18.\log N_B=O(m_2\log M)=O(B^{0.14}\log B)=o(M/B^{0.14}),\qquad M/B^{0.14}\asymp B^{0.18}.

The original full-sample event Ωdeg\Omega_{\mathrm{deg}} implies GR′\mathcal{G}_{R}^{\prime} for every subset J′J^{\prime} simultaneously, by nonnegativity of the envelope. Consequently

P(Ωdeg and some assignment has ZJ′,σ,SJ′′>E0+ϵ2/2)≤NBexp⁡(−cϵ2M/B0.14)=o(1).\mathbb{P}\left(\Omega_{\mathrm{deg}}\text{ and some assignment has }Z_{J^{\prime},\sigma,S_{J^{\prime}}}^{\prime}>E_{0}+\epsilon_{2}/2\right)\le N_{B}\exp\left(-c_{\epsilon_{2}}M/B^{0.14}\right)=o(1).

Only the exponential extension tails have been union-bounded; the probability of Ωdegc\Omega_{\mathrm{deg}}^{c} is counted once.

With probability 1−o(1)1-o(1), the squared-mass event also holds, and Lemma 9.3 now bounds the fully adaptive cut norm by E0+ϵ2/2+O(B−0.01)+O(B−0.11)E_{0}+\epsilon_{2}/2+O(B^{-0.01})+O(B^{-0.11}). On the exceptional event use ∥E∥□≤AB\lVert\mathcal{E}\rVert_{\square}\le A_{B}. Lemma 9.2 makes its expected contribution o(1)o(1). This proves (74). All probabilities and errors used here were uniform for each fixed ss in the compact range, which proves the stated uniformity without constructing a simultaneous event over ss.

The deterministic row estimate (73) gives

∑i,k∣Mik∣≪η,m1M=O(B0.32),\sum_{i,k}\lvert\mathcal{M}_{ik}\rvert\ll_{\eta,m_{1}}M=O(B^{0.32}),

uniformly in the site types and in ss. Thus M\mathcal{M} satisfies the O(B10)O(B^{10}) total-mass hypothesis of the independent-root comparison.

Finally, combine (74) with the independent-site comparison (43). For the actual sets Si(n)={p∈P:p∣n+i}S_{i}(n)=\{p\in\mathcal{P}:p\mid n+i\} and fixed BB, the upper limit as x→∞x\to\infty of the absolute error in replacing the kernel in (37) by Mik(Si(n),Sk(n);n/(Tx))\mathcal{M}_{ik}(S_{i}(n),S_{k}(n);n/(Tx)) is at most

Clab(E0+ϵ2)+o(1).C_{\mathrm{lab}}(E_{0}+\epsilon_{2})+o(1).

Here xx tends along the fixed bad subsequence, and the o(1)o(1) tends to zero as B→∞B\to\infty after this inner upper limit. The constant ClabC_{\mathrm{lab}} depends only on the fixed label bound and the compact ss-range. Indeed bounded complex label energies are controlled by a fixed multiple of the real cut norm. For each fixed BB, the error norm is a finite maximum of finite-residue functions continuous in ss (piecewise continuity would also suffice), so ordinary scale counting gives its Haar expectation integrated over ss, as explained after (37). The uniform expectation bound above can then be integrated over that fixed compact interval. This uses no independence between the actual labels and site types, and all matrices retain zero entries on nonallowed pairs.

Vanishing of the main energy

The main kernel in (71) has finitely many bounded site features. We first approximate these features by multiplicative weights and the lag factor by a function with a fixed period. Subdividing the position block then reduces its energy to products of the short averages in Lemma 2.3. In the vanishing argument, the regularity parameters, C∗C_{*}, η\eta, C6C_{6}, and the coarse partition are fixed; further approximation parameters are chosen before the limits. We take x→∞x\to\infty along the bad subsequence fixed in Section 2, and only then B→∞B\to\infty.

Choose a compact interval Is=[s0,s1]⊂(0,∞)\mathcal{I}_{s}=[s_{0},s_{1}]\subset(0,\infty) containing the supports in ss of all the kernels under consideration. It depends only on the fixed smooth functions: the support conditions ψ(u1)ψ(u2)≠0\psi(u_{1})\psi(u_{2})\ne0 and ϕ(s/u1)ϕ(s/u2)≠0\phi(s/u_{1})\phi(s/u_{2})\ne0 put ss in such a fixed interval. For the actual site sets

Si(n)={p∈P:p∣n+i},S_{i}(n)=\{p\in\mathcal{P}:p\mid n+i\},

write the main-kernel energy as

QB,xmain=1Tx∑n:n/(Tx)∈Is1M∑1≤i,k≤MηT≤∣k−i∣≤TFx(n+i)‾Fx(n+k)Mik(Si(n),Sk(n);n/(Tx)).(76)Q^{\mathrm{main}}_{B,x}=\frac{1}{T_x}\sum_{n:n/(T_x)\in I_s}\frac{1}{M}\sum_{\substack{1\leq i,k\leq M\\ \eta T\leq|k-i|\leq T}}\overline{F_x(n+i)}F_x(n+k)\mathcal{M}_{ik}\left(S_i(n),S_k(n);n/(T_x)\right). \tag*{(76)}

The summands are zero away from the original smooth supports. In particular, the choice of the enclosing interval introduces no new term.

Lemma 10.1 (Approximation of the site features). Fix the coarse partition I=⋃l=1m1IlI=\bigcup_{l=1}^{m_1}I_l used in (71). For every ζ>0\zeta>0 there are real polynomials

pl(t)=∑ν=0dlαl,νtν,1≤l≤m1,p_l(t)=\sum_{\nu=0}^{d_l}\alpha_{l,\nu}t^\nu,\qquad1\leq l\leq m_1,

independent of BB, such that the functions

V~l,B(S)=∑ν=0dlαl,ν∏p∈S1+p−ν/B2\widetilde{V}_{l,B}(S)=\sum_{\nu=0}^{d_l}\alpha_{l,\nu}\prod_{p\in S}\frac{1+p^{-\nu/B}}{2}

satisfy, for all sufficiently large BB,

∥Vl−V~l,B∥L2(S)≤ζ.\lVert V_l-\widetilde{V}_{l,B}\rVert_{L^2(S)}\leq\zeta.

Here the norm uses the independent site law with inclusion probabilities 1/p1/p. For 0<ζ≤10<\zeta\leq1, the approximants can be chosen with a common bound depending only on the fixed coarse partition. At actual integer sites they are the fixed finite combinations

V~l,B(Si(n))=∑ν=0dlαl,νGB,ν(n+i),\widetilde{V}_{l,B}(S_i(n))=\sum_{\nu=0}^{d_l}\alpha_{l,\nu}G_{B,\nu}(n+i),

where GB,νG_{B,\nu} is the function in (9) with PB=PP_B=\mathcal{P}.

Proof. Write xb=(log⁡b)/Bx_b=(\log b)/B for the normalized logarithm of the fair split product bb, and put ql0(y)=1Il(y)/∣Il∣q_l^0(y)=1_{I_l}(y)/|I_l| for y≥0y\geq0. The endpoints of the coarse cells lie in [1/2,3][1/2,3], where (23) bounds the marginal split-product densities uniformly for large BB. Choose small endpoint neighborhoods, still in a fixed compact subinterval of (0,3.2)(0,3.2), and replace ql0q_l^0 by a continuous compactly supported function qlq_l. We may arrange 0≤ql≤1/∣Il∣0\leq q_l\leq1/|I_l|, with equality to ql0q_l^0 outside these neighborhoods. By (23), their ν1/2\nu_{1/2} mass is at most a constant times their total length plus O(B−80)O(B^{-80}). Consequently the neighborhoods can be fixed so that

lim sup⁡B→∞Eν1/2∣ql0(xb)−ql(xb)∣2\limsup_{B\to\infty}\mathbb{E}_{\nu_{1/2}}\left|q_l^0(x_b)-q_l(x_b)\right|^2

is as small as desired, simultaneously for the finitely many cells. The function t↦ql(−log⁡t)t\mapsto q_l(-\log t) on (0,1](0,1], assigned value zero at t=0t=0, is continuous on [0,1][0,1]. Indeed qlq_l is zero for all sufficiently large arguments. Uniform polynomial approximation on [0,1][0,1] therefore gives a real polynomial plp_l for which pl(e−y)p_l(e^{-y}) approximates ql(y)q_l(y) uniformly for every y≥0y\geq0. Both this approximation and the preceding endpoint smoothing can be chosen so that

∥ql0(xb)−pl(e−xb)∥L2(ν1/2)≤ζ\lVert q_l^0(x_b)-p_l(e^{-x_b})\rVert_{L^2(\nu_{1/2})}\leq\zeta

for all sufficiently large BB. Taking the uniform polynomial error at most one also ensures sup⁡0≤t≤1∣pl(t)∣≤1/∣Il∣+1\sup_{0\le t\le1}|p_l(t)| \le1/|I_l|+1.

Conditional on a site set SS, fair splitting gives

Vl(S)=Esplit[ql0(xb)∣S],Esplit[e−vxb∣S]=∏p∈S(12+12p−v/B).V_l(S)=\mathbb{E}_{\mathrm{split}}[q_l^0(x_b)\mid S],\qquad\mathbb{E}_{\mathrm{split}}[e^{-v x_b}\mid S]=\prod_{p\in S}\left(\frac{1}{2}+\frac{1}{2}p^{-v/B}\right).

The second identity follows by making the fair split choices separately for each prime. Conditional expectation is a contraction in L2L^2, which proves the asserted error bound and the uniform bound on the approximants. The last identity applies to the distinct primes dividing n+in+i, irrespective of their multiplicities, and hence gives exactly GB,v(n+i)G_{B,v}(n+i) at the integer sites.

We spell out how the feature error is used against the labels. For fixed BB, any function of Si(n)S_i(n) is a function of finitely many residues of n+in+i. Unweighted progression counting gives, for r=1,2r=1,2,

lim⁡x→∞1Tx∑n:n/(Tx)∈Is∣Vl(Si(n))−V~l,B(Si(n))∣r=∣Is∣E∣Vl(S)−V~l,B(S)∣r.(77)\lim_{x\to\infty}\frac{1}{Tx}\sum_{\substack{n:\,n/(Tx)\in I_s}}\left|V_l(S_i(n))-\widetilde V_{l,B}(S_i(n))\right|^r=|I_s|\mathbb{E}\left|V_l(S)-\widetilde V_{l,B}(S)\right|^r. \tag*{(77)}

There are only finitely many ii at fixed BB; neither a bound on the size of the residue modulus nor uniformity with respect to growing BB is required for this inner limit.

The singular series in eq:8.13 is nonnegative and at most j/φ(j)j/\varphi(j). Its ordinary mean is bounded. Thus the lag majorant has deterministic row bounds

sup⁡1≤i≤M1T∑1≤k≤M0<∣k−i∣≤TS(∣k−i∣)≪1.(78)\sup_{1\le i\le M}\frac{1}{T}\sum_{\substack{1\le k\le M\\0<|k-i|\le T}}\mathfrak{S}(|k-i|)\ll1. \tag*{(78)}

For completeness, the identity

jφ(j)=∑d∣jμMob2(d)φ(d)\frac{j}{\varphi(j)}=\sum_{d\mid j}\frac{\mu_{\mathrm{Mob}}^2(d)}{\varphi(d)}

and the convergence of ∑dμMob2(d)/(dφ(d))\sum_d\mu_{\mathrm{Mob}}^2(d)/(d\varphi(d)) give this bound by summing the divisor expansion up to TT. Since ∣Fx∣≤2|F_x|\le2, the features and their approximants are bounded, and κB\kappa_B, ww and the finitely many hlh_l are bounded for the present fixed parameters, replacing both features in each term of (76) costs at most CζC\zeta in iterated upper limit. Indeed

∣Vl(Si)Vl(Sk)−V~l,B(Si)V~l,B(Sk)∣≤Cm1(∣Vl(Si)−V~l,B(Si)∣+∣Vl(Sk)−V~l,B(Sk)∣),|V_l(S_i)V_l(S_k)-\widetilde V_{l,B}(S_i)\widetilde V_{l,B}(S_k)|\le C_{m_1}\left(|V_l(S_i)-\widetilde V_{l,B}(S_i)|+|V_l(S_k)-\widetilde V_{l,B}(S_k)|\right),

and (78) reduces the sum to the single-site L1L^1 errors in (77). This uses no independence between either feature error and the labels or the other endpoint.

Lemma 10.2 (Periodic approximation of the lag factor). For a fixed prime cutoff P1≥2P_1\ge2, define on all integers jj

SP1(j)=∏p≤P1(1−1p)−1∏p≤P1p∣j(1−1(p−1)2).\mathfrak{S}_{P_1}(j)=\prod_{p\le P_1}\left(1-\frac{1}{p}\right)^{-1}\prod_{\substack{p\le P_1\\p\mid j}}\left(1-\frac{1}{(p-1)^2}\right).

This is a bounded nonnegative function of period D=∏p≤P1pD=\prod_{p\le P_1}p. Moreover,

lim⁡P1→∞sup⁡T≥11T∑1≤j≤T∣S(j)−SP1(j)∣=0,\lim_{P_1\to\infty}\sup_{T\ge1}\frac{1}{T}\sum_{1\le j\le T}|\mathfrak{S}(j)-\mathfrak{S}_{P_1}(j)|=0,

with the supremum taken over positive integers TT.

Proof. Write

a(j)=∏p∣j(1−1p)−1,aP1(j)=∏p≤P1p∣j(1−1p)−1.a(j)=\prod_{p\mid j}\left(1-\frac{1}{p}\right)^{-1},\qquad a_{P_1}(j)=\prod_{\substack{p\le P_1\\p\mid j}}\left(1-\frac{1}{p}\right)^{-1}.

The omitted nondividing-prime factors form a product of numbers in (0,1](0,1]. Since P1≥2P_1\ge2, its difference from one is at most

∑p>P11(p−1)2≪P1−1,\sum_{p>P_1}\frac{1}{(p-1)^2}\ll P_1^{-1},

uniformly in jj. All the retained nondividing-prime factors are in [0,1][0,1], including the possible zero factor at p=2p=2. Consequently

∣S(j)−SP1(j)∣≤a(j)−aP1(j)+CP1−1a(j).\left|\mathcal{S}(j)-\mathcal{S}_{P_1}(j)\right|\le a(j)-a_{P_1}(j)+CP_1^{-1}a(j).

The positive divisor expansion gives

a(j)−aP1(j)=∑d∣jd has a prime divisor greater than P1μMob2(d)φ(d).a(j)-a_{P_1}(j)=\sum_{\substack{d\mid j\\d\text{ has a prime divisor greater than }P_1}}\frac{\mu_{\mathrm{Mob}}^2(d)}{\varphi(d)}.

Its mean up to TT is bounded by

∑d≥1d has a prime divisor greater than P1μMob2(d)dφ(d).\sum_{\substack{d\ge1\\d\text{ has a prime divisor greater than }P_1}}\frac{\mu_{\mathrm{Mob}}^2(d)}{d\varphi(d)}.

The unrestricted positive series equals ∏p(1+1/(p(p−1)))<∞\prod_p(1+1/(p(p-1)))<\infty, so this tail tends to zero. The same expansion bounds the mean of a(j)a(j) by that product, uniformly in TT. The required conclusion follows.

Proposition 10.3 (Vanishing of the main energy). With C∗C_*, η\eta, C6C_6, and the coarse partition fixed as above,

lim⁡B→∞lim sup⁡x→∞∣QB,xmain∣=0.\lim_{B\to\infty}\limsup_{x\to\infty}\left|Q_{B,x}^{\mathrm{main}}\right|=0.

Proof. Let γ>0\gamma>0 be arbitrary. First choose the approximants in Lemma 10.1 with sufficiently small ζ\zeta that the feature replacement has iterated upper-limit error at most γ/3\gamma/3. Their degrees and coefficients are now fixed, and the approximants have the stated bounds depending only on the fixed coarse partition. Next choose a fixed P1P_1 by Lemma 10.2. Replacing S(∣k−i∣)\mathcal{S}(|k-i|) by SP1(k−i)\mathcal{S}_{P_1}(k-i) in the energy costs at most γ/3\gamma/3 in upper limit. To check the normalization, the absolute error is bounded by a fixed constant times

1MT∑i=1M∑1≤k≤M0<∣k−i∣≤T∣S(∣k−i∣)−SP1(k−i)∣≤2T∑j=1T∣S(j)−SP1(j)∣.\frac{1}{MT}\sum_{i=1}^{M}\sum_{\substack{1\le k\le M\\0<|k-i|\le T}}\left|\mathcal{S}(|k-i|)-\mathcal{S}_{P_1}(k-i)\right| \le\frac{2}{T}\sum_{j=1}^{T}\left|\mathcal{S}(j)-\mathcal{S}_{P_1}(j)\right|.

The label and origin bounds have been absorbed in that fixed constant. Choose a fixed 0<δ<η/40<\delta<\eta/4, to be made small after P1P_1 has been fixed. Partition {1,…,M}\{1,\ldots,M\} into a fixed number R0R_0 of consecutive blocks JaJ_a, 1≤a≤R01\le a\le R_0, of lengths LaL_a differing by at most one. Choosing R0R_0 sufficiently large in terms of δ\delta, C6C_6, for all sufficiently large BB we have

cδ,C6T≤La≤δT(1≤a≤R0).c_{\delta,C_6}T\le L_a\le\delta T\qquad(1\le a\le R_0).

with a positive fixed lower constant. In particular all these lengths tend to infinity.

Call an ordered block pair fully allowed when every (i,k)∈Ja×Jb(i,k) \in J_a \times J_b satisfies ηT≤∣k−i∣≤T\eta T \le|k-i| \le T. Discard any block pair containing both an allowed and a nonallowed pair of positions. The values of ∣k−i∣|k-i| on a block pair vary by at most 2δT2\delta T. Thus each discarded allowed pair has lag within 2δT2\delta T of ηT\eta T or TT. For each row there are only O(δT+1)O(\delta T+1) such positions. The factors now present, including the periodic lag factor, are bounded for the fixed P1P_1 and coarse partition. With the normalization 1/(MT)1/(MT), this discarding costs O(δ)+O(1/T)O(\delta)+O(1/T), with a constant allowed to depend on these already fixed parameters.

Choose a representative position ia∈Jai_a \in J_a for every block, and for every fully allowed block pair put λab=∣ib−ia∣/T∈[η,1]\lambda_{ab}=|i_b-i_a|/T \in[\eta,1]. Smoothness of ww on the fixed compact ss-range and on η≤λ≤1\eta\le\lambda\le1 gives

sup⁡s∈Is∣w(s,∣k−i∣/T)−w(s,λab)∣≪ηδ(i∈Ja,k∈Jb).\sup_{s\in I_s} |w(s,|k-i|/T)-w(s,\lambda_{ab})| \ll_{\eta} \delta\qquad(i\in J_a,k\in J_b).

Replacing the weight by this representative value therefore has another O(δ)O(\delta) cost. We fix δ\delta sufficiently small that the two block errors together have upper limit at most γ/3\gamma/3. Neither the representatives nor their scaled lags have to converge as BB grows; only this uniform bound is used.

It remains to show that the fully factored expression tends to zero. For 0≤r<D0 \le r < D and 1≤l≤m11 \le l \le m_1, define

Aa,r,l(n)=1La∑i∈Jai≡r(modD)Fx(n+i)V~l,B(Si(n)).A_{a,r,l}(n)=\frac{1}{L_a}\sum_{\substack{i\in J_a\\i\equiv r\pmod D}} F_x(n+i)\widetilde{V}_{l,B}(S_i(n)).

By the feature identity, this is a fixed finite linear combination of

1La∑i∈Jai≡r(modD)Fx(n+i)GB,v(n+i).\frac{1}{L_a}\sum_{\substack{i\in J_a\\i\equiv r\pmod D}} F_x(n+i)G_{B,v}(n+i).

The lengths LaL_a tend to infinity and the starting indices are fixed for each BB. The modulus DD, monomials vv, and polynomial coefficients are all fixed before BB grows. Lemma 2.3 applies with origins n∈[s0Tx,s1Tx]n\in[s_0T_x,s_1T_x], and yields

lim⁡B→∞lim sup⁡x→∞1Tx∑n:n/(Tx)∈Is∣Aa,r,l(n)∣2=0(79)\lim_{B\to\infty}\limsup_{x\to\infty}\frac{1}{T_x}\sum_{\substack{n:n/(T_x)\in I_s}} |A_{a,r,l}(n)|^2=0 \tag*{(79)}

for every one of the finitely many indices.

Set cr,r′=SP1(r′−r)c_{r,r'}=\mathcal{S}_{P_1}(r'-r), using the periodic extension in Lemma 10.2, including at the zero residue. On a block pair, splitting each position by its residue class gives the factored expression

kBTx∑n:n/(Tx)∈Is∑l=1m1hl∑(a,b)fully allowedLaLbMTw(n/(Tx),λab)∑r,r′=0D−1cr,r′Aa,r,l(n)‾Ab,r′,l(n).(80)\frac{k_B}{T_x}\sum_{\substack{n:n/(T_x)\in I_s}}\sum_{l=1}^{m_1}h_l\sum_{\substack{(a,b)\\\text{fully allowed}}}\frac{L_aL_b}{MT}w\left(n/(T_x),\lambda_{ab}\right)\sum_{r,r'=0}^{D-1}c_{r,r'}\overline{A_{a,r,l}(n)}A_{b,r',l}(n). \tag*{(80)}

The feature polynomials are real, so the conjugation here matches exactly the conjugation in the original energy. The factors LaLb/(MT)L_aL_b/(MT) are uniformly bounded. The numbers of cells, blocks, and residue classes are fixed, and kBk_B and the representative smooth weights are uniformly bounded. Cauchy–Schwarz in the origin variable and (79) make every term in (80) tend to zero in the asserted iterated upper-limit sense. This remains true when the bounded coefficients depend on n/(Tx)n/(T_x) or on BB.

The original main energy consequently has iterated upper limit at most γ\gamma in absolute value. Since the approximation parameters were fixed before the limits and γ>0\gamma>0 was arbitrary, the proposition follows.

□\square

Proof of Proposition 2.2. Suppose mixed decorrelation fails for some fixed JJ and two fixed phase vectors. Fix the resulting subsequence, the limiting profile WW and bump ϕ\phi from Section 2, and the smooth functions ρ0\rho_0, ψ\psi from Section 4. Choose the regularity grid and tolerances as required for the estimates of Sections 3 and 6. The amplification then supplies the positive constant c5c_5 in (33). This constant is independent of every sufficiently large fixed C∗C_*.

We collect the upper bounds, recording the dependencies needed to choose parameters without a cycle. The diagonal contribution to I2,x/(BT)I_{2,x}/(BT) tends to zero in the prescribed order. Equations (34)–(37) express the remaining contribution as the block energy, up to errors bounded in iterated upper limit by

C0η+C0/C6.C_0\eta+ C_0/C_6.

Here C0C_0 can be chosen independently of C∗C_*, η\eta and C6C_6, by the untruncated mean bound (36).

The root comparison (43) and the sampling bound (74) allow replacement of this block kernel by M\mathcal{M}. The resulting upper-limit error is bounded by

C1e−c3C∗+Cη,C6ϵ1+C2ϵ2.C_1e^{-c_3C_*} + C_{\eta,C_6}\epsilon_1 + C_2\epsilon_2.

The constants C1C_1, C2C_2 do not depend on the later choices η\eta, C6C_6 or the coarse partition. Indeed the truncation term in (72) has a coefficient independent of the lag range and block length, and converting a real cut-norm bound to an energy of labels of absolute value at most two costs only an absolute factor. For example, splitting each label into its real and imaginary parts gives a factor at most 16. Integrating over IsI_s costs only its fixed length. The finite-residue counting passage and the uniformity in ss were established in Sections 5 and 9. The remaining o(1)o(1) errors vanish as B→∞B \to\infty after the inner xx-limit, with all the displayed parameters fixed.

For any such fixed choices, Proposition 10.3 makes the main energy zero in upper limit. Consequently

lim sup⁡B→∞lim sup⁡x→∞I2,xBT≤C1e−c3C∗+C0η+C0/C6+Cη,C6ϵ1+C2ϵ2.(81)\limsup_{B\to\infty}\limsup_{x\to\infty}\frac{I_{2,x}}{BT} \le C_1e^{-c_3C_*}+C_0\eta+C_0/C_6+C_{\eta,C_6}\epsilon_1+C_2\epsilon_2. \tag*{(81)}

First choose a sufficiently large fixed C∗C_* so that the first term is less than c5/10c_5/10. Next choose a sufficiently small fixed η\eta and a sufficiently large fixed C6C_6 so that the next two terms are each less than c5/10c_5/10. Then choose positive fixed ϵ1\epsilon_1, ϵ2\epsilon_2 so that their terms are each less than c5/10c_5/10, and choose the corresponding coarse partition in (54). All feature, periodic and block approximations in Proposition 10.3 are made subsequently with these parameters fixed. Thus (81) is strictly smaller than c5c_5, contradicting the lower bound (33).

Finally, the bad subsequence was extracted from an arbitrary sequence of integer scales on which x−1∑1≤n<xgx(n)Fx(n+1)x^{-1}\sum_{1\le n<x}g_x(n)F_x(n+1) stays a fixed positive distance from zero. Every inner limit above uses the final subsequence selected for the profile. At each fixed BB, all multipliers, residue moduli and shifts are finite, so the argument never requires their invariance or counting estimates while BB and xx grow simultaneously. The contradiction rules out every original sequence of this kind. Thus the mixed correlation tends to zero along the full sequence of integer scales, for every fixed pair of phase vectors.

Marginals, the joint law, and the ordering corollary

Proposition 2.2 separates every pair of bin characters. Finite Fourier inversion therefore separates bin events, including smoothness at rational powers of the counting scale. The factorial-moment limits from Section 2 identify their marginals with the Dickman law. We then obtain the fixed-scale joint law through every real counting endpoint, pass to the moving thresholds in Theorem 1.1, and deduce the comparison corollary from the continuous limiting distribution.

Finite Fourier inversion

Adding back the centering in Proposition 2.2 gives

lim⁡x→∞1x∑1≤n<xgx(n)‾fx(n+1)=μg‾μ.(82)\lim_{x\to\infty}\frac{1}{x}\sum_{1\le n<x}\overline{g_x(n)}f_x(n+1)=\overline{\mu_g}\mu. \tag*{(82)}

Here and throughout this subsection, xx runs through positive integers. Indeed, the mean of gx(n)‾\overline{g_x(n)} tends to μg‾\overline{\mu_g} by Lemma 2.1 and its initial-segment extension. For m≤xm\le x, write

bx(m)=(ΩBk,x(m))1≤k<J.b_x(m)=(\Omega_{B_{k,x}}(m))_{1\le k<J}.

Every coordinate lies in {0,…,J}\{0,\ldots,J\}: each counted prime factor exceeds x1/Jx^{1/J}. Thus reduction modulo J+1J+1 is injective on the set of count vectors. Put ζ0=exp⁡(2πi/(J+1))\zeta_0=\exp(2\pi i/(J+1)). For r∈{0,…,J}J−1r\in\{0,\ldots,J\}^{J-1} and such a count vector bb, the exact identity

1{b=r}=1(J+1)J−1∑h∈{0,…,J}J−1ζ0h⋅(b−r)\mathbf{1}_{\{b=r\}}=\frac{1}{(J+1)^{J-1}}\sum_{h\in\{0,\ldots,J\}^{J-1}}\zeta_0^{h\cdot(b-r)}

expresses each bin event as a finite linear combination of characters. The two phase vectors in (82) are independent choices, and complex conjugation simply permutes the available root-of-unity phases. Applying this identity at nn and n+1n+1 proves factorization of the limiting expectations of every pair of fixed functions of their count vectors. Both marginal limits exist by Lemma 2.1; the shift of the averaging range changes a bounded marginal average by O(1/x)O(1/x).

If c,d∈(0,1)c,d\in(0,1) are rational, choose JJ so that both are grid points. For m≤xm\le x, the condition P+(m)≤xcP^+(m)\le x^c is exactly the vanishing of all counts in bins with lower endpoint at least xcx^c. Hence the joint smoothness event with thresholds xc,xdx^c,x^d has a limiting density equal to the product of its two marginal limits. The next calculation identifies those limits.

The Dickman marginal

We recover the classical marginal law, originating with Dickman and developed by Ramaswami and de Bruijn [4, 6, 22], from the factorial moments already computed for the bin counts.

Lemma 11.1 (Dickman marginal). For every fixed c∈(0,1)c\in(0,1),

lim⁡x→∞1x#{1≤m≤x:P+(m)≤xc}=ρ(1/c).(83)\lim_{x\to\infty}\frac{1}{x}\#\{1\le m\le x:P^+(m)\le x^c\}=\rho(1/c). \tag*{(83)}

The limit is through all real xx.

Proof. First suppose that c=k0/Jc=k_0/J for integers J≥2J\ge2 and 1≤k0<J1\le k_0<J. For m≤xm\le x, set

Nc,x(m)=∑k=k0J−1ΩBk,x(m).N_{c,x}(m)=\sum_{k=k_0}^{J-1}\Omega_{B_{k,x}}(m).

Then P+(m)≤xcP^+(m)\le x^c exactly when Nc,x(m)=0N_{c,x}(m)=0, and 0≤Nc,x(m)≤J0\le N_{c,x}(m)\le J. For a nonnegative integer zz, write (z)h=z(z−1)⋯(z−h+1)(z)_h=z(z-1)\cdots(z-h+1), with (z)0=1(z)_0=1. The finite inclusion–exclusion identity is

1{Nc,x(m)=0}=∑h=0J(−1)hh!(Nc,x(m))h.\mathbf{1}_{\{N_{c,x}(m)=0\}}=\sum_{h=0}^{J}\frac{(-1)^h}{h!}(N_{c,x}(m))_h.

For each 0≤h≤J0 \le h \le J, the falling-factorial multinomial identity gives

(Nc,x(m))h=∑rk0,…,rJ−1≥0rk0+⋯+rJ−1=hh!∏k=k0J−1rk!∏k=k0J−1(ΩBk,x(m))rk.(N_{c,x}(m))_h=\sum_{\substack{r_{k_0},\ldots,r_{J-1}\ge0\\r_{k_0}+\cdots+r_{J-1}=h}}\frac{h!}{\prod_{k=k_0}^{J-1}r_k!}\prod_{k=k_0}^{J-1}(\Omega_{B_{k,x}}(m))_{r_k}.

Apply the initial-segment form of (2.9) to every term, with rk=0r_k=0 for k<k0k<k_0. The multinomial coefficient counts the assignments of the hh ordered variables to the bins. Summing the resulting simplex integrals therefore joins those bins into (c,1](c,1] and gives

lim⁡x→∞1x∑1≤m≤x(Nc,x(m))h=∫s1,…,sh∈(c,1]s1+⋯+sh≤1∏i=1hdsisi.\lim_{x\to\infty}\frac{1}{x}\sum_{1\le m\le x}(N_{c,x}(m))_h=\int_{\substack{s_1,\ldots,s_h\in(c,1]\\s_1+\cdots+s_h\le1}}\prod_{i=1}^{h}\frac{ds_i}{s_i}.

For h=0h=0, the empty integral is one, and the left-hand limit is also one. The initial-segment limit used here holds through real xx, as does the factorial-moment limit in Section 2.

In this integral put si=cvis_i=cv_i. The upper bound on each viv_i is then implied by v1+⋯+vh≤1/cv_1+\cdots+v_h\le1/c, and changing the lower boundary from vi>1v_i>1 to vi≥1v_i\ge1 does not change the integral. Averaging the inclusion–exclusion identity shows that the rational smoothness density is H(1/c)H(1/c), where for u≥0u\ge0 we put

H(u)=1+∑h≥1(−1)hh!Ih(u),Ih(u)=∫v1,…,vh≥1v1+⋯+vh≤u∏i=1hdvivi.(84)H(u)=1+\sum_{h\ge1}\frac{(-1)^h}{h!}I_h(u),\qquad I_h(u)=\int_{\substack{v_1,\ldots,v_h\ge1\\v_1+\cdots+v_h\le u}}\prod_{i=1}^{h}\frac{dv_i}{v_i}. \tag*{(84)}

The sum is locally finite because Ih(u)=0I_h(u)=0 when h>uh>u. In particular, 1/c≤J1/c\le J ensures that the sum at u=1/cu=1/c agrees with the finite inclusion–exclusion sum above. Put I0(u)=1I_0(u)=1 for u≥0u\ge0.

For h≥1h\ge1 and y>z≥1y>z\ge1, symmetry and the identity 1=(v1+⋯+vh)/(v1+⋯+vh)1=(v_1+\cdots+v_h)/(v_1+\cdots+v_h) give

Ih(y)−Ih(z)=h∫zyIh−1(t−1)dtt.I_h(y)-I_h(z)=h\int_z^y I_{h-1}(t-1)\frac{dt}{t}.

To verify this, cancel the last denominator after using symmetry and set t=v1+⋯+vht=v_1+\cdots+v_h. The remaining variables satisfy v1+⋯+vh−1≤t−1v_1+\cdots+v_{h-1}\le t-1. The same formula holds for h=1h=1. Thus the IhI_h are continuous, H=1H=1 on [0,1][0,1], and

uH′(u)=−H(u−1)(u>1).uH'(u)=-H(u-1)\qquad(u>1).

Successive integration on the intervals [k,k+1][k,k+1] uniquely determines a continuous solution from its values on [0,1][0,1]. Hence H=ρH=\rho, proving (11.2) for rational cc.

For an arbitrary real c∈(0,1)c\in(0,1), choose rational c−<c<c+c_-<c<c_+ in (0,1)(0,1). The smoothness event at xc−x^{c_-} is contained in the event at xcx^c, which is contained in the event at xc+x^{c_+}. The rational limits bound the lower and upper limits of the middle normalized count by ρ(1/c−)\rho(1/c_-) and ρ(1/c+)\rho(1/c_+). Letting c−,c+c_-,c_+ approach cc and using continuity of ρ\rho at the finite argument 1/c1/c proves (11.2).

The marginal law also supplies the endpoint behavior of the limiting distribution. It gives 0≤ρ(u)≤10\le\rho(u)\le1 for u>1u>1, and the same bounds hold at u=1u=1 by definition. The delay equation makes ρ\rho nonincreasing on [1,∞)[1,\infty). Its limit at infinity must be zero: if that limit were L>0L>0, integrating ρ′(u)=−ρ(u−1)/u≤−L/u\rho'(u)=-\rho(u-1)/u\le-L/u would contradict nonnegativity. Consequently

D(t)={0,t≤0,ρ(1/t),0<t<1,1,t≥1.(85)D(t)= \begin{cases} 0, & t\le0,\\ \rho(1/t), & 0<t<1,\\ 1, & t\ge1. \end{cases} \tag*{(85)}

is a continuous distribution function, including at zero and one, and its probability measure has no atoms.

Fixed and moving thresholds

Combining the Fourier inversion with Lemma 11.1 gives, for every pair of rational c,d∈(0,1)c,d \in(0,1),

lim⁡x→∞1x#{1≤n<x:P+(n)≤xc, P+(n+1)≤xd}=ρ(1/c)ρ(1/d).(86)\lim_{x\to\infty}\frac{1}{x}\#\{1\le n<x:P^{+}(n)\le x^{c},\ P^{+}(n+1)\le x^{d}\}=\rho(1/c)\rho(1/d). \tag*{(86)}

We first extend this fixed-scale law to arbitrary real exponents and real counting endpoints. Fix c,d∈(0,1)c,d\in(0,1) and rational exponents c−<c<c+c_{-}<c<c_{+} and d−<d<d+d_{-}<d<d_{+} in (0,1)(0,1). Given real XX, put x=⌊X⌋+1x=\lfloor X\rfloor+1. Then x/X→1x/X\to1, the integer ranges n≤Xn\le X and n<xn<x agree, and for all sufficiently large XX,

xc−≤Xc≤xc+,xd−≤Xd≤xd+.x^{c_{-}}\le X^{c}\le x^{c_{+}},\qquad x^{d_{-}}\le X^{d}\le x^{d_{+}}.

Thus the normalized count over 1≤n≤X1\le n\le X with thresholds Xc,XdX^{c},X^{d} lies between x/Xx/X times the two normalized counts in (86) at the lower and upper rational exponents. Taking lower and upper limits through real XX, then letting the rational exponents approach c,dc,d, proves

lim⁡X→∞1X#{2≤n≤X:P+(n)≤Xc, P+(n+1)≤Xd}=D(c)D(d)=ρ(1/c)ρ(1/d)(0<c,d<1).(87)\begin{aligned} \lim_{X\to\infty}\frac{1}{X}\#\{2\le n\le X:P^{+}(n)\le X^{c},\ P^{+}(n+1)\le X^{d}\} \\ &=D(c)D(d)=\rho(1/c)\rho(1/d)\qquad(0<c,d<1). \tag*{(87)} \end{aligned}

The omission of n=1n=1 changes at most one summand. In (87), the exponents are fixed and the limit is through all real XX.

This also gives the fixed-scale upper-tail independence law discussed by Erdős and Pomerance [7]:

lim⁡X→∞1X#{2≤n≤X:P+(n)>Xc, P+(n+1)>Xd}=(1−D(c))(1−D(d))(0≤c,d≤1).(88)\begin{aligned} \lim_{X\to\infty}\frac{1}{X}\#\{2\le n\le X:P^{+}(n)>X^{c},\ P^{+}(n+1)>X^{d}\} \\ &=(1-D(c))(1-D(d))\qquad(0\le c,d\le1). \tag*{(88)} \end{aligned}

For interior exponents this follows by inclusion–exclusion from (87) and the two marginal laws. The marginal for n+1n+1 is the one in Lemma 11.1, since shifting its counting range changes only O(1)O(1) terms. Replacing either strict comparison in (88) by a weak one has the same limit. Indeed, for fixed c>0c>0, equality P+(m)=XcP^{+}(m)=X^{c} is impossible unless XcX^{c} is a prime pp; in that case every such mm is divisible by pp, giving only O(X1−c+1)=o(X)O(X^{1-c}+1)=o(X) possibilities for m≤X+1m\le X+1. At exponent zero, the upper-tail condition holds for every m≥2m\ge2, and equality can occur only at m=1m=1. At exponent one, an upper-tail condition on m≤X+1m\le X+1 has only O(1)O(1) possible integers. These observations give the endpoint cases of (88) as well. Its right-hand side is continuous on [0,1]2[0,1]^2 by (11.5).

Proof of Theorem 1.1. Fix a,b∈(0,1)a,b\in(0,1) and choose fixed exponents c−<a<c+c_{-}<a<c_{+} and d−<b<d+d_{-}<b<d_{+} in (0,1)(0,1). For each fixed ε∈(0,1)\varepsilon\in(0,1), all sufficiently large real XX satisfy, uniformly for εX≤n≤X\varepsilon X\le n\le X,

Xc−≤na≤Xc+,Xd−≤nb≤Xd+.X^{c_{-}}\le n^{a}\le X^{c_{+}},\qquad X^{d_{-}}\le n^{b}\le X^{d_{+}}.

On this interval, the fixed-scale event at the lower exponents is contained in the moving-threshold event of Theorem 1.1, which is contained in the fixed-scale event at the upper exponents. The discarded initial interval costs at most ε+O(1/X)\varepsilon+O(1/X) in normalized counting. The limit through all real XX in (87) therefore bounds the lower and upper limits of the desired count by the corresponding products at the lower and upper exponents, with that error. Let ε↓0\varepsilon\downarrow0, then let the exponents approach a,ba,b. Continuity of DD gives the claimed product through all real endpoints, with ordinary, unweighted counting.

The ordering probability

Proof of Corollary 1.2. For n≥2n \ge2, put

Un=log⁡P+(n)log⁡n,Vn=log⁡P+(n+1)log⁡n.U_n = \frac{\log P^{+}(n)}{\log n}, \qquad V_n = \frac{\log P^{+}(n+1)}{\log n}.

The empirical probability law means uniform sampling from 2≤n≤⌊X⌋2 \le n \le\lfloor X \rfloor; its normalization 1/(⌊X⌋−1)1/(\lfloor X \rfloor- 1) differs from 1/X1/X by a factor tending to one. Recall the continuous distribution function DD from (85).

Although 0≤Un≤10 \le U_n \le1, the second coordinate can slightly exceed one. This overshoot vanishes and does not alter the limiting law: for every n≥N≥2n \ge N \ge2,

0≤Vn≤log⁡(n+1)log⁡n≤1+1Nlog⁡N.0 \le V_n \le\frac{\log(n+1)}{\log n} \le1 + \frac{1}{N\log N}.

Consequently the empirical laws are tight and every limiting law is supported on [0,1]2[0,1]^2. More explicitly, clamp VnV_n to V~n=min⁡(Vn,1)\widetilde V_n = \min(V_n,1). For 0<a,b<10 < a,b < 1, the joint distribution functions of (Un,V~n)(U_n,\widetilde V_n) are exactly those of (Un,Vn)(U_n,V_n), hence converge by Theorem 1.1 to D(a)D(b)D(a)D(b). These interior rectangles determine the limiting probability law: their masses approach one as a,b↑1a,b \uparrow1, and finite differences give the masses of all interior grid cells. Approximating continuous functions on [0,1]2[0,1]^2 by such grids proves weak convergence to the product law with marginal distribution function DD. Removing the clamping does not affect this convergence, by the displayed bound and the vanishing proportion of n<Nn < N.

Let U,VU,V be independent with distribution function DD. Continuity gives P(U=V)=0\mathbb{P}(U=V)=0, and exchangeability of this limiting pair gives

P(U<V)=P(V<U)=12.\mathbb{P}(U<V)=\mathbb{P}(V<U)=\frac{1}{2}.

The boundary of the set {(u,v):u<v}\{(u,v):u<v\} is the diagonal, which has zero product mass. Weak convergence therefore yields

lim⁡X→∞1X#{2≤n≤X:Un<Vn}=12.\lim_{X\to\infty}\frac{1}{X}\#\{2\le n\le X:U_n<V_n\}=\frac{1}{2}.

Since both logarithmic sizes have the same positive denominator, Un<VnU_n<V_n is exactly P+(n)<P+(n+1)P^{+}(n)<P^{+}(n+1). Applying the same continuity-set argument to v<uv<u gives the reverse ordering limit as well. The symmetry used here belongs to the independent limiting law; no symmetry of the finite consecutive-integer pairs is assumed. □

References

References

  1. [1]Noga Alon, W. Fernandez de la Vega, Ravi Kannan, and Marek Karpinski. Random sampling and approximation of MAX-CSPs. Journal of Computer and System Sciences, 67(2):212–243, 2003. doi: 10.1016/S0022-0000(03)00008-4.DOI
  2. [2]Christian Borgs, Jennifer T. Chayes, László Lovász, Vera T. Sós, and Katalin Vesztergombi. Convergent sequences of dense graphs I: Subgraph frequencies, metric properties and testing. Advances in Mathematics, 219(6):1801–1851, 2008. doi: 10.1016/j.aim.2008.07.008.DOI
  3. [3]Telma Caputti. A note on the extension of Lipschitz functions. Revista de la Unión Matemática Argentina, 31:122–129, 1984. URL https://inmabb.criba.edu.ar/revuma/pdf/v31n3/p122-129.pdf.
  4. [4]N. G. de Bruijn. On the number of positive integers ≤ x and free of prime factors > y. Proceedings of the Koninklijke Nederlandse Akademie van Wetenschappen, Series A, 54(1):50–60, 1951. URL https://research.tue.nl/en/publications/on-the-number-of-positive-integers-leq-x-and-free-of-prime-factor/.DOI
  5. [5]Régis de la Bretèche, Carl Pomerance, and Gérald Tenenbaum. Products of ratios of consecutive integers. The Ramanujan Journal, 9(1–2):131–138, 2005. URL https://tenenb.perso.math.cnrs.fr/PPP/AB.pdf.DOI
  6. [6]Karl Dickman. On the frequency of numbers containing prime factors of a certain relative magnitude. Arkiv för Matematik, Astronomi och Fysik, 22A(10):1–14, 1930.
  7. [7]Paul Erdős and Carl Pomerance. On the largest prime factors of n and n + 1. Aequationes Mathematicae, 17:311–321, 1978. URL https://www.renyi.hu/~p_erdos/1978-29.pdf.
  8. [8]Kevin Ford. Zero-free regions for the Riemann zeta function. In M. A. Bennett, B. C. Berndt, N. Bost, H. G. Diamond, A. J. Hildebrand, and W. Philipp, editors, Number Theory for the Millennium, II, pages 25–56. A K Peters, Ltd., Natick, MA, 2002. URL https://arxiv.org/abs/1910.08205v5. Corrected version: arXiv:1910.08205v5, 18 February 2025.
  9. [9]Kevin Ford. Sieve methods lecture notes, spring 2023. Lecture notes, University of Illinois Urbana–Champaign, 2023. URL https://ford126.web.illinois.edu/sieve2023.pdf.
  10. [10]Alan Frieze and Ravi Kannan. Quick approximation to matrices and applications. Combinatorica, 19:175–220, 1999. doi: 10.1007/s004930050052.DOI
  11. [11]Andrew Granville and Dimitris Koukoulopoulos. Beyond the LSD method for the partial sums of multiplicative functions. The Ramanujan Journal, 49(2):287–319, 2019. doi: 10.1007/s11139-018-0119-3. URL https://dms.umontreal.ca/~koukoulo/documents/publications/LSD.pdf.DOI
  12. [12]Harald Andrés Helfgott and Maksym Radziwiłł. Expansion, divisibility and parity, 2021. URL https://arxiv.org/abs/2103.06853v2. Version 2, 13 April 2021.
  13. [13]Yujiao Jiang, Guangshi Lü, and Zhiwei Wang. Averaged forms of two conjectures of Erdős and Pomerance, and their applications. Advances in Mathematics, 409:108592, 2022. doi: 10.1016/j.aim.2022.108592. Part A, 42 pp.DOI
  14. [14]Dimitris Koukoulopoulos. The Distribution of Prime Numbers, volume 203 of Graduate Studies in Mathematics. American Mathematical Society, Providence, RI, 2019. ISBN 978-1-4704-4754-0. doi: 10.1090/gsm/203. URL https://dms.umontreal.ca/~koukoulo/documents/publications/primes.pdf. Theorem locators refer to the author’s publicly available preliminary version.
  15. [15]Xiaodong Lü and Zhiwei Wang. On the largest prime factors of consecutive integers. Monatshefte für Mathematik, 206(2):403–418, 2025. URL https://hal.science/hal-01797939.DOI
  16. [16]Kaisa Matomäki and Maksym Radziwiłł. Multiplicative functions in short intervals. Annals of Mathematics, 183(3):1015–1056, 2016. doi: 10.4007/annals.2016.183.3.6. URL https://arxiv.org/abs/1501.04585v4.
  17. [17]Kaisa Matomäki, Maksym Radziwiłł, and Terence Tao. An averaged form of Chowla’s conjecture. Algebra & Number Theory, 9(9):2167–2196, 2015. doi: 10.2140/ant.2015.9.2167. URL https://arxiv.org/abs/1503.05121v3. Appendix A is used in the corrected version, arXiv:1503.05121v3, 1 March 2022.
  18. [18]Colin McDiarmid. On the method of bounded differences. In Johannes Siemons, editor, Surveys in Combinatorics, 1989, volume 141 of London Mathematical Society Lecture Note Series, pages 148–188. Cambridge University Press, Cambridge, 1989. doi: 10.1017/CBO9781107359949.008. URL https://www.cambridge.org/core/books/abs/surveys-in-combinatorics-1989/on-the-method-of-bounded-differences/AABA597B562BDA7D89C6077E302694FB.DOI
  19. [19]E. J. McShane. Extension of range of functions. Bulletin of the American Mathematical Society, 40 (12):837–842, 1934. doi: 10.1090/S0002-9904-1934-05978-0.DOI
  20. [20]H. L. Montgomery and R. C. Vaughan. Exponential sums with multiplicative coefficients. Inventiones Mathematicae, 43(1):69–82, 1977. doi: 10.1007/BF01390204. URL https://link.springer.com/article/10.1007/BF01390204.DOI
  21. [21]Cédric Pilatte. Improved bounds for the two-point logarithmic Chowla conjecture, 2023. URL https://arxiv.org/abs/2310.19357v3. Version 3, 25 August 2026.
  22. [22]V. Ramaswami. On the number of positive integers less than x and free of prime divisors greater than xᶜ. Bulletin of the American Mathematical Society, 55(12):1122–1127, 1949. doi: 10.1090/S0002-9904-1949-09337-0.DOI
  23. [23]Terence Tao. The logarithmically averaged Chowla and Elliott conjectures for two-point correlations. Forum of Mathematics, Pi, 4:e8, 2016. doi: 10.1017/fmp.2016.6. 36 pp.DOI
  24. [24]Terence Tao and Joni Teräväinen. The structure of correlations of multiplicative functions at almost all scales, with applications to the Chowla and Elliott conjectures. Algebra & Number Theory, 13(9):2103–2150, 2019. doi: 10.2140/ant.2019.13.2103. URL https://arxiv.org/abs/1809.02518v2.
  25. [25]Terence Tao and Joni Teräväinen. Quantitative correlations and some problems on prime factors of consecutive integers, 2026. URL https://arxiv.org/abs/2512.01739v2. Version 2, 25 April 2026.
  26. [26]Joni Teräväinen. On binary correlations of multiplicative functions. Forum of Mathematics, Sigma, 6:e10, 2018. doi: 10.1017/fms.2018.10. URL https://arxiv.org/abs/1710.01195v2. 41 pp.
  27. [27]Zhiwei Wang. On the largest prime factors of consecutive integers in short intervals. Proceedings of the American Mathematical Society, 145(8):3211–3220, 2017.DOI
  28. [28]Zhiwei Wang. Sur les plus grands facteurs premiers d’entiers consécutifs. Mathematika, 64(2):343–379, 2018. doi: 10.1112/S0025579317000547. URL https://arxiv.org/abs/1706.02980v1.
  29. [29]Zhiwei Wang. Three conjectures on P⁺(n) and P⁺(n + 1) hold under the Elliott–Halberstam conjecture for friable integers. Journal of Number Theory, 223:1–11, 2021. doi: 10.1016/j.jnt.2020.12.013.DOI
  30. [30]Z hiyuan Yang. An improvement on the largest prime factors of consecutive integers. Preprint, arXiv:2607.16032v1, 2026. URL https://arxiv.org/abs/2607.16032v1. 17 July 2026.

Paper details

Contents