Introduction

The prime factors of a typical integer, viewed on a logarithmic scale, form a random partition of unit mass. We prove that the same limiting partition occurs for the predecessor of a uniformly chosen prime.

For a prime p≥3p \ge3, list the prime factors of p−1p-1, with multiplicity, in decreasing order,

q1(p)≥q2(p)≥⋯q_1(p) \ge q_2(p) \ge\cdots

and put qj(p)=1q_j(p)=1 after the list is exhausted. Define

Vj(p)=log⁡qj(p)log⁡(p−1).V_j(p)=\frac{\log q_j(p)}{\log(p-1)}.

Then Vj(p)≥0V_j(p) \ge0 and ∑jVj(p)=1\sum_j V_j(p)=1.

Let U1,U2,…U_1,U_2,\ldots be independent uniform random variables on (0,1)(0,1), and form the stick fragments

B1=1−U1,Bj=(∏i<jUi)(1−Uj)(j≥2).B_1=1-U_1,\qquad B_j=\left(\prod_{i<j}U_i\right)(1-U_j)\quad(j\ge2).

Their decreasing rearrangement (L1,L2,…)(L_1,L_2,\ldots) has the Poisson–Dirichlet distribution with parameter one, denoted PD⁡(1)\operatorname{PD}(1). Write π(x)\pi(x) for the number of primes at most $x.

Theorem 1.1. For every fixed k≥1k \ge1 and every bounded continuous function F:[0,1]k→RF : [0,1]^k \to\mathbb{R},

lim⁡x→∞1π(x)−1∑3≤p≤xp primeF(V1(p),…,Vk(p))=EF(L1,…,Lk).(1)\lim_{x \to\infty} \frac{1}{\pi(x)-1} \sum_{\substack{3 \le p \le x \\ p\ \mathrm{prime}}} F\bigl(V_1(p),\ldots,V_k(p)\bigr) = \mathbb{E}F(L_1,\ldots,L_k). \tag*{(1)}

The limit holds through all real xx and uses ordinary equal weighting of the primes.

Theorem 1.1 resolves positively the conjecture of Ford, Konyagin and Luca [13], Section 6, Conjecture 5. Their formulation uses the same multiplicity convention, normalization, and ordinary counting distribution on primes. The assertion includes all finite joint distributions, strengthening the largest-factor prediction alone.

History and significance. For ordinary uniformly sampled integers, the largest-factor law begins with Dickman [7] and the smooth-number estimates of de Bruijn [6]. Billingsley proved the joint limiting law of the ordered large prime divisors [5]. Donnelly and Grimmett [8], Theorem 1 and Corollaries 2–3 gave a size-biased treatment, including multiplicities, leading directly to the stick-breaking and Poisson–Dirichlet descriptions. Arratia, Kochman and Miller [2] express this identification through the joint intensities of distinct factor tuples. Our final passage from interior factor statistics to ranked factors uses the same size-biased approach; that passage is proved explicitly in Section 8.

The shifted-prime problem has a different arithmetic difficulty: requiring a product of large primes to divide p−1p-1 places pp in a progression with that product as modulus. Standard distribution estimates give useful information for products below a fixed power of xx, whereas the full factor partition involves products throughout the range below xx. The smoothness question arose in Erdős’s work on Euler’s function [9, 10], and Pomerance [23] developed its connection with large totient fibers. Baker and Harman [3], Theorem 1 and Lichtman [18], Theorem 1.1 obtained lower bounds of the form x/(log⁡x)Cx/(\log x)^C for primes with smooth predecessors, at thresholds x0.2961x^{0.2961} for p≤xp \le x and x0.2844x^{0.2844} for x<p≤2xx < p \le2x, respectively.

For shifted primes, Granville states the conjectural asymptotic [15], Section 5.3, (99)

#{p≤x:P+(p−1)≤x1/u}∼π(x)ρ(u)(u≥1 fixed),(2)\#\{p \le x : P^{+}(p-1) \le x^{1/u}\} \sim\pi(x)\rho(u) \qquad(u \ge1\ \mathrm{fixed}), \tag*{(2)}

where P+(n)P^{+}(n) is the largest prime factor, with P+(1)=1P^{+}(1)=1, and the Dickman function is determined by ρ(u)=1\rho(u)=1 for 0≤u≤10 \le u \le1 and uρ′(u)+ρ(u−1)=0u\rho'(u)+\rho(u-1)=0 for u>1u>1. Theorem 1.1 resolves this fixed-uu conjecture positively and proves the complete joint limiting law. The use of xx rather than p−1p-1 in the threshold in (1.3) causes no change: restrict first to εx<p≤x\varepsilon x < p \le x, where log⁡(p−1)/log⁡x→1\log(p-1)/\log x \to1 uniformly, and then let ε↓0\varepsilon\downarrow0. The largest part of PD⁡(1)\operatorname{PD}(1) has the continuous Dickman distribution, by the ordinary-integer limit in [8], Section 1 and the smooth-number asymptotic in [15], Section 1, (1.1); hence the corresponding threshold sandwich applies.

The closest general distribution theorem is due to Bharadwaj and Rodgers [4], Theorem 7. Their Poisson–Dirichlet theorem applies to nonnegative sequences satisfying their regularity and congruence-uniformity hypotheses, with level of distribution one. For shifted primes this gives the full law under the Elliott–Halberstam conjecture, as anticipated in [13], Section 6. Unconditionally, [4], Proposition 2 and Lemma 8 give level one half and the corresponding factor correlations for tests supported where the sum of the logarithmic coordinates is less than one half. Our extraction reaches every compact subset of the open simplex with coordinate sum less than one, and the final probability argument controls its boundary. All smoothness parameters in (1.3) remain fixed as xx grows.

There is also strong recent information about the small prime divisors of shifted primes. Ford [12] proves a total-variation approximation for their valuations on subpower ranges; Gorodetsky [14] obtains such approximations within a general sieve model. The large-factor law studied here concerns the logarithmic mass carried by prime factors of polynomial size.

Ford, Konyagin and Luca formulated their conjecture in studying prime chains and Pratt trees. Theorem 1.1 supplies the factorization law at the root of their probabilistic model; the recursive independence assumptions used for later generations are additional hypotheses. Smooth predecessors also underlie the Carmichael-number construction of Alford, Granville and Pomerance [1], together with a separate progression-distribution input.

The companion article Weighted dilation graphs, smooth shifted primes and totient fibers [21], Theorem 1.2, denoted S, proves that, for every fixed positive smoothness exponent, there are x1−o(1)x^{1-o(1)} primes in its stated interval with smooth predecessors. Such a lower bound does not imply a positive limiting proportion, much less (1.2). We use S’s dilation-graph operator theorems in a different extraction argument. Section 6 states their precise contracts and verifies the new endpoint hypotheses. The long-prime polynomial and logarithmic-phase estimates needed in both analytic branches are proved in Appendix A; Section 2 records the ranges in which we use them.

The new analytic step. The main new estimate is Theorem 3.1: a marked Type II estimate in which only the predecessor side carries divisor marks, namely weights recording selected small prime divisors. It permits a prime factor on the other side to be replaced by a rough integer, free of primes below a prescribed cutoff, with its local density. Cauchy’s inequality, applied after dividing out one tuple of marks, leads to a small-determinant relation between primitive lattice vectors. We average its long signed moments using a root-residue formula and an exact memory expansion. The memory records a prime across gaps between its uses, so its divisibility probability is charged only once.

The root coordinates are restricted by quantitative separation conditions. These give both a balanced lattice box for fresh-label minor arcs and a separation rule for retrieving remembered labels. Global distinctness is then restored by grouping equalities between newly drawn prime labels according to their rank, the number of independent equality constraints. High ranks supply many small point-probability factors, whereas low ranks alter only a vanishing fraction of the signed edge contractions. The resulting estimate has arbitrarily strong fixed logarithmic precision, with a number of marking bands independent of their particular fixed exponents.

The signed-moment method has precedents in divisibility graphs. Helfgott and Radziwiłł center prime-divisibility weights by subtracting 1/p1/p and analyze high traces through prime-labelled walks [16], Sections 1.3, 1.5 and 5. Their proposed composite shifts [16], Section 9.1 are developed by Pilatte using products of primes and a non-backtracking trace estimate [22], Definition 3.2 and Proposition 5.3. Recurring prime labels create dependencies in these walk expansions. Here the state contains active prime lists and a primitive lattice vector. At each prime pp, root averaging gives a uniform projective line, so a specified line has probability 1/(p+1)1/(p+1). The memory preserves that line across gaps between active runs. Sections 3–4 develop this determinant geometry and prove the exact memory identity; Section 5 supplies the complementary major-term estimate.

From marked correlations to the full law. The second estimate, Theorem 6.1, uses marks on both endpoints. We apply the general graph inputs of S to a centered long-prime divisor statistic and prove the required new endpoint Fourier estimates. A small mark restricts difficult frequencies to a sparse set; a long-prime polynomial controls the positive tuple term there, and a finite divisor model controls the constant term.

A first-moment presieve and the one-sided estimate replace all prime slots of the composite terms by rough slots. Successful marking groups then bring these composite terms within the two-sided estimate. Averaging over disjoint fixed marking arrays removes the marks from the prime average. This proves interior factorial-moment asymptotics for ordinary prime dyads.

Finally, sequential size-biased sampling turns the interior factorial measures into probability densities of total mass one. This excludes escape to the boundary and yields the stick fragments in (1.1). The expected undrawn mass controls sorting, and a finite dyadic decomposition gives all real upper endpoints. We provide this probability argument explicitly, since restricted-support moment formulas alone are not the statement of Theorem 1.1. Figure 1 summarizes the dependence of these steps. The two-sided endpoint analysis is in Section 6, the prime extraction and removal of marks in Section 7, and the probability argument in Section 8.

Flowchart of the proof's analytic inputs and probabilistic argument

Figure 1. The proof has two analytic inputs. The one-sided estimate replaces prime slots in composite terms by rough slots. Successful marking brings the resulting centered statistic under the two-sided estimate. Averaging the fixed marks then gives interior moments; the probability argument recovers the full law. Appendix A supplies the long-prime and logarithmic-phase estimates used in the analytic branches.

Conventions and analytic inputs

Throughout the proof xx is a real parameter tending to infinity, and

L=log⁡x,W=exp⁡(L0.24),V(W)=∏p≤W(1−1p),e(t)=exp⁡(2πit).L = \log x,\qquad W = \exp(L^{0.24}),\qquad V(W) = \prod_{p \le W}\left(1-\frac{1}{p}\right),\qquad e(t) = \exp(2\pi i t).

A positive integer is rough if none of its prime factors is at most WW; we set P−(1)=∞P^{-}(1)=\infty. Prime variables in explicitly indicated prime sums range only over primes. The functions τ(n)\tau(n) and τj(n)\tau_j(n) count divisors and ordered factorizations into jj positive integers, respectively. Dirichlet characters are zero on nonunits. We use n≍Hn \asymp H for a fixed bounded enlargement of [H,2H][H,2H].

Every unspecified exponent of LL is fixed before xx tends to infinity. Constants may depend on previously fixed data. Whenever an exponent must be independent of a marking parameter, this is stated explicitly. The number of determinant pads in Sections 3–5 grows slowly with LL; the padding parameter in Section 6 is instead a sufficiently large fixed integer. These are separate constructions.

We use the dilation-graph and ideal-kernel theorems of the companion Weighted dilation graphs, smooth shifted primes and totient fibers [21], denoted S. Its complete text accompanies this article. Section 6 states the exact operator contracts and verifies their endpoint hypotheses. The long-prime and logarithmic-phase estimates stated below are proved in Appendix A. Our presieving cutoff is always the WW defined above; it is not S’s cutoff exp⁡(log⁡x)\exp(\sqrt{\log x}).

Prime estimates and mean squares

Lemma 2.1 (Prime estimates). The prime number theorem holds with an error smaller than every fixed negative power of the logarithm. In particular, as y→∞y \to\infty,

∑p≤y1p=log⁡log⁡y+cM+o(1),∏p≤y(1−1p)∼e−γElog⁡y.\sum_{p \le y} \frac{1}{p} = \log\log y + c_M + o(1), \qquad\prod_{p \le y} \left(1-\frac{1}{p}\right) \sim\frac{e^{-\gamma_E}}{\log y}.

where cMc_M and γE\gamma_E are constants. For fixed A,C>0A,C>0, uniformly for r≤(log⁡y)Cr \le(\log y)^C and (a,r)=1(a,r)=1,

#{p≤y:p≡a(modr)}=Li⁡(y)φ(r)+OA,C(y(log⁡y)A).\#\{p \le y : p \equiv a \pmod r\} = \frac{\operatorname{Li}(y)}{\varphi(r)} + O_{A,C}\left(\frac{y}{(\log y)^A}\right).

The constants need not be effective.

These classical estimates follow from the quantitative prime number theorem, Siegel–Walfisz, and Mertens’ theorems; see [27], Corollary 39 and Exercises 40, 64 and [26], Theorems 15, 26. The harmonic asymptotic also follows by partial summation of the prime number theorem. Thus, for every fixed a>0a>0,

∑exp⁡(La)≤p≤exp⁡(2La)1p=log⁡2+o(1).(3)\sum_{\exp(La)\le p\le\exp(2La)} \frac{1}{p} = \log2 + o(1). \tag*{(3)}

We use prime asymptotics only on fixed positive-power ranges or on intervals for which taking differences of the stated estimates gives the required absolute error.

For the classical, stronger coefficient-independent mean-value theorem, see [20], Theorem 2 and Corollary 3. We prove the weaker bounds needed here directly.

Lemma 2.2 (Divisor moments and Dirichlet mean squares). For every fixed nonnegative integer kk there is CkC_k such that

∑n≤Yτ(n)k≪kY(log⁡(2Y))Ck,∑n≤Yτ(n)n≪k(log⁡(2Y))Ck.\sum_{n\le Y} \tau(n)^k \ll_k Y(\log(2Y))^{C_k}, \qquad\sum_{n\le Y} \frac{\tau(n)}{n} \ll_k (\log(2Y))^{C_k}.

Let P(t)=∑n≤YcnnitP(t)=\sum_{n\le Y}c_n n^{it}. On every real interval II of length T≥0T\ge0,

∫I∣P(t)∣2 dt≪(T+Ylog⁡(2Y))∑n∣cn∣2,(4)\int_I |P(t)|^2\,dt \ll(T+Y\log(2Y))\sum_n |c_n|^2, \tag*{(4)}
∑t∈T∣P(t)∣2≪(T+1+Ylog⁡(2Y))(log⁡(2Y))2∑n∣cn∣2,(5)\sum_{t\in\mathcal{T}} |P(t)|^2 \ll(T+1+Y\log(2Y))(\log(2Y))^2\sum_n |c_n|^2, \tag*{(5)}

whenever T⊂I\mathcal{T}\subset I has mutual spacing at least one. The constants in these two inequalities are absolute and do not depend on how the coefficients were formed.

Proof. Unique factorization and positivity bound the harmonic divisor sum by

∏p≤Y∑a≥0(a+1)kpa=∏p≤Y(1+2kp+Ok(p−2))≪k(log⁡(2Y))Ck.\prod_{p \le Y}\sum_{a \ge0}\frac{(a+1)^k}{p^a} =\prod_{p \le Y}\left(1+\frac{2^k}{p}+O_k(p^{-2})\right)\ll_k(\log(2Y))^{C_k}.

Multiplication by YY gives the counting bound. For the continuous mean square, expand the square and integrate. The diagonal is T∑n∣cn∣2T\sum_n|c_n|^2. For m≠nm\ne n, the integral has absolute value at most 2/∣log⁡(n/m)∣≪Y/∣n−m∣2/|\log(n/m)|\ll Y/|n-m|. The inequality 2∣cncm∣≤∣cn∣2+∣cm∣22|c_nc_m|\le|c_n|^2+|c_m|^2 and the harmonic sum over ∣n−m∣|n-m| give (2.2). On the unit interval centered at each point of TT, the one-dimensional Sobolev inequality bounds ∣P(t)∣2|P(t)|^2 by a constant times the integral of ∣P∣2+∣P′∣2|P|^2+|P'|^2. These intervals have bounded overlap and lie in an interval of length T+1T+1. Apply the continuous estimate to PP and P′P', whose coefficients are i(log⁡n)cni(\log n)c_n, to obtain (2.3). □

The coefficient-independent formulation in (2.2)–(2.3) is important: later a small-prime polynomial is raised to a growing power, whose coefficient norm is bounded separately by a factorial estimate. Fixed divisor moments alone are not applied at that growing order.

Prime polynomials and logarithmic phases

Lemma 2.3 (Long prime polynomial). Fix 0<τ<η<10<\tau<\eta<1 and C,A>0C,A>0. There is B0=B0(τ,η,C,A)B_0=B_0(\tau,\eta,C,A) such that, uniformly for

xτ2≤N≤xη,q≤LC,LB0≤∣t∣≤x2,\frac{x^\tau}{2}\le N\le x^\eta,\qquad q\le L^C,\qquad L^{B_0}\le|t|\le x^2,

every Dirichlet character χ\chi modulo qq and every interval I⊂[N,2N]I\subset[N,2N] satisfy

∣∑p∈Iχ(p)p−1+it∣≪τ,η,C,AL−A.(6)\left|\sum_{p\in I}\chi(p)p^{-1+it}\right|\ll_{\tau,\eta,C,A}L^{-A}. \tag*{(6)}

The complete proof is Lemma A.1 in Appendix A, with exactly the displayed ranges. Partial summation converts the reciprocal weight in (2.4) to normalization by NN on a dyadic interval.

Lemma 2.4 (Logarithmic phases on progressions). Fix C>0C>0. There is an absolute constant C3>0C_3>0 such that, for sufficiently large xx in terms of CC, the following holds uniformly:

exp⁡(L(log⁡L)2)≤N≤2x5,q≤LC,exp⁡(L2(log⁡L)2)≤∣t∣≤4x3.\exp\left(\frac{L}{(\log L)^2}\right)\le N\le2x^5,\qquad q\le L^C,\qquad\exp\left(\frac{L}{2(\log L)^2}\right)\le|t|\le4x^3.

For every residue a(modq)a\pmod q and every interval I⊂[N,2N]I\subset[N,2N],

∣∑n∈In≡a(modq)nit∣≪Nqexp⁡(−L(log⁡L)C3).(7)\left|\sum_{\substack{n\in I\\ n\equiv a\pmod q}}n^{it}\right| \ll\frac{N}{q}\exp\left(-\frac{L}{(\log L)^{C_3}}\right). \tag*{(7)}

The proof is Lemma A.5 in Appendix A. The estimate applies to arbitrary residue classes, not only units. At lower frequencies we use the elementary sum–integral comparison on a progression. For N≥xδN\ge x^\delta, q≤LCq\le L^C and 1≤∣t∣≤exp⁡(L/(log⁡L)2)1\le|t|\le\exp(L/(\log L)^2), it gives

qN∣∑n∈In≡a(modq)nit∣≪1∣t∣+x−cδ.(8)\frac{q}{N}\left|\sum_{\substack{n\in I\\ n\equiv a\pmod q}}n^{it}\right| \ll\frac{1}{|t|}+x^{-c_\delta}. \tag*{(8)}

after decreasing cδ>0c_\delta> 0 if necessary. Indeed the integral is O(N/(q∣t∣))O(N/(q|t|)), and the endpoint/total-variation error is O(1+∣t∣)O(1+|t|). The same estimate holds after partial summation against a smooth factor with fixed logarithmic derivative costs, with those costs included in the implied logarithmic power.

Elementary rough counts

Lemma 2.5 (Rough integers in long intervals). Fix δ>0\delta> 0, C>0C > 0, 0<a<10 < a < 1 and A>0A > 0. Suppose xδ≤H≤xCx^\delta\le H \le x^C and I⊂[H,2H]I \subset[H,2H] is an interval. Then

#{n∈I:P−(n)>W}=V(W)∣I∣+O(HL−A).(9)\#\{n \in I : P^{-}(n) > W\} = V(W)|I| + O(HL^{-A}). \tag*{(9)}

If d≤exp⁡(La)d \le\exp(L^a) and every prime divisor of dd exceeds WW, then for every residue bb (mod dd),

#{n∈I:n≡b(modd), P−(n)>W}=V(W)∣I∣d+O(HdL−A).(10)\#\{n \in I : n \equiv b \pmod d,\ P^{-}(n) > W\} = \frac{V(W)|I|}{d} + O\left(\frac{H}{d}L^{-A}\right). \tag*{(10)}

For every character χ\chi modulo q≤LCq \le L^C*,

∑n∈IP−(n)>Wχ(n)=1χ principalV(W)∣I∣+O(HL−A).(11)\sum_{\substack{n \in I \\ P^{-}(n)>W}} \chi(n) = \mathbf{1}_{\chi\ \mathrm{principal}}V(W)|I| + O(HL^{-A}). \tag*{(11)}

All statements are uniform in the indicated intervals and residues.

Proof. Use consecutive even and odd Bonferroni truncations for the events p∣np \mid n, p≤Wp \le W, at depths hh and h+1h+1, with h≍CAlog⁡Lh \asymp C_A \log L. Their divisors are at most

Wh+1=exp⁡(OA(L0.24log⁡L))=xo(1).W^{h+1} = \exp\left(O_A(L^{0.24}\log L)\right) = x^{o(1)}.

The number of divisor terms and their total absolute integer coefficient mass are also xo(1)x^{o(1)}. Counting each progression by its length divided by the modulus costs at most one per term. Writing SW=∑p≤W1/p=O(log⁡L)S_W = \sum_{p\le W}1/p = O(\log L), the difference of the two model bounds is at most

SWh+1(h+1)!,\frac{S_W^{h+1}}{(h+1)!},

which is smaller than any prescribed power of L−1L^{-1} on choosing the fixed constant CAC_A sufficiently large. The full model product is V(W)V(W).

For (2.8), each sieve divisor is coprime to dd, so the Chinese remainder theorem gives main term ∣I∣/(de)|I|/(de) for the divisor ee. The same omitted-degree bound is multiplied by H/dH/d. The total endpoint error remains xo(1)x^{o(1)}, which is smaller than (H/d)L−A(H/d)L^{-A}, since d=xo(1)d=x^{o(1)} and H≥xδH \ge x^\delta. This proves both counting formulas.

For (2.9), first count a fixed unit residue class modulo qq. Primes dividing qq are automatically absent; the model density in that progression is

1q∏p≤Wp∤q(1−1p)=V(W)φ(q).\frac{1}{q}\prod_{\substack{p\le W\\p\nmid q}}\left(1-\frac{1}{p}\right) = \frac{V(W)}{\varphi(q)}.

The endpoint and Bonferroni errors have arbitrary logarithmic precision, uniformly for q≤LCq \le L^C. Sum with the character over the unit residue classes, choosing the precision first to absorb their number. Character orthogonality gives the displayed answer.

A sieve for nonnegative weighted objects

Lemma 2.6 (Block sieve). Consider finitely many objects with nonnegative weights and a bad condition at each of some designated primes p≤zp \le z. Suppose that the total weight on which all conditions indexed by the primes dividing a squarefree dd hold is

Xg(d)+r(d),Xg(d) + r(d),

including d=1d = 1, where X≥0X \ge0, gg is multiplicative, and fixed η0>0\eta_0 > 0, C0C_0 satisfy

0≤g(p)≤1−η0,∑w<p≤w2g(p)≤C0(w>1).0 \le g(p) \le1 - \eta_0,\qquad\sum_{w<p\le w^2} g(p) \le C_0 \quad(w > 1).

For every sufficiently large even integer hh, the weight avoiding all designated bad conditions is

X∏p≤z(1−g(p))(1+O(e−h))+O(∑d≤z4h+2d squarefree∣r(d)∣).(12)X\prod_{p\le z}(1-g(p))(1+O(e^{-h}))+ O\left(\sum_{\substack{d\le z^{4h+2}\\d\ \mathrm{squarefree}}}|r(d)|\right). \tag*{(12)}

Products and sums use only designated primes. Constants depend only on the density bounds and are uniform in hh. For h=2h = 2, an upper bound holds with a fixed constant times the main term and the same remainder sum.

This is the block sieve of S [21], Lemma 2.9]. Its upper and lower polynomials have coefficients of absolute value at most one, supported on squarefree d≤z4h+2d \le z^{4h+2}; the upper polynomial is nonnegative. We will supply all needed remainder estimates for our particular weights. No assertion detecting primes from congruence information alone is included in this input.

A determinant estimate with marks on one side

The first correlation estimate removes prime conditions from factors of the unmarked endpoint. Throughout this section and the next two sections, put

L=log⁡x,W=exp⁡(L0.24),V(W)=∏p≤W(1−p−1).(13)L=\log x,\qquad W=\exp(L^{0.24}),\qquad V(W)=\prod_{p\le W}(1-p^{-1}). \tag*{(13)}

An integer is rough if it has no prime factor at most WW. Fix disjoint groups consisting of all the primes in the bands

Pi={p:exp⁡(Lai)≤p≤exp⁡(2Lai)},0.1<a1<⋯<aK<0.2.(14)\mathcal{P}_i=\{p:\exp(L^{a_i})\le p\le\exp(2L^{a_i})\},\qquad0.1<a_1<\cdots<a_K<0.2. \tag*{(14)}

Here KK and the aia_i are fixed before xx tends to infinity. Write

Vi=∑p∈Pi1p,μi(p)=1pVi,V_i=\sum_{p\in\mathcal{P}_i}\frac{1}{p},\qquad\mu_i(p)=\frac{1}{pV_i},
ωi(h)=∑p∈Pi1p∣h,ω(h)=∑i=1Kωi(h).(15)\omega_i(h)=\sum_{p\in\mathcal{P}_i}\mathbf{1}_{p\mid h},\qquad\omega(h)=\sum_{i=1}^{K}\omega_i(h). \tag*{(15)}

Lemma 2.1 and partial summation give Vi=log⁡2+o(1)V_i=\log2+o(1). For a fixed q∈(0,1)q\in(0,1), the marked weight is

W(h)=qω(h)−K∏i=1Kωi(h)Vi.(16)\mathcal{W}(h)=q^{\omega(h)-K}\prod_{i=1}^{K}\frac{\omega_i(h)}{V_i}. \tag*{(16)}

The weight is zero if some group supplies no divisor, and is bounded by a constant depending only on K,qK,q. Indeed, qj−1jq^{j-1}j is bounded for positive integers jj.

Theorem 3.1 (One-sided marked Type II estimate). Fix δ,C,D∗>0\delta,C,D_*>0 and q∈(0,1)q\in(0,1). Suppose Hm,Hn≥xδH_m,H_n\ge x^\delta and X=HmHn≍xX=H_mH_n\asymp x. Let FF satisfy ∣F(h)∣≤1|F(h)|\le1 and F(ph)=F(h)F(ph)=F(h) for every prime in the groups (3.2). Let αm\alpha_m be supported on an arbitrary interval in [Hm,2Hm][H_m,2H_m], where it has the value

αm=miv(1m prime−1m roughV(W)log⁡m),∣v∣≤LC.\alpha_m=m^{iv}\left(1_{m\ {\rm prime}}-\frac{1_{m\ {\rm rough}}}{V(W)\log m}\right),\qquad|v|\le L^C.

Let βn\beta_n be supported on rough integers in [Hn,2Hn][H_n,2H_n], with ∣βn∣≤LC|\beta_n|\le L^C. If KK is sufficiently large in terms of δ,C,D∗,q\delta,C,D_*,q, then

∣∑m,nαmβnF(mn−1)W(mn−1)∣≪XL−D∗.(17)\left|\sum_{m,n}\alpha_m\beta_nF(mn-1)\mathcal{W}(mn-1)\right|\ll XL^{-D_*}. \tag*{(17)}

The required lower bound on KK is independent of the particular aia_i. The implicit constant and the threshold for xx may depend on all the fixed parameters, including the aia_i.

We first give the reduction and geometric construction. The signed moment estimate is proved in Proposition 4.1; Proposition 5.1 estimates the resulting major term and completes the proof of the theorem. Exponents denoted OC(1)O_C(1) below can be chosen independently of KK. A factor depending on a fixed KK is harmless, but a loss LcKL^{cK} would not be harmless at the final choice of KK; we keep this distinction explicit.

Cauchy's inequality and the normalization of the square

Expand the product of the ωi\omega_i by selecting one prime from each group, and denote their product by b0b_0. For mn−1=b0hmn-1=b_0h, invariance gives F(mn−1)=F(h)F(mn-1)=F(h). Except when some selected prime has square dividing mn−1mn-1, one also has ω(h)=ω(mn−1)−K\omega(h)=\omega(mn-1)-K. The replacement of the original damping by qω(h)q^{\omega(h)} costs

OA(XL−A)for every fixed A>0.O_A(XL^{-A})\qquad\text{for every fixed }A>0.

To verify this, the sum over labels of either nonnegative weight is OK,q(1)O_{K,q}(1): removing KK distinct prime factors can lower ω\omega by at most KK, so the new weight is at most the old one. For fixed mm and a group prime pp, the congruence mn≡1(modp2)mn\equiv1\pmod{p^2} either has no solution or specifies one residue class of nn. Since Hn≥xδ≫p2H_n\ge x^\delta\gg p^2, its count is O(Hn/p2)O(H_n/p^2). Finally ∑i,p∈Pip−2≪Kexp⁡(−L0.1)\sum_{i,p\in\mathcal{P}_i}p^{-2}\ll_K\exp(-L^{0.1}). The coefficient bounds absorb only a fixed power of LL, proving (3.6).

Choose a real smooth dyadic partition with factors η(b0/Y)\eta(b_0/Y), where η\eta is supported in [1,4][1,4] and has bounded derivatives. There are OK(L)O_K(L) relevant YY, and

cL0.1≤log⁡Y≤CKL0.2.(18)cL^{0.1}\le\log Y\le C_KL^{0.2}. \tag*{(18)}

For a fixed factor of this partition the sum becomes

∑h≪X/YF(h)qω(h)∑mn−1=b0hb0=∏ipi, pi∈Pi(∏iVi−1)η(b0/Y)αmβn.\sum_{h\ll X/Y}F(h)q^{\omega(h)} \sum_{\substack{mn-1=b_0h\\ b_0=\prod_i p_i,\ p_i\in\mathcal{P}_i}} \left(\prod_iV_i^{-1}\right)\eta(b_0/Y)\alpha_m\beta_n.

Cauchy's inequality bounds its square by O(X/Y)O(X/Y) times the nonnegative expanded square

QY=(∏iVi−2)∑a,b,m,n,r,smn−1=ah, rs−1=bh for some hη(a/Y)η(b/Y)αmαr‾βnβs‾.(19)Q_Y=\left(\prod_iV_i^{-2}\right) \sum_{\substack{a,b,m,n,r,s\\ mn-1=ah,\ rs-1=bh\ {\rm for\ some}\ h}} \eta(a/Y)\eta(b/Y)\alpha_m\overline{\alpha_r}\beta_n\overline{\beta_s}. \tag*{(19)}

Both aa and bb select one prime per group. In particular, proving QY≪XYL−D1Q_Y \ll XYL^{-D_1} for arbitrarily large fixed D1D_1 is sufficient; the required D1D_1 depends on C,D∗C,D_\ast but not on KK.

We may first discard pairs a,ba,b sharing a label. The elementary bound used here, uniform for l≍Yl \asymp Y, is

∑h≪X/Yτ(1+lh)2≪(X/Y)L3.\sum_{h\ll X/Y}\tau(1+lh)^2\ll(X/Y)L^3.

Indeed, τ(n)2≤τ4(n)\tau(n)^2\leq\tau_4(n). In an ordered four-factor decomposition of an integer at most O(X)O(X), deleting a largest factor leaves a product d≪X3/4d\ll X^{3/4}. There are at most 4τ3(d)4\tau_3(d) choices for the three retained factors and their positions. If d∣1+lhd\mid1+lh, then (d,l)=1(d,l)=1, and hh lies in one progression modulo dd. Consequently the left side of (3.9) is at most a constant times

∑d≪X3/4τ3(d)(XYd+1)≪XYL3+X3/4L2≪XYL3.\sum_{d\ll X^{3/4}}\tau_3(d)\left(\frac{X}{Yd}+1\right)\ll\frac{X}{Y}L^3+X^{3/4}L^2\ll\frac{X}{Y}L^3.

These divisor estimates also follow from Lemma 2.2. For fixed a,ba,b, Cauchy’s inequality and (3.9) bound the absolute hh-sum of the two representation counts by (X/Y)LOC(1)(X/Y)L^{O_C(1)}. The number of pairs sharing a prime is

≪KY2∑i,p∈Pip−2≪KY2exp⁡(−L0.1);\ll K Y^2\sum_{i,p\in\mathcal P_i}p^{-2}\ll_K Y^2\exp(-L^{0.1});

here the number of integers in [Y,4Y][Y,4Y] divisible by pp is O(Y/p)O(Y/p). Thus the discarded part is OA(XYL−A)O_A(XYL^{-A}).

For the remaining pairs (a,b)=1(a,b)=1. Their two equations in (3.8) are equivalent to

t=bmn−ars=b−a.t=bmn-ars=b-a.

In fact, the displayed equality gives a∣mn−1a\mid mn-1 and b∣rs−1b\mid rs-1, and the quotients agree and are positive.

Fix a real compactly supported smooth ψ\psi, equal to one on [−4,4][-4,4]. For a fixed exponent A0>0A_0>0, let M\mathfrak M be the union, on R/Z\mathbb R/\mathbb Z, of the arcs

∣θ−hk∣≤2LA0Y,1≤k≤LA0,(h,k)=1.\left|\theta-\frac{h}{k}\right|\leq\frac{2L^{A_0}}{Y},\qquad1\leq k\leq L^{A_0},\qquad(h,k)=1.

They are disjoint for large xx. The replacement for (3.10) is

HM(t;a,b)=ψ(t/Y)∫Me(θ(t−b+a)) dθ.(20)\mathcal H_{\mathfrak M}(t;a,b)=\psi(t/Y)\int_{\mathfrak M}e(\theta(t-b+a))\,d\theta. \tag*{(20)}

In the ultimate major sum a,ba,b need not be disjoint. The signed error in this replacement is controlled by the lift constructed next.

Physical states, good positions, and endpoint vectors

Use the growing parameters

J=⌊L0.01⌋,M=J+1,R=⌈L0.5/2⌉,N=2R.J=\lfloor L^{0.01}\rfloor,\qquad M=J+1,\qquad R=\lceil L^{0.5}/2\rceil,\qquad N=2R.

An auxiliary list has JJ primes per group; its product is denoted by DD. Partition all possible DD into ordinary dyads [d0,2d0)[d_0,2d_0). There are OK(L)O_K(L) such dyads and log⁡(2d0)≤CKL0.21\log(2d_0)\leq C_KL^{0.21}. Fix one and put

H=d0Y,U=HHm,V=Hn,Ω=[U,16U]×[V,2V].H=d_0Y,\qquad U=HH_m,\qquad V=H_n,\qquad\Omega=[U,16U]\times[V,2V].

A physical state is a primitive integer vector P=(P1,P2)∈ΩP=(P_1,P_2)\in\Omega together with an ordered list pi=(pi,1,…,pi,M)\mathfrak{p}_i=(p_{i,1},\ldots,p_{i,M}) of distinct primes from Pi\mathcal{P}_i for each ii, all dividing P1P_1. Let H\mathcal{H} be the Hilbert space on these states with measure

dσ(P,p)=∏i=1KVi−Mtimes counting measure.(21)\mathrm{d}\sigma(P,\mathfrak{p})=\prod_{i=1}^{K}V_i^{-M}\quad\text{times counting measure.} \tag*{(21)}

Averaging independently over permutations of each list defines an orthogonal projection S\mathcal{S} on H\mathcal{H}.

We now specify the row operation before symmetrization. Copy slots 1,…,J1,\ldots,J in every group from source to target; their product is DD. The product of the source’s last slots is bb. The target’s last slots are new labels with product aa. Each new slot is summed with weight Vi−1V_i^{-1}; target positions are summed with counting measure. The labels shared across the edge are exactly the copied slots. Thus all unshared labels are distinct from one another and from every shared label, including across the two endpoints. This restriction will be called the cross-edge ban. Require D∈[d0,2d0)D\in[d_0,2d_0) and t=det⁡(P,Q)/D∈Zt=\det(P,Q)/D\in\mathbb{Z}, and use the multiplier

K(P,Q;D,a,b)=d0Dη(b/Y)η(a/Y)ψ(t/Y){1t=b−a−∫Re(θ(t−b+a)) dθ}q(ω(P1)−KM)/2+(ω(Q1)−KM)/2.(22)\mathcal{K}(P,Q;D,a,b)=\frac{d_0}{D}\eta(b/Y)\eta(a/Y)\psi(t/Y) \left\{\mathbf{1}_{t=b-a}-\int_{\mathbb{R}}e(\theta(t-b+a))\,\mathrm{d}\theta\right\} q^{(\omega(P_1)-KM)/2+(\omega(Q_1)-KM)/2}. \tag*{(22)}

Let T\mathcal{T} be this row operation. All sums are finite. The adjoint reverses the row rule and conjugates its multiplier. Indeed, the joint measure of the two lists in an edge pairing is ∏iVi−(M+1)\prod_i V_i^{-(M+1)} in either direction: one factor Vi−MV_i^{-M} comes from (3.14), and one Vi−1V_i^{-1} from the new slot.

After slot symmetrization, the copied products DD range over all choices obtained by omitting one prime from each full list and lying in the chosen pad dyad [d0,2d0)[d_0,2d_0). Each of the MKM^K omission tuples has weight M−KM^{-K} under the source symmetrization.

We next restrict the positions at which this row operation is used. The first restriction will keep the lattice determined by the shared labels from becoming too thin. The second separates phases associated with the omission choices; in the repeated-prime argument it will leave at most one compatible omission tuple.

For a primitive PP, complete it to an integral matrix of determinant one, and let rPr_P be the first coordinate of the second column divided by P1P_1, considered modulo one. Changing the completion adds an integer, so rPr_P is well defined. A position with its full unordered lists is good if the following hold.

(i) For every D∈[d0,2d0)D\in[d_0,2d_0) obtained by omitting one label per group, there is no integer 1≤l≤Y0.21\le l\le Y^{0.2} such that ∥lDrP∥R/Z≤Y−0.7\lVert lDr_P\rVert_{\mathbb{R}/\mathbb{Z}}\le Y^{-0.7}.

(ii) Let II be any nonempty subset of the groups and i∗=max⁡Ii_*=\max I. Draw one fresh prime in each group outside II, independently according to μi\mu_i, and let TfT_f be their product. Except for a set of these draws of probability at most exp⁡(−14Lai∗)\exp(-\frac{1}{4}L^{a_{i_*}}), the points DZTfrPDZT_fr_P corresponding to all distinct integers DZDZ are pairwise separated in circle distance by 100exp⁡(−Lai∗)100\exp(-L^{a_{i_*}}). Here DD ranges over the omission products just described, and ZZ selects one prime from each group in I∖{i∗}I\setminus\{i_*\}.

Let G\mathcal{G} be the projection onto good states. The definition is symmetric in the ordered slots, so G\mathcal{G} commutes with S\mathcal{S}. We use the restricted, symmetrized operator

A=GSTSG.\mathcal{A}=\mathcal{G}\mathcal{S}\mathcal{T}\mathcal{S}\mathcal{G}.

The first test is used in the lattice construction leading to (61); the second is used to prove the unique omission choice in the argument following (70).

Lemma 3.2. For any fixed lists, the set of r∈R/Zr \in\mathbb{R}/\mathbb{Z} failing goodness has Lebesgue measure at most exp⁡(−cL0.1)\exp(-cL^{0.1}), for some c>0c > 0 and all sufficiently large xx.

Proof. There are at most MKM^K omission products. The union bound for the first test is at most 2MKY−0.52M^K Y^{-0.5}, which has the required size by (18). For the second test fix I,TfI,T_f. The number of trial integers DZDZ is

MK∏i∈I∖{i∗}∣Pi∣=exp⁡(o(Lai∗)).M^K \prod_{i\in I\setminus\{i_*\}} |\mathcal{P}_i|=\exp(o(L^{a_{i_*}})).

For distinct DZ,D′Z′DZ,D'Z', the integer (DZ−D′Z′)Tf(DZ-D'Z')T_f is nonzero. Multiplication by this integer preserves uniform measure on the circle. The measure on which this pair violates the required separation is therefore at most 200exp⁡(−Lai∗)200\exp(-L^{a_{i_*}}). Summing over pairs and then averaging over TfT_f gives exp⁡(−Lai∗+o(Lai∗))\exp(-L^{a_{i_*}}+o(L^{a_{i_*}})). Markov’s inequality at the threshold exp⁡(−14Lai∗)\exp(-\frac{1}{4}L^{a_{i_*}}) gives exp⁡(−34Lai∗+o(Lai∗))\exp(-\frac{3}{4}L^{a_{i_*}}+o(L^{a_{i_*}})) for the exceptional set of rr. There are only 2K−12^K-1 possible II.

We record precisely the vectors that recover the Cauchy square. At an endpoint require the complete part of P1P_1 supported on the group primes to be squarefree, with exactly MM primes from each group. Denote that part by G(P1)G(P_1). Put

f(P)=αP1/G(P1)‾βP2‾,g(P)=αP1/G(P1)‾βP2‾,(23)f(P)=\overline{\alpha_{P_1/G(P_1)}}\overline{\beta_{P_2}},\qquad g(P)=\overline{\alpha_{P_1/G(P_1)}}\overline{\beta_{P_2}}, \tag*{(23)}

and set these vectors to zero when the stated conditions or the coefficient supports fail. The vectors are constant on lists. In an edge contributing to their pairing, the coordinates are

P=(Dbm,s),Q=(Dar,n),det⁡(P,Q)/D=bmn−ars.P=(Dbm,s),\qquad Q=(Dar,n),\qquad\det(P,Q)/D=bmn-ars.

No new restriction is imposed on the original coefficients: every m,r,n,sm,r,n,s with nonzero coefficient is rough, and every group prime is below WW. The damping in (22) equals one at these endpoints. Their norms have the uniform bounds

∥f∥,∥g∥≪(UV)1/2LOC(1),sup⁡∣f∣+sup⁡∣g∣≪LOC(1).\|f\|,\|g\|\ll(UV)^{1/2}L^{O_C(1)},\qquad\sup|f|+\sup|g|\ll L^{O_C(1)}.

For example, at a fixed ordered list with product GG, the number of possible positions is O(UV/G)O(UV/G), since G=xo(1)≪UG=x^{o(1)}\ll U. Summing its reciprocal product with the state normalization gives

∏iVi−M∑pi,1,…,pi,M∈Pidistinct in each group∏i,j1pi,j≤1.\prod_i V_i^{-M} \sum_{\substack{p_{i,1},\ldots,p_{i,M}\in\mathcal{P}_i\\\text{distinct in each group}}} \prod_{i,j}\frac{1}{p_{i,j}}\le1.

This proves (3.18) without a logarithmic exponent depending on KK or JJ.

Conversely, positions in (3.17) on the support of (22) are automatically primitive. A prime dividing both DbmDbm and ss cannot divide DbDb, by roughness. It would therefore divide m,sm,s and be larger than WW. It then divides t=bmn−arst=bmn-ars. The cutoff gives ∣t∣≪Y<W|t|\ll Y<W, so t=0t=0. But bmn=arsbmn=ars is impossible: any prime in bb divides none of a,r,sa,r,s. The same proof applies to QQ. The ranges in (3.17) lie in (3.13) because D∈[d0,2d0]D\in[d_0,2d_0] and a,b∈[Y,4Y]a,b\in[Y,4Y].

From the moment to endpoint pairings

For a primitive P∈ΩP \in\Omega, let uP\mathbf{u}_P be one on every list at PP and zero at other positions. These vectors are not normalized. The estimate proved in Proposition 4.1 is, for any prescribed E0>0E_0 > 0,

1UV∑P∈Ωprimitive⟨uP,(AA∗)RuP⟩≤L−E0N.(24)\frac{1}{UV}\sum_{\substack{P\in\Omega\\\text{primitive}}}\langle\mathbf{u}_P,(AA^*)^R\mathbf{u}_P\rangle\le L^{-E_0N}. \tag*{(24)}

First A0A_0 and then KK are chosen sufficiently large; thresholds in absolute errors may subsequently depend on the fixed group exponents. We prove here the consequence

∣⟨f,Ag⟩∣≪UVL−E0+OC(1).(25)|\langle f,Ag\rangle| \ll UVL^{-E_0+O_C(1)}. \tag*{(25)}

Partition positions into half-open intervals of the slope P2/P1P_2/P_1 of length H/U2H/U^2. A nonzero edge has ∣det⁡(P,Q)∣≪H|\det(P,Q)| \ll H, so only a bounded number of neighboring intervals can interact. Each interval contains O(H)O(H) primitive positions. To see this, choose a primitive pivot in it. Its determinant with any other position in the interval is an integer of size O(H)O(H). With that determinant fixed, all integer solutions differ by integral multiples of the pivot; the first coordinate ranges over [U,16U][U,16U], permitting O(1)O(1) multiples.

Let fjf_j be ff restricted to one interval, and gjg_j the restriction of gg to its interacting neighbors. Set B0=AA∗B_0=AA^* and

dj=∑P in interval j⟨uP,B0RuP⟩.d_j=\sum_{P\text{ in interval }j}\langle\mathbf{u}_P,B_0^R\mathbf{u}_P\rangle.

The matrix CPQ(j)=⟨uP,B0RuQ⟩C_{PQ}^{(j)}=\langle\mathbf{u}_P,B_0^R\mathbf{u}_Q\rangle is positive semidefinite on the ordinary coefficient space. Its largest eigenvalue is at most its trace djd_j, irrespective of the norms or orthogonality of the testing vectors. Since fj=∑Pf(P)uPf_j=\sum_P f(P)\mathbf{u}_P, this gives

⟨fj,B0Rfj⟩≤dj∑P∣f(P)∣2≪Hsup⁡∣f∣2dj.(26)\langle f_j,B_0^R f_j\rangle\le d_j\sum_P|f(P)|^2\ll H\sup|f|^2d_j. \tag*{(26)}

Spectral Hölder for the positive operator B0B_0, followed by Cauchy’s inequality, now gives

∣⟨fj,Agj⟩∣≤∥gj∥∥fj∥1−1/R(CHsup⁡∣f∣2dj)1/(2R).|\langle f_j,Ag_j\rangle|\le\|g_j\|\|f_j\|^{1-1/R}(CH\sup|f|^2d_j)^{1/(2R)}.

Hölder’s inequality in jj, with exponents 22, 2R/(R−1)2R/(R-1), 2R2R, and bounded overlap of the gjg_j imply

∣⟨f,Ag⟩∣≪∥g∥∥f∥1−1/R(CHsup⁡∣f∣2)1/(2R)(∑jdj)1/(2R).|\langle f,Ag\rangle|\ll\|g\|\|f\|^{1-1/R}(CH\sup|f|^2)^{1/(2R)}\left(\sum_jd_j\right)^{1/(2R)}.

By (3.7) and (3.13), log⁡H=OK(L0.21)\log H=O_K(L^{0.21}), so H1/N=O(1)H^{1/N}=O(1). Equations (3.19) and (3.18) prove (3.20). This argument uses the unnormalized trace in (3.19); no normalization by the number of lists has been introduced.

Uniform residues at a primitive root

The root average in (3.19) will be replaced by independent projective residue lines. We give an elementary count that works even when the two coordinate scales are comparable. For a primitive (u,v)∈Ω(u,v)\in\Omega there is a unique completion

g=(ucvd),ud−vc=1,0≤c<u.(27)g=\begin{pmatrix}u&c\\v&d\end{pmatrix},\qquad ud-vc=1,\qquad0\le c<u. \tag*{(27)}

Put r=c/ur=c/u.

Lemma 3.3 (Root residues). There is cδ>0c_\delta>0 such that the following holds. Let SS be squarefree with S≤exp⁡(CKL0.98)S\leq\exp(C_K L^{0.98}), and prescribe g0∈SL2(Z/SZ)g_0\in\mathrm{SL}_2(\mathbb{Z}/S\mathbb{Z}). If I1,I2,I3I_1,I_2,I_3 are intervals in [1,16][1,16], [1,2][1,2], [0,1][0,1], respectively, the number of primitive roots satisfying

u/U∈I1,v/V∈I2,r∈I3,g≡g0(modS)u/U\in I_1,\qquad v/V\in I_2,\qquad r\in I_3,\qquad g\equiv g_0\pmod S

is

UV∣I1∣∣I2∣∣I3∣ζ(2)∣SL2(Z/SZ)∣+O(UVx−cδ).(28)\frac{UV\lvert I_1\rvert\lvert I_2\rvert\lvert I_3\rvert}{\zeta(2)\lvert\mathrm{SL}_2(\mathbb{Z}/S\mathbb{Z})\rvert}+O(UVx^{-c_\delta}). \tag*{(28)}

The error is uniform in the intervals and in g0g_0.

Proof. We first prove the elementary exponential-sum estimate needed for the count. For m≥1m\geq1 put

Km(h,k)=∑y∈(Z/mZ)×e(hy+ky−1m).K_m(h,k)=\sum_{y\in(\mathbb{Z}/m\mathbb{Z})^\times}e\left(\frac{hy+ky^{-1}}{m}\right).

Orthogonality shows that ∑h,k mod m∣Km(h,k)∣4=m2Tm\sum_{h,k\bmod m}\lvert K_m(h,k)\rvert^4=m^2T_m, where TmT_m counts unit quadruples (y1,y2,z1,z2)(y_1,y_2,z_1,z_2) with equal sums and equal inverse sums. The first relation determines z2=y1+y2−z1z_2=y_1+y_2-z_1. After multiplying the inverse-sum relation by the product of the four units, one obtains

m∣(y1+y2)(z1−y1)(z1−y2).m\mid(y_1+y_2)(z_1-y_1)(z_1-y_2).

The map from the three free residues to these three linear forms has determinant of absolute value two, and its kernel modulo mm has size at most two. Distribute, prime by prime, the valuation of mm among the three factors. There are τ3(m)\tau_3(m) distributions d1d2d3=md_1d_2d_3=m; for each, the number of triples of forms with the prescribed divisibilities is m3/(d1d2d3)=m2m^3/(d_1d_2d_3)=m^2. Consequently Tm≤2m2τ3(m)T_m\leq2m^2\tau_3(m).

The value of Km(h,k)K_m(h,k) is constant on the unit-scaling orbit (h,k)↦(hs,ks−1)(h,k)\mapsto(hs,ks^{-1}). If g=(h,k,m)g=(h,k,m), its stabilizer consists of the units s≡1(modm/g)s\equiv1\pmod{m/g}, so the orbit has φ(m/g)≥(m/g)m−o(1)\varphi(m/g)\geq(m/g)m^{-o(1)} elements. Dividing the fourth moment bound by this orbit size gives

∣Km(h,k)∣≪m3/4+o(1)(h,k,m)1/4.(29)\lvert K_m(h,k)\rvert\ll m^{3/4+o(1)}(h,k,m)^{1/4}. \tag*{(29)}

Only the elementary bounds τ3(m)=mo(1)\tau_3(m)=m^{o(1)} and φ(m)≥m1−o(1)\varphi(m)\geq m^{1-o(1)} enter here; both follow by separating the finitely many small primes from the remaining prime factors.

Suppose first U≤VU\leq V. Write the entries of the prescribed matrix as u0,c0,v0,d0′u_0,c_0,v_0,d'_0; the prime on d0′d'_0 distinguishes it from the pad scale. Fix u≡u0(modS)u\equiv u_0\pmod S. Modulo m0=uSm_0=uS, the required pairs (c,v)(c,v) satisfy

c≡c0(modS),v≡v0(modS),cv≡ud0′−1(moduS).(30)c\equiv c_0\pmod S,\qquad v\equiv v_0\pmod S,\qquad cv\equiv ud'_0-1\pmod{uS}. \tag*{(30)}

There are exactly

M(u)=u∏p∣up∤S(1−p−1)M(u)=u\prod_{\substack{p\mid u\\p\nmid S}}(1-p^{-1})

such pairs. Indeed, the admissible cc must be coprime to uu. At primes common to u,Su,S this is forced by u0d0′−c0v0=1u_0d'_0-c_0v_0=1 modulo SS; at other primes the indicated Euler factors count the admissible cc in its progression. For any such cc, writing v=v0+Sjv=v_0+Sj reduces the last congruence to a linear congruence modulo uu with invertible coefficient cc, and hence determines exactly one j(modu)j\pmod u.

These M(u)M(u) pairs have discrepancy

O(u0.99S2)O(u^{0.99}S^2)

for rectangles in the torus (c/(uS),v/(uS))(c/(uS), v/(uS)). Here are details of the uniformity in SS. Put S1=∏p∣(S,u)pS_1 = \prod_{p\mid(S,u)}p and S2=S/S1S_2 = S/S_1. The Chinese remainder theorem fixes (c,v)(c,v) modulo S2S_2 and leaves an inverse graph modulo m1=uS1m_1=uS_1. Write a′=ud0′−1a'=ud_0'-1 and γ=S2−1(modm1)\gamma=S_2^{-1}\pmod{m_1}; both are units modulo m1m_1. On the inverse graph cv=a′cv=a' the condition c≡c0(modS1)c\equiv c_0\pmod{S_1} already forces v≡v0(modS1)v\equiv v_0\pmod{S_1}, since c0c_0 is a unit there. Additive orthogonality thus gives the Fourier sum, up to a constant of absolute value one, as

1S1∑j mod S1e(−jc0/S1)Km1(γh+ju,γka′).(31)\frac{1}{S_1}\sum_{j\bmod S_1}e(-jc_0/S_1)K_{m_1}(\gamma h+ju,\gamma ka'). \tag*{(31)}

For 0<max⁡(∣h∣,∣k∣)≤u0.020<\max(|h|,|k|)\le u^{0.02},

(γh+ju,γka′,m1)≤S1(h,k,u)≤S1u0.02.(\gamma h+ju,\gamma ka',m_1)\le S_1(h,k,u)\le S_1u^{0.02}.

Using the exponent 3/4+1/2003/4+1/200 in (29), each nonzero sum (31) is therefore O(u0.76S2)O(u^{0.76}S^2).

For completeness, sandwich each interval indicator between smooth upper and lower functions obtained by enlarging or contracting the interval by 2u−0.012u^{-0.01} and convolving with a probability bump of width u−0.01u^{-0.01}. Their zeroth coefficients differ from the interval length by O(u−0.01)O(u^{-0.01}). Their Fourier coefficients are bounded. There are O(u0.04)O(u^{0.04}) retained frequency pairs, whose total cost is O(u0.80S2)O(u^{0.80}S^2). For explicit control of the tail, if Δ=u−0.01\Delta=u^{-0.01} and T1=u0.02T_1=u^{0.02}, ten integrations by parts give a double Fourier tail outside max⁡(∣h∣,∣k∣)≤T1\max(|h|,|k|)\le T_1 bounded by O(Δ−11T1−9)=O(u−0.07)O(\Delta^{-11}T_1^{-9})=O(u^{-0.07}). The trivial bound M(u)≤uM(u)\le u makes its contribution O(u0.93)O(u^{0.93}). The error from the zeroth coefficient is O(u0.99)O(u^{0.99}). This proves (3.27).

The cc interval has length u∣I3∣u|I_3|, and the vv interval has length V∣I2∣V|I_2|. If the latter passes through several periods of length uSuS, apply the rectangle count to each; the number of pieces is O(V/u+1)O(V/u+1). Equations (3.26)–(3.27) give the fixed-uu main term

V∣I2∣∣I3∣S2∏p∣up∤S(1−p−1)\frac{V|I_2||I_3|}{S^2}\prod_{\substack{p\mid u\\p\nmid S}}(1-p^{-1})

and error O((V/u+1)u0.99S2)O((V/u+1)u^{0.99}S^2). For u≍Uu\asymp U the sum of these errors is O(UVU−0.01S2)O(UVU^{-0.01}S^2).

It remains to average the Euler factor in the progression of uu. Expanding it as ∑l∣u,(l,S)=1μ(l)/l\sum_{l\mid u,(l,S)=1}\mu(l)/l gives

∑u/U∈I1u≡u0(modS)∏p∣up∤S(1−p−1)=U∣I1∣S∑(l,S)=1μ(l)l2+O(log⁡(2U)).\sum_{\substack{u/U\in I_1\\u\equiv u_0\pmod S}}\prod_{\substack{p\mid u\\p\nmid S}}(1-p^{-1}) = \frac{U|I_1|}{S}\sum_{(l,S)=1}\frac{\mu(l)}{l^2}+O(\log(2U)).

The tail beyond 16U16U contributes O(1/S)O(1/S) and is covered by the error. The series equals ζ(2)−1∏p∣S(1−p−2)−1\zeta(2)^{-1}\prod_{p\mid S}(1-p^{-2})^{-1}. Counting first columns and their completions over each field shows

∣SL2(Z/SZ)∣=S3∏p∣S(1−p−2).(32)|\mathrm{SL}_2(\mathbb{Z}/S\mathbb{Z})|=S^3\prod_{p\mid S}(1-p^{-2}). \tag*{(32)}

This proves the stated main term. Since U≥xδU\ge x^\delta and S2=xo(1)S^2=x^{o(1)}, all the errors admit a saving x−cδx^{-c_\delta}; for example, cδ=δ/1000c_\delta=\delta/1000 is sufficient.

If V<UV<U, fix vv instead and apply the same argument to the congruence ud≡1+vc0(modvS)ud\equiv1+vc_0\pmod{vS}, with the roles of the coordinates reversed and determinant sign reversed. Use d/vd/v as the third coordinate. For u,v>1u,v>1 both c/uc/u and d/vd/v belong to [0,1)[0,1), and

dv−cu=1uv.\frac{d}{v}-\frac{c}{u}=\frac{1}{uv}.

Enlarging and contracting the prescribed interval by O((UV)−1)O((UV)^{-1}) changes its main mass by O(1)O(1) and its count by the already available discrepancy. The same formula follows. □

Independent residue lines and the archimedean error budget

Expand (3.19) as a path with NN edges, returning to its initial position but with no return condition on the lists. In the basis (3.22), each position is gzgz with primitive z∈Z2z \in\mathbb{Z}^2, and the initial and final vectors are e1=(1,0)e_1=(1,0). Put τ(z)=z1+rz2\tau(z)=z_1+rz_2. The cutoff and the physical box imply

∣z2∣≪NH,∣z1∣≪NH,τ(z)≍1.(33)|z_2| \ll NH,\qquad|z_1| \ll NH,\qquad\tau(z)\asymp1. \tag*{(33)}

Indeed, one edge changes slope by O(H/U2)O(H/U^2); summing at most NN changes gives z2=det⁡((u,v),gz)=O(NH)z_2=\det((u,v),gz)=O(NH). Also (gz)1=uτ(z)≍U(gz)_1=u\tau(z)\asymp U, which proves the remaining claims. For a group prime pp, the condition p∣(gz)1p\mid(gz)_1 means that the projective line [z]p[z]_p equals the kernel line of the first row of gg modulo pp.

Lemma 3.4 (Replacement of the root average). In the normalized path expansion of (3.19), one may replace the root average by

1ζ(2)∫[1,16]×[1,2]×[0,1]d(u/U) d(v/V) dr\frac{1}{\zeta(2)}\int_{[1,16]\times[1,2]\times[0,1]} \mathrm{d}(u/U)\,\mathrm{d}(v/V)\,\mathrm{d}r

and by independent uniformly distributed lines in P1(Fp)\mathbb{P}^1(\mathbb{F}_p) for all the group primes. The real matrix at an archimedean root is

g(u,v,r)=(uurvvr+1/u).g(u,v,r)= \begin{pmatrix} u & ur\\ v & vr+1/u \end{pmatrix}.

All box and good-state conditions are retained. The total error is OA(L−AN)O_A(L^{-AN}) for every fixed A>0A>0. The same assertion, with polynomial logarithmic coefficient bounds, holds for one edge and for failure of either endpoint’s goodness.

Proof. We specify both the combinatorial and the archimedean budgets. There are exp⁡(OK(L0.73))\exp(O_K(L^{0.73})) choices for paths zz satisfying the size bounds in (3.30), labels of all visits, and slot permutations. For example, the logarithm of the number of position paths is O(Nlog⁡(NH))=OK(L0.71)O(N\log(NH))=O_K(L^{0.71}); the label choices have logarithm at most OK(NML0.2)=OK(L0.71)O_K(NML^{0.2})=O_K(L^{0.71}); and the permutations cost OK(NMlog⁡M)O_K(NM\log M). These also bound the total absolute coefficients after dropping congruences and bounded damping. The product of the primes active anywhere on a path has logarithm OK(L0.72)O_K(L^{0.72}).

Keep the exact factors of the primes active at some visit. At the other primes the joint damping can be written 1−Xp1-X_p, where 0≤Xp≤10\le X_p\le1. Here the damping exponent at an endpoint is qj=qq_j=\sqrt{q} and at an interior visit is qj=qq_j=q. In the independent-line model,

∑pEXp≤∑p∑j=0N1−qjp+1≪KN+1.\sum_p \mathbb{E}X_p\le\sum_p\sum_{j=0}^{N}\frac{1-q_j}{p+1}\ll_K N+1.

Upper and lower Bonferroni polynomials of consecutive degrees near T=⌈L0.76⌉T=\lceil L^{0.76}\rceil sandwich ∏p(1−Xp)\prod_p(1-X_p) pointwise. Their expected difference is at most

(CK(N+1))TT!=exp⁡(−Ω(L0.76log⁡L)).\frac{(C_K(N+1))^T}{T!}=\exp\bigl(-\Omega(L^{0.76}\log L)\bigr).

The product inequality is ordinary inclusion–exclusion, applied to numbers in [0,1][0,1]; it requires no independence for the pointwise sandwich. Independence between different residue lines is used only to estimate its gap. This gap absorbs the exp⁡(OK(L0.73))\exp(O_K(L^{0.73})) path and coefficient budget.

Each monomial in the truncated expressions requires residues at the active primes and at most TT further primes. Its squarefree modulus satisfies log⁡S=OK(L0.96)\log S=O_K(L^{0.96}), and hence lies in the range of

Lemma 3.3. The total number of monomials, prime choices, and residue matrices is exp⁡(o(L))\exp(o(L)). The uniform distribution on SL2(Z/SZ)\mathrm{SL}_2(\mathbb{Z}/S\mathbb{Z}) factors over the primes by the Chinese remainder theorem. At each prime its first row is uniform among nonzero rows, and its kernel is uniform among the p+1p+1 projective lines. This is exactly the asserted independent model, with primitive-root density 1/ζ(2)1/\zeta(2).

We next make the archimedean conditions compatible with interval counts. Choose an integral complement ww to each primitive zz, with det⁡(z,w)=1\det(z,w)=1 and ∣w∣≪1+∣z∣|w|\ll1+|z|. The ratio for the position gzgz is

rgz=τ(w)τ(z)(mod1).(34)r_{gz}=\frac{\tau(w)}{\tau(z)}\pmod{1}. \tag*{(34)}

On τ(z)≍1\tau(z)\asymp1 its derivative in rr is 1/τ(z)21/\tau(z)^2. The good tests are therefore constant except at exp⁡(o(L))\exp(o(L)) values of rr: enumerate every product and fresh draw in the tests, every pair of distinct integers DZDZ, and the integer translates in each circle inequality. All their numbers and sizes are exp⁡(o(L))\exp(o(L)). The exceptional-probability condition in the second good test is also constant between these endpoints, because it is a finite weighted sum of such indicators.

The physical coordinates are

(gz)1=uτ(z),(gz)2=vτ(z)+z2/u.(gz)_1=u\tau(z),\qquad(gz)_2=v\tau(z)+z_2/u.

The normalized box faces consequently have derivatives bounded by exp⁡(o(L))\exp(o(L)). Near a first-coordinate face the derivative in u/Uu/U is bounded below, and near a second-coordinate face the derivative in v/Vv/V is bounded below, because τ(z)≍1\tau(z)\asymp1. One may first exclude τ(z)\tau(z) near zero, which is incompatible with the first-coordinate box condition. Thus the union of cubes meeting any face or good-test boundary has volume at most x−ϵδ+o(1)x^{-\epsilon_\delta+o(1)} on a grid of side x−ϵδx^{-\epsilon_\delta} in (u/U,v/V,r)(u/U,v/V,r).

Choose the grid exponent only after the saving in Lemma 3.3; concretely take ϵδ=cδ/8\epsilon_\delta=c_\delta/8. The grid has O(x3ϵδ)O(x^{3\epsilon_\delta}) cubes, not a subpower number. Summing all root-count errors, including the exp⁡(o(L))\exp(o(L)) discrete choices, costs at most

UVx−cδ+3ϵδ+o(1)=UVx−5cδ/8+o(1).UVx^{-c_\delta+3\epsilon_\delta+o(1)}=UVx^{-5c_\delta/8+o(1)}.

The main mass of boundary cubes is at most

UVx−ϵδ+o(1)=UVx−cδ/8+o(1).UVx^{-\epsilon_\delta+o(1)}=UVx^{-c_\delta/8+o(1)}.

Their discrepancy is already included in (3.33). On each remaining cube all archimedean indicator conditions are fixed. The edge factors in (22) depend, after fixing labels, on det⁡(z,z′)/D\det(z,z')/D and hence do not vary with the root; the finite residue factors have already been handled. This proves the replacement with a fixed-power error after normalization. Such an error is OA(L−AN)O_A(L^{-AN}), since N≍L0.5N\asymp L^{0.5}. The one-edge versions have smaller enumeration budgets and use the same proof. □

It remains to bound the signed independent-line path expansion. The next section keeps the precise correlations caused by a prime that leaves an active list and later returns, proves (24), and then removes the good-state projections and the auxiliary pads. This yields the replacement of (19) by (20) with error OA(XYL−A)O_A(XYL^{-A}).

Signed memory and the determinant moment

We retain the notation, physical state space, good-state tests, and operator A\mathcal{A} of the preceding section. In particular, the number J=⌊L0.1⌋J=\lfloor L^{0.1}\rfloor of pads per group grows with xx, M=J+1M=J+1, and N=2R≍L0.5N=2R\asymp L^{0.5}. All constants depending on the fixed number KK of groups are allowed to affect the threshold for xx. We will explicitly distinguish those constants from powers of LL.

Proposition 4.1 (The minor-arc moment). For every fixed E0>0E_0 > 0, one can choose A0A_0 sufficiently large and then KK sufficiently large, in terms of E0E_0, qq and the fixed parameters of Theorem 3.1, so that the operator with the good-state projections satisfies

1UV∑P∈ΩP primitive⟨uP,(AA∗)RuP⟩≤L−E0N.(35)\frac{1}{UV}\sum_{\substack{P\in\Omega\\ P\ \mathrm{primitive}}}\langle u_P,(\mathcal{A}\mathcal{A}^{*})^{R}u_P\rangle\le L^{-E_0N}. \tag*{(35)}

The choices of A0,KA_0,K do not depend on the particular fixed band exponents a1<⋯<aKa_1<\cdots<a_K; the threshold for xx may depend on them. Consequently, for any prescribed fixed A>0A>0, the determinant indicator in the expanded square of the preceding section can be replaced by (3.11), with error OA(XYL−A)O_A(XYL^{-A}). In the resulting major term the disjointness restriction on the two unshared products a,ba,b can be omitted. Equivalently, with the notation of (3.8) and (5.1),

∣QY−QYmaj∣≪AXYL−A.(36)\left|Q_Y-Q_Y^{\mathrm{maj}}\right|\ll_A XYL^{-A}. \tag*{(36)}

The precision needed in the pairing estimate (3.20) for this consequence is independent of KK.

The proof constructs operators which remember a prime between two of its uses. The construction has a resemblance to the lifespan expansion in [21], Section 3; the primewise identity and every operator estimate needed here are proved below for the determinant graph.

The exact primewise identity

Apply Lemma 3.4 to the expansion of the left side of (4.1). We work for now at a fixed archimedean root and with independent uniform kernel lines at the group primes. Write the path as z0,…,zNz_0,\ldots,z_N, where z0=zN=e1z_0=z_N=e_1; all the zjz_j are primitive and satisfy the tube and box conditions of (3.30). Put

qj={q,j=0,N,q,0<j<N,ηj′=1−qj,bp′=1−∑j=0Nηj′p+1,νi(p)=1Vi(p+1)bp′(p∈Pi).(37)q_j= \begin{cases} \sqrt{q}, & j=0,N,\\ q, & 0<j<N, \end{cases} \qquad \eta'_j=1-q_j, \qquad b'_p=1-\frac{\sum_{j=0}^{N}\eta'_j}{p+1}, \qquad \nu_i(p)=\frac{1}{V_i(p+1)b'_p}\quad(p\in\mathcal{P}_i). \tag*{(37)}

Thus bp′>0b'_p>0 for large xx. With pmin⁡=min⁡⋃iPip_{\min}=\min\bigcup_i\mathcal{P}_i, we have

νi(p)μi(p)=1+O(N+1p),si:=∑p∈Piνi(p)=1+O(N+1pmin⁡),sup⁡i,pνi(p)≤CKe−L.1.(38)\frac{\nu_i(p)}{\mu_i(p)}=1+O\left(\frac{N+1}{p}\right), \qquad s_i:=\sum_{p\in\mathcal{P}_i}\nu_i(p)=1+O\left(\frac{N+1}{p_{\min}}\right), \qquad \sup_{i,p}\nu_i(p)\le C_K e^{-L^{.1}}. \tag*{(38)}

The estimates are uniform in the path. Here and below CKC_K can also depend on qq and on fixed smooth cutoffs.

For one prime pp, let ApA_p be its set of active visits, namely the visits where pp occurs in the active list. For a line ℓ∈P1(Fp)\ell\in\mathbb{P}^1(\mathbb{F}_p) set Hj(ℓ)=1ℓ=[zj]pH_j(\ell)=1_{\ell=[z_j]_p}. Before active-list normalizations, its factor in the independent-line expectation is exactly

1p+1∑ℓ∈P1(Fp)∏j∈ApHj(ℓ)∏j∉Ap(1−ηj′Hj(ℓ)).(39)\frac{1}{p+1}\sum_{\ell\in\mathbb{P}^1(\mathbb{F}_p)} \prod_{j\in A_p}H_j(\ell)\prod_{j\notin A_p}(1-\eta'_jH_j(\ell)). \tag*{(39)}

Indeed the physical damping charges qjq_j precisely when the prime divides the first coordinate at visit jj without belonging to its active list.

First suppose Ap=∅A_p=\varnothing. Expand the product in (4.5). The empty and singleton subsets contribute bp′b'_p. In every larger subset fix its first and last indices a<ca<c; the sum over all choices of internal indices gives

bp′+1p+1∑0≤a<c≤Nηa′ηc′1[za]p=[zc]p∏a<j<c(1−ηj′1[zj]p=[za]p).(40)b'_p+\frac{1}{p+1}\sum_{0\le a<c\le N}\eta'_a\eta'_c1_{[z_a]_p=[z_c]_p}\prod_{a<j<c}(1-\eta'_j1_{[z_j]_p=[z_a]_p}). \tag*{(40)}

After division by bp′b'_p, a summand is a lifespan beginning with a ghost at aa and ending with a ghost at cc. It has one factor ((p+1)bp′)−1((p+1)b'_p)^{-1}, one factor −ηj′-\eta'_j at each endpoint, and a factor qjq_j at each strictly internal visit that hits its line. There is no lifespan with just one ghost: all singleton terms have already been included in the baseline.

Next suppose Ap≠∅A_p \ne\varnothing. If its active lines disagree, the factor is zero. Otherwise let ℓ\ell be their common line and put a=min⁡Apa=\min A_p, c=max⁡Apc=\max A_p. This line is chosen with probability 1/(p+1)1/(p+1). The inactive visits between aa and cc retain their damping factors. The products outside this interval have the exact identities

∏j<a(1−ηj′Hj(ℓ))=1−∑h<aηh′Hh(ℓ)∏h<j<a(1−ηj′Hj(ℓ)),(41)\prod_{j<a}(1-\eta'_jH_j(\ell))=1-\sum_{h<a}\eta'_hH_h(\ell)\prod_{h<j<a}(1-\eta'_jH_j(\ell)), \tag*{(41)}
∏j>c(1−ηj′Hj(ℓ))=1−∑h>cηh′Hh(ℓ)∏c<j<h(1−ηj′Hj(ℓ)).(42)\prod_{j>c}(1-\eta'_jH_j(\ell))=1-\sum_{h>c}\eta'_hH_h(\ell)\prod_{c<j<h}(1-\eta'_jH_j(\ell)). \tag*{(42)}

These follow by telescoping a finite product, in opposite orders at the two ends. They attach either no ghost or one ghost extension to each end of the active interval. Every inactive internal hit then pays qjq_j, and each chosen ghost endpoint pays −ηj′-\eta'_j.

Extract ∏pbp′≤1\prod_p b'_p\le1 over all group primes. Equations (4.6)–(4.8) show that every term of the expansion assigns at most one lifespan to each prime. All active visits and ghost endpoints of a lifespan must have the same line. Its probability 1/(p+1)1/(p+1), together with division by bp′b'_p, contributes the factor ((p+1)bp′)−1((p+1)b'_p)^{-1} exactly once. The remaining factors are the damping at inactive internal visits and the active-list normalization. An active run is a maximal interval of consecutive active visits. The ban in the physical edge ensures that a prime active at two consecutive visits is continued in a shared slot. Thus a lifespan in group ii with rr active runs has factor

Vi−r(p+1)bp′(43)\frac{V_i^{-r}}{(p+1)b'_p} \tag*{(43)}

before its ghost signs and internal damping. The power of Vi−1V_i^{-1} counts active runs; the line probability is paid only once.

A lifespan retains its prime and line while the prime is outside the active list; call such a retained prime pending. At a strictly internal inactive visit it is a hit if [zj]p[z_j]_p equals that line, and a miss otherwise. The prime keeps the same line across both kinds of visit. Figure 2 shows how one lifespan can join two active runs across a pending hit and a pending miss.

A schematic lifespan of a prime from group $i$

Figure 2. A schematic lifespan of a prime from group ii. All active visits and both ghost endpoints have its common line; the pending visit at 4 misses that line. The line contributes 1/(p+1)1/(p+1) once. After extracting bp′b'_p, the ghost signs, pending-hit damping, and active-run normalizations give the weight η0′η7′q3q6Vi−2/((p+1)bp′)\eta'_0\eta'_7q_3q_6V_i^{-2}/((p+1)b'_p). The two powers of Vi−1V_i^{-1} count active runs, not active visits.

The symmetric memory space and its operations

We realize the preceding expansion by storing each pending prime together with its line. A later active run retrieves that stored pair, instead of paying for a new line. The following measures and transfer factors are chosen to reproduce (4.9).

Let Z\mathcal{Z} be the finite set of primitive integer vectors in the tube which satisfy the fixed-root box condition; it has counting measure. An active list is an ordered MM-tuple in each group, endowed with measure ⨂iνi⊗M\bigotimes_i \nu_i^{\otimes M}. The underlying list space allows repetitions. Distinctness within physical lists and the other local restrictions will be imposed in the operators.

A memory particle in group ii consists of a prime and an anchor,

Xi={(p,ℓ):p∈Pi, ℓ∈P1(Fp)},λi(p,ℓ)=νi(p).(44)\mathcal{X}_i=\{(p,\ell):p\in\mathcal{P}_i,\ \ell\in\mathbb{P}^1(\mathbb{F}_p)\},\qquad\lambda_i(p,\ell)=\nu_i(p). \tag*{(44)}

The line coordinate has counting measure, not uniform probability measure. In particular, anchoring a particle at [z]p[z]_p leaves the prime measure νi(p)\nu_i(p), without another factor 1/(p+1)1/(p+1). A particle hits zz when its anchor equals [z]p[z]_p.

For a vector of memory sizes k=(k1,…,kK)\mathbf{k}=(k_1,\ldots,k_K), use symmetric functions of the kik_i particles in each group and the measure λi⊗ki/ki!\lambda_i^{\otimes k_i}/k_i!. The Hilbert space is

H=ℓ2(Z)⊗L2(⨂iνi⊗M)⊗⨂i=1K(⨁ki≥0Lsym2(Xiki,λi⊗kiki!)).(45)\mathcal{H}=\ell^2(\mathcal{Z})\otimes L^2\left(\bigotimes_i \nu_i^{\otimes M}\right)\otimes\bigotimes_{i=1}^K\left(\bigoplus_{k_i\geq0}L^2_{\mathrm{sym}}\left(\mathcal{X}_i^{k_i},\frac{\lambda_i^{\otimes k_i}}{k_i!}\right)\right). \tag*{(45)}

Products and sums of the operators below use the row convention: the input state is fixed, choices leading to output states are summed, and the resulting operator acts on a test function at the output. Let b\mathbf{b} be one when z=e1z=e_1 and memory is empty, and zero otherwise; it is independent of the active lists. By (4.4),

∥b∥2=∏iSiM=exp⁡(OK(M(N+1)pmin⁡))=1+o(1).(46)\lVert\mathbf{b}\rVert^2=\prod_i S_i^M=\exp\left(O_K\left(\frac{M(N+1)}{p_{\min}}\right)\right)=1+o(1). \tag*{(46)}

Fix once and for all

0<ρ<1,ρ2>q.(47)0<\rho<1,\qquad\rho^2>\sqrt{q}. \tag*{(47)}

The ghost operation Gj\mathcal{G}_j leaves the position and active list unchanged. Independently in each group it first chooses a subset of the indexed pending particles which hit the current position and deletes them, paying −ηj′/ρ-\eta'_j/\rho per deletion. Each surviving hit pays qj/ρ2q_j/\rho^2. It then appends an ordered batch of ℓ≥0\ell\geq0 primes, integrated with measure νj⊗ℓ/ℓ!\nu_j^{\otimes\ell}/\ell!, anchored at the current lines, and with coefficient (−ηj′Vi/ρ)ℓ(-\eta'_jV_i/\rho)^\ell. The divisor ℓ!\ell! makes the batch an unordered birth set when the primes are distinct; the ordered integral will also be useful after that restriction is removed.

The edge Ej\mathcal{E}_j retains the geometric multiplier in (3.15), or its adjoint according to the alternating physical path, but removes its ω\omega-damping and its active divisibility tests. It replaces target unshared counting and normalization by the following operations between the same endpoint slot symmetrizations.

  • (i) Pay ρ\rho for every old pending particle which hits the source zz. Choose a subset of the unshared source slots to store in memory, anchored at [z]p[z]_p, and drop the other unshared source slots. Copy the shared JJ slots in each group.

  • (ii) Choose a target position z′z' and its unshared label in each group. A subset of these slots is filled by promoting indexed old particles which hit z′z', removing them from memory and paying 1/Vi1/V_i per promotion in group ii. Particles just stored on this edge cannot be promoted on the same edge. Every other unshared target slot is filled by an independent νi\nu_i draw.

  • (iii) Pay ρ\rho for every particle remaining in memory which hits z′z', including newly stored particles. Impose the physical endpoint list restrictions, the cross-edge ban, the pad dyad, all good-state tests and geometric cutoffs, using the actual source and target label products.

Positions are always summed with counting measure. Since the real root matrix has determinant one, its edge determinant is det⁡(z,z′)\det(z,z').

For the moment impose also that all performed births have globally distinct prime values. Births comprise initial active entries, fresh target entries, and ghost creations. This is a restriction on the fully expanded history, and is not claimed to be a local projection on H\mathcal{H}. Then the independent-line path sum, divided by its extracted baseline, is exactly

⟨b,G0E0G1⋯EN−1GNb⟩,with global birth distinctness imposed.(48)\langle b,\mathcal{G}_0\mathcal{E}_0\mathcal{G}_1\cdots\mathcal{E}_{N-1}\mathcal{G}_N b\rangle, \qquad\text{with global birth distinctness imposed.} \tag*{(48)}

Here equality refers to the choice-by-choice expansion of the right side. To verify it, a shared prime continues with the same line: p∣Dp\mid D implies det⁡(z,z′)=0(modp)\det(z,z')=0\pmod p, and primitivity implies [z]p=[z′]p[z]_p=[z']_p. A prime leaving activity either terminates there or is stored. If it is used again, global birth distinctness forces its promotion from that stored particle. Ghost creations and deletions are exactly the optional endpoints of (4.6)–(4.8).

At an internal pending hit, the incoming edge, ghost survival, and outgoing edge give

ρ(qj/ρ2)ρ=qj.(49)\rho\left(q_j/\rho^2\right)\rho=q_j. \tag*{(49)}

At a ghost birth the creation coefficient and outgoing hit factor give −ηj′Vi-\eta'_jV_i; at termination the incoming hit factor and deletion give −ηj′-\eta'_j. There are no unmatched factors at 0,N0,N, because memory starts and ends empty. An active birth has measure νi(p)=((p+1)bp′)−1Vi−1\nu_i(p)=((p+1)b'_p)^{-1}V_i^{-1}, a ghost birth has prime measure Viνi(p)=((p+1)bp′)−1V_i\nu_i(p)=((p+1)b'_p)^{-1} after this cancellation, and each promotion pays Vi−1V_i^{-1}. These are exactly the factors in (4.9): each new active run contributes Vi−1V_i^{-1}, while the lifespan has only one line-probability factor. Conversely a compatible primewise lifespan determines all these choices: store on leaving any nonfinal active run, promote on entering the next, and create or terminate at its specified ghost endpoints. The factorial birth integral counts its unordered ghost birth set once. This proves the identity, including primes which leave and later re-enter activity.

Truncation and adjoints

Let B=⌈L2⌉B=\lceil L^2\rceil and let ΠB\Pi_B restrict total memory size to at most BB. We may insert ΠB\Pi_B before and after every operation in (4.14), with error

OA(L−AN)for every fixed A>0.(50)O_A(L^{-AN}) \qquad\text{for every fixed } A>0. \tag*{(50)}

This assertion is made while births are still globally distinct. There are at most KNKN stores in a history. A history reaching memory size greater than BB has therefore had at least B−KNB-KN ghost births. Insert a factor two at each ghost birth in an absolute majorant. For a prime active somewhere, there are at most (N+2)2(N+2)^2 possible ghost endpoint pairs. There are OK(M+N)O_K(M+N) such primes along a fixed active path. For a never-active prime, fixing its earlier ghost endpoint and summing its possible later endpoints costs Oq(1)O_q(1): successive later endpoints with the required line acquire successive internal factors at most q<1\sqrt q<1. Summing the earlier endpoint gives a total Oq((N+1)/p)O_q((N+1)/p) for the absolute ghost-only contribution, including the inserted factor two. Consequently the product of all never-active prime costs is exp⁡(OK(N+1))\exp(O_K(N+1)).

The extracted baseline and its reciprocal have logarithms OK(N+1)O_K(N+1), by (4.3) and the bounded reciprocal prime sums. The path and label enumeration used in the root replacement costs exp⁡(OK(L.73))\exp(O_K(L^{.73})); multiplying by the active endpoint choices and the preceding absolute costs gives, with room to spare, exp⁡(OK(L.74))\exp(O_K(L^{.74})). Removing the inserted factors on the omitted histories gives the upper bound

2−B+KNexp⁡(OK(L.74)),2^{-B+KN}\exp\left(O_K(L^{.74})\right),

which proves (4.16).

We next omit global birth distinctness and work on HB=∏BH\mathcal{H}_{B}=\prod_{B}\mathcal{H}. The resulting signed norm estimates apply only to the unrestricted history; we will restore distinctness by an exact expansion in equality constraints. Repeated pending values are now permitted. All deletion and promotion choices remain choices of indices, so that the following adjoint identities hold also in their presence.

Suppose aa particles are deleted from a memory list of size k+ak+a. For an unordered ghost deletion batch its indexed choices and the factorial memory measure give

(k+aa)(k+a)!=1k!a!.(51)\frac{\binom{k+a}{a}}{(k+a)!}=\frac{1}{k!a!}. \tag*{(51)}

Its anchors are forced by the hit condition, leaving exactly the νi\nu_i integrations. A fresh ghost batch has the same factorial integral. Interchanging the deleted and appended batches therefore gives the adjoint ghost operation: its deletion coefficient is −ηj′Vi/ρ-\eta'_jV_i/\rho and its creation coefficient is −ηj′/ρ-\eta'_j/\rho; the surviving-hit multiplier stays qj/ρ2q_j/\rho^2.

For retrieval into aa prescribed ordered active slots the corresponding identity is

(k+aa)a!(k+a)!=1k!.(52)\frac{\binom{k+a}{a}a!}{(k+a)!}=\frac{1}{k!}. \tag*{(52)}

The retrieved anchors again force one line and leave νi\nu_i prime integrations. Newly stored particles get those same prime measures from their source active entries. Thus at two fixed positions the joint measure of all source and target active entries and all memory coordinates is unchanged on exchanging stores and promotions. The factor 1/Vi1/V_i on a forward promotion stays on that transfer, becoming a reversed store factor; this is a bounded factor. The two endpoint ρ\rho factors exchange roles. Summing over the two positions preserves the equality because their measure is counting measure. Slot symmetrizations are self-adjoint averages, and all endpoint and edge restrictions transpose with the kernel. These observations prove the adjoint rules even when list restrictions depend jointly on both positions.

For later use, a local choice multiplier of modulus at most one may be inserted provided it is equivariant under simultaneous permutations of the old memory coordinates and the selected deletion or promotion indices. The indexed choice sum then preserves symmetry in the old particles. In a ghost birth integral, a multiplier depending on the ordered fresh batch is averaged over permutations of that batch when acting on symmetric test functions. This leaves the integral unchanged, keeps its modulus at most one, and retains the factor 1/l!1/l!. These modified operations therefore act on the same symmetric spaces. In both adjoint directions they are dominated in absolute value by the positive kernels obtained by summing the absolute values of the unmodified operation-choice coefficients. An arbitrary multiplier depending on an old particle’s index alone need not preserve symmetry and is not allowed here.

Absolute bounds for ghosts and edges

Call an edge clean if it stores and promotes no particle, and dirty otherwise. A clean edge draws all unshared target labels freshly, and its signed minor-arc kernel will give cancellation. For a dirty edge the good-state tests instead make an absolute row or column sum small. We first bound the ghosts and record crude edge bounds, then prove these two kinds of small edge estimate.

For a state with total pending size ktotk_{\mathrm{tot}} use the positive Schur weight

w=v0ktot,(53)w=v_0^{k_{\mathrm{tot}}}, \tag*{(53)}

where v0v_0 is a sufficiently large fixed constant. A weighted row bound RR means that the absolute row sum against output weight is at most RR times input weight; the analogous bound for the adjoint is the weighted column bound CC. Weighted Schur gives norm at most RC\sqrt{RC}. Restrictions by ΠB\Pi_B can only decrease these absolute sums.

In one group with hh pending hits, the ghost row bound, divided by the input weight, is

exp⁡(v0ηj′Visiρ)(qjρ2+ηj′ρv0)h.(54)\exp\left(\frac{v_0\eta'_j V_i s_i}{\rho}\right)\left(\frac{q_j}{\rho^2}+\frac{\eta'_j}{\rho v_0}\right)^h . \tag*{(54)}

The exponential sums fresh births; the power sums deletion or survival of each old hit. The adjoint bound is

exp⁡(v0ηj′siρ)(qjρ2+ηj′Viρv0)h.(55)\exp\left(\frac{v_0\eta'_j s_i}{\rho}\right)\left(\frac{q_j}{\rho^2}+\frac{\eta'_j V_i}{\rho v_0}\right)^h . \tag*{(55)}

Since qj≤q<ρ2q_j \le\sqrt{q}<\rho^2 and the ViV_i are bounded above and below, one fixed v0v_0 makes both parentheses at most one, for all i,ji,j and large xx. Multiplication over the fixed groups proves

∥Gj∥HB→HB≤CK,(56)\|\mathcal{G}_j\|_{\mathcal{H}_B\to\mathcal{H}_B}\le C_K, \tag*{(56)}

with the same absolute row and column bounds under the bounded choice modifications just described.

For the crude edge bounds fix the source zz and choose an integral complement wzw_z with det⁡(z,wz)=1\det(z,w_z)=1. Write

z′=j1wz+hz,j1=det⁡(z,z′),∣h+j1τ(wz)τ(z)∣≤16.(57)z'=j_1w_z+hz,\qquad j_1=\det(z,z'),\qquad\left|h+j_1\frac{\tau(w_z)}{\tau(z)}\right|\le16. \tag*{(57)}

The last inequality follows from τ(z′)/τ(z)\tau(z')/\tau(z) lying in the fixed box range. Enlarging its absolute constant would have no effect on any argument below. For fixed source, labels, and determinant j1j_1, this permits O(1)O(1) targets. In the raw determinant part j1=D(b−a)j_1=D(b-a) is fixed. In the comparison part j1=Dtj_1=Dt has O(Y)O(Y) possibilities, and its multiplier is at most

C∣M∣≪L3A0Y.(58)C|\mathcal{M}|\ll\frac{L^{3A_0}}{Y}. \tag*{(58)}

The masses of all fresh label draws are bounded by CKC_K. There are at most CKBKC_KB^K indexed promotion choices. Store subsets, Schur weight ratios and ViV_i factors cost CKC_K. Source and target symmetrizations are probability averages. Using the adjoint calculation for columns, we obtain a constant C=C(K,A0,q)C=C(K,A_0,q) such that

R(Ej), C(Ej)≤CKBK(1+L3A0)≤LC.(59)R(\mathcal{E}_j),\ C(\mathcal{E}_j)\le C_KB^K(1+L^{3A_0})\le L^C. \tag*{(59)}

This holds for arbitrary multipliers of modulus at most one on the choices. It also holds after fixing any or all fresh prime values, with their measures left outside the bound. This uniformity will preserve the cost of a forced prime atom in the birth-distinctness argument.

The lattice box and the clean signed norm

We first extract a geometric consequence of the first good-state test. Partition Z\mathbb{Z} by intervals of length H=d0YH=d_0Y for z2/τ(z)z_2/\tau(z). The identity

z2′τ(z′)−z2τ(z)=det⁡(z,z′)τ(z)τ(z′)\frac{z'_2}{\tau(z')}-\frac{z_2}{\tau(z)} =\frac{\det(z,z')}{\tau(z)\tau(z')}

shows that an edge connects only cells whose indices differ by a bounded amount. At fixed shared list, hence fixed squarefree product DD, partition further by the tuple ([z]p)p∣D([z]_p)_{p\mid D}. An edge preserves this tuple. The resulting operator is a sum over a bounded number of cell-index shifts of direct sums of blocks. It suffices to bound each block uniformly.

Choose a relevant good source zz in one such pair of cells and a complement wzw_z as in (4.23). All vectors in either cell with the specified line tuple belong to the lattice

Λ=Zz+ZDwz.\Lambda= \mathbb{Z}z + \mathbb{Z}Dw_z.

Indeed, on writing a vector as hz+jwzhz + jw_z, equality of the projective lines modulo every p∣Dp\mid D implies p∣jp\mid j; squarefreeness gives D∣jD\mid j. Its coordinates (h,l)(h,l) in this lattice satisfy

∣l∣≪Y,∣h+lα∣≪1,α=Dτ(wz)τ(z).(60)|l| \ll Y,\qquad|h+l\alpha| \ll1,\qquad\alpha= D^{\frac{\tau(w_z)}{\tau(z)}}. \tag*{(60)}

The ll bound follows from the cell separation and D≍d0D\asymp d_0; the other follows from the box condition. The image of this integer lattice under (h,l)↦(h+lα,l/Y)(h,l)\mapsto(h+l\alpha,l/Y) has covolume 1/Y1/Y. Let λ\lambda be the length of its shortest nonzero vector. If λ<Y−.8\lambda<Y^{-.8}, its second coordinate gives 0<∣l∣<Y.20<|l|<Y^{.2}, and its first gives ∥lα∥<Y−.8<Y−.7\|l\alpha\|<Y^{-.8}<Y^{-.7}; l=0l=0 is impossible for a nonzero vector this short. This contradicts the first good-state test, since τ(wz)/τ(z)\tau(w_z)/\tau(z) is a lift of rgzr_{g_z}. Thus λ≥Y−.8\lambda\ge Y^{-.8}. Dirichlet’s elementary pigeonhole approximation, using denominators at most ⌈Y⌉\lceil\sqrt{Y}\rceil, gives λ≪Y−.5\lambda\ll Y^{-.5}.

A shortest vector is primitive in the lattice, so complete it to a basis. Subtract a multiple of the first vector from the second to make its parallel component at most λ/2\lambda/2. Its perpendicular component is 1/(Yλ)1/(Y\lambda); since λ2≪1/Y\lambda^2\ll1/Y, its length is O(1/(Yλ))O(1/(Y\lambda)). Inverting this basis on the bounded region in (4.26) puts all relevant vectors in a centered coordinate box with side parameters O(1/λ)O(1/\lambda) and O(Yλ)O(Y\lambda). Choose its orientation positively. Enlarging constants, the box has parameters A′,B′A',B' satisfying

A′B′≪Y,Y.1≤A′,B′≤Y.9.(61)A'B' \ll Y,\qquad Y^{.1}\le A',B'\le Y^{.9}. \tag*{(61)}

The change of integer basis has determinant one, so that det⁡(z,z′)/D\det(z,z')/D is the ordinary determinant of their new integer coordinates. This proof applies unchanged for a continuous root: the tested quantity rgzr_{g_z} is still τ(wz)/τ(z)\tau(w_z)/\tau(z) mod 1.

For a clean edge, before the slot symmetrizations, condition on the shared list and all pending particles. They are unchanged by the central operation. Its remaining pending-hit factors are diagonal contractions at the two endpoints. We may omit the extra cross-edge ban at a norm cost smaller than any power of L−1L^{-1}. In fact, keeping distinctness of each full active list, a forbidden coincidence requires one of the fresh labels to equal a fixed source label, costing at most CKMsup⁡p,ννi(p)C_KM\sup_{p,\nu}\nu_i(p). After the labels are fixed, (4.23) and (4.24) give an absolute row bound CKL3A0C_KL^{3A_0}; the same argument for the adjoint gives the column bound. (38) proves the claimed negligible norm cost.

Inside a lattice block, the unshared ordered list, containing one prime from each disjoint band, is uniquely determined by its product bb. Its measure is

m(b)=∏i=1Kνi(pi)≤CKYwhen η(b/Y)≠0.(62)m(b)=\prod_{i=1}^{K}\nu_i(p_i)\le\frac{C_K}{Y}\qquad\text{when }\eta(b/Y)\ne0. \tag*{(62)}

Conjugation from L2(m)L^2(m) to counting measure inserts m(b)m(a)\sqrt{m(b)m(a)}. Consequently the clean norm is at most CKC_K times the norm of the following operator on the integer box in (4.27) and an unrestricted integer product variable:

1Yψ(det⁡(h,l)Y)∫T\Me(θ(det⁡(h,l)−b+a)) dθ.(63)\frac{1}{Y}\psi\left(\frac{\det(h,l)}{Y}\right)\int_{\mathbb{T}\backslash\mathcal{M}}e\bigl(\theta(\det(h,l)-b+a)\bigr)\,d\theta. \tag*{(63)}

Here T=R/Z\mathbb{T}=\mathbb{R}/\mathbb{Z}. Extending the product variable means composing with extension by zero and restriction. All actual product supports, full-list restrictions, goodness tests, cutoff factors, and the bounded square-root factors from (4.28) are endpoint diagonal multipliers. The factor u0/Du_0/D is bounded. Thus none of these operations increases the bound except by CKC_K. We use the transposed conjugate kernel on a reversed edge.

Fourier transformation in the unrestricted integer product variable diagonalizes its convolution: for each θ∈T∖M\theta\in\mathbb{T}\setminus\mathcal{M} the fiber is the determinant exponential matrix with its smooth cutoff, divided by YY. If ψ^(u)=∫Rψ(t)e(−ut) dt\widehat{\psi}(u)=\int_{\mathbb{R}}\psi(t)e(-ut)\,dt, then

ψ(det⁡(h,l)Y)e(θdet⁡(h,l))=∫Rψ^(u)e((θ+u/Y)det⁡(h,l)) du.(64)\psi\left(\frac{\det(h,l)}{Y}\right)e(\theta\det(h,l))=\int_{\mathbb{R}}\widehat{\psi}(u)e\left((\theta+u/Y)\det(h,l)\right)\,du. \tag*{(64)}

The density ψ^\widehat{\psi} decreases faster than any inverse power.

For ∣u∣≤LA0/2|u|\le L^{A_0}/2, put ϑ=θ+u/Y\vartheta=\theta+u/Y. Dirichlet approximation with Q=⌈Y/LA0⌉Q=\lceil Y/L^{A_0}\rceil gives a reduced c/dc/d such that

LA0<d≤Q,∣ϑ−cd∣≤1dQ≤1d2.(65)L^{A_0}<d\le Q,\qquad\left|\vartheta-\frac{c}{d}\right|\le\frac{1}{dQ}\le\frac{1}{d^2}. \tag*{(65)}

Indeed d≤LA0d\le L^{A_0} would put θ\theta within 32LA0/Y\frac{3}{2}L^{A_0}/Y of a rational of allowed denominator, contrary to the definition of M\mathcal{M} with radius 2LA0/Y2L^{A_0}/Y.

For completeness, let TA′,B′(ϑ)T_{A',B'}(\vartheta) have entries e(ϑhk)e(\vartheta hk) for integer ∣h∣≪A′|h|\ll A', ∣k∣≪B′|k|\ll B'. Its Gram matrix and the geometric-sum estimate give

∥TA′,B′(ϑ)∥2≪∑∣h∣≪A′min⁡(B′,1∥hϑ∥)≪(A′/d+1)(B′+dlog⁡(2d)).(66)\|T_{A',B'}(\vartheta)\|^2\ll\sum_{|h|\ll A'}\min\left(B',\frac{1}{\|h\vartheta\|}\right)\ll(A'/d+1)(B'+d\log(2d)). \tag*{(66)}

In the term h=0h=0 the minimum is interpreted as B′B'. To justify the last inequality, split the hh interval into blocks of length at most d/2d/2. For two different indices in a block, reduction of c/dc/d and the error in (4.31) separate their fractional parts by at least 1/(2d)1/(2d). Ordering their distances to the nearest integer bounds each block by O(B′+d∑1≤j≤dj−1)O(B'+d\sum_{1\le j\le d}j^{-1}).

Up to a permutation of columns, the determinant matrix e(ϑ(h1l2−h2l1))e(\vartheta(h_1l_2-h_2l_1)) is the tensor product of TA′,B′(ϑ)T_{A',B'}(\vartheta) and its transposed conjugate. Its norm is therefore the squared norm in (4.32). After division by YY, (4.27) implies the bound

∥(e(ϑdet⁡(h,l)))∥Y≪1d+Y−1log⁡(2Y)+dlog⁡(2d)Y≪KL.2−A0+Y−.1log⁡(2Y).(67)\frac{\|(e(\vartheta\det(h,l)))\|}{Y}\ll\frac{1}{d}+Y^{-1}\log(2Y)+\frac{d\log(2d)}{Y}\ll_K L^{.2-A_0}+Y^{-.1}\log(2Y). \tag*{(67)}

For the complementary Fourier tail in (4.30), the trivial matrix norm is O(A′B′)=O(Y)O(A'B')=O(Y), so rapid decay of ψ^\widehat{\psi} gives any required negative power of LL. Combining the direct sums of cell blocks and the endpoint symmetrizations, we conclude that, for any fixed G>0G>0, choosing A0A_0 sufficiently large in terms of GG gives

∥Ejclean∥≤L−G(68)\|\mathcal{E}^{\mathrm{clean}}_j\|\le L^{-G} \tag*{(68)}

for large xx, with an arbitrarily fixed margin in the exponent. Constants depending on the later fixed value of KK are absorbed by that margin and the threshold for xx.

Dirty raw edges: the two Schur sides

Fix the set SS of stored groups and the set II of promoted groups. There are at most 4K4^K choices, a constant. All store and promotion coefficients and the ratio of Schur weights cost CKC_K. If I=∅I=\varnothing, the weighted raw row bound is CKC_K, independently of BB. For each source omission tuple, sample the fresh target label in every group and use (57) with its fixed determinant. There are no pending-index choices. The source symmetrization averages the omission tuples and thus does not multiply their number. This argument allows arbitrary SS.

Suppose I≠∅I\ne\varnothing and put i∗=max⁡Ii_*=\max I. Fix the source state and sample the fresh target labels in the groups outside II, with product TfT_f. Their product law ν\nu is boundedly comparable to the law μ\mu in the second good-state test, by (38). Outside an exceptional mass CKexp⁡(−14Lai∗)C_K\exp\left(-\frac{1}{4}L^{a i_*}\right), that test supplies its asserted phase separation. Fix one old index for promotion in group i∗i_*; there are at most BB such indices, and write its prime and anchor as (p∗,ℓ∗)(p_*,\ell_*). Let ZZ be the product of the other promoted numeric primes and let PfullP_{\mathrm{full}} be the product of the complete source active list. The raw determinant equation is

j1=Pfull−Dp∗ZTf.(69)j_1=P_{\mathrm{full}}-D p_* ZT_f. \tag*{(69)}

The ban excludes p∗p_* from the entire source list. Thus j1≡Pfull≢0(modp∗)j_1\equiv P_{\mathrm{full}}\not\equiv0\pmod{p_*}, independently of the omission tuple and of ZZ.

In (57) the line condition [j1wz+hz]p∗=ℓ∗[j_1w_z+h_z]_{p_*}=\ell_* is either impossible or fixes a single residue h=h∗h=h_* (mod p∗p_*). To see uniqueness, in the basis (z,wz)(z,w_z) the second coordinate j1j_1 is nonzero modulo p∗p_*; therefore its projective line determines the first coordinate uniquely. This residue depends on the fixed index and PfullP_{\mathrm{full}}, but not on D,ZD,Z. Substitution in (57), using the real lift rz∗=τ(wz)/τ(z)r_z^*=\tau(w_z)/\tau(z), yields

∥DZTfrz∗−Pfullrz∗+h∗p∗∥≤16p∗.(70)\left\lVert DZT_f r_z^*-\frac{P_{\mathrm{full}}r_z^*+h_*}{p_*}\right\rVert\le\frac{16}{p_*}. \tag*{(70)}

The phase separation is 100e−Lai∗100e^{-L^{a i_*}}, whereas the diameter of the interval in (70) is at most 32/p∗≤32e−Lai∗32/p_*\le32e^{-L^{a i_*}}. Hence at most one distinct integer DZDZ is possible.

This also determines DD and ZZ separately. All prime factors of DD are source labels and all prime factors of ZZ are excluded from the source list by the ban. Their prime factorizations therefore separate unambiguously. Because the source list is distinct, DD determines which single label was omitted in every group. Under source symmetrization this one numeric omission tuple has mass exactly M−KM^{-K}. It is an average over tuples, including over their internal orders, rather than an unnormalized sum. The disjoint prime bands also make ZZ determine each other promoted numeric prime.

The labels now fix j1j_1, so there are O(1)O(1) target positions. At any one of them, multiplicities among the remaining promoted indices are absorbed by the outgoing pending damping. If a group has hh old particles hitting that target, the choice of one eligible index and the damping of all other old hits cost at most

hρh−1≤sup⁡n≥1nρn−1<∞.(71)h\rho^{h-1}\le\sup_{n\ge1}n\rho^{n-1}<\infty. \tag*{(71)}

This remains true for repeated numeric values and repeated anchors. Additional newly stored hits only decrease the factor. Applying this in each of the other promoted groups costs CKC_K. For the exceptional fresh draws use the crude BKB^K count and (57). We obtain the weighted raw row bound

εK=CK(BMK+BKe−cL1)(I≠∅).(72)\varepsilon_K=C_K\left(\frac{B}{M^K}+B^K e^{-cL^1}\right)\qquad(I\ne\varnothing). \tag*{(72)}

Reversal exchanges SS and II by (52), changes only bounded ViV_i factors, and retains both good-state tests and the ban. Its raw determinant equation, with the reversed sign, again has source full product minus DD times the target unshared product. The preceding argument therefore gives the following complete table. Every entry includes the fixed CKC_K factors. The first row is only an absolute estimate; its signed estimate is (68). Once KK is large enough that εK≤1\varepsilon_K \le1, weighted Schur bounds the sum of the three dirty raw cases by

Stored groups SSPromoted groups IIWeighted rowWeighted column
∅\varnothing∅\varnothingCKC_KCKC_K
∅\varnothingnonemptyεK\varepsilon_KCKC_K
nonempty∅\varnothingCKC_KεK\varepsilon_K
nonemptynonemptyεK\varepsilon_KεK\varepsilon_K

Table 1.

CKεK≤CKL1−.005K+o(1)+OA(L−A)for every fixed A.(73)C_K\sqrt{\varepsilon_K} \le C_K L^{1-.005K+o(1)} + O_A(L^{-A}) \quad\text{for every fixed } A. \tag*{(73)}

The bounded entry opposite a small Schur entry is essential here: it contains no factor BKB^K. Thus making .005K.005K larger than a prescribed constant purchases that constant in log decay. All remaining KK-dependence in this estimate is a fixed multiplicative constant.

Dirty comparison edges

Use the supremum (58), discarding its dependence on the target unshared labels. Fix a source omission tuple, hence DD, and suppose I≠∅I \ne\varnothing. A promoted old index with prime pp requires one specified projective line at the target. By the ban, p∤Dp \nmid D. The lattice basis used to obtain (61) has determinant DD in the original integer coordinates and is therefore invertible modulo pp. The prescribed line is consequently one nonzero homogeneous linear congruence in the two box coordinates.

For a box of sides A′,B′A', B', the number of solutions of such a congruence is

O(A′B′p+max⁡(A′,B′))≪Y(1p+1min⁡(A′,B′)).(74)O\left(\frac{A'B'}{p}+\max(A',B')\right)\ll Y\left(\frac{1}{p}+\frac{1}{\min(A',B')}\right). \tag*{(74)}

If the coefficient of the shorter coordinate is nonzero, fix the other coordinate and count its solutions in that shorter interval. If that coefficient vanishes, the congruence restricts the longer coordinate instead; the same displayed upper bound follows. We have included vectors divisible by pp in this count, which only increases it.

For each source only a bounded number of neighboring cell blocks occurs. Take the union over at most BB old indices in one promoted group. The other index choices can be bounded by BK−1B^{K-1}, or by (71). Sum the fresh prime masses and the source omission average. Equations (58), (61), and (74) give a row bound

ζK:=CKL3A0BK(pmin⁡−1+Y−1),ζK=OA(L−A)for every fixed A.(75)\zeta_K := C_K L^{3A_0} B^K\left(p_{\min}^{-1}+Y^{-1}\right), \qquad\zeta_K = O_A(L^{-A}) \quad\text{for every fixed } A. \tag*{(75)}

All parameters K,A0K, A_0 are fixed in the last assertion. The opposite Schur side is at most LCL^C by (59). For a store-only piece apply the same argument to the adjoint; with both stores and promotions it applies to both sides. For clarity the comparison counterpart of the dirty table is

Stored groups SSPromoted groups IIWeighted rowWeighted column
∅\varnothingnonemptyζK\zeta_KLCL^C
nonempty∅\varnothingLCL^CζK\zeta_K
nonemptynonemptyζK\zeta_KζK\zeta_K

Table 2.

Weighted Schur makes every dirty comparison piece smaller than every fixed negative power of LL.

Choose G>E0+3G > E_0 + 3. First choose A0A_0 so that the clean argument gives more than GG powers of decay, and then choose KK so that .005K>G+2.005K > G + 2. Combining the clean, dirty raw, and dirty comparison estimates, with their fixed finite sums, proves

∥Ej∥HB→HB≤L−G(76)\lVert\mathcal{E}_j \rVert_{\mathcal{H}_B\to\mathcal{H}_B} \le L^{-G} \tag*{(76)}

before global birth distinctness is restored. If desired all strict inequalities here can be enlarged by one to absorb the constants. The crude exponent C(K,A0,q)C(K,A_0,q) in (4.25) has not entered the choice needed in (4.39).

Restoring global birth distinctness

The contractions just proved did not require globally distinct births. We now restore that condition exactly. A union bound over one repeated prime would not suffice against the absolute cost of a long path; we instead separate equality constraints by their rank.

Allocate a finite ordered set B\mathcal{B} of potential birth addresses as follows: the KMKM initial slots; the KK canonical target slots at each edge, before its final symmetrization; and BB ordered fresh-batch indices per group at each ghost operation. The truncation ensures that a ghost batch has length at most BB. Each address has a local performed flag: at an edge it is performed exactly if that slot is filled freshly, and in a ghost batch exactly if its index does not exceed the batch length. Initial addresses are always performed. Thus

Qbirth:=∣B∣≤KM+KN+K(N+1)B=LO(1).(77)Q_{\mathrm{birth}} := |\mathcal{B}| \le KM + KN + K(N+1)B = L^{O(1)}. \tag*{(77)}

Order these addresses chronologically, with a fixed order within each operation.

Let C\mathcal{C} be a collection of disjoint nonsingleton subsets of B\mathcal{B}. For each block C∈CC \in\mathcal{C} require that every address is performed and that their prime values are equal. Denote this event by ECE_C. Then the exact distinctness indicator is

1distinct performed birth values=∑C∏C∈C(−1)∣C∣−1(∣C∣−1)!∏C∈C1EC.(78)\mathbf{1}_{\text{distinct performed birth values}} = \sum_{\mathcal{C}} \prod_{C\in\mathcal{C}} (-1)^{|C|-1}(|C|-1)! \prod_{C\in\mathcal{C}}\mathbf{1}_{E_C}. \tag*{(78)}

One proof expands the right side as the sum of signs of all permutations of the performed addresses which preserve their values: a nontrivial cycle on CC has sign (−1)∣C∣−1(-1)^{|C|-1} and there are (∣C∣−1)!(|C|-1)! such cycles. In each equal-value class of size at least two the signs of permutations sum to zero, and for a singleton they sum to one. Unperformed addresses cannot occur in a nontrivial cycle, which is exactly enforced by the flags.

Define the rank

r(C)=∑C∈C(∣C∣−1).(79)r(\mathcal{C}) = \sum_{C\in\mathcal{C}} (|C|-1). \tag*{(79)}

A rank-rr collection involves at most 2r2r addresses. Its total absolute coefficient mass, summed over all collections of that rank, is at most

Qbirth2r=LO(r).(80)Q_{\mathrm{birth}}^{2r} = L^{O(r)}. \tag*{(80)}

Indeed the coefficients count permutations with that cycle rank, and each such permutation has a fixed canonical decomposition into rr transpositions. Its ordered list of transpositions, drawn from at most Qbirth2Q_{\mathrm{birth}}^2 possibilities, determines the permutation.

We first bound a fixed system of equality constraints absolutely. In each block choose its earliest address as pivot, and condition chronologically on the entire earlier history and all known pivot values. Each nonpivot performed birth must draw one prescribed prime. By (38) this costs an atom at most

Δ:=sup⁡i,pνi(p)≤CKe−L.1.(81)\Delta:= \sup_{i,p} \nu_i(p) \le C_K e^{-L^{.1}}. \tag*{(81)}

If its prescribed value belongs to a wrong group, the term is zero. No later endpoint conditioning is imposed in this absolute bound: we drop the final return requirement and use weighted row iteration. Thus dependence of intermediate positions on the earlier pivot values cannot change the atom cost.

Here is the precise uniform estimate justifying that iteration. For an edge with ss prescribed nonpivot fresh values, fix all fresh values first. The possible target positions still number O(1)O(1) in the raw part or O(Y)O(Y) in the comparison part. The latter count is compensated by (58). Consequently its weighted absolute row is at most

CKBK(1+L3A0)Δs,(82)C_K B^K(1+L^{3A_0})\Delta^s, \tag*{(82)}

uniformly in the input state and all fixed values. The remaining fresh values have bounded total mass, and all local restrictions may be discarded for this upper bound.

For a ghost in group ii, set ci=v0ηj′Vi/ρc_i=v_0\eta_j'V_i/\rho. If ss specified ordered coordinates of its new batch are prescribed, its birth row sum is at most

∑l≥scilΔssil−sl!≤(ciΔ)sexp⁡(cisi).(83)\sum_{l\ge s}\frac{c_i^l\Delta^s s_i^{l-s}}{l!}\le(c_i\Delta)^s\exp(c_i s_i). \tag*{(83)}

The actual requirement that the batch reaches its largest named index can only reduce this sum. Within a batch the earlier pivot coordinates are integrated before the prescribed later ones, which gives the same estimate. Deletion and survival have weighted factor at most one by (54). Anchors add no factor p+1p+1, as their measure is counting measure. Initial slots have the same single-atom estimate. Thus for each fixed rank-rr constraint system, the absolute boundary contribution is at most

LC(N+1)(CKe−L.1)r,(84)L^{C(N+1)}(C_K e^{-L^{.1}})^r, \tag*{(84)}

for a fixed C=C(K,A0,q)C=C(K,A_0,q). The boundary weight is one because memory is empty; after dropping return, its terminal indicator is at most the positive Schur weight, so the chronological row bounds apply.

Choose, after CC is fixed,

r0=⌈C′Nlog⁡LL.1⌉,(85)r_0=\left\lceil\frac{C' N\log L}{L^{.1}}\right\rceil, \tag*{(85)}

with C′C' sufficiently large. Combining (80) and (84), and using log⁡Qbirth=O(log⁡L)\log Q_{\mathrm{birth}}=O(\log L), bounds all ranks r≥r0r\ge r_0 by

LC(N+1)∑r≥r0(CKQbirth2e−L.1)r≤L−(E0+2)N.(86)L^{C(N+1)}\sum_{r\ge r_0}(C_K Q_{\mathrm{birth}}^2e^{-L^{.1}})^r\le L^{-(E_0+2)N}. \tag*{(86)}

The series is geometric for large xx. For example its ratio is at most e−L.1/2e^{-L^{.1}/2}, so choosing C′/2>C+E0+3C'/2>C+E_0+3 suffices after enlarging the threshold for xx.

For ranks r<r0r<r_0, use signed contraction on most edges. If C={a0,a1,…,as}C=\{a_0,a_1,\ldots,a_s\} is one block, with earliest address a0a_0, write its equality event by ordinary circle orthogonality as

1EC=(∏a∈C1a performed)∫Ts∏h=1se(θh(pah−pa0)) dθ1⋯dθs.(87)\mathbf{1}_{E_C}=\left(\prod_{a\in C}\mathbf{1}_{a\ \mathrm{performed}}\right)\int_{\mathbb{T}^s}\prod_{h=1}^{s}e\bigl(\theta_h(p_{a_h}-p_{a_0})\bigr)\,d\theta_1\cdots d\theta_s. \tag*{(87)}

When an address is unperformed the integrand is interpreted as zero; no value is assigned to an unperformed birth. After the circle variables are fixed, every phase is attached only to the operation where its prime was created. In particular the pivot receives the single local multiplier e(−pa0∑hθh)e(-p_{a_0}\sum_h\theta_h). Its value need not be remembered through later operations. Flags are local choice restrictions of modulus at most one. An edge flag tests whether its slot is freshly drawn rather than promoted; it does not distinguish old promotion indices. A ghost flag tests whether its fresh batch reaches the specified address. Thus these flags and birth-value phases satisfy the old-memory equivariance condition stated after (52). For fixed circle variables, backward induction through the later operations gives a symmetric continuation function. Averaging the ghost-batch multiplier over its new coordinates consequently leaves the full integral unchanged, without increasing its absolute kernel, changing its factor 1/l!1/l!, or modifying any later operation.

It follows that a rank-rr system modifies at most 2r2r operations, and hence at most 2r2r edges. Initial phases change bb by a bounded multiplier and preserve (46). Every ghost, modified or not, has norm at most CKC K; a modified edge has norm at most LCL^C by (59); all other edges retain (76). For each fixed set of phases the boundary matrix element is therefore at most

CKN+1∥b∥2L−G(N−2r)L2Cr.(88)C_K^{N+1}\lVert b\rVert^2 L^{-G(N-2r)}L^{2Cr}. \tag*{(88)}

Integration of phases does not increase this bound. As r0/N=O(log⁡L/L.1)=o(1)r_0/N=O(\log L/L^{.1})=o(1), summing (88) with (80) gives

∑r<r0Qbirth2rCKN+1∥b∥2L−G(N−2r)+2Cr≤L(−G+o(1))N.(89)\sum_{r<r_0} Q_{\mathrm{birth}}^{2r} C_K^{N+1}\lVert b\rVert^2 L^{-G(N-2r)+2Cr}\le L^{(-G+o(1))N}. \tag*{(89)}

This remains true even though CC depends on KK: KK was fixed before letting xx grow, and its crude exponent is paid on only o(N)o(N) edges.

Equations (86) and (89) bound (48) with exact global birth distinctness, uniformly in all archimedean root variables. Multiply back the baseline, whose modulus is at most one. The root integral has bounded mass. The error (50) and the root-replacement error are OA(L−AN)O_A(L^{-AN}) for arbitrary fixed AA. Because G>E0+3G>E_0+3, these margins absorb all constants and prove (35). The passage from this moment to (25) was proved in the preceding section.

Removing goodness and stripping the pads

We complete the second assertion of Proposition 4.1. First remove the good-state projections from the endpoint pairing. The part where the source is bad, and separately the part where the target is bad, may be majorized absolutely; bound the endpoint coefficients by their log-power suprema. Apply the single-edge version of the root replacement. In the independent-line model, a shared prime requires the same line at both positions, or contributes zero. Every distinct prime in the union of the two active lists is charged exactly once by 1/(p+1)1/(p+1). The ban makes the union consist of M+1M+1 distinct primes per group. Its list normalization is Vi−(M+1)V_i^{-(M+1)}, so summing these probabilities over ordered union lists is at most

∏i=1K(∑p∈Pi1Vi(p+1))M+1≤1.(90)\prod_{i=1}^{K}\left(\sum_{p\in\mathcal{P}_i}\frac{1}{V_i(p+1)}\right)^{M+1}\le1. \tag*{(90)}

We have dropped any incompatibilities for an upper bound. Damping is also at most one on the physical states and may be discarded.

For fixed labels, (57) counts O(1)O(1) raw targets; the comparison target count is O(Y)O(Y) and is compensated by (58). For every fixed source list, the set of root ratios failing either good-state test has measure O(e−cL.1)O(e^{-cL^{.1}}) by Lemma 3.2. This estimate is uniform in the other root coordinates. The preceding target counts hold for every ratio, so they may be integrated over that bad set. Reversal proves the identical assertion for failure of target goodness. Equation (4.56) then gives the normalized pairing error

O(UVLC1+3A0e−cL.1)+OA(UVL−A),(91)O(UVL^{C_1+3A_0}e^{-cL^{.1}})+O_A(UVL^{-A}), \tag*{(91)}

where C1C_1 comes from the endpoint suprema. The root-replacement error can here be made smaller than every log power by its single-edge estimate; this also handles the indicators of goodness failure. Therefore removing the projections costs OA(UVL−A)O_A(UVL^{-A}) for every fixed AA.

The endpoint vectors are invariant under permutations of their lists. Hence the two slot symmetrizations disappear when the pairing is expanded. Write the ordered pads as pi,1,…,pi,Jp_{i,1},\ldots,p_{i,J} in group ii, with total product DD. The remaining source and target slots have products bb and aa. The common state normalization and the fresh-target normalization are ∏iVi−(J+2)\prod_i V_i^{-(J+2)}. Divide the pairing for the dyad d0d_0 by d0d_0. Its scalar multiplier becomes 1/D1/D, and the exact measure identity is

1D∏i=1KVi−(J+2)=(∏i=1K∏h=1Jμi(pi,h))∏i=1KVi−2.(92)\frac{1}{D}\prod_{i=1}^{K}V_i^{-(J+2)} = \left(\prod_{i=1}^{K}\prod_{h=1}^{J}\mu_i(p_{i,h})\right)\prod_{i=1}^{K}V_i^{-2}. \tag*{(92)}

Thus the pads are independent ordered harmonic draws, and the remaining factors are exactly the two unshared-list normalizations in the expanded square. There is no factorial or MM-dependent normalization left over: the pads occupy prescribed ordered shared slots and the unshared label occupies its prescribed remaining slot.

The change of variables is P=(Dbm,s)P=(Dbm,s), Q=(Dar,n)Q=(Dar,n), so det⁡(P,Q)/D=bmn−ars\det(P,Q)/D=bmn-ars. The support and roughness properties of the endpoint vectors give the original coefficients and make the physical damping equal to one. Primitivity and the box conditions are automatic on their nonzero support, as established in the preceding section. On summing the ordinary pad dyads every ordered pad tuple occurs once. Therefore, apart from the pad collision restrictions, (4.58) gives exactly the signed difference between the determinant condition and Hm\mathcal{H}_m in the expanded square.

For fixed disjoint unshared products a,ba,b, independent pad draws violate pad distinctness or the ban with probability at most

CKM2sup⁡i,pμi(p)≪KM2e−L.1.(93)C_KM^2\sup_{i,p}\mu_i(p)\ll_K M^2e^{-L^{.1}}. \tag*{(93)}

This follows by a union bound over pairs of pad addresses and over coincidences with the two fixed unshared primes in each group; for two independent draws their equality probability is at most the largest atom. To remove this probability loss from a signed sum, we need absolute bounds for each of its two parts.

The absolute raw expanded square is

≪KXYLC2,(94)\ll_K XYL^{C_2}, \tag*{(94)}

with C2C_2 depending on the coefficient bounds and fixed divisor moments, independently of KK. Indeed for each pair a,ba,b its original hh representation has h≪X/Yh\ll X/Y and contributes at most a log-power coefficient bound times ∑h≪X/Yτ(1+ah)τ(1+bh)\sum_{h\ll X/Y}\tau(1+ah)\tau(1+bh). Cauchy and the divisor moment estimate from the preceding section bound this by (X/Y)LC2(X/Y)L^{C_2}. There are O(Y2)O(Y^2) possible numeric products a,ba,b, and their normalization costs only CKC_K.

For the absolute comparison sum, fix a,b≍Ya,b\asymp Y. If k=mnk=mn and l=rsl=rs, the cutoff imposes

∣bk−al∣≪Y,k,l≪X.|bk-al|\ll Y,\qquad k,l\ll X.

For each kk there are O(1)O(1) choices of ll, and conversely, because a,b≍Ya,b\asymp Y. Thus Cauchy on this bounded-degree relation and the ordinary divisor second moment give

∑k,l≪X∣bk−al∣≪Yτ(k)τ(l)≪∑k≪Xτ(k)2≪XLC3.(95)\sum_{\substack{k,l\ll X\\ |bk-al|\ll Y}} \tau(k)\tau(l)\ll\sum_{k\ll X}\tau(k)^2\ll XL^{C_3}. \tag*{(95)}

The endpoint coefficients cost a fixed log power, and (4.24) contributes O(L3A0/Y)O(L^{3A_0}/Y). Summing the O(Y2)O(Y^2) possible a,ba,b proves

≪KXYLC4+3A0,(96)\ll_K XYL^{C_4+3A_0}, \tag*{(96)}

where C3,C4C_3,C_4 are again independent of KK. Multiplying (4.60) and (4.62) by (4.59) gives OA(XYL−A)O_A(XYL^{-A}) for every fixed AA. This strips all pads.

The comparison estimate also allows the two unshared products to have a common prime. The number of pairs a,b≲Ya,b\lesssim Y sharing any group prime is at most

CKY2∑p∈⋃iPi1p2≪KY2e−L.1.(97)C_KY^2\sum_{p\in\bigcup_i\mathcal{P}_i}\frac{1}{p^2}\ll_K Y^2e^{-L^{.1}}. \tag*{(97)}

For fixed pp there are at most O(Y/p)O(Y/p) possible numeric products divisible by pp, and each product determines its prime list. Apply (4.61) to each such pair and then (4.24); their total is again smaller than XYXY times every fixed negative log power.

Finally, if C5C_5 denotes the log-power loss in (3.20), its endpoint calculation gives C5C_5 depending only on the fixed coefficient bounds, independently of KK. Since UV=d0YXUV=d_0YX, division by d0d_0 bounds each pad dyad by XYL−E0+C5XYL^{-E_0+C_5}. There are OK(L)O_K(L) pad dyads, so their sum is

OK(XYL−E0+C5+1)+OA(XYL−A).(98)O_K(XYL^{-E_0+C_5+1})+O_A(XYL^{-A}). \tag*{(98)}

Given a desired replacement precision AA, choose E0>A+C5+2E_0>A+C_5+2 first, then choose A0A_0, then KK as above. All exponentially small errors have already been bounded after these parameters are fixed and impose no further condition on E0E_0. This proves the replacement assertion of Proposition 4.1 and leaves the unrestricted major term for the next section.

The major term of the determinant estimate

We retain the notation of Theorem 3.1, in particular X=HmHn≍xX=H_mH_n\asymp x and W=exp⁡(L0.24)W=\exp(L^{0.24}). The exponent A0A_0 defining M\mathcal{M} is fixed throughout this section. Constants may depend on KK and on its fixed prime bands; the powers of LL used before choosing KK will not depend on KK.

Proposition 5.1 (The determinant major term). For every fixed D>0D>0, the unrestricted major sum

QYmaj=(∏i=1KVi−2)∑a,b,m,n,r,sη(a/Y)η(b/Y)αmαr‾βnβs‾HM(bmn−ars;a,b)(99)Q_Y^{\mathrm{maj}}=\left(\prod_{i=1}^{K}V_i^{-2}\right)\sum_{a,b,m,n,r,s}\eta(a/Y)\eta(b/Y)\alpha_m\overline{\alpha_r}\beta_n\overline{\beta_s}\mathcal{H}_{\mathcal{M}}(bmn-ars;a,b) \tag*{(99)}

satisfies

∣QYmaj∣≪DXYL−D.(100)\left|Q_Y^{\mathrm{maj}}\right|\ll_D XYL^{-D}. \tag*{(100)}

*Here a,ba,b run independently over products of one prime from each Pi\mathcal{P}_i, with no disjointness condition, and HM\mathcal{H}_{\mathcal{M}} denotes the kernel in (3.11). The estimate is uniform in the coefficient intervals, in ∣v∣≤LC|v|\leq L^C, and in the admissible scale YY.

The major-arc decomposition will separate three factors: the centered prime-minus-rough coefficient α\alpha, the arbitrary rough coefficient β\beta, and the product of small prime labels. Uniform cancellation of the first factor alone does not control an integral over a frequency range of length comparable to XX. We isolate one small prime label. Where its polynomial is small, a mean-square bound for the product of the two long polynomials suffices; where it is large, the frequencies are sparse enough to control the energy of the β\beta polynomial. The next two lemmas supply the uniform cancellation and the sparse-frequency bound, respectively. All characters below are extended by zero away from the units.

Lemma 5.2 (The long coefficient). Fix A0,B,A>0A_0,B,A>0. For every character χ\chi of modulus at most LA0L^{A_0}, put

Mα(t)=1Hm∑mαmχ(m)mit.(101)M_\alpha(t)=\frac{1}{H_m}\sum_m \alpha_m\chi(m)m^{it}. \tag*{(101)}

Then, uniformly for ∣t∣≤2XLB|t|\le2XL^B,

∣Mα(t)∣≪AL−A.|M_\alpha(t)|\ll_A L^{-A}.

The assertion also holds with conjugated coefficients and characters.

Proof. Write u=t+vu=t+v and let I⊂[Hm,2Hm]I\subset[H_m,2H_m] be the coefficient interval. We give the approximation of rough numbers explicitly. Set P(W)={p:p≤W}\mathcal{P}(W)=\{p:p\le W\} and SW=∑p≤Wp−1=0.24log⁡L+O(1)S_W=\sum_{p\le W}p^{-1}=0.24\log L+O(1). For an integer r≍clog⁡Lr\asymp c\log L, define

Tr(m)=∑d∣m, d squarefreep∣d⇒p≤W, #{p:p∣d}≤rμ(d).T_r(m)=\sum_{\substack{d\mid m,\ d\ \mathrm{squarefree}\\ p\mid d\Rightarrow p\le W,\ \#\{p:p\mid d\}\le r}}\mu(d).

The two consecutive Bonferroni truncations bracket 1P−(m)>W\mathbf{1}_{P_-(m)>W}. Their difference is supported on products of r+1r+1 distinct primes. Consequently, on every subinterval I′⊂II'\subset I,

∑m∈I′∣1P−(m)>W−Tr(m)∣≪HmSWr+1(r+1)!+Wr+1.(102)\sum_{m\in I'}\left|\mathbf{1}_{P_-(m)>W}-T_r(m)\right|\ll H_m\frac{S_W^{r+1}}{(r+1)!}+W^{r+1}. \tag*{(102)}

where changing rr by one, if necessary, is harmless. This follows by summing 1d∣m\mathbf{1}_{d\mid m} over the first omitted degree, using #{m∈I′:d∣m}=∣I′∣/d+O(1)\#\{m\in I':d\mid m\}=|I'|/d+O(1). Since r!≥(r/e)rr!\ge(r/e)^r, choosing cc sufficiently large makes the first term O(HmL−D′)O(H_mL^{-D'}) for any prescribed fixed D′D'. The second is xo(1)x^{o(1)}, since rlog⁡W=O(L0.24log⁡L)r\log W=O(L^{0.24\log L}). Moreover, the absolute reciprocal coefficient mass of every truncation satisfies

∑d used1d≤∏p≤W(1+p−1)≪L0.24.(103)\sum_{\text{$d$ used}}\frac{1}{d}\le\prod_{p\le W}(1+p^{-1})\ll L^{0.24}. \tag*{(103)}

In particular, increasing the depth to improve the error does not increase the exponent in this bound.

First consider large ∣u∣|u|. For each surviving divisor dd with (d,k)=1(d,k)=1, where kk is the character modulus, write m=dzm=dz and split zz into residues modulo kk. Its scale is N′=Hm/d=xΩ(1)N'=H_m/d=x^{\Omega(1)}. For 1≤∣u∣≤T∗=exp⁡(L/(log⁡L)2)1\le|u|\le T_*=\exp(L/(\log L)^2), summation against an integral on a progression gives

∑z∈J⊂[N′,2N′]z≡a(modk)ziu=1k∫Jyiu dy+O(1+∣u∣),∣∫Jyiu dy∣≪N′∣u∣.(104)\begin{aligned} \sum_{\substack{z\in J\subset[N',2N']\\ z\equiv a\pmod{k}}}z^{iu} &=\frac{1}{k}\int_J y^{iu}\,dy+O(1+|u|),& \left|\int_J y^{iu}\,dy\right|&\ll\frac{N'}{|u|}. \tag*{(104)} \end{aligned}

Indeed the total variation of yiuy^{iu} on this dyad is O(∣u∣)O(|u|); the integral formula follows by evaluating y1+iu/(1+iu)y^{1+iu}/(1+iu) at the endpoints. The error, after summing the at most kk residues and all d≤Wrd\le W^r, is x−cδx^{-c\delta} after division by HmH_m, uniformly up to T∗T_*. At T∗≤∣u∣<x2T_* \le|u| < x^2, apply Lemma 2.4: its lower scale condition holds because N′≥xδ/2N' \ge x^{\delta/2} for sufficiently large xx, and its lower frequency condition is weaker than ∣u∣≥T∗|u| \ge T_*. Together with (5.6), these bounds and partial summation for 1/log⁡m1/\log m imply

∣1HmV(W)∑m∈IP−(m)>Wχ(m)miulog⁡m∣≪L−D′+{LC0(∣u∣−1+x−cδ),1≤∣u∣≤T∗,LC0exp⁡(−L/(log⁡L)C3),T∗≤∣u∣<x2.(105)\left|\frac{1}{H_mV(W)}\sum_{\substack{m\in I\\P^-(m)>W}}\frac{\chi(m)m^{iu}}{\log m}\right| \ll L^{-D'}+ \begin{cases} L^{C_0}(|u|^{-1}+x^{-c\delta}), & 1\le|u|\le T_*,\\ L^{C_0}\exp(-L/(\log L)^{C_3}), & T_*\le|u|<x^2. \end{cases} \tag*{(105)}

Here C0C_0 is fixed after A0,δA_0,\delta; it need not grow with the chosen Bonferroni depth. The factor 1/V(W)≪L0.241/V(W)\ll L^{0.24} and the bounded variation of 1/log⁡m1/\log m have been included in C0C_0.

For the prime part, choose fixed 0<τ<δ0<\tau<\delta and 1−δ<η<11-\delta<\eta<1. The scale conditions imply xτ/2≤Hm≤xηx^{\tau/2}\le H_m\le x^\eta for large xx. Lemma 2.3, followed by partial summation to replace p−1p^{-1} by Hm−1H_m^{-1}, gives an arbitrary negative power of LL when LB0′≤∣u∣≤x2L^{B'_0}\le|u|\le x^2. Choose B2>C+2B_2>C+2 sufficiently large after A,A0,BA,A_0,B and the exponent C0C_0. Then LB2≤∣t∣≤2XLBL^{B_2}\le|t|\le2XL^B implies ∣u∣≍∣t∣|u|\asymp|t|, ∣u∣≥LB0′|u|\ge L^{B'_0}, and ∣u∣<x2|u|<x^2. The prime estimate and (5.8) prove (5.4) on this range.

It remains to treat ∣t∣≤LB2|t|\le L^{B_2}, hence ∣u∣≤2LB2|u|\le2L^{B_2}. Choose the counting precision D′D' after B2B_2. Lemma 2.5 gives, on every subinterval,

∑m∈I′χ(m)1P−(m)>W=1χ principalV(W)∣I′∣+O(HmL−D′).(106)\sum_{m\in I'}\chi(m)\mathbf{1}_{P^-(m)>W} =\mathbf{1}_{\chi\ {\rm principal}}V(W)|I'|+O(H_mL^{-D'}). \tag*{(106)}

The principal character equals one on rough numbers because all primes dividing kk are at most WW; for nonprincipal characters the count has zero main term by complete-period cancellation. For primes, Lemma 2.1 and character orthogonality give, to any prescribed precision,

∑p∈I′χ(p)=1χ principal∫I′dylog⁡y+O(HmL−D′).\sum_{p\in I'}\chi(p)=\mathbf{1}_{\chi\ {\rm principal}}\int_{I'}\frac{dy}{\log y}+O(H_mL^{-D'}).

Partial summation introduces at most a fixed power of LL on this low-frequency range. Dividing the rough main term by V(W)log⁡yV(W)\log y produces exactly dy/log⁡ydy/\log y; its twisted main term therefore cancels the prime main term. Increasing D′D' gives (5.4). Complex conjugation changes only the signs of the real frequencies and the character, so the proof is unchanged.

Lemma 5.3 (A small prime factor and sparse frequencies). Let PP satisfy cL0.1≤log⁡P≤C1L0.2cL^{0.1}\le\log P\le C_1L^{0.2}, and let P\mathcal{P} be any set of primes in [P,2P][P,2P]. For any real ξ\xi and any character χ\chi, put

Ps(t)=P−1∑p∈Pχ(p)pi(t+ξ).P_s(t)=P^{-1}\sum_{p\in\mathcal{P}}\chi(p)p^{i(t+\xi)}.

Let Nβ(t)=Hn−1∑nβnχn(n)nitN_\beta(t)=H_n^{-1}\sum_n\beta_n\chi_n(n)n^{it}, with ∣βn∣≤LC|\beta_n|\le L^C on [Hn,2Hn][H_n,2H_n] and xδ≤Hn≪xx^\delta\le H_n\ll x. For every fixed B,As>0B,A_s>0, the unit intervals meeting

{t:∣t∣≤2XLB, ∣Ps(t)∣>L−As}\{t: |t|\le2XL^B,\ |P_s(t)|>L^{-A_s}\}

form a family I\mathcal{I} satisfying

∣I∣≤exp⁡(O(L0.91)),∑I∈Isup⁡t∈I∣Nβ(t)∣2≪L2C+1.(107)|\mathcal{I}|\le\exp(O(L^{0.91})),\qquad \sum_{I\in\mathcal{I}}\sup_{t\in I}|N_\beta(t)|^2\ll L^{2C+1}. \tag*{(107)}

Both assertions are uniform in ξ\xi and in the prime set.

Proof. Put Z=4XLBZ = 4XL^{B} and j=⌊log⁡Z/log⁡(2P)⌋j = \lfloor\log Z/\log(2P)\rfloor. Write Ps(t)j=∑n≤(2P)jcnnitP_s(t)^j = \sum_{n\le(2P)^j} c_n n^{it}. For each product nn, unique factorization bounds the number of ordered prime tuples producing it by j!j!. Hence

∑n∣cn∣2≤j!P−2j∣P∣j≤j!(C2/P)j.(108)\sum_n |c_n|^2 \le j!P^{-2j}|P|^j \le j!(C_2/P)^j. \tag*{(108)}

This is a coefficient estimate for the actual growing degree jj; no fixed-divisor-moment constant is used for it.

Choose an exceeding point from each member of I\mathcal{I}, and split these points into three classes according to the integer left endpoint modulo three. Each class is separated by at least one. The separated mean-square estimate of Lemma 2.2, whose constant is independent of the coefficients, gives

∣I∣≪ZLO(1)L2Asjj!(C2/P)j.(109)|\mathcal{I}| \ll ZL^{O(1)}L^{2A_sj}j!(C_2/P)^j. \tag*{(109)}

Since log⁡Z−jlog⁡P≪log⁡P+j\log Z-j\log P\ll\log P+j and j≪L0.9j\ll L^{0.9}, the logarithm of its right-hand side is at most

O(log⁡P+jlog⁡(j+1)+Asjlog⁡L+log⁡L)=O(L0.9log⁡L+L0.2)=O(L0.91).O(\log P+j\log(j+1)+A_sj\log L+\log L)=O(L^{0.9}\log L+L^{0.2})=O(L^{0.91}).

The factor piξp^{i\xi} affects no coefficient absolute value, so this estimate is uniform for all real ξ\xi.

For the second assertion choose a maximizing point tIt_I on the closure of each unit interval, and again use three separated classes. In one class consider the evaluation matrix on the full integer interval [Hn,2Hn][H_n,2H_n]. Its Gram entries are

GI,J=∑Hn≤n≤2Hnni(tI−tJ).G_{I,J}=\sum_{H_n\le n\le2H_n} n^{i(t_I-t_J)}.

Set T∗=exp⁡(L/(log⁡L)2)T_*=\exp(L/(\log L)^2). The diagonal is O(Hn)O(H_n). For 1≤∣Δ∣≤T∗1\le|\Delta|\le T_*, (104) with modulus one gives

Hn−1∣GI,J∣≪∣Δ∣−1+x−cδ.(110)H_n^{-1}|G_{I,J}|\ll|\Delta|^{-1}+x^{-c\delta}. \tag*{(110)}

For T∗<∣Δ∣≤4XLB+2<4x3T_*<|\Delta|\le4XL^{B}+2<4x^3, Lemma 2.4 gives

Hn−1∣GI,J∣≪exp⁡(−L/(log⁡L)C3).H_n^{-1}|G_{I,J}|\ll\exp(-L/(\log L)^{C_3}).

All its scale hypotheses hold because Hn≥xδH_n\ge x^\delta and Hn≪xH_n\ll x. Separation implies that the sum of 1/∣Δ∣1/|\Delta| in any row is O(log⁡(2Z))=O(L)O(\log(2Z))=O(L). The other row sums tend to zero, since

∣I∣x−cδ=o(1),∣I∣exp⁡(−L/(log⁡L)C3)=o(1).|\mathcal{I}|x^{-c\delta}=o(1),\qquad|\mathcal{I}|\exp(-L/(\log L)^{C_3})=o(1).

Thus the absolute Gram row sums are O(HnL)O(H_nL), and the same is true of its operator norm. The coefficient vector of NβN_\beta has squared norm O(L2C/Hn)O(L^{2C}/H_n). Applying the Gram bound and adding the three classes proves (5.10).

Proof of Proposition 5.1. For large xx the arcs are disjoint: distinct reduced rationals of denominator at most LA0L^{A_0} are at distance at least L−2A0L^{-2A_0}, whereas their radii are 2LA0/Y2L^{A_0}/Y and Y≥exp⁡(cL0.1)Y\ge\exp(cL^{0.1}). On an arc write θ=h/k+λ/Y\theta=h/k+\lambda/Y, with ∣λ∣≤2LA0|\lambda|\le2L^{A_0} and dθ=dλ/Yd\theta=d\lambda/Y. Every variable in (5.1) is a unit modulo kk: the group primes and all prime factors of the rough variables exceed kk. The function

(a,b,m,n,r,s)⟼e(h(bmn−ars−b+a)/k)(a,b,m,n,r,s)\longmapsto e\left(h(bmn-ars-b+a)/k\right)

on (Z/kZ)×6(\mathbb{Z}/k\mathbb{Z})^{\times6} has an exact Fourier expansion in products of six multiplicative characters. Each Fourier coefficient has absolute value at most one, so their total absolute mass is at most φ(k)6≤k6\varphi(k)^6\le k^6. There are O(L2A0)O(L^{2A_0}) arcs. It is therefore enough to estimate each separated character term, uniformly in λ\lambda, with arbitrarily large logarithmic saving.

Set w=bmn/(YX)w=bmn/(YX) and w′=ars/(YX)w'=ars/(YX). Both lie in a fixed compact subinterval of (0,∞)(0,\infty). Apart from the factors e(−λb/Y)e(λa/Y)e(-\lambda b/Y)e(\lambda a/Y), the remaining common kernel is

ψλ(X(w−w′)),ψλ(y)=ψ(y)e(λy).\psi_\lambda(X(w-w')), \qquad\psi_\lambda(y)=\psi(y)e(\lambda y).

Here is the Mellin separation with its normalization. In coordinates z=log⁡w′z=\log w' and ζ=Xlog⁡(w/w′)\zeta=X\log(w/w'), it is

ψλ(Xez(eζ/X−1)).\psi_\lambda\left(Xe^z(e^{\zeta/X}-1)\right).

Insert fixed smooth cutoffs equal to one on the possible values of w,w′w,w'. The resulting function is supported on a fixed compact set of (ζ,z)(\zeta,z): the support of ψ\psi forces ∣ζ∣≪1|\zeta|\ll1. Its derivatives of order r+sr+s are Or,s(LB0(r+s))O_{r,s}(L^{B_0(r+s)}) for some fixed B0B_0 chosen after A0A_0. Fourier inversion, with inverse kernel ei(vζ+s′z)e^{i(v\zeta+s'z)}, and the change of variable t=Xvt=Xv consequently give the exact identity

ψλ(X(w−w′))=1X∬R2Fλ(t/X,s′)wit(w′)i(s′−t) dt ds′(111)\psi_\lambda(X(w-w'))=\frac{1}{X}\iint_{\mathbb{R}^2}F_\lambda(t/X,s')w^{it}(w')^{i(s'-t)}\,dt\,ds' \tag*{(111)}

on the coefficient support. Repeated integration by parts yields, for every fixed integer j≥1j\geq1,

∣Fλ(v,s′)∣≪j(1+∣v∣/LB0)−j(1+∣s′∣/LB0)−j.|F_\lambda(v,s')|\ll_j (1+|v|/L^{B_0})^{-j}(1+|s'|/L^{B_0})^{-j}.

Define the normalized endpoint polynomials

Bλ(t)=1Y∑bη(b/Y)e(−λb/Y)(∏iVi−1)χb(b)bit,B_\lambda(t)=\frac{1}{Y}\sum_b \eta(b/Y)e(-\lambda b/Y)\left(\prod_i V_i^{-1}\right)\chi_b(b)b^{it},
Nβ(t)=1Hn∑nβnχn(n)nit,Q(t)=Bλ(t)Mα(t)Nβ(t).(112)N_\beta(t)=\frac{1}{H_n}\sum_n\beta_n\chi_n(n)n^{it},\qquad Q(t)=B_\lambda(t)M_\alpha(t)N_\beta(t). \tag*{(112)}

The other endpoint gives a polynomial Q′Q' of the same type with conjugations and sign changes. Each unnormalized endpoint sum is YXYX times its polynomial. Combining these two factors with the X−1X^{-1} in (111) and the arc measure Y−1Y^{-1} gives

(YX)2⋅X−1⋅Y−1=XY.(YX)^2\cdot X^{-1}\cdot Y^{-1}=XY.

The additional factor (YX)−is′(YX)^{-is'} has absolute value one. Thus, after division by XYXY, the separated expression is an integral of Fλ(t/X,s′)Q(t)Q′(s′−t)F_\lambda(t/X,s')Q(t)Q'(s'-t), followed by the λ\lambda integral.

The bound ∣Bλ(t)∣≪1|B_\lambda(t)|\ll1 follows directly from b≍Yb\asymp Y and

∑b1b∏iVi−1≤∏i(Vi−1∑p∈Pip−1)=1.\sum_b\frac{1}{b}\prod_i V_i^{-1}\leq\prod_i\left(V_i^{-1}\sum_{p\in\mathcal{P}_i}p^{-1}\right)=1.

The coefficient of index u=mnu=mn in MαNβM_\alpha N_\beta is O(LCτ(u)/X)O(L^{C}\tau(u)/X) and is supported on u≍Xu\asymp X. Lemma 2.2 therefore gives a squared coefficient norm O(LC4/X)O(L^{C_4}/X) and, on every real interval II of length Z≥1Z\geq1,

∫I∣Mα(t)Nβ(t)∣2 dt≪LC4(1+Z/X).(113)\int_I |M_\alpha(t)N_\beta(t)|^2\,dt\ll L^{C_4}(1+Z/X). \tag*{(113)}

The fixed exponent C4C_4 depends only on the original coefficient bounds. The estimate is uniform in translates of II, by twisting the coefficients, and also bounds the mean square of QQ and Q′Q'.

Fix B1>B0+1B_1>B_0+1. Equations (5.16) and (113) permit us, to arbitrary logarithmic precision, to restrict to

∣s′∣≤LB1,∣t∣≤XLB1.|s'|\le L^{B_1},\qquad|t|\le XL^{B_1}.

For completeness, on a dyadic tt interval of length ZZ, Cauchy’s inequality bounds the integral of ∣Q(t)Q′(s′−t)∣|Q(t)Q'(s'-t)| by LC4(1+Z/X)L^{C_4}(1+Z/X), uniformly in s′s'. For Z≥XLB1Z\ge XL^{B_1} multiply this by (1+Z/(XLB0))−j(1+Z/(XL^{B_0}))^{-j} and sum the geometric tail. For the s′s' tail, the full weighted tt integral is O(LC4+B0+1)O(L^{C_4+B_0+1}), uniformly in s′s', and

∫∣s′∣>LB1(1+∣s′∣/LB0)−j ds′≪jLB0−(B1−B0)(j−1).\int_{|s'|>L^{B_1}}\left(1+|s'|/L^{B_0}\right)^{-j}\,ds'\ll_j L^{B_0-(B_1-B_0)(j-1)}.

Taking jj large makes both errors as small as required, including the character and arc costs. This argument also covers translated frequencies s′−ts'-t without a pointwise tail assumption.

On (5.20), Cauchy’s inequality in tt reduces the claim to proving, for every prescribed A′>0A'>0,

∫∣t∣≤2XLB1∣Bλ(t)Mα(t)Nβ(t)∣2 dt≪A′L−A′.(114)\int_{|t|\le2XL^{B_1}}|B_\lambda(t)M_\alpha(t)N_\beta(t)|^2\,dt\ll_{A'}L^{-A'}. \tag*{(114)}

Indeed the retained s′s' integral has length 2LB12L^{B_1} and all remaining character and λ\lambda costs are fixed powers of LL.

Write b=pb′b=pb' with p∈P1p\in\mathcal{P}_1 and split this prime band into O(L)O(L) ordinary dyads [P,2P)[P,2P). The cutoff on bb permits restriction to Y/(2P)≤b′≤4Y/PY/(2P)\le b'\le4Y/P. Mellin inversion of fλ(y)=η(y)e(−λy)f_\lambda(y)=\eta(y)e(-\lambda y) gives

fλ(y)=∫Rf^λ(ξ)yiξ dξ,∫R∣f^λ(ξ)∣ dξ≪LC5,f_\lambda(y)=\int_{\mathbb{R}}\widehat f_\lambda(\xi)y^{i\xi}\,d\xi,\qquad\int_{\mathbb{R}}|\widehat f_\lambda(\xi)|\,d\xi\ll L^{C_5},

where C5C_5 is fixed after A0A_0. This last bound follows, for example, by twice integrating by parts outside ∣ξ∣≤1|\xi|\le1; the first two derivatives of fλ(eu)f_\lambda(e^u) have fixed compact support and size O((1+∣λ∣)2)O((1+|\lambda|)^2).

The remaining factor

Qb′(u)=PY∑Y/(2P)≤b′≤4Y/P(∏i=2KVi−1)χb(b′)(b′)iuQ_{b'}(u)=\frac{P}{Y}\sum_{Y/(2P)\le b'\le4Y/P}\left(\prod_{i=2}^{K}V_i^{-1}\right)\chi_b(b')(b')^{iu}

has absolute value O(1)O(1) by its reciprocal coefficient mass. Precisely, the dyad contribution to Bλ(t)B_\lambda(t) is

V1−1∫Rf^λ(ξ)Y−iξ(P−1∑p∈P1∩[P,2P)χb(p)pi(t+ξ))Qb′(t+ξ) dξ.V_1^{-1}\int_{\mathbb{R}}\widehat f_\lambda(\xi)Y^{-i\xi}\left(P^{-1}\sum_{p\in\mathcal{P}_1\cap[P,2P)}\chi_b(p)p^{i(t+\xi)}\right)Q_{b'}(t+\xi)\,d\xi.

Minkowski’s integral inequality shows that it suffices to prove

∫∣t∣≤2XLB1∣Ps(t)Mα(t)Nβ(t)∣2 dt≪L−A′′(115)\int_{|t|\le2XL^{B_1}}|P_s(t)M_\alpha(t)N_\beta(t)|^2\,dt\ll L^{-A''} \tag*{(115)}

with arbitrary fixed A′′A'', uniformly in every real shift ξ\xi. The O(L)O(L) dyads and the Mellin L1L^1 norm are then paid by increasing A′′A''. In particular there is no unaccounted Fourier tail in this factor separation.

On ∣Ps(t)∣≤L−As|P_s(t)|\le L^{-A_s}, (113) bounds (5.22) by O(L−2As+C4+B1)O(L^{-2A_s+C_4+B_1}). Choose AsA_s large after A′′A''. On the remaining set use the unit intervals of Lemma 5.3. We have ∣Ps(t)∣≪1|P_s(t)|\ll1, and Lemma 5.2 gives ∣Mα(t)∣≪L−A|M_\alpha(t)|\ll L^{-A} everywhere in the retained range, including its low frequencies. Therefore its contribution is

≪L−2A∑I∈Isup⁡t∈I∣Nβ(t)∣2≪L−2A+2C+1.\ll L^{-2A}\sum_{I\in\mathcal{I}}\sup_{t\in I}|N_\beta(t)|^2\ll L^{-2A+2C+1}.

Taking AA sufficiently large proves (5.22), then (5.21), and finally (5.2).

Completion of Theorem 3.1. Fix the requested precision D∗D_\ast and choose the expanded-square precision D1D_1 sufficiently large in terms of D∗D_\ast and the original coefficient bound CC. These choices precede KK. The squarefree-label and coincident-label errors from the Cauchy reduction are smaller than every fixed logarithmic power. Proposition 4.1, including its restoration of good states and harmonic pads, replaces the remaining determinant equality by (20) with error O(XYL−D1)O(XY L^{-D_1}).

To record the parameter order, first choose E0E_0 after D1,CD_1,C, then the arc exponent A0A_0, and finally KK, as in that proposition. For a pad dyad, division of its pairing bound by d0d_0 uses UV=d0YXUV=d_0YX and gives XYL−E0+OC(1)XYL^{-E_0+O_C(1)}. There are OK(L)O_K(L) pad dyads; thus E0E_0 can absorb this loss with an exponent independent of KK. Constants involving the now fixed KK and its band exponents only affect the threshold for xx. The analytic accuracies in Proposition 5.1 are chosen last, after A0A_0. That proposition bounds the unrestricted major term by O(XYL−D1)O(XYL^{-D_1}).

Consequently the Cauchy square in each YY dyad is O(XYL−D1)O(XYL^{-D_1}). Multiplication by its Cauchy factor O(X/Y)O(X/Y) and taking square roots gives O(XL−D1/2)O(XL^{-D_1/2}) for that dyad. The OK(L)O_K(L) dyads of YY, and all preceding fixed coefficient losses, are absorbed by the original choice of D1D_1. This proves (17), with KK depending on the stated fixed data and not on the particular exponents aia_i.

Two-sided marked correlations

We now allow marks on both endpoints of the shift. The integer JJ used in this section is sufficiently large and fixed as x→∞x \to\infty; it is different from the growing determinant parameter in the preceding sections. Likewise the active damping parameter q0q_0 below need not equal the parameter qq in the fixed-band weight WW.

Retain the fixed KK bands defining WW. Independently choose pairwise disjoint active groups Pg\mathcal{P}_g, indexed by S⊔B\mathcal{S}\sqcup\mathcal{B}, such that, for fixed positive constants c0,C0,v−,v+c_0,C_0,v_-,v_+,

c0log⁡L≤∣S∣,∣B∣≤C0log⁡L,v−≤Vg:=∑p∈Pgp−1≤v+,(116)c_0\log L \leq|\mathcal{S}|,|\mathcal{B}| \leq C_0\log L,\qquad v_- \leq V_g:=\sum_{p\in\mathcal{P}_g}p^{-1}\leq v_+, \tag*{(116)}
Pg⊂[exp⁡(L.27),exp⁡(L.36)](g∈S),Pg⊂[exp⁡(L.39),exp⁡(L.46)](g∈B).\mathcal{P}_g\subset[\exp(L^{.27}),\exp(L^{.36})]\quad(g\in\mathcal{S}),\qquad\mathcal{P}_g\subset[\exp(L^{.39}),\exp(L^{.46})]\quad(g\in\mathcal{B}).

Every big group is the set of all primes in an interval. Write P=⋃gPg\mathcal{P}=\bigcup_g\mathcal{P}_g, sP=∣S∣+∣B∣s_{\mathcal{P}}=|\mathcal{S}|+|\mathcal{B}|, and define

n∗:=∏p∉Ppvp(n),ωg(n):=∑p∈Pg1p∣n,ωP(n):=∑gωg(n),μg(p):=1Vgp.n_\ast:=\prod_{p\notin\mathcal{P}}p^{v_p(n)},\qquad\omega_g(n):=\sum_{p\in\mathcal{P}_g}\mathbf{1}_{p\mid n},\qquad\omega_{\mathcal{P}}(n):=\sum_g\omega_g(n),\qquad\mu_g(p):=\frac{1}{V_gp}.

Thus n∗n_\ast removes the entire active part, including its multiplicities. For a vector t=(tg)\mathbf{t}=(t_g) of nonnegative integers, a represented list at nn consists of tgt_g ordered distinct prime divisors of nn from each group. For a function b\mathbf{b} of these lists put

Wtb(n):=q0ωP(n)−∑gtg∏gVg−tg∑represented lists at nb(p),0<q0<1.(117)W_{\mathbf{t}}^{\mathbf{b}}(n):=q_0^{\omega_{\mathcal{P}}(n)-\sum_g t_g}\prod_g V_g^{-t_g}\sum_{\text{represented lists at }n}\mathbf{b}(\mathbf{p}),\qquad0<q_0<1. \tag*{(117)}

The weight is zero if no list exists. Write WtW_{\mathbf{t}} when b=1\mathbf{b}=\mathbf{1}, and W1=W(1,…,1)W_1=W_{(1,\ldots,1)}. For every fixed m′m',

∣Wtb(n)∣≤Wt(n)≤LC(m′)(tg≤m′, ∣b∣≤1).(118)|W_{\mathbf{t}}^{\mathbf{b}}(n)|\leq W_{\mathbf{t}}(n)\leq L^{C(m')}\qquad(t_g\leq m',\ |\mathbf{b}|\leq1). \tag*{(118)}

Write (u)t=u(u−1)⋯(u−t+1)(u)_t=u(u-1)\cdots(u-t+1), with (u)0=1(u)_0=1. Indeed, each factor q0u−t(u)t/Vgtq_0^{u-t}(u)_t/V_g^t, for u≥tu\geq t, is bounded by a constant depending only on m′,q0,v−m',q_0,v_-, and there are O(log⁡L)O(\log L) factors.

Theorem 6.1 (Two-sided correlation). Fix d≥1d \ge1, ε>0\varepsilon> 0, and intervals I1,…,IdI_1,\ldots,I_d whose lower endpoints are at least xεx^\varepsilon and whose upper endpoints have product at most x1−εx^{1-\varepsilon}. Let l=(l1,…,ld)\mathbf{l}=(l_1,\ldots,l_d) run over ordered tuples of distinct primes with li∈Iil_i \in I_i, and put

l:=l1⋯ld,Cx:=∑l1l,Fx(n):=W(n)(∑l∣n1l−Cx).(119)l := l_1 \cdots l_d,\qquad C_x := \sum_{\mathbf{l}} \frac{1}{l},\qquad F_x(n) := W(n)\left(\sum_{\mathbf{l}\mid n}\frac{1}{l}-C_x\right). \tag*{(119)}

Suppose G(n)=G(n∗)G(n)=G(n_*) and ∣G(n)∣≤LCτ(n∗)C|G(n)|\le L^{C}\tau(n_*)^C, for fixed CC. For every fixed smooth compactly supported Ψ:(0,∞)→C\Psi:(0,\infty)\to\mathbb{C} and every fixed A>0A>0,

∑w≥1Ψ(w/x)wFx(w)G(w+1)W1(w)W1(w+1)≪AL−A.(120)\sum_{w\ge1}\frac{\Psi(w/x)}{w}F_x(w)G(w+1)W_1(w)W_1(w+1)\ll_A L^{-A}. \tag*{(120)}

The constants and sufficiently-large-xx threshold may depend on all the fixed data, including the fixed inert bands, but the estimate is uniform over the active groups and slot intervals satisfying the stated bounds. The slots have no joint restriction other than distinctness.

The harmonic prime estimates give Cx=Oε(1)C_x=O_\varepsilon(1). On every fixed positive-power range for nn, there are only Oε(1)O_\varepsilon(1) prime divisors in the slot ranges. Since WW is bounded for fixed K,qK,q, FxF_x is bounded there. Also Fx(pn)=Fx(n)F_x(pn)=F_x(n) for every active prime pp: the active groups are disjoint both from the inert bands and from the long-prime slots. These two facts will give the stronger endpoint hypotheses required by the graph argument.

The proof has two stages. First, the graph inputs from S give arbitrary logarithmic savings for a signed combination of correlations. After common prime labels have been divided out of both endpoints, one term, called the raw term, is the unit-shift correlation in (120); the other terms have shifts that are products of independently sampled big-group primes. We call these other terms the comparison correlations. We then bound each comparison separately, using the centered first endpoint, and recover the unit-shift term by subtraction. We state the graph inputs with their exact measures before making this reduction. The new endpoint estimates begin once the comparison correlations have been identified.

The exact graph inputs

For clarity we state the general graph contracts, rather than using a correlation theorem with more restrictive endpoint hypotheses. In this subsection only, replace .27,.36,.39,.46.27,.36,.39,.46 in (6.1) by arbitrary fixed 0<a<b<c<d0<.470<a<b<c<d_0<.47, and fix an integer ℓ≥1\ell\ge1. Extend the divisibility definitions to all n∈Zn\in\mathbb{Z}, with every group prime dividing zero. Put M=J+ℓM=J+\ell. A pattern ν\nu chooses sets

Cg⊂{1,…,J},Cg=∅ (g∈S),∣Cg∣≤m0 (g∈B),tg=ℓ+∣Cg∣.C_g\subset\{1,\ldots,J\},\qquad C_g=\varnothing\ (g\in S),\qquad|C_g|\le m_0\ (g\in B),\qquad t_g=\ell+|C_g|.

In each of the first JJ slots outside CgC_g, the source and target lists have the same prime label. The product of these shared labels is DRD_R. In each slot of CgC_g take an independent auxiliary label of law μg\mu_g; their product is DCD_C. The dilation D=DRDCD=D_RD_C therefore contains JJ labels from each group, counted with multiplicity. A coefficient KνK_\nu may depend only on the ordered big-group source lists, target lists, and auxiliary free labels.

The physical Hilbert space consists of states (n,p)(n,\mathbf{p}) with n∈Zn\in\mathbb{Z} and MM ordered distinct divisors from every group. Each state has mass ∏gVg−M\prod_g V_g^{-M}, with counting measure in nn. Fix an integer kk with 0<∣k∣≤LC20<|k|\le L^{C_2}. A row operation samples the auxiliary labels, sets n′=n+kDn'=n+kD, and retains the shared labels while summing over physical target lists with coefficient ∏gVg−tg\prod_g V_g^{-t_g}. Every auxiliary free label is forbidden from both endpoint lists; auxiliary labels may coincide with one another. The row multiplier is

KνDiζq0(ωP(n)−MsP)/2q0(ωP(n′)−MsP)/2.(121)K_{\nu}D^{i\zeta}\frac{q_0^{(\omega_P(n)-M s_P)/2}}{q_0^{(\omega_P(n')-M s_P)/2}}. \tag*{(121)}

Add the patterns and average independent permutations of the MM slots in every group at both ends. This defines AζA_\zeta. In particular, the joint measure of the two lists in group gg is Vg−M−tgV_g^{-M-t_g} in either direction. A row operator integrates the target function and returns a function at the source. The ideal Hilbert space is

Hid=L2(∏g∈BPgM,⨂g∈Bμg⊗M).\mathcal{H}_{\mathrm{id}}=L^2\left(\prod_{g\in\mathcal{B}}\mathcal{P}_g^M,\bigotimes_{g\in\mathcal{B}}\mu_g^{\otimes M}\right).

Repetitions are allowed here. Retain the shared coordinates and sample the unshared target and auxiliary coordinates independently. If DBD_{\mathcal{B}} is the big-group part of DD, use multiplier KνDBiζε(ΘDB)K_\nu D_{\mathcal{B}}^{i\zeta}\varepsilon(\Theta D_{\mathcal{B}}). The sum, symmetrized at both ends, is Tζ(Θ)T_\zeta(\Theta).

Proposition 6.2 (General graph transference input). Assume ∑νsup⁡∣Kν∣≤LC1\sum_\nu\sup|K_\nu| \le L^{C_1}, where m0,C1m_0,C_1 are fixed independently of JJ. For every fixed E>0E>0 there is Eid>0E_{\mathrm{id}}>0, depending only on EE and the fixed group data, such that

sup⁡Θ∈R∥Tζ(Θ)∥2→2≤L−Eid\sup_{\Theta\in\mathbb{R}}\|T_\zeta(\Theta)\|_{2\to2}\le L^{-E_{\mathrm{id}}}

has the following consequence for all sufficiently large fixed JJ. Let I=[X,2X)I=[X,2X), with xγ≤X≤x1/γx^\gamma\le X\le x^{1/\gamma} for fixed γ>0\gamma>0. If ff is supported on positions in II, is independent of the chosen marks, and satisfies

sup⁡∣f∣≤exp⁡(O(L)),∥f∥,∥g∥≤X1/2LC3,\sup|f|\le\exp(O(\sqrt{L})),\qquad\|f\|,\|g\|\le X^{1/2}L^{C_3},

then

∣⟨f,Aζg⟩∣≪XL−E+2C3.(122)|\langle f,A_\zeta g\rangle|\ll XL^{-E+2C_3}. \tag*{(122)}

The lower bound on JJ may also depend on m0,C1m_0,C_1. The threshold for xx may depend on J,C2,γJ,C_2,\gamma. No further restriction on ζ\zeta is imposed beyond the ideal norm hypothesis. Taking absolute values of the individual transition coefficients gives row and column sums at most LCabsL^{C_{\mathrm{abs}}}, where CabsC_{\mathrm{abs}} is independent of JJ.

This is the combination of the local transference theorem, endpoint pairing corollary, and physical Schur bound of SS [21] (Theorem 3.5, Corollary 3.11, and Lemma 3.4). Its parameter qq is our q0q_0; its laws, physical state masses, two-sided symmetrization, and ban on free endpoint labels are exactly those defined above. In the absolute bound, if a target has y≥My\ge M divisors in a group, its unshared list sum is bounded by

sup⁡z≥0Vg−tg(z+tg)tgq0z/2=Om0,ℓ(1).\sup_{z\ge0}V_g^{-t_g}(z+t_g)_{t_g}q_0^{z/2}=O_{m_0,\ell}(1).

The source damping is at most one. Products over O(log⁡L)O(\log L) groups and the pattern mass give the asserted exponent independent of JJ; the symmetric joint normalization gives the column bound as well.

Proposition 6.3 (Residual ideal family input). For every Eid>0E_{\mathrm{id}}>0 and fixed C4≥0C_4\ge0, there are fixed m0,J0,C1m_0,J_0,C_1 such that, for each fixed J≥J0J\ge J_0, a family consisting of the raw pattern Cg=∅C_g=\varnothing, K0=1K_0=1, and signed comparison patterns satisfies

∑νsup⁡∣Kν∣≤LC1,sup⁡Θ∈R∣ζ∣≤LC4∥Tζ(Θ)∥2→2≤L−Eid.(123)\sum_\nu\sup|K_\nu|\le L^{C_1},\qquad \sup_{\substack{\Theta\in\mathbb{R}\\|\zeta|\le L^{C_4}}}\|T_\zeta(\Theta)\|_{2\to2}\le L^{-E_{\mathrm{id}}}. \tag*{(123)}

The family is independent of Θ\Theta, ζ\zeta; its defining parameters are independent of JJ.

Split the big groups into two blocks. Every comparison frees nonempty sets of probes AA, BB in the respective blocks, with coefficient

−(−1)∣A∣+∣B∣KA(DA,XA)KB(DB,XB).(124)-(-1)^{|A|+|B|}K_A(D_A,X_A)K_B(D_B,X_B). \tag*{(124)}

Here DA,DBD_A,D_B are the free products, XAX_A is the corresponding target-mark product, and XBX_B the corresponding source-mark product. No shared or final ℓ\ell label enters these kernels. For a nonempty slot set Λ\Lambda with at most m0m_0 slots in each group,

KΛ(D,X)=∑j:νΛj≥ξ1log⁡D∈Ij1log⁡X∈IjνΛj∑χ primitivecond⁡(χ)≤Qχ(D)‾χ(X),Ij=[jh,(j+1)h).(125)K_\Lambda(D,X)=\sum_{j:\nu_{\Lambda j}\ge\xi}\frac{\mathbf{1}_{\log D\in I_j}\mathbf{1}_{\log X\in I_j}}{\nu_{\Lambda j}}\sum_{\substack{\chi\ \mathrm{primitive}\\ \operatorname{cond}(\chi)\le Q}}\overline{\chi(D)}\chi(X),\qquad I_j=[jh,(j+1)h). \tag*{(125)}

The number νΛj\nu_{\Lambda j} is the cell probability for the independent slot laws. The parameters QQ, h−1h^{-1}, ξ−1\xi^{-1} are fixed powers of LL, with h≤1h\le1. There are at most Q2Q^2 characters and Om0(1+Ld0log⁡L/h)O_{m_0}(1+L^{d_0}\log L/h) nonempty cells, and

sup⁡∣KΛ∣≤Q2ξ−1,sup⁡XED∣KΛ(D,X)∣≤Q2.\sup|K_\Lambda|\le Q^2\xi^{-1},\qquad\sup_X\mathbb{E}_D|K_\Lambda(D,X)|\le Q^2.

More precisely, for any prescribed fixed A0>0A_0>0, the parameters may be chosen, independently of JJ, so that for arbitrary L2L^2 tuple tests f(XB),g(XA)f(X_B),g(X_A) and all θ∈R\theta\in\mathbb{R}, ∣ζ∣≤LC4|\zeta|\le L^{C_4},

∣E f(XB)‾g(XA)[(XAXB)iζe(θXAXB)−KA(DA,XA)KB(DB,XB)(DADB)iζe(θDADB)]∣≤L−A0∥f∥2∥g∥2.(126)\left|\mathbb{E}\,\overline{f(X_B)}g(X_A)\left[(X_AX_B)^{i\zeta}e(\theta X_AX_B)-K_A(D_A,X_A)K_B(D_B,X_B)(D_AD_B)^{i\zeta}e(\theta D_AD_B)\right]\right|\le L^{-A_0}\|f\|_2\|g\|_2. \tag*{(126)}

The four label tuples in this expectation are independent; the tests need not depend only on their products. Expanding the two kernels, including all patterns, has total coefficient mass and number of terms bounded by fixed powers of LL.

This is S [21], Theorem 4.1 and Lemma 4.3). The kernels in (6.10) are two global cell and character expansions; there is not a separate character expansion for every group. We use these two imported propositions with ℓ=k=1\ell=k=1 and the exponents in (6.1).

Endpoint bounds and removal of shared labels

On a physical state ωP(n)≥MsP\omega_P(n)\ge M s_P. Use endpoint vectors with values Fx(n)‾\overline{F_x(n)} and G(n′)G(n'), respectively, and with one extra half-damping factor at each endpoint. The conjugation compensates for that in the Hilbert-space pairing. These extra factors are at most one and are independent of the lists. For a fixed ordered physical list with squarefree product PP, n∗∣n/Pn_*\mid n/P. The divisor moment bound of Lemma 2.2 gives

∑n≍XP∣n∣G(n)∣2≤L2C∑v≪X/Pτ(v)2C≪(X/P)LC∗.\sum_{\substack{n\asymp X\\ P\mid n}}|G(n)|^2\le L^{2C}\sum_{v\ll X/P}\tau(v)^{2C}\ll(X/P)L^{C_*}.

The same statement holds for FxF_x, which is bounded. Here P≤exp⁡(OJ(L.46log⁡L))=xo(1)P\le\exp(O_J(L^{.46}\log L))=x^{o(1)}, and the normalized list sum satisfies

∏gVg−M∑physical lists1P≤∏g(∑p∈Pg1Vgp)M=1.(127)\prod_g V_g^{-M}\sum_{\text{physical lists}}\frac{1}{P}\le\prod_g\left(\sum_{p\in\mathcal{P}_g}\frac{1}{V_gp}\right)^M=1. \tag*{(127)}

Consequently both endpoint norms on an enlarged dyad are at most X1/2LC3X^{1/2}L^{C_3}, where C3C_3 is independent of JJ and m0m_0. The first vector has bounded supremum. It is essential here to keep q0(ωP(n)−MsP)/2≤1q_0^{(\omega_P(n)-M s_P)/2} \le1 intact; separating its two powers would introduce an unnecessary exponent depending on JJ.

Insert Ψ(n/(xD))/n\Psi(n/(xD))/n into the physical edge pairing. Its support has n≍n′≍xD=x1+o(1)n \asymp n' \asymp xD = x^{1+o(1)}. Partition nn into O(L)O(L) dyads [X,2X)[X,2X), restricting n′n' to fixed enlargements. Let ϕ(u)=Ψ(eu)\phi(u)=\Psi(e^u) and use the convention ϕ^(ζ)=∫Rϕ(u)e−iζu du\widehat{\phi}(\zeta)=\int_{\mathbb{R}}\phi(u)e^{-i\zeta u}\,du. Then

Ψ(nxD)=12π∫Rϕ^(ζ)(n/x)iζD−iζ dζ.(128)\Psi\left(\frac{n}{xD}\right)=\frac{1}{2\pi}\int_{\mathbb{R}}\widehat{\phi}(\zeta)(n/x)^{i\zeta}D^{-i\zeta}\,d\zeta. \tag*{(128)}

Put X/nX/n into the first vector and 1/X1/X outside. The endpoint bounds just proved are unchanged. Fix the cutoff ∣ζ∣≤L|\zeta|\le L now, so C4=1C_4=1 suffices in Proposition 6.3.

The choices have the following order. After the desired saving AA and the original data, fix C3C_3, then a transfer saving EE large enough to cover AA, 2C32C_3, and the fixed dyadic costs. Choose EidE_{\mathrm{id}} from Proposition 6.2, obtain m0m_0, J0J_0, C1C_1 from Proposition 6.3, and finally choose a fixed JJ satisfying both lower bounds. For the tail of (6.13), the absolute Schur estimate and the endpoint norms cost at most a fixed power of LL. Since

∫∣ζ∣>L∣ϕ^(ζ)∣ dζ≪NL−N\int_{|\zeta|>L}|\widehat{\phi}(\zeta)|\,d\zeta\ll_N L^{-N}

for every fixed NN, its differentiation order can be chosen after the family is fixed. The cutoff exponent does not have to change. Thus the total residual edge pairing is OA(L−A)O_A(L^{-A}), with as much extra fixed saving as will be needed below.

For a single pattern put n=DRwn=D_Rw, n′=DRw′n'=D_Rw'. The shift becomes w′=w+DCw'=w+D_C. Away from shared-label collisions,

ωg(DRw)−M=ωg(w)−tg,Vg−M−tg=Vg−(M−tg)Vg−2tg,1n=1DRw.(129)\omega_g(D_Rw)-M=\omega_g(w)-t_g,\qquad V_g^{-M-t_g}=V_g^{-(M-t_g)}V_g^{-2t_g},\qquad\frac{1}{n}=\frac{1}{D_Rw}. \tag*{(129)}

The factor DR−1D_R^{-1} changes each shared sum into its probability law μg\mu_g, and the remaining list factors and damping are exactly Wt(w)Wt(w′)W_t(w)W_t(w'). The endpoint cores are invariant under these labels. The comparison coefficients involve no shared labels. Consequently the stripped expression is

EDC∑w≥1Ψ(w/(xDC))wFx(w)G(w+DC)Wt(w)Wt(w+DC)Ep(w),p(w+DC)Kν,(130)\mathbb{E}_{D_C}\sum_{w\ge1}\frac{\Psi(w/(xD_C))}{w}F_x(w)G(w+D_C)W_t(w)W_t(w+D_C)\mathbb{E}_{\mathbf p(w),\mathbf p(w+D_C)}K_\nu, \tag*{(130)}

where the last expectation is uniform on each endpoint’s ordered unshared lists. The independent permutations have total mass one and only rename the ordered slots, so they introduce no factorial factor.

We justify restoring all the restrictions suppressed in this formula. There are OJ(log⁡L)O_J(\log L) shared or free slots. Every group atom has size O(exp⁡(−L.27))O(\exp(-L^{.27})). Conditional on ww, w′w' and the free labels, the harmonic mass of a shared/shared coincidence, a shared/free coincidence, or a shared prime dividing ww′ww' is therefore

OJ(LCJexp⁡(−L.27)).O_J(L^{C_J}\exp(-L^{.27})).

For the last assertion use that an integer of size xO(1)x^{O(1)} has at most O(L)O(L) distinct prime factors. On actual overlap exceptions the unshared labels still divide w,w′w,w'. Changing the damping loses at most q0−2MsP=LO(J)q_0^{-2M s_P}=L^{O(J)}. The marked bounds and Cauchy with fixed divisor moments on the translated dyads bound the remaining harmonic sum by a fixed log power. Thus (6.16) also controls these actual exceptions.

If a free prime occurs in either unshared endpoint list, it divides DCD_C and one of w,w+DCw,w+D_C, hence both. It is at least exp⁡(L.39)\exp(L^{.39}). On w≍xDCw\asymp xD_C the count of its multiples is O(xDC/p)O(xD_C/p); Cauchy and a fixed higher divisor moment bound the corresponding weighted harmonic mass by LO(1)p−1/2L^{O(1)}p^{-1/2}. Summing over the Om0(log⁡L)O_{m_0}(\log L) free slots costs LO(1)exp⁡(−cL.39)L^{O(1)}\exp(-cL^{.39}). This also restores any coincidence between the two unshared endpoint lists, since such a label must divide their difference DCD_C. All errors remain smaller than every negative log power after the fixed pattern cost. They are incurred after JJ is fixed and do not change the exponent C3C_3 used in transference. For the raw pattern, DC=1D_C=1 and tg=1t_g=1, so (6.15) is precisely the left side of (120).

The transferred estimate therefore controls the desired correlation plus the signed comparison correlations (130). It remains to bound the latter separately. Their shifts contain two nonempty products of free big-group labels; these products will give a bilinear Fourier multiplier. All cancellation required at an endpoint will come from the centered function FxF_x.

Comparison multipliers and local energy

Expand the two kernels in (124). Their contract supplies only a fixed log-power total cost. In each term the two nonempty free products are independent and lie in log cells of width at most one, say DA≍U1D_A \asymp U_1, DB≍U2D_B \asymp U_2. Put

H′=U1U2,Y=xH′,Ui≥exp⁡(L.39−O(1)),H′≤exp⁡(Om0(L.46log⁡L)).(131)H' = U_1U_2,\qquad Y=xH',\qquad U_i \ge\exp(L^{.39}-O(1)),\qquad H' \le\exp\left(O_{m_0}(L^{.46}\log L)\right). \tag*{(131)}

The cell and character expansion factors into separate bounded tests on the two free tuples and on each endpoint’s big marks. Insert fixed smooth cutoffs at scale YY at both endpoints. Fourier inversion in logarithmic size separates Ψ(w/(xDADB))/w\Psi(w/(xD_AD_B))/w as in (128). After a factor 1/Y1/Y, it is enough to estimate combinations of

E b1(DA)b2(DB)∑wU(w)V(w+DADB),∣bi∣≤1.(132)\mathbb{E}\,b_1(D_A)b_2(D_B)\sum_w \mathcal{U}(w)\mathcal{V}(w+D_AD_B),\qquad|b_i|\le1. \tag*{(132)}

Tuple tests may be used in place of product tests; conditioning on the product produces the notation here. The endpoint sequences have smooth cutoffs at YY, fixed log-power twists, and the form of their respective invariant cores times WtbW_t^b, with ∣b∣≤1|b|\le1 depending only on big marks. Fixed divisor moments give

∑w∣U(w)∣2+∑w∣V(w)∣2≪YLC5.(133)\sum_w|\mathcal{U}(w)|^2+\sum_w|\mathcal{V}(w)|^2\ll YL^{C_5}. \tag*{(133)}

These estimates also bound the separated Fourier tails absolutely: on any translated time interval, the corresponding collapsed Dirichlet polynomials obey the coefficient mean-square estimate. Smooth Fourier decay can therefore give any required precision. All exponents chosen from now on may depend on the fixed ideal family.

The additive multiplier for (132) is

m(θ):=E b1(DA)b2(DB)e(θDADB),∣m(θ)∣≤1.m(\theta):=\mathbb{E}\,b_1(D_A)b_2(D_B)e(\theta D_AD_B),\qquad|m(\theta)|\le1.

The product probability of a tuple using kg≤m0k_g\le m_0 slots per group is, by unique factorization,

ρ(u)=1u∏gkg!Vgkg∏pap!≤LC7u.\rho(u)=\frac{1}{u}\prod_g\frac{k_g!}{V_g^{k_g}\prod_p a_p!}\le\frac{L^{C_7}}{u}.

Since its total mass is at most one, each collapsed sequence on u≍Uiu\asymp U_i has squared coefficient norm at most LC7/UiL^{C_7}/U_i. The elementary bilinear Fourier estimate (Cauchy followed by the spacing estimate for ∥θu∥−1\|\theta_u\|^{-1}) consequently gives, if (h,r)=1(h,r)=1 and ∣θ−h/r∣≤r−2|\theta-h/r|\le r^{-2},

∣m(θ)∣≪LC7((U1/r+1)(U2+rlog⁡(2r))H′)1/2.(134)|m(\theta)|\ll L^{C_7}\left(\frac{(U_1/r+1)(U_2+r\log(2r))}{H'}\right)^{1/2}. \tag*{(134)}

For completeness, Cauchy in the variable on the interval of length O(U2)O(U_2), expansion, and the geometric-sum bound give an unnormalized square at most

∥a∥22∥b∥22(U2+∑1≤v≪U1min⁡{U2,∥θv∥R/Z−1}).\|a\|_2^2\|b\|_2^2\left(U_2+\sum_{1\le v\ll U_1}\min\{U_2,\|\theta v\|_{\mathbb{R}/\mathbb{Z}}^{-1}\}\right).

For r≥2r\ge2, split the vv-range into O(1+U1/r)O(1+U_1/r) blocks of diameter at most r/2r/2. Distinct points in one block have residues θv\theta v separated by at least 1/(2r)1/(2r), because ∣θ−h/r∣≤r−2|\theta-h/r|\le r^{-2}. The sum in a block is at most O(U2+rlog⁡(2r))O(U_2+r\log(2r)). The case r=1r=1 is immediate from the trivial bound U2U_2 on each summand. This proves (134) with the two coefficient norms above.

Given any required saving, choose C6C_6 sufficiently large and set

H=H′/LC6,MH=⋃k≤LC6(h,k)=1{θ:∣θ−h/k∣≤H−1}.(135)H=H'/L^{C_6},\qquad\mathfrak{M}_H=\bigcup_{\substack{k\le L^{C_6}\\(h,k)=1}}\{\theta:|\theta-h/k|\le H^{-1}\}. \tag*{(135)}

Dirichlet approximation with denominator bound ⌊H′/LC6−1⌋\lfloor H'/L^{C_6-1}\rfloor gives a fraction h/rh/r with ∣θ−h/r∣≤r−2|\theta-h/r|\le r^{-2}. If r≤LC6r\le L^{C_6}, this places θ\theta in MH\mathfrak{M}_H. Otherwise LC6<r≪H′/LC6−1L^{C_6}<r\ll H'/L^{C_6-1}, and the expression under the square root in (134) is bounded by

1r+log⁡(2r)U2+1U1+rlog⁡(2r)H′.\frac{1}{r}+\frac{\log(2r)}{U_2}+\frac{1}{U_1}+\frac{r\log(2r)}{H'}.

The middle terms save every log power by (131), and the other terms give any desired fixed saving by increasing C6C_6. Thus mm is arbitrarily log-power small off MH\mathfrak{M}_H.

Use the Fourier convention a^(θ)=∑wawe(θw)\widehat{a}(\theta)=\sum_w a_w e(\theta w). Parseval and (133) handle the complement of MH\mathfrak{M}_H. On its O(L2C6)O(L^{2C_6}) arcs, Cauchy and the global energy of V\mathcal{V} reduce the problem to

∫∣β∣≤H−1∣u^(h/k+β)∣2 dβ≪A′YL−A′(136)\int_{|\beta|\le H^{-1}}|\widehat{u}(h/k+\beta)|^2\,\mathrm{d}\beta\ll_{A'}YL^{-A'} \tag*{(136)}

for every fixed A′A'. The scale conversion we need is as follows.

Lemma 6.4 (Local Mellin energy). Let awa_w be supported on w≍Yw\asymp Y, where Y=x1+o(1)Y=x^{1+o(1)}, and suppose H=xo(1)→∞H=x^{o(1)}\to\infty. For every fixed integer j≥1j\ge1,

∫∣β∣≤H−1∣a^(β)∣2 dβ≪j1Y∫R(1+H∣t∣/Y)−j∣∑wawwit∣2 dt+(HY)2∑w∣aw∣2.(137)\int_{|\beta|\le H^{-1}}|\widehat{a}(\beta)|^2\,\mathrm{d}\beta \ll_j \frac{1}{Y}\int_{\mathbb{R}}(1+H|t|/Y)^{-j}\left|\sum_w a_w w^{it}\right|^2\,\mathrm{d}t+\left(\frac{H}{Y}\right)^2\sum_w|a_w|^2. \tag*{(137)}

Proof. Choose a smooth bump KK of integral one, supported close enough to zero that its additive Fourier transform has modulus at least 1/21/2 on [−1,1][-1,1]. Plancherel gives

∫∣β∣≤H−1∣a^(β)∣2 dβ≪H−2∫R∣∑wawK((w−v)/H)∣2 dv.\int_{|\beta|\le H^{-1}}|\widehat{a}(\beta)|^2\,\mathrm{d}\beta \ll H^{-2}\int_{\mathbb{R}}\left|\sum_w a_wK((w-v)/H)\right|^2\,\mathrm{d}v.

Only v≍Yv\asymp Y contributes. On the union of the relevant supports, replace K((w−v)/H)K((w-v)/H) by K(vlog⁡(w/v)/H)K(v\log(w/v)/H): their arguments differ by O(H/Y)O(H/Y). Cauchy over the O(H)O(H) available ww, and integration over the O(H)O(H) centers for each ww, give normalized squared error O((H/Y)2∑∣aw∣2)O((H/Y)^2\sum|a_w|^2).

Write v=Yeuv=Ye^u, h0=H/Yh_0=H/Y, and Pa(t)=∑wawwitP_a(t)=\sum_w a_ww^{it}. Fourier inversion for the new kernel reads

∑wawK(vlog⁡(w/v)/H)=h02π∫Re−uK^(h0e−ut/(2π))v−itPa(t) dt.\sum_w a_wK(v\log(w/v)/H)=\frac{h_0}{2\pi}\int_{\mathbb{R}}e^{-u}\widehat{K}(h_0e^{-u}t/(2\pi))v^{-it}P_a(t)\,\mathrm{d}t.

Insert a fixed smooth cutoff in uu for v≍Yv \asymp Y. Two integrations by parts in uu and the Schwartz bounds for K^\widehat{K} bound the Fourier transform of this amplitude by Cj(1+∣s∣)−2(1+h0∣t∣)−jC_j(1+|s|)^{-2}(1+h_0|t|)^{-j}. Expand in ss, apply Minkowski and Plancherel in uu, and use dv≪Y dudv \ll Y\,du. The outside normalization is H−2Yh02=Y−1H^{-2}Yh_0^2=Y^{-1}. Increasing the decay order proves (137).

Apply the lemma to aw=U(w)e(hw/k)a_w=\mathcal{U}(w)e(hw/k). The additive error is negligible by (133). The Dirichlet mean-square bound in Lemma 2.2 gives, on every dyadic time interval of length RR,

Y−2∫R≤∣t∣≤2R∣Pa(t)∣2 dt≪LC5+1(1+R/Y).Y^{-2}\int_{R\le|t|\le2R}|P_a(t)|^2\,dt\ll L^{C_5+1}(1+R/Y).

For T0=LY/HT_0=LY/H, the tail in (137) is thus at most

YLC5+1∑r≥0(2rL)−j(1+2rL/H),YL^{C_5+1}\sum_{r\ge0}(2^rL)^{-j}(1+2^rL/H),

which is as small as required on choosing jj large. It remains to prove, for every fixed A′>0A'>0,

∫∣t∣≤T0∣Y−1∑wU(w)e(hw/k)wit∣2dt≪A′L−A′,T0=LYH=xLC6+1.(138)\int_{|t|\le T_0}\left|Y^{-1}\sum_w\mathcal{U}(w)e(hw/k)w^{it}\right|^2dt\ll_{A'}L^{-A'},\qquad T_0=\frac{LY}{H}=xL^{C_6+1}. \tag*{(138)}

Notice that the first term in (137) is YY times the integral in (138), as required by (136).

We now prove this one endpoint estimate. At low Mellin frequencies, a finite divisor model makes the centering in FxF_x cancel the main term. At high frequencies, one small-prime mark supplies a polynomial that is small outside a sparse set of times. On that sparse set we treat the positive tuple sum and its subtracted constant separately: the tuple sum has a long-prime factor, whereas the constant term requires cancellation on the divisor progressions of the finite model.

A divisor model and low Mellin frequencies

Lemma 6.5 (Truncated divisor model). Let Hmarks(n)=W(n)Wb(n)H_{\mathrm{marks}}(n)=\mathcal{W}(n)\mathcal{W}^{\mathbf b}(n), where 0≤tg≤m′0\le t_g\le m' for fixed m′m' and ∣b∣≤1|\mathbf b|\le1. In particular, zero values of tgt_g are allowed; these occur when a small mark is removed below. There is a divisor sum

H0(n)=∑d′∣nAd′,d′≤exp⁡(O(L.47)),∑d′∣Ad′∣d′≤LC8,(139)H_0(n)=\sum_{d'\mid n}A_{d'},\qquad d'\le\exp(O(L^{.47})),\qquad\sum_{d'}\frac{|A_{d'}|}{d'}\le L^{C_8}, \tag*{(139)}

whose prime divisors all belong to the active or inert groups, such that, on every positive integer dyad n≍Rn\asymp R,

∑n≍R∣Hmarks(n)−H0(n)∣2≪NRL−N(140)\sum_{n\asymp R}|H_{\mathrm{marks}}(n)-H_0(n)|^2\ll_N RL^{-N} \tag*{(140)}

for every fixed NN. The assertion is uniform in the bounded tuple test b\mathbf b, with no regularity assumption on that test.

Proof. Expand both weights into their represented lists. The inert list has one prime per band, and the active list has tgt_g per group. For a fixed represented list, the damping is exactly the product of qq or q0q_0 over the remaining distinct group primes that divide nn. Expand this product as

∏pp not represented(1−(1−qp)1p∣n),qp∈{q,q0},\prod_{\substack{p\\p\ \mathrm{not\ represented}}}\left(1-(1-q_p)1_{p\mid n}\right),\qquad q_p\in\{q,q_0\},

and truncate after h0=⌊L.01⌋h_0=\lfloor L^{.01}\rfloor additional primes. All represented primes are distinct within their groups, and the additional primes are disjoint from them. Each resulting divisor has at most Om′(log⁡L)+K+h0O_{m'}(\log L)+K+h_0 prime factors, all at most exp⁡(L.46)\exp(L^{.46}). This proves the support assertion in (6.25).

In the reciprocal coefficient sum the normalized represented lists have total mass at most one, by their harmonic definitions. The additional subset sums are at most

∏p in all groups(1+(1−qp)/p)≤exp⁡(O(log⁡L))=LO(1).\prod_{p\text{ in all groups}}\left(1+(1-q_p)/p\right)\leq\exp(O(\log L))=L^{O(1)}.

This proves the coefficient bound, even after absolute values.

Let ωtot\omega_{\mathrm{tot}} count hits in all active and inert groups. For fixed list sizes, their normalized counts are at most LC92ωtotL^{C_9}2^{\omega_{\mathrm{tot}}}: use (u)t=t!(ut)≤t!2u(u)_t=t!{u\choose t}\leq t!2^u in each group. The discarded binomial expansion is at most 22ωtot1ωtot≥h02^{2\omega_{\mathrm{tot}}}1_{\omega_{\mathrm{tot}}\geq h_0}. Hence the absolute pointwise error is bounded by LC94ωtot1ωtot≥h0L^{C_9 4^{\omega_{\mathrm{tot}}}}1_{\omega_{\mathrm{tot}}\geq h_0}. For any fixed B>1B>1, expanding BωtotB^{\omega_{\mathrm{tot}}} into squarefree divisors and counting their multiples gives

∑n≍RBωtot(n)≪R∏p in all groups(1+(B−1)/p)≪RLC(B).(141)\sum_{n\asymp R}B^{\omega_{\mathrm{tot}}(n)}\ll R\prod_{p\text{ in all groups}}\left(1+(B-1)/p\right)\ll RL^{C(B)}. \tag*{(141)}

Indeed only divisors at most a constant times RR can occur, and each has O(R/d)O(R/d) multiples in the dyad. Since 16u1u≥h0≤2−h032u16^u1_{u\geq h_0}\leq2^{-h_0}32^u, the square of the error has total at most RLO(1)2−h0RL^{O(1)}2^{-h_0}. This proves (6.26) for every fixed NN. □

Write the first endpoint as

u(w)=ψ1(w/Y)wiσ(∑l1l∣w−Cx)Hmarks(w),(142)u(w)=\psi_1(w/Y)w^{i\sigma}\left(\sum_l1_{l\mid w}-C_x\right)H_{\mathrm{marks}}(w), \tag*{(142)}

where ψ1\psi_1 has fixed compact support and its rescaled derivatives, as well as ∣σ∣|\sigma|, have fixed log-power bounds. For every fixed B′>0B'>0, we claim uniformly for ∣t∣≤LB′|t|\leq L^{B'} that the normalized sum in (6.24) saves any prescribed log power.

In a tuple term put w=lnw=ln. All long primes lie outside the marked groups, so Hmarks(ln)=Hmarks(n)H_{\mathrm{marks}}(ln)=H_{\mathrm{marks}}(n). Apply Lemma 6.5 on the scales Y/lY/l and YY in the tuple and constant terms respectively. Cauchy and (6.26) make the total normalized error at most L−N(1+∑1/l)L^{-N}(1+\sum1/l), after renaming NN. For a fixed divisor d′d', all primes in ld′ld' exceed kk for large xx, so gcd⁡(ld′,k)=1\gcd(ld',k)=1. Smooth sum–integral comparison on the progressions modulo kk gives

1Y∑v≥1ψ1(ld′v/Y)(ld′v)i(t+σ)e(hld′v/k)=1ld′(1k∑a mod ke(ha/k))∫0∞ψ1(z)(Yz)i(t+σ) dz+O(Y−1LC(B′)).(143)\begin{aligned} \frac{1}{Y}\sum_{v\geq1}\psi_1(ld'v/Y)(ld'v)^{i(t+\sigma)}e(hld'v/k) \\ &=\frac{1}{ld'}\left(\frac{1}{k}\sum_{a\bmod k}e(ha/k)\right)\int_0^\infty\psi_1(z)(Yz)^{i(t+\sigma)}\,dz+O(Y^{-1}L^{C(B')}). \tag*{(143)} \end{aligned}

The sum has scale Y/(ld′)≥xε−o(1)Y/(ld')\geq x^{\varepsilon-o(1)}. The error follows by summing the endpoint and total-variation errors over kk residue classes; all derivative and modulus costs are fixed log powers. Moreover

∑l1≤x1−εCx,∑d′∣Ad′∣≤exp⁡(O(L.47))LC8=xo(1).\sum_l1\leq x^{1-\varepsilon}C_x,\qquad\sum_{d'}|A_{d'}|\leq\exp(O(L^{.47}))L^{C_8}=x^{o(1)}.

The summed errors in (6.29) are therefore O(x−ε+o(1))O(x^{-\varepsilon+o(1)}), since Y≥x1−o(1)Y\geq x^{1-o(1)}. The main term depends on ll only through 1/l1/l. Its tuple sum is exactly CxC_x times the main term for the subtracted constant. They cancel, including the complete residue average. Choosing NN after B′B' proves (6.24) on every fixed log-power time interval.

High frequencies: factorization and exceptional times

Choose a small group gsg_s. It has tgs=1t_{g_s}=1 in every comparison, and the tuple test depends only on big marks. Except when a prime of this group divides ww twice,

Wtb(w)=1Vgs∑ps∣wps∈PgsWt−egsb(wps).(144)W_{\mathrm{t}}^{\mathrm{b}}(w)=\frac{1}{V_{g_s}}\sum_{\substack{p_s\mid w\\p_s\in\mathcal{P}_{g_s}}}W_{\mathrm{t-e_{g_s}}}^{\mathrm{b}}\left(\frac{w}{p_s}\right). \tag*{(144)}

This follows directly by removing the represented small prime; the remaining damping exponent is unchanged. Both sides are bounded by fixed log powers. The square exceptions have relative count at most LO(1)exp⁡(−L.27)L^{O(1)}\exp(-L^{.27}) on the relevant dyads. In the positive tuple part we may also allow repeated slot primes: the extra terms have a square of a prime at least xεx^\varepsilon dividing ww, a relative count O(x−ε)O(x^{-\varepsilon}). Slot multiplicities at each w≍Yw\asymp Y remain bounded. Lemma 2.2, applied to these coefficient errors, makes their normalized integrated square over ∣t∣≤T0|t|\le T_0 smaller than every negative log power. Indeed their squared coefficient norm is at most YLO(1)(exp⁡(−L.27)+x−ε)YL^{O(1)}(\exp(-L^{.27})+x^{-\varepsilon}), and T0/Y=L/H=o(1)T_0/Y=L/H=o(1).

In the tuple part factor w=pspluw=p_sp_lu, where pl=l1p_l=l_1 lies in the first long-prime slot. After allowing repeated slots the coefficient on uu is exactly

Wt−egsb(u)W(u)∑l2,…,ld1l2⋯ld∣u.(145)W_{\mathrm{t-e_{g_s}}}^{\mathrm{b}}(u)\mathcal{W}(u)\sum_{l_2,\ldots,l_d}\mathbf{1}_{l_2\cdots l_d\mid u}. \tag*{(145)}

For d=1d=1 the final factor is one. This is bounded by a fixed log power times a fixed divisor power. In the constant part write w=psnw=p_sn; its coefficient is just the remaining marked weight times W(n)\mathcal{W}(n).

Separate the rational phase by residue classes modulo kk. In the tuple part, ps,plp_s,p_l are units modulo kk; on fixed classes for psp_s and uu, expand the remaining test in characters of plp_l by orthogonality on the units. In the constant part fix classes for ps,np_s,n. Both operations have polynomial-in-kk total cost and do not require uu or nn to be units. Divide all factors into dyads and Fourier-separate the smooth product cutoff in logarithmic size. There are only a fixed log-power number of boxes. Their factor scales have product comparable to YY, and every collapsed coefficient sequence of normalized scale RR has

∑v≍R∣cv∣2≪LC11R,(146)\sum_{v\asymp R}|c_v|^2\ll\frac{L^{C_{11}}}{R}, \tag*{(146)}

by fixed divisor moments. This applies also to the uncut product of all factors, so Dirichlet mean squares uniformly on translated time intervals justify discarding the Fourier tails with arbitrary log saving. Retain only shifts of fixed log-power size. Enlarge the low-time exponent B′B' beyond these shifts and any later prime thresholds. The low-time argument already proved is available for this enlarged B′B'. At the remaining times all retained shifts of tt are comparable in magnitude to tt.

It remains to treat products of normalized polynomials at a common time, with LB′≤∣t∣≤2T0L^{B'}\le|t|\le2T_0. In every such product there is a small-prime factor

Ps(t)=P−1∑ps<Pps∈Pgsbs(ps)psit,cL.27≤log⁡P≤L.36,∣bs∣≤LC12.(147)P_s(t)=P^{-1}\sum_{\substack{p_s<P\\p_s\in\mathcal{P}_{g_s}}}b_s(p_s)p_s^{it},\qquad cL^{.27}\le\log P\le L^{.36},\qquad|b_s|\le L^{C_{12}}. \tag*{(147)}

The coefficients include residue restrictions and retained twists. In the tuple part the remaining factors are a long-prime polynomial Pl(t)P_l(t) and a polynomial R(t)R(t) at scales NlN_l and

R0≳YPNl≥xε−o(1).(148)R_0\gtrsim\frac{Y}{PN_l}\ge x^{\varepsilon-o(1)}. \tag*{(148)}

In the constant part the remaining polynomial N(t)N(t) has scale R0≍Y/PR_0\asymp Y/P.

The threshold split and exceptional-frequency treatment follow the Matomäki–Radziwiłł strategy; see [19], §2.1 and Lemmas 8–9 of arXiv v4. The residual-scale sparse Gram estimate below is adapted here using Appendix A, while the factorial coefficient estimate retains the Soundararajan credit stated at its use.

On the set ∣Ps(t)∣≤L−As|P_s(t)| \le L^{-A_s}, collapse all remaining factors. Their scale is Y/PY/P, and

Y/P2T0=H2LP⟶∞,(149)\frac{Y/P}{2T_0} = \frac{H}{2LP} \longrightarrow\infty, \tag*{(149)}

because H≥exp⁡(2L.39−O(log⁡L))H \ge\exp(2L^{.39} - O(\log L)) and P≤exp⁡(L.36)P \le\exp(L^{.36}). (146) and the mean-square bound give a fixed log-power integrated square for this collapsed polynomial. Taking AsA_s sufficiently large handles this part.

Let J\mathcal{J} be the integer unit intervals meeting {∣t∣≤2T0:∣Ps(t)∣>L−As}\{|t| \le2T_0 : |P_s(t)| > L^{-A_s}\}. Then

∣J∣≤exp⁡(O(L.36+L.73log⁡L))≤exp⁡(O(L.9)).(150)|\mathcal{J}| \le\exp(O(L^{.36} + L^{.73}\log L)) \le\exp(O(L^{.9})). \tag*{(150)}

For this step we use the factorial prime-polynomial coefficient estimate, as in [24], Lemma 3 and its proof, together with a coefficient-independent mean square. Put Z=8T0+8Z = 8T_0 + 8 and j0=⌊log⁡Z/log⁡(2P)⌋j_0 = \lfloor\log Z/\log(2P) \rfloor. Unique factorization gives squared coefficient norm for Psj0P_s^{j_0} at most

j0!(P−2∑ps≍P∣bs(ps)∣2)j0≤j0!(LC13/P)j0.j_0! \left(P^{-2}\sum_{p_s \asymp P}|b_s(p_s)|^2\right)^{j_0} \le j_0!(L^{C_{13}}/P)^{j_0}.

Select a point exceeding the threshold from each occupied interval, and split their integer indices modulo three to get separated sets. The indices of this powered polynomial are at most ZZ. The separated-point mean-square bound of Lemma 2.2 therefore gives

∣J∣≪ZLO(1)j0!(LC13+2As/P)j0.|\mathcal{J}| \ll ZL^{O(1)}j_0!(L^{C_{13}+2A_s}/P)^{j_0}.

Since ZP−j0≤2P2j0ZP^{-j_0} \le2P^{2j_0} and j0=O(L.73)j_0 = O(L^{.73}), taking logarithms proves (6.36).

We will also use the following sparse bound for any normalized polynomial R(t)=∑n≍R0cnnitR(t) = \sum_{n\asymp R_0} c_n n^{it} with xε/2≤R0≤x2x^{\varepsilon/2} \le R_0 \le x^2 and ∑∣cn∣2≤LC14/R0\sum|c_n|^2 \le L^{C_{14}}/R_0:

∑I∈Jsup⁡t∈I∣R(t)∣2≪LC15.(151)\sum_{I\in\mathcal{J}}\sup_{t\in I}|R(t)|^2 \ll L^{C_{15}}. \tag*{(151)}

Choose maximizing points and again split into three separated classes. For 1≤∣Δ∣≤Tmod:=exp⁡(L/(log⁡L)2)1 \le|\Delta| \le T_{\mathrm{mod}} := \exp(L/(\log L)^2), sum–integral comparison on the full integer dyad gives

R0−1∣∑n≍R0niΔ∣≪∣Δ∣−1+x−cε.R_0^{-1}\left|\sum_{n\asymp R_0}n^{i\Delta}\right| \ll|\Delta|^{-1}+x^{-c\varepsilon}.

For larger differences, Lemma 2.4 gives O(exp⁡(−L/(log⁡L)C16))O(\exp(-L/(\log L)^{C_{16}})) instead. Its hypotheses hold: R0≥xε/2>exp⁡(L/(log⁡L)2)R_0 \ge x^{\varepsilon/2} > \exp(L/(\log L)^2), all relevant scales are at most 2x52x^5, and the differences are between TmodT_{\mathrm{mod}} and 4T0+O(1)<4x34T_0 + O(1) < 4x^3. The Gram matrix of the vectors (nit)n≍R0(n^{it})_{n\asymp R_0} consequently has absolute row sum

≪R0(1+log⁡(2T0)+∣J∣{x−cε+exp⁡(−L/(log⁡L)C16)})≪R0L.\ll R_0\left(1+\log(2T_0)+|\mathcal{J}|\{x^{-c\varepsilon}+\exp(-L/(\log L)^{C_{16}})\}\right) \ll R_0L.

Here separation bounds the reciprocal-difference sum by O(log⁡T0)O(\log T_0), and (6.36) absorbs both small errors. The Gram operator bound multiplied by ∑∣cn∣2\sum|c_n|^2 proves (6.37).

The tuple and constant contributions at high times

In the tuple term the long-prime polynomial is, after partial summation, a normalized prime sum with a character modulo kk and a fixed log-power twist. Its dyad lies between xε/2x^{\varepsilon}/2 and x1−εx^{1-\varepsilon}. Choose fixed 0<τ<ε0<\tau<\varepsilon and 1−ε<η<11-\varepsilon<\eta<1 to apply Lemma 2.3. For any prescribed AlA_l, it gives

∣Pl(t)∣≪L−Al(LB′≤∣t∣≤2T0),|P_l(t)| \ll L^{-A_l}\qquad(L^{B'}\leq|t|\leq2T_0),

after increasing B′B' beyond its threshold and all retained shifts. The upper frequencies are <x2<x^2. The remaining coefficient (6.31) has scale (6.34) and obeys (6.32), so (6.37) applies. The small polynomial has absolute bound LO(1)L^{O(1)}. Hence on the exceptional intervals the integrated square of PsPlRP_sP_lR is

≪LO(1)−2Al∑I∈Isup⁡t∈I∣R(t)∣2,\ll L^{O(1)-2A_l}\sum_{I\in\mathcal{I}}\sup_{t\in I}|R(t)|^2,

which has arbitrary log saving on choosing AlA_l large enough. Together with the ordinary-time estimate this handles the tuple part.

For the constant part set R0≍Y/PR_0\asymp Y/P, and apply Lemma 6.5 to the remaining marked coefficient at this scale. Keep its residue-class restriction separately. If E(n)E(n) is the coefficient error, then for every fixed NN, ∑n≍R0∣E(n)∣2≪R0L−N\sum_{n\asymp R_0}|E(n)|^2\ll R_0L^{-N}. The mean-square bound on the entire retained time interval gives

∫∣t∣≤2T0∣R0−1∑n≍R0E(n)nit∣2 dt≪L−N(T0R0+log⁡R0)≪L−N(LPH+O(L)).(152)\int_{|t|\leq2T_0}\left|R_0^{-1}\sum_{n\asymp R_0}E(n)n^{it}\right|^2\,dt \ll L^{-N}\left(\frac{T_0}{R_0}+\log R_0\right) \ll L^{-N}\left(\frac{LP}{H}+O(L)\right). \tag*{(152)}

By (6.35), LP/H=o(1)LP/H=o(1). Thus the known small-prime suprema and all fixed decomposition costs can be paid by choosing NN last. In particular the model error is controlled before restricting to the exceptional times.

For the truncated model, a divisor d′d' leaves a progression modulo kk at scale R′=R0/d′R'=R_0/d'. Its normalization in N(t)N(t) is Ad′/d′A_{d'}/d' times normalization at R′R'. Since d′≤exp⁡(O(L.47))d'\leq\exp(O(L^{.47})), one has R′=x1−o(1)R'=x^{1-o(1)}. The divisor is a unit modulo kk, so the original residue restriction becomes a single residue restriction at this scale. For LB′≤∣t∣≤TmodL^{B'}\leq|t|\leq T_{\mathrm{mod}}, comparison with the integral of the logarithmic phase gives

∣N0(t)∣≪LC17(∣t∣−1+x−c′).|N_0(t)|\ll L^{C_{17}}\left(|t|^{-1}+x^{-c'}\right).

Indeed the normalized integral over a dyad is O(1/∣t∣)O(1/|t|) and the progression endpoint and variation error is O(LO(1)(1+∣t∣)/R′)=O(x−c′)O(L^{O(1)}(1+|t|)/R')=O(x^{-c'}); sum using ∑∣Ad′∣/d′≤LC8\sum|A_{d'}|/d'\leq L^{C_8}. Retained twists may either be incorporated in tt or absorbed by increasing B′B'. The square of (6.40) is integrable with

∫LB′≤∣t∣≤Tmod∣N0(t)∣2 dt≪L2C17−B′+x−c′′.(153)\int_{L^{B'}\leq|t|\leq T_{\mathrm{mod}}}|N_0(t)|^2\,dt\ll L^{2C_{17}-B'}+x^{-c''}. \tag*{(153)}

This has arbitrary log saving after enlarging B′B', even after the small-prime absolute bound. We used square integration of 1/∣t∣1/|t|, not a pointwise log saving multiplied by the full interval length.

For Tmod<∣t∣≤2T0T_{\mathrm{mod}}<|t|\leq2T_0, apply Lemma 2.4 to each such progression. Its modulus is at most LC6L^{C_6}, its scale is x1−o(1)x^{1-o(1)}, and, including retained shifts, its frequency is at least 12Tmod>exp⁡(L/(2(log⁡L)2))\frac12T_{\mathrm{mod}}>\exp(L/(2(\log L)^2)) and at most x1+o(1)<4x3x^{1+o(1)}<4x^3. After summing the reciprocal divisor coefficients this yields

∣N0(t)∣≪LC18exp⁡(−L(log⁡L)C19).(154)|N_0(t)|\ll L^{C_{18}}\exp\left(-\frac{L}{(\log L)^{C_{19}}}\right). \tag*{(154)}

We now integrate only on the exceptional unit intervals. By (150) and the small-prime absolute bound, their total contribution is at most

LO(1)exp⁡(O(L.9)−2L(log⁡L)C19).L^{O(1)}\exp\left(O(L^{.9})-\frac{2L}{(\log L)^{C_{19}}}\right).

which saves every fixed log power. No length-dominated mean-square estimate was required on an individual divisor progression: R′R' may be smaller than T0T_0, and the bounds (6.40)–(154) are pointwise. Equations (152), (153), and (154) complete the constant-term estimate.

We have proved (138) to every fixed log precision. The local Mellin inequality gives (136); the comparison multiplier estimate and Parseval then bound each term in (132), divided by YY, by an arbitrarily large negative log power. Choose that precision after the known fixed kernel, dyad, arc, and separation costs. Their total is therefore negligible. Subtracting these comparison terms from the residual pairing leaves its raw pattern, already identified with (120). This proves Theorem 6.1.

Extraction of the prime statistic

Fix d≥1d \ge1 and ε>0\varepsilon> 0. For each i≤di \le d, let Ii\mathcal{I}_i be an interval of primes with lower endpoint at least xεx^\varepsilon, and suppose that the product of the upper endpoints is at most x1−εx^{1-\varepsilon}. Define

fx(u)=∑ℓi∈Iiℓ1,…,ℓd distinct1ℓ1⋯ℓd∣u,Cx=∑ℓi∈Iiℓ1,…,ℓd distinct1ℓ1⋯ℓd.(155)f_x(u)=\sum_{\substack{\ell_i\in\mathcal{I}_i\\ \ell_1,\ldots,\ell_d\ \mathrm{distinct}}}\mathbf{1}_{\ell_1\cdots\ell_d\mid u}, \qquad C_x=\sum_{\substack{\ell_i\in\mathcal{I}_i\\ \ell_1,\ldots,\ell_d\ \mathrm{distinct}}}\frac{1}{\ell_1\cdots\ell_d}. \tag*{(155)}

Lemma 2.1 gives Cx=Od,ε(1)C_x=O_{d,\varepsilon}(1). On any fixed range u≍xu\asymp x, the same bound holds for fx(u)f_x(u): there are only Oε(1)O_\varepsilon(1) prime divisors of uu exceeding xεx^\varepsilon. All cutoffs below are supported in a fixed compact subset of (0,∞)(0,\infty).

Theorem 7.1 (Interior prime statistic). For every nonnegative Φ∈Cc∞((0,∞))\Phi\in C_c^\infty((0,\infty)),

∑p primeΦ((p−1)/x)(fx(p−1)−Cx)=o(x/log⁡x).(156)\sum_{p\ \mathrm{prime}}\Phi((p-1)/x)\left(f_x(p-1)-C_x\right)=o(x/\log x). \tag*{(156)}

The limit holds through all real x→∞x\to\infty for the slot intervals just described.

Throughout the proof write L=log⁡xL=\log x, W=exp⁡(L.24)W=\exp(L^{.24}), and V(W)=∏p≤W(1−1/p)V(W)=\prod_{p\le W}(1-1/p). Fix q,q0∈(0,1)q,q_0\in(0,1), for example q=q0=1/2q=q_0=1/2. A marking array will mean KK disjoint bands from Theorem 3.1, with their weight

W(u)=∏i=1Kqωi(u)−1ωi(u)Vi,Vi=∑p∈Pi1p.(157)\mathcal{W}(u)=\prod_{i=1}^{K}q^{\omega_i(u)-1}\frac{\omega_i(u)}{V_i}, \qquad V_i=\sum_{p\in\mathcal{P}_i}\frac{1}{p}. \tag*{(157)}

Each band consists of all primes between exp⁡(Lai)\exp(L^{a_i}) and exp⁡(2Lai)\exp(2L^{a_i}), with .1<ai<.2.1<a_i<.2. The number KK and the exponents are fixed in each limit. Since Vi=log⁡2+o(1)V_i=\log2+o(1) and sup⁡n≥0nqn−1<∞\sup_{n\ge0}nq^{n-1}<\infty, these weights are bounded by a constant depending only on K,qK,q, for large xx.

Presieving u+1u+1 leaves both primes and composites in the average of W(u)(fx(u)−Cx)\mathcal{W}(u)(f_x(u)-C_x). We will show that the composite contribution is negligible in two stages. The one-sided estimate first replaces its prime factors by rough variables; additional divisor marks then allow the two-sided estimate to bound the resulting centered sum. This isolates a prime average carrying W\mathcal{W}. Finally, a mean-square bound for averages of disjoint marking arrays removes that weight.

Progressions and the independent divisor model

For finitely many disjoint arrays, define the independent divisor model by replacing every 1p∣u\mathbf{1}_{p\mid u} at a band prime with an independent Bernoulli variable of mean 1/p1/p. For a function HH of these indicators write Mx(H)M_x(H) for its expectation in this model. In particular put mW=Mx(W)m_{\mathcal{W}}=M_x(\mathcal{W}). There are constants 0<cK<CK<∞0<c_K<C_K<\infty such that

cK≤mW≤CK,Mx(W2)≤CK.(158)c_K\le m_{\mathcal{W}}\le C_K,\qquad M_x(\mathcal{W}^2)\le C_K. \tag*{(158)}

For the lower bound, in each band the event of exactly one hit has probability ∏p(1−1/p)∑p1/(p−1)\prod_p(1-1/p)\sum_p1/(p-1), bounded below since the reciprocal sum tends to log⁡2\log2 and the largest 1/p1/p tends to zero. The weights have a fixed positive value on the event of one hit per band. Independence between bands proves the lower bound; boundedness proves the other two assertions. These constants can be chosen independently of the particular fixed exponents.

Lemma 7.2 (Summed progression remainder). Choose 0<c0<ε/20<c_0<\varepsilon/2. Let HH be a weight from a fixed finite collection of disjoint arrays, a product of two such weights (allowing a repeated array), a constant, or a bounded linear combination of these functions. Then, for every fixed A>0A>0,

∣∑r≤xc0 ∑u≥1r∣u+1Φ(u/x)fx(u)H(u)−xr(∫0∞Φ(t) dt)CxMx(H)∣≪AxL−A.(159)\left|\sum_{r\le x^{c_0}}\ \sum_{\substack{u\ge1\\r\mid u+1}}\Phi(u/x)f_x(u)H(u)-\frac{x}{r}\left(\int_0^\infty\Phi(t)\,dt\right)C_xM_x(H)\right|\ll_A xL^{-A}. \tag*{(159)}

The same assertion holds with fx,Cxf_x,C_x replaced by 1,11,1. The constants may depend on the fixed arrays and on the bounded coefficients in the linear combination.

Proof. We give the expansion including the repeated-array case needed later. It suffices to treat one product of weights. In any band gg which occurs, let ag∈{1,2}a_g\in\{1,2\} be its multiplicity in the product. Its factor is

Vg−agqag(ωg(u)−1)ωg(u)ag.V_g^{-a_g}q^{a_g}( \omega_g(u)-1)\omega_g(u)^{a_g}.

Expand ωgag\omega_g^{a_g} as the sum over ordered aga_g-tuples of divisor primes, permitting coincidences. For a chosen tuple let SgS_g be its set of distinct primes. On the event Sg∣uS_g\mid u, where this notation means that every prime in SgS_g divides uu, the contribution is

Vg−agqag(∣Sg∣−1)1Sg∣u∏p∈Pg∖Sg(1−(1−qag)1p∣u).(160)V_g^{-a_g}q^{a_g}(|S_g|-1)\mathbf{1}_{S_g\mid u}\prod_{p\in\mathcal{P}_g\setminus S_g}\left(1-(1-q^{a_g})\mathbf{1}_{p\mid u}\right). \tag*{(160)}

Thus repetitions change the damping to q2q^2 on the unrepresented primes. For instance, the reciprocal mass of the represented unions when ag=2a_g=2 is exactly

Vg−2{Vg+q2(Vg2−∑p∈Pgp−2)},V_g^{-2}\left\{V_g+q^2\left(V_g^2-\sum_{p\in\mathcal{P}_g}p^{-2}\right)\right\},

and is bounded. For ag=1a_g=1 that mass is one. Multiplying over the fixed number of bands shows that all represented unions have total reciprocal mass O(1)O(1), with their nonnegative coefficients and their multiplicities included.

Expand the remaining product in (7.6) through degree h0=⌊L.01⌋h_0=\lfloor L^{.01}\rfloor in the unrepresented primes. Bonferroni inequalities, valid for factors 1−tp1-t_p with 0≤tp≤10\le t_p\le1, bound the error in absolute value by the nonnegative term of degree h0+1h_0+1. If C∗C_* bounds the sum of the reciprocal masses of all bands involved, the total reciprocal coefficient mass at degree tt is at most a fixed constant times C∗t/t!C_*^t/t!. Hence the truncated expansion has reciprocal coefficient mass O(1)O(1) and its error envelope has reciprocal mass at most

O(C∗h0+1(h0+1)!)=OA(L−A)O\left(\frac{C_*^{h_0+1}}{(h_0+1)!}\right)=O_A(L^{-A})

for every fixed AA. {#eq:7.7}

Every divisor in the truncation or the error envelope satisfies

d′≤exp⁡(OK((h0+1)L.2))≤exp⁡(OK(L.22))=xo(1).(161)d' \le\exp\left(O_K\left((h_0+1)L^{.2}\right)\right)\le\exp\left(O_K\left(L^{.22}\right)\right)=x^{o(1)}. \tag*{(161)}

It follows also that the ordinary absolute coefficient mass is xo(1)x^{o(1)}: multiply its reciprocal mass by the largest divisor. All these facts hold as well for the constant function.

Fix an ordered long-prime tuple, put ℓ=∏iℓi\ell=\prod_i\ell_i, and write u=ℓyu=\ell y. Every band prime is coprime to ℓ\ell. Moreover, every ℓi>xc0\ell_i>x^{c_0}, so (ℓ,r)=1(\ell,r)=1 for r≤xc0r\le x^{c_0}. The congruence r∣ℓy+1r\mid\ell y+1 therefore fixes one reduced residue class of yy modulo rr. For a divisor d′d' of band primes the additional condition d′∣yd'\mid y is impossible if (d′,r)>1(d',r)>1. Otherwise the Chinese remainder theorem and smooth summation in one progression give

∑y≥1ℓy≡−1(modr)d′∣yΦ(ℓy/x)=xℓrd′∫0∞Φ(t) dt+OΦ(1).(162)\sum_{\substack{y\ge1\\ \ell y\equiv-1\pmod r\\ d'\mid y}}\Phi(\ell y/x)=\frac{x}{\ell r d'}\int_0^\infty\Phi(t)\,\mathrm{d}t+O_\Phi(1). \tag*{(162)}

The error follows, for example, by comparing each sampling interval with its integral and using the bounded total variation of Φ\Phi. It is uniform in ℓ,r,d′\ell,r,d'.

Inserting the expansion into (162) recovers the model in which the band primes dividing rr are forced absent. The same calculation for the positive Bonferroni envelope bounds the truncation error by eq:7.7 times x/(ℓr)x/(\ell r), plus xo(1)x^{o(1)} in rounding errors. There are at most d!x1−εd!x^{1-\varepsilon} ordered long-prime tuples, since a squarefree product determines at most d!d! orders. Consequently the sum of all rounding errors over tuples and r≤xc0r\le x^{c_0} is

O(x1−ε+c0+o(1))=O(x1−ε/3),(163)O\left(x^{1-\varepsilon+c_0+o(1)}\right)=O\left(x^{1-\varepsilon/3}\right), \tag*{(163)}

after increasing the threshold for xx. When fx=1f_x=1, use ℓ=1\ell=1 and the stronger bound xc0+o(1)x^{c_0+o(1)} instead. The principal Bonferroni errors sum to at most O(xCxLC∗h0+1/(h0+1)!)O(xC_xLC_*^{h_0+1}/(h_0+1)!) and save every fixed log power.

Finally, coupling the forced-absence model to the unconditional one and using the boundedness of HH bounds their mean difference by

O(∑p∣rp in the arrays1p)≪Lexp⁡(−L1).O\left(\sum_{\substack{p\mid r\\ p\ \mathrm{in\ the\ arrays}}}\frac{1}{p}\right)\ll L\exp(-L^1).

Its contribution after summing 1/ℓ1/\ell and 1/r1/r is smaller than every xL−AxL^{-A}. This proves eq:7.5; bounded linear combinations follow by the triangle inequality.

Presieving and replacement of composite prime slots

Fix a large even integer hh, and set

γ=c04h+3,z=xγ.(164)\gamma= \frac{c_0}{4h+3}, \qquad z=x^\gamma. \tag*{(164)}

Apply the block sieve, Lemma 2.6, to the bad conditions r′∣u+1r' \mid u+1 at primes r′≤zr' \le z, with density g(r′)=1/r′g(r')=1/r'. Use separately the two nonnegative weights Φ(u/x)W(u)fx(u)\Phi(u/x)W(u)f_x(u) and Φ(u/x)W(u)\Phi(u/x)W(u). The density is at most 1/21/2, and its reciprocal mass in every block (b,b2](b,b^2] is bounded by an absolute constant by Lemma 2.1. The sieve remainder is bounded by the sum in Lemma 7.2, because

z4h+2=xc0(4h+2)/(4h+3)<xc0.(165)z^{4h+2}=x^{c_0(4h+2)/(4h+3)}<x^{c_0}. \tag*{(165)}

Subtract the second sieve formula times CxC_x from the first. Their main terms cancel, while each relative sieve error is O(e−h)O(e^{-h}). Since the main terms both contain mWmW and ∏r′≤z(1−1/r′)≍1/(γL)\prod_{r'\le z}(1-1/r')\asymp1/(\gamma L), we obtain

lim sup⁡x→∞LxmW∣∑u≥1P−(u+1)>zΦ(u/x)W(u)(fx(u)−Cx)∣≪e−hγ.(166)\limsup_{x\to\infty}\frac{L}{xmW}\left|\sum_{\substack{u\ge1\\P^-(u+1)>z}}\Phi(u/x)W(u)(f_x(u)-C_x)\right|\ll\frac{e^{-h}}{\gamma}. \tag*{(166)}

The implied constant depends on the slot and cutoff data but is independent of KK, hh and of the fixed array. The progression remainders may have constants depending on these choices; their normalized limits are zero. In particular, no reciprocal of a rare-mark mean occurs in the surviving error in (7.13).

For integers in this sum that are composite, write their ordered prime factorization as u+1=p1⋯pju+1=p_1\cdots p_j, with pi>zp_i>z. For large xx, 2≤j≤2/γ2\le j\le2/\gamma. Integers divisible by p2p^2 for a prime p>zp>z number

≪∑p>zxp2≪x/z.\ll\sum_{p>z}\frac{x}{p^2}\ll x/z.

The weights and the possible ordered factorization counts are bounded for fixed γ,d,K\gamma,d,K. We may therefore discard these integers at error o(x/L)o(x/L), count the others with weight 1/j!1/j!, and restore repeated factors when convenient at the same cost. For each fixed jj we will prove

∑p1,…,pj>zΦ((p1⋯pj−1)/x)W(p1⋯pj−1)(fx(p1⋯pj−1)−Cx)=o(x/L).(167)\sum_{p_1,\ldots,p_j>z}\Phi((p_1\cdots p_j-1)/x)W(p_1\cdots p_j-1)(f_x(p_1\cdots p_j-1)-C_x)=o(x/L). \tag*{(167)}

Let

κ(m)=1V(W)log⁡m(m≥2),a(m)=1m>z1P−(m)>Wκ(m).(168)\kappa(m)=\frac{1}{V(W)\log m}\quad(m\ge2),\qquad a(m)=1_{m>z}1_{P^-(m)>W}\kappa(m). \tag*{(168)}

Set κ(1)=0\kappa(1)=0; values at one are always excluded by the positive-power threshold tests. We first replace every prime variable in (7.14) by this coefficient. Telescope one variable mm at a time and write nn for the product of the other j−1j-1 variables. Dyadically decompose m,nm,n on the support of Φ((mn−1)/x)\Phi((mn-1)/x). There are O(L2)O(L^2) boxes, with HmHn≍xH_mH_n\asymp x, and both scales are at least a fixed positive power of xx. The restriction m>zm>z only shortens its dyadic interval, as permitted in Theorem 3.1. The companion coefficient βn\beta_n is supported on WW-rough integers and is a fixed convolution of prime indicators and proxy coefficients. It is bounded by LCjτj−1(n)L^{C_j}\tau_{j-1}(n) for a fixed exponent CjC_j.

The divisor moment bound in Lemma 2.2 permits truncation of βn\beta_n at LBL^B, for a sufficiently large fixed BB, with an arbitrarily large log saving in its ℓ1\ell^1 norm. More explicitly, for any fixed integer t>1t>1,

∑n≍Hn∣βn∣1∣βn∣>LB≤L−B(t−1)∑n≍Hn∣βn∣t≪HnLCjt−B(t−1).\sum_{n\asymp H_n}|\beta_n|1_{|\beta_n|>L^B}\le L^{-B(t-1)}\sum_{n\asymp H_n}|\beta_n|^t\ll H_nL^{C_jt-B(t-1)}.

Multiplying by the number and maximum size of the mm coefficients and by the bounded endpoint factor shows that the discarded terms are O(xL−A)O(xL^{-A}) for any prescribed AA, on taking BB large enough. This truncation does not change rough support.

Replacing Φ((mn−1)/x)\Phi((mn-1)/x) by Φ(mn/x)\Phi(mn/x) costs O(x−1)O(x^{-1}) per coefficient, hence only a fixed log power in the whole sum; it is negligible compared with x/Lx/L. Fourier inversion of the smooth function in log⁡(mn/x)\log(mn/x) separates the two variables into factors mivm^{iv} and nivn^{iv}, with a rapidly decreasing Fourier density. Its tails may be discarded past a fixed log power, with any desired saving, using the same coefficient bounds. On the remaining frequencies the discrepancy in the mm slot is exactly the one in Theorem 3.1.

To meet its invariance hypothesis globally, clip fx(u)−Cxf_x(u)-C_x to a fixed interval containing all its values on the support under consideration, and divide by a fixed bound so that its modulus is at most one. Every long-slot prime lies outside the array bands, so fx(pu)=fx(u)f_x(pu)=f_x(u) for each band prime pp; clipping preserves this identity. Thus the Type II theorem applies with FF equal to this normalized clipped function. Choose its log-saving target after all the fixed dyadic, Fourier, and truncation losses. We then choose KK sufficiently large for that target, simultaneously for all 2≤j≤2/γ2 \le j \le2/\gamma. This choice is made after hh, γ\gamma, and is independent of the particular fixed band exponents. Telescoping all slots gives

∑p1,…,pj>zΦ((p1⋯pj−1)/x)W(p1⋯pj−1)(fx(p1⋯pj−1)−Cx)=∑m1,…,mjΦ((m1⋯mj−1)/x)W(m1⋯mj−1)(fx(m1⋯mj−1)−Cx)∏i=1ja(mi)+o(x/L).(169)\sum_{p_1,\ldots,p_j>z}\Phi((p_1\cdots p_j-1)/x)W(p_1\cdots p_j-1)(f_x(p_1\cdots p_j-1)-C_x) = \sum_{m_1,\ldots,m_j}\Phi((m_1\cdots m_j-1)/x)W(m_1\cdots m_j-1)(f_x(m_1\cdots m_j-1)-C_x)\prod_{i=1}^{j}a(m_i)+o(x/L). \tag*{(169)}

Removing candidate prime parts

Fix jj in the preceding finite range. Construct candidate small groups from all primes in consecutive bands [exp⁡(y),exp⁡(2y))[\exp(y),\exp(2y)), starting at y=L.28y=L^{.28} and doubling yy while 2y≤L.352y\le L^{.35}. Construct candidate big groups in the same way, from y=L.40y=L^{.40} while 2y≤L.452y\le L^{.45}. There are between positive constant multiples of log⁡L\log L groups of each kind. They are disjoint, their reciprocal sums VgV_g are bounded above and below by positive constants, and they lie in the allowed small and big ranges of Theorem 6.1. The big groups are full prime intervals. All candidate primes exceed WW.

We will group factorizations that differ only in the allocation of candidate primes among their jj factors. Their coefficients must agree, so we first remove all candidate prime parts from the threshold and logarithm in a(m)a(m). For any integer mm, let mˉ\bar m be obtained by removing the entire prime-power parts belonging to all candidate groups. Replace a(m)a(m) by

a~(m)=1P−(m)>W1mˉ>xγκ(mˉ).(170)\widetilde{a}(m)=\mathbf{1}_{P_-(m)>W}\mathbf{1}_{\bar m>x^\gamma}\kappa(\bar m). \tag*{(170)}

We show that this changes the right-hand side of (7.16) by o(x/L)o(x/L) in absolute value. In particular, the estimate will remain valid before any cancellation is used.

For these error estimates, decompose tuples into full product dyads mi≍Him_i\asymp H_i, with ∏iHi≍x\prod_i H_i\asymp x and Hi≥xγ/2H_i\ge x^{\gamma/2}. There are O(Lj−1)O(L^{j-1}) such boxes. By Lemma 2.5, the number of WW-rough integers in each dyad is O(HiV(W))O(H_iV(W)). On an unconditioned integer dyad,

∑m≍Hilog⁡(m/mˉ)≤∑p candidate∑a≥1(log⁡p)⌊2Hipa⌋≪Hi∑p≤exp⁡(L.45)log⁡pp−1≪HiL.45.(171)\sum_{m\asymp H_i}\log(m/\bar m) \le\sum_{p\ \mathrm{candidate}}\sum_{a\ge1}(\log p)\left\lfloor\frac{2H_i}{p^a}\right\rfloor \ll H_i\sum_{p\le\exp(L^{.45})}\frac{\log p}{p-1}\ll H_iL^{.45}. \tag*{(171)}

Thus the number with log⁡(m/m‾)>L.9\log(m/\overline{m}) > L^{.9} is O(HiL−.45)O(H_iL^{-.45}). On the support of either coefficient in the comparison its size is O((LV(W))−1)O((LV(W))^{-1}). Taking the bad count in one coordinate and the rough counts in all others bounds the bad-tuple mass in a box by

O(xLjL−.45V(W))=O(xL−j−.21),(172)O\left(\frac{x}{L^j}\frac{L^{-.45}}{V(W)}\right)=O(xL^{-j-.21}), \tag*{(172)}

since V(W)≍L−.24V(W)\asymp L^{-.24}. After summing boxes this is O(xL−1−.21)O(xL^{-1-.21}). For the other tuples, a change in the threshold test requires γL<log⁡mi≤γL+L.9\gamma L < \log m_i \leq\gamma L+L^{.9} for at least one coordinate. There are O(L.9Lj−2)O(L^{.9}L^{j-2}) product dyads meeting such a strip, because j≥2j\geq2. Their positive mass is O(x/Lj)O(x/L^j) per box, so the total is O(xL−1.1)O(xL^{-1.1}). Away from those strips the supports coincide, and

κ(mi‾)κ(mi)log⁡milog⁡mi‾=1+Oγ(L−1).\frac{\kappa(\overline{m_i})}{\kappa(m_i)}\frac{\log m_i}{\log\overline{m_i}}=1+O_\gamma(L^{-1}).

The total positive tuple mass is O(x/L)O(x/L), by the same rough counts and box count. Together with the boundedness of W(fx−Cx)\mathcal{W}(f_x-C_x) this proves the required absolute o(x/L)o(x/L) error.

Many successes before the signed expansion

Write v=m1⋯mjv=m_1\cdots m_j and u=v−1u=v-1. For each candidate group gg, independently conditional on the tuple, toss a coin whose success probability is

sg(u,v)=c1q0ωg(u)−1ωg(u)Vg(q0j)ωg(v)−1ωg(v)Vg.(173)s_g(u,v)=c_1q_0^{\omega_g(u)-1}\frac{\omega_g(u)}{V_g}\left(\frac{q_0}{j}\right)^{\omega_g(v)-1}\frac{\omega_g(v)}{V_g}. \tag*{(173)}

The constant c1>0c_1>0 is fixed sufficiently small, depending on jj, q0q_0 and the bounds for VgV_g, that these probabilities are at most one. The factors with no hit are interpreted as zero. The division by jj in the product-side damping anticipates the jj possible allocations of each candidate prime among the factors when the candidate part of vv is squarefree. Summing those allocations will recover the q0q_0 damping of the two-sided estimate. We claim that for some fixed c2>0c_2>0, outcomes with fewer than c2log⁡Lc_2\log L successes in either family have total positive weighted mass o(x/L)o(x/L) under the coefficient ∏ia~(mi)\prod_i\widetilde{a}(m_i).

We prove the uniformity needed for this claim on full rough product dyads. Select each mim_i independently and uniformly among WW-rough integers in its full dyad. For any fixed number of distinct candidate primes, their product QQ satisfies log⁡Q=O(L.45)\log Q=O(L^{.45}). Each prescribed residue class modulo QQ is relatively equidistributed among these rough integers, with error o(1)o(1) uniformly in the classes, primes, and dyads. To verify this, apply Bonferroni to divisibility by primes p≤Wp\leq W at depth Dlog⁡LD\log L, where DD is a sufficiently large fixed constant. The reciprocal sum of those primes is O(log⁡L)O(\log L), so taking DD large makes the omitted reciprocal mass smaller than any prescribed log power times V(W)V(W). The moduli in the retained terms are at most

QWO(log⁡L)=exp⁡(O(L.45+L.24log⁡L))=xo(1).QW^{O(\log L)}=\exp(O(L^{.45}+L^{.24}\log L))=x^{o(1)}.

Since every factor dyad has positive-power length, CRT counting errors divided by its expected class count tend to zero uniformly. Normalizing by the similarly evaluated total rough count proves the assertion, with arbitrary fixed logarithmic precision if needed.

In the independent residue model, for a candidate prime pp the events p∣up\mid u and p∣vp\mid v are disjoint, with probabilities

ap=(p−1)j−1pj,bp=1−(1−1/p)j.(174)a_p=\frac{(p-1)^{j-1}}{p^j},\qquad b_p=1-(1-1/p)^j. \tag*{(174)}

respectively. For the first formula all residues must be nonzero, and the first j−1j-1 determine the last. For the second, at least one of the jj residues must be zero. Distinct primes are independent in this model. In each group, ∑pap=Vg+o(1)\sum_p a_p = V_g + o(1) and ∑pbp=jVg+o(1)\sum_p b_p = jV_g + o(1), while the largest individual probability tends to zero. Summing the uniform residue estimates proves convergence of every fixed joint factorial moment of the two hit counts, for any one or two groups, uniformly over their choices and over product dyads. These moments are bounded by CrC^r at total order rr, with CC depending only on the fixed jj and the reciprocal group bounds.

Here is why fixed moments suffice for the bounded probabilities in (173), without any growing-moment assumption. First cut each of the at most four counts at a fixed bound aa. The bounded first moments give tail probability O(1/a)O(1/a) uniformly. For exact counts below this bound, list the required hits and impose the absence of further hits by Bonferroni. At a fixed depth bb, the limit superior of the remainder is at most CaCb/b!C_a C^b/b!, by the corresponding factorial moment estimate. Take x→∞x \to\infty, then b→∞b \to\infty, and then a→∞a \to\infty. It follows that the expectations of the bounded one-group tests and their two-group products agree with the model up to a uniform o(1)o(1).

Explicitly, for a count vector N=(N1,…,Nr)N=(N_1,\ldots,N_r) and a fixed target kk, list the kik_i required hits of each type. Conditional on N≥kN \ge k coordinatewise, inclusion–exclusion for no other hit is the alternating binomial sum in Ntot−∣k∣N_{\mathrm{tot}}-|k|, multiplied by ∏i(Niki)\prod_i \binom{N_i}{k_i}. The first omitted term at depth bb is bounded by

(Ntot)∣k∣+b+1(b+1)!∏iki!,\frac{(N_{\mathrm{tot}})_{|k|+b+1}}{(b+1)!\prod_i k_i!},

where (n)r=n(n−1)⋯(n−r+1)(n)_r=n(n-1)\cdots(n-r+1). Its expected limit superior is at most C∣k∣+b+1/((b+1)!∏iki!)C^{|k|+b+1}/((b+1)!\prod_i k_i!). This supplies the stated uniform remainder for every exact count used after the fixed truncation.

In the model, the probability of exactly one hit of each type in a group is bounded below. Indeed it equals

∏p∈g(1−ap−bp)∑p≠qp,q∈gap1−ap−bpbq1−aq−bq,\prod_{p\in g}(1-a_p-b_p)\sum_{\substack{p\ne q\\p,q\in g}}\frac{a_p}{1-a_p-b_p}\frac{b_q}{1-a_q-b_q},

whose product and sum are bounded below using the displayed reciprocal bounds and the vanishing largest atom. On this event the conditional success probability is c1/Vg2c_1/V_g^2, also bounded below. Distinct groups are independent in the model. Therefore, if IgI_g denotes the actual coin indicator, there are β>0\beta>0 and ϵx→0\epsilon_x\to0, uniform in the dyads and groups, such that

EIg≥β,∣Cov⁡(Ig,Ig′)∣≤ϵx(g≠g′).(175)\mathbb{E}I_g\ge\beta,\qquad\left|\operatorname{Cov}(I_g,I_{g'})\right|\le\epsilon_x\quad(g\ne g'). \tag*{(175)}

Conditional independence of the coins identifies the second quantity with the covariance of their success probabilities. For a family of N≳log⁡LN\gtrsim\log L groups, Chebyshev gives

P(∑gIg<βN/2)≤4β2N+4ϵxβ2=o(1).(176)\mathbb{P}\left(\sum_g I_g<\beta N/2\right)\le\frac{4}{\beta^2N}+\frac{4\epsilon_x}{\beta^2}=o(1). \tag*{(176)}

Only two families occur. Notice that a uniform o(1)o(1) covariance suffices; no rate relative to 1/log⁡L1/\log L is required.

Finally transfer this full-dyad probability to the actual weighted tuples. A rough product box contains O(xV(W)j)O(xV(W)^j) tuples, and ∏ia~(mi)≪(LV(W))−j\prod_i\tilde a(m_i)\ll(LV(W))^{-j} on its support. Thus its failure mass is o(x/L)o(x/L) uniformly. There are O(Lj−1)O(L^{j-1}) boxes, proving the claimed o(x/L)o(x/L) total. The bounded factor Φ(u/x)ν(u)(fx(u)−Cx)\Phi(u/x)\nu(u)(f_x(u)-C_x) does not alter this conclusion. We discard these failed outcomes now, while their weights form a nonnegative probability partition.

Allocation and the two-sided correlation

Let C\mathcal{C} be the full candidate set of groups. The coin partition is

1=∑S⊆C∏g∈Ssg∏g∉S(1−sg).1=\sum_{S\subseteq\mathcal{C}}\prod_{g\in S}s_g\prod_{g\notin S}(1-s_g).

Retain only SS containing at least c2log⁡Lc_2\log L groups of each kind, at the error just proved. Now expand the failure factors in each retained term. Each expanded term has the form

±∏g∈Psg,P=S∪T,T⊆C∖S.(177)\pm\prod_{g\in P}s_g,\qquad P=S\cup T,\quad T\subseteq\mathcal{C}\setminus S. \tag*{(177)}

There are at most 3∣C∣=LO(1)3^{|\mathcal{C}|}=L^{O(1)} terms; their coefficients, including c1∣P∣c_1^{|P|}, have fixed log-power bounds. Every active set PP has between fixed positive multiples of log⁡L\log L groups of each kind and satisfies all the active-group hypotheses of Theorem 6.1. This expansion is performed after the failed mass has been removed, so it never multiplies the preceding qualitative o(x/L)o(x/L) error.

For a fixed active set let sP=∣P∣s_P=|P|, let v∗v_* be obtained from vv by removing the entire parts of its active primes, and use W1,q0W_{1,q_0} to denote the active one-mark weight with damping q0q_0. We may first restrict to vv not divisible by the square of any candidate prime. Such square exceptions have integer count

O(x∑p candidatep−2)≪xexp⁡(−cL.28).(178)O\left(x\sum_{p\text{ candidate}}p^{-2}\right)\ll x\exp(-cL^{.28}). \tag*{(178)}

Their weighted tuple sums remain negligible after all fixed log-power losses. Indeed the number of ordered factor tuples of vv is at most τj(v)\tau_j(v); Cauchy–Schwarz and a fixed divisor moment bound give an exponential saving in a power of LL for sums of any fixed divisor power over this exceptional set. All active marked factors are bounded by fixed log powers, because there are O(log⁡L)O(\log L) groups with exponential damping in each. The same argument will allow these exceptions to be restored at the end of the calculation.

Define the core

GP(v)=∑e1⋯ej=v∗∏i=1j[1P−(ei)>W1ei‾>xγκ(ei‾)].(179)G_P(v)=\sum_{e_1\cdots e_j=v_*}\prod_{i=1}^{j}\left[1_{P_-(e_i)>W}1_{\overline{e_i}>x^\gamma}\kappa(\overline{e_i})\right]. \tag*{(179)}

Here the bar still removes all candidate prime parts, including the inactive ones. This definition implies GP(v)=GP(v∗)G_P(v)=G_P(v_*) and

∣GP(v)∣≤(C/(LV(W)))jτj(v∗)≤LCjτ(v∗)Cj.(180)|G_P(v)|\leq(C/(L V(W)))^j\tau_j(v_*)\leq L^{Cj}\tau(v_*)^{Cj}. \tag*{(180)}

The second bound follows from the elementary inequality τj(n)≤τ(n)j−1\tau_j(n)\leq\tau(n)^{j-1}, obtained prime by prime.

Off the square exceptions, fix a factorization e1⋯ej=v∗e_1\cdots e_j=v_*. Every active prime dividing vv can be assigned independently to any of the jj slots. These are all the original factorizations with that stripped factorization, and all have the same coefficient in (7.17): active primes exceed WW and the tests and logarithms depend only on the bars. Thus there are jωP(v)j^{\omega_P(v)} equal contributions. The exact identity

jωP(v)(q0/j)ωP(v)−sP=jsPq0ωP(v)−sP(181)j^{\omega_P(v)}(q_0/j)^{\omega_P(v)-s_P}=j^{s_P}q_0^{\omega_P(v)-s_P} \tag*{(181)}

shows that summing the vv-side factor in (7.24) over factorizations gives

jsPW1,q0(v)GP(v).j^{s_P}W_{1,q_0}(v)G_P(v).

This is an identity of weights, including the mark factors ∏gωg(v)/Vg\prod_g\omega_g(v)/V_g. Restoring the square exceptions by (178) and its divisor-moment bound, each expanded term is, up to a negligible error, a constant of fixed log-power size times

∑u≥1Φ(u/x)W(u)(fx(u)−Cx)W1,q0(u)W1,q0(u+1)GP(u+1).(182)\sum_{u \ge1} \Phi(u/x)\mathcal{W}(u)(f_x(u)-C_x)\mathcal{W}_{1,q_0}(u)\mathcal{W}_{1,q_0}(u+1)G_P(u+1). \tag*{(182)}

Take Ψ(y)=yΦ(y)\Psi(y)=y\Phi(y). Since Φ(u/x)=xΨ(u/x)/u\Phi(u/x)=x\Psi(u/x)/u, Theorem 6.1 bounds (182) by OA(xL−A)O_A(xL^{-A}) for every fixed AA. Its invariant core hypothesis is exactly (179), and its growth hypothesis is (180). Its other endpoint is exactly Fx(u)=W(u)(fx(u)−Cx)F_x(u)=\mathcal{W}(u)(f_x(u)-C_x). Choose AA after the fixed costs from 3∣C∣3^{|\mathcal{C}|}, jSPj^{SP}, and the coefficients. They are all absorbed. Summing the expanded terms and adding back the earlier, unamplified, discarded masses proves that the proxy sum in (169) is o(x/L)o(x/L). This proves (167).

Removal of the small-band marks

Subtract the composite contribution from (166). All primes in the support exceed zz for large xx, and we have proved

lim sup⁡x→∞Lx∣∑p primeΦ((p−1)/x)W(p−1)mW(fx(p−1)−Cx)∣≪e−h/γ.(183)\limsup_{x\to\infty}\frac{L}{x}\left|\sum_{p\ \mathrm{prime}}\Phi((p-1)/x)\frac{\mathcal{W}(p-1)}{m_{\mathcal{W}}}(f_x(p-1)-C_x)\right|\ll e^{-h}/\gamma. \tag*{(183)}

The constant here remains independent of K,hK,h.

Fix this hh and its sufficient fixed KK. For any finite b≥1b\ge1 choose bb disjoint arrays, each with KK distinct exponents, using distinct exponents across all arrays inside (.1,.2)(.1,.2). This is possible for every finite bb, and the Type II choice of KK applies to every one of them. Write Wa,ma\mathcal{W}_a,m_a for their weights and model means, and set

Zb(u)=(1b∑a=1bWa(u)ma−1)2.(184)Z_b(u)=\left(\frac{1}{b}\sum_{a=1}^{b}\frac{\mathcal{W}_a(u)}{m_a}-1\right)^2. \tag*{(184)}

Disjoint arrays are independent in the divisor model. Their normalized weights have mean one and uniformly bounded second moments by eq:7.4. Hence

Mx(Zb)=1b2∑a=1bVar⁡M(Wa/ma)≤CK/b.(185)M_x(Z_b)=\frac{1}{b^2}\sum_{a=1}^{b}\operatorname{Var}_{M}\left(\mathcal{W}_a/m_a\right)\le C_K/b. \tag*{(185)}

The square and all its cross terms lie within Lemma 7.2, whose coefficient bounds are uniform for these normalized weights when b,Kb,K are fixed.

Apply the nonnegative upper bound of the block sieve at the fixed depth h′=2h'=2, with z′=xc0/20z'=x^{c_0/20}, to the weight Φ(u/x)Zb(u)\Phi(u/x)Z_b(u). Its remainder range is

(z′)4h′+2=xc0/2<xc0,(z')^{4h'+2}=x^{c_0/2}<x^{c_0},

so Lemma 7.2 without tuple factors applies. Every prime u+1u+1 in the support exceeds z′z'. The sieve main term is bounded by a fixed constant times xMx(Zb)/log⁡z′xM_x(Z_b)/\log z', and its remainder is o(x/L)o(x/L). It follows that

lim sup⁡x→∞Lx∑p primeΦ((p−1)/x)Zb(p−1)≪CK/b.(186)\limsup_{x\to\infty}\frac{L}{x}\sum_{p\ \mathrm{prime}}\Phi((p-1)/x)Z_b(p-1)\ll C_K/b. \tag*{(186)}

The implied constant here depends on the fixed c0c_0 and cutoff, but not on b,Kb,K; the latter dependence is recorded in CKC_K.

Average (7.30) over the bb arrays. By Cauchy–Schwarz, the difference between this averaged weighted sum and the unweighted sum is at most

(∑pΦ((p−1)/x)Zb(p−1))1/2⋅(∑pΦ((p−1)/x)(fx(p−1)−Cx)2)1/2.\left(\sum_p \Phi((p-1)/x)Z_b(p-1)\right)^{1/2}\cdot\left(\sum_p \Phi((p-1)/x)(f_x(p-1)-C_x)^2\right)^{1/2}.

The second factor is O((x/L)1/2)O((x/L)^{1/2}) by the boundedness of fx−Cxf_x-C_x and the prime number theorem. We conclude that

lim sup⁡x→∞Lx∣∑pΦ((p−1)/x)(fx(p−1)−Cx)∣≤Ce−hγ+CCK/b.(187)\limsup_{x\to\infty}\frac{L}{x}\left|\sum_p\Phi((p-1)/x)(f_x(p-1)-C_x)\right|\leq C\frac{e^{-h}}{\gamma}+C\sqrt{C_K/b}. \tag*{(187)}

All constants in the preceding remainder estimates were allowed to depend on the finite bb; they have disappeared in this real xx limit. First let bb be arbitrarily large with hh, KK fixed. Then let the even integer hh be arbitrarily large, choosing its new sufficient fixed KK each time. Since γ−1=(4h+3)/c0\gamma^{-1}=(4h+3)/c_0, the surviving bound tends to zero. This proves Theorem 7.1.

In particular, the order of choices is: fix the interior data and c0c_0; fix hh, then γ\gamma and the finitely many composite lengths; choose the Type II precision and KK; fix finitely many disjoint exponent arrays; take the limit in real xx; then remove the arrays by b→∞b\to\infty and the sieve error by h→∞h\to\infty. No number of arrays grows with xx in an application of either correlation theorem or the progression lemma.

From interior statistics to the full law

We now pass from Theorem 7.1 to ordinary prime averages of every finite collection of ranked factors. The argument also shows why no additional assertion about very small factors or the boundary of a simplex is needed. The probability argument uses the size-biased viewpoint of Donnelly and Grimmett [8] (Section 2) and the factorial-measure characterization of Arratia, Kochman and Miller [2] (Lemma 2 and Section 3.3). We give the passage, including boundary control and sorting, in full.

Factorial measures in the open simplex

Choose a prime pp uniformly from x<p≤2xx<p\leq2x, and regard the prime factors of p−1p-1, with their multiplicities, as distinct labelled balls. A ball corresponding to a prime qq has mass

t=log⁡qlog⁡(p−1).t=\frac{\log q}{\log(p-1)}.

The masses of all the balls sum to one. For d≥1d\geq1, let μx,d\mu_{x,d} be the expected counting measure of ordered dd-tuples of distinct balls, mapped to their masses. Distinct balls may correspond to the same numerical prime. Set

Dd={(t1,…,td):ti>0, ∑i=1dti<1}.(188)D_d=\left\{(t_1,\ldots,t_d):t_i>0,\ \sum_{i=1}^{d}t_i<1\right\}. \tag*{(188)}

We first prove the local convergence

∫Ddφ dμx,d⟶∫Ddφ(t)∏i=1ddtiti(φ∈Cc(Dd)).(189)\int_{D_d}\varphi\,\mathrm{d}\mu_{x,d}\longrightarrow\int_{D_d}\varphi(t)\prod_{i=1}^{d}\frac{\mathrm{d}t_i}{t_i}\qquad(\varphi\in C_c(D_d)). \tag*{(189)}

Start with a rectangle

R=∏i=1d(ai,bi],0<ai<bi,∑ibi<1.(190)R = \prod_{i=1}^{d} (a_i,b_i], \qquad0 < a_i < b_i, \qquad\sum_i b_i < 1. \tag*{(190)}

Use in Theorem 7.1 the prime slots xai<ℓi≤xbix^{a_i} < \ell_i \leq x^{b_i}. They satisfy its hypotheses for a fixed positive ε\varepsilon. The reciprocal sum over distinct numerical primes obeys

Cx=∑ℓi in their slotsℓi distinct1ℓ1⋯ℓd⟶∏i=1dlog⁡biai.(191)C_x = \sum_{\substack{\ell_i\ \text{in their slots}\\ \ell_i\ \text{distinct}}} \frac{1}{\ell_1 \cdots\ell_d} \longrightarrow\prod_{i=1}^{d} \log\frac{b_i}{a_i}. \tag*{(191)}

Indeed, Lemma 2.1 gives the limit in each slot. Terms with a coincidence in two slots have total reciprocal mass O(∑q≥xmin⁡aiq−2)=o(1)O(\sum_{q \geq x^{\min a_i}} q^{-2}) = o(1), times bounded reciprocal sums from the other slots.

To remove the smooth cutoff in Theorem 7.1, approximate the indicator of the prime dyad from above and below by fixed smooth functions of (p−1)/x(p-1)/x. The tuple count is bounded by a constant depending on dd and min⁡ai\min a_i, because an integer of size O(x)O(x) has at most O(1/min⁡ai)O(1/\min a_i) prime factors exceeding xmin⁡aix^{\min a_i}. The prime number theorem bounds the contribution of endpoint strips of relative width η\eta by O(ηx/log⁡x)+o(x/log⁡x)O(\eta x/\log x) + o(x/\log x). First let real xx tend to infinity with the cutoffs fixed, and then let η\eta decrease to zero. Since π(2x)−π(x)∼x/log⁡x\pi(2x) - \pi(x) \sim x/\log x, this proves the rectangle formula with masses initially normalized by log⁡x\log x.

The required denominator is log⁡(p−1)\log(p-1). Uniformly on the dyad,

log⁡(p−1)log⁡x=1+O(1log⁡x).(192)\frac{\log(p-1)}{\log x} = 1 + O\left(\frac{1}{\log x}\right). \tag*{(192)}

For fixed sufficiently small η>0\eta> 0, a rectangle with endpoints ai+η,bi−ηa_i+\eta,b_i-\eta on the log⁡x\log x scale is therefore contained in the test in (190) on the exact scale; the latter is contained in the rectangle with endpoints ai−η,bi+ηa_i-\eta,b_i+\eta. Choose η\eta small enough that all endpoints stay positive and the enlarged upper endpoints still sum to less than one. Apply the preceding rectangle limits, then let η\eta decrease to zero. This proves the same formula with the exact denominator, still for distinct numerical primes.

The distinction between numerical primes and labelled balls is negligible locally. If every tested coordinate is at least a>0a > 0, a difference requires q2∣p−1q^2 \mid p-1 for some q≥xa/2q \geq x^{a/2}, for all sufficiently large xx. The number of possible integers p−1p-1 in the dyad is at most

∑q≥xa/2⌊2xq2⌋≪x1−a/2.\sum_{q \geq x^{a/2}} \left\lfloor\frac{2x}{q^2} \right\rfloor\ll x^{1-a/2}.

Here and below a sum over primes may be bounded by the corresponding sum over integers. Each such integer contributes at most a constant number of tested tuples: there are at most 1/a1/a balls of mass at least aa. Division by π(2x)−π(x)\pi(2x) - \pi(x) makes (8.6) negligible. We have proved

μx,d(R)⟶∏ilog⁡(bi/ai)=∫R∏idtiti.(193)\mu_{x,d}(R) \longrightarrow\prod_i \log(b_i/a_i) = \int_R \prod_i \frac{\mathrm{d}t_i}{t_i}. \tag*{(193)}

For completeness, these restricted rectangles determine all the local limits claimed in (8.2). A compact set in Dd\mathcal{D}_d has all coordinates at least some a>0a > 0 and its coordinate sum at most 1−a1-a, after decreasing aa if necessary. Its μx,d\mu_{x,d}-mass is bounded uniformly by a−da^{-d}, since the underlying masses sum to one. A sufficiently fine rectangular grid on a slightly larger compact set consists of boxes with positive lower endpoints and upper endpoints summing to less than one. Apply (193) to these finitely many boxes. Upper and lower step approximations to a continuous function have an error bounded by its modulus of continuity times the uniformly bounded local mass. Refining the fixed grid after taking the xx-limit proves (189). This argument uses no bound for the total factorial mass near zero.

Size-biased sampling and the boundary

Given the balls, sample them successively without replacement, choosing at each step a remaining ball with probability equal to its mass divided by the total remaining mass. Write Tx,iT_{x,i} for the mass selected at step ii, and set all subsequent values to zero after the finite collection is exhausted. Let νx,d\nu_{x,d} be the law of (Tx,1,…,Tx,d)(T_{x,1},\ldots,T_{x,d}) on [0,1]d[0,1]^d. For an ordered tuple lying in DdD_d, its conditional probability of being the first dd draws is

wd(t)=∏i=1dti1−t1−⋯−ti−1.(194)w_d(t)=\prod_{i=1}^{d}\frac{t_i}{1-t_1-\cdots-t_{i-1}}. \tag*{(194)}

Consequently, for φ∈Cc(Dd)\varphi\in C_c(D_d),

∫φ dνx,d=∫φwd dμx,d.\int\varphi\,\mathrm{d}\nu_{x,d}=\int\varphi w_d\,\mathrm{d}\mu_{x,d}.

The function wdw_d is continuous and bounded on every compact subset of DdD_d, so (189) gives the limiting interior density

hd(t)=∏i=1d(1−t1−⋯−ti−1)−1.(195)h_d(t)=\prod_{i=1}^{d}(1-t_1-\cdots-t_{i-1})^{-1}. \tag*{(195)}

This density has total mass one, as can be checked directly. Put

Ai=ti1−t1−⋯−ti−1,ti=Ai∏r<i(1−Ar).A_i=\frac{t_i}{1-t_1-\cdots-t_{i-1}},\qquad t_i=A_i\prod_{r<i}(1-A_r).

These formulas give inverse bijections between DdD_d and (0,1)d(0,1)^d. The inverse map is triangular and its Jacobian is

∏i=1d∏r<i(1−Ar)=∏i=1d(1−t1−⋯−ti−1).\prod_{i=1}^{d}\prod_{r<i}(1-A_r)=\prod_{i=1}^{d}(1-t_1-\cdots-t_{i-1}).

It follows that hd(t) dt=dA1⋯dAdh_d(t)\,\mathrm{d}t=\mathrm{d}A_1\cdots\mathrm{d}A_d. In particular,

∫Ddhd(t) dt=1.(196)\int_{D_d}h_d(t)\,\mathrm{d}t=1. \tag*{(196)}

Taking Ui=1−AiU_i=1-A_i identifies this measure with the first dd stick fragments Bi=(∏r<iUr)(1−Ui)B_i=(\prod_{r<i}U_r)(1-U_i) in Theorem 1.1, where the UiU_i are independent uniform random variables.

Equation (196) also excludes any escaped boundary mass. Given η>0\eta>0, choose 0≤χ≤10\le\chi\le1 in Cc(Dd)C_c(D_d) with ∫χhd>1−η\int\chi h_d>1-\eta. Interior convergence yields ∫χ dνx,d>1−2η\int\chi\,\mathrm{d}\nu_{x,d}>1-2\eta for large xx. Thus at most 2η2\eta of the probability lies outside supp⁡χ\operatorname{supp}\chi. This includes all boundary faces, exhaustion events, and configurations with padded zeros. For any continuous HH on [0,1]d[0,1]^d, apply interior convergence to χH\chi H; the integrals of (1−χ)H(1-\chi)H under the two probability measures have total absolute value at most 3η∥H∥∞3\eta\lVert H\rVert_\infty in the limit superior. Letting η\eta decrease to zero proves

(Tx,1,…,Tx,d)⟹(B1,…,Bd)on [0,1]d.(197)(T_{x,1},\ldots,T_{x,d})\Longrightarrow(B_1,\ldots,B_d)\quad\text{on }[0,1]^d. \tag*{(197)}

Sorting and ordinary prime averages

Let

Rx,d=1−∑i=1dTx,i,Rd=1−∑i=1dBi=∏i=1dUi.R_{x,d}=1-\sum_{i=1}^{d}T_{x,i},\qquad R_d=1-\sum_{i=1}^{d}B_i=\prod_{i=1}^{d}U_i.

By (197),

ERx,d⟶ERd=2−d.(198)\mathbb{E}R_{x,d}\longrightarrow\mathbb{E}R_d=2^{-d}. \tag*{(198)}

The decreasing sequence RdR_d has a limit whose expectation is zero, by monotone convergence applied to 1−Rd1-R_d. Hence ∑i≥1Bi=1\sum_{i\ge1}B_i=1 almost surely. Its decreasing rearrangement is therefore well-defined and has total mass one.

For a finite or summable nonnegative collection with total mass one, let ViV_i be its decreasing rearrangement. Sort any selected dd members and pad the resulting list with zeros, writing Sd,iS_{d,i}. If the unselected mass is rr, then

0≤Vi−Sd,i≤r(i≥1).(199)0\le V_i-S_{d,i}\le r\qquad(i\ge1). \tag*{(199)}

The first inequality follows because removing members cannot increase any order statistic. To see the second, if Vi>rV_i>r, every member at least ViV_i must have been selected, since each unselected member is at most rr. Thus Sd,i=ViS_{d,i}=V_i in that case. If Vi≤rV_i\le r, the asserted bound is immediate.

Fix the target dimension kk and a bounded continuous function F:[0,1]k→RF:[0,1]^k\to\mathbb{R}. Let ωF(η)\omega_F(\eta) be its modulus of continuity for the maximum norm. Applying (199) and Markov’s inequality gives

∣EF(V1,…,Vk)−EF(Sd,1,…,Sd,k)∣≤ωF(η)+2∥F∥∞Erη.(200)\left|\mathbb{E}F(V_1,\ldots,V_k)-\mathbb{E}F(S_{d,1},\ldots,S_{d,k})\right|\le\omega_F(\eta)+2\lVert F\rVert_\infty\frac{\mathbb{E}r}{\eta}. \tag*{(200)}

For fixed dd, sorting dd coordinates and padding is a continuous map; for example, each order statistic is a finite maximum of finite minima. Equation (197) therefore gives convergence of the finite sorting test. Apply (200) both to the factor balls and to the stick fragments. First let real xx tend to infinity, use (198), then let dd tend to infinity for fixed η\eta, and finally let η\eta decrease to zero. We obtain

1π(2x)−π(x)∑x<p≤2xF(V1(p),…,Vk(p))⟶EF(L1,…,Lk).(201)\frac{1}{\pi(2x)-\pi(x)}\sum_{x<p\le2x}F(V_1(p),\ldots,V_k(p))\longrightarrow\mathbb{E}F(L_1,\ldots,L_k). \tag*{(201)}

Every step used limits through all real xx.

Completion of the proof of Theorem 1.1. For a real upper endpoint XX, fix an integer J0≥1J_0\ge1 and partition (X/2J0,X](X/2^{J_0},X] into the J0J_0 dyads (X/2j,X/2j−1](X/2^j,X/2^{j-1}]. Each dyad scale tends to infinity with XX, so (201), applied a finite number of times, shows that the weighted sum over these dyads has the asserted limiting average. The discarded primes below X/2J0X/2^{J_0} have proportion

π(X/2J0)π(X)−1=2−J0+o(1)\frac{\pi(X/2^{J_0})}{\pi(X)-1}=2^{-J_0}+o(1)

by the prime number theorem. Their effect on the discrepancy from the limiting expectation is at most 2∥F∥∞(2−J0+o(1))2\lVert F\rVert_\infty(2^{-J_0}+o(1)). Take X→∞X\to\infty, then J0→∞J_0\to\infty. Removing the prime 2 changes the normalization and sum by o(1)o(1). This proves the statement for ordinary equal weighting of 3≤p≤X3\le p\le X and for every real upper endpoint tending to infinity.

Prime polynomials and logarithmic phases

This appendix proves the two estimates used in Section 2: cancellation in prime polynomials up to height x2x^2, and cancellation of logarithmic phases in arithmetic progressions. We retain the quantitative dependence on the degree in the classical Vinogradov mean-value iteration; compare Stechkin [25] and Ford [11], discussion preceding Theorem 3. The logarithmic-phase method and its application to zero-free regions are classical; stronger estimates are available in [11], Theorem 2 and Corollary 2A and [17], Theorem 1.1. The weaker forms below suffice and will be proved in full. Throughout this appendix, L=log⁡xL = \log x and T=log⁡LT = \log L.

Lemma A.1 (Long prime polynomial). Fix 0<τ<η<10 < \tau< \eta< 1, C>0C > 0, and A>0A > 0. There is a constant B0=B0(τ,η,C,A)B_0 = B_0(\tau,\eta,C,A) such that, uniformly for

xτ2≤N≤xη,q≤LC,LB0≤∣t∣≤x2,\frac{x^\tau}{2} \le N \le x^\eta,\qquad q \le L^C,\qquad L^{B_0} \le|t| \le x^2,

every Dirichlet character χ\chi modulo qq and every interval I⊂[N,2N]I \subset[N,2N] satisfy

∣∑p∈Iχ(p)p−1+it∣≪τ,η,C,AL−A.(202)\left|\sum_{p\in I}\chi(p)p^{-1+it}\right| \ll_{\tau,\eta,C,A} L^{-A}. \tag*{(202)}

The proof proceeds from a quantitative power-sum estimate to cancellation for a logarithmic phase, then to a high-height zero-free strip, and finally to primes by Mellin inversion.

Elementary estimates and quantitative power sums

We use two elementary estimates before the mean-value iteration. The first is an immediate consequence of the prime number theorem already recorded in Lemma 2.1.

Lemma A.2 (Primes for the congruence iteration). There is an absolute constant A1>0A_1 > 0 such that, for every integer k≥2k \ge2 and every real y≥A1k6y \ge A_1 k^6, the interval [y,2y][y,2y] contains at least 2k3+52k^3 + 5 primes, all greater than kk.

Proof. The prime number theorem gives an absolute y0y_0 such that [y,2y][y,2y] contains at least y/(2log⁡y)y/(2\log y) primes for y≥y0y \ge y_0. For a sufficiently large absolute A1A_1, this lower bound exceeds 2k3+52k^3 + 5 whenever y≥A1k6y \ge A_1 k^6, uniformly for k≥2k \ge2. Increasing A1A_1 also ensures y>ky > k.

Lemma A.3 (Monotone first derivative estimate). Let ff be real-valued and continuously differentiable on an interval. Suppose f′f' is monotone and takes values in [j+λ,j+1−λ][j+\lambda,j+1-\lambda] for an integer jj and 0<λ≤1/20 < \lambda\le1/2. Then, on every subinterval,

∣∑ne(f(n))∣≪λ−1,\left|\sum_n e(f(n))\right| \ll\lambda^{-1},

where the sum is over the integers in that subinterval and the implied constant is absolute.

Proof. The assertion is immediate for at most one integer. Otherwise write zn=e(f(n))z_n=e(f(n)) and δn=f(n+1)−f(n)=∫nn+1f′(u) du\delta_n=f(n+1)-f(n)=\int_n^{n+1}f'(u)\,\mathrm{d}u between consecutive summation indices. The numbers δn\delta_n are monotone and lie in [j+λ,j+1−λ][j+\lambda,j+1-\lambda]. Put

wn=(e(δn)−1)−1=12−i2cot⁡(πδn).w_n=(e(\delta_n)-1)^{-1}=\frac{1}{2}-\frac{i}{2}\cot(\pi\delta_n).

Then zn=wn(zn+1−zn)z_n=w_n(z_{n+1}-z_n). Summation by parts bounds the sum, including its final term, by 1+2sup⁡n∣wn∣+∑n∣wn+1−wn∣1+2\sup_n|w_n|+\sum_n|w_{n+1}-w_n|. The supremum is O(λ−1)O(\lambda^{-1}). Since the cotangent is real and monotone on the indicated interval modulo integers, the variation has the same bound.

We first establish a power-sum estimate with sufficient uniformity in its degree. For integers k≥2k \ge2, s≥1s \ge1, and M≥1M \ge1, let Js,k(M)J_{s,k}(M) count the solutions of

∑i=1suij=∑i=1svij(1≤j≤k),1≤ui,vi≤M.\sum_{i=1}^{s} u_i^j=\sum_{i=1}^{s} v_i^j\quad(1\le j\le k),\qquad1\le u_i,v_i\le M.

Writing K=k(k+1)/2K=k(k+1)/2 and

f(α)=∑n=1Me(∑j=1kαjnj),f(\boldsymbol{\alpha})=\sum_{n=1}^{M}e\left(\sum_{j=1}^{k}\alpha_j n^j\right),

orthogonality gives Js,k(M)=∫[0,1]k∣f(α)∣2s dαJ_{s,k}(M)=\int_{[0,1]^k}|f(\boldsymbol{\alpha})|^{2s}\,\mathrm{d}\boldsymbol{\alpha}. The argument uses the classical Vinogradov mean-value method in Linnik’s pp-adic form; see Wooley [28], Section 2, pp. 1583–1585 for an exposition. We derive the required degree-uniform quantitative estimate below, without invoking the modern efficient-congruencing theorem of that paper.

Lemma A.4 (A quantitative power-sum bound). There are absolute constants C1,C2>0C_1,C_2>0 such that, for every integer k≥2k\ge2, some integer ss with k≤s≤C1k4k\le s\le C_1k^4 satisfies

Js,k(M)≤exp⁡(C1kC2)M2s−K+1/100(M≥1).J_{s,k}(M)\le\exp(C_1k^{C_2})M^{2s-K+1/100}\qquad(M\ge1).

Proof. We iterate estimates

Js,k(M)≤CsMEs,Es=2s−K+ϵs,J_{s,k}(M)\le C_sM^{E_s},\qquad E_s=2s-K+\epsilon_s,

starting with s=ks=k, Ck=1C_k=1, and ϵk=K\epsilon_k=K. The iteration sends ss to s+ks+k and ϵs\epsilon_s to (1−1/k)ϵs(1-1/k)\epsilon_s.

Choose a sufficiently large absolute constant A1A_1. If M1/k<A1k6M^{1/k}<A_1k^6, then log⁡M≪klog⁡(2k)\log M\ll k\log(2k), and the trivial bound Ju,k(M)≤M2uJ_{u,k}(M)\le M^{2u} costs at most

MK≤exp⁡(O(k3log⁡(2k)))M^K\le\exp(O(k^3\log(2k)))

relative to every proposed estimate with exponent 2u−K+ϵ2u-K+\epsilon, ϵ≥0\epsilon\ge0. Thus it suffices to treat M1/k≥A1k6M^{1/k}\ge A_1k^6. Lemma A.2, after enlarging A1A_1, supplies a fixed list of 2k3+52k^3+5 primes rr in [M1/k,2M1/k][M^{1/k},2M^{1/k}], all greater than kk.

At moment s+ks+k, call a tuple degenerate if it has fewer than kk distinct coordinates. There are at most ks+kMk−1k^{s+k}M^{k-1} such tuples. If GG is their exponential sum, its contribution on one side of the equations is at most

∫[0,1]k∣G∣∣f∣s+k≤ks+kMk−1Js+k,k(M)1/2.\int_{[0,1]^k}|G||f|^{s+k}\le k^{s+k}M^{k-1}J_{s+k,k}(M)^{1/2}.

The contribution with a degenerate tuple on either side is therefore at most twice this quantity.

For a nondegenerate solution, select kk distinct coordinates on each side and permute them to the first kk positions. This costs at most (s+k)2k(s+k)^{2k}. The product of the two Vandermonde products is a nonzero integer of absolute value at most Mk2(k−1)M^{k^2(k-1)}. Fewer than k2(k−1)+1k^2(k-1)+1 primes of size at least M1/kM^{1/k} can divide it. Hence some prime rr in our list makes these first kk coordinates distinct modulo rr on each side.

For such a prime put

Fr(α)=∑1≤z1,…,zk≤Mzi≢zj(modr) (i≠j)e(∑j=1kαj∑i=1kzij).F_r(\boldsymbol{\alpha})= \sum_{\substack{1\le z_1,\ldots,z_k\le M\\ z_i\not\equiv z_j\pmod r\ (i\ne j)}} e\left(\sum_{j=1}^{k}\alpha_j\sum_{i=1}^{k}z_i^j\right).

The relevant number of solutions is bounded by

(s+k)2k∑r∫[0,1]k∣Fr∣2∣f∣2s.(s+k)^{2k}\sum_{r}\int_{[0,1]^k}|F_r|^2|f|^{2s}.

The integrands are nonnegative. Write f=∑a mod rfaf=\sum_{a\bmod r}f_a, where faf_a is restricted to n≡a(modr)n\equiv a\pmod r. Hölder’s inequality gives

∣f∣2s≤r2s−1∑a mod r∣fa∣2s.|f|^{2s}\le r^{2s-1}\sum_{a\bmod r}|f_a|^{2s}.

Fix aa. In the system counted by ∫∣Fr∣2∣fa∣2s\int|F_r|^2|f_a|^{2s}, translate all variables by −a-a. Translation preserves the equations for the first kk powers by the binomial formula. The remaining ss variables on either side are multiples of rr, so the first lists z,z′\boldsymbol{z},\boldsymbol{z}' obey

∑i=1kzij≡∑i=1k(zi′)j(modrj)(1≤j≤k).(203)\sum_{i=1}^{k}z_i^j\equiv\sum_{i=1}^{k}(z_i')^j\pmod{r^j}\qquad(1\le j\le k). \tag*{(203)}

There are at most MkM^k choices for the first list. For each such list, there are at most k!rK−kk!r^{K-k} choices for the second. To see this, lift the prescribed jjth sum modulo rjr^j to a residue modulo rkr^k. The number of choices for all lifts is ∏j=1krk−j=rK−k\prod_{j=1}^{k}r^{k-j}=r^{K-k}. For each full vector of sums modulo rkr^k, Newton’s identities determine the multiset of roots modulo rr, since r>kr>k. There are at most k!k! orderings. The Jacobian of the power sums has determinant

k!∏i<j(zj′−zi′),k!\prod_{i<j}(z_j'-z_i'),

up to sign, and is invertible modulo rr. Each ordering therefore lifts uniquely from modulus rr to modulus rkr^k: at each stage the next digits are the unique solution of the corresponding linear system modulo rr. Finally, an interval of MM integers contains at most one representative of any residue modulo rkr^k, because M≤rkM\le r^k.

After both first lists are fixed, divide the remaining variables by rr. They range over a common interval of at most ⌈M/r⌉\lceil M/r\rceil integers and have prescribed differences of power sums. Translation to an initial interval changes only these prescribed differences. The count is a Fourier coefficient of the nonnegative function ∣f∣2s|f|^{2s} at that shorter length, and consequently is at most Js,k(⌈M/r⌉)J_{s,k}(\lceil M/r\rceil). Summing over aa accounts for the final factor rr, and gives

Js+k,k(M)≤2ks+kMk−1Js+k,k(M)1/2+(s+k)2k(2k3+5)k!max⁡rr2s+K−kMkJs,k(⌈M/r⌉).(204)J_{s+k,k}(M)\le2k^{s+k}M^{k-1}J_{s+k,k}(M)^{1/2} +(s+k)^{2k}(2k^3+5)k!\max_r r^{2s+K-k}M^kJ_{s,k}(\lceil M/r\rceil). \tag*{(204)}

Insert (A.3). Since Es≥0E_s\ge0, rounding costs at most 2Es2^{E_s}, and the exponent of rr becomes

2s+K−k−Es=k2−ϵs≥0.2s+K-k-E_s=k^2-\epsilon_s\ge0.

Replacing rr by at most 2M1/k2M^{1/k} therefore gives the exponent

k+Es+k2−ϵsk=2(s+k)−K+(1−1/k)ϵs.k+E_s+\frac{k^2-\epsilon_s}{k}=2(s+k)-K+(1-1/k)\epsilon_s.

The inequality J≤aJ+bJ\le a\sqrt{J}+b implies J≤2a2+2bJ\le2a^2+2b. The exponent 2k−22k-2 from a2a^2 is admissible: the target exponent starts at 2k2k, and at each step its increase is 2k−ϵs/k≥2k−K/k>02k-\epsilon_s/k\ge2k-K/k>0. Thus the asserted iteration holds. Its constants can be chosen with

log⁡Cs+k≤log⁡(1+Cs)+O((s+k)log⁡(2k)+klog⁡(s+k)+k3log⁡(2k)),\log C_{s+k}\le\log(1+C_s)+O\bigl((s+k)\log(2k)+k\log(s+k)+k^3\log(2k)\bigr),
在天天中彩票在天天中彩票

including (A.4). After O(klog⁡(2k))O(k\log(2k)) steps, ϵs=K(1−1/k)(s−k)/k≤1/100\epsilon_s=K(1-1/k)^{(s-k)/k}\le1/100. Then s=O(k2log⁡(2k))≤C1k4s=O(k^2\log(2k))\le C_1k^4, and summing the displayed costs gives log⁡Cs≤C1kC2\log C_s\le C_1k^{C_2} for absolute constants. Increasing the exponent from ϵs\epsilon_s to 1/1001/100 proves the Lemma.

Logarithmic phases on progressions

The quantitative moment bound now gives cancellation for a logarithmic phase even when the Taylor degree grows with xx. This step yields the second estimate from Section 2 and will also control Dirichlet series near the line of absolute convergence.

Lemma A.5 (A logarithmic phase on progressions). Fix C>0C>0. There is an absolute constant C3>0C_3>0 such that, for sufficiently large xx in terms of CC, the following holds uniformly:

exp⁡(L/T2)≤N′≤2x5,q≤LC,exp⁡(L/(2T2))≤∣v∣≤4x3.\exp(L/T^2)\le N'\le2x^5,\qquad q\le L^C,\qquad\exp(L/(2T^2))\le|v|\le4x^3.

For every residue a(modq)a\pmod q and every interval I⊂[N′,2N′]\mathcal{I}\subset[N',2N'],

∣∑n∈In≡a(modq)niv∣≪N′qexp⁡(−L/TC3).(205)\left|\sum_{\substack{n\in\mathcal{I}\\ n\equiv a\pmod q}}n^{iv}\right|\ll\frac{N'}{q}\exp(-L/T^{C_3}). \tag*{(205)}

Proof. Put H=N′/qH=N'/q and y=log⁡∣v∣/log⁡Hy=\log|v|/\log H. Then

log⁡H≥L/T2−CT≫L/T2,y≪T2.\log H\ge L/T^2-CT\gg L/T^2,\qquad y\ll T^2.

Writing n=q(b+a/q)n=q(b+a/q), with 0≤a<q0\le a<q, reduces the sum, up to a factor of modulus one, to ∑b∈I(b+a/q)iv\sum_{b\in\mathcal{I}}(b+a/q)^{iv}, where I\mathcal{I} is an interval of integers and H≤b+a/q≤2HH\le b+a/q\le2H.

If y<4/5y<4/5, the derivative of vlog⁡(b+a/q)/(2π)v\log(b+a/q)/(2\pi) is monotone, has magnitude comparable to ∣v∣/H|v|/H, and has magnitude less than 1/21/2. The monotone first derivative estimate in Lemma A.3 therefore bounds the sum by O(H/∣v∣)O(H/|v|). The lower bound on ∣v∣|v| makes this smaller than (A.7).

Suppose now that y≥4/5y\ge4/5. Set

M=⌊H3/4⌋,k=⌈4y+8⌉.M=\lfloor H^{3/4}\rfloor,\qquad k=\lceil4y+8\rceil.

Averaging the sum over forward shifts 1≤h≤M1\le h\le M changes it by O(M)O(M): a shift changes an interval at only O(h)O(h) endpoints. For b∈Ib\in\mathcal{I}, Taylor expansion gives

v2πlog⁡(b+h+a/q)=v2πlog⁡(b+a/q)+∑j=1kαj(b)hj+O(H−2),αj(b)=(−1)j−1v2πj(b+a/q)j.\frac{v}{2\pi}\log(b+h+a/q)=\frac{v}{2\pi}\log(b+a/q)+\sum_{j=1}^{k}\alpha_j(b)h^j+O(H^{-2}),\qquad\alpha_j(b)=\frac{(-1)^{j-1}v}{2\pi j(b+a/q)^j}.

Indeed the remainder is O(∣v∣(M/H)k+1)=O(Hy−(k+1)/4)=O(H−9/4)O(|v|(M/H)^{k+1})=O(H^{y-(k+1)/4})=O(H^{-9/4}), uniformly even as kk grows. Thus, with

S(α)=∑h=1Me(∑j=1kαjhj),S(\boldsymbol{\alpha})=\sum_{h=1}^{M}e\left(\sum_{j=1}^{k}\alpha_jh^j\right),

the original sum is bounded by

1M∑b∈I∣S(α(b))∣+O(M+H−1).(206)\frac{1}{M}\sum_{b\in\mathcal{I}}|S(\boldsymbol{\alpha}(b))|+O(M+H^{-1}). \tag*{(206)}

Partition the coefficient torus [0,1)k[0,1)^k into boxes with side length M−jM^{-j} in coordinate jj. Each box contains at most

exp⁡(O(k))H4/5\exp(O(k))H^{4/5}

of the vectors α(b)\boldsymbol{\alpha}(b) (mod 11). For this it suffices to use the coordinate j=⌈y+3/20⌉j=\lceil y+3/20\rceil, which lies between 11 and kk. Before reduction modulo one, this coordinate has magnitude O(H−3/20)O(H^{-3/20}), is monotone, and has derivative of magnitude at least

∣v∣2π(2H)j+1≥exp⁡(−O(k))Hy−j−1.\frac{|v|}{2\pi(2H)^{j+1}}\geq\exp(-O(k))H^{y-j-1}.

A coordinate interval of length M−jM^{-j} on the torus pulls back to at most two intervals in bb. Their total length is at most exp⁡(O(k))H1+j/4−y\exp(O(k))H^{1+j/4-y}. Here the floor in MM costs exp⁡(O(k))\exp(O(k)), while

y−j4≥3y4−2380≥516>15.y-\frac{j}{4}\geq\frac{3y}{4}-\frac{23}{80}\geq\frac{5}{16}>\frac{1}{5}.

Counting integer points proves (A.10).

Let ss be supplied by Lemma A.4. We also need the following bound for suprema over boxes QQ:

∑Qsup⁡α∈Q∣S(α)∣2s≤exp⁡(O(sk+k))MKJs,k(M).(207)\sum_Q\sup_{\boldsymbol{\alpha}\in Q}|S(\boldsymbol{\alpha})|^{2s}\leq\exp(O(sk+k))M^KJ_{s,k}(M). \tag*{(207)}

To prove it, rescale each box to [0,1]k[0,1]^k. Iterating the one-dimensional fundamental theorem of calculus gives, for any smooth function FF on that cube,

sup⁡∣F∣≤∑ϵ∈{0,1}k∫[0,1]k∣∂ϵF∣.\sup|F|\leq\sum_{\boldsymbol{\epsilon}\in\{0,1\}^k}\int_{[0,1]^k}|\partial^{\boldsymbol{\epsilon}}F|.

Hölder’s inequality bounds the 2s2sth power of the right-hand side by 2k(2s−1)2^{k(2s-1)} times the sum of the corresponding 2s2sth moments. For the rescaled SS, every mixed derivative has coefficients bounded in modulus by (2π)k(2\pi)^k: each differentiation in coordinate jj introduces the factor 2πihj/Mj2\pi i h^j/M^j. On summing over the boxes, change of variables contributes MKM^K. Orthogonality then bounds the full-torus moment of every such derivative by (2π)2skJs,k(M)(2\pi)^{2sk}J_{s,k}(M). This proves (A.11), with constants controlled at the growing degree.

Apply Hölder’s inequality to the sum in (A.9), and use (A.10), (A.11), and (A.2). Since ∣I∣≪H|\mathcal{I}|\ll H, the result is

1M∑b∈I∣S(α(b))∣≪H(exp⁡(O(kC4))H−1/5M1/100)1/(2s)(208)\frac{1}{M}\sum_{b\in\mathcal{I}}|S(\boldsymbol{\alpha}(b))|\ll H\left(\exp(O(k^{C_4}))H^{-1/5}M^{1/100}\right)^{1/(2s)} \tag*{(208)}

for an absolute constant C4C_4. We have k≪T2k\ll T^2 and s≪T8s\ll T^8, whereas

−15log⁡H+1100log⁡M≤77400log⁡H.-\frac{1}{5}\log H+\frac{1}{100}\log M\leq\frac{77}{400}\log H.

Every fixed power of TT is o(L/T2)o(L/T^2). Thus the right-hand side of (A.12) is at most Hexp⁡(−c1L/T10)H\exp(-c_1L/T^{10}) for some absolute c1>0c_1>0, once xx is sufficiently large. The errors in (A.9) are smaller. Increasing a fixed exponent C3>10C_3>10 absorbs c1c_1 and proves the Lemma.

A zero-free strip from the phase estimate

We next convert progression cancellation into a bound for a Dirichlet LL-function near Re⁡s=1\operatorname{Re} s = 1. A local logarithmic-derivative formula and the classical positive trigonometric polynomial then exclude zeros in the narrower strip needed for Mellin inversion.

We write D(s,χ)D(s,\chi) for the Dirichlet LL-function, to distinguish it from L=log⁡xL = \log x. Characters need not be primitive. Periodicity gives ∑n≤uχ(n)=cχu+O(q)\sum_{n\le u}\chi(n)=c_\chi u+O(q), where cχ=q−1∑a mod qχ(a)c_\chi=q^{-1}\sum_{a\bmod q}\chi(a). Partial summation therefore continues D(s,χ)D(s,\chi) meromorphically to Re⁡s>0\operatorname{Re}s>0, with only the possible simple pole at s=1s=1, and gives, for σ=Re⁡s\sigma=\operatorname{Re}s in a fixed compact subinterval of (0,∞)(0,\infty),

D(s,χ)=∑n≤Yχ(n)ns+cχY1−ss−1+O(q(1+∣s∣)Y−σ).D(s,\chi)=\sum_{n\le Y}\frac{\chi(n)}{n^s}+\frac{c_\chi Y^{1-s}}{s-1}+O\left(q(1+|s|)Y^{-\sigma}\right).

Lemma A.6 (A bound near the line Re⁡s=1\operatorname{Re}s=1). Fix C>0C>0 and put r∗=T4/Lr_\ast=T^4/L. Uniformly for q≤LCq\le L^C,

log⁡∣D(σ+iv,χ)∣≪CT2(∣σ−1∣≤10r∗,12≤∣v∣≤3x3).(209)\log|D(\sigma+iv,\chi)|\ll_C T^2\qquad\left(|\sigma-1|\le10r_\ast,\quad\frac{1}{2}\le|v|\le3x^3\right). \tag*{(209)}

Proof. First suppose ∣v∣≤exp⁡(L/(2T2))|v|\le\exp(L/(2T^2)), and use Y=exp⁡(2L/T2)Y=\exp(2L/T^2) in (A.13). Absolute summation gives

∑n≤Yn−σ≪(1+log⁡Y)max⁡(1,Y1−σ)≤exp⁡(O(T2)).\sum_{n\le Y}n^{-\sigma}\ll(1+\log Y)\max(1,Y^{1-\sigma})\le\exp(O(T^2)).

Since ∣s−1∣≥1/2|s-1|\ge1/2, the pole term has the same bound. The error is at most

exp⁡(−3L2T2+O(T2+CT)),\exp\left(-\frac{3L}{2T^2}+O(T^2+CT)\right),

and hence is negligible.

For exp⁡(L/(2T2))<∣v∣≤3x3\exp(L/(2T^2))<|v|\le3x^3, take Y=x5Y=x^5. The terms with n≤exp⁡(L/T2)n\le\exp(L/T^2) again contribute exp⁡(O(T2))\exp(O(T^2)) in absolute value. On a dyadic interval above this threshold, sum (A.7), with phase −v-v, over the residue classes modulo qq, including their character coefficients. The factors qq and 1/q1/q cancel. Partial summation with n−σn^{-\sigma} bounds that dyad by

exp⁡(−L/TC3)exp⁡(O(T4)).\exp(-L/T^{C_3})\exp(O(T^4)).

There are O(L)O(L) dyads, and the same bound applies to a final partial dyad. Their total is negligible, since L/TC3L/T^{C_3} exceeds every fixed power of TT. Finally, the pole term and remainder in (A.13) are bounded respectively by

exp⁡(−L2T2+O(T4)),exp⁡(−2L+O(T4+CT)).\exp\left(-\frac{L}{2T^2}+O(T^4)\right),\qquad\exp(-2L+O(T^4+CT)).

This proves (A.14).

Lemma A.7 (Local logarithmic derivative). Fix C>0C>0, let q≤LCq\le L^C, and put s0=1+r∗+ivs_0=1+r_\ast+iv, where 1≤∣v∣≤52x31\le|v|\le\frac{5}{2}x^3. There is a radius RR with 4r∗≤R≤5r∗4r_\ast\le R\le5r_\ast, with no zero on its boundary, for which

D′D(s,χ)=∑∣ρ−s0∣<R1s−ρ+OC(T2/r∗)(∣s−s0∣≤2r∗),(210)\frac{D'}{D}(s,\chi)=\sum_{|\rho-s_0|<R}\frac{1}{s-\rho}+O_C(T^2/r_\ast)\qquad(|s-s_0|\le2r_\ast), \tag*{(210)}

away from zeros. The sum counts zeros with multiplicity and has OC(T2)O_C(T^2) terms.

Proof. For large xx, the disk ∣s−s0∣≤8r∗|s-s_0|\le8r_* avoids the possible pole at s=1s=1 and lies in the region of Lemma A.6. At its center the Euler product gives

∣D(s0,χ)∣≥∏p(1+p−1−r∗)−1=ζ(2+2r∗)ζ(1+r∗)≫r∗.|D(s_0,\chi)|\ge\prod_p(1+p^{-1-r_*})^{-1}=\frac{\zeta(2+2r_*)}{\zeta(1+r_*)}\gg r_*.

Thus log⁡∣D(s0,χ)∣≥−O(T)\log|D(s_0,\chi)|\ge-O(T). Jensen’s formula, using the radius 8r∗8r_* and the upper bound OC(T2)O_C(T^2), shows that the disk of radius 5r∗5r_* contains OC(T2)O_C(T^2) zeros. Choose R∈[4r∗,5r∗]R\in[4r_*,5r_*] so that none lies on its boundary.

Use centered coordinates z=s−s0z=s-s_0, and write a=ρ−s0a=\rho-s_0 for each zero in ∣z∣<R|z|<R. Divide D(s0+z,χ)D(s_0+z,\chi) by the disk Blaschke factors

Ba(z)=R(z−a)R2−a‾z.B_a(z)=\frac{R(z-a)}{R^2-\overline{a}z}.

repeated with multiplicity. The quotient GG is holomorphic and nonvanishing on the closed disk, and has the same boundary modulus as DD. Maximum modulus gives log⁡∣G∣≤M0\log|G|\le M_0 throughout the disk for some M0=OC(T2)M_0=O_C(T^2). Since ∣Ba(0)∣<1|B_a(0)|<1, log⁡∣G(0)∣≥−O(T)\log|G(0)|\ge-O(T).

Let hh be an analytic logarithm of GG. The positive harmonic function M0−Re⁡hM_0-\operatorname{Re}h has value OC(T2)O_C(T^2) at the center. The Poisson formula, or its derivative together with Harnack’s inequality on concentric disks, gives

∣h′(z)∣≪CT2r∗(∣z∣≤2r∗).|h'(z)|\ll_C \frac{T^2}{r_*}\qquad(|z|\le2r_*).

Restoring the factors uses

Ba′(z)Ba(z)=1z−a+a‾R2−a‾z.\frac{B_a'(z)}{B_a(z)}=\frac{1}{z-a}+\frac{\overline{a}}{R^2-\overline{a}z}.

The second term is O(1/r∗)O(1/r_*) on ∣z∣≤2r∗|z|\le2r_*, uniformly in ∣a∣<R|a|<R. There are OC(T2)O_C(T^2) such terms, giving exactly (A.15).

Lemma A.8 (A zero-free strip at large height). For each fixed C>0C>0 there is c=c(C)>0c=c(C)>0 such that every character of modulus at most LCL^C has no zero in

Re⁡s≥1−cT2/L,1≤∣Im⁡s∣≤x3.\operatorname{Re}s\ge1-cT^2/L,\qquad1\le|\operatorname{Im}s|\le x^3.

Moreover,

∣D′D(s,χ)∣≪CL(1−cT22L≤Re⁡s≤1+1L,2≤∣Im⁡s∣≤x32).(211)\left|\frac{D'}{D}(s,\chi)\right|\ll_C L\left(1-\frac{cT^2}{2L}\le\operatorname{Re}s\le1+\frac{1}{L},\quad2\le|\operatorname{Im}s|\le\frac{x^3}{2}\right). \tag*{(211)}

Proof. For σ>1\sigma>1, the Euler products and the inequality 3+4cos⁡θ+cos⁡(2θ)=2(1+cos⁡θ)2≥03+4\cos\theta+\cos(2\theta)=2(1+\cos\theta)^2\ge0 give

0≤−3ζ′ζ(σ)−4Re⁡D′D(σ+iv,χ)−Re⁡D′D(σ+2iv,χ2).(212)0\le-3\frac{\zeta'}{\zeta}(\sigma)-4\operatorname{Re}\frac{D'}{D}(\sigma+iv,\chi)-\operatorname{Re}\frac{D'}{D}(\sigma+2iv,\chi^2). \tag*{(212)}

Indeed this follows term by term in the absolutely convergent prime-power expansions; primes dividing the modulus contribute only the positive zeta term. Also −ζ′/ζ(σ)=1/(σ−1)+O(1)-\zeta'/\zeta(\sigma)=1/(\sigma-1)+O(1) near 11.

Suppose ρ=β+iv\rho=\beta+iv were a zero in (A.16), and set σ=1+20cT2/L\sigma=1+20cT^2/L. Apply Lemma A.7 at heights vv and 2v2v. The evaluation points lie in the respective inner disks, and ρ\rho lies in the first zero sum, because T2/L=o(r∗)T^2/L=o(r_*). No zero has real part greater than 11, by the absolutely convergent Euler product. All terms in the zero sums therefore contribute nonpositively to the negative real logarithmic derivatives. Retaining the term at ρ\rho, and using 0<σ−β≤21cT2/L0 < \sigma- \beta\le21cT^{2}/L, bounds the right-hand side of (212) by

(320c−421c+OC(1))LT2=(−17420c+OC(1))LT2.\left(\frac{3}{20c}-\frac{4}{21c}+O_C(1)\right)\frac{L}{T^{2}}=\left(-\frac{17}{420c}+O_C(1)\right)\frac{L}{T^{2}}.

Here the errors are OC(T2/r∗)=OC(L/T2)O_C(T^{2}/r_{\ast})=O_C(L/T^{2}). Choosing a sufficiently small fixed c>0c>0 gives a contradiction.

For ss in the region of (211), apply the local formula centered at 1+r∗+iIm⁡s1+r_{\ast}+i\operatorname{Im}s. Every zero in its sum has absolute imaginary part between 11 and x3x^{3}, since 2≤∣Im⁡s∣≤x3/22\leq|\operatorname{Im}s|\leq x^{3/2} and r∗=o(1)r_{\ast}=o(1). By (A.16), its horizontal distance from ss is at least cT2/(2L)cT^{2}/(2L). The OC(T2)O_C(T^{2}) zero terms and the local error therefore total OC(L)O_C(L), proving (211).

Mellin inversion and the prime polynomial

The zero-free strip permits a short contour shift while keeping the imaginary part away from zero. Smoothing the interval first makes the horizontal edges and discarded Mellin tails uniformly negligible.

Proof of Lemma A.1. We first establish the analogous bound with the von Mangoldt weight. Let δ=L−A−4\delta=L^{-A-4}. Smooth the indicator of I/N⊂[1,2]I/N\subset[1,2] by convolution with a nonnegative smooth kernel of width δ\delta. This gives 0≤g≤10\leq g\leq1, supported in [1/2,3][1/2,3], whose difference from the indicator is supported within O(δ)O(\delta) of its endpoints, and with ∥g(j)∥∞≪jδ−j\|g^{(j)}\|_{\infty}\ll_j\delta^{-j}. The construction also applies when II is shorter than δN\delta N. Since Λ(n)≤log⁡n\Lambda(n)\leq\log n, the error in replacing the interval by g(n/N)g(n/N) is

O((δN+1)log⁡(3N)N)=O(L−A−2),(213)O\left((\delta N+1)\frac{\log(3N)}{N}\right)=O(L^{-A-2}), \tag*{(213)}

uniformly in II.

Define the Mellin transform by

g^(z)=∫0∞g(w)wzdww.\widehat{g}(z)=\int_{0}^{\infty}g(w)w^{z}\frac{dw}{w}.

On every fixed bounded real-part strip, integration by parts gives

∣g^(u+iv)∣≪jL(j+1)(A+5)(1+∣v∣)−j.(214)|\widehat{g}(u+iv)|\ll_j L^{(j+1)(A+5)}(1+|v|)^{-j}. \tag*{(214)}

Also ∣g^(u+iv)∣≪1|\widehat{g}(u+iv)|\ll1 there, by absolute integration. Mellin inversion and the absolutely convergent logarithmic derivative on Re⁡z=1/L\operatorname{Re}z=1/L yield

∑nΛ(n)χ(n)n−1+itg(n/N)=12πi∫1/L−i∞1/L+i∞g^(z)Nz(−D′D(1−it+z,χ)) dz.(215)\sum_n\Lambda(n)\chi(n)n^{-1+it}g(n/N)=\frac{1}{2\pi i}\int_{1/L-i\infty}^{1/L+i\infty}\widehat{g}(z)N^z\left(-\frac{D'}{D}(1-it+z,\chi)\right)\,dz. \tag*{(215)}

On this full line, absolute convergence gives

∣D′D(1+1/L+iu,χ)∣≤∑nΛ(n)n1+1/L=−ζ′ζ(1+1/L)≪L,\left|\frac{D'}{D}(1+1/L+iu,\chi)\right|\leq\sum_n\frac{\Lambda(n)}{n^{1+1/L}}=-\frac{\zeta'}{\zeta}(1+1/L)\ll L,

at every height uu.

Choose a fixed C5>A+6C_5>A+6 and put H0=LC5H_0=L^{C_5}. Then choose a fixed integration-by-parts order jj large enough that (214) makes the two tails with ∣Im⁡z∣>H0|\operatorname{Im}z|>H_0 in (215) O(L−A−2)O(L^{-A-2}). Explicitly their bound is

Oj(L1+(j+1)(A+5)H01−j).O_j\left(L^{1+(j+1)(A+5)}H_0^{1-j}\right).

since N1/L≪1N^{1/L} \ll1. Fix B0>C5+2B_0 > C_5 + 2. For LB0≤∣t∣≤x2L^{B_0} \le|t| \le x^2, every point in the rectangle

−cT22L≤Re⁡z≤1L,∣Im⁡z∣≤H0-\frac{cT^2}{2L} \le\operatorname{Re} z \le\frac{1}{L}, \qquad|\operatorname{Im} z| \le H_0

has

2≤∣−t+Im⁡z∣≤x2+H0<x32.(216)2 \le|-t+\operatorname{Im} z| \le x^2+H_0 < \frac{x^3}{2}. \tag*{(216)}

for large xx. Thus Lemma A.8 applies throughout the rectangle. There are no zeros or poles of the logarithmic derivative inside it; in particular, the possible principal-character pole at z=itz=it is outside it.

Shift the truncated contour to Re⁡z=−cT2/(2L)\operatorname{Re} z=-cT^2/(2L). The horizontal edges are O(L−A−2)O(L^{-A-2}), by (211) and (214), increasing the already fixed order jj if necessary. On the new vertical segment,

∣Nz∣=exp⁡(−cT2log⁡N2L)≪exp⁡(−cτT2/2).|N^z|=\exp\left(-\frac{cT^2\log N}{2L}\right)\ll\exp(-c\tau T^2/2).

Its remaining factors and length cost at most a fixed power of LL. Since exp⁡(−cτT2/2)\exp(-c\tau T^2/2) is smaller than every prescribed fixed power of L−1L^{-1}, the shifted integral is O(L−A−2)O(L^{-A-2}). Together with (213), this proves, uniformly for all subintervals I⊂[N,2N]I\subset[N,2N],

∑n∈IΛ(n)χ(n)n−1+it≪L−A−2.\sum_{n\in I}\Lambda(n)\chi(n)n^{-1+it}\ll L^{-A-2}.

Prime powers of exponent at least two contribute in absolute value at most O(N−1/2(log⁡(3N))2)O(N^{-1/2}(\log(3N))^2): there are O(Nlog⁡(3N))O(\sqrt{N}\log(3N)) possible bases and exponents with N≤pa≤2NN\le p^a\le2N, and every summand has size O(log⁡(3N)/N)O(\log(3N)/N). This is smaller than every fixed power of L−1L^{-1} in the present range. Removing them from (A.23) gives the same uniform bound for ∑p∈I(log⁡p)χ(p)p−1+it\sum_{p\in I}(\log p)\chi(p)p^{-1+it}. Finally, partial summation against 1/log⁡u1/\log u, whose value and total variation on [N,2N][N,2N] are Oτ(1/L)O_\tau(1/L), removes the logarithmic weight and proves (202).

References

References

  1. [1]W. R. Alford, Andrew Granville, and Carl Pomerance. There are infinitely many Carmichael numbers. Annals of Mathematics, 139(3):703–722, 1994.DOI
  2. [2]Richard Arratia, Fred Kochman, and Victor S. Miller. Extensions of Billingsley’s theorem via multi-intensities. https://arxiv.org/abs/1401.1555v1, 2014. Version 1, January 8, 2014.
  3. [3]R. C. Baker and Glyn Harman. Shifted primes without large prime factors. Acta Arithmetica, 83(4):331–361, 1998.DOI
  4. [4]Abhishek Bharadwaj and Brad Rodgers. Large prime factors of well-distributed sequences. Canadian Mathematical Bulletin, pages 1–17, 2026. First View, published online April 17, 2026; arXiv:2402.11884v4, April 9, 2026.arxiv.org/abs/2402.11884
  5. [5]Patrick Billingsley. On the distribution of large prime divisors. Periodica Mathematica Hungarica, 2:283–289, 1972.DOI
  6. [6]N. G. de Bruijn. On the number of positive integers ≤ x and free of prime factors > y. Proceedings of the Koninklijke Nederlandse Akademie van Wetenschappen, Series A, 54(1):50–60, 1951.DOI
  7. [7]K. Dickman. On the frequency of numbers containing prime factors of a certain relative magnitude. Arkiv för Matematik, Astronomi och Fysik, 22A(10):1–14, 1930.
  8. [8]Peter Donnelly and Geoffrey Grimmett. On the asymptotic distribution of large prime factors. Journal of the London Mathematical Society, 47(3):395–404, 1993.DOI
  9. [9]Paul Erdős. On the normal number of prime factors of p − 1 and some related problems concerning Euler’s φ-function. The Quarterly Journal of Mathematics, os-6(1):205–213, 1935.
  10. [10]Paul Erdős. On pseudoprimes and Carmichael numbers. Publicationes Mathematicae Debrecen, 4:201–206, 1956.
  11. [11]Kevin Ford. Vinogradov’s integral and bounds for the Riemann zeta function. Proceedings of the London Mathematical Society, 85(3):565–633, 2002. Author version: arXiv:1910.08209v1.
  12. [12]Kevin Ford. Poisson approximation of prime divisors of shifted primes. International Mathematics Research Notices, 2025(7):rnaf079, 2025. 16 pages.arxiv.org/abs/2408.03803
  13. [13]Kevin Ford, Sergei V. Konyagin, and Florian Luca. Prime chains and Pratt trees. Geometric and Functional Analysis, 20(5):1231–1258, 2010. Author version: arXiv:0904.0473v4, September 15, 2010.arxiv.org/abs/0904.0473
  14. [14]Ofir Gorodetsky. A Kubilius model for sieve-theoretic sequences. Analysis Mathematica, 2026. Published online August 14, 2026; arXiv:2608.14190v1.DOI
  15. [15]Andrew Granville. Smooth numbers: computational number theory and beyond. In J. P. Buhler and P. Stevenhagen, editors, Algorithmic Number Theory: Lattices, Number Fields, Curves and Cryptography, volume 44 of Mathematical Sciences Research Institute Publications, pages 267–323. Cambridge University Press, Cambridge, 2008.DOI
  16. [16]Harald Andrés Helfgott and Maksym Radziwiłł. Expansion, divisibility and parity. https://arxiv.org/abs/2103.06853v2, 2021. Version 2, April 13, 2021.
  17. [17]Tanmay Khale. An explicit Vinogradov–Korobov zero-free region for Dirichlet L-functions. The Quarterly Journal of Mathematics, 75(1):299–332, 2024.arxiv.org/abs/2210.06457
  18. [18]Jared Duker Lichtman. Primes in arithmetic progressions to large moduli, and shifted primes without large prime factors. https://arxiv.org/abs/2211.09641v1, 2022. Version 1, November 14, 2022.
  19. [19]Kaisa Matomäki and Maksym Radziwiłł. Multiplicative functions in short intervals. https://arxiv.org/abs/1501.04585v4, 2017. Version 4, October 15, 2017.
  20. [20]Hugh L. Montgomery and Robert C. Vaughan. Hilbert’s inequality. Journal of the London Mathematical Society, 8:73–82, 1974.DOI
  21. [21]OpenAI. Weighted dilation graphs, smooth shifted primes and totient fibers. OpenAI Math Release preprint OAI:Weighted-Dilation-Graphs-Smooth-Shifted-Primes-and-Totient-Fibers-September-24-2026, 2026.
  22. [22]Cédric Pilatte. Improved bounds for the two-point logarithmic Chowla conjecture. https://arxiv.org/abs/2310.19357v3, 2026. Version 3, August 25, 2026; first submitted in 2023.
  23. [23]Carl Pomerance. Popular values of Euler’s function. Mathematika, 27(1):84–89, 1980.DOI
  24. [24]Kannan Soundararajan. Moments of the Riemann zeta function. Annals of Mathematics, 170(2):981–993, 2009.
  25. [25]S. B. Stechkin. Mean values of the modulus of a trigonometric sum. Trudy Matematicheskogo Instituta imeni V. A. Steklova, 134:283–309, 1975. English translation: Proceedings of the Steklov Institute of Mathematics 134 (1977), 321–350.
  26. [26]Terence Tao. 254A, notes 1: Elementary multiplicative number theory. https://terrytao.wordpress.com/2014/11/23/254a-notes-1-elementary-multiplicative-number-theory/, 2014. November 23, 2014. Theorems 15 and 26.
  27. [27]Terence Tao. 254A, notes 2: Complex-analytic multiplicative number theory. https://terrytao.wordpress.com/2014/12/09/254a-notes-2-complex-analytic-multiplicative-number-theory/, 2014. December 9, 2014. Corollary 39, Exercise 40, and Exercise 64.
  28. [28]Trevor D. Wooley. Vinogradov’s mean value theorem via efficient congruencing. Annals of Mathematics, 175(3):1575–1627, 2012. doi:10.4007/annals.2012.175.3.12.DOI

Paper details

Contents