Introduction

The result

For an integer q≥2q \ge2 and an integer aa with (a,q)=1(a,q)=1, let P(a,q)P(a,q) be the least prime congruent to aa modulo qq. Dirichlet’s theorem [31] ensures that this prime exists. Linnik [77, 78] proved that there are absolute constants CC and LL such that

P(a,q)≤CqL(1)P(a,q) \le Cq^{L} \tag*{(1)}

for every reduced residue class. An exponent LL for which such a constant CC exists is called admissible; Linnik’s constant is the infimum of the admissible exponents.

Theorem 1.1 (Least prime in a reduced residue class). There is an absolute constant q0≥2q_{0} \ge2 such that, for every integer q≥q0q \ge q_{0} and every integer aa with (a,q)=1(a,q)=1, there is a prime pp satisfying

p≡a(modq),p<q3.99.p \equiv a \pmod{q}, \qquad p < q^{3.99}.

Consequently, there is an absolute constant C>0C>0 such that P(a,q)≤Cq3.99P(a,q) \le Cq^{3.99} for all q≥2q \ge2 and (a,q)=1(a,q)=1. In particular, Linnik’s constant is at most 3.99.

We do not compute q0q_{0} or CC; the issue of effectivity is discussed at the end of Section 13. Except in the presence of an extremely close real zero, the proof finds a prime in the smaller interval

q3.156342<p<q3.99.q^{3.156342} < p < q^{3.99}.

The lower endpoint comes from the support of the weight used to detect primes. When a real zero is extremely close to 1, Heath-Brown’s theorem on Siegel zeros [50] instead gives P(a,q)≤q3.5P(a,q) \le q^{3.5}.

The argument combines published analytic estimates, new lemmas proved here, and exact finite computations. The theorem requires all three parts. In particular, checking a numerical certificate proves a bound for its specified configuration of zeros; the analytic argument must also show that every modulus gives a configuration covered by one of the certificates. We make this passage explicit in Sections 11 and 13.

Earlier work

The admissible exponents in Table 1 record the main numerical developments. Heath-Brown [51] obtained 5.5, and Xylouris [143, 144, 145, 146] subsequently obtained 5.2, 5.18, and 5. Before them, Pan [98, 99] gave the first numerical values, and Chen, Jutila, Graham and Wang lowered them in a long series of papers, which includes Graham’s introduction of Selberg sieve weights into zero-density estimates [39, 41]. Our proof uses the framework of these works. It rests on the three principles that Heath-Brown isolates at the start of his paper: a zero-free region with at most one exceptional zero [27, 43, 74, 97], the repulsion of other zeros by an exceptional zero, known as the Deuring–Heilbronn phenomenon [30, 53, 78], and a log-free zero-density estimate [77]. The last means that the density bound has no additional power of log⁡q\log q, a feature needed when studying zeros at distance O(1/log⁡q)O(1/\log q) from 1. Simpler proofs of these principles, and of Linnik’s theorem, were given by Rodosskiĭ [111], Turán [133, 135], Knapowski [71], Fogels [32], Gallagher [37], Jutila [65], Motohashi [91, 92] and Bombieri [10], among others; accounts are in the books [110, 84, 62, 34].

LLYearAuthor(s)Reference
100001957Pan (announced)[98]
54481958Pan[99]
7771965Chen[18]
6301971Jutila (reported by Turán)[134]
5501970Jutila[64]
1681977Chen[19]
801977Jutila[66]
361977Graham[39]
201981Graham[41]
171979Chen[20]
161986Wang[139]
13.51989Chen and Liu[21, 22]
11.51991Chen and Liu[23]
81991Wang[140]
5.51992Heath-Brown[51]
5.22009Xylouris[143]
5.182011Xylouris[144]
52011Xylouris[145, 146]
3.992026this paper

Table 1. Admissible values of Linnik’s constant, following the tables in [51] and [144], in order of the bound. Graham’s value 20 was submitted before Chen’s 17 appeared.

TypeFirst character and zeronnα\alpha
rrχ1\chi_{1} real, ρ1\rho_{1} real1111
rcχ1\chi_{1} real, ρ1\rho_{1} nonreal1122
complexχ1\chi_{1} nonreal2211

Table 1.

The conjectured scale is much smaller. Chowla [24] conjectured P(a,q)≪εq1+εP(a,q)\ll_{\varepsilon}q^{1+\varepsilon} for every ε>0\varepsilon>0; the maximum over reduced classes is conjecturally of order φ(q)(log⁡q)2\varphi(q)(\log q)^2 [138, 109, 42, 76]. The Generalized Riemann Hypothesis gives admissible exponents 2+ε2+\varepsilon; see [4, 73, 15] for explicit conditional bounds. Explicit [9] and uniform [128] unconditional estimates for primes in progressions are also known. In the other direction, an exceptional (Siegel) zero helps: Heath-Brown [50] showed that a sufficiently strong exceptional zero forces P(a,q)≤q3+δP(a,q)\leq q^{3+\delta}, in an effective form that we use below, and q2+δq^{2+\delta} ineffectively; see also [33, 61, 147]. Other proofs of Linnik’s theorem use pretentious methods [114], sieve methods [36], or avoid LL-functions [80], but give larger exponents. For moduli with special multiplicative structure there are stronger zero-free regions [7, 38, 59], and much smaller exponents are known [16].

From zeros to primes

Write L=log⁡q\mathcal{L}=\log q and express a zero of a Dirichlet LL-function as

ρ=1−λL+iμL.\rho=1-\frac{\lambda}{\mathcal{L}}+i\frac{\mu}{\mathcal{L}}.

Thus small λ\lambda means that the zero is close to the line σ=1\sigma=1; μ\mu is its height multiplied by log⁡q\log q. We group each nonprincipal character with its complex conjugate and call this a family. The parameters λ1,λ2,λ3\lambda_{1},\lambda_{2},\lambda_{3} record the first zero of successive families in a specified rectangle near 1. The parameter λ′\lambda^{\prime} records the next zero occurrence in the first family, after removing the distinguished zero and its conjugate where appropriate. Precise definitions, including multiplicities and ties, are given in Section 2.

A nonnegative weight hh, supported in [A,L][A,L] with A=L−2TA=L-2T and T=0.416829T=0.416829, detects primes between qAq^{A} and qLq^{L}. Let HH be its Laplace transform. Xylouris’s positivity criterion [145] (3.57) gives

∑p≡a (q)log⁡pph(log⁡pL)≥LH(0)φ(q)(1−W−η),\sum_{p\equiv a\ (q)}\frac{\log p}{p}h\left(\frac{\log p}{\mathcal{L}}\right)\geq\frac{\mathcal{L}H(0)}{\varphi(q)}(1-W-\eta),

where η=10−6\eta=10^{-6} and

W=1H(0)∑χ≠χ0∑ρ∈RP∣H((1−ρ)L)∣.W=\frac{1}{H(0)}\sum_{\chi\ne\chi_{0}}\sum_{\rho\in R_{P}}\left|H((1-\rho)\mathcal{L})\right|.

Here RPR_{P} is a fixed square in the normalized coordinates, and zeros are counted with multiplicity. It is therefore enough to prove W≤1−2ηW\leq1-2\eta. The contribution of an individual zero decays exponentially with λ\lambda, so zeros close to 1 require the most accurate estimates.

We combine three kinds of information about these zeros. First, a per-character bound estimates the total contribution of one LL-function from a lower bound for the parameters of its zeros. Second, a far-density bound limits a weighted sum over characters, with weight depending on a selected zero of each character. Third, near-density bounds constrain the characters whose selected zeros lie closest to 1. Zero-location estimates for λ1\lambda_{1}, λ′\lambda', λ2\lambda_{2}, and λ3\lambda_{3} determine which of these bounds are available in a given case.

The new near-density estimate.

Heath-Brown’s near-density estimate [51] bounds the number N(λ)N(\lambda) of characters with a zero in σ≥1−λ/L\sigma\ge1-\lambda/\mathcal{L}, ∣t∣≤1|t|\le1. Its proof applies the Cauchy–Schwarz inequality to a prime sum at the single point β1=1−λ1/L\beta_{1}=1-\lambda_{1}/\mathcal{L}, and compares the resulting Gram form of the characters with the diagonal. Because every character is tested at the same point, a zero at distance λ\lambda counts only through the single number N(λ)N(\lambda), and the bound is weaker than the density estimate of Heath-Brown’s §11 as soon as λ≥1.1\lambda\ge1.1 and void soon after. That density estimate, proved with Selberg sieve weights after Graham, has the form eq:11.4 of a sum over characters of a weight that depends on each character’s own distance λ(k)\lambda^{(k)}, and it contains the count eq:11.5 of the characters with λ(k)≤λ\lambda^{(k)}\le\lambda. In his closing list of possible improvements Heath-Brown writes [51]:

It would be nice to have a weighted version of Lemma 12.1, in the way that eq:11.4 is a weighted version of eq:11.5. Unfortunately, no neat way of achieving this seems available.

The graded near-density lemma (Theorem 8.4) supplies such an estimate. We test each character at its own anchor, a real parameter determining the point at which a smoothed prime sum is evaluated. Cauchy–Schwarz relates these prime sums to a Gram form of the characters. We estimate its off-diagonal terms by the local explicit formula and, where sieve weights are used, Burgess’s bound; Graham’s sieve asymptotic controls the diagonal terms. Allowing the anchors to vary is the grading in the name of the lemma.

After normalization, the resulting constraint has the form

(∑jajvj)+2≤∑jDjaj2+d(∑jaj)2(aj≥0).\left(\sum_{j} a_{j}v_{j}\right)_{+}^{2} \le\sum_{j} D_{j}a_{j}^{2} + d\left(\sum_{j} a_{j}\right)^{2} \qquad(a_{j} \ge0).

Here x+=max⁡(x,0)x_{+}=\max(x,0), the index jj runs over character–height entries, vjv_{j} is a lower bound for the response of an entry, and Dj>0D_{j}>0 and d>0d>0 bound its diagonal and common correlation costs. Most entries correspond to one selected zero; some retain a second zero of the same character, with an explicit correction for their correlation.

Proposition 8.8 turns this inequality into a form suited to computation: there is a number τ∈[0,d]\tau\in[0,\sqrt{d}] such that

∑j(vj−τ)+2Dj≤1−τ2d.\sum_{j}\frac{(v_{j}-\tau)_{+}^{2}}{D_{j}} \le1-\frac{\tau^{2}}{d}.

A single constraint now distinguishes zeros at different distances from 1. The sieve weighting has precedents in Graham’s work [39, 41], Heath-Brown’s far-density argument, and the sieve-weighted large-sieve inequalities of Motohashi [91].

Where the new lemma enters. In Heath-Brown’s proof the near-density estimate enters at the very end, in the assembly of [51], §15. There the characters with a zero below a level Λ\Lambda are sorted into bins (Λr+1,Λr](\Lambda_{r+1},\Lambda_{r}], with Λr=Λ−0.025r\Lambda_{r}=\Lambda-0.025r; each character is charged the cost of the end of its bin nearer to s=1s=1; and the numbers of characters in the bins are controlled only through the counts N(Λr)N(\Lambda_{r}), bounded one level at a time by Lemma 12.1 (his Table 13). The far-density lemma controls the characters beyond Λ\Lambda. Xylouris’s assembly [145], §6 has the same structure, with a sharper form of Lemma 12.1 (his Lemma 5.3). A count N(λ)N(\lambda) treats all characters with a zero below λ\lambda alike, so it cannot tell a few zeros very close to s=1s=1 from many zeros further away. In our proof this place is taken by the near rows of the leaf programs described below. Each near row is one instance of Theorem 8.4: a single quadratic constraint in which every bin of characters enters with its own feature, computed at its own anchor. The weighted version of Lemma 12.1 that Heath-Brown asked for is thus used exactly where his count was used, in every leaf program, and it is the main source of the improvement (Section 1.6).

The finite case analysis.

We divide the middle range 0.1≤λ1<1.50.1\le\lambda_{1}<1.5 into cases determined by the type of the first family and intervals for the first few zero parameters. There are four levels in the computation, serving different purposes.

  1. A root case specifies a range of possible zero configurations. The 58 parent rows taken from published zero-location tables lead to 4453 root cases, covering the middle range.

  2. A leaf is a terminal case after further subdivision and valid zero-location exclusions. There are 4788 leaves that require numerical bounds.

  3. A leaf program groups characters into bins according to their selected zero parameters. Its nonnegative variables are the numbers of characters in these bins, together with weighted masses for the tails. The objective bounds WW, and its constraints express the far-density bound, restrictions on the first families, and the near-density inequalities. For fixed thresholds these constraints are linear.

  4. A threshold box bounds the auxiliary numbers τ\tau in the near rows. On each box, a linear relaxation is either excluded or bounded by exact dual multipliers. Subdividing these boxes covers every possible choice of thresholds.

The case tree partitions information about zeros; the threshold boxes partition the auxiliary parameters used to bound them. Keeping these two subdivisions distinct is essential to the coverage argument.

Each leaf, or each subcase of a leaf, has its own certificate, a binary tree of threshold boxes. The 4788 leaves give 5591 such certificate trees, because some leaves are subdivided further, and these trees contain 4,196,879 threshold boxes. On each box the near rows are replaced by linear relaxations, some with two alternative forms, and every choice of alternatives, a relaxation case, is certified separately: the boxes carry 29,397,336 relaxation cases, of which 9430 are shown to be infeasible and the others are bounded by their own integer dual multipliers. The largest certified value is 0.9999998223…<10.9999998223\ldots< 1; it bounds the zero sum together with the required error allowance. All certificate inequalities are checked in integer arithmetic. The subdivisions and zero-location arguments of the tree do not depend on the target exponent; only the leaf programs do. Table 2 compares these numbers with the case analyses of Heath-Brown and Xylouris, and Table 10 lists every level of the analysis.

[51][145]This paper
L=5.5L=5.5 (1992)L=5L=5 (2011)L=3.99L=3.99
Near-density inputcounts N(λ)N(\lambda) from Lemma 12.1counts N(λ)N(\lambda) from a sharper Lemma 12.1the weighted (graded) lemma, Theorem 8.4
Numerical tables1211 more58 rows from theirs, plus over 100 new
Final case analysis14 intervals of λ1\lambda_121 cases, each split by two zero counts4453 roots, 4788 leaves
How a case is closedfloating point, without rounding analysisMaple, 10-digit floating pointexact certificates on 4.2 million boxes
Largest value of WW (must be below 1)0.99430.9980.99999982

Table 2. Three proofs of an explicit Linnik constant compared, in the middle range of the first zero (λ1≥0.348\lambda_1 \ge0.348 for Heath-Brown and Xylouris, 0.1≤λ1<1.50.1 \le\lambda_1 < 1.5 here). The last row is the largest bound for the normalized zero sum WW over all cases of the final assembly: for Heath-Brown the largest total he reports, for Xylouris his stated bound, and for this paper the largest certified value, a bound for W+2ηW+2\eta. The entry “without rounding analysis” paraphrases Heath-Brown’s own remark [51], §1, quoted above.

LevelNumberWhat it partitions or representsWhere definedHow it is checked
Parent rows58Published implications: type t and λ1∈[Aj,Bj]\lambda_{1} \in[A_{j}, B_{j}] give λ′≥pj\lambda' \ge p_{j}, λ2≥rj\lambda_{2} \ge r_{j} (33 rr, 13 rc, 12 complex)Proposition 7.5Printed tables; Heath-Brown’s Tables 4 and 7 recomputed
Base cells340Parent intervals cut into cells of width 0.01 (145, 89, 106)Section 11.3Validated in the replay
First-zero cells478Tiling of the base cells (195, 115, 168), with bounds p≤λ′p \le\lambda' and l2≤λ2l_{2} \le\lambda_{2}Section 11.3Validated in the replay
Gap cases684422 cells without a gap; 262 gaps for λ′\lambda' partitioning 56 complex cellsSection 11.3Validated in the replay
Specifications2768Gap cases with a second-family record: 1083 rr (195 unreserved, 444 + 444 reserved), 115 rc, 1570 complexDefinition 11.2Validated in the replay
Root cases44532768 specifications inside, and the 1685 not of type rr also outsideTheorem 11.4Hand proof; replay
Refinement nodes19591029 splits, 236 restricting location updates, 694 exclusions (75 by real location updates, 422 by lower bounds for λ2\lambda_{2} or λ′\lambda', 197 by positivity); 4453+1029=4788+6944453 + 1029 = 4788 + 694Section 11.4, Table 9Replay; some location rows by separate programs
Leaves47883146 inside (233 supplementary) and 1642 outside; 343 inside and 70 outside roots have noneSection 11.4Replay; Lean numeric checker
Certificate trees5591One per leaf or subcase of a leaf: 3949 inside (349 for supplementary leaves), 1642 outsideSection 11.6Replay; independent checker; Lean checker
Threshold boxes4,196,879Terminal boxes of the trees: 4,109,455 inside (1,042,363 for supplementary leaves), 87,424 outsideDefinition 10.6As above
Relaxation cases29,397,3362k2^{k} in a box with kk tangent rows; 9430 closed by an exclusion, the rest by integer dualsDefinition 10.6As above, in exact arithmetic
Rows and exclusions
Near rows16,842All family, shifted and graded rows (13,528 inside, 3314 outside), with 11,135,323 entriesSection 9.5Interval arithmetic; Lean row checker
Two-test rows37On 37 inside roots: 31 single, 2 mixture, 4 pairedSection 9.6Replay; exact comparison
Complex location rows9425 for λ2\lambda_{2} and 69 for λ′\lambda' (Propositions 7.12 and 7.13)Appendix AInterval arithmetic; separate program
Degree-5 real rows22Cells covering [0.700,0.755][0.700, 0.755], used in 300 location nodesProposition 7.10, Table 6Interval arithmetic; replay
Positivity exclusions197Reserved nodes of type rr; least normalized margin 7.8⋅10−57.8 \cdot10^{-5}Proposition 7.14Interval arithmetic; replay

Triples of numbers are counts for the types rr, rc, complex; “supplementary” refers to the supplementary leaves of Section 11.4. Each split adds one terminal node; hence the identity in the refinement row. The printed tables were compared with the pages of [51, 145], and Heath-Brown’s Tables 4 and 7 were recomputed in floating point (Section 7.1). The replay does not regenerate the 94 complex rows or the conditions [145] (4.29), (4.34); separate interval programs check them. The Lean row checker covers all near rows except the 37 two-test rows.

Table 10. The finite analysis of the middle range 0.1≤λ1<1.50.1 \le\lambda_{1} < 1.5 at a glance. The case tree, from the parent rows to the leaves, partitions configurations of zeros; the boxes and relaxation cases partition the thresholds of the near rows. The checks are described in Section 14.

NodeInsideOutside
first-zero / second-family / gap splits1 / 408 / 5930 / 27 / 0
location updates, real (restrict / exclude)225 / 75–
location updates, complex (restrict / exclude)11 / 0–
exclusions by lower bounds for λ2\lambda_{2} or λ′\lambda'35270
exclusions by positivity1970
leaves31461642

Table 9. Nodes of the refinement trees.

The exterior ranges are shorter arguments. An exceptionally small λ1\lambda_{1} is covered by Heath-Brown’s theorem on Siegel zeros. The rest of λ1≤0.1\lambda_{1} \le0.1 is handled by his quantitative Deuring–Heilbronn estimates and zero-location tables. A single further certificate treats λ1≥1.5\lambda_{1} \ge1.5. If the relevant rectangle contains no zero, then W=0W = 0 and the positivity criterion applies immediately.

Comparison with Heath-Brown and Xylouris

Our framework is Heath-Brown’s [51] with Xylouris’s refinements [145, 146]. We use the following from their work:

  • the positivity criterion [145], (3.57);

  • the per-character bound [145], Lemma 3.10;

  • the far-density lemma [145], Lemma 5.1;

  • the zero-free regions and Deuring–Heilbronn estimates, as printed in their tables;

  • the local explicit formulas [145], Lemmas 3.1, 3.2, which are [51], Lemmas 5.3, 5.2.

The case analyses compared. Both earlier proofs are computer-assisted, and both are organized, like ours, as a case analysis on the first zeros; Table 2 compares the three. Heath-Brown’s paper has twelve numerical tables of zero-location and density bounds. Its final assembly [51], §15 treats λ1≥0.348\lambda_{1} \ge0.348 in fourteen intervals, mostly of length 0.02, each by a chain of inequalities evaluated numerically; the largest total it reports is 0.9943. Heath-Brown writes [51], §1 that “the arguments rely heavily on numerical calculations. These we do not reproduce in full, nor have we attempted any rigorous analysis of the rounding and truncation errors in the computer algorithms employed.” Xylouris’s dissertation adds eleven tables. Its final assembly [145], §6.4, pp. 86–87 has 21 cases: three intervals of λ1\lambda_{1} for a real zero of a real character, two for a complex zero of a real character, and sixteen for a nonreal character. Each case is split further according to the possible values of two zero-counting functions [145], (6.46), and the computations were done in Maple with ten digits [145], p. 14; the largest value is W<0.998W < 0.998. In both, each case is closed by a chain of inequalities in which the unknown zeros are estimated separately, one after another.

In our proof the same role is played by 4453 root cases and 4788 leaves, and each leaf is closed by a linear program with an exact certificate instead of a floating-point evaluation. The 4.2 million boxes have no counterpart in the earlier work. They do not subdivide the configurations of zeros. They subdivide the auxiliary thresholds of the near rows, and they arise because the graded lemma is a quadratic constraint rather than a count. The largest certified value, 0.99999982 against 0.9943 and 0.998, shows how much finer the analysis must be at L=3.99L = 3.99.

Where the gain comes from. The gain from 5 to 3.99 has three sources:

  • the graded near-density lemma, which replaces the counts N(λ)N(\lambda) and applies to every range of λ\lambda at once;

  • the linear-programming combination, which lets all constraints act together on each case instead of chaining worst cases;

  • a fine case analysis of the first zeros, possible because every case is certified by machine.

The last source alone does not go far. Xylouris estimates that his method, with much larger computations (smaller intervals and finer grids), would give about 4.96 [146], p. 82.

The certified exponent and numerical evidence.

We state the result at 3.99 because this is the exponent for which the complete cover has been certified. The tightest cases have a nonreal first character and λ1∈[0.64,0.80]\lambda_1 \in[0.64,0.80]. In this range the available lower bounds for λ2\lambda_2 are comparatively small, which restricts the anchors in the near-density estimates. Several of the corresponding leaf bounds are within 10−710^{-7} of 1 after the error allowance, and require fine subdivisions.

Exploratory floating-point computations suggest that a somewhat smaller exponent may be accessible. One difficult leaf, with the reserved-family columns of Section 11.5, has computed values 0.992 at L=3.97L=3.97 and 1.020 at L=3.95L=3.95. These are values of particular relaxations, not certified global bounds or an obstruction to other methods. Similarly, the positive numerical margins for the exterior ranges at smaller exponents do not constitute a complete proof there. The verification reported in this paper is at L=3.99L=3.99.

Verification and a guide to the proof.

The certificates have been checked by independent programs. Parts of the analytic argument and the soundness of the checkers have also been formalized in Lean, with the published analytic inputs stated as hypotheses. This formalization includes a theorem that a valid certified leaf yields a prime when the modulus satisfies that leaf’s hypotheses about zeros. It does not include the full passage from arbitrary moduli to leaves, or the exterior arguments. Section 14 gives the verification records and separates these claims precisely.

Sections 2 and 3 give the notation and published inputs. Sections 4 to 6 reduce prime detection to bounds for the zero sum, and Section 7 supplies the zero-location estimates. The main new analytic argument is in Section 8; Section 9 derives the constraints used in the computation.

Section 10 proves certificate soundness, while Section 11 shows that the zero configurations are covered and satisfy the programs. The exterior cases and the final choice of constants are treated in Sections 12 and 13. A reader interested first in the new density estimate may read Sections 2 and 3 and then Section 8, returning to the numerical construction afterwards.

Figure 1 shows how the main results depend on one another, and which of them are published inputs, new results, computations or formally verified statements. The terms of the case analysis are introduced in Section 2.4, and the fixed constants and the error budget in Section 2.7. Table 10 in Section 11 lists all levels of the finite analysis, and Table 12 in Section 14 records, for each component, its source and how it was checked.

Proof map for Theorem 1.1, showing published inputs, results proved in the paper, computations, and statements formally verified in Lean

Figure 1. Proof map for Theorem 1.1. An arrow points from a result to a result whose proof uses it; the joined lines into (R3) are the ingredients of Theorem 11.7. The regimes (R1)–(R4) of Section 13 are defined by the first zero λ1\lambda_{1}, and (R2) splits at the constant uexcu_{\mathrm{exc}} of Section 12. A double border marks a statement formally verified in Lean, with the published inputs as hypotheses, or, for the certificates, a computation accepted by a checker proved sound in Lean. The formalization also covers the far budget and parts of Propositions 9.7 and 11.6; the zero-location results, the case tree, the costs, the two-test row, the exterior regimes and the assembly are not formally verified (Section 14.3). Some arrows are omitted, for instance from the explicit formulas to Sections 5 and 7 and into Theorem 12.6.

ComponentSourceKindMethodChecked by
Positivity criterion and detection (Input 4.2 and Proposition 4.4)Published [145]Analytic–Lean, with the criterion as a hypothesis
Envelope and first-family bounds (Section 5)New, after [145]AnalyticInterval arithmeticHand proof; a hypothesis in Lean, which checks only the enclosures
Far density (Input 6.1)Published [145]Analytic–Lean hypothesis; profile conditions proved in Lean
Far constants VV, ww (Section 6)New profilesNumericalInterval arithmeticMultiple-precision audit; Lean numeric checker
Printed zero-location tables (Proposition 7.5)Published [51, 145]AnalyticAs printedCompared with the pages; Heath-Brown’s Table 4 and 7 recomputed in floating point
New location rows (Section 7)NewBothInterval arithmeticReplay; the 94 complex rows by a separate program
Conditions [145], (4.29), (4.34)Published; new checkNumericalInterval arithmeticSeparate programs
Conductor refinement 1108\frac{1}{108} (Section 7.6)NewAnalytic–Hand proof
Graded near lemma, response lemma, threshold form (Section 8)NewAnalytic–Lean (Comparator, two kernels)
Family, shifted and graded rows (Section 9)NewNumericalInterval arithmeticLean row checker (16,842 rows); multiple-precision audit
Two-test rows (Section 9.6)New, after [145]BothInterval arithmeticHand proof; replay; exact comparison with a separate implementation on every certificate interval (constants: that implementation’s interval code only)
Certificate soundness (Theorem 10.7)NewAnalytic–Lean (Comparator, two kernels)
Leaf certificates (Definition 10.6)NewNumericalExact integerReplay; independent checker; Lean checker
Leaf data (Section 11.5)NewNumericalInterval arithmeticMultiple-precision audits; Lean numeric checker (count-type integers: audit only)
Case tree and realization (Section 11)NewAnalytic–Hand proof; replay
Tiny exceptional zero (Proposition 12.1); empty rectanglePublished [50]; W=0W=0Analytic–As printed
Small exceptional zero (Theorem 12.5)New, from lemmas of [51]BothInterval arithmeticHand proof; exterior replay; margins recomputed independently
Large first zero (Theorem 12.6)NewNumericalExact integerReplay’s exact checker; independent checker; Lean certificate (native and kernel), numeric and row checkers
Assembly (Section 13)NewAnalytic–Hand proof

Table 12. The interval arithmetic is described at the beginning of this section. Floating point is used only to propose parameters, dual multipliers and sieve heights, which need no justification, and to recompute published tables. TABLE 12. Status of the components. Published inputs are used as printed, and those used in the Lean theorems enter them as hypotheses; “Both” means an analytic argument with numerical margins; “hand proof” means a conventional proof in this paper. The checks are described in Section 14.

Notation and conventions

We use standard notation for Dirichlet characters and their LL-functions. For background on Dirichlet LL-functions and their zeros we refer to [57, 131, 69, 85, 26, 62, 87, 124], and for the history of the subject to [94]. The distinctions between a zero and a zero occurrence, and between normalized and physical height, will be used throughout the proof.

Moduli, characters and zeros

Throughout, q≥2q \ge2 is an integer modulus and

L=log⁡q.\mathcal{L}=\log q.

Characters are Dirichlet characters modulo qq; χ0\chi_{0} is the principal character. For a nonprincipal character χ\chi we consider the zeros ρ=β+iγ\rho=\beta+i\gamma of L(s,χ)L(s,\chi) with β>0\beta>0, counted with multiplicity, and we write

ρ=1−λL+iμL,λ=(1−β)L,μ=γL.(2)\rho=1-\frac{\lambda}{\mathcal{L}}+i\frac{\mu}{\mathcal{L}},\qquad\lambda=(1-\beta)\mathcal{L},\qquad\mu=\gamma\mathcal{L}. \tag*{(2)}

We call λ=λρ\lambda=\lambda_{\rho} the parameter of ρ\rho and μ=μρ\mu=\mu_{\rho} its normalized height; γ=γρ\gamma=\gamma_{\rho} is its physical height. Thus a height bound always concerns ∣γ∣|\gamma| or ∣μ∣|\mu|. Smaller λ\lambda means a zero farther to the right. The multiplicity of ρ\rho as a zero of the specified L(s,χ)L(s,\chi) is mρm_{\rho}. A zero occurrence is one copy of the pair (χ,ρ)(\chi,\rho): a zero of multiplicity mρm_{\rho} gives mρm_{\rho} occurrences. Unless stated otherwise, sums over zeros count multiplicity. A zero of an imprimitive character is a zero of the LL-function of the primitive character inducing it, with the same multiplicity, as long as β>0\beta>0; all zeros considered below have β>1/2\beta>1/2.

Regions

Following [145] (3.7) we put, for x>0x>0,

R(x)={σ+it: 1−log⁡log⁡L3L≤σ≤1, ∣t∣≤x}.R(x)=\left\{\sigma+it:\ 1-\frac{\log\log\mathcal{L}}{3\mathcal{L}}\leq\sigma\leq1,\ |t|\leq x\right\}.

There is an integer l=l(q)l=l(q) with 1≤l≤L/101\leq l\leq\mathcal{L}/10 such that no zero of any L(s,χ)L(s,\chi) lies in R(10l)∖R(l)R(10l)\setminus R(l) (Input 3.5). We fix one such choice, for example the least admissible integer. The distinguished zeros (Definition 2.2) and the representatives (Section 2.4) are selected in R(l)R(l). The local explicit formula also involves disc zeros, which need not all belong to this rectangle. Three further regions have separate roles:

  • the prime window RP={0≤λ≤CP, ∣μ∣≤CP}R_{P}=\{0\leq\lambda\leq C_{P},\ |\mu|\leq C_{P}\}, whose zeros are the only ones that enter the zero sum;

  • the buffer RB={0≤λ≤CB, ∣μ∣≤CB}R_{B}=\{0\leq\lambda\leq C_{B},\ |\mu|\leq C_{B}\}, with CB>CP+MC_{B}>C_{P}+M;

  • the height-one region ∣γ∣≤1|\gamma|\leq1, in which the far-density lemma operates.

The constants CPC_{P}, MM and CBC_{B} are fixed independently of qq, in the order described in Section 13. For sufficiently large qq, RP⊂RB⊂R(l)R_{P}\subset R_{B}\subset R(l), and the buffer has physical height CB/log⁡q<1C_{B}/\log q<1. The word inside means that the selected first zero belongs to RBR_{B}; outside means that it does not. It does not refer to R(l)R(l).

Families and the first zeros.

Definition 2.1 (Families and successive minima). A family is the set F(χ)={χ,χ‾}\mathcal{F}(\chi)=\{\chi,\overline{\chi}\} for a nonprincipal character χ\chi. It has one member when χ\chi is real and two otherwise. If there is a nonprincipal zero in R(l)R(l), select a pair (χ1,ρ1)(\chi_{1},\rho_{1}) with least parameter and write λ1=λρ1\lambda_{1}=\lambda_{\rho_{1}}, μ1=μρ1\mu_{1}=\mu_{\rho_{1}}, and F1=F(χ1)\mathcal{F}_{1}=\mathcal{F}(\chi_{1}). Remove all occurrences belonging to F1\mathcal{F}_{1} and repeat to define (χ2,ρ2)(\chi_{2},\rho_{2}), λ2\lambda_{2}, and F2\mathcal{F}_{2}; continue in this way. Thus λ1≤λ2≤λ3≤⋯\lambda_{1}\leq\lambda_{2}\leq\lambda_{3}\leq\cdots. We call F1\mathcal{F}_{1} the first family, F2\mathcal{F}_{2} the second family and F3\mathcal{F}_{3} the third family, and (χ1,ρ1)(\chi_{1},\rho_{1}) the first zero. If a subsequent minimum is over an empty set, its value is +∞+\infty. Ties may be resolved by any admissible choice; each result used below holds for all such choices. If the initial set is empty, we use the empty-region argument of Section 12.

These are the selections of Xylouris [145]. They order families, not the individual zeros of one LL-function.

Definition 2.2 (The additional first-family zero). The distinguished occurrences are one copy of ρ1\rho_{1} as a zero of χ1\chi_{1}, together with one copy of ρ1‾\overline{\rho_{1}} as a zero of χ1\chi_{1} if χ1\chi_{1} is real and ρ1\rho_{1} is not, or as a zero of χ1‾\overline{\chi_{1}} if χ1\chi_{1} is not real. We write λ′\lambda' for the least parameter of a zero occurrence of χ1\chi_{1} or χ1‾\overline{\chi_{1}} in R(l)R(l) that is not distinguished; thus λ′=λ1\lambda'=\lambda_{1} when ρ1\rho_{1} is a multiple zero. If no occurrence remains, put λ′=+∞\lambda'=+\infty.

When λ′\lambda' is finite, it can always be realized by a zero of χ1\chi_{1} itself. Indeed, zeros of L(s,χ1‾)L(s,\overline{\chi_{1}}) are conjugates of zeros of L(s,χ1)L(s,\chi_{1}), with the same multiplicities. Replace any minimizing occurrence of χ1‾\overline{\chi_{1}} by its conjugate occurrence of χ1\chi_{1}. We call every resulting non-distinguished occurrence ρ′\rho' with parameter λ′\lambda' admissible (a second copy of ρ1\rho_{1} is admissible when ρ1\rho_{1} is a multiple zero), and write γ′\gamma' for its height; as ρ′∈R(l)\rho'\in R(l), ∣γ′∣≤l|\gamma'|\leq l. The arguments below hold for every admissible choice; in the second-zero split of Section 11.6 the choice is the one used there.

The following table fixes the three types and the two counts used in the first-family bounds. Here n=∣F1∣n=|\mathcal{F}_{1}|, while α\alpha counts distinguished occurrences per character.

For type rc we choose the conjugate with μ1>0\mu_{1}>0.

The vocabulary of the case analysis.

In the middle range 0.1≤λ1<1.50.1\leq\lambda_{1}<1.5 the proof is a finite case analysis, defined precisely in Sections 10 and 11. Its terms are needed in the analytic sections that precede those, and we introduce them here.

  • Anchor. The local explicit formulas of Section 3 are applied at points s=1−σ/L+iγs=1-\sigma/\mathcal{L}+i\gamma; the real number σ\sigma is the anchor. An inequality anchored at σ\sigma is useful only if the zeros near ss with parameter below σ\sigma are absent or are retained explicitly, so anchors are chosen at or below known lower bounds for zero parameters.

  • Specification. A specification s\mathfrak{s} (Definition 11.2) records the type of the first family; an interval [a,b][a,b] containing λ1\lambda_{1}, the first-zero cell; a lower bound p≤λ′p\leq\lambda' and possibly an interval [lo′,hi′][\mathrm{lo}',\mathrm{hi}'] containing λ′\lambda', called a gap; a lower bound l2≤λ2l_{2}\leq\lambda_{2}; a lower bound rr for the parameters of the zeros of height at most 11 of the families other than the first; possibly a reservation (next item); and whether ρ1\rho_{1} lies inside or outside the buffer. Its configuration set C(s)C(\mathfrak{s}) consists of the zero configurations compatible with these data.

  • Reserved and ordinary characters. A reservation concerns a family other than the first whose least parameter among zeros of height at most 11 is as small as possible. It fixes the

number n2∈{1,2}n_{2}\in\{1,2\} of its characters and an interval [lo⁡2,hi⁡2][\operatorname{lo}_{2},\operatorname{hi}_{2}] for that parameter. This reserved family need not be F2F_{2}. The characters outside the first and the reserved family are ordinary.

  • Representatives. The height-one representative of a nonprincipal character χ\chi is a zero of L(s,χ)L(s,\chi) in R(l)R(l) with ∣γ∣≤1|\gamma|\leq1 and least parameter; its T∗T^{*}-representative is defined in the same way with ∣γ∣≤T∗|\gamma|\leq T^{*}, for a height T∗∈[12,1]T^{*}\in[\frac{1}{2},1] fixed in Lemma 6.3 (Definition 6.4). The leaf programs place a character by the parameter of one of its representatives.

  • The case tree. The roots are 4453 specifications whose configuration sets cover the middle range. Each root is refined by splitting intervals and by applying zero-location results. This gives a finite tree of specifications, the root’s refinement tree; its vertices are nodes, and together these trees form the case tree. A node at which a zero-location result sharpens one of the bounds is a location update; a node whose configuration set is shown to be empty is excluded. The terminal nodes that are not excluded are the leaves. Some leaves are divided further into finitely many subcases.

  • Leaf programs, bins and columns. Each leaf or subcase has a leaf program (Definition 10.1) whose value bounds WW for every configuration in it. The parameters of the ordinary characters (and in some leaves those of the reserved characters) are divided into finitely many intervals, the bins. The column of a bin is the variable of the program that counts the characters whose representative zero (Definition 6.4) has parameter in the bin. The ordinary characters beyond the last bin form the tail, whose column is a weighted mass rather than a count. In an outside leaf the first family is represented by hidden columns, which bin the least parameter tχt_{\chi} of the zeros in RPR_{P} of each first-family character (Section 5.7). The constant first bounds the cost of the first family in an inside leaf (0 in an outside leaf), plus any fixed charge for the reserved family, and final is an allowance for the error terms.

  • Far budget. The far-density lemma (Input 6.1) bounds a sum over all characters of a positive decreasing function ww of the parameter of a selected zero of height at most 1. We call w(λ)w(\lambda) the far weight of such a zero, and the bound (9) the far budget; in a leaf program it is one linear constraint.

  • Near rows. The other constraints of a leaf program are conditions on counts and the near rows. A near row is, with one exception (Section 9.6), an instance of the graded lemma (Theorem 8.4) in the form of Proposition 8.8. Its terms are entries, one for each character (occasionally each zero) that it includes. Each entry has a feature, a lower bound for how strongly its zero is detected, and a diagonal, what it costs on its own, and one correlation term bounds the interaction between distinct entries (Proposition 8.8). The entries of a family whose zeros are known are grouped into family terms. A row is built from two test functions, a detector and a Gram test (Definition 8.1); by the choice of its anchors and sieve weight it is a family, shifted, graded or two-test row (Table 8). It is linear in the columns once an auxiliary number, its threshold τ\tau, is fixed.

  • Certificates. A certificate of a leaf or subcase is a binary tree, the certificate tree, whose nodes are boxes of thresholds, starting from a box that contains every admissible threshold; its leaves are the terminal boxes. On a terminal box each near row is replaced by a linear relaxation, and each choice among the alternatives of these relaxations, a relaxation case, is either shown to be infeasible (an exclusion) or bounded below 1 by integer dual multipliers, in exact integer arithmetic (Section 10).

RowReference anchor ssSieve weightsEntries, insideEntries, outside
familys1s_{1}none (Ξ=0)(\Xi=0)(a), (f) and (b), (c) or (d)(a), (f), (g)
shiftedmin⁡(1.9,p∗,l2)\min(1.9,p^{*},l_{2})none (Ξ=0)(\Xi=0)(a), (e), (f)not used
graded1.5, 1.6, 1.9, 2.1 or 2.3present(a), (b), (f)(a), (f)
two-testthe shift of the leaftwo test functionsTheorem 9.4, on 37 rootsnot used

Table 8. The kinds of near rows. The shifted row is used only if s>s1s>s_{1} and s≥as\geq a.

Elementary notation

For real xx we write x+=max⁡(x,0)x_{+}=\max(x,0), and 1E\mathbf{1}_{E} denotes the indicator of a condition or set EE. The symbol φ(q)\varphi(q) is Euler’s totient; the separately defined φ(χ)\varphi(\chi) is a conductor coefficient (Section 3). A cost means an upper bound for a contribution to the normalized zero sum WW; a budget is the right-hand side of a constraint on these contributions. All finite decimals used as parameters are exact rational numbers unless an approximation is explicitly indicated.

Asymptotic conventions

Statements about zeros below are understood for all sufficiently large qq, unless a different range is stated. The threshold q0q_{0} depends on finitely many fixed choices (the exponent, test functions, weights, constants and certificates, and the tolerances of the published results we use); the order in which these are chosen is spelled out in Section 13. We write o(1)o(1) for a quantity tending to 00 as q→∞q \to\infty, uniformly in everything except the fixed choices. The number

η=10−6\eta= 10^{-6}

is a fixed tolerance used throughout, and the exponent is

L=3.99.L = 3.99.

Certificate coefficients use the scale S=1016S = 10^{16}. Counts of characters are unscaled; the precise conversion between coefficients and real quantities is given in Section 10.1.

Constants and the error budget

The fixed constants of the proof, with their values, their roles and the stage of Section 13.1 at which each is chosen, are listed in Table 15 in Appendix B. The tolerance η=10−6\eta= 10^{-6} is spent as follows; the allowance 5η5\eta in each leaf value covers the analytic errors with room to spare.

SymbolValueRoleWhere fixedStage
Exponent and kernel
LL3.993.99The exponentSection 2(S1)
TT0.4168290.416829Support of ψ\psi is [0,T][0,T]; hh is supported in [L−2T,L][L-2T,L]Section 4.1(S1)
κ\kappaT/16=0.0260518125T/16=0.0260518125Width of the 16 steps of ψ\psiSection 4.1(S1)
b0,…,b15b_0,\ldots,b_{15}Table 3Step heights of ψ\psiTable 3(S1)
AAL−2T=3.156342L-2T=3.156342Lower end of the support of hh; primes are found in (qA,qL)(q^A,q^L); A>3A>3 and A>2xA>2xSection 4.1, Lemma 6.2–
Ψ(0)\Psi(0)κ∑ibi=0.1936422928…\kappa\sum_i b_i=0.1936422928\ldotsIntegral of ψ\psiSection 4.1–
H0H_0Ψ(0)2=0.0374973375…\Psi(0)^2=0.0374973375\ldotsH(0)H(0); normalizes the zero sum WW(4.1)–
Tolerances
η\eta10−610^{-6}Detection needs W≤1−2ηW\leq1-2\eta; far budget (1+η)V(1+\eta)V; allowance 5η5\etaSection 2, Proposition 4.4(S1)
η′\eta'10−610^{-6}Tolerance of the near rows: in every feature, diagonal and correlation termSection 9.2(S1)
SymbolValueRoleWhere fixedStage
ε1\varepsilon_{1}η/23\eta/23Envelope error ε1e−Aλ\varepsilon_{1}e^{-A\lambda} per character; 23=19+423=19+4, where 19>Kfar19>K_{\mathrm{far}} bounds ∑χe−Aλχ\sum_{\chi}e^{-A\lambda_{\chi}} over the ordinary characters and 4 counts the first and reserved charactersSection 5.8(S1)
final/S\mathrm{final}/S5η5\etaAllowance included in every leaf valueSection 11.5(S1)
ε\varepsilon in Input 4.2ηH0\eta H_{0}Tolerance of the positivity criterionProposition 4.4(S5)
ε\varepsilon in Input 6.1ηJ2\eta J^{2}Tolerance of the far-density lemma; gives (9)Section 6(S5)
Analytic constants
λ11\lambda_{11}0.050.05Least anchor of the envelope; the anchor set A\mathcal{A} lies in [λ11,CP][\lambda_{11},C_{P}]Section 5.3(S1)
smax⁡s_{\max}33Largest anchor of an ordinary entry; K0K_{0} counts zeros with λ≤smax⁡\lambda\leq s_{\max}Section 6.3(S1)
CPC_{P}≥max⁡(C0(ηH0),3)\geq\max(C_{0}(\eta H_{0}),3)Size of the prime window RPR_{P}; C0C_{0} is the constant of Input 4.2(5)(S2)
NNN(CP)N(C_{P})Bound for the zeros of one character in RPR_{P}Lemma 5.3(S3)
hA,Ah_{\mathcal{A}},\mathcal{A}Mesh with ∣Bϕ(x)−Bϕ(y)∣≤ε1/4\lvert B_{\phi}(x)-B_{\phi}(y)\rvert\leq\varepsilon_{1}/4 for ∣x−y∣≤hA\lvert x-y\rvert\leq h_{\mathcal{A}}Finite anchor set: the grid λ11+hAZ\lambda_{11}+h_{\mathcal{A}}\mathbb{Z} and the anchors a,p∗,p∘a,p^{*},p^{\circ} of the leavesSection 5.3(S3)
θ\theta (smoothing)In (0,1)(0,1), with (5.1)L1L^{1}-distance of the smooth step function ψθ\psi^{\theta} from ψ\psiSection 5.2(S3)
coutc_{\mathrm{out}}CPfpθ(0)+∣(fpθ)′(0)∣+eCPT∥(fpθ)′′∥1C_{P}f_{p}^{\theta}(0)+\lvert(f_{p}^{\theta})'(0)\rvert+e^{C_{P}T}\lVert(f_{p}^{\theta})''\rVert_{1}, maximized over the anchors ppCost of the distinguished terms of an outside first familyProposition 5.9(S3)
K0K_{0}Independent of qqBound for the zeros with λ≤smax⁡\lambda\leq s_{\max} and ∣γ∣≤2\lvert\gamma\rvert\leq2Section 6.3(S4)
MMLargeHeight separation for Lemmas 6.3 and 8.7; depends on K0K_{0} and ηmin⁡′′\eta_{\min}''Section 13.1(S4)
CBC_{B}>max⁡(CP+M,1.5)>\max(C_{P}+M,1.5), with 4cout/(H0CB2)≤η4c_{\mathrm{out}}/(H_{0}C_{B}^{2})\leq\etaSize of the buffer RBR_{B}, which separates inside from outsideSection 2, Section 13.1(S4)
δ0\delta_{0}In (0,13](0,\frac{1}{3}]Common disc radius of the explicit formulasSection 5.3(S5)
q0q_{0}Not computedMaximum of the finitely many thresholdsSection 13.1(S6)
Far density (profile I; profile II)
J,ε0J,\varepsilon_{0}10, 10−710,\,10^{-7}Number of weights αi\alpha_{i}; offset in w0w_{0} (both profiles)Section 6(S1)
c1c_{1}0.09035; 0.08219220.09035;\ 0.0821922Parameter of Input 6.1Section 6(S1)
c2c_{2}0.235968; 0.21709030.235968;\ 0.2170903Parameter of Input 6.1Section 6(S1)
θfar\theta_{\mathrm{far}}1.28683; 1.49642741.28683;\ 1.4964274Exponent in w0(t)2w_{0}(t)^{2}Section 6(S1)
α1,…,α10\alpha_{1},\ldots,\alpha_{10}Table 4Weights of Input 6.1, of sum 1Table 4(S1)
xx1.1736846…; 1.1303335…1.1736846\ldots;\ 1.1303335\ldotsUpper end 23+3c1+c2\frac{2}{3}+3c_{1}+c_{2} of the profile; 2x<A2x<ALemma 6.2–
SymbolValueRoleWhere fixedStage
VV175.26640331…175.26640331\ldots; 243.32098105…243.32098105\ldotsRight side of Input 6.1 with ε=0\varepsilon=0; far budget (1+η)V(1+\eta)VSection 6–
w(0)w(0)10.7054…10.7054\ldots; 13.1730…13.1730\ldotsFar weight at λ=0\lambda=0Section 6–
KfarK_{\mathrm{far}}(1+η)V/w(0)(1+\eta)V/w(0): <16.38<16.38; <18.48<18.48Bound for ∑χe−Aλχ\sum_{\chi}e^{-A\lambda_{\chi}}(10)–
Leaf programs and certificates
SS101610^{16}Scale of the stored coefficientsSection 2, Section 10.1(S1)
TST_{S}10610^{6}Thresholds are measured in ticks of 1/TS1/T_{S}Section 10(S1)
DSD_{S}101210^{12}Scale of the dual multipliersSection 10(S1)
den\mathrm{den}One of 200, 400, 500, 800, 2000 per leafOrdinary bins between rr and R=max⁡(3,r)R=\max(3,r) on the grid 1/den1/\mathrm{den}Section 11.5(S1)
Reserved grid1/4001/400Grid of the columns of a reserved second familySection 11.5(S1)
Near rows
ε′\varepsilon'1/20001/2000Sieve levels ϑk=12(tk−13)=ε′\vartheta_{k}=\frac{1}{2}(t_{k}-\frac{1}{3})=\varepsilon'Section 9.2(S1)
Cells2000 sieve cells; 4000 cells for IuI_{u}Equal cells t0<⋯<t2000=2γft_{0}<\cdots<t_{2000}=2\gamma_{f}; upper Riemann sum for IBI_{B}Section 9.2(S1)
hk,μrowh_{k},\mu_{\mathrm{row}}Rational, denominators at most 101210^{12}; μrow=109\mu_{\mathrm{row}}=10^{9} in family and shifted rowsSieve heights; they vanish in family and shifted rowsSection 9.2(S1)
⌊δ⌋1/200\lfloor\delta\rfloor_{1/200}Grid 1/2001/200Offsets at which the diagonal D(δ)D(\delta) is evaluatedSection 9.2(S1)
s1s_{1}⌊50a⌋/50\lfloor50a\rfloor/50Safe anchor of a leaf with first-zero cell [a,b][a,b]Section 9.3(S1)
ss (shift)min⁡(1.9,p∗,l2)\min(1.9,p^{*},l_{2})Anchor of the shifted row and of the two-test rowSection 9.5, Section 11.5(S1)
Graded anchors1.51.5, 1.61.6, 1.91.9, 2.12.1, 2.32.3Reference anchors ss of the graded rows usedSection 9.5(S1)
Exterior regimes
uexcu_{\mathrm{exc}}min⁡(1/ηHB(12),0.08)\min(1/\eta_{\mathrm{HB}}(\frac{1}{2}),0.08)Below it, Proposition 12.1 (P(a,q)≤q3.5)(P(a,q)\le q^{3.5}); not computedSection 12(S1)
K,AsK,A_{s}0.18210.1821; L−2K=3.6258L-2K=3.6258Single triangle hL,Kh_{L,K} of the small branchTheorem 12.5(S1)
c1,c2,ϕc_{1},c_{2},\phi0.0570.057, 0.15540.1554, 13\frac{1}{3}Parameters of Input 12.4Theorem 12.5(S1)
ad,bd,Vda_{d},b_{d},V_{d}37241875\frac{3724}{1875}, 671750\frac{671}{750}, 1847506993\frac{184750}{6993}Exponents and constant of Input 12.4Theorem 12.5–
α\alpha1.09<12111.09<\frac{12}{11}Exponent in the bounds λ′,λ2≥αlog⁡(1/λ1)\lambda',\lambda_{2}\ge\alpha\log(1/\lambda_{1}) on [uexc,0.08][u_{\mathrm{exc}},0.08]Theorem 12.5(S1)
m∗m^{*}αlog⁡12.5=2.7530442…\alpha\log12.5=2.7530442\ldotsAnchor for the other characters on [uexc,0.08][u_{\mathrm{exc}},0.08]Theorem 12.5–
m2,m2′m_{2},m'_{2}2.8292.829, 4.9594.959Bounds λ2≥2.83−ε\lambda_{2}\ge2.83-\varepsilon, λ′≥4.96−ε\lambda'\ge4.96-\varepsilon with ε=0.001\varepsilon=0.001, on [0.08,0.1][0.08,0.1]Theorem 12.5(S1)
B1B_{1}K2+K4K^{2}+\frac{K}{4}Bound for B1/4B_{1}/4 for the first characterTheorem 12.5–
Large branchs=s1=32s=s_{1}=\frac{3}{2}, γf=1\gamma_{f}=1, γg=32\gamma_{g}=\frac{3}{2}; t0=1t_{0}=1One near row; bins of width 1100\frac{1}{100} on [1.5,3)[1.5,3) and a tailTheorem 12.6(S1)

Table 15. Fixed constants and parameters. The last column gives the stage of the order of choices of Section 13.1 in which the constant is fixed; a dash marks a quantity computed from earlier ones. Apart from the tolerances ε=ηH0\varepsilon=\eta H_0 and ε=ηJ2\varepsilon=\eta J^2, the constants of stages (S2)–(S6) are not computed; they exist by the arguments cited. Decimals are exact rational numbers unless they end with an ellipsis.

  1. Detection. With ε=ηH0\varepsilon= \eta H_{0} in Input 4.2, the weighted prime sum is at least LH0φ(q)(1−W−η)\frac{\mathcal{L}H_{0}}{\varphi(q)}(1-W-\eta); so W≤1−2ηW \le1-2\eta suffices (Proposition 4.4).

  2. Analytic error. In (8) the error EE has two parts (Section 5.8). The envelope errors, which include the smoothing error by (7), total less than 19ε119\varepsilon_{1} for the ordinary characters, since ∑χe−Aλχ≤Kfar<19\sum_{\chi}e^{-A\lambda_{\chi}} \le K_{\mathrm{far}} < 19 by (10), and at most 4ε14\varepsilon_{1} for the at most two first-family and two reserved characters; together less than 23ε1=η23\varepsilon_{1}=\eta. The outside first-family bound (Proposition 5.9) adds at most 4cout/(H0CB)≤η4c_{\mathrm{out}}/(H_{0}C_{B}) \le\eta, by the choice of CBC_{B} in (S4). Hence E≤2ηE \le2\eta.

  3. Allowance. Every leaf value includes final/S=5η\mathrm{final}/S=5\eta (Section 11.5); so does the program of Theorem 12.6.

  4. Conclusion. By the proof of Proposition 11.6, W≤V(x)−final/S+E≤V(x)−3ηW \le\mathcal{V}(x)-\mathrm{final}/S+E \le\mathcal{V}(x)-3\eta, that is, W+3η≤V(x)<1W+3\eta\le\mathcal{V}(x)<1 by Theorem 10.7. Hence W<1−3ηW<1-3\eta, and Proposition 4.4 applies with a margin η\eta.

  5. Other tolerances. They enter the constraints rather than EE: ε=ηJ22\varepsilon=\eta J_{2}^{2} in Input 6.1 gives the far budget (1+η)V(1+\eta)V of (9), the tolerance η′\eta' is built into every feature, diagonal and correlation term of a near row (Corollary 8.10), and every stored number is rounded in the safe direction (Section 11.5).

Published inputs

This section records the analytic estimates used in the proof. Statements labelled Input are taken from the literature, with notation adapted to this paper. Our main sources are Heath-Brown [51] and Xylouris’s dissertation [145]; the dissertation, rather than Xylouris’s shorter article, is the source of the lemma and equation numbers cited below. Any additional hypothesis needed from a source proof, or extension of a printed statement, is identified explicitly.

Test functions

Definition 3.1 (Conditions 1 and 2). Let x0>0x_{0}>0 and let f:[0,∞)→Rf:[0,\infty)\to\mathbb{R} be continuous with f(t)=0f(t)=0 for t≥x0t\ge x_{0}.

  • ff satisfies Condition 1 if it is twice continuously differentiable on (0,x0)(0,x_{0}) with ∣f′′∣≤B|f''|\le B there for some constant BB [51, 145].

  • ff satisfies Condition 2 if f≥0f\ge0 and its Laplace transform F(z)=∫0∞f(t)e−zt dtF(z)=\int_{0}^{\infty}f(t)e^{-zt}\,\mathrm{d}t satisfies Re⁡F(z)≥0\operatorname{Re}F(z)\ge0 for Re⁡z≥0\operatorname{Re}z\ge0 [51, 145].

Condition 1 supplies the regularity needed for the explicit formula. Condition 2 ensures that zero terms with nonnegative real argument can be discarded in an upper bound. The two conditions have different roles, so we specify each one when it is needed.

A function satisfying Condition 1 is bounded, its derivative is bounded on (0,x0)(0,x_{0}), and its transform FF is entire. The conductor coefficient of a nonprincipal character χ\chi modulo qq is [145]

φ(χ)={14,q cube-free, or ord⁡χ≤L,13,otherwise.\varphi(\chi)= \begin{cases} \frac{1}{4}, & q\ \text{cube-free, or }\operatorname{ord}\chi\le\mathcal{L},\\ \frac{1}{3}, & \text{otherwise}. \end{cases}

In particular φ(χ)=14\varphi(\chi)=\frac{1}{4} for every real character when q≥8q\ge8 (its order is 2≤L2\le\mathcal{L}), and φ(χ)≤13\varphi(\chi)\le\frac{1}{3} always.

Input 3.2 (Comparison lemma [51, 145]). Let F1,F2F_{1},F_{2} be holomorphic in {Re⁡z≥0}\{\operatorname{Re}z\ge0\}, with Re⁡F1(z)≥∣F2(z)∣\operatorname{Re}F_{1}(z)\ge|F_{2}(z)| on Re⁡z=0\operatorname{Re}z=0, and suppose that F1F_{1} and F2F_{2} tend to 0 uniformly as ∣z∣→∞|z|\to\infty in Re⁡z≥0\operatorname{Re}z\ge0. Then Re⁡F1(z)≥∣F2(z)∣\operatorname{Re}F_{1}(z)\ge|F_{2}(z)| for all Re⁡z≥0\operatorname{Re}z\ge0.

This is a consequence of the maximum principle [101, 130]. All transforms below are Laplace transforms of bounded functions of compact support [142].

Explicit formulas

The following two lemmas are Heath-Brown’s Lemmas 5.3 and 5.2, as stated by Xylouris [145]. In them s=σ+its=\sigma+it with

∣σ−1∣≤(log⁡L)1/2L,∣t∣≤L.(3)|\sigma-1|\le\frac{(\log\mathcal{L})^{1/2}}{\mathcal{L}},\qquad|t|\le\mathcal{L}. \tag*{(3)}

Input 3.3 (Principal character). Let ff satisfy Condition 1 and let ss satisfy (3). Then

∑n=1∞Λ(n)χ0(n)nsf(log⁡nL)=LF((s−1)L)+O(Llog⁡L),\sum_{n=1}^{\infty}\Lambda(n)\frac{\chi_{0}(n)}{n^{s}}f\left(\frac{\log n}{\mathcal{L}}\right) =\mathcal{L}F\left((s-1)\mathcal{L}\right)+O\left(\frac{\mathcal{L}}{\log\mathcal{L}}\right),

with an implied constant depending only on ff.

Input 3.4 (Local explicit formula). Let χ≠χ0\chi\ne\chi_{0}, let s=σ+its=\sigma+it satisfy (3), and let ff satisfy Condition 1 with f(0)≥0f(0)\ge0. For every ε>0\varepsilon>0 there are δ=δ(f,ε)∈(0,1)\delta=\delta(f,\varepsilon)\in(0,1) and q0=q0(f,ε)q_{0}=q_{0}(f,\varepsilon), independent of χ\chi and ss, such that for q≥q0q\ge q_{0}

∑n=1∞Λ(n)Re⁡χ(n)nsf(log⁡nL)≤−L∑∣1+it−ρ∣≤δRe⁡F((s−ρ)L)+f(0)2φ(χ)L+εL,\sum_{n=1}^{\infty}\Lambda(n)\operatorname{Re}\frac{\chi(n)}{n^{s}}f\left(\frac{\log n}{\mathcal{L}}\right) \le-\mathcal{L}\sum_{|1+it-\rho|\le\delta}\operatorname{Re}F\left((s-\rho)\mathcal{L}\right) +\frac{f(0)}{2}\varphi(\chi)\mathcal{L}+\varepsilon\mathcal{L},

where the sum runs over the nontrivial zeros of L(s,χ)L(s,\chi) in the disc, with multiplicity.

These are smoothed forms of the explicit formula [141, 26, 87, 62]; the conductor coefficient φ(χ)\varphi(\chi) comes from Burgess’s bounds, through Heath-Brown’s growth estimate for L(s,χ)L(s,\chi) and a Jensen-type formula [51]. Heath-Brown’s printed Lemma 5.2 omits the factor Λ(n)\Lambda(n) on the left, a misprint that Xylouris’s statement corrects. Xylouris notes [145] that the radius δ\delta and the threshold q0q_{0} of Input 3.4, and the implied constant of Input 3.3, depend only on ε\varepsilon and on upper bounds for x0x_0, x0−1x_0^{-1}, sup⁡∣f∣\sup|f| and sup⁡(0,x0)(∣f′∣+∣f′′∣)\sup_{(0,x_0)}(|f'|+|f''|); so both inputs hold uniformly for a family of test functions with such common bounds. In the proof of Heath-Brown’s Lemma 3.1, on which Lemma 5.2 rests, the disc radius is δ=min⁡(1/(2k),ε0/(3c0k2))\delta=\min(1/(2k),\varepsilon_0/(3c_0k^2)) with k≥3k\geq3, so δ≤1/6\delta\leq1/6. The same proof works for every smaller radius, and Heath-Brown remarks after Lemma 5.2 that δ=1/log⁡L\delta=1/\log\mathcal{L} may be taken. We may and do assume δ<1/2\delta<1/2.

Input 3.5 (The height ll [51, 145]). There are q0q_0 and a number l=l(q)∈Nl=l(q)\in\mathbb{N} with l≤L/10l\leq\mathcal{L}/10 such that for q≥q0q\geq q_0 the function ∏χL(s,χ)\prod_\chi L(s,\chi) has no zeros in R(10l)∖R(l)R(10l)\setminus R(l).

The principal LL-function has no zeros in R(l)R(l) for large qq [145]. All results we quote from the tables of [51, 145] concern zeros in R(l)R(l) for this ll; Heath-Brown’s proof of Lemma 6.1 produces ll and the later sections use it only through 1≤l≤L/101\leq l\leq\mathcal{L}/10 and the zero-free annulus.

Input 3.6 (All-zero envelope [145]). Let C0C_0 be the size of the square in Input 4.2, and fix ε,λ11>0\varepsilon,\lambda_{11}>0 and an integer m≥1m\geq1. For λ∈(λ11,C0]\lambda\in(\lambda_{11},C_0], suppose that:

(i) each f1iλf_{1i}^{\lambda}, 1≤i≤m1\leq i\leq m, satisfies Condition 1 with support parameter x0i>0x_{0i}>0 and f1iλ(0)≥0f_{1i}^{\lambda}(0)\geq0;

(ii) the quantities x0ix_{0i}, x0i−1x_{0i}^{-1}, sup⁡∣f1iλ∣\sup|f_{1i}^{\lambda}|, and sup⁡(0,x0i)(∣(f1iλ)′∣+∣(f1iλ)′′∣)\sup_{(0,x_{0i})}(|(f_{1i}^{\lambda})'|+|(f_{1i}^{\lambda})''|) have common bounds depending only on λ11\lambda_{11} and C0C_0, and f1λ=∑if1iλ≥0f_1^\lambda=\sum_i f_{1i}^\lambda\geq0;

(iii) H2H_2 is holomorphic in Re⁡z≥0\operatorname{Re}z\geq0, ∣H2(λ+it)∣≤Re⁡F1λ(it)|H_2(\lambda+it)|\leq\operatorname{Re}F_1^\lambda(it) for every real tt, and both transforms tend uniformly to 00 as ∣z∣→∞|z|\to\infty in that half-plane;

(iv) χ≠χ0\chi\neq\chi_0 and L(s,χ)L(s,\chi) has no zero with 1−λ/L<β≤11-\lambda/\mathcal{L}<\beta\leq1 and ∣γ∣≤1|\gamma|\leq1.

Here F1λF_1^\lambda is the Laplace transform of f1λf_1^\lambda. Then there is an effectively computable q0q_0, depending on ε,m,λ11,C0\varepsilon,m,\lambda_{11},C_0 but not on λ\lambda, such that for q≥q0q\geq q_0

∑ρ′∣H2((1−ρ)L)∣≤F1λ(−λ)+f1λ(0)6+ε,\sum_{\rho}'\left|H_2\left((1-\rho)\mathcal{L}\right)\right| \leq F_1^\lambda(-\lambda)+\frac{f_1^\lambda(0)}{6}+\varepsilon,

the sum running over the zeros of L(s,χ)L(s,\chi) in the square of Input 4.2.

We have included the hypothesis f1iλ(0)≥0f_{1i}^{\lambda}(0)\geq0 in (i), although it is omitted from the printed statement. The proof [145] applies Input 3.2, then Input 3.4 at s=1−λ/Ls=1-\lambda/\mathcal{L} to each f1iλf_{1i}^{\lambda}, then Input 3.3. It therefore uses f1iλ(0)≥0f_{1i}^{\lambda}(0)\geq0 for each ii; every application below has this property. The zero-free hypothesis is used only to know that the zeros of L(s,χ)L(s,\chi) in the disc of Input 3.4 about 11 have β≤1−λ/L\beta\leq1-\lambda/\mathcal{L}. We use the lemma in this localized form (Section 5).

Character sums and the Selberg sieve.

Input 3.7 (Burgess [51], k=3k=3). Let q≥1q\geq1 and let χ\chi be a primitive character modulo qq. Let N≥1N\geq1 and 1≤H≤q1\leq H\leq q. Then for every ε>0\varepsilon>0

∑N<n≤N+Hχ(n)≪εq1/9+εH2/3.\sum_{N<n\leq N+H}\chi(n)\ll_{\varepsilon}q^{1/9+\varepsilon}H^{2/3}.

This is the case k=3k=3 of Heath-Brown’s statement, which is Burgess’s theorem [12, 11, 13, 14]; see [52] for an account. It improves on the Pólya–Vinogradov inequality [108, 136] for short intervals. For special moduli stronger bounds are known [7, 16, 6, 70, 100], but we do not use them.

The weights ξd\xi_d in the next input are Selberg’s sieve weights [116, 117] in the form used for zero-density estimates by Graham and Heath-Brown; see [45, 93, 34] for the Selberg sieve in general.

Input 3.8 (Graham [40], p. 84; [51], (11.13), U=1U=1). Let z≥2z \ge2 and ξd=μ(d)log⁡(z/d)/log⁡z\xi_d=\mu(d)\log(z/d)/\log z for d≤zd \le z, ξd=0\xi_d=0 otherwise. Then for N≥zN \ge z

∑n≤N(∑d∣nξd)2=Nlog⁡z+O(Nlog⁡2z).\sum_{n \le N}\left(\sum_{d \mid n}\xi_d\right)^2=\frac{N}{\log z}+O\left(\frac{N}{\log^2 z}\right).

Log-free zero density.

Input 3.9 (Log-free density [66], Theorem 1, as quoted in the proof of [51], Lemma 6.1). For T≥1T \ge1, 54≤σ≤1\frac{5}{4} \le\sigma\le1 and every ε>0\varepsilon>0,

∑χ mod qN(σ,T,χ)≪ε(qT)(2+ε)(1−σ),\sum_{\chi\bmod q}N(\sigma,T,\chi)\ll_{\varepsilon}(qT)^{(2+\varepsilon)(1-\sigma)},

where N(σ,T,χ)N(\sigma,T,\chi) counts the nontrivial zeros β+iγ\beta+i\gamma of L(s,χ)L(s,\chi), with multiplicity, in β≥σ\beta\ge\sigma, ∣γ∣≤T|\gamma| \le T.

Estimates of this kind go back to Linnik [77, 78]; see [133, 32, 37, 84, 66, 91, 10] and [51], (1.4), [145], Prinzip 3, p. 11; explicit versions are in [127], and related estimates in [58, 56, 65, 49, 120, 105]. We use it only to see that the number of zeros of all L(s,χ)L(s,\chi) modulo qq, counted with multiplicity, with λ≤3\lambda\le3 and ∣γ∣≤T|\gamma| \le T is bounded independently of qq, for T=2T=2 and for T=LT=\mathcal{L}: take 1−σ=3/L1-\sigma=3/\mathcal{L}; then (qT)(2+ε)(1−σ)≪(qL)7/L≪1(qT)^{(2+\varepsilon)(1-\sigma)}\ll(q\mathcal{L})^{7/\mathcal{L}}\ll1.

Exceptional zeros.

Input 3.10 (Siegel zeros [50], Corollary 1, p. 406). Suppose L(β0,χ)=0L(\beta_0,\chi)=0 for a real character χ\chi modulo qq, not necessarily primitive, and a real β0≥1−1/(3log⁡q)\beta_0 \ge1-1/(3\log q), and put η0=((1−β0)log⁡q)−1\eta_0=((1-\beta_0)\log q)^{-1}. For every δ>0\delta>0 there is an effectively computable constant ηHB(δ)\eta_{\mathrm{HB}}(\delta) such that η0≥ηHB(δ)\eta_0 \ge\eta_{\mathrm{HB}}(\delta) implies P(a,q)≤q3+δP(a,q)\le q^{3+\delta} for every aa coprime to qq.

The zero-location results of Heath-Brown and Xylouris that we use are stated in Section 7, where they are applied, and the inputs for the small first-zero regime in Section 12.

Prime detection

Our goal is to make a nonnegative weighted sum over primes in a reduced residue class strictly positive. The explicit formula expresses this as a comparison between a main term and a sum over zeros. We first fix the weight, then state the precise zero-sum bound that will suffice.

The kernel.

Fix

T=0.416829,κ=T/16,T=0.416829,\qquad\kappa=T/16,

and the sixteen step heights b0,…,b15b_0,\ldots,b_{15} of Table 3. Put

iibib_iiibib_iiibib_iiibib_i
00.0231874140.2323715580.47877833120.75172165
10.0715933750.2908158390.54505502130.82236428
20.1226508560.35147264100.61279917140.89340300
30.1762770470.41418551110.68177389150.96451862

Table 3. The step heights of the kernel ψ\psi. They are exact decimal numbers.

ψ(t)=∑i=015bi1[iκ,(i+1)κ)(t),Ψ(z)=∫0Tψ(t)e−zt dt.\psi(t)=\sum_{i=0}^{15}b_i\mathbf{1}_{[i\kappa,(i+1)\kappa)}(t),\qquad \Psi(z)=\int_0^T\psi(t)e^{-zt}\,\mathrm{d}t.

For the target exponent LL put A=L−2TA=L-2T; at L=3.99L=3.99 we have A=3.156342A=3.156342. The prime weight is

h(t)=(ψ∗ψ)(t−A),h(t)=(\psi*\psi)(t-A),

which is supported in [A,L][A,L]. Its Laplace transform is

H(z)=∫h(t)e−zt dt=e−AzΨ(z)2,H0=H(0)=Ψ(0)2.(4)H(z)=\int h(t)e^{-zt}\,\mathrm{d}t=e^{-Az}\Psi(z)^2,\qquad H_0=H(0)=\Psi(0)^2. \tag*{(4)}

Numerically Ψ(0)=κ∑ibi=0.193642…\Psi(0)=\kappa\sum_i b_i=0.193642\ldots and H0=0.037497…H_0=0.037497\ldots.

Xylouris [145], following Heath-Brown [51], uses the triangular weights

hL′,K(t)={t−(L′−2K),L′−2K≤t≤L′−K,L′−t,L′−K≤t≤L′,0,otherwise,h_{L',K}(t)= \begin{cases} t-(L'-2K), & L'-2K\le t\le L'-K,\\ L'-t, & L'-K\le t\le L',\\ 0, & \text{otherwise}, \end{cases}

for L′>2K>0L'>2K>0; the transform of hL′,Kh_{L',K} is e−(L′−2K)z((1−e−Kz)/z)2e^{-(L'-2K)z}\left((1-e^{-Kz})/z\right)^2.

Lemma 4.1 (Decomposition into positive triangular weights). For the step function ψ\psi and shifted convolution hh defined above,

h=∑k=030ckhA+(k+2)κ,κ,ck=∑0≤i,j≤15i+j=kbibj>0.h=\sum_{k=0}^{30}c_k h_{A+(k+2)\kappa,\kappa}, \qquad c_k=\sum_{\substack{0\le i,j\le15\\i+j=k}}b_i b_j>0.

Every triangle in this sum has left endpoint A+kκ≥AA+k\kappa\ge A.

Proof. For i,j≥0i,j\ge0 the convolution 1[iκ,(i+1)κ)∗1[jκ,(j+1)κ)\mathbf{1}_{[i\kappa,(i+1)\kappa)}*\mathbf{1}_{[j\kappa,(j+1)\kappa)} is the tent function of height κ\kappa supported on [(i+j)κ,(i+j+2)κ][(i+j)\kappa,(i+j+2)\kappa], which is h(i+j+2)κ,κh_{(i+j+2)\kappa,\kappa}. Expanding ψ∗ψ\psi*\psi and shifting by AA gives the formula. Each ckc_k is a sum of products of positive numbers, and the sum over i+j=ki+j=k is nonempty for 0≤k≤300\le k\le30. □\square

The positivity criterion

The following is Xylouris’s form of the explicit formula for a finite combination of triangles [145]; it refines Heath-Brown’s Lemmas 13.1 and 13.2 [51].

Input 4.2 (Xylouris’s criterion). Let h=∑iαihLi,Kih=\sum_i\alpha_i h_{L_i,K_i} be a finite real combination of triangles with Li>2Ki+3L_i>2K_i+3 for every ii, and let HH be its Laplace transform. For every ε>0\varepsilon>0 there are C0=C0(ε)C_0=C_0(\varepsilon) and q0=q0(ε)q_0=q_0(\varepsilon) such that for every q≥q0q\ge q_0 and every aa coprime to qq,

∑p≡a (q)log⁡pph(log⁡pL)≥Lφ(q)[H(0)−∑χ≠χ0∑ρ′∣H((1−ρ)L)∣−ε],\sum_{\substack{p\equiv a\ (q)}}\frac{\log p}{p} h\left(\frac{\log p}{\mathcal{L}}\right) \ge \frac{\mathcal{L}}{\varphi(q)} \left[ H(0)-\sum_{\chi\ne\chi_0}\sum_{\rho}^{\prime} \left|H\left((1-\rho)\mathcal{L}\right)\right|-\varepsilon \right],

where ∑ρ′\sum_{\rho}^{\prime} runs over the zeros ρ=β+iγ\rho=\beta+i\gamma of L(s,χ)L(s,\chi), with multiplicity, in the square 1−C0/L≤β≤11-C_0/\mathcal{L}\le\beta\le1, ∣γ∣≤C0/L|\gamma|\le C_0/\mathcal{L}.

Xylouris states the extension to combinations without a separate proof, and it follows from Heath-Brown’s proof of his Lemma 13.2 [51]. That proof starts from the explicit formula [51], which is linear in the weight, and bounds the zeros outside the square by summing over rectangles 1−m+1L≤β≤1−mL1-\frac{m+1}{\mathcal{L}}\le\beta\le1-\frac{m}{\mathcal{L}}, n2L≤∣γ∣≤nL\frac{n}{2\mathcal{L}}\le|\gamma|\le\frac{n}{\mathcal{L}} with max⁡(m,n)≥R\max(m,n)\ge R, using a zero-density bound ≪e3m(1+n/L)3/2\ll e^{3m}(1+n/\mathcal{L})^{3/2} for each rectangle and ∣F((1−ρ)L)∣≪e−(L′−2K)mn−2|F((1-\rho)\mathcal{L})|\ll e^{-(L'-2K)m}n^{-2} for each zero. For a combination with positive coefficients ckc_k one has ∣H∣≤∑kck∣Hk∣|H|\le\sum_k c_k|H_k|, so the same estimate holds with L′−2KL'-2K replaced by the smallest left endpoint, which exceeds 3; and the zeros in the square contribute at most ∑∣H((1−ρ)L)∣\sum|H((1-\rho)\mathcal{L})|.

In the normalized coordinates (2) the square is {0≤λ≤C0, ∣μ∣≤C0}\{0 \le\lambda\le C_{0},\, |\mu| \le C_{0}\}. Enlarging it only adds nonnegative terms to the zero sum, so the conclusion holds for every square {λ≤CP, ∣μ∣≤CP}\{\lambda\le C_{P},\, |\mu| \le C_{P}\} with CP≥C0C_{P} \ge C_{0}. We fix

CP≥max⁡(C0(ηH0),3),(5)C_{P} \ge\max(C_{0}(\eta H_{0}),3), \tag*{(5)}

which fixes the prime window RP={0≤λ≤CP, ∣μ∣≤CP}R_{P}=\{0\le\lambda\le C_{P},\,|\mu|\le C_{P}\} of Section 2. Since CPC_{P} is fixed and log⁡log⁡L→∞\log\log\mathcal{L}\to\infty, for large qq every zero in RPR_{P} lies in the rectangle R(l)R(l) of Section 2.

Definition 4.3 (Normalized zero sum). For the kernel HH in (4) and the prime window fixed by (5), put

W=W(q)=1H0∑χ≠χ0∑ρ∈RP∣H((1−ρ)L)∣,W=W(q)=\frac{1}{H_{0}}\sum_{\chi\ne\chi_{0}}\sum_{\rho\in R_{P}}\left|H((1-\rho)\mathcal{L})\right|,

where each zero is counted with its multiplicity.

Proposition 4.4 (A zero-sum bound that guarantees a prime). Let L>3+2TL>3+2T. There is q0q_{0} such that for every q≥q0q\ge q_{0} with W(q)≤1−2ηW(q)\le1-2\eta and every aa coprime to qq there is a prime pp with

p≡a(modq),qA<p<qL,A=L−2T.p\equiv a\pmod q,\qquad q^{A}<p<q^{L},\qquad A=L-2T.

Here W(q)W(q) is formed from the kernel and prime window for this fixed LL; the threshold q0q_{0} is independent of aa and of the zero configuration of qq.

Proof. By Lemma 4.1, hh is a finite positive combination of triangles hLi,κh_{L_{i},\kappa} with Li−2κ=A+kκ≥A>3L_{i}-2\kappa=A+k\kappa\ge A>3. So Input 4.2 applies with ε=ηH0\varepsilon=\eta H_{0}; it gives a q0q_{0} such that for every q≥q0q\ge q_{0} with W(q)≤1−2ηW(q)\le1-2\eta,

∑p≡a (q)log⁡pph(log⁡pL)≥LH0φ(q)(1−W−η)≥LH0ηφ(q)>0.\sum_{p\equiv a\ (q)}\frac{\log p}{p}h\left(\frac{\log p}{\mathcal{L}}\right) \ge\frac{\mathcal{L}H_{0}}{\varphi(q)}(1-W-\eta) \ge\frac{\mathcal{L}H_{0}\eta}{\varphi(q)}>0.

The sum is therefore nonempty. Since hh vanishes outside (A,L)(A,L), some prime p≡a(modq)p\equiv a\pmod q satisfies A<log⁡p/L<LA<\log p/\mathcal{L}<L. □\square

At L=3.99L=3.99 the condition L>3+2T=3.833658L>3+2T=3.833658 holds. Two features of Input 4.2 matter below.

(i) Occurrence form. The zero sum charges each zero occurrence separately, through the absolute value of its own term. Every bound in Section 5 is a bound for a sum of such individual terms. In particular we never need the cancellation between zeros that a grouped form of the explicit formula would record.

(ii) Size of a term. Since ∣Ψ(λ−iμ)∣≤Ψ(λ)|\Psi(\lambda-i\mu)|\le\Psi(\lambda) for λ≥0\lambda\ge0,

∣H((1−ρ)L)∣=e−Aλ∣Ψ(λ−iμ)∣2≤e−AλΨ(λ)2≤H0e−Aλ.(6)\left|H((1-\rho)\mathcal{L})\right| =e^{-A\lambda}|\Psi(\lambda-i\mu)|^{2} \le e^{-A\lambda}\Psi(\lambda)^{2} \le H_{0}e^{-A\lambda}. \tag*{(6)}

A zero at distance λ\lambda from σ=1\sigma=1 costs at most e−Aλe^{-A\lambda} times H0H_{0}. The difficulty is entirely with the number of zeros close to σ=1\sigma=1.

If no nonprincipal L(s,χ)L(s,\chi) has a zero in R(l)R(l), then for large qq it has none in RPR_{P}, and W=0W=0. This is the first exterior regime (Section 12).

Zero costs

The zero sum WW of Section 4 is a sum over characters. This section bounds the contribution of a single character, in terms of an anchor (Section 2.4): a number λ∗\lambda_{*} such that no zero of L(s,χ)L(s,\chi) near s=1s=1 has parameter below λ∗\lambda_{*}, at which the explicit formula is evaluated. For most characters the anchor is the parameter of a representative zero (Definition 6.4); the first family needs a finer treatment. The argument is Xylouris’s proof of his Lemma 3.10 [145], pp. 35–36, which we reproduce in the localized form we need.

The envelope functions

The function Gϕ(λ)\mathcal{G}_{\phi}(\lambda) below bounds the total cost of one character. The auxiliary functions BϕB_{\phi}, h~\widetilde{h}, and CC separate the common exponential decay from the correction for a distinguished zero. The parameter ϕ\phi will be ϕ(χ)\phi(\chi), or an upper bound for it.

For real λ\lambda let ψλ(u)=ψ(u)e−λu\psi_{\lambda}(u)=\psi(u)e^{-\lambda u} and

fλ(t)=2∫0∞ψλ(u)ψλ(u+t) du=2∫0T−tψ(u)ψ(u+t)e−λ(2u+t) du(t≥0),f_{\lambda}(t)=2\int_{0}^{\infty}\psi_{\lambda}(u)\psi_{\lambda}(u+t)\,\mathrm{d}u =2\int_{0}^{T-t}\psi(u)\psi(u+t)e^{-\lambda(2u+t)}\,\mathrm{d}u \qquad(t\geq0),

so that fλ(t)=0f_{\lambda}(t)=0 for t≥Tt\geq T. Let FλF_{\lambda} be its Laplace transform. For 0≤ϕ≤130\leq\phi\leq\frac{1}{3}, λ≥0\lambda\geq0 and p≥a≥0p\geq a\geq0 put

Bϕ(λ)=1H0[2∫0Tψ(u)e−2λu∫uTψ(v) dv du+ϕ∫0Tψ(u)2e−2λu du],B_{\phi}(\lambda)=\frac{1}{H_{0}}\left[2\int_{0}^{T}\psi(u)e^{-2\lambda u}\int_{u}^{T}\psi(v)\,\mathrm{d}v\,\mathrm{d}u +\phi\int_{0}^{T}\psi(u)^{2}e^{-2\lambda u}\,\mathrm{d}u\right],
Gϕ(λ)=e−AλBϕ(λ),G=G1/3,h~(λ)=Ψ(λ)2H0,\mathcal{G}_{\phi}(\lambda)=e^{-A\lambda}B_{\phi}(\lambda), \qquad \mathcal{G}=\mathcal{G}_{1/3}, \qquad \widetilde{h}(\lambda)=\frac{\Psi(\lambda)^{2}}{H_{0}},
C(p,a)=2H0∫0Tψ(u)e−2pu∫uTψ(v)e−a(v−u) dv du.C(p,a)=\frac{2}{H_{0}}\int_{0}^{T}\psi(u)e^{-2pu}\int_{u}^{T}\psi(v)e^{-a(v-u)}\,\mathrm{d}v\,\mathrm{d}u.

Lemma 5.1 (Properties of the envelope functions). With the functions just defined, the following hold. Parts (i)–(iii) hold for every real λ\lambda; in (iv)–(v), take 0≤ϕ≤1/30\leq\phi\leq1/3, λ≥0\lambda\geq0, and p≥a≥0p\geq a\geq0.

(i) fλ≥0f_{\lambda}\geq0, fλf_{\lambda} is Lipschitz on [0,T][0,T], and fλ(T)=0f_{\lambda}(T)=0.

(ii) Re⁡Fλ(iy)=∣Ψ(λ+iy)∣2\operatorname{Re}F_{\lambda}(iy)=|\Psi(\lambda+iy)|^{2} for real yy.

(iii) Re⁡Fλ(z)≥∣Ψ(λ+z)∣2\operatorname{Re}F_{\lambda}(z)\geq|\Psi(\lambda+z)|^{2} for Re⁡z≥0\operatorname{Re}z\geq0. In particular Re⁡Fλ(z)≥0\operatorname{Re}F_{\lambda}(z)\geq0 for Re⁡z≥0\operatorname{Re}z\geq0.

(iv) H0Bϕ(λ)=Fλ(−λ)+ϕ2fλ(0)H_{0}B_{\phi}(\lambda)=F_{\lambda}(-\lambda)+\frac{\phi}{2}f_{\lambda}(0) and H0C(p,a)=Fp(a−p)H_{0}C(p,a)=F_{p}(a-p). Moreover C(a,a)=h~(a)C(a,a)=\widetilde{h}(a).

(v) BϕB_{\phi}, Gϕ\mathcal{G}_{\phi} and h~\widetilde{h} are positive and decreasing on [0,∞)[0,\infty), and so is Gϕ/w\mathcal{G}_{\phi}/w, where ww is either far weight of Section 6.

Proof. (i) Nonnegativity is clear. For 0≤t<t′≤T0\leq t<t'\leq T, ∣fλ(t)−fλ(t′)∣≤2∥ψλ∥∞∫∣ψλ(u+t)−ψλ(u+t′)∣ du|f_{\lambda}(t)-f_{\lambda}(t')|\leq2\lVert\psi_{\lambda}\rVert_{\infty}\int|\psi_{\lambda}(u+t)-\psi_{\lambda}(u+t')|\,\mathrm{d}u, which is O(t′−t)O(t'-t) because ψλ\psi_{\lambda} has bounded variation.

(ii) Let aλ(t)=∫ψλ(u)ψλ(u+∣t∣) dua_{\lambda}(t)=\int\psi_{\lambda}(u)\psi_{\lambda}(u+|t|)\,\mathrm{d}u for t∈Rt\in\mathbb{R}; it is even and integrable, and its Fourier transform is ∣ψ^λ(y)∣2=∣Ψ(λ+iy)∣2|\widehat{\psi}_{\lambda}(y)|^{2}=|\Psi(\lambda+iy)|^{2}. Hence Re⁡Fλ(iy)=∫0∞fλ(t)cos⁡(yt) dt=∫Raλ(t)e−iyt dt=∣Ψ(λ+iy)∣2\operatorname{Re}F_{\lambda}(iy)=\int_{0}^{\infty}f_{\lambda}(t)\cos(yt)\,\mathrm{d}t=\int_{\mathbb{R}}a_{\lambda}(t)e^{-iyt}\,\mathrm{d}t=|\Psi(\lambda+iy)|^{2}.

(iii) Apply Input 3.2 to F1=FλF_{1}=F_{\lambda} and F2(z)=Ψ(λ+z)2F_{2}(z)=\Psi(\lambda+z)^{2}. Both are entire. By (i), Fλ(z)=z−1(fλ(0)+∫0Tfλ′(t)e−zt dt)F_{\lambda}(z)=z^{-1}(f_{\lambda}(0)+\int_{0}^{T}f_{\lambda}'(t)e^{-zt}\,\mathrm{d}t), so ∣Fλ(z)∣≤(fλ(0)+∥fλ′∥1)/∣z∣|F_{\lambda}(z)|\leq(f_{\lambda}(0)+\lVert f_{\lambda}'\rVert_{1})/|z| for Re⁡z≥0\operatorname{Re}z\geq0; similarly ∣Ψ(λ+z)∣≤(ψλ(0+)+V(ψλ))/∣z∣|\Psi(\lambda+z)|\leq(\psi_{\lambda}(0+)+V(\psi_{\lambda}))/|z| with VV the total variation. So both tend to 0 uniformly, and on Re⁡z=0\operatorname{Re}z=0 we have equality by (ii).

(iv) Substitute v=u+tv=u+t: Fλ(−λ)=2∬ψ(u)ψ(v)e−λ(u+v)eλ(v−u) dv duF_{\lambda}(-\lambda)=2\iint\psi(u)\psi(v)e^{-\lambda(u+v)}e^{\lambda(v-u)}\,\mathrm{d}v\,\mathrm{d}u over 0≤u≤v≤T0\leq u\leq v\leq T, which is the first term of H0BϕH_{0}B_{\phi}; and fλ(0)=2∫ψ2e−2λuf_{\lambda}(0)=2\int\psi^{2}e^{-2\lambda u}. The same substitution gives Fp(a−p)=2∫0Tψ(u)e−2pu∫uTψ(v)e−a(v−u) dv duF_{p}(a-p)=2\int_{0}^{T}\psi(u)e^{-2pu}\int_{u}^{T}\psi(v)e^{-a(v-u)}\,\mathrm{d}v\,\mathrm{d}u. Finally H0C(a,a)=2∬u≤vψ(u)ψ(v)e−a(u+v)=Ψ(a)2H_{0}C(a,a)=2\iint_{u\leq v}\psi(u)\psi(v)e^{-a(u+v)}=\Psi(a)^{2}.

(v) All integrands are positive and decrease in λ\lambda. For the last claim,

Gϕ(λ)w(λ)=Bϕ(λ)∫u0xw0(t)2e(2t−A)λ dt,\frac{\mathcal{G}_{\phi}(\lambda)}{w(\lambda)} = B_{\phi}(\lambda)\int_{u_{0}}^{x}w_{0}(t)^{2}e^{(2t-A)\lambda}\,\mathrm{d}t,

and 2t−A<02t-A<0 on [u0,x][u_{0},x] (Lemma 6.2). □\square

The number G(λ)\mathcal{G}(\lambda) is the cost, in units of H0H_{0}, that the argument below assigns to a character whose zeros near s=1s=1 all have parameter at least λ\lambda. For a real character the factor ϕ=13\phi=\frac{1}{3} may be replaced by 14\frac{1}{4} (Proposition 5.5).

Smoothing

Since ψ\psi is a step function, fλf_{\lambda} has corners at the multiples of κ\kappa and does not satisfy Condition 1. We therefore work with smooth approximations of ψ\psi in the auxiliary functions only; the prime weight hh and its transform HH are never changed. For θ∈(0,1)\theta\in(0,1) fix a function ψθ∈C∞(R)\psi^{\theta} \in C^{\infty}(\mathbb{R}) with support in (0,T)(0,T) such that

0≤ψθ≤1,∥ψ−ψθ∥L1≤θ,0 \le\psi^{\theta} \le1,\qquad\lVert\psi-\psi^{\theta}\rVert_{L^{1}} \le\theta,

for instance a mollification of ψ1[θ/4,T−θ/4]\psi\mathbf{1}_{[\theta/4,T-\theta/4]}. Let Ψθ\Psi^{\theta}, fλθf_{\lambda}^{\theta}, FλθF_{\lambda}^{\theta}, BϕθB_{\phi}^{\theta}, h~θ\tilde{h}^{\theta} and CθC^{\theta} be defined as above with ψθ\psi^{\theta} in place of ψ\psi, and H0H_{0} kept.

Lemma 5.2 (Uniform smooth approximation). Fix θ∈(0,1)\theta\in(0,1) and Λ≥1\Lambda\ge1, and choose ψθ\psi^{\theta} as above.

(i) For every λ∈[0,Λ]\lambda\in[0,\Lambda], fλθf_{\lambda}^{\theta} satisfies Condition 1 and Condition 2 with x0=Tx_{0}=T. The quantities x0x_{0}, x0−1x_{0}^{-1}, sup⁡∣fλθ∣\sup|f_{\lambda}^{\theta}| and sup⁡(0,T)(∣(fλθ)′∣+∣(fλθ)′′∣)\sup_{(0,T)}(|(f_{\lambda}^{\theta})'|+|(f_{\lambda}^{\theta})''|) are bounded independently of λ∈[0,Λ]\lambda\in[0,\Lambda]; the bounds may depend on θ\theta and Λ\Lambda. Lemma 5.1(i)–(iv) hold for ψθ\psi^{\theta}.

(ii) For Re⁡z≥0\operatorname{Re} z \ge0, ∣Ψ(z)2−Ψθ(z)2∣≤3θ|\Psi(z)^{2}-\Psi^{\theta}(z)^{2}| \le3\theta.

(iii) As θ→0\theta\to0, Bϕθ→BϕB_{\phi}^{\theta} \to B_{\phi}, h~θ→h~\tilde{h}^{\theta} \to\tilde{h} and Cθ→CC^{\theta} \to C uniformly on {0≤a≤p≤Λ}\{0 \le a \le p \le\Lambda\} and 0≤ϕ≤130 \le\phi\le\frac{1}{3}.

Proof. (i) The function (t,λ)↦fλθ(t)(t,\lambda) \mapsto f_{\lambda}^{\theta}(t) is smooth on R2\mathbb{R}^{2}, and it vanishes for t≥Tt \ge T because ψθ\psi^{\theta} is supported in a compact subset of (0,T)(0,T); this gives Condition 1 and the uniform bounds. The proof of Lemma 5.1 applies verbatim to ψθ\psi^{\theta}, and Condition 2 is Lemma 5.1(i) and (iii) for ψθ\psi^{\theta}. (ii) For Re⁡z≥0\operatorname{Re} z \ge0 we have ∣Ψ(z)∣≤∥ψ∥1≤T<12|\Psi(z)| \le\lVert\psi\rVert_{1} \le T < \frac{1}{2}, ∣Ψθ(z)∣≤∥ψ∥1+θ|\Psi^{\theta}(z)| \le\lVert\psi\rVert_{1}+\theta and ∣Ψ(z)−Ψθ(z)∣≤θ|\Psi(z)-\Psi^{\theta}(z)| \le\theta; hence ∣Ψ2−(Ψθ)2∣≤θ(2T+θ)≤3θ|\Psi^{2}-(\Psi^{\theta})^{2}| \le\theta(2T+\theta) \le3\theta. (iii) Each quantity is a bilinear expression in (ψ,ψ)(\psi,\psi) with a kernel bounded by 11 on [0,T]2[0,T]^{2} for the parameters in question, and ∥ψ∥∞,∥ψθ∥∞≤1\lVert\psi\rVert_{\infty},\lVert\psi^{\theta}\rVert_{\infty} \le1. □

The disc and a uniform zero count

Throughout this section λ11=0.05\lambda_{11}=0.05, the smoothing parameter θ\theta is the one fixed in stage (S3) of Section 13.1, and ε1=η/23\varepsilon_{1}=\eta/23. Let A⊂[λ11,CP]\mathcal{A}\subset[\lambda_{11},C_{P}] be the finite set consisting of the points of the grid λ11+hAZ\lambda_{11}+h_{\mathcal{A}}\mathbb{Z} in [λ11,CP][\lambda_{11},C_{P}] together with the anchors of the first-family bounds of all leaves (Sections 5.6 and 5.7). For a leaf with first-zero cell [a,b][a,b] and lower bound p∗p^{*} for λ′\lambda' (the lower end of its gap if it has one, and pp otherwise; Section 2.4), these are aa and p∗p^{*} if the leaf is inside and p∘=max⁡(a,p∗)p^{\circ}=\max(a,p^{*}) if it is outside (Section 11.5); the mesh hA>0h_{\mathcal{A}}>0 is chosen so that ∣Bϕ(x)−Bϕ(y)∣≤ε1/4|B_{\phi}(x)-B_{\phi}(y)|\le\varepsilon_{1}/4 for ϕ∈{14,13}\phi\in\{\frac{1}{4},\frac{1}{3}\} and x,y∈[0,CP]x,y\in[0,C_{P}] with ∣x−y∣≤hA|x-y|\le h_{\mathcal{A}}. In stage (S5) we fix a radius δ0∈(0,13]\delta_{0}\in(0,\frac{1}{3}] such that Input 3.4 holds with radius δ0\delta_{0} for each of the finitely many pairs of a test function and a tolerance used in this paper: the function f(1)f^{(1)} below with tolerance 11, the functions fσθf_{\sigma}^{\theta} (σ∈A)(\sigma\in\mathcal{A}) with tolerance H0ε1/8H_{0}\varepsilon_{1}/8, and the detectors (Definition 8.1) of the near rows with the tolerances of Lemma 8.7. This is possible because Input 3.4 remains true for every smaller radius (see the remark after it). The disc of a character χ\chi is the set of zeros ρ\rho of L(s,χ)L(s,\chi) with ∣1−ρ∣≤δ0|1-\rho|\le\delta_{0}. For fixed CC and large qq the square {λ≤C,∣μ∣≤C}\{\lambda\le C,|\mu|\le C\} lies in the disc, since there ∣1−ρ∣≤2C/L|1-\rho|\le\sqrt{2}C/\mathcal{L}.

Lemma 5.3 (A uniform bound for zeros in a fixed window). Let C≥1C\ge1. There are N=N(C)N=N(C) and q0q_{0} such that for q≥q0q\ge q_{0} the following holds. If χ≠χ0\chi\ne\chi_{0} and every zero in the disc of χ\chi has parameter at least λ11\lambda_{11}, then L(s,χ)L(s,\chi) has at most NN zeros, counted with multiplicity, with λ≤C\lambda\le C and ∣μ∣≤C|\mu|\le C.

Proof. Let K1=1/(4C)K_{1}=1/(4C) and ϑ(z)=(1−e−K1z)/z\vartheta(z)=(1-e^{-K_{1}z})/z, the transform of 1[0,K1]\mathbf{1}_{[0,K_{1}]}. Let f(1)(t)=2∫0K1−te−λ11(2u+t) duf^{(1)}(t)=2\int_{0}^{K_{1}-t}e^{-\lambda_{11}(2u+t)}\,\mathrm{d}u for 0≤t≤K10\le t\le K_{1}, and f(1)(t)=0f^{(1)}(t)=0 for t≥K1t\ge K_{1}. This is the function fλ11f_{\lambda_{11}} of Section 5.1 for the step 1[0,K1]\mathbf{1}_{[0,K_{1}]}; it equals (e−λ11t−e−λ11(2K1−t))/λ11(e^{-\lambda_{11}t}-e^{-\lambda_{11}(2K_{1}-t)})/\lambda_{11} on [0,K1][0,K_{1}], so it satisfies Condition 1, and f(1)(0)>0f^{(1)}(0)>0. As in Lemma 5.1(iii), Re⁡F(1)(z)≥∣ϑ(λ11+z)∣2≥0\operatorname{Re}F^{(1)}(z)\ge|\vartheta(\lambda_{11}+z)|^{2}\ge0 for Re⁡z≥0\operatorname{Re}z\ge0.

Apply Input 3.4 to f(1)f^{(1)} at s=1−λ11/Ls=1-\lambda_{11}/\mathcal{L}, with tolerance 11 and a radius at most δ0\delta_{0}. For large qq, the fixed normalized square lies in this smaller disc. All its zeros still have parameter at least λ11\lambda_{11}. In the remainder of this proof, “the disc” refers to this smaller disc. The point ss lies in the range (3.1) for large qq, and (s−ρ)L=(λρ−λ11)−iμρ(s-\rho)\mathcal{L}=(\lambda_{\rho}-\lambda_{11})-i\mu_{\rho}. Since −Re⁡χ(n)≤χ0(n)-\operatorname{Re}\chi(n)\leq\chi_{0}(n) for every nn and f(1)≥0f^{(1)}\geq0, Input 3.3 gives

−∑nΛ(n)Re⁡χ(n)nsf(1)(log⁡nL)≤∑nΛ(n)χ0(n)nsf(1)(log⁡nL)=LF(1)(−λ11)+o(L).-\sum_{n}\Lambda(n)\operatorname{Re}\frac{\chi(n)}{n^{s}}f^{(1)}\left(\frac{\log n}{\mathcal{L}}\right) \leq \sum_{n}\Lambda(n)\frac{\chi_{0}(n)}{n^{s}}f^{(1)}\left(\frac{\log n}{\mathcal{L}}\right) = \mathcal{L}F^{(1)}(-\lambda_{11})+o(\mathcal{L}).

Hence ∑ρ in the discRe⁡F(1)((λρ−λ11)−iμρ)≤F(1)(−λ11)+16f(1)(0)+2=:c1\sum_{\rho\ \mathrm{in\ the\ disc}}\operatorname{Re}F^{(1)}((\lambda_{\rho}-\lambda_{11})-i\mu_{\rho})\leq F^{(1)}(-\lambda_{11})+\frac{1}{6}f^{(1)}(0)+2=:c_{1} for large qq. Every term is nonnegative, because λρ≥λ11\lambda_{\rho}\geq\lambda_{11}. For a zero with λ≤C\lambda\leq C, ∣μ∣≤C|\mu|\leq C put z=λ−iμz=\lambda-i\mu; then ∣K1z∣≤2/4<12|K_{1}z|\leq\sqrt{2}/4<\frac{1}{2} and ∣1−e−w∣≥∣w∣−∣w∣2e∣w∣/2≥∣w∣/2|1-e^{-w}|\geq|w|-|w|^{2}e^{|w|}/2\geq|w|/2 for ∣w∣≤12|w|\leq\frac{1}{2}, so the term is at least ∣ϑ(z)∣2≥K12/4|\vartheta(z)|^{2}\geq K_{1}^{2}/4. The number of such zeros is therefore at most N=⌊4c1/K12⌋N=\lfloor4c_{1}/K_{1}^{2}\rfloor. □\square

The anchored envelope

The following inequality is the common source of all the costs below. For σ∈A\sigma\in\mathcal{A}, a character χ≠χ0\chi\neq\chi_{0} and a zero ρ\rho put

Φσθ(ρ)=Re⁡Fσθ((λρ−σ)−iμρ).\Phi_{\sigma}^{\theta}(\rho)=\operatorname{Re}F_{\sigma}^{\theta}\left((\lambda_{\rho}-\sigma)-i\mu_{\rho}\right).

By Lemma 5.1(iii), Φσθ(ρ)≥∣Ψθ(λρ−iμρ)∣2≥0\Phi_{\sigma}^{\theta}(\rho)\geq|\Psi^{\theta}(\lambda_{\rho}-i\mu_{\rho})|^{2}\geq0 whenever λρ≥σ\lambda_{\rho}\geq\sigma.

Lemma 5.4 (Anchored inequality). Fix the smoothing, finite anchor set A\mathcal{A}, and common disc radius as in Section 5.3. There is q0q_{0}, uniform over χ≠χ0\chi\neq\chi_{0} and σ∈A\sigma\in\mathcal{A}, such that for q≥q0q\geq q_{0},

∑ρ in the disc of χΦσθ(ρ)≤H0Bφ(χ)θ(σ)+14H0ε1.\sum_{\rho\ \mathrm{in\ the\ disc\ of}\ \chi}\Phi_{\sigma}^{\theta}(\rho) \leq H_{0}B_{\varphi(\chi)}^{\theta}(\sigma)+\frac{1}{4}H_{0}\varepsilon_{1}.

Proof. Apply Input 3.4 to f=fσθf=f_{\sigma}^{\theta} at s=1−σ/Ls=1-\sigma/\mathcal{L} (so t=0t=0), with tolerance H0ε1/8H_{0}\varepsilon_{1}/8 and radius δ0\delta_{0}, and bound the prime sum as in the proof of Lemma 5.3, using Input 3.3:

L∑ρΦσθ(ρ)≤fσθ(0)2φ(χ)L+LFσθ(−σ)+18H0ε1L+O(Llog⁡L).\mathcal{L}\sum_{\rho}\Phi_{\sigma}^{\theta}(\rho) \leq \frac{f_{\sigma}^{\theta}(0)}{2}\varphi(\chi)\mathcal{L} +\mathcal{L}F_{\sigma}^{\theta}(-\sigma) +\frac{1}{8}H_{0}\varepsilon_{1}\mathcal{L} +O\left(\frac{\mathcal{L}}{\log\mathcal{L}}\right).

By Lemma 5.1(iv) for ψθ\psi^{\theta} the right side is at most L(H0Bφ(χ)θ(σ)+14H0ε1)\mathcal{L}(H_{0}B_{\varphi(\chi)}^{\theta}(\sigma)+\frac{1}{4}H_{0}\varepsilon_{1}) for large qq. Only the finitely many functions fσθf_{\sigma}^{\theta}, σ∈A\sigma\in\mathcal{A}, occur, so one q0q_{0} serves them all. □\square

Proposition 5.5 (Localized envelope). For q≥q0q\geq q_{0} the following holds. Let χ≠χ0\chi\neq\chi_{0} and λ∗∈[λ11,CP]\lambda_{*}\in[\lambda_{11},C_{P}], and suppose that every zero in the disc of χ\chi has parameter at least λ∗\lambda_{*}. Then

1H0∑ρ∈RP∣H((1−ρ)L)∣≤Gφ(χ)(λ∗)+ε1e−Aλ∗,\frac{1}{H_{0}}\sum_{\rho\in R_{P}}\left|H((1-\rho)\mathcal{L})\right| \leq G_{\varphi(\chi)}(\lambda_{*})+\varepsilon_{1}e^{-A\lambda_{*}},

the sum running over the zeros of L(s,χ)L(s,\chi) in RPR_{P} with multiplicity. In particular the left side is at most G(λ∗)+ε1e−Aλ∗G(\lambda_{*})+\varepsilon_{1}e^{-A\lambda_{*}}, and at most G1/4(λ∗)+ε1e−Aλ∗G_{1/4}(\lambda_{*})+\varepsilon_{1}e^{-A\lambda_{*}} if χ\chi is real.

Proof. Let N=N(CP)N=N(C_{P}) be as in Lemma 5.3. The parameter θ\theta of stage (S3) satisfies

3(N+6)θ+H0sup⁡∣Bϕθ−Bϕ∣≤14H0ε1,(7)3(N+6)\theta+H_{0}\sup\left|B_{\phi}^{\theta}-B_{\phi}\right| \leq\frac{1}{4}H_{0}\varepsilon_{1}, \tag*{(7)}

the supremum over [0,CP][0,C_{P}] and ϕ∈{14,13}\phi\in\{\frac{1}{4},\frac{1}{3}\} (Lemma 5.2(iii)). Let σ∈A\sigma\in\mathcal{A} be the largest grid point with σ≤λ∗\sigma\leq\lambda_{*}, so that λ∗−σ≤hA\lambda_{*}-\sigma\leq h_{\mathcal{A}}. Let ρ∈RP\rho\in R_{P}. Then ρ\rho is in the disc, so λρ≥λ∗≥σ\lambda_{\rho}\geq\lambda_{*}\geq\sigma, and by (4.3), Lemma 5.2(ii) and Lemma 5.1(iii) for ψθ\psi^{\theta},

∣H((1−ρ)L)∣=e−Aλρ∣Ψ(λρ−iμρ)∣2≤e−Aλ∗(Φσθ(ρ)+3θ).\left|H((1-\rho)\mathcal{L})\right| = e^{-A\lambda_{\rho}}\left|\Psi(\lambda_{\rho}-i\mu_{\rho})\right|^{2} \leq e^{-A\lambda_{*}}\left(\Phi_{\sigma}^{\theta}(\rho)+3\theta\right).

The zeros of the disc that are not in RPR_{P} have Φσθ≥0\Phi_{\sigma}^{\theta}\geq0. Summing over the at most NN zeros in RPR_{P} and using Lemma 5.4 and (7),

∑ρ∈RP∣H((1−ρ)L)∣≤e−Aλ∗(H0Bφ(χ)θ(σ)+14H0ε1+3Nθ)≤e−Aλ∗H0(Bφ(χ)(σ)+12ε1),\sum_{\rho\in R_{P}}\left|H((1-\rho)\mathcal{L})\right| \leq e^{-A\lambda_{*}}\left(H_{0}B_{\varphi(\chi)}^{\theta}(\sigma)+\frac{1}{4}H_{0}\varepsilon_{1}+3N\theta\right) \leq e^{-A\lambda_{*}}H_{0}\left(B_{\varphi(\chi)}(\sigma)+\frac{1}{2}\varepsilon_{1}\right),

and Bφ(χ)(σ)≤Bφ(χ)(λ∗)+14ε1B_{\varphi(\chi)}(\sigma) \le B_{\varphi(\chi)}(\lambda_*) + \frac{1}{4}\varepsilon_1 by the choice of hAh_{\mathcal A}. Finally φ(χ)≤13\varphi(\chi) \le\frac{1}{3} always, φ(χ)=14\varphi(\chi) = \frac{1}{4} for real χ\chi (whose order is 2≤L2 \le\mathcal L), and BϕB_\phi increases with ϕ\phi. □

Remark 5.6. Proposition 5.5 is a localized form of Xylouris’s Lemma 3.10 [145], pp. 35–36, applied with m=1m=1, f1=fσθf_1=f_\sigma^\theta and H2=(Ψθ)2H_2=(\Psi^\theta)^2, for which his condition (3.59) holds with equality by Lemma 5.1(ii); the prime weight HH is then compared with (Ψθ)2(\Psi^\theta)^2 through Lemma 5.2(ii). We require only that the zeros in the disc have parameter at least λ∗\lambda_*, where Xylouris requires that L(s,χ)L(s,\chi) have no zero with 1−λ∗/L<β≤11-\lambda_*/\mathcal L<\beta\le1 and ∣γ∣≤1|\gamma|\le1; his proof uses this hypothesis only for the zeros in the disc, which have ∣γ∣≤δ0<1|\gamma|\le\delta_0<1. His printed statement allows f1f_1 to be a sum of mm pieces; for m>1m>1 his proof needs in addition that f1i(0)≥0f_{1i}(0)\ge0 for each piece, and that the sum over the square be enlarged to a sum over a disc common to all pieces before Input 3.4 is applied to them. With m=1m=1 neither point arises. Finally, we use finitely many anchors, so that no uniformity in the anchor is needed. The coefficient φ(χ)/2\varphi(\chi)/2 comes from Input 3.4; the printed Lemma 3.10 uses φ(χ)≤13\varphi(\chi)\le\frac{1}{3}.

Ordinary and reserved characters

Let χ\chi be a nonprincipal character (in the case tree, one outside the first family; in Theorem 12.6, any), and let λχ\lambda_\chi be the parameter of its T∗T^*-representative or of its height-one representative (Definition 6.4). If L(s,χ)L(s,\chi) has a zero in RPR_P, then that zero lies in R(l)R(l) and has ∣γ∣≤CP/L<T∗|\gamma|\le C_P/\mathcal L<T^*, so the representative exists and λχ≤CP\lambda_\chi\le C_P. In the regimes in which we use these costs λ1≥0.1\lambda_1\ge0.1, so λχ≥λ11\lambda_\chi\ge\lambda_{11}.

Corollary 5.7 (Cost of a character with a representative). Suppose λ1≥0.1\lambda_1\ge0.1. Let χ≠χ0\chi\ne\chi_0 have a zero in RPR_P, and let λχ\lambda_\chi be the parameter of either its T∗T^*-representative or its height-one representative. For sufficiently large qq, its contribution to WW is at most

G(λχ)+ε1e−Aλχ,ε1=η/23.G(\lambda_\chi)+\varepsilon_1 e^{-A\lambda_\chi},\qquad\varepsilon_1=\eta/23.

If χ\chi is real, GG may be replaced by G1/4G_{1/4}. The threshold is uniform over these choices of character and representative.

Proof. We check the hypothesis of Proposition 5.5 with λ∗=λχ\lambda_*=\lambda_\chi. A zero in the disc with parameter at most CPC_P lies in R(l)R(l) for large qq, since CP<13log⁡log⁡LC_P<\frac{1}{3}\log\log\mathcal L and ∣γ∣≤δ0<1≤l|\gamma|\le\delta_0<1\le l; it has ∣γ∣≤δ0<12≤T∗|\gamma|\le\delta_0<\frac{1}{2}\le T^*, so its parameter is at least λχ\lambda_\chi by the minimality of the representative. A zero in the disc with parameter above CPC_P has parameter above λχ\lambda_\chi. □

The first family inside the buffer

Suppose that ρ1∈RB\rho_1\in R_B, that λ1∈[a,b]\lambda_1\in[a,b] with a≥λ11a\ge\lambda_{11}, and that every zero occurrence of the first family in R(l)R(l) other than the distinguished ones has parameter at least pp, where p≤CPp\le C_P; equivalently, p≤λ′p\le\lambda'. We assume that a,p∈Aa,p\in\mathcal A (in the leaves they are the numbers aa and p∗p^* of Section 5.3). The distinguished occurrences and their number α\alpha per character are those of Definition 2.2 and the table after it. Let ϕ1=14\phi_1=\frac{1}{4} if χ1\chi_1 is real and ϕ1=13\phi_1=\frac{1}{3} otherwise, and put

Jold=n[e−Ap(Bϕ1(a)−αh~(a))++αe−Aah~(a)],J_{\mathrm{old}} = n\left[e^{-Ap}\left(B_{\phi_1}(a)-\alpha\widetilde h(a)\right)_{+} +\alpha e^{-Aa}\widetilde h(a)\right],
Jnew=n[αe−Aah~(a)+e−Ap(Bϕ1(p)−αC(p,a))](p≥b).J_{\mathrm{new}} = n\left[\alpha e^{-Aa}\widetilde h(a) +e^{-Ap}\left(B_{\phi_1}(p)-\alpha C(p,a)\right)\right] \qquad(p\ge b).

Proposition 5.8 (Cost of the first family inside the buffer). Suppose ρ1∈RB\rho_1\in R_B, λ1∈[a,b]\lambda_1\in[a,b], a,p∈Aa,p\in\mathcal A, a≥λ11a\ge\lambda_{11}, and p≤min⁡(λ′,CP)p\le\min(\lambda',C_P). For sufficiently large qq, the contribution of all nn characters of the first family to WW is at most Jold+nε1J_{\mathrm{old}}+n\varepsilon_1. If also p≥bp\ge b, it is at most Jnew+nε1J_{\mathrm{new}}+n\varepsilon_1. Here JoldJ_{\mathrm{old}} and JnewJ_{\mathrm{new}} are the displayed functions of a,p,n,αa,p,n,\alpha, and ϕ1\phi_1.

The two bounds use different anchors. The bound JoldJ_{\mathrm{old}} uses the lower endpoint aa of the first-zero cell. The bound JnewJ_{\mathrm{new}} uses pp, the lower bound for the remaining zeros, and estimates the distinguished zero together with its correction in the explicit formula. Their combined expression is smaller than a separate estimate of the two terms.

Proof. It suffices to treat one first-family character χ\chi; the bound for its conjugate is the same, because RPR_P and the disc are symmetric under conjugation. A zero of the disc of χ\chi lies in R(l)R(l) or has parameter above 13log⁡log⁡L>CP\frac{1}{3}\log\log\mathcal{L}>C_P. So every zero of the disc has parameter at least λ1≥a\lambda_1\geq a, and every non-distinguished one has parameter at least pp. Write SS for the distinguished occurrences of χ\chi; they lie in RBR_B, hence in the disc.

The bound JoldJ_{\mathrm{old}}. Every non-distinguished occurrence has parameter at least max⁡(p,λ1)≥max⁡(p,a)\max(p,\lambda_1)\geq\max(p,a), and JoldJ_{\mathrm{old}} decreases as pp increases; so we may and do assume p≥ap\geq a. For a non-distinguished ρ∈RP\rho\in R_P we then have λρ≥p≥a\lambda_\rho\geq p\geq a, and ∣H((1−ρ)L)∣≤e−Ap∣Ψ(λρ−iμρ)∣2≤e−Ap(Φaθ(ρ)+3θ)|H((1-\rho)\mathcal{L})|\leq e^{-Ap}|\Psi(\lambda_\rho-i\mu_\rho)|^2\leq e^{-Ap}(\Phi_a^\theta(\rho)+3\theta). By Lemma 5.4 at σ=a\sigma=a, since every term Φaθ\Phi_a^\theta in the disc is nonnegative and there are at most NN zeros in RPR_P,

∑ρ∈RP∖S∣H((1−ρ)L)∣≤e−Ap(H0Bϕ1θ(a)+14H0ε1+3Nθ−∑ρ∈SΦaθ(ρ)).\sum_{\rho\in R_P\setminus S}|H((1-\rho)\mathcal{L})| \leq e^{-Ap}\left(H_0B_{\phi_1}^\theta(a)+\frac{1}{4}H_0\varepsilon_1+3N\theta-\sum_{\rho\in S}\Phi_a^\theta(\rho)\right).

For ρ∈S\rho\in S put Xρ=∣Ψ(λ1−iμρ)∣2X_\rho=|\Psi(\lambda_1-i\mu_\rho)|^2; then Φaθ(ρ)≥Xρ−3θ\Phi_a^\theta(\rho)\geq X_\rho-3\theta and ∣H((1−ρ)L)∣=e−Aλ1Xρ|H((1-\rho)\mathcal{L})|=e^{-A\lambda_1}X_\rho. We add e−Aλ1Xρe^{-A\lambda_1}X_\rho for every ρ∈S\rho\in S; for those in RPR_P this is their contribution, and for the others it is a nonnegative addition. By (5.1) the total is at most

e−ApH0Bϕ1θ(a)+∑ρ∈SXρ(e−Aλ1−e−Ap)+H0ε1.e^{-Ap}H_0B_{\phi_1}^\theta(a)+\sum_{\rho\in S}X_\rho\left(e^{-A\lambda_1}-e^{-Ap}\right)+H_0\varepsilon_1.

If λ1≤p\lambda_1\leq p, each summand increases with XρX_\rho, and Xρ≤Ψ(λ1)2≤Ψ(a)2=H0h~(a)X_\rho\leq\Psi(\lambda_1)^2\leq\Psi(a)^2=H_0\widetilde{h}(a); so the total is at most H0[e−Ap(Bϕ1θ(a)−αh~(a))+αe−Aλ1h~(a)+ε1]H_0[e^{-Ap}(B_{\phi_1}^\theta(a)-\alpha\widetilde{h}(a))+\alpha e^{-A\lambda_1}\widetilde{h}(a)+\varepsilon_1], which is at most H0(Jold/n+ε1)H_0(J_{\mathrm{old}}/n+\varepsilon_1). If λ1>p\lambda_1>p, each summand is ≤0\leq0, and the total is at most H0[e−ApBϕ1θ(a)+ε1]≤H0[e−Ap(Bϕ1θ(a)−αh~(a))++αe−Aph~(a)+ε1]H_0[e^{-Ap}B_{\phi_1}^\theta(a)+\varepsilon_1]\leq H_0[e^{-Ap}(B_{\phi_1}^\theta(a)-\alpha\widetilde{h}(a))_++\alpha e^{-Ap}\widetilde{h}(a)+\varepsilon_1], which is again at most H0(Jold/n+ε1)H_0(J_{\mathrm{old}}/n+\varepsilon_1) because p≥ap\geq a.

The bound JnewJ_{\mathrm{new}}. Now p≥b≥λ1p\geq b\geq\lambda_1. Apply Lemma 5.4 at σ=p\sigma=p. The non-distinguished zeros of the disc have Φpθ≥0\Phi_p^\theta\geq0, and for those in RPR_P, ∣H((1−ρ)L)∣≤e−Ap(Φpθ(ρ)+3θ)|H((1-\rho)\mathcal{L})|\leq e^{-Ap}(\Phi_p^\theta(\rho)+3\theta). Therefore the contribution of χ\chi is at most

e−Ap(H0Bϕ1θ(p)+14H0ε1+3Nθ)+∑ρ∈S(e−Aλ1∣Ψ(λ1−iμρ)∣21ρ∈RP−e−ApΦpθ(ρ)).e^{-Ap}\left(H_0B_{\phi_1}^\theta(p)+\frac{1}{4}H_0\varepsilon_1+3N\theta\right) +\sum_{\rho\in S}\left(e^{-A\lambda_1}|\Psi(\lambda_1-i\mu_\rho)|^2\mathbf{1}_{\rho\in R_P}-e^{-Ap}\Phi_p^\theta(\rho)\right).

The indicator may be replaced by 11, since the term it multiplies is nonnegative, and ∣Ψ(λ1−iμρ)∣2≤∣Ψθ(λ1−iμρ)∣2+3θ|\Psi(\lambda_1-i\mu_\rho)|^2\leq|\Psi^\theta(\lambda_1-i\mu_\rho)|^2+3\theta. For a real number yy consider

D(y)=e−Aλ1∣Ψθ(λ1+iy)∣2−e−ApRe⁡Fpθ(λ1−p+iy)=∫0∞k(t)cos⁡(yt) dt,D(y)=e^{-A\lambda_1}|\Psi^\theta(\lambda_1+iy)|^2-e^{-Ap}\operatorname{Re}F_p^\theta(\lambda_1-p+iy) =\int_0^\infty k(t)\cos(yt)\,\mathrm{d}t,

where, by Lemma 5.1(ii) for ψθ\psi^\theta and the definition of fpθf_p^\theta,

k(t)=e−Aλ1fλ1θ(t)−e−Ape(p−λ1)tfpθ(t)=2∫ψθ(u)ψθ(u+t)e−λ1t(e−λ1(A+2u)−e−p(A+2u)) du.k(t)=e^{-A\lambda_1}f_{\lambda_1}^\theta(t)-e^{-Ap}e^{(p-\lambda_1)t}f_p^\theta(t) =2\int\psi^\theta(u)\psi^\theta(u+t)e^{-\lambda_1t}\left(e^{-\lambda_1(A+2u)}-e^{-p(A+2u)}\right)\,\mathrm{d}u.

Since λ1≤p\lambda_1\leq p, the integrand is nonnegative, so k≥0k\geq0 and D(y)≤D(0)=∫0∞k(t) dtD(y)\leq D(0)=\int_0^\infty k(t)\,\mathrm{d}t. Moreover k(t)k(t) decreases as λ1\lambda_1 increases, for each tt, because both e−λ1te^{-\lambda_1t} and the bracket decrease. Hence

D(y)≤∫0∞k(t)∣λ1=a dt=e−AaΨθ(a)2−e−ApFpθ(a−p)=H0(e−Aah~θ(a)−e−ApCθ(p,a)),D(y)\leq\int_0^\infty k(t)\big|_{\lambda_1=a}\,\mathrm{d}t =e^{-Aa}\Psi^\theta(a)^2-e^{-Ap}F_p^\theta(a-p) =H_0\left(e^{-Aa}\widetilde{h}^\theta(a)-e^{-Ap}C^\theta(p,a)\right),

by Lemma 5.1(iv). By Lemma 5.2(ii),(iii), H0∣h~θ(a)−h~(a)∣≤3θH_0|\widetilde{h}^\theta(a)-\widetilde{h}(a)|\leq3\theta and H0∣Cθ(p,a)−C(p,a)∣≤3θH_0|C^\theta(p,a)-C(p,a)|\leq3\theta. So each of the α≤2\alpha\leq2 distinguished occurrences contributes at most H0(e−Aah~(a)−e−ApC(p,a))+9θH_0(e^{-Aa}\widetilde{h}(a)-e^{-Ap}C(p,a))+9\theta, and with (5.1) the contribution of χ\chi is at most H0(Jnew/n+ε1)H_0(J_{\mathrm{new}}/n+\varepsilon_1). □\square

We write J1=min⁡(Jold,Jnew)J_{1}=\min(J_{\mathrm{old}},J_{\mathrm{new}}), where JnewJ_{\mathrm{new}} is omitted when p<bp<b. In every leaf of the case tree the anchor pp is at most 2.293<3≤CP2.293<3\le C_{P}, so the hypothesis p≤CPp\le C_{P} holds. In JnewJ_{\mathrm{new}} the quantity Bϕ1(p)−αC(p,a)B_{\phi_{1}}(p)-\alpha C(p,a) may be negative; the argument does not require it to be positive.

The first family outside the buffer

Suppose now that ρ1∉RB\rho_{1}\notin R_{B}. Since λ1<CB\lambda_{1}<C_{B} in the regimes where this occurs, this means ∣μ1∣>CB|\mu_{1}|>C_{B}. The global zero ρ1\rho_{1} is not in RPR_{P} and contributes nothing to WW, but the first-family characters may have other zeros in RPR_{P}. As before, the distinguished occurrences are ρ1\rho_{1} (and ρ1‾\overline{\rho_{1}}) for χ1\chi_{1} and ρ1‾\overline{\rho_{1}} for χ1‾\overline{\chi_{1}}, and every other zero occurrence of the first family in R(l)R(l) has parameter at least p≥ap\ge a, where p∈Ap\in A (in the leaves, p=p∘p=p^{\circ}).

For a first-family character χ\chi with a zero in RPR_{P}, let tχt_{\chi} be the least parameter of its zeros in RPR_{P}. Then tχ≥pt_{\chi}\ge p, and the corresponding zero has ∣γ∣≤CP/L≤1|\gamma|\le C_{P}/\mathcal{L}\le1.

Proposition 5.9 (Cost of the first family outside the buffer). Suppose ρ1∉RB\rho_{1}\notin R_{B}, λ1≥a≥λ11\lambda_{1}\ge a\ge\lambda_{11}, and p∈Ap\in A satisfies a≤p≤λ′a\le p\le\lambda'. For a first-family character with a zero in RPR_{P}, let tχt_{\chi} be the least parameter of its zeros in RPR_{P}. Let α\alpha be the number of distinguished occurrences per character, and put ϕ1=1/4\phi_{1}=1/4 for a real χ1\chi_{1} and ϕ1=1/3\phi_{1}=1/3 otherwise. There is a constant coutc_{\mathrm{out}}, depending only on the fixed choices of stages (S1)–(S3) of Section 13.1 (in particular on θ\theta), such that for q≥q0q\ge q_{0} the following holds. Each first-family character χ\chi with a zero in RPR_{P} contributes at most

e−AtχBϕ1(p)+ε1+αcoutH0CB2e^{-At_{\chi}}B_{\phi_{1}}(p)+\varepsilon_{1}+\frac{\alpha c_{\mathrm{out}}}{H_{0}C_{B}^{2}}

to WW, and a first-family character without a zero in RPR_{P} contributes nothing.

Proof. Apply Lemma 5.4 at σ=p∈A\sigma=p\in A. For ρ∈RP\rho\in R_{P} we have λρ≥tχ≥p\lambda_{\rho}\ge t_{\chi}\ge p and ∣H((1−ρ)L)∣≤e−Atχ(Φpθ(ρ)+3θ)|H((1-\rho)\mathcal{L})|\le e^{-At_{\chi}}(\Phi_{p}^{\theta}(\rho)+3\theta). The non-distinguished zeros of the disc have Φpθ≥0\Phi_{p}^{\theta}\ge0, and so does a distinguished occurrence with λρ≥p\lambda_{\rho}\ge p. A distinguished occurrence ρ\rho in the disc with λρ<p\lambda_{\rho}<p has λρ−p∈[a−p,0)⊆[−CP,0]\lambda_{\rho}-p\in[a-p,0)\subseteq[-C_{P},0] and ∣μρ∣>CB|\mu_{\rho}|>C_{B}. Since fpθf_{p}^{\theta} is smooth on [0,T][0,T] and vanishes near TT, two integrations by parts give, for x∈[−CP,0]x\in[-C_{P},0] and real Y≠0Y\ne0,

∣Re⁡Fpθ(x+iY)∣=∣Re⁡(fpθ(0)x+iY+(fpθ)′(0)(x+iY)2+1(x+iY)2∫0T(fpθ)′′(t)e−(x+iY)t dt)∣≤coutY2,\left|\operatorname{Re}F_{p}^{\theta}(x+iY)\right| = \left|\operatorname{Re}\left( \frac{f_{p}^{\theta}(0)}{x+iY} + \frac{(f_{p}^{\theta})'(0)}{(x+iY)^{2}} + \frac{1}{(x+iY)^{2}}\int_{0}^{T}(f_{p}^{\theta})''(t)e^{-(x+iY)t}\,\mathrm{d}t \right)\right| \le \frac{c_{\mathrm{out}}}{Y^{2}},

where cout=CPfpθ(0)+∣(fpθ)′(0)∣+eCPT∥(fpθ)′′∥1c_{\mathrm{out}}=C_{P}f_{p}^{\theta}(0)+|(f_{p}^{\theta})'(0)|+e^{C_{P}T}\|(f_{p}^{\theta})''\|_{1}, maximized over the finitely many anchors pp in use; the first term is bounded by fpθ(0)∣x∣/Y2f_{p}^{\theta}(0)|x|/Y^{2}. So the distinguished terms subtracted in Lemma 5.4 total at least −αcout/CB2-\alpha c_{\mathrm{out}}/C_{B}^{2}, and summing as in Proposition 5.5, with (5.1), gives the claim. □\square

In the leaf programs these contributions are the hidden columns: at most one per first-family character, carrying the value tχt_{\chi}, the objective e−AtχBϕ1(p)e^{-At_{\chi}}B_{\phi_{1}}(p) and, because the zero realizing tχt_{\chi} has ∣γ∣≤1|\gamma|\le1, the far weight w(tχ)w(t_{\chi}) (Section 11.5).

The objective and the error allowance

In a leaf of the case tree (Section 2.4) the unknown characters are described as follows (Section 11.5): the first family, possibly a reserved second family of n2∈{1,2}n_{2}\in\{1,2\} characters with height-one representatives in a given interval [l02,h2][l_{02},h_{2}], and the ordinary characters, described by their T∗T^{*}-representatives. By Corollary 5.7 and Propositions 5.8 and 5.9,

W≤J1+∑χ reservedGϕ2(νχ)+∑χ ordinaryG(λχ)+E,(8)W\le J_{1} +\sum_{\chi\ \mathrm{reserved}}\mathcal{G}_{\phi_{2}}(\nu_{\chi}) +\sum_{\chi\ \mathrm{ordinary}}\mathcal{G}(\lambda_{\chi}) +E, \tag*{(8)}

where ϕ2=14\phi_{2}=\frac{1}{4} if the reserved family is a single real character and ϕ2=13\phi_{2}=\frac{1}{3} otherwise, νχ\nu_{\chi} is the parameter of the height-one representative of a reserved character. The first-family term J1J_{1} is J1J_{1} for an inside leaf and ∑χe−AtχBϕ1(p)\sum_{\chi}e^{-At_{\chi}}B_{\phi_{1}}(p) for an outside leaf. In the latter case it is represented by hidden columns, rather than by the constant part of the leaf objective (Section 11.5).

There are two sources of error in (8). The envelope errors total less than 19ε119\varepsilon_{1} for the ordinary characters, by (10), and at most 4ε14\varepsilon_{1} for the first and reserved families. The outside first-family bound has the additional error 4cout/(H0CB2)4c_{\mathrm{out}}/(H_{0}C_{B}^{2}). We choose

ε1=η23,4coutH0CB2≤η,\varepsilon_{1}=\frac{\eta}{23},\qquad\frac{4c_{\mathrm{out}}}{H_{0}C_{B}^{2}}\leq\eta,

so E≤2ηE\leq2\eta. The error η\eta in the prime-detection criterion is accounted for separately in Proposition 4.4.

Every leaf objective includes an allowance of 5η5\eta. Consequently its value bounds W+3ηW+3\eta, and therefore also W+2ηW+2\eta, as required by the positivity criterion. These choices are made uniformly over the finite collection of leaves in Section 13.1.

Far density and representatives

The per-character costs must be summed over many characters. A weighted density estimate supplies a common budget for that sum. Its zeros have physical height at most 1, so we also specify how to choose representatives within this height range.

Xylouris’s weighted density lemma

The following is Xylouris’s improvement [145], Lemma 5.1, (5.19), p. 66 of Heath-Brown’s Lemma 11.1 [51]; the latter is the case J=1J=1, w0≡1w_{0}\equiv1. (Xylouris writes MM for our JJ; we reserve MM for a separation constant.)

Input 6.1 (Far density). Let ε,c1,c2>0\varepsilon,c_{1},c_{2}>0, J∈NJ\in\mathbb{N}, and α1,…,αJ≥0\alpha_{1},\ldots,\alpha_{J}\geq0 with ∑iαi=1\sum_{i}\alpha_{i}=1. Put

x=23+3c1+c2,ui=13+2c1+ic2J(0≤i≤J),x=\frac{2}{3}+3c_{1}+c_{2},\qquad u_{i}=\frac{1}{3}+2c_{1}+\frac{ic_{2}}{J} \qquad(0\leq i\leq J),

and let w0:[u0,x]→Rw_{0}:[u_{0},x]\to\mathbb{R} be continuous and continuously differentiable except at finitely many points, with 1≪w0≪11\ll w_{0}\ll1 and w0′≪1w_{0}'\ll1. Choose for each nonprincipal character χ\chi modulo qq at most one zero ρχ\rho_{\chi} of L(s,χ)L(s,\chi) with ∣Im⁡ρχ∣≤1|\operatorname{Im}\rho_{\chi}|\leq1 and λχ≤λ0=13log⁡log⁡L\lambda_{\chi}\leq\lambda_{0}=\frac{1}{3}\log\log\mathcal{L}. Then for q≥q0q\geq q_{0}, with q0q_{0} depending on all the parameters,

∑χ(∫u0xw0(t)2e2λχt dt)−1≤J2+εc1c22∑i=1Jαi2∫ui−1xw0(t)−2min⁡{t−ui−1,ui−ui−1} dt.\sum_{\chi}\left(\int_{u_{0}}^{x}w_{0}(t)^{2}e^{2\lambda_{\chi}t}\,\mathrm{d}t\right)^{-1} \leq \frac{J^{2}+\varepsilon}{c_{1}c_{2}^{2}} \sum_{i=1}^{J}\alpha_{i}^{2} \int_{u_{i-1}}^{x}w_{0}(t)^{-2} \min\{t-u_{i-1},u_{i}-u_{i-1}\}\,\mathrm{d}t.

In [145] the statement is made for the zeros ρ(k)\rho^{(k)} that are chosen for the characters counted by N(λ0)N(\lambda_{0}), one zero per character [145], §3.2.2, pp. 25–26, [51], §11; any such choice is allowed, and the estimates in the proof [145], pp. 67–72 are uniform over these choices. This uniformity is needed because the selected zeros, including the T∗(q)T^{*}(q)-representatives below, depend on qq. A nonreal character and its conjugate are different characters, and each contributes its own term; the principal character is excluded, since L(s,χ0)L(s,\chi_{0}) has no zero in the relevant region for large qq [145], p. 21. The hypotheses on w0w_{0} are satisfied by our profiles, which are those of Xylouris’s own application [145], (6.20)–(6.21), p. 80.

The two weight profiles

We use two profiles of the shape [145], (6.20), both with J=10J=10 and ε0=10−7\varepsilon_{0}=10^{-7}:

w0(t)2=e−θfartmin⁡(t−u0+ε0, c2+ε0)(u0≤t≤x).w_{0}(t)^{2}=e^{-\theta_{\mathrm{far}}t}\sqrt{\min(t-u_{0}+\varepsilon_{0},\,c_{2}+\varepsilon_{0})} \qquad(u_{0}\leq t\leq x).

Profile I has c1=0.09035c_{1}=0.09035, c2=0.235968c_{2}=0.235968 and θfar=1.28683\theta_{\mathrm{far}}=1.28683, and profile II has c1=0.0821922c_{1}=0.0821922, c2=0.2170903c_{2}=0.2170903 and θfar=1.4964274\theta_{\mathrm{far}}=1.4964274. The weights αi\alpha_{i} are the exact decimals of Table 4; in each profile they sum to 1. Both profiles satisfy the hypotheses of Input 6.1: w0w_{0} is continuous and positive on [u0,x][u_{0},x], and it is smooth except at t=u0+c2t=u_{0}+c_{2}.

iiprofile I αi\alpha_iprofile II αi\alpha_i
10.07888272180.0806195583
20.08493861480.0862705300
30.08956297790.0905489307
40.09382315160.0944644867
50.09797104910.0982543287
60.10212849160.1020315591
70.10637329390.1058668825
80.11076518690.1098131140
90.11535654920.1139152197
100.12019796320.1182153903

Table 4. The weights αi\alpha_i of the two far profiles.

Put

w(λ)=(∫u0xw0(t)2e2λt dt)−1,w(\lambda)=\left(\int_{u_0}^{x} w_0(t)^2 e^{2\lambda t}\,\mathrm{d}t\right)^{-1},

a positive decreasing function of λ\lambda. Let VV be the right-hand side of Input 6.1 with ε=0\varepsilon=0. Rigorous enclosures give

V=175.26640331… (profile I),V=243.32098105… (profile II).V=175.26640331\ldots\ \text{(profile I)},\qquad V=243.32098105\ldots\ \text{(profile II)}.

Taking ε=ηJ2\varepsilon=\eta J^2 in Input 6.1, we obtain for q≥q0q\geq q_0 and every choice of zeros as in Input 6.1:

∑χw(λχ)≤(1+η)V.(9)\sum_{\chi} w(\lambda_\chi)\leq(1+\eta)V. \tag*{(9)}

Each leaf of the case tree uses one of the two profiles, recorded in its data. Both are fixed finite choices within Input 6.1.

Lemma 6.2 (Comparison of the prime weight and the far weight). For both profiles, x≤1.1737x\leq1.1737. Hence A>2xA>2x for L=3.99L=3.99, and for every λ≥0\lambda\geq0 the function λ↦e−Aλ/w(λ)\lambda\mapsto e^{-A\lambda}/w(\lambda) is decreasing. In particular e−Aλ≤w(λ)/w(0)e^{-A\lambda}\leq w(\lambda)/w(0).

Proof. The values are x=1.1736846…x=1.1736846\ldots and x=1.1303335…x=1.1303335\ldots, so 2x<2.348<A=3.1563422x<2.348<A=3.156342. Now e−Aλ/w(λ)=∫u0xw0(t)2e(2t−A)λ dte^{-A\lambda}/w(\lambda)=\int_{u_0}^{x}w_0(t)^2e^{(2t-A)\lambda}\,\mathrm{d}t, and 2t−A<02t-A<0 on [u0,x][u_0,x]. □\square

With w(0)=10.7054…w(0)=10.7054\ldots (profile I) and w(0)=13.1730…w(0)=13.1730\ldots (profile II), Lemma 6.2 and (9) give, for any selection satisfying Input 6.1,

∑χe−Aλχ≤(1+η)Vw(0)=Kfar,Kfar<16.38 (profile I),Kfar<18.48 (profile II).(10)\sum_{\chi}e^{-A\lambda_\chi}\leq\frac{(1+\eta)V}{w(0)}=K_{\mathrm{far}},\qquad K_{\mathrm{far}}<16.38\ \text{(profile I)},\qquad K_{\mathrm{far}}<18.48\ \text{(profile II)}. \tag*{(10)}

This bound controls errors that are summed over unboundedly many characters.

Representatives

The leaf programs describe an ordinary character through one parameter, the parameter of a representative zero. For the near-density argument of Section 8 the representative must be separated in height from any zero with a smaller parameter, and we arrange this by choosing the height of the strip in which representatives are taken.

Fix smax⁡=3s_{\max}=3; every anchor used in the near rows is at most 2.32.3 (Section 9), and every ordinary entry uses an anchor at most smax⁡s_{\max}. Let K0K_0 be an upper bound for the number of zeros, of all nonprincipal characters modulo qq and counted with multiplicity, with λ≤smax⁡\lambda\leq s_{\max} and ∣γ∣≤2|\gamma|\leq2. By a log-free zero-density estimate (Input 3.9), K0K_0 may be taken independent of qq. Next fix MM as in Section 13; it depends on K0K_0.

Lemma 6.3 (A strip boundary separated from low-parameter zeros). Let K0K_{0} bound the number of zeros with λ≤smax⁡=3\lambda\le s_{\max}=3 and ∣γ∣≤2|\gamma|\le2, as above, and fix M>0M>0. If L>4M(K0+1)\mathcal{L}>4M(K_{0}+1), there is T∗∈[1/2,1]T^{*}\in[1/2,1] such that no nonprincipal L(s,χ)L(s,\chi) has a zero with λ≤smax⁡\lambda\le s_{\max} and ∣∣μ∣−T∗L∣<M\bigl||\mu|-T^{*}\mathcal{L}\bigr|<M. We take T∗=T∗(q)T^{*}=T^{*}(q) to be the least such number.

Proof. The at most K0K_{0} zeros in question exclude at most K0K_{0} open intervals of TT of length 2M/L2M/\mathcal{L} each. The interval [1/2,1][1/2,1] has length 1/2>2MK0/L1/2>2MK_{0}/\mathcal{L}, so the remaining set is a nonempty finite union of closed intervals, and it has a least element. □\square

Definition 6.4 (Representatives). For a nonprincipal character χ\chi, its T∗T^{*}-representative is a zero of L(s,χ)L(s,\chi) in R(ℓ)R(\ell) with ∣γ∣≤T∗|\gamma|\le T^{*} and least parameter, if there is one; its parameter is λχ\lambda_{\chi}. Representatives are used for the characters outside the first family; in Theorem 12.6 all characters use their height-one representatives. The height-one representative is defined in the same way with the bound ∣γ∣≤1|\gamma|\le1. The reserved family of a leaf (Section 2.4 and Definition 11.2) and an inside first family use height-one representatives.

Every representative has ∣γ∣≤1|\gamma|\le1, so the far bound (9) applies to any selection of representatives, one per character. Representatives for conjugate characters may be chosen as conjugate zeros, with the same parameter, because every selection region is symmetric in height. The T∗T^{*}-representative of χ\chi satisfies λχ≥\lambda_{\chi}\ge the parameter of its height-one representative. Shrinking the strip can only increase the minimum parameter; in terms of real parts, the selected zero can only move to the left.

The far constraint of a leaf

In a leaf program (Section 2.4), the characters other than the first family are placed in bins by the parameters of their representatives, and each bin has a column (Sections 10 and 11). The column of a bin [lo,hi][\mathrm{lo},\mathrm{hi}] has far weight Wc=⌊Sw(hi)⌋W_{c}=\lfloor Sw(\mathrm{hi})\rfloor; since ww is decreasing, each character in the bin contributes at least w(hi)w(\mathrm{hi}) to the left side of (9). The tail column collects the ordinary characters χ\chi with a zero in RPR_{P} and λχ≥R\lambda_{\chi}\ge R, where R=max⁡(3,r)R=\max(3,r) and rr is the leaf’s lower bound for ordinary representatives. Its value is the mass xtail=∑w(λχ)x_{\mathrm{tail}}=\sum w(\lambda_{\chi}) over these characters, with far weight SS. The far budget of the leaf is

F=⌊S(1+η)V⌋−1insiden⌊Sw(b)⌋.F=\lfloor S(1+\eta)V\rfloor-\mathbf{1}_{\mathrm{inside}}n\lfloor Sw(b)\rfloor.

Here the last term is present for an inside leaf, that is when ρ1∈RB\rho_{1}\in R_{B}. Then ∣γ1∣≤CB/L≤1|\gamma_{1}|\le C_{B}/\mathcal{L}\le1, so ρ1\rho_{1} and ρ1‾\overline{\rho_{1}} are the height-one representatives of χ1\chi_{1} and χ1‾\overline{\chi_{1}}, and they contribute nw(λ1)≥nw(b)nw(\lambda_{1})\ge nw(b) to (9). A reserved second family is charged either through its own columns, which carry its far weight, or by subtracting the fixed amount n2⌊Sw(hi2)⌋n_{2}\lfloor Sw(\mathrm{hi}_{2})\rfloor from FF (Section 11.5).

Zero location

This section collects the lower bounds for λ1\lambda_{1}, λ′\lambda', λ2\lambda_{2} and λ3\lambda_{3} that the case tree uses. Most are printed results of Heath-Brown [51] and Xylouris [145]; we state each in the form used, with its scope. The others are new consequences of the same explicit formula, proved here by the positivity method; one of them uses a sharpened conductor coefficient, which we prove in Section 7.6. None of the results of this section depends on the exponent LL.

Conventions

We fix once and for all an admissible height function l(q)l(q) as in Input 3.5, for instance the least one. All the zero-location results we quote are proved for an arbitrary integer ll with 1≤l≤L/101\le l\le\mathcal{L}/10 for which R(10l)∖R(l)R(10l)\setminus R(l) contains no zero; their proofs use ll only through these two properties, and their thresholds depend only on fixed data. So they hold simultaneously for our l(q)l(q).

The zeros ρ1,ρ2,ρ3\rho_{1},\rho_{2},\rho_{3} are selected as in [145], (3.11)–(3.12), p. 21 (which follows [51], §6). At step kk, remove the characters in the families F1,…,Fk−1\mathcal{F}_{1},\ldots,\mathcal{F}_{k-1} and choose a zero of maximal real part in R(l)R(l) among the remaining nonprincipal LL-functions. Each real character is removed once. This is the successive-minimum convention of Definition 2.1. The additional zero ρ′\rho' of the first family is selected as in [145]: ρ′=ρ1\rho'=\rho_{1} if ρ1\rho_{1} is multiple, and otherwise ρ′\rho' has maximal real part among the zeros of L(s,χ1)L(s,\chi_{1}) in R(l)R(l) other than ρ1\rho_{1} (and other than ρ‾1\overline{\rho}_{1} if χ1\chi_{1} is real). Its parameter is the λ′\lambda' of Section 2 (Heath-Brown also writes λ′\lambda'; Xylouris indexes this zero by 0). When several zeros qualify (ties in the real part, or several characters attaining a minimum), any of them may be selected: the proofs of the quoted results use only the extremal property of the selection, so they hold for every such choice. This is why the configuration sets of Definition 11.2 quantify over the admissible choices. All the results of this section concern zeros in R(l)R(l) at any height up to ll, so they apply whether ρ1\rho_{1} is inside or outside the buffer.

Printed tables. Every printed bound is used as an exact implication for q≥q0q\geq q_{0}. Each of them was obtained by showing that a certain inequality fails with a positive margin, which persists when the error term ε\varepsilon is small enough [145]; for Heath-Brown’s Tables 2 and 3 the printed values lie a little below the computed roots [51]. Heath-Brown states no convention for his Tables 4 and 7, and we have recomputed their roots from his printed parameters; every recomputed root exceeds the printed value, and each printed value is the root rounded down. The closest case in Table 4 is row 1.05, with root 1.43904 against 1.439, and the closest in Table 7 are rows 0.20 and 0.55, with 2.01015 against 2.01 and 1.00036 against 1.00. The rows of Table 4 are derived under the side conditions [51]; we have checked from the printed parameters that these hold for every row from 0.35 on (the tightest is (8.6), with slacks 3.8⋅10−63.8\cdot10^{-6}, 3.8⋅10−63.8\cdot10^{-6} and 2.3⋅10−62.3\cdot10^{-6} for the rows 1.05, 1.15 and 1.294). For the row (0.3,2.293)(0.3,2.293) they need not hold, but that row also follows from Heath-Brown’s Table 3, which gives λ′≥2.43\lambda'\geq2.43 for λ1≤0.30\lambda_{1}\leq0.30 [51]. These checks, and the recomputation of the roots, were made in floating point, separately from the replay. A separate program compares every printed number that the proof uses, in the code and in this paper, with the text of the source pages and with the printed scope of its row; each agrees with the printed value or is weaker than it in the safe direction. Results stated in the form c−εc-\varepsilon are used with an explicit reduction: 0.857<670.857<\frac{6}{7} in Input 7.4, and 1.09<12111.09<\frac{12}{11} and ε=0.001\varepsilon=0.001 in Section 12.

Types and the first zero.

Proposition 7.1 (Lower bounds for a nonreal first zero). For all sufficiently large qq, the first family satisfies

λ1>0.440if χ1 is nonreal,λ1>0.628if χ1 is real and ρ1 is nonreal.\lambda_{1}>0.440\quad\text{if }\chi_{1}\text{ is nonreal},\qquad \lambda_{1}>0.628\quad\text{if }\chi_{1}\text{ is real and }\rho_{1}\text{ is nonreal}.

Thus λ1≤0.44\lambda_{1}\leq0.44 forces both χ1\chi_{1} and ρ1\rho_{1} to be real. These bounds are [145].

Xylouris’s lemma gives λ1>0.440\lambda_{1}>0.440, 0.493, 0.478, 0.498 and 0.628 for characters of order ≥6\geq6, =5=5, =4=4, =3=3 and =2=2 respectively; the minimum over orders at least 3 is 0.440.

The test function.

The location inequalities use one family of test functions. We define it here so that all parameters in the statements below are specified before they occur. For γ>0\gamma>0, put

fγ(t)={16γ515(1−u)3(1+3u+u2),0≤t≤2γ, u=t/(2γ),0,t≥2γ,(11)f_{\gamma}(t)= \begin{cases} \frac{16\gamma^{5}}{15}(1-u)^{3}(1+3u+u^{2}), & 0\leq t\leq2\gamma,\ u=t/(2\gamma),\\ 0, & t\geq2\gamma, \end{cases} \tag*{(11)}

The function fγf_{\gamma} is the autocorrelation of x↦(γ2−x2)+x\mapsto(\gamma^{2}-x^{2})_{+}, the test function of [145]; it satisfies Conditions 1 and 2 (Lemma 9.1), and its transform FF is positive and decreasing on R\mathbb{R}.

Printed zero-location inputs.

Input 7.2 (Real first zero). Let χ1\chi_{1} and ρ1\rho_{1} be real. For q≥q0q \ge q_{0}:

(a) [51], Tables 3 and 4, Lemmas 8.2 and 8.3. If λ1≤t\lambda_{1} \le t then λ′≥c\lambda' \ge c, for the 24 pairs (t,c)(t,c) from (0.3,2.293)(0.3,2.293) to (1.294,1.294)(1.294,1.294) of the table. For a real ρ′\rho' this follows from λ′≥2.427\lambda' \ge2.427 (Lemma 8.2).

(b) [51], Lemma 8.4. For every ε>0\varepsilon> 0: λ′≥(2−ε)log⁡λ1−1\lambda' \ge(2-\varepsilon)\log\lambda_{1}^{-1} if λ1≤0.2\lambda_{1} \le0.2 and q≥q(ε)q \ge q(\varepsilon); and λ′≥max⁡(32log⁡λ1−1,1.294)\lambda' \ge\max(\frac{3}{2}\log\lambda_{1}^{-1},1.294) for every λ1\lambda_{1}.

(c) [51], Table 7 and Lemma 8.7. If λ1≤t\lambda_{1} \le t then λ2≥c\lambda_{2} \ge c, for the 16 pairs (t,c)(t,c) from (0.12,2.56)(0.12,2.56) to (0.745,0.745)(0.745,0.745) of the table; these hold whether or not χ24=χ0\chi_{2}^{4}=\chi_{0} (the row t=0.10t=0.10 is not used).

EntryAnchor σ\sigmaKept set SSResponse bound rrDiagonal numeratorRows
(a) ordinary χ\chi, bin [lo,hi)[\mathrm{lo},\mathrm{hi}), at its T∗T^{*}-representativemin⁡(lo,s)\min(\mathrm{lo},s){ρχ}\{\rho_{\chi}\}F(hi−σ)−16f(0)F(\mathrm{hi}-\sigma)-\frac{1}{6}f(0)D(s−σ)−d1D(s-\sigma)-d_{1}all three
(b) inside first family: χ1\chi_{1} at γ1\gamma_{1}, and χ‾1\overline{\chi}_{1} at −γ1-\gamma_{1}min⁡(a,s)\min(a,s){ρ1}\{\rho_{1}\}, resp. {ρ‾1}\{\overline{\rho}_{1}\}F(b−σ)−16f(0)F(b-\sigma)-\frac{1}{6}f(0)D(s−σ)−d1D(s-\sigma)-d_{1}family, graded
(c) type rc: χ1\chi_{1} at ±γ1\pm\gamma_{1}s1s_{1}{ρ1,ρ‾1}\{\rho_{1},\overline{\rho}_{1}\}F(b−s1)+Rlo−16f(0)F(b-s_{1})+R_{\mathrm{lo}}-\frac{1}{6}f(0)D(0)−d1+(chi−d1)+D(0)-d_{1}+(c^{\mathrm{hi}}-d_{1})_{+}family
(d) second zero, type complex: χ1\chi_{1} at γ1,γ′\gamma_{1},\gamma', χ‾1\overline{\chi}_{1} at −γ1,−γ′-\gamma_{1},-\gamma's1s_{1}{ρ1,ρ′}\{\rho_{1},\rho'\} and conjugatesF(b−s1)+Rp−16f(0)F(b-s_{1})+R_{p}-\frac{1}{6}f(0) at γ1\gamma_{1}; F(hi′−s1)+Rq−16f(0)F(\mathrm{hi}'-s_{1})+R_{q}-\frac{1}{6}f(0) at γ′\gamma'D(0)−d1+(chi−d1)+D(0)-d_{1}+(c^{\mathrm{hi}}-d_{1})_{+}family
(e) shifted first family: χ1\chi_{1} at γ1\gamma_{1}, and χ‾1\overline{\chi}_{1} at −γ1-\gamma_{1}ss{ρ1}\{\rho_{1}\}; {ρ1,ρ‾1}\{\rho_{1},\overline{\rho}_{1}\} for rcF(b−s)−1rcCZ−16f(0)F(b-s)-\mathbf{1}_{\mathrm{rc}}C_{Z}-\frac{1}{6}f(0)D(0)−d1D(0)-d_{1}shifted
(f) reserved family, at height-one representativesσ2\sigma_{2}the representativeF(hi2−σ2)−16f(0)F(\mathrm{hi}_{2}-\sigma_{2})-\frac{1}{6}f(0)D(s−σ2)−d1D(s-\sigma_{2})-d_{1}all three
(g) outside first family: χ1\chi_{1} at γ1\gamma_{1} (and χ‾1\overline{\chi}_{1}), and one hidden entry per characters1s_{1}{ρ1}\{\rho_{1}\}; {ρh}\{\rho_{h}\}F(b−s1)−16f(0)F(b-s_{1})-\frac{1}{6}f(0); F(hih−s1)−16f(0)F(\mathrm{hi}_{h}-s_{1})-\frac{1}{6}f(0)D(0)−d1D(0)-d_{1}family

Table 7. The entries of the near rows, justified in the list below. The anchor σ\sigma enters through the offset s−σs-\sigma, where ss is the reference anchor of the row; σ2\sigma_{2} and the constants RloR_{\mathrm{lo}}, RpR_{p}, RqR_{q}, chic^{\mathrm{hi}} and CZC_{Z} are defined in the list. “All three” means the family, shifted and graded rows of Section 9.5.

(d) [51], Lemma 8.8. λ2≥max⁡((1211−ε)log⁡λ1−1,0.745)\lambda_{2} \ge\max((\frac{12}{11}-\varepsilon)\log\lambda_{1}^{-1},0.745) for q≥q(ε)q \ge q(\varepsilon).

Heath-Brown remarks that the estimates of his §8 have not been proved for extremely small values of λ1\lambda_{1} [51], §8. We use (a)–(d) only for λ1\lambda_{1} bounded below by a fixed positive constant: λ1≥0.1\lambda_{1} \ge0.1 in the case tree, and, in Theorem 12.5, λ1\lambda_{1} at least the fixed constant uexc>0u_{\mathrm{exc}}>0 of Section 12.

Input 7.3 (Nonreal first character or zero). Let χ1\chi_{1} or ρ1\rho_{1} be nonreal. For q≥q0q \ge q_{0}:

(a) [145], Table 2′, p. 62. If λ1≤t\lambda_{1} \le t then λ′>c\lambda' > c, for the pairs (0.46,1.85),…,(0.827,0.827)(0.46,1.85),\ldots,(0.827,0.827) of the table.

(b) [145], Table 3, p. 45. If ord⁡χ1∈{2,3,4}\operatorname{ord}\chi_{1} \in\{2,3,4\} and λ1≤t\lambda_{1} \le t, then λ′>c\lambda' > c, for the pairs (0.38,2.53),…,(1.099,1.099)(0.38,2.53),\ldots,(1.099,1.099); in particular λ1≤0.86\lambda_{1} \le0.86 implies λ′>1.35\lambda' > 1.35.

(c) [145], Tables 6 and 7, p. 53. If λ1≤t\lambda_{1} \le t then λ2>c\lambda_{2}>c; Table 6 concerns a real χ1\chi_{1} (Fall 7), Table 7 all cases.

(d) [51], Lemma 9.4 and Table 10. λ2≥0.702\lambda_{2} \ge0.702 always, and λ1≤0.70\lambda_{1} \le0.70 implies λ2≥0.704\lambda_{2} \ge0.704.

Input 7.4 (The third family). For q≥q0q \ge q_{0}:

(a) [51], Lemma 10.3. λ3≥67−ε\lambda_{3} \ge\frac{6}{7}-\varepsilon; we use λ3>0.857\lambda_{3}>0.857.

(b) [145], Table 8, column “alle Fälle”, p. 55. If χ1\chi_{1} or ρ1\rho_{1} is nonreal and λ1≤t\lambda_{1} \le t, then λ3>c\lambda_{3}>c, for (t,c)=(0.52,1.320)(t,c)=(0.52,1.320), (0.54,1.243)(0.54,1.243), (0.56,1.160)(0.56,1.160), (0.58,1.079)(0.58,1.079), (0.60,1.001)(0.60,1.001), (0.62,0.933)(0.62,0.933).

(c) [145], Lemma 4.4 and Table 10, p. 55. If χ1\chi_{1} and ρ1\rho_{1} are real, then λ1∈[0.44,0.60]\lambda_{1} \in[0.44,0.60], [0.60,0.70][0.60,0.70], [0.70,0.80][0.70,0.80] imply λ3>1.176\lambda_{3}>1.176, 1.0551.055, 0.9520.952.

(d) [145], (4.28)–(4.29), pp. 56–57. Let χ1\chi_{1} be nonreal and λ1∈[0.44,0.85]\lambda_{1} \in[0.44,0.85], and let ff be the test function (11) with γ=54\gamma=\frac{5}{4}. Then

0≤F(−λ1)−F(λ3−λ1)−F(λ2−λ1)−F(0)+76f(0)+ε,0 \le F(-\lambda_{1})-F(\lambda_{3}-\lambda_{1})-F(\lambda_{2}-\lambda_{1})-F(0)+\frac{7}{6}f(0)+\varepsilon,

provided that sup⁡y∈RRe⁡{F(−λ1+iy)−2F(iy)}≤16f(0)\sup_{y \in\mathbb{R}}\operatorname{Re}\{F(-\lambda_{1}+iy)-2F(iy)\} \le\frac{1}{6}f(0).

(e) [145], (4.31) and (4.34), pp. 58–59. Let χ1\chi_{1} and ρ1\rho_{1} be real, λ1∈[0.44,0.80]\lambda_{1} \in[0.44,0.80] and λ2∈[0.44,1.176]\lambda_{2} \in[0.44,1.176], and let ff be (11) with γ=1.04\gamma=1.04. Then

0≤F(−λ2)−F(λ3−λ2)−F(0)−F(λ1−λ2)+98f(0)+ε,0 \le F(-\lambda_{2})-F(\lambda_{3}-\lambda_{2})-F(0)-F(\lambda_{1}-\lambda_{2})+\frac{9}{8}f(0)+\varepsilon,

provided that sup⁡t∈RRe⁡{F(−λ2+it)−F(λ1−λ2+it)−F(it)}≤548f(0)\sup_{t \in\mathbb{R}}\operatorname{Re}\{F(-\lambda_{2}+it)-F(\lambda_{1}-\lambda_{2}+it)-F(it)\} \le\frac{5}{48}f(0) on this box.

In (d), the side condition is [145], (4.29), needed when a product character in his argument is principal; he verified it numerically, and we have verified it with interval arithmetic for all λ1∈[0.44,0.85]\lambda_{1} \in[0.44,0.85]: the supremum is at most 0.003350.00335, while 16f(0)=0.5425…\frac{1}{6}f(0)=0.5425\ldots. In (e), Xylouris’s argument [145], (4.32)–(4.34), p. 59 needs

sup⁡t∈RRe⁡{F(−λ2+it)−F(λ1−λ2+it)−F(it)}≤548f(0)\sup_{t \in\mathbb{R}}\operatorname{Re}\{F(-\lambda_{2}+it)-F(\lambda_{1}-\lambda_{2}+it)-F(it)\} \le\frac{5}{48}f(0)

on the whole box, and he states that the supremum is below 0.10, while 548f(0)=0.1351…\frac{5}{48}f(0)=0.1351\ldots. We have verified this with interval arithmetic as well: over λ1∈[0.44,0.80]\lambda_{1}\in[0.44,0.80], λ2∈[0.44,1.18]\lambda_{2}\in[0.44,1.18] and all real tt the supremum is at most 0.0031. (Both checks enclose the transforms on a grid of heights with step 1100\frac{1}{100} up to 30, with second-derivative interpolation errors in all variables, and use a closed-form bound beyond height 30.) (Xylouris derives eq:4.28 from Heath-Brown’s (19)–eq:10.6, whose derivation does not use the assumption λ3≤67\lambda_{3}\le\frac{6}{7} made in [51], §10; he himself applies it to prove bounds above 67\frac{6}{7}.)

Proposition 7.5 (Parent rows). The 58 parent rows (t,Aj,Bj,pj,rj)(t,A_{j},B_{j},p_{j},r_{j}) of Table 5, from which the roots of the case tree are built (Section 11.3), are valid implications: for q≥q0q\ge q_{0}, every configuration of type tt with λ1∈[Aj,Bj]\lambda_{1}\in[A_{j},B_{j}] has λ′≥pj\lambda'\ge p_{j} and λ2≥rj\lambda_{2}\ge r_{j}.

Type[Aj,Bj][A_j,B_j]pjp_jrjr_jType[Aj,Bj][A_j,B_j]pjp_jrjr_j
rr[0.1,0.12][0.1,0.12]2.2932.2932.562.56rc[0.628,0.66][0.628,0.66]1.671.670.930.93
rr[0.12,0.14][0.12,0.14]2.2932.2932.392.39rc[0.66,0.7][0.66,0.7]1.591.590.930.93
rr[0.14,0.16][0.14,0.16]2.2932.2932.252.25rc[0.7,0.74][0.7,0.74]1.521.520.930.93
rr[0.16,0.18][0.16,0.18]2.2932.2932.122.12rc[0.74,0.78][0.74,0.78]1.461.460.930.93
rr[0.18,0.2][0.18,0.2]2.2932.2932.012.01rc[0.78,0.82][0.78,0.82]1.401.400.820.82
rr[0.2,0.25][0.2,0.25]2.2932.2931.771.77rc[0.82,0.86][0.82,0.86]1.351.350.820.82
rr[0.25,0.3][0.25,0.3]2.2932.2931.581.58rc[0.86,0.9][0.86,0.9]1.301.300.820.82
rr[0.3,0.348][0.3,0.348]2.1952.1951.421.42rc[0.9,0.94][0.9,0.94]1.251.250.820.82
rr[0.348,0.35][0.348,0.35]2.1952.1951.421.42rc[0.94,0.98][0.94,0.98]1.211.210.820.82
rr[0.35,0.4][0.35,0.4]2.1082.1081.291.29rc[0.98,1.02][0.98,1.02]1.171.170.820.82
rr[0.4,0.45][0.4,0.45]2.032.031.181.18rc[1.02,1.06][1.02,1.06]1.131.130.820.82
rr[0.45,0.5][0.45,0.5]1.9581.9581.081.08rc[1.06,1.099][1.06,1.099]1.0991.0990.820.82
rr[0.5,0.55][0.5,0.55]1.8931.8931.01.0rc[1.099,1.5][1.099,1.5]1.0991.0990.820.82
rr[0.55,0.6][0.55,0.6]1.8321.8320.920.92complex[0.44,0.58][0.44,0.58]1.361.361.041.04
rr[0.6,0.65][0.6,0.65]1.7761.7760.850.85complex[0.58,0.64][0.58,0.64]1.151.150.850.85
rr[0.65,0.7][0.65,0.7]1.7241.7240.790.79complex[0.64,0.66][0.64,0.66]1.081.080.790.79
rr[0.7,0.745][0.7,0.745]1.6761.6760.7450.745complex[0.66,0.68][0.66,0.68]1.021.020.740.74
rr[0.745,0.75][0.745,0.75]1.6761.6760.7450.745complex[0.68,0.70][0.68,0.70]0.960.960.7040.704
rr[0.75,0.8][0.75,0.8]1.631.630.7450.745complex[0.70,0.72][0.70,0.72]0.930.930.7020.702
rr[0.8,0.85][0.8,0.85]1.5871.5870.80.8complex[0.72,0.74][0.72,0.74]0.910.910.720.72
rr[0.85,0.9][0.85,0.9]1.5471.5470.80.8complex[0.74,0.76][0.74,0.76]0.890.890.740.74
rr[0.9,0.95][0.9,0.95]1.5091.5090.80.8complex[0.76,0.78][0.76,0.78]0.860.860.740.74
rr[0.95,1][0.95,1]1.4731.4730.80.8complex[0.78,0.80][0.78,0.80]0.840.840.780.78
rr[1,1.05][1,1.05]1.4391.4390.80.8complex[0.80,0.82][0.80,0.82]0.830.830.780.78
rr[1.05,1.1][1.05,1.1]1.4061.4060.80.8complex[0.82,1.5][0.82,1.5]0.8270.8270.820.82
rr[1.1,1.15][1.1,1.15]1.3751.3750.80.8
rr[1.15,1.175][1.15,1.175]1.361.360.80.8
rr[1.175,1.2][1.175,1.2]1.3461.3460.80.8
rr[1.2,1.225][1.2,1.225]1.3311.3310.80.8
rr[1.225,1.25][1.225,1.25]1.3181.3180.80.8
rr[1.25,1.275][1.25,1.275]1.3041.3040.80.8
rr[1.275,1.294][1.275,1.294]1.2941.2940.80.8
rr[1.294,1.5][1.294,1.5]1.2941.2940.80.8

Table 5. Parent implications: for the stated type and λ1∈[Aj,Bj]\lambda_1\in[A_j,B_j], λ′≥pj\lambda'\geq p_j and λ2≥rj\lambda_2\geq r_j. Each panel is read downwards, and each row is valid on the whole closed interval. The left panel gives the 33 real-first-zero rows; the right gives the 13 real-character/nonreal-zero rows and 12 nonreal-character rows.

Proof. Each entry is at most the largest of the following bounds, each valid on the whole interval [Aj,Bj][A_{j},B_{j}] for the type tt: (i) the applicable rows (t,c)(t,c) of Inputs 7.2 and 7.3 with t≥Bjt\ge B_{j}; (ii) the bound cc of an applicable row (t,c)(t,c) with c=tc=t (for λ1≤t\lambda_{1}\le t the row gives it, and for λ1>t\lambda_{1}>t the trivial bound gives λ′,λ2≥λ1>t=c\lambda',\lambda_{2}\ge\lambda_{1}>t=c), as for the rows (0.827,0.827)(0.827,0.827) of Table 2′2' and (1.099,1.099)(1.099,1.099) of Table 3 of [145] and (0.745,0.745)(0.745,0.745) of Table 7 of [51]; (iii) the bounds of Input 7.2(b),(d) and Input 7.3(d), which hold for every λ1\lambda_{1}; and (iv) the trivial bounds λ′,λ2≥λ1≥Aj\lambda',\lambda_{2}\ge\lambda_{1}\ge A_{j}. □\square

The positivity method

The use of nonnegative trigonometric polynomials to locate zeros goes back to de la Vallée Poussin [28]; it underlies all zero-free regions for Dirichlet LL-functions [43, 74, 97, 82, 68] and for ζ\zeta [67, 90], and, combined with smoothed explicit formulas, the tables of Heath-Brown and Xylouris and their refinements [79]. All the new zero-location results come from the following consequence of [145], Lemmas 3.1 and 3.4. Lemma 3.4 there is a “working version” of Input 3.4: if ff satisfies Conditions 1 and 2, χ≠χ0\chi\ne\chi_{0} and s∈R(9l)s\in R(9l), and A1A_{1} is the set of zeros ρ∈R(l)\rho\in R(l) of L(s,χ)L(s,\chi) with Re⁡ρ>Re⁡s\operatorname{Re}\rho>\operatorname{Re}s (with multiplicity, at most NN of them) and A2A_{2} is any set of at most NN zeros ρ∈R(l)\rho\in R(l) with Re⁡ρ≤Re⁡s\operatorname{Re}\rho\le\operatorname{Re}s, then for every ε>0\varepsilon>0 and q≥q0(f,ε,N)q\ge q_{0}(f,\varepsilon,N)

∑nΛ(n)Re⁡χ(n)nsf(log⁡nL)≤−L∑ρ∈A1∪A2Re⁡F((s−ρ)L)+ϕ(χ)2f(0)L+εL.(12)\sum_{n}\Lambda(n)\operatorname{Re}\frac{\chi(n)}{n^{s}}f\left(\frac{\log n}{\mathcal{L}}\right) \le -\mathcal{L}\sum_{\rho\in A_{1}\cup A_{2}}\operatorname{Re}F\left((s-\rho)\mathcal{L}\right) +\frac{\phi(\chi)}{2}f(0)\mathcal{L} +\varepsilon\mathcal{L}. \tag*{(12)}

Lemma 7.6 (Positivity method). Let ff satisfy Conditions 1 and 2, with Laplace transform FF. Fix a nonnegative anchor aa, and suppose a≤λ1a\le\lambda_{1}. Let finitely many triples (cj,χ(j),tj)(c_{j},\chi^{(j)},t_{j}) be given, with cj>0c_{j}>0, χ(j)\chi^{(j)} a character modulo qq and ∣tj∣≤9l|t_{j}|\le9l, such that

∑jcjRe⁡(χ(j)(n)n−itj)≥0for all n coprime to q.\sum_{j}c_{j}\operatorname{Re}\left(\chi^{(j)}(n)n^{-it_{j}}\right)\ge0 \qquad\text{for all }n\text{ coprime to }q.

For each nonprincipal χ(j)\chi^{(j)} let A2(j)A_{2}(j) be a set of at most NN zeros of L(s,χ(j))L(s,\chi^{(j)}) in R(l)R(l), with multiplicity. Then for every ε>0\varepsilon>0 and q≥q0q\ge q_{0},

∑χ(j)=χ0cjRe⁡F(−a+itjL)−∑χ(j)≠χ0cj∑ρ∈A2(j)Re⁡F((1+itj−ρ)L−a)+f(0)2∑χ(j)≠χ0cjϕ(χ(j))+ε≥0.\sum_{\chi^{(j)}=\chi_{0}}c_{j}\operatorname{Re}F(-a+it_{j}\mathcal{L}) - \sum_{\chi^{(j)}\ne\chi_{0}}c_{j} \sum_{\rho\in A_{2}(j)} \operatorname{Re}F\left((1+it_{j}-\rho)\mathcal{L}-a\right) + \frac{f(0)}{2} \sum_{\chi^{(j)}\ne\chi_{0}}c_{j}\phi\left(\chi^{(j)}\right) +\varepsilon \ge0.

Proof. Multiply the hypothesis by Λ(n)n−(1−a/L)f(log⁡n/L)≥0\Lambda(n)n^{-(1-a/\mathcal{L})}f(\log n/\mathcal{L})\ge0 and sum over nn (for (n,q)>1(n,q)>1 every term vanishes, the principal ones included). For principal χ(j)\chi^{(j)} use Input 3.3 at 1−a/L+itj1-a/\mathcal{L}+it_{j}. For the others use (12) at sj=1−a/L+itjs_{j}=1-a/\mathcal{L}+it_{j}, which lies in R(9l)R(9l); there A1=∅A_{1}=\varnothing, because every zero in R(l)R(l) of a nonprincipal LL-function has parameter at least λ1≥a\lambda_{1}\ge a. Divide by L\mathcal{L}. □\square

In every application the retained zeros satisfy λρ≥a\lambda_{\rho}\ge a, the transform FF is decreasing on R\mathbb{R}, and the terms that we cannot evaluate are bounded over all real heights with interval arithmetic. A necessary inequality between the parameters of the retained zeros results, and a case is excluded when that inequality fails with a positive margin. Since there are finitely many such inequalities, each with a fixed positive margin, one q0q_0 serves all of them.

A sharper conductor coefficient

Heath-Brown remarks [51] that his estimate for λ2\lambda_2 in the case of a real first character can be improved by replacing the conductor coefficient 1124\frac{11}{24} by 97216\frac{97}{216}; he describes the ingredients but gives no proof. We need a weighted form of this refinement, and we prove it here. For a modulus qq put

v=∏pe∥q,e≥3pe,θq=log⁡vlog⁡q∈[0,1].v=\prod_{\substack{p^e\lVert q,\,e\geq3}}p^e,\qquad\theta_q=\frac{\log v}{\log q}\in[0,1].

We also write rad⁡q=∏p∣qp\operatorname{rad}q=\prod_{p\mid q}p.

Lemma 7.7 (Growth bounds with conductor dependence). Let k≥3k \ge3 and ε>0\varepsilon> 0. For every nonprincipal character χ\chi modulo qq,

∣L(σ+it,χ)∣≪ε,kqϕq(χ)(1−σ)(1+k−1)+ε(1+∣t∣)(1−1k≤σ≤1+log⁡LL),\lvert L(\sigma+it,\chi)\rvert\ll_{\varepsilon,k} q^{\phi_q(\chi)(1-\sigma)(1+k^{-1})+\varepsilon}(1+\lvert t\rvert) \qquad \left(1-\frac{1}{k}\le\sigma\le1+\frac{\log\mathcal{L}}{\mathcal{L}}\right),

where ϕq(χ)=14log⁡(4rad⁡q)log⁡q\phi_q(\chi)=\frac{1}{4}\frac{\log(4\operatorname{rad}q)}{\log q} if χ\chi is real and ϕq(χ)=min⁡(13,14+34θq)\phi_q(\chi)=\min\left(\frac{1}{3},\frac{1}{4}+\frac{3}{4}\theta_q\right) otherwise.

Proof. Let χ\chi be induced by the primitive character χ∗\chi^{*} modulo q∗q^{*}; then q∗∣qq^{*}\mid q and ∣L(s,χ)∣≪qε∣L(s,χ∗)∣\lvert L(s,\chi)\rvert\ll q^{\varepsilon}\lvert L(s,\chi^{*})\rvert for σ≥12\sigma\ge\frac{1}{2} [51]. Heath-Brown proves ([51], (2.2)–(2.4)) that a bound ∑N<n≤N+Hχ∗(n)≪ε,kQ(k+1)/(4k2)+εq∗εH1−1/k\sum_{N<n\le N+H}\chi^{*}(n)\ll_{\varepsilon,k}Q^{(k+1)/(4k^{2})+\varepsilon}q^{*\varepsilon}H^{1-1/k} for 1≤H≤q∗1\le H\le q^{*} gives, by partial summation with the cut q1=Q(k+1)/(4k)q_1=Q^{(k+1)/(4k)} and the Pólya–Vinogradov inequality beyond q∗q^{*}, L(s,χ∗)≪Q(1−σ)(1+k−1)/4q∗ε(1+∣t∣)L(s,\chi^{*})\ll Q^{(1-\sigma)(1+k^{-1})/4}q^{*\varepsilon}(1+\lvert t\rvert) in the stated range (if q1>q∗q_1>q^{*} the trivial bound up to q∗q^{*} is already stronger). For σ≤1\sigma\le1 a bound of this form with a smaller QQ is stronger; for 1<σ≤1+log⁡L/L1<\sigma\le1+\log\mathcal{L}/\mathcal{L} every factor Q(1−σ)cQ^{(1-\sigma)c} with 1≤Q≤q41\le Q\le q^{4} and 0≤c≤10\le c\le1 lies between L−4\mathcal{L}^{-4} and 11, so all such bounds agree up to a factor qεq^{\varepsilon}.

If χ\chi is real, then χ∗\chi^{*} is a real primitive character, so q∗q^{*} is the absolute value of a fundamental discriminant and q∗≤4rad⁡(q∗)≤4rad⁡(q)q^{*}\le4\operatorname{rad}(q^{*})\le4\operatorname{rad}(q). Its order is 22, so by [51], which bounds the sums with the factor l3/(2k)l^{3/(2k)} for a character of order ll, the character-sum bound holds with Q=q∗Q=q^{*} for every kk. This gives the real case.

In general, [51] gives the bound with Q(k+1)/(4k2)=v∗(3k−1)/(4k2)q∗(k+1)/(4k2)≤(v∗3q∗)(k+1)/(4k2)Q^{(k+1)/(4k^{2})}=v^{*(3k-1)/(4k^{2})}q^{*(k+1)/(4k^{2})}\le(v^{*3}q^{*})^{(k+1)/(4k^{2})}, where v∗v^{*} is the cube-full part of q∗q^{*}; since q∗∣qq^{*}\mid q we have v∗∣vv^{*}\mid v and v∗3q∗≤v3qv^{*3}q^{*}\le v^{3}q. This gives the exponent 14(1+3log⁡vlog⁡q)=14+34θq\frac{1}{4}\left(1+\frac{3\log v}{\log q}\right)=\frac{1}{4}+\frac{3}{4}\theta_q. The exponent 13\frac{1}{3} is [51]. Taking the smaller bound gives the general case. (Bounds that are uniform in qq and tt together are studied in [48]; here the factor 1+∣t∣1+\lvert t\rvert is harmless, since ∣t∣≤L+1\lvert t\rvert\le\mathcal{L}+1.) □\square

Lemma 7.8 (The refined coefficient in the explicit formula). For every ε>0\varepsilon>0, Input 3.4 and (12) remain true, with a suitable q0q_0, when ϕ(χ)\phi(\chi) is replaced by ϕq(χ)+ε\phi_q(\chi)+\varepsilon, with the radius δ\delta that the proof of Input 3.4 provides for ϕ=13\phi=\frac{1}{3} (which serves all characters).

Proof. Both are deduced from Heath-Brown’s Lemma 3.1 [51, 145], whose proof uses the bound (2.5) of his Lemma 2.5 only through the inequality

log⁡∣L(s0+Reiϑ,χ)∣≤{ϕ(1+k−1)(1−σ0−Rcos⁡ϑ)+2ε}L\log\lvert L(s_0+Re^{i\vartheta},\chi)\rvert \le \left\{\phi(1+k^{-1})(1-\sigma_0-R\cos\vartheta)+2\varepsilon\right\}\mathcal{L}

for π2≤ϑ≤3π2\frac{\pi}{2}\le\vartheta\le\frac{3\pi}{2}, R≤1/kR\le1/k and large qq. By Lemma 7.7 this holds with ϕ=ϕq(χ)\phi=\phi_q(\chi), with an implied constant independent of qq and χ\chi. The remaining choices in that proof (k=3+[3ϕ/(2ε0)]k=3+\left[3\phi/(2\varepsilon_0)\right], R=1/kR=1/k, and the radius) may be made with the upper bound ϕ≤13\phi\le\frac{1}{3} in place of ϕ\phi, so they do not depend on qq. The deduction of Lemma 5.2 of [51] (Input 3.4) and of (12) from Lemma 3.1 is unchanged. □\square

Proposition 7.9 (A coupled saving in the conductor term). Let χ1\chi_1 be a real nonprincipal character and let χ2\chi_2 be a character with χ2k≠χ0\chi_2^{k}\ne\chi_0 and χ1χ2k≠χ0\chi_1\chi_2^{k}\ne\chi_0 for k∈Kk\in K, a finite set of positive integers. Let ck>0c_k>0 (k∈K)(k\in K) with Σ=∑kck≥19\Sigma=\sum_k c_k\ge\frac{1}{9}. Suppose the polynomial 1+∑k∈Kckcos⁡(kϑ)1+\sum_{k\in K}c_k\cos(k\vartheta) is nonnegative for every real ϑ\vartheta. In an application of Lemma 7.6 to (1+χ1(n))(1+∑k∈KckRe⁡((χ2(n)n−iγ2)k))(1+\chi_1(n))(1+\sum_{k\in K}c_k\operatorname{Re}((\chi_2(n)n^{-i\gamma_2})^k)), the conductor term may, for every fixed ε>0\varepsilon>0 and sufficiently large qq, be taken to be

(18+Σ31108)f(0)+ε.\left(\frac{1}{8}+\frac{\Sigma}{3}\frac{1}{108}\right)f(0)+\varepsilon.

Proof. The nonprincipal characters are χ1\chi_1 with coefficient 11, and χ2k\chi_2^{k} and χ1χ2k\chi_1\chi_2^{k}, each with coefficient ckc_k. By Lemma 7.8 the conductor term is at most f(0)2(ϕq(χ1)+2Σmax⁡χϕq(χ))+ε\frac{f(0)}{2}\left(\phi_q(\chi_1)+2\Sigma\max_{\chi}\phi_q(\chi)\right)+\varepsilon, the maximum being over the characters χ2k\chi_2^k and χ1χ2k\chi_1\chi_2^k. Since p≤(pe)1/3p\le(p^e)^{1/3} when e≥3e\ge3, rad⁡q≤qv−2/3\operatorname{rad}q\le qv^{-2/3}, so ϕq(χ1)≤14−θq6+o(1)\phi_q(\chi_1)\le\frac14-\frac{\theta_q}{6}+o(1), and ϕq(χ)≤min⁡(13,14+34θq)+o(1)\phi_q(\chi)\le\min(\frac13,\frac14+\frac34\theta_q)+o(1) for every nonprincipal χ\chi (for a real χ\chi because 14−θq6≤14\frac14-\frac{\theta_q}{6}\le\frac14). The conductor term is therefore at most K(θq)f(0)+2ε\mathcal{K}(\theta_q)f(0)+2\varepsilon for large qq, with

K(x)=18−x12+Σmin⁡(13,14+3x4)(0≤x≤1).\mathcal{K}(x)=\frac18-\frac{x}{12}+\Sigma\min\left(\frac13,\frac14+\frac{3x}{4}\right)\qquad(0\le x\le1).

On [0,19][0,\frac19] the slope of K\mathcal{K} is 34Σ−112≥0\frac34\Sigma-\frac1{12}\ge0, and on [19,1][\frac19,1] it is −112-\frac1{12}. Hence K(θq)≤K(19)=18+Σ3−1108\mathcal{K}(\theta_q)\le\mathcal{K}(\frac19)=\frac18+\frac{\Sigma}{3}-\frac1{108}. □\square

For Σ=1\Sigma=1 this is Heath-Brown’s 97216=1124−1108\frac{97}{216}=\frac{11}{24}-\frac1{108}.

Lower bounds for λ2\lambda_2 when the first zero is real.

Proposition 7.10 (Second-family bounds for a real first zero). Let χ1,ρ1\chi_1,\rho_1 be real with λ1∈[a,b]\lambda_1\in[a,b], where 0≤a≤b0\le a\le b, and let h≥bh\ge b. Let P(x)=1+c1x+c2(2x2−1)+c5(16x5−20x3+5x)P(x)=1+c_1x+c_2(2x^2-1)+c_5(16x^5-20x^3+5x), that is P(cos⁡ϑ)=1+c1cos⁡ϑ+c2cos⁡2ϑ+c5cos⁡5ϑP(\cos\vartheta)=1+c_1\cos\vartheta+c_2\cos2\vartheta+c_5\cos5\vartheta, with

(c1,c2,c5)=(1.41866466, 0.43287503, 0.01421037)1.000001,(c_1,c_2,c_5)=\frac{(1.41866466,\ 0.43287503,\ 0.01421037)}{1.000001},

so that P≥0P\ge0 on [−1,1][-1,1] and Σ=c1+c2+c5≥19\Sigma=c_1+c_2+c_5\ge\frac19. Choose f=fγf=f_\gamma from (11), with transform FF and fixed γ>0\gamma>0. If λ2≤h\lambda_2\le h, then for every ε>0\varepsilon>0 and all sufficiently large qq, the inequality corresponding to the character relation below must hold. The generic case means that χ2\chi_2 is nonreal and none of the relations in (ii)–(iii) holds.

(i) (generic) F(b−a)+c1F(h−a)≤F(−a)+(18+Σ3−1108)f(0)+εF(b-a)+c_1F(h-a)\le F(-a)+(\frac18+\frac{\Sigma}{3}-\frac1{108})f(0)+\varepsilon;

(ii) (χ22=χ1)(\chi_2^2=\chi_1) F(b−a)+c1F(h−a)≤F(−a)+(1+c28+Σ3+1c1+c5)f(0)+sup⁡ΔR(Δ)+εF(b-a)+c_1F(h-a)\le F(-a)+(\frac{1+c_2}{8}+\frac{\Sigma}{3}+\frac1{c_1+c_5})f(0)+\sup_\Delta R(\Delta)+\varepsilon, where R(Δ)=c2{Re⁡F(−a+iΔ)−Re⁡F(λ1−a+iΔ)}−c1Re⁡F(λ2−a+iΔ)R(\Delta)=c_2\{\operatorname{Re}F(-a+i\Delta)-\operatorname{Re}F(\lambda_1-a+i\Delta)\}-c_1\operatorname{Re}F(\lambda_2-a+i\Delta), the supremum over Δ∈R\Delta\in\mathbb{R}, λ1∈[a,b]\lambda_1\in[a,b] and λ2∈[a,h]\lambda_2\in[a,h];

(iii) (χ25∈{χ0,χ1})(\chi_2^5\in\{\chi_0,\chi_1\}) F(b−a)+c1F(h−a)≤(1+c5)F(−a)+(1+c58+c1+c24)f(0)+εF(b-a)+c_1F(h-a)\le(1+c_5)F(-a)+(\frac{1+c_5}{8}+\frac{c_1+c_2}{4})f(0)+\varepsilon;

(iv) (χ2 real)(\chi_2\ \mathrm{real}) F(b−a)+F(h−a)≤F(−a)+38f(0)+εF(b-a)+F(h-a)\le F(-a)+\frac38f(0)+\varepsilon, with a second test function (11).

Here χ2\chi_2 is a character of the second family. The test parameter in (iv) may be chosen independently of the one in (i)–(iii). If all four possible cases are excluded with a positive margin at ε=0\varepsilon=0, then λ2>h\lambda_2>h for all sufficiently large qq.

Proof. Put zn=χ2(n)n−iγ2z_n=\chi_2(n)n^{-i\gamma_2} and apply Lemma 7.6 with anchor aa to (1+χ1(n))P(Re⁡zn)≥0(1+\chi_1(n))P(\operatorname{Re}z_n)\ge0 (with Re⁡znk=Tk(Re⁡zn)\operatorname{Re}z_n^k=T_k(\operatorname{Re}z_n) for ∣zn∣=1|z_n|=1), retaining ρ1\rho_1 in the term of χ1\chi_1 and ρ2\rho_2 in the term c1Re⁡znc_1\operatorname{Re}z_n. The heights are at most 5l≤9l5l\le9l. The nonzero frequencies are 1, 2 and 5, and χ2≠χ1\chi_2\ne\chi_1. A product character is principal only if χ2\chi_2 is real, or χ22=χ1\chi_2^2=\chi_1 (order 4), or χ25∈{χ0,χ1}\chi_2^5\in\{\chi_0,\chi_1\} (order 5 or 10); the last two cannot occur together, since together they would give χ22=χ0\chi_2^2=\chi_0. Generic case: there is no principal product, and Proposition 7.9 gives (i), using that FF is decreasing, λ1≤b\lambda_1\le b and λ2≤h\lambda_2\le h. Case χ22=χ1\chi_2^2=\chi_1: then χ1χ22=χ0\chi_1\chi_2^2=\chi_0 at height 2γ22\gamma_2, χ22=χ1\chi_2^2=\chi_1 at height 2γ22\gamma_2 retains ρ1\rho_1, and χ1χ2=χ2‾\chi_1\chi_2=\overline{\chi_2} retains ρ2‾\overline{\rho_2}; with Δ=2μ2\Delta=2\mu_2 these give R(Δ)R(\Delta), while the characters of frequencies 1 and 5 are nonprincipal and are charged 13\frac13 (conservatively, since some of them have order 4). Case χ25∈{χ0,χ1}\chi_2^5\in\{\chi_0,\chi_1\}: exactly one term is principal, bounded by c5F(−a)c_5F(-a) since f≥0f\ge0, and every other character has order at most 10≤L10\le L, so ϕ=14\phi=\frac14. Real χ2\chi_2: apply the method to (1+χ1(n))(1+Re⁡zn)≥0(1+\chi_1(n))(1+\operatorname{Re}z_n)\ge0, whose three nonprincipal characters are real. □\square

The rows used are given in Table 6: 22 cells of width 0.0025 covering [0.700,0.755][0.700,0.755]. For each cell [a,b][a,b] we give hh and the test parameter γ\gamma of the generic case; each of the four inequalities fails by the margin shown (normalized by f(0)f(0)), so that λ2>h\lambda_2>h throughout the cell. The margins were computed with interval arithmetic: the correlation term R(Δ)R(\Delta) is enclosed on a grid in Δ∈[0,30]\Delta\in[0,30] with second-derivative interpolation errors in all variables, and by a closed-form bound for ∣Δ∣≥30|\Delta| \ge30. The nonnegativity of PP was verified at the points x=j/2000x=j/2000 with a second-derivative bound, and independently by exact Bernstein expansion. In addition, on cells inside [0.700,0.7025][0.700,0.7025] we use the degree-2 row P(cos⁡ϑ)=(78+cos⁡ϑ)2/8164P(\cos\vartheta)=(\frac{7}{8}+\cos\vartheta)^2/\frac{81}{64} without the conductor refinement, which gives λ2>0.762\lambda_2>0.762 (Proposition 7.11).

cell [a,b][a,b]hhγ\gammagenericorder 4order 5, 10real
[.7000,.7025][.7000,.7025].7793.77931.096241.096241.02⋅10−41.02\cdot10^{-4}.0804.0804.1280.1280.0289.0289
[.7025,.7050][.7025,.7050].7783.77831.095221.095225.86⋅10−55.86\cdot10^{-5}.0803.0803.1279.1279.0287.0287
[.7050,.7075][.7050,.7075].7772.77721.094191.094198.66⋅10−58.66\cdot10^{-5}.0804.0804.1280.1280.0285.0285
[.7075,.7100][.7075,.7100].7762.77621.093171.093174.92⋅10−54.92\cdot10^{-5}.0803.0803.1279.1279.0284.0284
[.7100,.7125][.7100,.7125].7751.77511.092161.092168.35⋅10−58.35\cdot10^{-5}.0804.0804.1280.1280.0282.0282
[.7125,.7150][.7125,.7150].7741.77411.091151.091155.21⋅10−55.21\cdot10^{-5}.0803.0803.1279.1279.0280.0280
[.7150,.7175][.7150,.7175].7730.77301.090141.090149.25⋅10−59.25\cdot10^{-5}.0804.0804.1279.1279.0278.0278
[.7175,.7200][.7175,.7200].7720.77201.089141.089146.71⋅10−56.71\cdot10^{-5}.0804.0804.1279.1279.0276.0276
[.7200,.7225][.7200,.7225].7710.77101.088141.088144.46⋅10−54.46\cdot10^{-5}.0803.0803.1279.1279.0275.0275
[.7225,.7250][.7225,.7250].7699.76991.087151.087159.41⋅10−59.41\cdot10^{-5}.0804.0804.1279.1279.0273.0273
[.7250,.7275][.7250,.7275].7689.76891.086151.086157.75⋅10−57.75\cdot10^{-5}.0804.0804.1279.1279.0271.0271
[.7275,.7300][.7275,.7300].7679.76791.085171.085176.39⋅10−56.39\cdot10^{-5}.0804.0804.1279.1279.0269.0269
[.7300,.7325][.7300,.7325].7669.76691.084181.084185.31⋅10−55.31\cdot10^{-5}.0804.0804.1279.1279.0267.0267
[.7325,.7350][.7325,.7350].7659.76591.083211.083214.53⋅10−54.53\cdot10^{-5}.0804.0804.1279.1279.0265.0265
[.7350,.7375][.7350,.7375].7649.76491.082231.082234.03⋅10−54.03\cdot10^{-5}.0804.0804.1278.1278.0263.0263
[.7375,.7400][.7375,.7400].7639.76391.081261.081263.83⋅10−53.83\cdot10^{-5}.0804.0804.1278.1278.0261.0261
[.7400,.7425][.7400,.7425].7629.76291.080291.080293.91⋅10−53.91\cdot10^{-5}.0804.0804.1278.1278.0260.0260
[.7425,.7450][.7425,.7450].7619.76191.079331.079334.28⋅10−54.28\cdot10^{-5}.0804.0804.1278.1278.0258.0258
[.7450,.7475][.7450,.7475].7609.76091.078371.078374.93⋅10−54.93\cdot10^{-5}.0804.0804.1278.1278.0256.0256
[.7475,.7500][.7475,.7500].7599.75991.077411.077415.87⋅10−55.87\cdot10^{-5}.0804.0804.1278.1278.0254.0254
[.7500,.7525][.7500,.7525].7589.75891.076451.076457.09⋅10−57.09\cdot10^{-5}.0804.0804.1278.1278.0252.0252
[.7525,.7550][.7525,.7550].7579.75791.075501.075508.60⋅10−58.60\cdot10^{-5}.0804.0804.1278.1278.0250.0250

Table 6. The degree-5 rows of Proposition 7.10: on each cell, λ2>h\lambda_2>h. The values of γ\gamma are rounded to five decimals (the exact rational values are in the data). The margins are normalized by f(0)f(0). The real case uses γ=0.82\gamma=0.82.

Proposition 7.11 (A second-family bound near 0.70). If χ1,ρ1\chi_1,\rho_1 are real and 0.700≤λ1≤0.70250.700\le\lambda_1\le0.7025, then λ2>0.762\lambda_2>0.762 for q≥q0q\ge q_0.

Proof. As in Proposition 7.10 with c1=11281c_1=\frac{112}{81}, c2=3281c_2=\frac{32}{81}, c5=0c_5=0, the printed conductor coefficients and γ=1.09\gamma=1.09 (and γ=0.82\gamma=0.82 for a real χ2\chi_2); the normalized margins of the generic, order-4 and real cases are at least 3.6⋅10−43.6\cdot10^{-4}, 0.0820.082 and 0.0330.033. □\square

Lower bounds for λ2\lambda_2 and λ′\lambda' when the first character is nonreal

For t>0t>0 put

Pt(eiϑ)=(t+cos⁡ϑ)2t2+12=1+c1cos⁡ϑ+c2cos⁡2ϑ,c1=2tt2+12,c2=12t2+12,P_t(e^{i\vartheta})=\frac{(t+\cos\vartheta)^2}{t^2+\frac{1}{2}} =1+c_1\cos\vartheta+c_2\cos2\vartheta,\qquad c_1=\frac{2t}{t^2+\frac{1}{2}},\qquad c_2=\frac{\frac{1}{2}}{t^2+\frac{1}{2}},

so that c1/c2=4tc_1/c_2=4t.

Proposition 7.12 (Second-family bounds for a nonreal first character). Let χ1\chi_1 be nonreal, let λ1∈[a,b]\lambda_1\in[a,b] with 0<a<b0<a<b, and let h>bh>b. Fix t>0t>0, let c1,c2c_1,c_2 be the coefficients of PtP_t above, and let f=fγf=f_\gamma with γ>0\gamma>0 and transform FF. Suppose λ2≤h\lambda_2\le h. For every ε>0\varepsilon>0 and sufficiently large qq, the following necessary inequalities hold:

(i) if χ2\chi_2 is nonreal and

sup⁡λ∈[a,h], y∈RRe⁡{F(−a+iy)−4tF(λ−a+iy)}≤16f(0),(13)\sup_{\lambda\in[a,h],\,y\in\mathbb{R}} \operatorname{Re}\{F(-a+iy)-4tF(\lambda-a+iy)\} \le\frac{1}{6}f(0), \tag*{(13)}

then

c1{F(b−a)+F(h−a)}≤F(−a)+(1+c1+c2)2−16f(0)+ε;c_1\{F(b-a)+F(h-a)\} \le F(-a)+\frac{(1+c_1+c_2)^2-1}{6}f(0)+\varepsilon;

(ii) if χ2\chi_{2} is real, choose t′>0t' > 0 and let c1′,c2′c'_{1}, c'_{2} be the coefficients of Pt′P_{t'}. If (13) holds with t′t' in place of tt and [a,b][a,b] in place of [a,h][a,h], then c1′F(b−a)+F(h−a)≤F(−a)+(18+c1′+c2′3)f(0)+εc'_{1}F(b-a)+F(h-a)\leq F(-a)+\left(\frac{1}{8}+\frac{c'_{1}+c'_{2}}{3}\right)f(0)+\varepsilon.

Proof. (i) Apply Lemma 7.6 to Pt(χ1(n)n−iγ1)Pt(χ2(n)n−iγ2)≥0P_{t}(\chi_{1}(n)n^{-i\gamma_{1}})P_{t}(\chi_{2}(n)n^{-i\gamma_{2}})\geq0, expanded as

∑w∈{−2,…,2}2a∣w1∣a∣w2∣χw(n)n−iw⋅γ,a0=1,a1=c12,a2=c22,χw=χ1w1χ2w2.\sum_{w\in\{-2,\ldots,2\}^{2}} a_{\lvert w_{1}\rvert}a_{\lvert w_{2}\rvert}\chi^{w}(n)n^{-iw\cdot\gamma}, \qquad a_{0}=1,\quad a_{1}=\frac{c_{1}}{2},\quad a_{2}=\frac{c_{2}}{2},\quad \chi^{w}=\chi_{1}^{w_{1}}\chi_{2}^{w_{2}}.

The indices (±1,0)(\pm1,0) retain ρ1\rho_{1} (or ρ1‾\overline{\rho_{1}}) and (0,±1)(0,\pm1) retain ρ2\rho_{2} (or ρ2‾\overline{\rho_{2}}). If no nonzero index is principal, the conductor term is at most 16((1+c1+c2)2−1)f(0)\frac{1}{6}((1+c_{1}+c_{2})^{2}-1)f(0). A nonzero principal index ww is not on an axis, since χ1\chi_{1} and χ2\chi_{2} are nonreal, and does not have ∣w1∣,∣w2∣≤1\lvert w_{1}\rvert,\lvert w_{2}\rvert\leq1, since χ2∉{χ1,χ1‾}\chi_{2}\notin\{\chi_{1},\overline{\chi_{1}}\}. So some ∣wi∣=2\lvert w_{i}\rvert=2; choose such ii, with i=1i=1 in case of a tie, and put T(w)=w−sgn⁡(wi)eiT(w)=w-\operatorname{sgn}(w_{i})e_{i}. Then χT(w)=χi−sgn⁡wi\chi^{T(w)}=\chi_{i}^{-\operatorname{sgn}w_{i}} is nonprincipal and not an axis index of length one, its coefficient is 4t4t times that of ww, and the zero of χi−sgn⁡wi\chi_{i}^{-\operatorname{sgn}w_{i}} selected at the axis index sits at relative height w⋅μw\cdot\mu. The map TT is injective: a collision could only come from (2s,σ)(2s,\sigma) and (s,2σ)(s,2\sigma) with s,σ=±1s,\sigma=\pm1, and if both were principal then χ1s=χ2σ\chi_{1}^{s}=\chi_{2}^{\sigma}, which is excluded. Retaining at T(w)T(w) the zero of χi−sgn⁡wi\chi_{i}^{-\operatorname{sgn}w_{i}}, the principal term at ww and the extra retained term together change the inequality by at most aw{Re⁡F(−a+iy)−16f(0)−4tRe⁡F(λi−a+iy)}≤0a_{w}\{\operatorname{Re}F(-a+iy)-\frac{1}{6}f(0)-4t\operatorname{Re}F(\lambda_{i}-a+iy)\}\leq0, by (13). (ii) Use Pt′(χ1(n)n−iγ1)(1+Re⁡χ2(n)n−iγ2)≥0P_{t'}(\chi_{1}(n)n^{-i\gamma_{1}})(1+\operatorname{Re}\chi_{2}(n)n^{-i\gamma_{2}})\geq0. Now χ2\chi_{2} is real, with coefficient 1 and φ=14\varphi=\frac{1}{4}; the only possible principal indices are (±2,±1)(\pm2,\pm1), which are matched to (±1,±1)(\pm1,\pm1) with target ρ1‾\overline{\rho_{1}} or ρ1\rho_{1}, as before. □\square

Proposition 7.13 (Additional-zero bounds for a nonreal first character). Let χ1\chi_{1} be nonreal of order at least 5, let λ1∈[a,b]\lambda_{1}\in[a,b] with 0≤a≤b0\leq a\leq b, and suppose λ′≤h\lambda'\leq h. Fix t,γ>0t,\gamma>0, let c1,c2c_{1},c_{2} be the coefficients of PtP_{t}, and let FF be the transform of f=fγf=f_{\gamma}. Choose an admissible ρ′\rho' of χ1\chi_{1} with parameter λ′\lambda' and height γ′\gamma'. Put x=λ1−ax=\lambda_{1}-a, y=λ′−ay=\lambda'-a and Δ=L(γ1−γ′)\Delta=\mathcal{L}(\gamma_{1}-\gamma'), and

Mt=(1+c1+c2)2−1−12(c12+c22),p1=12c12,p2=12c22,q1=c1(1+12c2),q2=12c1c2.M_{t}=(1+c_{1}+c_{2})^{2}-1-\frac{1}{2}(c_{1}^{2}+c_{2}^{2}),\qquad p_{1}=\frac{1}{2}c_{1}^{2},\qquad p_{2}=\frac{1}{2}c_{2}^{2},\qquad q_{1}=c_{1}\left(1+\frac{1}{2}c_{2}\right),\qquad q_{2}=\frac{1}{2}c_{1}c_{2}.

Then for every ε>0\varepsilon>0 and all sufficiently large qq,

c1[F(x)+F(y)]≤F(−a)+Mt6f(0)+R(Δ)+ε,c_{1}[F(x)+F(y)]\leq F(-a)+\frac{M_{t}}{6}f(0)+R(\Delta)+\varepsilon,

where

R(Δ)=p1Re⁡F(−a+iΔ)+p2Re⁡F(−a+2iΔ)−q1Re⁡{F(x+iΔ)+F(y+iΔ)}−q2Re⁡{F(x+2iΔ)+F(y+2iΔ)}.\begin{aligned} R(\Delta)={}&p_{1}\operatorname{Re}F(-a+i\Delta)+p_{2}\operatorname{Re}F(-a+2i\Delta)\\ &-q_{1}\operatorname{Re}\{F(x+i\Delta)+F(y+i\Delta)\}\\ &-q_{2}\operatorname{Re}\{F(x+2i\Delta)+F(y+2i\Delta)\}. \end{aligned}

Proof. Apply Lemma 7.6 to Pt(χ1(n)n−iγ1)Pt(χ1(n)n−iγ′)≥0P_{t}(\chi_{1}(n)n^{-i\gamma_{1}})P_{t}(\chi_{1}(n)n^{-i\gamma'})\geq0, whose terms are χ1k+l\chi_{1}^{k+l} at heights kγ1+lγ′k\gamma_{1}+l\gamma', ∣k∣,∣l∣≤2\lvert k\rvert,\lvert l\rvert\leq2. Since the order of χ1\chi_{1} is at least 5 and ∣k+l∣≤4\lvert k+l\rvert\leq4, the principal terms are exactly those with k+l=0k+l=0; they have total mass p1p_{1} at height ±Δ\pm\Delta and p2p_{2} at ±2Δ\pm2\Delta. In every term with k+l=±1k+l=\pm1 retain both ρ1\rho_{1} and ρ′\rho' (or their conjugates); a multiple zero is retained with its multiplicity. The terms (1,0),(0,1)(1,0),(0,1) and their conjugates give −c1[F(x)+F(y)]−c1Re⁡[F(x+iΔ)+F(y+iΔ)]-c_{1}[F(x)+F(y)]-c_{1}\operatorname{Re}[F(x+i\Delta)+F(y+i\Delta)], and the terms (2,−1),(−1,2)(2,-1),(-1,2) and their conjugates give −12c1c2Re⁡[F(x+iΔ)+F(y+iΔ)+F(x+2iΔ)+F(y+2iΔ)]-\frac{1}{2}c_{1}c_{2}\operatorname{Re}[F(x+i\Delta)+F(y+i\Delta)+F(x+2i\Delta)+F(y+2i\Delta)]. □\square

For χ1\chi_{1} of order 3 or 4, Input 7.3(b) gives λ′>1.35\lambda'>1.35 whenever λ1≤0.86\lambda_{1}\leq0.86, which is stronger than every bound in Proposition 7.13 that we use.

The rows used are the following, each on a cell of width 0.0025 and each with a positive normalized margin. There are 25 rows of Proposition 7.12, on cells covering [0.66,0.7225][0.66,0.7225], with conclusions from λ2>0.763\lambda_{2}>0.763 (on [0.66,0.6625][0.66,0.6625]) to λ2>0.725\lambda_{2}>0.725 (on [0.72,0.7225][0.72,0.7225]), parameters γ∈{1.25,1.3}\gamma\in\{1.25,1.3\}, t∈{0.9,0.95}t\in\{0.9,0.95\} and, for a real χ2\chi_{2}, test parameter γ∈{1.1,1.15}\gamma\in\{1.1,1.15\} and t′=0.85t'=0.85; their generic margins are at least 8.2⋅10−48.2 \cdot10^{-4}, and the margins in (13) at least 0.163. There are 69 rows of Proposition 7.13, on 69 of the 70 cells of width 0.0025 in [0.68,0.855][0.68,0.855] (all except [0.6975,0.70][0.6975,0.70]), with conclusions from λ′>0.975\lambda' > 0.975 (on [0.68,0.6825][0.68,0.6825]) to λ′>0.855\lambda' > 0.855 (on [0.8525,0.855][0.8525,0.855]), parameters γ∈{1.05,1.1,1.15}\gamma\in\{1.05,1.1,1.15\} and t∈{0.9,0.95}t \in\{0.9,0.95\}, and margins at least 1.5⋅10−31.5 \cdot10^{-3}. For these rows, if λ′≤h\lambda' \le h, the left side of Proposition 7.13 is at least c1[F(b−a)+F(h−a)]c_1[F(b-a)+F(h-a)], since FF is decreasing, and R(Δ)R(\Delta) is replaced by its supremum over all real Δ\Delta, x∈[0,b−a]x \in[0,b-a] and y∈[p−a,h−a]y \in[p-a,h-a], where p=max⁡(pj,a)≤λ′p=\max(p_j,a)\le\lambda' comes from the bound pjp_j of the parent row of the cell (Proposition 7.5; it comes from [145], Table 2′) and the trivial bound λ′≥λ1≥a\lambda' \ge\lambda_1 \ge a: p=0.96p=0.96, 0.93, 0.91, 0.89, 0.86, 0.84, 0.83 on the cells in [0.68,0.70][0.68,0.70], [0.70,0.72][0.70,0.72], …\ldots, [0.80,0.82][0.80,0.82], p=0.827p=0.827 on [0.82,0.8275][0.82,0.8275], and p=ap=a on the cells beyond. So each row proves that λ1∈[a,b]\lambda_1 \in[a,b] and λ′≥p\lambda' \ge p imply λ′>h\lambda' > h, and since λ′≥p\lambda' \ge p holds for every configuration with λ1∈[a,b]\lambda_1 \in[a,b], λ1∈[a,b]\lambda_1 \in[a,b] implies λ′>h\lambda' > h. The complete list is in Appendix A. On 11 further cells the inequalities of Proposition 7.12 are evaluated on the actual cell of the case tree.

A positivity exclusion.

Proposition 7.14 (Excluding a nearby real second character). Let χ1,ρ1\chi_1,\rho_1 be real with λ1∈[a,b]\lambda_1 \in[a,b], and let χ2≠χ1\chi_2 \ne\chi_1 be a real nonprincipal character with a zero of height at most 1 and parameter ν≤h\nu\le h. If

F(b−a)+F(h−a)−F(−a)−38f(0)>0F(b-a)+F(h-a)-F(-a)-\frac{3}{8}f(0)>0

for a test function (11), then this situation does not occur for q≥q0q \ge q_0.

Proof. Apply Lemma 7.6 to (1+χ1(n))(1+Re⁡χ2(n)n−iγ2)≥0(1+\chi_1(n))(1+\operatorname{Re}\chi_2(n)n^{-i\gamma_2})\ge0. Its three nonprincipal characters χ1\chi_1, χ2\chi_2, χ1χ2\chi_1\chi_2 are real, so φ=14\varphi=\frac{1}{4}; retain ρ1\rho_1 and the given zero of χ2\chi_2, which lies in R(l)R(l). This gives F(λ1−a)+F(ν−a)≤F(−a)+38f(0)+εF(\lambda_1-a)+F(\nu-a)\le F(-a)+\frac{3}{8}f(0)+\varepsilon. □

It excludes 197 nodes of the case tree (Section 2.4), all of type rr\mathrm{rr} with a reserved real family; the smallest normalized margin is 7.8⋅10−57.8 \cdot10^{-5}.

Applying the rows.

Lemma 7.15 (Preservation of configurations under location updates). In the notation of Definition 11.2, suppose that a result of this section proves λ2>h\lambda_2>h for every configuration of type t\mathbf{t} with λ1∈[a′,b′]\lambda_1 \in[a',b'], and let s\mathfrak{s} be a specification of type t\mathbf{t} with cell [a,b]⊆[a′,b′][a,b]\subseteq[a',b']. Then replacing l2l_2 and rr by max⁡(l2,h)\max(l_2,h) and max⁡(r,h)\max(r,h) (and lo⁡2\operatorname{lo}_2 by the new rr) does not change C(s)C(\mathfrak{s}), and C(s)=∅C(\mathfrak{s})=\varnothing if s\mathfrak{s} is reserved with hi⁡2≤h\operatorname{hi}_2\le h. Similarly a proof of λ′>h\lambda'>h allows pp and lo⁡′\operatorname{lo}' to be replaced by max⁡(⋅,h)\max(\mathord{\cdot},h), and makes C(s)C(\mathfrak{s}) empty if hi⁡′≤h\operatorname{hi}'\le h.

Proof. Every family other than the first has all its zeros in R(l)R(l) at parameter at least λ2>h\lambda_2>h; in particular ν(F)>h\nu(\mathcal{F})>h. □

The third family.

Proposition 7.16 (Lower bounds for the third family). Consider a configuration satisfying a specification of type t\mathbf{t} with λ1∈[a,b]\lambda_1 \in[a,b]. Put h=hi⁡2h=\operatorname{hi}_2 if a family is reserved and h=∞h=\infty otherwise; then λ2≤h\lambda_2\le h by Lemma 11.3. Each applicable rule below gives a strict lower bound for λ3\lambda_3, for all sufficiently large qq. Let λ3lo\lambda_3^{\mathrm{lo}} be the maximum of the finitely many bounds selected for the leaf. Then λ3>λ3lo\lambda_3>\lambda_3^{\mathrm{lo}}.

  1. 0.857 (Input 7.4(a));

  2. for t=rr\mathbf{t}=\mathrm{rr} and [a,b][a,b] inside [0.44,0.60][0.44,0.60], [0.60,0.70][0.60,0.70] or [0.70,0.80][0.70,0.80]: 1.176, 1.055, 0.952 (Input 7.4(c));

  3. for t≠rr\mathbf{t}\ne\mathrm{rr} and b≤tb\le t: the value cc of the smallest applicable row of Input 7.4(b);

  1. for t=complext=\mathrm{complex} and 0.44≤a≤b≤0.850.44 \le a \le b \le0.85: any tt with

F(t−b)+F(min⁡(t,h)−a)−F(−b)+F(0)−76f(0)>10−5f(0)F(t-b)+F(\min(t,h)-a)-F(-b)+F(0)-\frac{7}{6}f(0)>10^{-5}f(0)

(γ=54)(\gamma=\frac{5}{4});

  1. for t=rrt=\mathrm{rr}, 0.44≤a≤b≤0.800.44 \le a \le b \le0.80 and t∈[0.952,1.176]t \in[0.952,1.176]: any tt such that [a,min⁡(t,h)][a,\min(t,h)] is covered by intervals [c,c′][c,c'] with

F(−c′)−F(t−c′)−F(0)−F(b−c)+98f(0)<−10−5f(0)F(-c')-F(t-c')-F(0)-F(b-c)+\frac{9}{8}f(0)<-10^{-5}f(0)

(γ=2625)(\gamma=\frac{26}{25}).

Proof. (1)–(3) are printed. For (4), suppose λ3≤t\lambda_{3} \le t; then λ2≤min⁡(t,h)\lambda_{2} \le\min(t,h). The function λ↦F(−λ)−F(t−λ)=∫f(u)eλu(1−e−tu) du\lambda\mapsto F(-\lambda)-F(t-\lambda)=\int f(u)e^{\lambda u}(1-e^{-tu})\,\mathrm{d}u is increasing, so F(−λ1)−F(λ3−λ1)≤F(−λ1)−F(t−λ1)≤F(−b)−F(t−b)F(-\lambda_{1})-F(\lambda_{3}-\lambda_{1}) \le F(-\lambda_{1})-F(t-\lambda_{1}) \le F(-b)-F(t-b), and F(λ2−λ1)≥F(min⁡(t,h)−a)F(\lambda_{2}-\lambda_{1}) \ge F(\min(t,h)-a). So the right side of Input 7.4(d) is below −10−5f(0)+ε<0-10^{-5}f(0)+\varepsilon<0, a contradiction; the side condition [145] (4.29) holds by our verification. For (5), suppose λ3≤t≤1.176\lambda_{3} \le t \le1.176. Then λ2∈[a,min⁡(t,h)]⊆[0.44,1.176]\lambda_{2} \in[a,\min(t,h)] \subseteq[0.44,1.176] and λ1∈[0.44,0.80]\lambda_{1} \in[0.44,0.80], so Input 7.4(e) applies. If λ2∈[c,c′]\lambda_{2} \in[c,c'], the same monotonicity gives F(−λ2)−F(λ3−λ2)≤F(−c′)−F(t−c′)F(-\lambda_{2})-F(\lambda_{3}-\lambda_{2}) \le F(-c')-F(t-c') and F(λ1−λ2)≥F(b−c)F(\lambda_{1}-\lambda_{2}) \ge F(b-c), and the right side of (e) is negative for small ε\varepsilon, a contradiction; the side condition [145] (4.34) holds by our verification. □\square

For example, on the rr\mathrm{rr} cell [0.7025,0.705][0.7025,0.705] rule (5) gives λ3>1.0502\lambda_{3}>1.0502 without a reservation, and λ3>1.0888\lambda_{3}>1.0888 with a reserved family below 0.850.85; Table 10 gives 0.9520.952. Rules (4) and (5) are evaluated with interval arithmetic, and a value tt is accepted only with a normalized margin exceeding 10−510^{-5}.

The graded near-density lemma

The purpose of this section is to constrain many characters at once while retaining the location of each selected zero. We first bound a quadratic form in smoothed prime sums, called responses (Theorem 8.4). We then bound each response from below using selected zeros (Lemma 8.7). Finally, an elementary optimization argument introduces the single threshold used in a near row (Proposition 8.8). The analytic inputs are stated in Section 3; the scope of the formalization is described in Section 14.3. Section 8.3 explains the idea of the proof, and Example 8.9 shows what the resulting constraint can do that separate counts cannot.

Data

The detector determines the response to a zero. The Gram test and sieve weight control correlations between characters. The anchors may vary between entries, but all test functions and allowed offsets are fixed before qq tends to infinity.

Definition 8.1 (Near-density data). A near datum consists of the following fixed objects; Conditions 1 and 2 are those of Definition 3.1.

  • A detector ff satisfying Condition 1, with transform FF.

  • A Gram test gg satisfying Conditions 1 and 2, with transform GG.

  • Real numbers s1s_{1} and ss, called the safe anchor and the reference anchor.

  • A sieve weight: cells t0<t1<⋯<tmt_{0}<t_{1}<\cdots<t_{m} with t0>13+2ε′t_{0}>\frac{1}{3}+2\varepsilon' for a fixed ε′>0\varepsilon'>0, and heights h0,…,hm−1≥0h_{0},\ldots,h_{m-1}\ge0. We put

Ξ(t)=∑k<mhk1[tk,tk+1)(t),ω(t)=g(t)es1t+Ξ(t),\Xi(t)=\sum_{k<m}h_{k}\mathbf{1}_{[t_{k},t_{k+1})}(t),\qquad \omega(t)=g(t)e^{s_{1}t}+\Xi(t),

and require that ω\omega be bounded below by a positive constant on the support of ff.

  • A finite set Δ⊂[0,∞)\Delta\subset[0,\infty) of offsets.

We put ϑk=12(tk−13)−ε′\vartheta_{k}=\frac{1}{2}(t_{k}-\frac{1}{3})-\varepsilon' (the sieve levels), so 0<ϑk<12tk0<\vartheta_{k}<\frac{1}{2}t_{k}, and

IB=∫0∞e2stf(t)2ω(t) dt,D(δ)=∫0∞g(t)e(s1−2δ)t dt+∑k<mhke−2δtktk+12−tk22ϑk,d1=g(0)6.I_{B}=\int_{0}^{\infty}\frac{e^{2st}f(t)^{2}}{\omega(t)}\,\mathrm{d}t,\qquad D(\delta)=\int_{0}^{\infty}g(t)e^{(s_{1}-2\delta)t}\,\mathrm{d}t+\sum_{k<m}h_{k}e^{-2\delta t_{k}}\frac{t_{k+1}^{2}-t_{k}^{2}}{2\vartheta_{k}},\qquad d_{1}=\frac{g(0)}{6}.

Here the integrand of IBI_B is taken to be 0 where f=0f=0. We require IB>0I_B>0.

The number IBI_B is the first Cauchy–Schwarz factor, D(δ)D(\delta) bounds a diagonal entry at offset δ\delta, and d1d_1 bounds the correlation between distinct characters. Heath-Brown’s Lemma 12.1 is the formal analogue of the case g=fg=f, s=s1=λ1s=s_1=\lambda_1, Ξ=0\Xi=0 and all offsets 0 (in that case ω=fes1t\omega=fe^{s_1t} is not bounded below on the support of ff, and Heath-Brown writes f=f⋅ff=\sqrt{f}\cdot\sqrt{f} instead).

Definition 8.2 (Entries and responses). An entry is a triple (χj,γj,δj)(\chi_j,\gamma_j,\delta_j) of a nonprincipal character modulo qq, a real height γj\gamma_j and an offset δj∈Δ\delta_j\in\Delta. Its response anchor is sj=s−δjs_j=s-\delta_j, and its response is

Rj=−1L∑n≥1Λ(n)Re⁡(χj(n)n−(1−sj/L+iγj))f(log⁡nL).R_j=-\frac{1}{\mathcal{L}}\sum_{n\geq1}\Lambda(n)\operatorname{Re}\left(\chi_j(n)n^{-(1-s_j/\mathcal{L}+i\gamma_j)}\right)f\left(\frac{\log n}{\mathcal{L}}\right).

For a family of entries, the pair excess of entry jj is

ej=∑k≠jχk=χj(Re⁡G(−σjk+i(γj−γk)L)−d1)+,σjk=s1−δj−δk.e_j=\sum_{\substack{k\ne j\\ \chi_k=\chi_j}}\left(\operatorname{Re}G\left(-\sigma_{jk}+i(\gamma_j-\gamma_k)\mathcal{L}\right)-d_1\right)_{+},\qquad\sigma_{jk}=s_1-\delta_j-\delta_k.

Definition 8.3 (Safe anchor). For Γ′>0\Gamma'>0, the safe-anchor hypothesis SA⁡(s1,Γ′)\operatorname{SA}(s_1,\Gamma') says that no nonprincipal L(s,χ)L(s,\chi) modulo qq has a zero ρ\rho with ∣Im⁡ρ∣≤Γ′|\operatorname{Im}\rho|\leq\Gamma' and Re⁡ρ>1−s1/L\operatorname{Re}\rho>1-s_1/\mathcal{L}.

The pair excess eje_j is needed only when two entries use the same character. It records the part of their correlation that exceeds the bound d1d_1 for distinct characters. The safe-anchor hypothesis ensures that discarded zero terms in the Gram estimate have the required sign.

In the application s1≤λ1s_1\leq\lambda_1. Then SA⁡(s1,2l+1)\operatorname{SA}(s_1,2l+1) holds for large qq, by the definition of λ1\lambda_1 and Input 3.5: a zero with ∣γ∣≤2l+1≤10l|\gamma|\leq2l+1\leq10l and Re⁡ρ>1−s1/L\operatorname{Re}\rho>1-s_1/\mathcal{L} lies in R(10l)R(10l), hence in R(l)R(l), so its parameter is at least λ1≥s1\lambda_1\geq s_1.

Statement

Theorem 8.4 (Graded near-density lemma). Assume Inputs 3.3, 3.4, 3.7 and 3.8, Mertens’ estimate, and the divisor bound (the number of divisors of qq is Oε(qε)O_\varepsilon(q^\varepsilon)). Fix a near datum as in Definition 8.1 and a tolerance η′>0\eta'>0. There is a threshold q0q_0, depending only on this fixed data, with the following property.

For q≥q0q\geq q_0, let JJ be a finite index set and let (χj,γj,δj)j∈J(\chi_j,\gamma_j,\delta_j)_{j\in J} be entries satisfying

  1. ∣γj∣≤Γ|\gamma_j|\leq\Gamma for some 0≤Γ≤L/30\leq\Gamma\leq\mathcal{L}/3;

  2. SA⁡(s1,2Γ+1)\operatorname{SA}(s_1,2\Gamma+1);

  3. the characters χj\chi_j are pairwise distinct whenever Ξ≠0\Xi\ne0.

Then, for every choice of real coefficients aj≥0a_j\geq0,

(∑jajRj)+2≤(1+η′)IB[∑j(D(δj)−d1+ej)aj2+(d1+η′)(∑jaj)2].\left(\sum_j a_jR_j\right)_{+}^{2}\leq(1+\eta')I_B\left[\sum_j\left(D(\delta_j)-d_1+e_j\right)a_j^{2}+(d_1+\eta')\left(\sum_j a_j\right)^{2}\right].

The responses RjR_j and pair excesses eje_j are those of Definition 8.2. The threshold q0q_0 is uniform in the number of entries, their characters and heights, and the coefficients aja_j.

Two special cases of the pair excess are used.

  1. Separated heights. If χk=χj\chi_k=\chi_j implies ∣γj−γk∣L≥M|\gamma_j-\gamma_k|\mathcal{L}\geq M for k≠jk\ne j, and MM is so large that Re⁡G(−σ+iY)≤d1\operatorname{Re}G(-\sigma+iY)\leq d_1 for ∣Y∣≥M|Y|\geq M and σ\sigma in the finite set of values σjk\sigma_{jk}, then ej=0e_j=0 (Lemma 8.6; this holds once g(0)>0g(0)>0).

  2. A controlled pair. If entry jj shares its character with exactly one other entry kk, with δj=δk=0\delta_j=\delta_k=0 and ∣(γj−γk)L∣|(\gamma_j-\gamma_k)\mathcal{L}| in a range on which Re⁡G(−s1+iy)≤chi\operatorname{Re}G(-s_1+iy)\leq c^{\mathrm{hi}}, then ej≤(chi−d1)+e_j\leq(c^{\mathrm{hi}}-d_1)_{+}.

The idea of the proof

Each response RjR_j is a smoothed sum over primes, twisted by χj(n)n−iγj\chi_j(n)n^{-i\gamma_j} and evaluated at the point 1−sj/L+iγj1-s_j/\mathcal{L}+i\gamma_j. By the explicit formula (Lemma 8.7 below), every zero ρ\rho of L(s,χj)L(s,\chi_j) near 1+iγj1+i\gamma_j contributes about Re⁡F((λρ−sj)+i(γjL−μρ))\operatorname{Re} F((\lambda_\rho-s_j)+i(\gamma_j\mathcal{L}-\mu_\rho)) to RjR_j, and this is large when λρ\lambda_\rho is small. So characters with zeros close to 1 have large responses.

On the other hand, all the responses are correlations of one vector, the weighted primes, with the twisted vectors (χj(n)n−iγj)n(\chi_j(n)n^{-i\gamma_j})_n. By the Cauchy–Schwarz inequality a combination ∑jajRj\sum_j a_jR_j is at most the length of the prime vector, which gives the factor IBI_B, times the length of ∑jajχj(n)n−iγj\sum_j a_j\chi_j(n)n^{-i\gamma_j} measured with the weight ω\omega. The latter is a Gram form in the characters. Its diagonal entries are the numbers D(δj)D(\delta_j), and distinct characters are nearly orthogonal on the primes: their correlation is at most d1=g(0)/6d_1=g(0)/6. Indeed, by the explicit formula the correlation is governed by the zeros of the quotient character χjχk‾\chi_j\overline{\chi_k} to the right of the safe anchor, and the safe-anchor hypothesis says that there are none. Hence only a bounded number of characters can have large responses.

Heath-Brown’s Lemma 12.1 is the case in which all entries are tested at the same anchor. The grading gives each entry its own anchor sj=s−δjs_j=s-\delta_j. An entry tested at a larger offset δj\delta_j has a smaller diagonal D(δj)D(\delta_j), because of the factor e−2δjte^{-2\delta_jt}, so characters whose zeros are further from 1 are charged less. The sieve weight Ξ\Xi adds Selberg-sieve majorants of the primes to ω\omega on the ranges [qtk,qtk+1)[q^{t_k},q^{t_{k+1}}). This makes ω\omega larger and the first factor IBI_B smaller. The price is a sieve contribution to the diagonal, computed with Graham’s asymptotic, and off-diagonal sums of the quotient characters over long intervals. These are small by Burgess’s estimate provided the quotient characters are nonprincipal, that is, provided the characters are distinct.

A Burgess bound for all nonprincipal characters

Lemma 8.5 (Burgess bound for imprimitive characters and long intervals). Assume Input 3.7. For every ε>0\varepsilon>0 there is CC such that for every nonprincipal character χ\chi modulo qq and all real N≥0N\geq0 and H≥1H\geq1,

∣∑N<n≤N+Hχ(n)∣≤Cq1/9+εmin⁡(H,q)2/3.\left|\sum_{N<n\leq N+H}\chi(n)\right| \leq Cq^{1/9+\varepsilon}\min(H,q)^{2/3}.

Proof. Let χ\chi be induced by the primitive character χ∗\chi^* modulo q∗>1q^*>1, and let rr be the product of the primes dividing qq but not q∗q^*. Then χ(n)=χ∗(n)∑e∣(n,r)μ(e)\chi(n)=\chi^*(n)\sum_{e\mid(n,r)}\mu(e), so the sum is ∑e∣rμ(e)χ∗(e)∑N/e<m≤(N+H)/eχ∗(m)\sum_{e\mid r}\mu(e)\chi^*(e)\sum_{N/e<m\leq(N+H)/e}\chi^*(m). The sum only depends on the integers in the range, so we may take the endpoints real. In each inner sum remove complete periods of χ∗\chi^*, over which χ∗\chi^* sums to 0 because χ∗\chi^* is nonprincipal; what remains is a sum over at most min⁡(H/e,q∗)\min(H/e,q^*) consecutive integers, which by periodicity may be shifted to start at a point ≥1\geq1. Input 3.7 bounds it by Cq∗1/9+εmin⁡(H/e,q∗)2/3≤Cq1/9+εmin⁡(H,q)2/3Cq^{*1/9+\varepsilon}\min(H/e,q^*)^{2/3}\leq Cq^{1/9+\varepsilon}\min(H,q)^{2/3} (for an interval of length less than 1 the bound is trivial). By the divisor bound, the number of ee is Oε(qε)O_\varepsilon(q^\varepsilon). □\square

Proof of Theorem 8.4

Write tn=log⁡n/Lt_n=\log n/\mathcal{L}. For (n,q)=1(n,q)=1 put

Yn=∑jaje−δjtnχj(n)n−iγj.Y_n=\sum_j a_je^{-\delta_jt_n}\chi_j(n)n^{-i\gamma_j}.

Since n−(1−sj/L+iγj)=n−1estne−δjtnn−iγjn^{-(1-s_j/\mathcal{L}+i\gamma_j)}=n^{-1}e^{st_n}e^{-\delta_jt_n}n^{-i\gamma_j}, and χj(n)=0\chi_j(n)=0 for (n,q)>1(n,q)>1,

∑jajRj=−Re⁡1L∑(n,q)=1Λ(n)nestnf(tn)Yn.\sum_j a_jR_j=-\operatorname{Re}\frac{1}{\mathcal{L}}\sum_{(n,q)=1}\frac{\Lambda(n)}{n}e^{st_n}f(t_n)Y_n.

Cauchy–Schwarz. Since ω>0\omega>0 on the support of ff and ω≥0\omega\ge0 everywhere,

(∑jajRj)+2≤(1L∑nΛ(n)ne2sntnf(tn)2ω(tn))(1L∑(n,q)=1Λ(n)nω(tn)∣Yn∣2).(14)\left(\sum_j a_jR_j\right)_{+}^{2} \le \left(\frac{1}{\mathcal{L}}\sum_n\frac{\Lambda(n)}{n}\frac{e^{2s_nt_n}f(t_n)^2}{\omega(t_n)}\right) \left(\frac{1}{\mathcal{L}}\sum_{(n,q)=1}\frac{\Lambda(n)}{n}\omega(t_n)|Y_n|^2\right). \tag*{(14)}

The first factor. The function φB(t)=e2s1tf(t)2/ω(t)\varphi_B(t)=e^{2s_1t}f(t)^2/\omega(t) is bounded, because ω\omega is bounded below on the support of ff; it vanishes outside the support of ff and is continuous except at finitely many points, so it is Riemann integrable. Given εB>0\varepsilon_B>0 choose a partition 0=u0<⋯<uK0=u_0<\cdots<u_K of an interval containing the support of ff whose upper Riemann sum is at most IB+εB/2I_B+\varepsilon_B/2. By Mertens’ estimate ∑n≤xΛ(n)/n=log⁡x+O(1)\sum_{n\le x}\Lambda(n)/n=\log x+O(1) (see [112] for explicit forms) we have L−1∑qui<n≤qui+1Λ(n)/n=ui+1−ui+O(1/L)\mathcal{L}^{-1}\sum_{q^{u_i}<n\le q^{u_{i+1}}}\Lambda(n)/n=u_{i+1}-u_i+O(1/\mathcal{L}), so the first factor is at most IB+εBI_B+\varepsilon_B for large qq.

The second factor. Split it as Q1+Q2Q_1+Q_2 according to ω=ges1t+Ξ\omega=ge^{s_1t}+\Xi. Expanding ∣Yn∣2|Y_n|^2 and writing χjk=χjχk‾\chi_{jk}=\chi_j\overline{\chi_k},

Q1=∑j,kajakXjk,Xjk=Re⁡1L∑(n,q)=1Λ(n)χjk(n)n−(1−σjk/L)−i(γj−γk)g(tn).Q_1=\sum_{j,k}a_ja_kX_{jk}, \qquad X_{jk}=\operatorname{Re}\frac{1}{\mathcal{L}} \sum_{(n,q)=1}\Lambda(n)\chi_{jk}(n)n^{-(1-\sigma_{jk}/\mathcal{L})-i(\gamma_j-\gamma_k)}g(t_n).

All the points wjk=1−σjk/L+i(γj−γk)w_{jk}=1-\sigma_{jk}/\mathcal{L}+i(\gamma_j-\gamma_k) involved satisfy ∣Re⁡wjk−1∣≤C/L|\operatorname{Re}w_{jk}-1|\le C/\mathcal{L} and ∣Im⁡wjk∣≤2Γ≤L|\operatorname{Im}w_{jk}|\le2\Gamma\le\mathcal{L}, so Inputs 3.3 and 3.4 apply uniformly. The error terms o(1)o(1) below are uniform in the characters, heights and entries.

  • Diagonal, j=kj=k. Here χjj=χ0\chi_{jj}=\chi_0 and the height is 0. By Input 3.3,

    Xjj=∫0∞g(t)e(s1−2δj)t dt+o(1).X_{jj}=\int_0^\infty g(t)e^{(s_1-2\delta_j)t}\,\mathrm{d}t+o(1).
  • Distinct characters. Here χjk\chi_{jk} is nonprincipal. By Input 3.4,

    Xjk≤−∑ρRe⁡G((wjk−ρ)L)+φ(χjk)2g(0)+o(1),X_{jk}\le-\sum_\rho\operatorname{Re}G((w_{jk}-\rho)\mathcal{L})+\frac{\varphi(\chi_{jk})}{2}g(0)+o(1),

    the sum running over the zeros of L(s,χjk)L(s,\chi_{jk}) in a disc of radius δ<1\delta<1 about 1+i(γj−γk)1+i(\gamma_j-\gamma_k). Such a zero has ∣Im⁡ρ∣≤2Γ+1|\operatorname{Im}\rho|\le2\Gamma+1, so by SA(s1,2Γ+1)\mathrm{SA}(s_1,2\Gamma+1) it has Re⁡ρ≤1−s1/L≤1−σjk/L=Re⁡wjk\operatorname{Re}\rho\le1-s_1/\mathcal{L}\le1-\sigma_{jk}/\mathcal{L}=\operatorname{Re}w_{jk}. Hence Re⁡(wjk−ρ)L≥0\operatorname{Re}(w_{jk}-\rho)\mathcal{L}\ge0, and each term Re⁡G(⋅)\operatorname{Re}G(\cdot) is ≥0\ge0 by Condition 2. Dropping them and using φ≤13\varphi\le\frac13, g(0)≥0g(0)\ge0, we get Xjk≤d1+o(1)X_{jk}\le d_1+o(1). When σjk<0\sigma_{jk}<0 the point wjkw_{jk} lies to the right of 1, and the same argument applies.

  • The same character, j≠kj\ne k. Here χjk=χ0\chi_{jk}=\chi_0 at the height γj−γk\gamma_j-\gamma_k. By Input 3.3,

    Xjk=Re⁡G(−σjk+i(γj−γk)L)+o(1).X_{jk}=\operatorname{Re}G(-\sigma_{jk}+i(\gamma_j-\gamma_k)\mathcal{L})+o(1).

For a≥0a\ge0 the off-diagonal part is therefore at most

∑j≠kajak(d1+ϵjk)+o(1)(∑jaj)2,ϵjk=1χj=χk(Re⁡G(−σjk+i(γj−γk)L)−d1)+.\sum_{j\ne k}a_ja_k(d_1+\epsilon_{jk}) +o(1)\left(\sum_j a_j\right)^2, \qquad \epsilon_{jk} =\mathbf{1}_{\chi_j=\chi_k} \left(\operatorname{Re}G(-\sigma_{jk}+i(\gamma_j-\gamma_k)\mathcal{L})-d_1\right)_{+}.

Since ϵjk=ϵkj\epsilon_{jk}=\epsilon_{kj} (Re⁡G\operatorname{Re}G is even in the imaginary part, gg being real) and 2ajak≤aj2+ak22a_ja_k\le a_j^2+a_k^2, we have ∑j≠kajakϵjk≤∑jejaj2\sum_{j\ne k}a_ja_k\epsilon_{jk}\le\sum_je_ja_j^2. Also ∑j≠kajak=(∑jaj)2−∑jaj2\sum_{j\ne k}a_ja_k=(\sum_ja_j)^2-\sum_ja_j^2. Hence

Q1≤∑j(∫0∞g(t)e(s1−2δj)t dt−d1+ej)aj2+(d1+o(1))(∑jaj)2.(15)Q_1\le \sum_j\left(\int_0^\infty g(t)e^{(s_1-2\delta_j)t}\,\mathrm{d}t-d_1+e_j\right)a_j^2 +(d_1+o(1))\left(\sum_j a_j\right)^2. \tag*{(15)}

The sieve part Q2Q_2. It is present only when Ξ≠0\Xi\ne0, so the characters are distinct and all pair excesses vanish. Fix a cell kk, let zk=qϑkz_k=q^{\vartheta_k} and let

ξd=μ(d)log⁡(zk/d)log⁡zk(d≤zk),νk(n)=(∑d∣n, d≤zkξd)2≥0.\xi_d=\mu(d)\frac{\log(z_k/d)}{\log z_k}\quad(d\le z_k), \qquad \nu_k(n)=\left(\sum_{\substack{d\mid n,\ d\le z_k}}\xi_d\right)^2\ge0.
  • Majorant. If pp is prime with tp≥tkt_p \ge t_k then p≥qtk>zkp \ge q^{t_k} > z_k, so the only divisor of pp up to zkz_k is 1 and νk(p)=ξ12=1\nu_k(p)=\xi_1^2=1. Hence Λ(p)=(log⁡p)νk(p)\Lambda(p)=(\log p)\nu_k(p). For nn not a prime power, Λ(n)=0≤(log⁡n)νk(n)\Lambda(n)=0\le(\log n)\nu_k(n). Prime powers pep^e with e≥2e\ge2 and tpe≥tkt_{p^e}\ge t_k contribute ≪Lq−tk/2(∑jaj)2=o(1)(∑jaj)2\ll\mathcal{L}q^{-t_k/2}(\sum_j a_j)^2=o(1)(\sum_j a_j)^2. So the contribution of cell kk to Q2Q_2 is at most

hkL∑tn∈[tk,tk+1)log⁡nnνk(n)∣Yn∣2+o(1)(∑jaj)2,\frac{h_k}{\mathcal{L}} \sum_{t_n\in[t_k,t_{k+1})} \frac{\log n}{n}\nu_k(n)|Y_n|^2 +o(1)\left(\sum_j a_j\right)^2,

where now ∣Yn∣2|Y_n|^2 is expanded without the coprimality condition, with χj(n)=0\chi_j(n)=0 for (n,q)>1(n,q)>1.

  • Diagonal. The terms j=k′j=k' of ∣Yn∣2|Y_n|^2 are at most aj2e−2δjtna_j^2e^{-2\delta_jt_n} with χjχj‾≤1\chi_j\overline{\chi_j}\le1. By Input 3.8, ∑n≤xνk(n)=x/log⁡zk+O(x/log⁡2zk)\sum_{n\le x}\nu_k(n)=x/\log z_k+O(x/\log^2z_k) for x≥zkx\ge z_k. Partial summation gives

1L∑tn∈[tk,tk+1)log⁡nnνk(n)e−2δjtn=∫tktk+1tϑke−2δjt dt+O(1L)≤e−2δjtktk+12−tk22ϑk+O(1L).\frac{1}{\mathcal{L}} \sum_{t_n\in[t_k,t_{k+1})} \frac{\log n}{n}\nu_k(n)e^{-2\delta_jt_n} = \int_{t_k}^{t_{k+1}}\frac{t}{\vartheta_k}e^{-2\delta_jt}\,\mathrm{d}t +O\left(\frac{1}{\mathcal{L}}\right) \le e^{-2\delta_jt_k}\frac{t_{k+1}^2-t_k^2}{2\vartheta_k} +O\left(\frac{1}{\mathcal{L}}\right).
  • Off-diagonal. For j≠k′j\ne k' the character χjk′=χjχk′‾\chi_{jk'}=\chi_j\overline{\chi_{k'}} is nonprincipal, since the characters are distinct. Put W(n)=(log⁡n)n−1e−(δj+δk′)tnn−i(γj−γk′)W(n)=(\log n)n^{-1}e^{-(\delta_j+\delta_{k'})t_n}n^{-i(\gamma_j-\gamma_{k'})}. Opening νk\nu_k, the term is

1L∑d,e≤zkξdξeχjk′([d,e])∑mχjk′(m)W([d,e]m),\frac{1}{\mathcal{L}} \sum_{d,e\le z_k}\xi_d\xi_e\chi_{jk'}([d,e]) \sum_m\chi_{jk'}(m)W([d,e]m),

where mm runs over an interval [M0,M1)[M_0,M_1) with M0=qtk/[d,e]≥qtk−2ϑk=q1/3+2ε′M_0=q^{t_k}/[d,e]\ge q^{t_k-2\vartheta_k}=q^{1/3+2\varepsilon'}. By Lemma 8.5, ∣∑M0≤m≤uχjk′(m)∣≪u2/3q1/9+ε|\sum_{M_0\le m\le u}\chi_{jk'}(m)|\ll u^{2/3}q^{1/9+\varepsilon} for every real u≥M0u\ge M_0. We also have ∣W(u[d,e])∣≪L(u[d,e])−1|W(u[d,e])|\ll\mathcal{L}(u[d,e])^{-1} and ∣∂uW(u[d,e])∣≪L(1+∣γj−γk′∣)u−2[d,e]−1≪L2u−2[d,e]−1|\partial_uW(u[d,e])|\ll\mathcal{L}(1+|\gamma_j-\gamma_{k'}|)u^{-2}[d,e]^{-1}\ll\mathcal{L}^2u^{-2}[d,e]^{-1}. Abel summation, with the bound for the partial sums ∑M0≤m≤uχjk′(m)\sum_{M_0\le m\le u}\chi_{jk'}(m) at every uu, therefore bounds the inner sum by

≪L2[d,e]−1q1/9+εM0−1/3=L2q1/9+ε−tk/3[d,e]−2/3.\ll\mathcal{L}^2[d,e]^{-1}q^{1/9+\varepsilon}M_0^{-1/3} = \mathcal{L}^2q^{1/9+\varepsilon-t_k/3}[d,e]^{-2/3}.

Since ∑d,e≤zk[d,e]−2/3≤∑g≤zkg−2/3(∑d′≤zk/gd′−2/3)2≤9ζ(43)zk2/3\sum_{d,e\le z_k}[d,e]^{-2/3}\le\sum_{g\le z_k}g^{-2/3}(\sum_{d'\le z_k/g}d'^{-2/3})^2\le9\zeta(\frac{4}{3})z_k^{2/3} and ∣ξd∣≤1|\xi_d|\le1, the term is

≪Lq1/9+ε−tk/3+2ϑk/3=Lqε−2ε′/3,\ll\mathcal{L}q^{1/9+\varepsilon-t_k/3+2\vartheta_k/3} = \mathcal{L}q^{\varepsilon-2\varepsilon'/3},

because 2ϑk=tk−13−2ε′2\vartheta_k=t_k-\frac{1}{3}-2\varepsilon'. With ε<ε′/2\varepsilon<\varepsilon'/2 this is o(1)o(1) uniformly, and the sum over j,k′j,k' is o(1)(∑jaj)2o(1)(\sum_j a_j)^2. (A bound for the partial sums only at the end point of the interval, combined with the total variation of WW, would lose a factor (M1/M0)2/3(M_1/M_0)^{2/3}, which is too large for the cells used.)

Summing over the cells,

Q2≤∑jaj2∑k<mhke−2δjtktk+12−tk22ϑk+o(1)(∑jaj)2.(16)Q_2\le \sum_j a_j^2\sum_{k<m}h_ke^{-2\delta_jt_k} \frac{t_{k+1}^2-t_k^2}{2\vartheta_k} +o(1)\left(\sum_j a_j\right)^2. \tag*{(16)}

Conclusion. By (15) and (16) the second factor is at most B+o(1)(∑jaj)2B+o(1)(\sum_j a_j)^2, where B=∑j(D(δj)−d1+ej)aj2+d1(∑jaj)2≥∑j(D(δj)+ej)aj2≥0B=\sum_j(D(\delta_j)-d_1+e_j)a_j^2+d_1(\sum_j a_j)^2\ge\sum_j(D(\delta_j)+e_j)a_j^2\ge0. With the first factor IB+o(1)I_B+o(1) and IB>0I_B>0, the right side of (14) is at most (1+η′)IB[B+η′(∑jaj)2](1+\eta')I_B[B+\eta'(\sum_j a_j)^2] for qq large. □\square

Responses and the threshold form.

Lemma 8.6 (Uniform decay in the imaginary direction). If ff satisfies Condition 1, then Re⁡F(x+iY)→0\operatorname{Re}F(x+iY)\to0 as ∣Y∣→∞|Y|\to\infty, uniformly for xx in compact sets.

Proof. Integrating by parts twice, F(z)=f(0)/z+OK(∣z∣−2)F(z)=f(0)/z+O_K(|z|^{-2}) for Re⁡z\operatorname{Re} z in a compact set KK. Moreover Re⁡(f(0)/z)=f(0)Re⁡z/∣z∣2=OK(∣z∣−2)\operatorname{Re}(f(0)/z)=f(0)\operatorname{Re} z/|z|^2=O_K(|z|^{-2}). □\square

Lemma 8.7 (Response lemma). Assume Input 3.4. Let ff satisfy Conditions 1 and 2, and fix σmax⁡>0\sigma_{\max}>0, an integer K0≥0K_0\geq0, and η′>0\eta'>0. There are M>0M>0, δ∈(0,1)\delta\in(0,1) and q0q_0 such that the following holds for q≥q0q\geq q_0, and it remains true when δ\delta is replaced by any smaller radius (with a suitable q0q_0). Let χ\chi be nonprincipal, ∣σ∣≤σmax⁡|\sigma|\leq\sigma_{\max} and ∣γ∣≤L|\gamma|\leq\mathcal{L}. Let SS be a finite set of distinct zeros of L(s,χ)L(s,\chi) in the disc ∣1+iγ−ρ∣≤δ|1+i\gamma-\rho|\leq\delta, each retained with its full multiplicity mρm_\rho. Suppose that every other zero of L(s,χ)L(s,\chi) in the disc with λρ<σ\lambda_\rho<\sigma satisfies ∣γL−μρ∣≥M|\gamma\mathcal{L}-\mu_\rho|\geq M, and that there are at most K0K_0 such zeros, counted with multiplicity. Then the response of (χ,γ,σ)(\chi,\gamma,\sigma), that is R=−L−1∑nΛ(n)Re⁡(χ(n)n−(1−σ/L+iγ))f(tn)R=-\mathcal{L}^{-1}\sum_n\Lambda(n)\operatorname{Re}(\chi(n)n^{-(1-\sigma/\mathcal{L}+i\gamma)})f(t_n), satisfies

R≥∑ρ∈SmρRe⁡F((λρ−σ)+i(γL−μρ))−f(0)6−η′.R\geq\sum_{\rho\in S}m_\rho\operatorname{Re}F\bigl((\lambda_\rho-\sigma)+i(\gamma\mathcal{L}-\mu_\rho)\bigr)-\frac{f(0)}{6}-\eta'.

Proof. By Lemma 8.6 choose MM with ∣Re⁡F(x+iY)∣≤η′/(2K0+2)|\operatorname{Re}F(x+iY)|\leq\eta'/(2K_0+2) for ∣Y∣≥M|Y|\geq M and x∈[−σmax⁡,0]x\in[-\sigma_{\max},0]. Apply Input 3.4 with ε=η′/2\varepsilon=\eta'/2 at the point w=1−σ/L+iγw=1-\sigma/\mathcal{L}+i\gamma, and let δ\delta be its disc radius, or any smaller radius (see the remark after Input 3.4). We have (w−ρ)L=(λρ−σ)+i(γL−μρ)(w-\rho)\mathcal{L}=(\lambda_\rho-\sigma)+i(\gamma\mathcal{L}-\mu_\rho), so

R≥∑∣1+iγ−ρ∣≤δRe⁡F((w−ρ)L)−φ(χ)2f(0)−η′2,R\geq\sum_{|1+i\gamma-\rho|\leq\delta}\operatorname{Re}F\bigl((w-\rho)\mathcal{L}\bigr)-\frac{\varphi(\chi)}{2}f(0)-\frac{\eta'}{2},

the sum counting multiplicity. Zeros in SS are kept. A zero outside SS with λρ≥σ\lambda_\rho\geq\sigma has Re⁡(w−ρ)L≥0\operatorname{Re}(w-\rho)\mathcal{L}\geq0 and a term ≥0\geq0 (Condition 2), which we drop. The remaining zeros have λρ∈[0,σ)\lambda_\rho\in[0,\sigma) and ∣γL−μρ∣≥M|\gamma\mathcal{L}-\mu_\rho|\geq M; there are at most K0K_0 of them, and each term is at least −η′/(2K0+2)-\eta'/(2K_0+2). Finally φ(χ)≤13\varphi(\chi)\leq\frac{1}{3} and f(0)≥0f(0)\geq0. □\square

Proposition 8.8 (Threshold form). Let JJ be a finite index set, let Dj>0D_j>0 and vj∈Rv_j\in\mathbb{R} for j∈Jj\in J, and let d>0d>0. Suppose that

(∑jajvj)+2≤∑jDjaj2+d(∑jaj)2for every (aj)j∈J∈[0,∞)J.\left(\sum_j a_jv_j\right)_+^2\leq\sum_jD_ja_j^2+d\left(\sum_j a_j\right)^2 \qquad\text{for every }(a_j)_{j\in J}\in[0,\infty)^J.

Then there is τ∈[0,d]\tau\in[0,\sqrt d] with

(T)∑j(vj−τ)+2Dj≤1−τ2d.\text{(T)}\qquad\sum_j\frac{(v_j-\tau)_+^2}{D_j}\leq1-\frac{\tau^2}{d}.

We call the vjv_j the features, the DjD_j the diagonals and dd the correlation term of the inequality, and τ\tau its threshold. A feature measures how strongly an entry is detected, a diagonal what the entry costs on its own, and dd bounds the correlation between distinct entries.

Example 8.9 (One level and two levels). (i) Suppose that NN entries have the same feature v>0v>0 and the same diagonal DD, with v2>dv^2>d. Taking every aj=1a_j=1 in the hypothesis of Proposition 8.8 gives (Nv)2≤ND+dN2(Nv)^2\leq ND+dN^2, that is,

N≤Dv2−d.N\leq\frac{D}{v^2-d}.

This is the kind of bound that Heath-Brown’s Lemma 12.1 provides: a bound for the number N(λ)N(\lambda) of characters with a zero at distance at most λ\lambda, all counted alike.

(ii) Now take D=1D=1 and d=14d=\frac{1}{4}, one entry with feature v1=1v_1=1 (a zero close to 1) and two entries with feature v2=34v_2=\frac{3}{4} (zeros further away). The count of (i), applied to each level separately, allows this: it permits up to D/(v12−d)=43D/(v_1^2-d)=\frac{4}{3} entries with feature at least v1v_1, and up to D/(v22−d)=165D/(v_2^2-d)=\frac{16}{5} with feature at least v2v_2. But (T) fails for every τ\tau: for 0≤τ≤120 \le\tau\le\frac{1}{2} its left side minus its right side is

(1−τ)2+2(34−τ)2−(1−4τ2)=7τ2−5τ+98,(1-\tau)^2+2\left(\frac{3}{4}-\tau\right)^2-\left(1-4\tau^2\right)=7\tau^2-5\tau+\frac{9}{8},

whose discriminant 25−63225-\frac{63}{2} is negative. (Equivalently, taking all three aj=1a_j=1 violates the hypothesis: (1+32)2=254>3+94\left(1+\frac{3}{2}\right)^2=\frac{25}{4}>3+\frac{9}{4}.) So this configuration is excluded. A single threshold couples the near entry and the farther ones, which separate counts cannot do. In the leaf programs this coupling acts on all the bins of characters at once.

Proof. The function J(τ)=τ2/d+∑j(vj−τ)+2/Dj\mathcal{J}(\tau)=\tau^2/d+\sum_j(v_j-\tau)_+^2/D_j is convex and C1C^1 on [0,∞)[0,\infty), and tends to ∞\infty; let τ∗\tau^* minimize it. If τ∗=0\tau^*=0 then J′(0)≥0\mathcal{J}'(0)\ge0, so all vj≤0v_j\le0 and J(0)=0\mathcal{J}(0)=0. Otherwise J′(τ∗)=0\mathcal{J}'(\tau^*)=0, that is τ∗=d∑jaj\tau^*=d\sum_j a_j with aj=(vj−τ∗)+/Dj≥0a_j=(v_j-\tau^*)_+/D_j\ge0. Then ∑jajvj=∑jaj(vj−τ∗)+τ∗∑jaj=∑jDjaj2+d(∑jaj)2=J(τ∗)\sum_j a_jv_j=\sum_j a_j(v_j-\tau^*)+\tau^*\sum_j a_j=\sum_jD_ja_j^2+d(\sum_j a_j)^2=\mathcal{J}(\tau^*). The hypothesis gives J(τ∗)2≤J(τ∗)\mathcal{J}(\tau^*)^2\le\mathcal{J}(\tau^*), so J(τ∗)≤1\mathcal{J}(\tau^*)\le1. Finally (T) forces τ2≤d\tau^2\le d. □\square

Corollary 8.10 (Zero form). Fix a near datum whose detector ff also satisfies Condition 2. Fix a common bound K0K_0 for the omitted zeros in Lemma 8.7, and choose σmax⁡≥max⁡δ∈Δ∣s−δ∣\sigma_{\max}\ge\max_{\delta\in\Delta}|s-\delta| with σmax⁡>0\sigma_{\max}>0. Choose η′>0\eta'>0, Iu≥IBI_u\ge I_B, and Du>0D_u>0, and put

η′′=η′min⁡{1,IBDu,Du}.\eta''=\eta'\min\{1,\sqrt{I_BD_u},D_u\}.

For sufficiently large qq, consider entries satisfying Theorem 8.4 with tolerance η′′\eta''. For each entry jj, suppose the retained set SjS_j satisfies Lemma 8.7 at anchor sjs_j, with the separation and disc radius chosen for tolerance η′′\eta''. Define

rj=∑ρ∈SjmρRe⁡F((λρ−sj)+i(γjL−μρ))−f(0)6,r_j=\sum_{\rho\in S_j}m_\rho\operatorname{Re}F\left((\lambda_\rho-s_j)+i(\gamma_j\mathcal{L}-\mu_\rho)\right)-\frac{f(0)}{6},

and choose Dj+>0D_j^+>0 with Dj+≥D(δj)−d1+ejD_j^+\ge D(\delta_j)-d_1+e_j for every jj. Then there is τ∈[0,d]\tau\in[0,\sqrt d] satisfying (T) with

vj=rjIuDu−η′,Dj=(1+η′)Dj+Du,d=(1+η′)(d1+η′Du)Du (=(1+η′)(d1/Du+η′)).v_j=\frac{r_j}{\sqrt{I_uD_u}}-\eta',\qquad D_j=\frac{(1+\eta')D_j^+}{D_u},\qquad d=\frac{(1+\eta')(d_1+\eta'D_u)}{D_u} \ \left(=(1+\eta')(d_1/D_u+\eta')\right).

The same threshold remains valid if features vjv_j are decreased, diagonals DjD_j or the correlation term dd are increased, or nonpositive features are replaced by 00. An entry with nonpositive feature may also be omitted.

Proof. Apply Theorem 8.4 and Lemma 8.7 with the tolerance η′′\eta'' fixed in the statement. The response lemma gives Rj≥kj≔rj−η′′R_j\ge k_j\coloneqq r_j-\eta'', so for a≥0a\ge0 we have (∑jajkj)+≤(∑jajRj)+(\sum_j a_jk_j)_+\le(\sum_j a_jR_j)_+. Dividing the inequality of Theorem 8.4 by IBDuI_BD_u, we get for all a≥0a\ge0

(∑jajkjIBDu)+2≤∑j(1+η′′)Dj+Duaj2+(1+η′′)(d1+η′′)Du(∑jaj)2,\left(\sum_j a_j\frac{k_j}{\sqrt{I_BD_u}}\right)_+^2 \le \sum_j\frac{(1+\eta'')D_j^+}{D_u}a_j^2 + \frac{(1+\eta'')(d_1+\eta'')}{D_u}\left(\sum_j a_j\right)^2,

where we also replaced each coefficient D(δj)−d1+ejD(\delta_j)-d_1+e_j by the larger Dj+D_j^+. Proposition 8.8 gives τ\tau satisfying (T) with these features, diagonals and correlation term; only the positivity of the Dj+D_j^+ is needed, not that of D(δj)−d1+ejD(\delta_j)-d_1+e_j. We compare with the displayed quantities.

  • If rj>0r_j>0 then vj≤kj/IBDuv_j\le k_j/\sqrt{I_BD_u}, because Iu≥IBI_u\ge I_B and η′′/IBDu≤η′\eta''/\sqrt{I_BD_u}\le\eta'. If rj≤0r_j\le0 then vj<0v_j<0.

  • (1+η′)≥(1+η′′)(1+\eta')\ge(1+\eta''), so the displayed diagonals are at least the ones above.

  • (1+η′)(d1+η′Du)≥(1+η′′)(d1+η′′)(1+\eta')(d_1+\eta'D_u)\ge(1+\eta'')(d_1+\eta''), since η′′≤η′Du\eta''\le\eta'D_u.

Finally, if (T) holds at τ\tau for features vjv_j, diagonals DjD_j and correlation term dd, it also holds at τ\tau for features vj′≤max⁡(vj,0)v'_j \le\max(v_j,0) with vj′≤vjv'_j \le v_j or vj′≤0v'_j \le0, diagonals Dj′≥DjD'_j \ge D_j and correlation term d′≥dd' \ge d. Indeed each term (vj′−τ)+2/Dj′(v'_j-\tau)_+^2/D'_j is at most (vj−τ)+2/Dj(v_j-\tau)_+^2/D_j (the former is 0 when vj′≤0≤τv'_j \le0 \le\tau), 1−τ2/d≤1−τ2/d′1-\tau^2/d \le1-\tau^2/d', and τ≤d≤d′\tau\le\sqrt d \le\sqrt{d'}. □\square

The near rows

A near row is a constraint obtained by applying the threshold inequality (T) to selected characters and zeros. In a leaf program (Section 2.4), characters are grouped by parameter into bins, each with its own column. Known entries from the first or reserved family may instead be recorded separately as family terms. There are four kinds of rows, the family, shifted, graded and two-test rows; Table 8 in Section 9.5 lists their anchors, sieve weights and entries.

We first specify the test functions and normalization. We then justify each permitted kind of entry and describe how to combine them into rows. Most rows follow from Corollary 8.10; the two-test row in Section 9.6 has a separate proof. Proposition 9.7 is the conclusion needed later: every actual zero configuration in a leaf satisfies each of its near rows at some threshold.

The parabolic test function

All detectors and Gram tests are normalized parabolic autocorrelations. For c>0c>0 put

P(u)=(1−u)3(1+3u+u2)=1−5u2+5u3−u5,f^c(t)={P(t/c),0≤t≤c,0,t≥c.P(u)=(1-u)^3(1+3u+u^2)=1-5u^2+5u^3-u^5, \qquad \widehat f_c(t)= \begin{cases} P(t/c), & 0\le t\le c,\\ 0, & t\ge c. \end{cases}

Thus f^2γ=fγ/fγ(0)\widehat f_{2\gamma}=f_\gamma/f_\gamma(0) with fγf_\gamma as in (7.1). Write F^c\widehat F_c for its transform.

Lemma 9.1 (Properties of the parabolic test functions). Let c>0c>0.

(i) f^c≥0\widehat f_c\ge0, f^c(0)=1\widehat f_c(0)=1, f^c\widehat f_c is decreasing on [0,c][0,c], and f^c\widehat f_c extended by 00 is twice continuously differentiable on [0,∞)[0,\infty) with ∣f^c′′∣≤10/c2|\widehat f_c''|\le10/c^2. So Condition 1 holds.

(ii) Re⁡F^c(iy)=60c(cycos⁡cy2−2sin⁡cy2)2/(cy)6≥0\operatorname{Re}\widehat F_c(iy)=60c(cy\cos\frac{cy}{2}-2\sin\frac{cy}{2})^2/(cy)^6\ge0 for y≠0y\ne0, and Condition 2 holds.

(iii) x↦F^c(x)x\mapsto\widehat F_c(x) is positive and strictly decreasing on R\mathbb R.

(iv) For z≠0z\ne0, F^c(z)=cΦ(cz)\widehat F_c(z)=c\Phi(cz), where

Φ(w)=1w−10w3+30(1+e−w)w4+120e−ww5−120(1−e−w)w6(w≠0).\Phi(w)=\frac{1}{w}-\frac{10}{w^3}+\frac{30(1+e^{-w})}{w^4}+\frac{120e^{-w}}{w^5}-\frac{120(1-e^{-w})}{w^6} \qquad(w\ne0).

The singularity at 00 is removable. Consequently, for d≥0d\ge0 and ∣y∣≥U>0|y|\ge U>0,

Re⁡F^c(−d+iy)≤cR(cU,ecd),−Re⁡F^c(−d+iy)≤dU2+cR(cU,ecd),\operatorname{Re}\widehat F_c(-d+iy)\le cR(cU,e^{cd}), \qquad -\operatorname{Re}\widehat F_c(-d+iy)\le\frac{d}{U^2}+cR(cU,e^{cd}),

where R(u,E)=10u3+30u4+120u6+E(30u4+120u5+120u6)R(u,E)=\frac{10}{u^3}+\frac{30}{u^4}+\frac{120}{u^6}+E\left(\frac{30}{u^4}+\frac{120}{u^5}+\frac{120}{u^6}\right).

(v) For d≥0d\ge0 put Cc(d)=sup⁡Re⁡z≥−d(−Re⁡F^c(z))C_c(d)=\sup_{\operatorname{Re}z\ge-d}(-\operatorname{Re}\widehat F_c(z)). It is finite,

Cc(d)=max⁡(0,sup⁡y∈R(−Re⁡F^c(−d+iy))),C_c(d)=\max\left(0,\sup_{y\in\mathbb R}(-\operatorname{Re}\widehat F_c(-d+iy))\right),

and Cc(0)=0C_c(0)=0.

Proof. (i) is elementary: P′(u)=−5u(2+u)(1−u)2P'(u)=-5u(2+u)(1-u)^2, and P(1)=P′(1)=P′′(1)=0P(1)=P'(1)=P''(1)=0. (ii) The function f^c\widehat f_c is a positive multiple of the autocorrelation of x↦(γ2−x2)+x\mapsto(\gamma^2-x^2)_+ with c=2γc=2\gamma, so its transform on the imaginary axis is a positive multiple of (∫0∞(γ2−x2)+cos⁡(xy) dx)2\left(\int_0^\infty(\gamma^2-x^2)_+\cos(xy)\,\mathrm dx\right)^2 [145], Lemma 3.6, which gives the closed form. Condition 2 then follows from Input 3.2 with F2=0F_2=0, since F^c(z)→0\widehat F_c(z)\to0 uniformly in Re⁡z>0\operatorname{Re}z>0. (iii) F^c′(x)=−∫tf^c(t)e−xt dt<0\widehat F_c'(x)=-\int t\widehat f_c(t)e^{-xt}\,\mathrm dt<0. (iv) Five integrations by parts, using the values of PP and its derivatives at 0 and 1. For z=−d+iyz=-d+iy we have cRe⁡1cz=−dd2+y2c\operatorname{Re}\frac{1}{cz}=\frac{-d}{d^{2}+y^{2}}, which is ≤0\le0 and ≥−d/U2\ge-d/U^{2}, and each other term is bounded by the corresponding term of cR(c∣z∣,ecd)cR(c|z|,e^{cd}), which decreases in ∣z∣≥U|z|\ge U. (v) −Re⁡F^c(z−d)-\operatorname{Re}\widehat{F}_{c}(z-d) is harmonic in Re⁡z>0\operatorname{Re}z>0, continuous up to the boundary and tends to 0 uniformly at infinity, so its supremum over Re⁡z≥0\operatorname{Re}z\ge0 is attained on the boundary or is 0; Cc(0)=0C_{c}(0)=0 by Condition 2. □\square

Row data

Every family, shifted and graded row fixes rational numbers γf>0\gamma_{f}>0 (detector), γg>0\gamma_{g}>0 (Gram test), s1≤ss_{1}\le s (safe and reference anchors), t0t_{0} and μrow\mu_{\mathrm{row}}, and uses the near datum of Section 8.1 with f=f^2γff=\widehat{f}_{2\gamma_{f}}, g=f^2γgg=\widehat{f}_{2\gamma_{g}}, ε′=12000\varepsilon'=\frac{1}{2000}, and 2000 equal cells t0<t1<⋯<t2000=2γft_{0}<t_{1}<\cdots<t_{2000}=2\gamma_{f}, where t0>13+2ε′t_{0}>\frac{1}{3}+2\varepsilon'. The heights hk≥0h_{k}\ge0 are rational numbers. They are chosen by a heuristic that brings ω\omega close to the Cauchy–Schwarz optimum ω∝estf\omega\propto e^{st}f: the height of cell kk is the positive part of esmkf(mk)/μrowκk−g(mk)es1mke^{sm_{k}}f(m_{k})/\sqrt{\mu_{\mathrm{row}}\kappa_{k}}-g(m_{k})e^{s_{1}m_{k}} at the midpoint mkm_{k}, where κk=(tk+12−tk2)/(2ϑk(tk+1−tk))\kappa_{k}=(t_{k+1}^{2}-t_{k}^{2})/(2\vartheta_{k}(t_{k+1}-t_{k})) is the mean sieve cost per unit length of the cell, computed in binary64 arithmetic and converted to a rational number with denominator at most 101210^{12}. Any nonnegative heights are admissible, so this choice plays no role in the proofs. In every family and shifted row μrow=109\mu_{\mathrm{row}}=10^{9} and all the computed heights are 0, so Ξ≡0\Xi\equiv0. The offsets are the finitely many numbers s−σs-\sigma that occur for the anchors σ\sigma of the entries. In every row the lower bound of ω\omega on each cell of the support of ff is checked to be positive; in the rows with Ξ≡0\Xi\equiv0 this holds because γg>γf\gamma_{g}>\gamma_{f}.

We use the normalization of Corollary 8.10: IuI_{u} is an upper Riemann sum for IBI_{B} over 4000 cells, computed with outward interval arithmetic; Du≥D(0)D_{u}\ge D(0); each entry jj receives the feature vj=rj/IuDu−η′v_{j}=r_{j}/\sqrt{I_{u}D_{u}}-\eta' (rounded down, and replaced by 0 if negative) and the diagonal (1+η′)Dj+/Du(1+\eta')D_{j}^{+}/D_{u} (rounded up), where Dj+D_{j}^{+} is an upper bound for D(δj)−d1+ejD(\delta_{j})-d_{1}+e_{j}; and the row has the correlation term d=(1+η′)(d1/Du+η′)d=(1+\eta')(d_{1}/D_{u}+\eta') (rounded up). Here η′=10−6\eta'=10^{-6}. For an entry with offset δ\delta the value D(⌊δ⌋1/200)≥D(δ)D(\lfloor\delta\rfloor_{1/200})\ge D(\delta) is used (here ⌊δ⌋1/200\lfloor\delta\rfloor_{1/200} is δ\delta rounded down to a multiple of 1200\frac{1}{200}, and DD is decreasing), and an entry is omitted if its diagonal numerator does not meet the positivity margin used in the construction. An omitted entry has feature 0 in the row and contributes nothing to (T). The rounding directions are justified by the last assertion of Corollary 8.10.

Facts available in a leaf

A leaf specifies intervals and lower bounds for the first few zeros. The full definition of a specification s\mathfrak{s} and its configuration set C(s)C(\mathfrak{s}) is in Definition 11.2; the information needed for this section is listed below. Write ν(F)\nu(\mathcal{F}) for the least parameter in a family among zeros of physical height at most 1, and νT∗(χ)\nu_{T^{*}}(\chi) for the corresponding minimum at height at most T∗T^{*}. A reserved family minimizes ν\nu outside the first family. It need not be the family attaining the global parameter λ2\lambda_{2}.

Fix a leaf with first-zero cell [a,b][a,b] and a configuration in C(s)C(\mathfrak{s}). For all sufficiently large qq, the following facts are available.

(Z1) λ1∈[a,b]\lambda_{1}\in[a,b]. The safe anchor of the leaf is s1=⌊50a⌋/50≤λ1s_{1}=\lfloor50a\rfloor/50\le\lambda_{1}.

(Z2) λ′≥p∗\lambda'\ge p^{*}, where p∗=lo′p^{*}=\mathrm{lo}' if there is a gap (an interval [lo′,hi′][\mathrm{lo}',\mathrm{hi}'] known to contain λ′\lambda') and p∗=pp^{*}=p otherwise; if there is a finite gap, λ′≤hi′\lambda'\le\mathrm{hi}'.

(Z3) Every zero in R(l)R(l) of a character outside the first family has parameter at least λ2≥l2\lambda_{2}\ge l_{2}.

(Z4) Every family F≠F1\mathcal{F}\ne\mathcal{F}_{1} has ν(F)≥r\nu(\mathcal{F})\ge r, and ordinary characters have νT∗≥ν≥r\nu_{T^{*}}\ge\nu\ge r.

(Z5) If s\mathfrak{s} is reserved, the reserved family has n2n_{2} characters with common height-one parameter ν∗∈[lo2,hi2]\nu_{*}\in[\mathrm{lo}_{2},\mathrm{hi}_{2}].

(Z6) Inside: ∣μ1∣≤CB|\mu_{1}|\le C_{B}. Outside: ∣μ1∣>CB|\mu_{1}|>C_{B}.

Lemma 9.2 (No zero to the right). Let χ\chi be nonprincipal, let ∣γ∣≤2l|\gamma|\le2l and 0<δ≤10<\delta\le1, and fix a real anchor σ\sigma. For sufficiently large qq, every zero ρ\rho of L(s,χ)L(s,\chi) in the disc ∣1+iγ−ρ∣≤δ|1+i\gamma-\rho|\le\delta has λρ≥σ\lambda_{\rho}\ge\sigma in either of the following cases:

(a) σ<λ1\sigma<\lambda_{1};

(b) χ∉F1\chi\notin\mathcal{F}_{1} and σ≤l2\sigma\le l_{2}, where l2≤λ2l_{2}\le\lambda_{2} is the leaf’s global second-family lower bound.

Equivalently, this disc contains no zero to the right of the line Re⁡s=1−σ/L\operatorname{Re}s=1-\sigma/\mathcal{L}.

Proof. Such a zero has ∣Im⁡ρ∣≤2l+δ<10l|\operatorname{Im}\rho|\le2l+\delta<10l. By Input 3.5 it lies in R(l)R(l) or has λρ>13log⁡log⁡L\lambda_{\rho}>\frac{1}{3}\log\log\mathcal{L}, which exceeds every fixed σ\sigma. In R(l)R(l) its parameter is at least λ1\lambda_{1}, and at least λ2≥l2\lambda_{2}\ge l_{2} if χ∉F1\chi\notin\mathcal{F}_{1} (Z3). □\square

Lemma 9.3 (Separation of zeros beyond the representative strip). Let χ\chi be an ordinary character with T∗T^{*}-representative ρχ\rho_{\chi}, of height γχ\gamma_{\chi} and parameter λχ\lambda_{\chi}, and let σ≤min⁡(λχ,smax⁡)\sigma\le\min(\lambda_{\chi},s_{\max}). Then every zero ρ\rho of L(s,χ)L(s,\chi) with ∣1+iγχ−ρ∣≤δ0|1+i\gamma_{\chi}-\rho|\le\delta_{0} and λρ<σ\lambda_{\rho}<\sigma satisfies ∣γχL−μρ∣≥M|\gamma_{\chi}\mathcal{L}-\mu_{\rho}|\ge M, and there are at most K0K_{0} such zeros, counted with multiplicity.

Proof. Such a zero has ∣Im⁡ρ∣≤T∗+δ0<2≤10l|\operatorname{Im}\rho|\le T^{*}+\delta_{0}<2\le10l and λρ<σ≤smax⁡<13log⁡log⁡L\lambda_{\rho}<\sigma\le s_{\max}<\frac{1}{3}\log\log\mathcal{L}, so it lies in R(10l)R(10l) and hence, by Input 3.5, in R(l)R(l). As λρ<σ≤λχ\lambda_{\rho}<\sigma\le\lambda_{\chi}, the minimality of the representative gives ∣Im⁡ρ∣>T∗|\operatorname{Im}\rho|>T^{*}; so the zero is counted in K0K_{0}. By Lemma 6.3, ∣Im⁡ρ∣≥T∗+M/L|\operatorname{Im}\rho|\ge T^{*}+M/\mathcal{L}, so L∣Im⁡ρ−γχ∣≥L(∣Im⁡ρ∣−T∗)≥M\mathcal{L}|\operatorname{Im}\rho-\gamma_{\chi}|\ge\mathcal{L}(|\operatorname{Im}\rho|-T^{*})\ge M. □\square

Entries

We now list the entries that occur in the rows, with their kept zeros. In each case the hypotheses of Lemma 8.7 hold with the stated kept set SS and anchor σ\sigma; we write rr for the resulting lower bound for the response, before normalization, so that the entry’s feature is r/IuDu−η′r/\sqrt{I_{u}D_{u}}-\eta'. Table 7 summarizes the entries. Recall that F(x)F(x) is positive and decreasing on R\mathbb{R}, and that a kept zero with λρ≥σ\lambda_{\rho}\ge\sigma has mρRe⁡F(⋅)≥Re⁡F(⋅)≥0m_{\rho}\operatorname{Re}F(\mathord{\cdot})\ge\operatorname{Re}F(\mathord{\cdot})\ge0.

(a) Ordinary character in the bin [lo,hi)[\mathrm{lo},\mathrm{hi}). Entry (χ,γχ,s−σ)(\chi,\gamma_{\chi},s-\sigma) at its T∗T^{*}-representative, with σ=min⁡(lo,s)\sigma=\min(\mathrm{lo},s); S={ρχ}S=\{\rho_{\chi}\}. If σ≤l2\sigma\le l_{2} there is no zero to the right (Lemma 9.2); otherwise Lemma 9.3 gives the hypotheses. Then r=F(hi−σ)−16f(0)r=F(\mathrm{hi}-\sigma)-\frac{1}{6}f(0), and the diagonal numerator is D(s−σ)−d1D(s-\sigma)-d_{1}.

(b) Inside first family, one entry per character. Entries (χ1,γ1,s−σ)(\chi_{1},\gamma_{1},s-\sigma) and, for type complex, (χ‾1,−γ1,s−σ)(\overline{\chi}_{1},-\gamma_{1},s-\sigma), with σ=min⁡(a,s)≤λ1\sigma=\min(a,s)\le\lambda_{1}; S={ρ1}S=\{\rho_{1}\} (resp. {ρ‾1}\{\overline{\rho}_{1}\}). No zero to the right. r=F(b−σ)−16f(0)r=F(b-\sigma)-\frac{1}{6}f(0), and the diagonal numerator is D(s−σ)−d1D(s-\sigma)-d_{1}.

(c) Conjugate pair, type rc (rows with s=s1s=s_{1} and Ξ=0\Xi=0). Entries (χ1,±γ1,0)(\chi_{1},\pm\gamma_{1},0), a pair of entries of the same character with height difference 2γ12\gamma_{1}; S={ρ1,ρ‾1}S=\{\rho_{1},\overline{\rho}_{1}\} for both. No zero to the right, and

r=F(b−s1)+Rlo−16f(0),Rlo≤inf⁡Re⁡F(x+2iμ)r=F(b-s_{1})+R_{\mathrm{lo}}-\frac{1}{6}f(0),\qquad R_{\mathrm{lo}}\le\inf\operatorname{Re}F(x+2i\mu)

over x∈[a−s1,b−s1]x\in[a-s_{1},b-s_{1}] and μ\mu in the height range of the leaf ([0,1][0,1] or [1,∞)[1,\infty)). Pair excess e≤(chi−d1)+e\le(c^{\mathrm{hi}}-d_{1})_{+} with chi≥sup⁡Re⁡G(−s1+2iμ)c^{\mathrm{hi}}\ge\sup\operatorname{Re}G(-s_{1}+2i\mu) over the range, so the diagonal numerator is D(0)−d1+(chi−d1)+D(0)-d_{1}+(c^{\mathrm{hi}}-d_{1})_{+}.

(d) Second zero of the first family (rows with s=s1s=s_{1} and Ξ=0\Xi=0; type complex with a finite gap [p∗,hi′][p^{*},\mathrm{hi}']). Let ρ′\rho' be an admissible zero of χ1\chi_{1} realizing λ′\lambda' (Section 2), of height γ′\gamma', namely the one that witnesses the piece of the second-zero split (Section 11.6) in which the configuration lies, and y=(γ1−γ′)Ly=(\gamma_{1}-\gamma')\mathcal{L}, restricted to that piece: ∣y∣∈[0,1]|y|\in[0,1], [1,32][1,\frac{3}{2}], [32,2][\frac{3}{2},2], [2,3][2,3], [3,5][3,5] or [5,∞)[5,\infty). Entries (χ1,γ1,0)(\chi_{1},\gamma_{1},0), (χ1,γ′,0)(\chi_{1},\gamma',0), (χ‾1,−γ1,0)(\overline{\chi}_{1},-\gamma_{1},0), (χ‾1,−γ′,0)(\overline{\chi}_{1},-\gamma',0), with kept sets {ρ1,ρ′}\{\rho_{1},\rho'\}, {ρ‾′,ρ‾1}\{\overline{\rho}',\overline{\rho}_{1}\} and their conjugates (a zero is kept only if it lies in the disc of the entry). No zero to the right. Then

rγ1=F(b−s1)+Rp−16f(0),rγ′=F(hi′−s1)+Rq−16f(0),r_{\gamma_{1}}=F(b-s_{1})+R_{p}-\frac{1}{6}f(0),\qquad r_{\gamma'}=F(\mathrm{hi}'-s_{1})+R_{q}-\frac{1}{6}f(0),

where Rp≤inf⁡Re⁡F(x+iy)R_{p}\le\inf\operatorname{Re}F(x+iy) over x∈[p∗−s1,hi′−s1]x\in[p^{*}-s_{1},\mathrm{hi}'-s_{1}] and Rq≤inf⁡Re⁡F(x+iy)R_{q}\le\inf\operatorname{Re}F(x+iy) over x∈[a−s1,b−s1]x\in[a-s_{1},b-s_{1}], with ∣y∣|y| in the piece. Here x≥0x\ge0 (as s1≤a≤p∗s_{1}\le a\le p^{*}), so Re⁡F(x+iy)≥0\operatorname{Re}F(x+iy)\ge0 by Condition 2 and we take Rp,Rq≥0R_{p},R_{q}\ge0; the multiplicities of the kept zeros can then only increase the sums. On the last piece Rp=Rq=0R_{p}=R_{q}=0, which also covers a zero outside the disc (its term is then absent). Each same-character pair is a controlled pair with excess (chi−d1)+(c^{\mathrm{hi}}-d_{1})_{+}, chi≥sup⁡Re⁡G(−s1+iy)c^{\mathrm{hi}}\ge\sup\operatorname{Re}G(-s_{1}+iy), so every diagonal numerator is D(0)−d1+(chi−d1)+D(0)-d_{1}+(c^{\mathrm{hi}}-d_{1})_{+}; an entry of χ1\chi_{1} and an entry of χ‾1\overline{\chi}_{1} form a pair of distinct characters, with the nonprincipal quotient χ12\chi_{1}^{2}, and carry no mutual excess. If ρ′=ρ1\rho'=\rho_{1} is a double zero, the two entries coincide and r≥mρ1F(λ1−s1)−16f(0)r\ge m_{\rho_{1}}F(\lambda_{1}-s_{1})-\frac{1}{6}f(0) with mρ1≥2m_{\rho_{1}}\ge2, which is at least both rγ1r_{\gamma_{1}} and rγ′r_{\gamma'} (here y=0y=0 and λ′=λ1\lambda'=\lambda_{1}).

(e) Shifted first family (the shifted row, s=min⁡(1.9,p∗,l2)s=\min(1.9,p^{*},l_{2}), used only if s>s1s>s_{1} and s≥as\ge a). Entries (χ1,γ1,0)(\chi_{1},\gamma_{1},0) and, for type complex, (χ‾1,−γ1,0)(\overline{\chi}_{1},-\gamma_{1},0), with anchor ss. The only zeros of χ1\chi_{1} with parameter below ss are ρ1\rho_{1} and, for type rc, ρ‾1\overline{\rho}_{1}, since every other occurrence has parameter at least λ′≥p∗≥s\lambda'\ge p^{*}\ge s. With S={ρ1}S=\{\rho_{1}\} (and ρ‾1\overline{\rho}_{1} for type rc) nothing remains in the hypotheses of Lemma 8.7. As conjugate zeros of a real character have the same multiplicity mm,

∑ρ∈SmρRe⁡F(⋅)=m(F(λ1−s)+1rcRe⁡F(λ1−s+2iμ1))≥m(F(b−s)−1rcCZ),\sum_{\rho\in S}m_{\rho}\operatorname{Re}F(\mathord{\cdot})=m\left(F(\lambda_{1}-s)+\mathbf{1}_{\mathrm{rc}}\operatorname{Re}F(\lambda_{1}-s+2i\mu_{1})\right)\ge m\left(F(b-s)-\mathbf{1}_{\mathrm{rc}}C_{Z}\right),

with CZ=Cc(s−a)C_{Z}=C_{c}(s-a) from Lemma 9.1(v) (00 if s≤as\le a). If F(b−s)−1rcCZ≥0F(b-s)-\mathbf{1}_{\mathrm{rc}}C_{Z}\ge0 this is at least F(b−s)−1rcCZF(b-s)-\mathbf{1}_{\mathrm{rc}}C_{Z}, and we take r=F(b−s)−1rcCZ−16f(0)r=F(b-s)-\mathbf{1}_{\mathrm{rc}}C_{Z}-\frac{1}{6}f(0). Otherwise the proposed bound F(b−s)−CZ−16f(0)F(b-s)-C_{Z}-\frac{1}{6}f(0) is negative and the entry receives the feature 00; this is admissible whatever the true response bound rjr_{j} is, by the last clause of Corollary 8.10 (with vj′=0v'_{j}=0). The diagonal numerator is D(0)−d1D(0)-d_{1}.

(f) Reserved family. Entries (χ,γχ,s−σ2)(\chi,\gamma_{\chi},s-\sigma_{2}) for its n2n_{2} characters at their height-one representatives, with σ2=min⁡(max⁡(a,min⁡(l2,lo2)),s)\sigma_{2}=\min(\max(a,\min(l_{2},\mathrm{lo}_{2})),s). Every zero of these characters in R(l)R(l) has parameter at least max⁡(λ1,λ2)≥max⁡(a,l2)≥σ2\max(\lambda_{1},\lambda_{2})\geq\max(a,l_{2})\geq\sigma_{2}, so there is no zero to the right, r=F(hi2−σ2)−16f(0)r=F(\mathrm{hi}_{2}-\sigma_{2})-\frac{1}{6}f(0), and the diagonal numerator is D(s−σ2)−d1D(s-\sigma_{2})-d_{1}. With second-family columns, the column [lo2,j,hi2,j][\mathrm{lo}_{2,j},\mathrm{hi}_{2,j}] uses σ=min⁡(max⁡(a,min⁡(l2,lo2,j)),s)\sigma=\min(\max(a,\min(l_{2},\mathrm{lo}_{2,j})),s), r=F(hi2,j−σ)−16f(0)r=F(\mathrm{hi}_{2,j}-\sigma)-\frac{1}{6}f(0) and the diagonal numerator D(s−σ)−d1D(s-\sigma)-d_{1}.

(g) Outside first family (family row, s=s1s=s_{1}, Ξ=0\Xi=0). Global entries (χ1,γ1,0)(\chi_{1},\gamma_{1},0) with S={ρ1}\mathcal{S}=\{\rho_{1}\} and, for type complex, (χˉ1,−γ1,0)(\bar{\chi}_{1},-\gamma_{1},0) with S={ρˉ1}\mathcal{S}=\{\bar{\rho}_{1}\} (for type rc there is the single global entry (χ1,γ1,0)(\chi_{1},\gamma_{1},0)); r=F(b−s1)−16f(0)r=F(b-s_{1})-\frac{1}{6}f(0). And a hidden entry (χ,γh,0)(\chi,\gamma_{h},0) for each first-family character χ\chi with a zero in RPR_{P}, at its zero ρh\rho_{h} of least parameter tχ∈[loh,hih)t_{\chi}\in[\mathrm{lo}_{h},\mathrm{hi}_{h}), with S={ρh}\mathcal{S}=\{\rho_{h}\} and r=F(hih−s1)−16f(0)r=F(\mathrm{hi}_{h}-s_{1})-\frac{1}{6}f(0). No zero to the right. Entries of the same character have normalized height differences at least CB−CP>MC_{B}-C_{P}>M, so their pair excess vanishes, and every diagonal numerator is D(0)−d1D(0)-d_{1}.

The rows

Each leaf program has one or more rows, of the kinds listed in Table 8. In the family row of an inside leaf the first family is entered by (b), by (c) for type rc (with the height split of Section 11.6), or by (d) in a piece of the second-zero split. In the shifted row every ordinary anchor is at most s≤l2s\leq l_{2}. In the graded rows all characters are distinct.

In every family, shifted and graded row the safe-anchor hypothesis holds by (Z1) and Input 3.5, as explained after Definition 8.3, and the characters are pairwise distinct whenever Ξ≢0\Xi\not\equiv0. All the entries have heights at most ll. The first family is entered once per character in a row with sieve weights (for type rc only ρ1\rho_{1} is kept); the pair and second-zero entries occur only in rows with Ξ=0\Xi=0.

The two-test row

On 37 inside roots the graded rows alone do not certify, and the leaf program contains in addition the following near row. It is not an instance of Theorem 8.4: its Gram form is at the shift ss of the leaf, which may exceed λ1\lambda_{1}, and it pays explicit corrections for the quotient characters χ1,χˉ1\chi_{1},\bar{\chi}_{1}. Its proof follows Xylouris’s proof of his Lemma 5.3 [145], (5.34)–(5.40), pp. 72–76, with two test functions.

Fix γG>γZ>0\gamma_{G}>\gamma_{Z}>0 and let f=fγZf=f_{\gamma_{Z}} and g=fγGg=f_{\gamma_{G}} as in (11), so that g>0g>0 on the support of ff. Let s≤min⁡(p∗,l2)s\leq\min(p^{*},l_{2}) be the shift of the leaf, and put

RG=G(−s),RZ=∫02γZf(t)2g(t)est dt,N=RGRZ,d0=g(0)6RG,CG=sup⁡Re⁡z≥a−s(−Re⁡G(z)),CF=sup⁡Re⁡z≥a−s(−Re⁡F(z)),cG=CGRG,kη=1+η.\begin{aligned} R_{G} &= G(-s), & R_{Z} &= \int_{0}^{2\gamma_{Z}}\frac{f(t)^{2}}{g(t)}e^{st}\,\mathrm{d}t, & \mathcal{N} &= \sqrt{R_{G}R_{Z}}, & d_{0} &= \frac{g(0)}{6R_{G}}, \\ C_{G} &= \sup_{\operatorname{Re}z\geq a-s}\left(-\operatorname{Re}G(z)\right), & C_{F} &= \sup_{\operatorname{Re}z\geq a-s}\left(-\operatorname{Re}F(z)\right), & c_{G} &= \frac{C_{G}}{R_{G}}, & k_{\eta} &= 1+\eta. \end{aligned}

The function f2/gf^{2}/g, extended by 00, satisfies Condition 1, since it vanishes to order 66 at 2γZ2\gamma_{Z}. Define the features

v(λ)=F(λ−s)−16f(0)N−η,v2=F(hi2−s)−ϕ22f(0)N−η,v(\lambda)=\frac{F(\lambda-s)-\frac{1}{6}f(0)}{\mathcal{N}}-\eta,\qquad v_{2}=\frac{F(\mathrm{hi}_{2}-s)-\frac{\phi_{2}}{2}f(0)}{\mathcal{N}}-\eta,
vf=F(b−s)−ϕ12f(0)−1rcCFN−η,v_{f}=\frac{F(b-s)-\frac{\phi_{1}}{2}f(0)-\mathbf{1}_{\mathrm{rc}}C_{F}}{\mathcal{N}}-\eta,

with ϕ1=14\phi_{1}=\frac{1}{4} for a real χ1\chi_{1} and ϕ1=13\phi_{1}=\frac{1}{3} otherwise, ϕ2=14\phi_{2}=\frac{1}{4} if n2=1n_{2}=1 and ϕ2=13\phi_{2}=\frac{1}{3} otherwise, and 1rc\mathbf{1}_{\mathrm{rc}} the indicator of type rc; the correlation term d=kη(d0+η)d=k_{\eta}(d_{0}+\eta); and the diagonals

typeDo/kηD_{o}/k_{\eta}Df/kηD_{f}/k_{\eta}
rr1−d0+cG1-d_{0}+c_{G}1−d01-d_{0}
rc1−d0+2cG1-d_{0}+2c_{G}1−d01-d_{0}
complex1−d0+2cG1-d_{0}+2c_{G}1−d0+cG1-d_{0}+c_{G}

caption:

:::

Theorem 9.4 (Two-test inequality). Use the two-test data defined above, with γG>γZ>0\gamma_{G}>\gamma_{Z}>0 and s≤min⁡(p∗,l2)s\leq\min(p^{*},l_{2}). For every sufficiently large qq and every configuration of an inside leaf, form the following entries:

  1. any finite set of distinct ordinary characters, each at a zero ρχ∈R(l)\rho_{\chi}\in R(l) with ∣Im⁡ρχ∣≤1|\operatorname{Im}\rho_{\chi}|\leq1;

  2. the first character at ρ1\rho_{1}, and also χ‾1\overline{\chi}_{1} at ρ‾1\overline{\rho}_{1} when χ1\chi_{1} is nonreal;

  3. the n2n_{2} characters of the reserved family at their height-one representatives, if a family is reserved.

For type rc, the single first-family entry retains both ρ1\rho_{1} and ρ‾1\overline{\rho}_{1}. Assign features v(λρχ)v(\lambda_{\rho_{\chi}}), vfv_{f}, and v2v_{2} to groups (i)–(iii), and diagonals DoD_{o}, DfD_{f}, and DoD_{o}, respectively. Then for every choice of real coefficients aj≥0a_{j}\geq0,

(∑jajvj)+2≤∑jDjaj2+d(∑jaj)2.\left(\sum_{j}a_{j}v_{j}\right)_{+}^{2} \leq \sum_{j}D_{j}a_{j}^{2} +d\left(\sum_{j}a_{j}\right)^{2}.

Proof. Put σs=1−s/L\sigma_{s}=1-s/\mathcal{L} and, in the Hilbert space of sequences indexed by nn,

Aj(n)=Λ(n)n−σsg(tn) χj(n)n−iγj,B(n)=χ0(n)Λ(n)n−σs f(tn)g(tn).A_{j}(n)=\sqrt{\Lambda(n)n^{-\sigma_{s}}g(t_{n})}\,\chi_{j}(n)n^{-i\gamma_{j}}, \qquad B(n)=\chi_{0}(n)\sqrt{\Lambda(n)n^{-\sigma_{s}}}\,\frac{f(t_{n})}{\sqrt{g(t_{n})}}.

Then −Re⁡⟨Aj,B⟩=LRj-\operatorname{Re}\langle A_{j},B\rangle=\mathcal{L}R_{j}, where RjR_{j} is the response of the entry at the anchor ss, and L∑jajRj≤∥∑jajAj∥∥B∥\mathcal{L}\sum_{j}a_{j}R_{j}\leq\left\|\sum_{j}a_{j}A_{j}\right\|\|B\|.

Responses. We use Lemma 8.7 (or directly Input 3.4). For an ordinary or reserved character every zero in the disc has parameter at least ss (Lemma 9.2(b), as s≤l2s\leq l_{2}), so Rj≥F(λρχ−s)−ϕ2f(0)−εR_{j}\geq F(\lambda_{\rho_{\chi}}-s)-\frac{\phi}{2}f(0)-\varepsilon. For the first family, the zeros of χ1\chi_{1} with parameter below ss are ρ1\rho_{1} and, for type rc, ρ‾1\overline{\rho}_{1} (the others have parameter at least λ′≥p∗≥s\lambda'\geq p^{*}\geq s); keeping them gives F(λ1−s)≥F(b−s)F(\lambda_{1}-s)\geq F(b-s) and, for type rc, Re⁡F(λ1−s+2iμ1)≥−CF\operatorname{Re}F(\lambda_{1}-s+2i\mu_{1})\geq-C_{F} (Lemma 9.1(v)). Conjugate zeros of a real character have the same multiplicity m≥1m\geq1, so the kept terms total at least m(F(b−s)−1rcCF)m(F(b-s)-\mathbf{1}_{\mathrm{rc}}C_{F}), which is at least F(b−s)−1rcCFF(b-s)-\mathbf{1}_{\mathrm{rc}}C_{F} when this is nonnegative (otherwise vf<0v_{f}<0, see below). Since φ(χ)=14\varphi(\chi)=\frac{1}{4} for real χ\chi and φ(χ)≤13\varphi(\chi)\leq\frac{1}{3} always (Input 3.4), the terms φ2f(0)\frac{\varphi}{2}f(0) are at most 16f(0)\frac{1}{6}f(0), ϕ12f(0)\frac{\phi_{1}}{2}f(0) and ϕ22f(0)\frac{\phi_{2}}{2}f(0) respectively. Thus Rj≥N(vj+η)−εR_{j}\geq\mathcal{N}(v_{j}+\eta)-\varepsilon for entries with positive feature. Entries with nonpositive feature may be omitted in estimating the positive part on the left; we justify restoring their coefficients to the right-hand side below.

Norms. By Input 3.3 applied to f2/gf^{2}/g, ∥B∥2=LRZ(1+o(1))\|B\|^{2}=\mathcal{L}R_{Z}(1+o(1)), and ∥Aj∥2=LRG(1+o(1))\|A_{j}\|^{2}=\mathcal{L}R_{G}(1+o(1)).

Off-diagonal terms. For j≠kj\neq k the characters are distinct and ⟨Aj,Ak⟩\langle A_{j},A_{k}\rangle is the gg-weighted sum for the nonprincipal character χjk=χjχ‾k\chi_{jk}=\chi_{j}\overline{\chi}_{k} at σs+i(γj−γk)\sigma_{s}+i(\gamma_{j}-\gamma_{k}). By Input 3.4 its real part is at most L(d0RG+ε)\mathcal{L}(d_{0}R_{G}+\varepsilon) plus L\mathcal{L} times the sum of −Re⁡G(⋅)-\operatorname{Re}G(\cdot) over the zeros of χjk\chi_{jk} in the disc with parameter below ss. If χjk∉{χ1,χ‾1}\chi_{jk}\notin\{\chi_{1},\overline{\chi}_{1}\} there are none, by Lemma 9.2(b), which applies because ∣γj−γk∣≤2≤2l|\gamma_j-\gamma_k|\le2\le2l. If χjk=χ1\chi_{jk}=\chi_1 (or χ‾1\overline{\chi}_1), they are ρ1\rho_1 (resp. ρ‾1\overline{\rho}_1), and for type rc both, each contributing at most CGC_G; these zeros are simple, since a multiple ρ1\rho_1 would give λ1=λ′≥s\lambda_1=\lambda'\ge s. For a given jj, the kk with χjχ‾k∈{χ1,χ‾1}\chi_j\overline{\chi}_k\in\{\chi_1,\overline{\chi}_1\} are χk∈{χjχ‾1,χjχ1}\chi_k\in\{\chi_j\overline{\chi}_1,\chi_j\chi_1\}: at most two, and at most one if χ1\chi_1 is real; a first-family character has none if χ1\chi_1 is real (χ1χ‾1=χ0\chi_1\overline{\chi}_1=\chi_0) and at most one (χ12\chi_1^2) otherwise. With 2ajak≤aj2+ak22a_ja_k\le a_j^2+a_k^2 for each such pair, these corrections give the diagonals of the table.

Assembly. For a≥0a\ge0, ∥∑jajAj∥2≤LRG[∑jaj2(1−d0+exc⁡j)+(d0+ε)(∑jaj)2]\|\sum_j a_jA_j\|^2\le LR_G[\sum_j a_j^2(1-d_0+\operatorname{exc}_j)+(d_0+\varepsilon)(\sum_j a_j)^2], and ∥B∥2≤(1+ε)LRZ\|B\|^2\le(1+\varepsilon)LR_Z. Dividing by L2N2L^2\mathcal{N}^2 and absorbing the errors in η\eta gives the claim for the positive-feature entries. To restore any omitted entries, expand the right-hand side as ∑j(Dj+d)aj2+2d∑j<kajak\sum_j(D_j+d)a_j^2+2d\sum_{j<k}a_ja_k. Its coefficients are nonnegative: the table gives Dj+d≥kη(1+η)>0D_j+d\ge k_\eta(1+\eta)>0 and d>0d>0. Restoring nonnegative aja_j therefore increases this bound, while a nonpositive feature can only decrease the positive part on the left. □

Proposition 9.5 (Detector mixture). Theorem 9.4 remains true when ff is replaced by fϵ=f+ϵgf_\epsilon=f+\epsilon g with ϵ≥0\epsilon\ge0, provided FF, f(0)f(0) and CFC_F are replaced by F+ϵGF+\epsilon G, f(0)+ϵg(0)f(0)+\epsilon g(0) and CF+ϵCGC_F+\epsilon C_G, and N\mathcal{N} by Nϵ=RG(RZ+2ϵF(−s)+ϵ2RG)\mathcal{N}_\epsilon=\sqrt{R_G(R_Z+2\epsilon F(-s)+\epsilon^2R_G)}; the Gram quantities are unchanged.

Proof. fϵf_\epsilon satisfies Conditions 1 and 2, and fϵ2/g=f2/g+2ϵf+ϵ2gf_\epsilon^2/g=f^2/g+2\epsilon f+\epsilon^2g is a sum of functions satisfying Condition 1, so ∥Bϵ∥2=L(RZ+2ϵF(−s)+ϵ2RG)(1+o(1))\|B_\epsilon\|^2=L(R_Z+2\epsilon F(-s)+\epsilon^2R_G)(1+o(1)). The rest of the proof is unchanged, using −Re⁡(F+ϵG)≤CF+ϵCG-\operatorname{Re}(F+\epsilon G)\le C_F+\epsilon C_G on Re⁡z≥a−s\operatorname{Re}z\ge a-s. □

Proposition 9.6 (Paired first family). Use the data and inside-leaf hypotheses of Theorem 9.4. Suppose χ1\chi_1 is nonreal and the specification s\mathfrak{s} has a finite gap λ′∈[p∗,h]\lambda'\in[p^*,h], where h=hi′h=h_{i'}. Fix ζ>0\zeta>0 and a number B\mathcal{B} with

Re⁡G(−s+iΔ)RG−ζRe⁡F(λ1−s+iΔ)+Re⁡F(λ′−s+iΔ)2N≤B(17)\frac{\operatorname{Re}G(-s+i\Delta)}{R_G} -\zeta\frac{\operatorname{Re}F(\lambda_1-s+i\Delta)+\operatorname{Re}F(\lambda'-s+i\Delta)}{2\mathcal{N}} \le\mathcal{B} \tag*{(17)}

for all real Δ\Delta, λ1∈[a,b]\lambda_1\in[a,b] and λ′∈[p∗,h]\lambda'\in[p^*,h]. Let e≥CF/(2N)e\ge C_F/(2\mathcal{N}), and put

m=F(b−s)+F(h−s)2N−f(0)6N−η−e,D∗=kη{1+B2−d0+cG−ζe2},ζ∗=kηζ2.m=\frac{F(b-s)+F(h-s)}{2\mathcal{N}}-\frac{f(0)}{6\mathcal{N}}-\eta-e,\qquad D_* = k_\eta\left\{\frac{1+\mathcal{B}}{2}-d_0+c_G-\frac{\zeta e}{2}\right\},\qquad \zeta_*=\frac{k_\eta\zeta}{2}.

Let m′≤mm'\le m and D′≥D∗D'\ge D_* with D′>0D'>0 and 2D′≥ζ∗max⁡(m′,0)2D'\ge\zeta_*\max(m',0). Then, for every configuration of the leaf, there is τ∈[0,d]\tau\in[0,\sqrt d] satisfying (T) for the entries of Theorem 9.4, with the first-family feature m′m' and diagonal D′D' in place of vfv_f and DfD_f.

Proof. For χ1\chi_1 use the averaged vector Af=12(Aχ1,γ1+Aχ1,γ′)A_f=\frac12(A_{\chi_1,\gamma_1}+A_{\chi_1,\gamma'}), with an admissible ρ′\rho' of height γ′\gamma' (Section 2; so ρ′\rho' is a zero of χ1\chi_1 and ∣γ′∣≤l|\gamma'|\le l), and its conjugate for χ‾1\overline{\chi}_1. Its response is the average of the two responses, which we bound by (12) at the points 1−s/L+iγ11-s/\mathcal{L}+i\gamma_1 and 1−s/L+iγ′1-s/\mathcal{L}+i\gamma'; they lie in R(9l)R(9l), and (12) does not require the kept zeros to lie in a disc. The set A1A_1 of zeros of χ1\chi_1 in R(l)R(l) with parameter below ss is contained in {ρ1}\{\rho_1\}, since every occurrence in R(l)R(l) other than the distinguished one has parameter at least λ′≥p∗≥s\lambda'\ge p^*\ge s. We put ρ′\rho' into A2A_2, and also ρ1\rho_1 if λ1≥s\lambda_1\ge s, with their multiplicities (if ρ′=ρ1\rho'=\rho_1 this is ρ1\rho_1 with multiplicity at least 2); by Input 3.9 their number is bounded independently of qq. The multiplicities can only increase the response: the terms of ρ′\rho' are ≥0\ge0 by Condition 2 because λ′≥s\lambda'\ge s, the term F(λ1−s)F(\lambda_1-s) is positive, and if ρ1\rho_1 is a multiple zero then λ1=λ′≥s\lambda_1=\lambda'\ge s, so its terms are ≥0\ge0 as well. With Δ=L(γ1−γ′)\Delta=\mathcal{L}(\gamma_1-\gamma'), P=Re⁡G(−s+iΔ)/RGP=\operatorname{Re}G(-s+i\Delta)/R_G and QQ the second fraction in (17), the feature of AfA_f is therefore at least F(b−s)+F(h−s)2N−f(0)6N+Q−η\frac{F(b-s)+F(h-s)}{2\mathcal{N}}-\frac{f(0)}{6\mathcal{N}}+Q-\eta, and ∥Af∥2=LRG1+P2(1+o(1))\|A_f\|^2=LR_G\frac{1+P}{2}(1+o(1)) by Input 3.3. Each off-diagonal term ⟨Af,Ak⟩\langle A_f,A_k\rangle is the average of two terms of the kind estimated in the proof of Theorem 9.4, at the heights γ1−γk\gamma_1-\gamma_k and γ′−γk\gamma'-\gamma_k, whose absolute values are at most l+1≤2ll+1\le2l; the term between AfA_f and its conjugate is an average of four terms, at heights 2γ12\gamma_1, γ1+γ′\gamma_1+\gamma' and 2γ′2\gamma', which are also at most 2l2l in absolute value. So Lemma 9.2 applies to all of them, averaging creates no new quotient characters, and the off-diagonal estimates are unchanged (for a cubic χ1\chi_1 the quotient χ12=χ1‾\chi_{1}^{2}=\overline{\chi_{1}} of the conjugate pair is paid by the cGc_{G} in D∗D_{*}). Put U=Q+eU=Q+e; then U≥0U\geq0, since λ′≥s\lambda'\geq s gives Q≥−CF/(2N)Q\geq-C_{F}/(2N).

By (17), P≤B+ζQP\leq B+\zeta Q, so the actual first-family feature is at least m+Um+U and the actual diagonal kη{1+P2−d0+cG}k_{\eta}\left\{\frac{1+P}{2}-d_{0}+c_{G}\right\} is at most D∗+ζ∗UD_{*}+\zeta_{*}U. Put U′=U+(m−m′)≥U≥0U'=U+(m-m')\geq U\geq0. Then the actual feature is at least m′+U′m'+U', and the actual diagonal is at most D′+ζ∗U′D'+\zeta_{*}U', which is positive. Replacing the first-family features and diagonals by these bounds preserves the quadratic inequality of Theorem 9.4, and Proposition 8.8 gives τ∈[0,d]\tau\in[0,\sqrt{d}] satisfying (T) with the first-family data (m′+U′,D′+ζ∗U′)(m'+U',D'+\zeta_{*}U'). Finally, for τ≥0\tau\geq0 and U′≥0U'\geq0,

(m′+U′−τ)+2D′+ζ∗U′≥(m′−τ)+2D′.\frac{(m'+U'-\tau)_{+}^{2}}{D'+\zeta_{*}U'}\geq\frac{(m'-\tau)_{+}^{2}}{D'}.

Indeed the right side is 0 if m′≤τm'\leq\tau. Otherwise put z=m′−τ∈(0,m′]z=m'-\tau\in(0,m']. The derivative of (z+u)2/(D′+ζ∗u)(z+u)^{2}/(D'+\zeta_{*}u) in u≥0u\geq0 is (z+u)(2D′−ζ∗z+ζ∗u)/(D′+ζ∗u)2≥0(z+u)(2D'-\zeta_{*}z+\zeta_{*}u)/(D'+\zeta_{*}u)^{2}\geq0, because 2D′≥ζ∗m′≥ζ∗z2D'\geq\zeta_{*}m'\geq\zeta_{*}z. So (T) holds at the same τ\tau with (m′,D′)(m',D'). □\square

In the leaf program the two-test row is the threshold form (T) of Theorem 9.4 (or of one of its two variants, Propositions 9.5 and 9.6):

∑ixi(Vi−τ)+2Do+n(vf−τ)+2Df+n2(v2−τ)+2Do≤1−τ2d,\sum_{i}x_{i}\frac{(V_{i}-\tau)_{+}^{2}}{D_{o}}+n\frac{(v_{f}-\tau)_{+}^{2}}{D_{f}}+n_{2}\frac{(v_{2}-\tau)_{+}^{2}}{D_{o}}\leq1-\frac{\tau^{2}}{d},

with Vi=v(hii)V_{i}=v(\mathrm{hi}_{i}) the feature of the right end of the ordinary bin ii (the tail and the reserved columns have feature 0 in this row), and negative features replaced by 0. The T∗T^{*}-representatives are admissible entries, since Theorem 9.4 allows any zero of height at most 1, and all of them have parameter at least l2≥sl_{2}\geq s (Z3). The reserved family appears as a fixed term with the feature v2v_{2} at hi2\mathrm{hi}_{2}, which bounds its feature over the whole reserved range, and with the diagonal DoD_{o}. The row is used on 37 roots, once on each: with a single test on 31 of them, with a mixture on 2, and in the paired form on 4 (all with a nonreal χ1\chi_{1}, as Proposition 9.6 requires). In the paired form the certified pair (m′,D′)(m',D') consists of a lower bound for mm and an upper bound for D∗D_{*}, and the condition 2D′≥ζ∗max⁡(m′,0)2D'\geq\zeta_{*}\max(m',0) is checked for this pair; the margins 2D′−ζ∗m′2D'-\zeta_{*}m' are between 0.18 and 0.30. The condition (17) is verified with interval arithmetic. Both of its terms are even in Δ\Delta, so Δ≥0\Delta\geq0 suffices. On the grid Δ∈1200Z∩[0,30]\Delta\in\frac{1}{200}\mathbb{Z}\cap[0,30] the left side is enclosed at the four corners of the parameter box [a,b]×[p∗,h][a,b]\times[p^{*},h]. As the left side is a sum of a function of λ1\lambda_{1} and a function of λ′\lambda', its maximum over the box exceeds the maximum over the corners by at most ζ2N((b−a)28∫t2f(t)e(s−a)+t dt+(h−p∗)28∫t2f(t) dt)\frac{\zeta}{2N}\left(\frac{(b-a)^{2}}{8}\int t^{2}f(t)e^{(s-a)+t}\,\mathrm{d}t+\frac{(h-p^{*})^{2}}{8}\int t^{2}f(t)\,\mathrm{d}t\right), the linear interpolation error in each variable, and this is added. Between grid points a second-derivative bound in Δ\Delta is used, and for Δ≥30\Delta\geq30 a closed-form bound from Lemma 9.1(iv). The value BB used is at least 0.

Validity of the rows.

Proposition 9.7 (Every configuration satisfies the near rows). Let ss specify a leaf or one of its subcases. For sufficiently large qq, take any configuration in C(s)C(s) and assign its characters to columns as in Proposition 11.6, obtaining column values xx. Then, for every near row kk, there is a threshold τk\tau_{k} such that

0≤τk≤dk/S,Φk(x,τk)≤1−τk2dk/S.0\leq\tau_{k}\leq\sqrt{d_{k}/S},\qquad\Phi_{k}(x,\tau_{k})\leq1-\frac{\tau_{k}^{2}}{d_{k}/S}.

Here dkd_{k} is the stored correlation term, S=1016S=10^{16}, and Φk\Phi_{k} is the row expression of Definition 10.1. Each row may have its own threshold.

Proof. For a graded, family or shifted row, take as entries the characters of the configuration that are placed in the columns and family terms of the row: the ordinary characters of OO at their T∗T^*-representatives, the reserved characters, and the first family, with the entries of Section 9.4. These form an admissible family of entries for Theorem 8.4, with the kept sets verified in Section 9.4, so Corollary 8.10 gives a threshold τ\tau at which (T)(T) holds with the normalized features, diagonals and correlation term. Each column’s feature is at most the feature of every character placed in it (features are taken at the right end of the column, and FF is decreasing), or is 00; each column’s diagonal is at least theirs, or the column has feature 00 (an entry whose diagonal numerator is not safely positive is omitted, and its column is given feature 00), in which case it contributes nothing for τ≥0\tau\ge0; each family term (n,v,D)(n,v,D) receives nn entries with feature at least vv and diagonal at most DD; and the characters in the tail column and in the hidden tail are simply omitted from the row. Since each term (v−τ)+2/D(v-\tau)_{+}^{2}/D increases with vv and decreases with DD for τ≥0\tau\ge0, the row of the leaf holds at τ\tau. For the two-test row the same argument applies to Theorem 9.4 and its variants, through Proposition 8.8. □\square

The numbers IuI_u, DuD_u, D(δ)D(\delta), the transforms and the special terms RloR_{\mathrm{lo}}, chic^{\mathrm{hi}}, RpR_p, RqR_q, CZC_Z and CFC_F are computed from the rational parameters of each row with outward interval arithmetic. The transforms are enclosed with the closed form of Lemma 9.1(iv) or with its power series near 00. The infima and suprema over heights in RloR_{\mathrm{lo}}, chic^{\mathrm{hi}}, RpR_p and RqR_q are enclosed on grids of step at most 150\frac{1}{50}, with second-derivative bounds ∫t2f^c(t)edt dt\int t^2\widehat{f}_c(t)e^{dt}\,\mathrm{d}t between grid points and Lipschitz bounds in the real parts, and by the tail bounds of Lemma 9.1(iv) beyond the grid. The constants CZC_Z, CFC_F and CGC_G are enclosed on the line Re⁡z=−d\operatorname{Re} z=-d (by Lemma 9.1(v)), for heights in [0,20][0,20] by adaptive linear interpolation, starting from 200200 cells of width 110\frac{1}{10} and bisecting cells until the largest cell bound is within a small fixed tolerance of the best value found, with the same second-derivative bound, and beyond height 2020 by a closed-form tail bound. The features, diagonals and correlation terms of all rows other than the 3737 two-test rows have been recomputed from their rational parameters by an independent program whose soundness is formally verified (Section 14).

Leaf programs and exact certificates

For each case of the zero analysis, we construct a family of linear programs indexed by the thresholds in the near rows. Their common objective bounds the zero sum. We cover the threshold domain by finitely many boxes and give a linear relaxation on each box. An exact certificate then bounds every feasible point, without having to find an optimal one.

The argument uses weak linear-programming duality [25, 115] with outward rounding of the data; compare [63, 95, 1]. This section is independent of the analytic number theory. Its conclusion is the certificate soundness theorem, Theorem 10.7, which is formally verified (Section 14.3).

Leaf data

The coefficients, budgets, and objective constants of a leaf are stored as integers at scale

S=1016.S=10^{16}.

Thus a stored coefficient vv represents v/Sv/S. Character counts such as nn, ngn_g, and nEn_E are ordinary unscaled integers. The variables xcx_c and thresholds τk\tau_k are also unscaled real numbers.

A leaf has a finite set of columns cc and a finite list of near rows k=1,…,Kk=1,\ldots,K. A column cc carries

  • an objective coefficient GcG_c and a far weight WcW_c;

  • a count coefficient CcC_c, a hidden-count coefficient NcN_c and a second-family indicator EcE_c (stored at scale SS: Cc,Ec∈{0,S}C_c,E_c\in\{0,S\}, and NcN_c is 00, SS or, for the hidden tail, ⌊S/w(end)⌋\lfloor S/w(\mathrm{end})\rfloor, with end as in Section 11.5);

  • for each row kk, a feature vckv_{ck} and a diagonal Dck>0D_{ck}>0.

A row kk carries a correlation term dk>0d_k>0 and finitely many family terms (n,v,D)(n,v,D), where n≥0n\ge0 is a count, vv a feature and D>0D>0 a diagonal. The leaf also has a far budget FF, counts ngn_g (the number nn of first-family characters, in outside leaves) and nEn_E (the number of reserved characters carried by second-family columns: the size of the reserved family if it is charged through columns, and 0 otherwise), and two constants: first\mathrm{first}, which bounds the cost of the first family of an inside leaf (0 for an outside leaf) plus any fixed charge for the reserved family, and final\mathrm{final}, the error allowance, with final/S=5η\mathrm{final}/S=5\eta (Section 11.5).

Definition 10.1 (Leaf program). For real column values x=(xc)x=(x_c) and real thresholds τ=(τk)\tau=(\tau_k), put

Φk(x,τk)=∑cxc(vck/S−τk)+2Dck/S+∑(n,v,D)∈fam(k)n(v/S−τk)+2D/S.\Phi_k(x,\tau_k)=\sum_c x_c\frac{(v_{ck}/S-\tau_k)_+^2}{D_{ck}/S} +\sum_{(n,v,D)\in\mathrm{fam}(k)}n\frac{(v/S-\tau_k)_+^2}{D/S}.

The pair (x,τ)(x,\tau) is feasible if all of the following hold:

  • xc≥0x_c\ge0 for every cc;

  • (far) ∑cWcxc≤F\sum_c W_cx_c\le F;

  • (count) ∑cCcxc≤2S\sum_c C_cx_c\le2S;

  • (hidden count) ∑cNcxc≤ngS\sum_c N_cx_c\le n_gS;

  • (second family) ∑cEcxc=nES\sum_c E_cx_c=n_ES;

  • (near rows) for every kk: 0≤τk0\le\tau_k, τk2≤dk/S\tau_k^2\le d_k/S and Φk(x,τk)≤1−τk2/(dk/S)\Phi_k(x,\tau_k)\le1-\tau_k^2/(d_k/S).

The value of xx is

V(x)=1S(first+final+∑cGcxc).\mathcal{V}(x)=\frac{1}{S}\left(\mathrm{first}+\mathrm{final}+\sum_c G_cx_c\right).

For fixed τ\tau, these are linear constraints on xx. Allowing τ\tau to vary gives the family of programs to be certified. The variables xcx_c are real and nonnegative: integer character counts are feasible values, but the relaxation also permits fractional counts and the weighted tail masses defined in Section 11.5.

In Section 11 we show that for every configuration of zeros the actual characters give a feasible pair whose value bounds the zero sum WW of Section 4. It is therefore enough to show that every feasible pair has value below 1.

Relaxations of a near row

Fix a row kk and write d=dk/Sd=d_k/S. The near-row constraint says Qk(τ)≤1Q_k(\tau)\le1, where

Qk(τ)=∑jcj(uj−τ)+2+τ2dQ_k(\tau)=\sum_j c_j(u_j-\tau)_+^2+\frac{\tau^2}{d}

with cj≥0c_j\ge0 collecting the coefficients xc/(Dck/S)x_c/(D_{ck}/S) and n/(D/S)n/(D/S) and uju_j the features. The function QkQ_k is convex and continuously differentiable, with Qk′(τ)=−2∑jcj(uj−τ)++2τ/dQ_k'(\tau)=-2\sum_j c_j(u_j-\tau)_++2\tau/d. In the next two lemmas, d>0d>0, all cj≥0c_j\ge0, and 0≤a≤b0\le a\le b.

Lemma 10.2 (First-order relaxation). If a≤τ≤ba\le\tau\le b and Qk(τ)≤1Q_k(\tau)\le1, then

∑jcj(uj−b)+2≤1−a2d.\sum_j c_j(u_j-b)_+^2\le1-\frac{a^2}{d}.

Proof. Each (uj−τ)+2≥(uj−b)+2(u_j-\tau)_+^2\ge(u_j-b)_+^2 since τ≤b\tau\le b, and τ2≥a2\tau^2\ge a^2 since 0≤a≤τ0\le a\le\tau. □\square

Lemma 10.3 (Tangent relaxation). Let m=(a+b)/2m=(a+b)/2 and h=(b−a)/2h=(b-a)/2. If a≤τ≤ba\leq\tau\leq b and Qk(τ)≤1Q_k(\tau)\leq1, then at least one of the following holds:

(case 0)∑jcj(((uj−m)++h)2−h2)≤1−(m−h)2−h2d,\text{(case 0)}\qquad\sum_j c_j\left(\left((u_j-m)_{+}+h\right)^2-h^2\right)\leq1-\frac{(m-h)^2-h^2}{d},
(case 1)∑jcjθ~j≤1−(m+h)2−h2d,θ~j={((uj−m)−h)2−h2,uj>m,0,uj≤m.\text{(case 1)}\qquad\sum_j c_j\widetilde{\theta}_j\leq1-\frac{(m+h)^2-h^2}{d},\qquad \widetilde{\theta}_j= \begin{cases} \left((u_j-m)-h\right)^2-h^2, & u_j>m,\\ 0, & u_j\leq m. \end{cases}

Proof. By convexity, 1≥Qk(τ)≥Qk(m)+Qk′(m)(τ−m)≥Qk(m)−h∣Qk′(m)∣1\geq Q_k(\tau)\geq Q_k(m)+Q'_k(m)(\tau-m)\geq Q_k(m)-h|Q'_k(m)|, so Qk(m)−hQk′(m)≤1Q_k(m)-hQ'_k(m)\leq1 or Qk(m)+hQk′(m)≤1Q_k(m)+hQ'_k(m)\leq1. Expanding with (u−m)+2±2h(u−m)+=((u−m)+±h)2−h2(u-m)_{+}^{2}\pm2h(u-m)_{+}=((u-m)_{+}\pm h)^2-h^2 and m2∓2hm=(m∓h)2−h2m^2\mp2hm=(m\mp h)^2-h^2 gives the two cases. In case 1 a term with uj≤mu_j\leq m is 0. □\square

The left-hand sides are linear in the cjc_j, hence in the column values xx. In case 1 the coefficients θ~j\widetilde{\theta}_j can be negative.

Certificates

Thresholds are measured in ticks of 1/TS1/T_S, where TS=106T_S=10^6, and dual multipliers are scaled by DS=1012D_S=10^{12}. In this subsection [a,b][a,b] denotes an interval of thresholds in ticks, not a first-zero cell, and m,h,pm,h,p and β\beta are local to the relaxation formulas.

Definition 10.4 (Root threshold box). For row kk, put

ek=⌈⌊dkTS2/S⌋+1⌉.e_k=\left\lceil\sqrt{\left\lfloor d_kT_S^2/S\right\rfloor+1}\right\rceil.

The root box is ∏k[0,ek]\prod_k[0,e_k] in integer-tick coordinates. Its corresponding domain for real thresholds is ∏k[0,ek/TS]\prod_k[0,e_k/T_S].

Since ek2S≥dkTS2e_k^2S\geq d_kT_S^2, every threshold with τk2≤dk/S\tau_k^2\leq d_k/S lies in [0,ek/TS][0,e_k/T_S].

Definition 10.5 (Integer costs and budgets). Let [a,b][a,b] be an interval of integer ticks with 0≤a<b0\leq a<b, and let vv and D>0D>0 be a stored feature and diagonal. The following costs are coefficients of column variables; the family terms are subtracted from the row budget.

(a) (First order.) The cost is ⌊(v−bS/TS)+2/D⌋\left\lfloor(v-bS/T_S)_{+}^{2}/D\right\rfloor. The budget of row kk is

S−⌊S2a2TS2dk⌋−∑(n,v,D)⌊n(v−bS/TS)+2D⌋.S-\left\lfloor\frac{S^2a^2}{T_S^2d_k}\right\rfloor-\sum_{(n,v,D)}\left\lfloor\frac{n(v-bS/T_S)_{+}^{2}}{D}\right\rfloor.

(b) (Tangent, case ι∈{0,1}\iota\in\{0,1\}.) Put m=(a+b)/(2TS)m=(a+b)/(2T_S), h=(b−a)/(2TS)h=(b-a)/(2T_S) and vˉ=v−mS\bar v=v-mS (mSmS and 2hS2hS are integers, since 2TS2T_S divides SS). The cost is ⌊SΘ/D⌋\left\lfloor S\Theta/D\right\rfloor, where Θ=0\Theta=0 if vˉ≤0\bar v\leq0, and otherwise Θ=⌊vˉ(vˉ+2hS)/S⌋\Theta=\left\lfloor\bar v(\bar v+2hS)/S\right\rfloor for ι=0\iota=0 and Θ=⌊vˉ(vˉ−2hS)/S⌋\Theta=\left\lfloor\bar v(\bar v-2hS)/S\right\rfloor for ι=1\iota=1. The budget of row kk is ⌈Sβ⌉\left\lceil S\beta\right\rceil, where β\beta is the exact rational number

β=1−p2−h2dk/S−∑(n,v,D)nΘι(v/S)D/S,\beta=1-\frac{p^2-h^2}{d_k/S}-\sum_{(n,v,D)}n\frac{\Theta_{\iota}(v/S)}{D/S},

with p=a/TSp=a/T_S for ι=0\iota=0 and p=b/TSp=b/T_S for ι=1\iota=1, and Θ0(u)=((u−m)++h)2−h2\Theta_0(u)=((u-m)_{+}+h)^2-h^2, Θ1(u)=((u−m)−h)2−h2\Theta_1(u)=((u-m)-h)^2-h^2 if u>mu>m and 0 otherwise.

The costs are the exact relaxation coefficients, multiplied by SS and rounded down. The budgets are at least the exact right-hand sides multiplied by SS: the tangent budget is rounded up, and in the first-order budget each subtracted term is rounded down.

Definition 10.6 (Certificate). A certificate for a leaf is a finite binary tree starting from the root box of Definition 10.4. An internal node splits one coordinate kk of its box [a,b][a,b] at an integer tick a<mid<ba<\mathrm{mid}<b, into [a,mid][a,\mathrm{mid}] and [mid,b][\mathrm{mid},b]. At a terminal box, each row uses either the first-order relaxation or the two alternatives of the tangent relaxation. A relaxation case chooses one alternative for each tangent row. All combinations must be certified, because the alternative supplied by Lemma 10.3 can depend on the feasible point.

For consistency with the stored data, the mode is encoded by ok=2o_k=2 for a first-order row and ok=1o_k=1 for a tangent row. The case vector has ιk=2\iota_k=2 when ok=2o_k=2, and ιk∈{0,1}\iota_k\in\{0,1\} when ok=1o_k=1. In the stored trees the halves of a split are listed in the order [a,mid][a,\mathrm{mid}], [mid,b][\mathrm{mid},b], the rows are numbered from 0, and the relaxation cases of a terminal box are listed in lexicographic order of (ι1,…,ιK)(\iota_1,\ldots,\iota_K). For every relaxation case the node records either

(i) an exclusion: a row kk whose costs are all ≥0\ge0 and whose budget is <0<0; or

(ii) integer duals YF,YC,YN,YE+,YE−≥0Y_F,Y_C,Y_N,Y_E^+,Y_E^-\ge0 and Z1,…,ZK≥0Z_1,\ldots,Z_K\ge0 such that for every column cc

DSGc≤YFWc+YCCc+YNNc+(YE+−YE−)Ec+∑kZk costck.(18)D_S G_c \le Y_F W_c+Y_C C_c+Y_N N_c+(Y_E^+-Y_E^-)E_c+\sum_k Z_k\,\mathrm{cost}_{ck}. \tag*{(18)}

The value of the relaxation case is then

⌊YFF+2SYC+ngSYN+(YE+−YE−)nES+∑kZk budgetkDS⌋+first+final.\left\lfloor\frac{Y_FF+2SY_C+n_gSY_N+(Y_E^+-Y_E^-)n_ES+\sum_k Z_k\,\mathrm{budget}_k}{D_S}\right\rfloor+\mathrm{first}+\mathrm{final}.

The value of a terminal box is the maximum of −1-1 and the values of its relaxation cases that carry duals (so it is −1-1 if every relaxation case is excluded), and the certificate is accepted if every terminal box has value <S<S. The data must be well formed: all diagonals and correlation terms are positive, and all family counts are nonnegative.

Theorem 10.7 (Soundness of the integer certificates). Let a leaf have positive diagonals and correlation terms and nonnegative family counts. If it has an accepted certificate in the sense of Definition 10.6, then every feasible pair (x,τ)(x,\tau) of Definition 10.1 satisfies V(x)<1\mathcal{V}(x)<1. This conclusion holds for all real nonnegative column values, including fractional ones.

Proof. Let (x,τ)(x,\tau) be feasible and put tk=τkTSt_k=\tau_kT_S. Since τk2≤dk/S\tau_k^2\le d_k/S we have tk≤ekt_k\le e_k, so tt lies in the root box. At every internal node tt lies in one of the two halves, so there is a terminal box ∏k[ak,bk]\prod_k[a_k,b_k] that contains tt.

For each row, Lemma 10.2 (if ok=2o_k=2) or Lemma 10.3 (if ok=1o_k=1) gives a case ιk\iota_k in which the exact relaxed inequality holds at xx. Take the relaxation case ι\iota so obtained. Multiply the relaxed inequality of row kk by SS. Each coefficient of xcx_c is then at least costck\mathrm{cost}_{ck} and the right side is at most budgetk\mathrm{budget}_k, by the rounding directions in Definition 10.5 and since xc≥0x_c\ge0 and n≥0n\ge0. Therefore

∑cxc costck≤budgetk(1≤k≤K).(19)\sum_c x_c\,\mathrm{cost}_{ck}\le\mathrm{budget}_k\qquad(1\le k\le K). \tag*{(19)}

Here we used S⋅(v/S−b/TS)+2/(D/S)=(v−bS/TS)+2/DS\cdot(v/S-b/T_S)_+^2/(D/S)=(v-bS/T_S)_+^2/D and the corresponding identities for the tangent terms.

If the relaxation case ι\iota is excluded through row kk, then the left side of (19) is ≥0\ge0 and the right side is <0<0, a contradiction. So ι\iota carries duals. Multiply (18) by xc≥0x_c\ge0 and sum over cc. Then use the far, count and hidden-count constraints, with multipliers YF,YC,YN≥0Y_F,Y_C,Y_N\ge0, the second-family equality, with the multiplier YE+−YE−Y_E^+-Y_E^- of either sign, and (19), with multipliers Zk≥0Z_k\ge0. This gives

DS∑cGcxc≤YFF+2SYC+ngSYN+(YE+−YE−)nES+∑kZk budgetk.D_S\sum_c G_cx_c\le Y_FF+2SY_C+n_gSY_N+(Y_E^+-Y_E^-)n_ES+\sum_k Z_k\,\mathrm{budget}_k.

Hence first ++ final +∑cGcxc+ \sum_{c} G_{c}x_{c} is at most the value of ι\iota, which is at most the value of the terminal box, which is <S< S. □\square

Remark 10.8. The formal proof of Theorem 10.7 brought out that the family counts nn must be nonnegative. With a negative count, the first-order relaxation can fail. The leaf builders already ensure n≥0n \ge0, and the checker requires it.

The case tree

This section describes the finite case analysis of the middle range 0.1≤λ1<1.50.1 \le\lambda_{1} < 1.5 and proves that it covers every configuration of zeros (Theorem 11.7). The analytic facts it uses are the zero-location results of Section 7, the costs of Section 5, the far budget of Section 6 and the near rows of Section 9. There are two tasks: showing that the cases cover every possible configuration, and showing that a configuration in a case gives a feasible point of its program. These are proved separately before being combined in Theorem 11.7. Table 10, after Section 11.4, lists the levels of the analysis, with their numbers and the places where they are defined and checked.

Configurations.

Definition 11.1 (Zero configuration). Fix a sufficiently large qq. Its zero configuration is the multiset Z\mathcal{Z} of pairs (χ,ρ)(\chi,\rho) with χ≠χ0\chi\ne\chi_{0} a character modulo qq, ρ∈R(l)\rho\in R(l) and L(ρ,χ)=0L(\rho,\chi)=0, each pair repeated according to the multiplicity of ρ\rho.

Minima over empty sets are +∞+\infty. In the middle range the first zero exists. The quantities of Section 2 are read off from Z\mathcal{Z}: the first zero (χ1,ρ1)(\chi_{1},\rho_{1}), the first family F1\mathcal{F}_{1} with n=∣F1∣n=|\mathcal{F}_{1}|, the parameters λ2,λ3\lambda_{2},\lambda_{3} and the family F2\mathcal{F}_{2} (Definition 2.1), and the distinguished occurrences, λ′\lambda' and an admissible ρ′\rho' (Definition 2.2); for type rc, μ1>0\mu_{1}>0. In addition we use the following.

  • The height-one parameter of χ\chi is ν(χ)=min⁡{λρ:(χ,ρ)∈Z, ∣γρ∣≤1}\nu(\chi)=\min\{\lambda_{\rho}:(\chi,\rho)\in\mathcal{Z},\ |\gamma_{\rho}|\le1\}, and the T∗T^{*}-parameter is νT∗(χ)\nu_{T^{*}}(\chi), defined in the same way with ∣γρ∣≤T∗|\gamma_{\rho}|\le T^{*} (Definition 6.4). Since L(ρ,χ)=0L(\rho,\chi)=0 if and only if L(ρˉ,χˉ)=0L(\bar{\rho},\bar{\chi})=0, both are functions of the family. We put ν∗=min⁡{ν(F):F≠F1}\nu_{*}=\min\{\nu(\mathcal{F}):\mathcal{F}\ne\mathcal{F}_{1}\}.

  • The configuration is inside if ρ1∈RB\rho_{1}\in R_{B} and outside otherwise. Since λ1<1.5<CB\lambda_{1}<1.5<C_{B} in the middle range, outside means ∣μ1∣>CB|\mu_{1}|>C_{B}. Type rr is always inside.

The definitions give λ1≤λ′\lambda_{1}\le\lambda', λ1≤λ2≤λ3\lambda_{1}\le\lambda_{2}\le\lambda_{3}, and νT∗(χ)≥ν(χ)≥λ2\nu_{T^{*}}(\chi)\ge\nu(\chi)\ge\lambda_{2} for χ∉F1\chi\notin\mathcal{F}_{1}.

Specifications.

A specification records interval information about a configuration. It may describe no actual configuration; in that event the corresponding case can be excluded. Its two lower bounds l2l_{2} and rr have different scopes: l2l_{2} applies throughout R(l)R(l), whereas rr applies only to the height-one representatives.

Definition 11.2 (Specification). A specification is a tuple

s=(t,[a,b],p,gap,l2,r,res,loc)\mathfrak{s}=(\mathrm{t},[a,b],p,\mathrm{gap},l_{2},r,\mathrm{res},\mathrm{loc})

with the following data:

  • a type t∈{rr,rc,complex}\mathrm{t}\in\{\mathrm{rr},\mathrm{rc},\mathrm{complex}\} and a first-zero cell [a,b][a,b], where a<ba<b;

  • a lower bound pp for λ′\lambda', and optionally an interval gap=[lo′,hi′]\mathrm{gap}=[\mathrm{lo}',\mathrm{hi}'] with p≤lo′<hi′≤∞p\le\mathrm{lo}'<\mathrm{hi}'\le\infty, called a gap (used only for type complex);

  • a global second-family lower bound l2l_{2} and a height-one lower bound rr, with l2≤r≤3l_{2}\le r\le3;

  • either no reservation, or a pair res=(hi2,n2)\mathrm{res}=(\mathrm{hi}_{2},n_{2}) with r<hi2≤2r<\mathrm{hi}_{2}\le2 and n2∈{1,2}n_{2}\in\{1,2\}; in the latter case put lo2=r\mathrm{lo}_{2}=r;

  • a height class loc∈{inside,outside}\mathrm{loc}\in\{\mathrm{inside},\mathrm{outside}\}.

The configuration set C(s)C(\mathfrak{s}) consists of the configurations for which some admissible choices of the first zero and the minimizing families satisfy all of the following:

  • (C1) the type is tt and λ1∈[a,b]\lambda_{1}\in[a,b];

  • (C2) λ′≥p\lambda'\ge p, and λ′∈[lo′,hi′]\lambda'\in[\mathrm{lo}',\mathrm{hi}'] if a gap is present (with λ′=∞\lambda'=\infty allowed when hi′=∞\mathrm{hi}'=\infty);

  • (C3) λ2≥l2\lambda_{2}\ge l_{2};

  • (C4) if res\mathrm{res} is none, then ν(F)≥r\nu(\mathcal{F})\ge r for every family F≠F1\mathcal{F}\ne\mathcal{F}_{1}; if res=(hi2,n2)\mathrm{res}=(\mathrm{hi}_{2},n_{2}), then ν∗∈[r,hi2]\nu_{*}\in[r,\mathrm{hi}_{2}] and some family Fres≠F1\mathcal{F}_{\mathrm{res}}\ne\mathcal{F}_{1} with ν(Fres)=ν∗\nu(\mathcal{F}_{\mathrm{res}})=\nu_{*} has exactly n2n_{2} characters (so it is a real character if n2=1n_{2}=1 and a pair of conjugate nonreal characters if n2=2n_{2}=2);

  • (C5) the configuration is inside if loc=inside\mathrm{loc}=\mathrm{inside} and outside if loc=outside\mathrm{loc}=\mathrm{outside}.

In the reserved case, (C4) implies ν(F)≥r\nu(\mathcal{F})\ge r for every F≠F1\mathcal{F}\ne\mathcal{F}_{1}. We call Fres\mathcal{F}_{\mathrm{res}} the reserved family and its characters reserved; the characters outside F1\mathcal{F}_{1} and Fres\mathcal{F}_{\mathrm{res}} are ordinary.

In a formula containing lo′\mathrm{lo}', that argument is omitted if no gap is present. Likewise, reserved-family terms are absent when there is no reservation.

The reserved family is a family with the least height-one parameter among the families other than the first; it is not assumed to be the family F2\mathcal{F}_{2} that attains λ2\lambda_{2}, which may have its nearest zero at a height above 1. The following three facts are used in every leaf.

Lemma 11.3 (Consequences of reserving a family). Let s\mathfrak{s} be a specification and fix a configuration and admissible choices witnessing its membership in C(s)C(\mathfrak{s}).

(i) (Count.) At most two characters χ∉F1\chi\notin\mathcal{F}_{1} have νT∗(χ)<λ3\nu_{T^{*}}(\chi)<\lambda_{3}.

(ii) (Drop.) If s\mathfrak{s} is reserved, then νT∗(χ)≥ν(χ)≥λ3\nu_{T^{*}}(\chi)\ge\nu(\chi)\ge\lambda_{3} for every ordinary character χ\chi.

(iii) (Cap.) If s\mathfrak{s} is reserved, then λ2≤ν∗≤hi2\lambda_{2}\le\nu_{*}\le\mathrm{hi}_{2}.

Proof. (i) Such a character has a zero in R(l)R(l) with parameter below λ3\lambda_{3}, so it lies in F2\mathcal{F}_{2}, and ∣F2∣≤2|\mathcal{F}_{2}|\le2. (ii) Let F\mathcal{F} be the family of an ordinary character. If F≠F2\mathcal{F}\ne\mathcal{F}_{2}, every zero of F\mathcal{F} in R(l)R(l) has parameter at least λ3\lambda_{3}. If F=F2\mathcal{F}=\mathcal{F}_{2}, then Fres∉{F1,F2}\mathcal{F}_{\mathrm{res}}\notin\{\mathcal{F}_{1},\mathcal{F}_{2}\}, so ν(Fres)≥λ3\nu(\mathcal{F}_{\mathrm{res}})\ge\lambda_{3}, and ν(F)≥ν∗=ν(Fres)\nu(\mathcal{F})\ge\nu_{*}=\nu(\mathcal{F}_{\mathrm{res}}) by minimality. (iii) The height-one representative of Fres\mathcal{F}_{\mathrm{res}} is a zero in R(l)R(l) of a character outside F1\mathcal{F}_{1}. □\square

The source cover

The roots of the case tree are 2768 specifications, each used with the height class inside, and the 1685 of them whose type is not rr used also with the height class outside: 4453 roots in all. They are constructed from the zero-location results of Section 7 in four steps.

  1. Parent rows. There are 58 parent rows (t,Aj,Bj,pj,rj)(t,A_{j},B_{j},p_{j},r_{j}): 33 of type rr tiling [0.1,1.5][0.1,1.5], 13 of type rc tiling [0.628,1.5][0.628,1.5], and 12 of type complex tiling [0.44,1.5][0.44,1.5]. Each is the statement

type t, λ1∈[Aj,Bj]⟹λ′≥pj  and  λ2≥rj,(20)\text{type }t,\ \lambda_{1}\in[A_{j},B_{j}] \quad\Longrightarrow\quad \lambda'\ge p_{j}\ \text{ and }\ \lambda_{2}\ge r_{j}, \tag*{(20)}

which is proved in Proposition 7.5 from the tables of Heath-Brown and Xylouris.

  1. Base cells. Each parent interval is cut into cells of width 0.01, giving 340 base cells.

  1. Cells and gaps. The base cells are tiled by 478 first-zero cells in all. A cell [lo,hi][\mathrm{lo},\mathrm{hi}] of the parent jj receives numbers p≤max⁡(pj,lo)p\le\max(p_{j},\mathrm{lo}) and l2≤max⁡(rj,lo)l_{2}\le\max(r_{j},\mathrm{lo}); this is valid since λ′,λ2≥λ1≥lo\lambda',\lambda_{2}\ge\lambda_{1}\ge\mathrm{lo}. A cell of type complex may be partitioned by gaps [p,m1],[m1,m2],…,[mk,∞][p,m_{1}],[m_{1},m_{2}],\ldots,[m_{k},\infty] for λ′\lambda'. The cells that are not partitioned, and the gaps of those that are, are the 684 gap cases.

  1. Records. Each gap case is either unreserved, with one specification with r=l2r=l_{2} and no reservation, or reserved with a chain l2=u0<u1<⋯<um=e≤2l_{2}=u_{0}<u_{1}<\cdots<u_{m}=e\le2, with specifications (res=(uj+1,n2), r=uj)(\mathrm{res}=(u_{j+1},n_{2}),\,r=u_{j}) for n2=1,2n_{2}=1,2 and 0≤j<m0\le j<m, followed by an unreserved tail specification with r=er=e.

By type there are 1083 specifications of type rr (195 unreserved, 444 reserved with n2=1n_{2}=1 and 444 with n2=2n_{2}=2), 115 of type rc (all unreserved and without gap) and 1570 of type complex.

Theorem 11.4 (Coverage by the root cases). Every configuration with 0.1≤λ1<1.50.1 \le\lambda_{1}<1.5 lies in C(s)C(s) for at least one of the 4453 root specifications ss.

Proof. Let t be the type. By Proposition 7.1, λ1>0.628\lambda_{1}>0.628 if t=rc\mathrm{t}=\mathrm{rc} and λ1>0.44\lambda_{1}>0.44 if t=complex\mathrm{t}=\mathrm{complex}, so λ1\lambda_{1} lies in some cell of type t (the cells are closed, and a boundary value lies in two cells). By [T2] for the parent of the cell, λ′≥p\lambda' \ge p and λ2≥l2\lambda_{2} \ge l_{2}. If the cell is partitioned, some gap contains λ′\lambda'. In the corresponding gap case, if it is unreserved, every family F≠F1\mathcal{F} \ne\mathcal{F}_{1} has ν(F)≥λ2≥l2=r\nu(\mathcal{F}) \ge\lambda_{2} \ge l_{2}=r. If it is reserved with chain u0<⋯<umu_{0}<\cdots<u_{m}, then ν∗≥λ2≥u0\nu_{*} \ge\lambda_{2} \ge u_{0}; if ν∗<e\nu_{*}<e we choose jj with ν∗∈[uj,uj+1]\nu_{*} \in[u_{j},u_{j+1}] and a minimizing family, which is a real character or a nonreal pair, and the configuration lies in the corresponding record; otherwise it lies in the tail record. Finally the height class is determined by ∣μ1∣|\mu_{1}|. □

Refinement trees

Each root is the root of a finite refinement tree. Every internal node carries a specification (computed from the root, never read from the data) and is of one of the following kinds. In each case every configuration of the node lies in the configuration set of one of its children, or the node has no configuration at all.

  • First-zero split at a<m<ba<m<b: children with cells [a,m][a,m] and [m,b][m,b], all other data kept. Exhaustive, since λ1≤m\lambda_{1} \le m or λ1≥m\lambda_{1} \ge m.

  • Second-family split at r<m<hi2r<m<\mathrm{hi}_{2} (reserved nodes): children (hi2≔m)(\mathrm{hi}_{2} \coloneqq m) and (r≔m, lo2≔m)(r \coloneqq m,\ \mathrm{lo}_{2} \coloneqq m). Exhaustive, since ν∗≤m\nu_{*} \le m or ν∗≥m\nu_{*} \ge m, and in the second case every family F≠F1\mathcal{F} \ne\mathcal{F}_{1} has ν(F)≥ν∗≥m\nu(\mathcal{F}) \ge\nu_{*} \ge m.

  • Gap split at lo′<m<hi′\mathrm{lo}'<m<\mathrm{hi}' (type complex): children with gaps [lo′,m][\mathrm{lo}',m] and [m,hi′][m,\mathrm{hi}']. Exhaustive.

  • Location update: a lower bound λ2>h\lambda_{2}>h valid on the whole cell (Propositions 7.10 and 7.12; Lemma 7.15) replaces l2l_{2} and rr by max⁡(l2,h)\max(l_{2},h) and max⁡(r,h)\max(r,h), and excludes the node if it is reserved with hi2≤h\mathrm{hi}_{2} \le h. Sound, since every family F≠F1\mathcal{F} \ne\mathcal{F}_{1} has ν(F)≥λ2\nu(\mathcal{F}) \ge\lambda_{2}.

  • Exclusion of a reserved interval or a gap by a lower bound for λ2\lambda_{2} or λ′\lambda' (Propositions 7.11 and 7.13), or of a whole node by the positivity argument of Proposition 7.14.

Before a leaf model is built, all applicable location updates are applied once more to its specification; this only adds valid implications. None of these nodes depends on LL. The terminal nodes are the leaves. The inside trees have 3146 leaves. Of these, 233 (in 199 roots) are stored in a separate file and are called supplementary leaves. Every leaf, supplementary or not, receives the same model. The outside trees have 1642 leaves. In all there are 4788 leaves; 343 inside roots and 70 outside roots have none, because all their configurations are excluded.

The leaf model

This subsection gives the real coefficients before integer scaling. An interval [lo,hi)[\mathrm{lo},\mathrm{hi}) is charged at its left endpoint in the objective, where the cost is largest, and at its right endpoint in density constraints, where the available weight is smallest. These choices enlarge the feasible set and give an upper bound for every actual distribution of characters in the interval.

Let s\mathfrak{s} be the specification of a leaf, after the location updates. Let ss denote the shift min⁡(1.9,p∗,l2)\min(1.9,p^{*},l_{2}), where p∗=lo′p^{*}=\mathrm{lo}' if s\mathfrak{s} has a gap and p∗=pp^{*}=p otherwise, and let λ3lo\lambda_{3}^{\mathrm{lo}} be the lower bound for λ3\lambda_{3} of Proposition 7.16, computed from [a,b][a,b], the type and, for a reserved s\mathfrak{s}, the cap hi2\mathrm{hi}_{2}. The leaf program (Definition 10.1) has the columns of Table 11. Every number is rounded in the safe direction at scale SS: far weights, features and the hidden-tail count coefficient 1/w(end)1/w(\mathrm{end}) down; objective coefficients, diagonals, correlation terms, the far budget and the constant first up. Let OO be the set of ordinary characters with a zero in RPR_{P}, and H\mathcal{H} the set of first-family characters with a zero in RPR_{P}.

ColumnValue xcx_{c}GcG_{c}WcW_{c}CcC_{c}NcN_{c}EcE_{c}Near rows
ordinary bin [ξi,ξi+1)[\xi_{i},\xi_{i+1})number of χ∈O\chi\in O with νT∗(χ)\nu_{T^{*}}(\chi) in the binG(ξi)\mathcal{G}(\xi_{i})w(ξi+1)w(\xi_{i+1})0 or 1a1^{\mathrm{a}}00entry (a)
tail∑χ∈O withνT∗(χ)≥Rw(νT∗(χ))\sum_{\substack{\chi\in O\ \mathrm{with}\\ \nu_{T^{*}}(\chi)\geq R}}w(\nu_{T^{*}}(\chi))G(R)/w(R)\mathcal{G}(R)/w(R)1000omitted
reserved columnb^{\mathrm{b}} [lo2,j,hi2,j][\mathrm{lo}_{2,j},\mathrm{hi}_{2,j}]zj=n2z_{j}=n_{2} if ν∗\nu_{*} is in the column, else 0Gϕ2(lo2,j)\mathcal{G}_{\phi_{2}}(\mathrm{lo}_{2,j})w(hi2,j)w(\mathrm{hi}_{2,j})001entry (f)
hidden binc^{\mathrm{c}} [loh,hih)[\mathrm{lo}_{h},\mathrm{hi}_{h})number of χ∈H\chi\in\mathcal{H} with tχt_{\chi} in the bine−AlohBϕ1(p∘)e^{-A_{\mathrm{lo}_{h}}B_{\phi_{1}}(p^{\circ})}w(hih)w(\mathrm{hi}_{h})010entry (g)
hidden tailc^{\mathrm{c}}∑w(tχ)\sum w(t_{\chi}) over χ∈H\chi\in\mathcal{H} with tχ≥endt_{\chi}\geq\mathrm{end}e−AendBϕ1(p∘)w(end)\dfrac{e^{-A_{\mathrm{end}}B_{\phi_{1}}(p^{\circ})}}{w(\mathrm{end})}101w(end)\dfrac{1}{w(\mathrm{end})}0omitted

Table 11. The columns of a leaf program and their real coefficients before scaling by SS: objective GcG_{c}, far weight WcW_{c}, count CcC_{c}, hidden count NcN_{c} and second-family indicator EcE_{c} (Definition 10.1). The last column gives the entry of Table 7 that supplies the feature and diagonal of the column in each near row. (a) 1 if s\mathfrak{s} is unreserved and ξi+1≤λ3lo\xi_{i+1}\leq\lambda_{3}^{\mathrm{lo}}, and 0 otherwise. (b) Only if the reserved family is charged through columns. (c) Outside leaves only.

(a) Ordinary bins. Let R=max⁡(3,r)R=\max(3,r) and let r=ξ0<ξ1<⋯<ξN=Rr=\xi_{0}<\xi_{1}<\cdots<\xi_{N}=R consist of rr, RR and the multiples of 1/den1/\mathrm{den} between them, where den∈{200,400,500,800,2000}\mathrm{den}\in\{200,400,500,800,2000\} is fixed per leaf. The bins are the intervals [ξi,ξi+1)[\xi_{i},\xi_{i+1}), except that in a reserved leaf the bins with ξi+1≤λ3lo\xi_{i+1}\leq\lambda_{3}^{\mathrm{lo}} are omitted. The ordinary characters beyond RR form the tail.

(b) Reserved family. Either (fixed terms) the family is charged n2Gϕ2(lo2)n_{2}\mathcal{G}_{\phi_{2}}(\mathrm{lo}_{2}) in first and n2w(hi2)n_{2}w(\mathrm{hi}_{2}) in the far budget, and it appears in the near rows as a family term with its feature at hi2\mathrm{hi}_{2} (there are then no second-family columns, and the count nEn_{E} of the second-family equality of Definition 10.1 is 0); or (columns) the interval [lo2,hi2][\mathrm{lo}_{2},\mathrm{hi}_{2}] is cut at the multiples of 1/4001/400 into reserved columns [lo2,j,hi2,j][\mathrm{lo}_{2,j},\mathrm{hi}_{2,j}], with values zjz_{j} and ∑jzj=n2=nE\sum_{j}z_{j}=n_{2}=n_{E}.

(c) Hidden columns (outside leaves only). Put p∘=max⁡(a,p,lo′)p^{\circ}=\max(a,p,\mathrm{lo}') and end=max⁡(3,p∘)\mathrm{end}=\max(3,p^{\circ}). The hidden bins [loh,hih)[\mathrm{lo}_{h},\mathrm{hi}_{h}) divide [p∘,end][p^{\circ},\mathrm{end}] on the grid 1/den1/\mathrm{den}, and the hidden tail collects the values tχ≥endt_{\chi}\geq\mathrm{end}. The hidden count is at most nn.

Before scaling and rounding, the far budget is (1+η)V−1insiden w(b)(1+\eta)V-\mathbf{1}_{\mathrm{inside}}n\,w(b), with a further subtraction of n2w(hi2)n_{2}w(\mathrm{hi}_{2}) for a fixed reserved-family term. The stored integer budget FF is formed by rounding the initial budget up and each subtraction down, as in Section 6. The count budget is 2, and the stored allowance satisfies final/S=5η\mathrm{final}/S=5\eta. For an inside leaf, first/S\mathrm{first}/S is rounded up from J1=min⁡(Jold,Jnew)J_{1}=\min(J_{\mathrm{old}},J_{\mathrm{new}}) with anchor p∗p^{*}, omitting JnewJ_{\mathrm{new}} when p∗<bp^{*}<b. For an outside leaf it starts at 0. The fixed reserved-family charge, if used, is added with upward rounding. The near rows are those of Section 9.

Subdivision of a leaf

A leaf may be certified through finitely many subcases, each of which is certified with its own leaf model. Besides the second-family split above, three kinds are used. The first replaces the specification by three specifications. The other two restrict one further quantity QQ of the configuration to one of finitely many intervals II whose union is the range of QQ; the configuration set of the subcase (s,I)(\mathfrak{s},I) is the set C(s,I)C(\mathfrak{s},I) of configurations for which, for some admissible choice of (χ1,ρ1)(\chi_{1},\rho_{1}), of the minimizers and of ρ′\rho', (C1)–(C5) hold and Q∈IQ\in I. Then C(s)C(\mathfrak{s}) is the union of the sets C(s,I)C(\mathfrak{s},I), and all the results of Sections 9 and 11 that are stated for C(s)C(\mathfrak{s}) hold for C(s,I)C(\mathfrak{s},I) with the same proofs.

Lemma 11.5 (Reservation split). Let s\mathfrak{s} be unreserved with ordinary lower bound rr, and let r<h≤2r<h\le2. Let s1\mathfrak{s}_{1} and s2\mathfrak{s}_{2} be s\mathfrak{s} with the reservation (h,1)(h,1) and (h,2)(h,2) (so lo2=r\mathrm{lo}_{2}=r and hi2=h\mathrm{hi}_{2}=h), and let s0\mathfrak{s}_{0} be s\mathfrak{s} with rr replaced by hh. Then C(s)⊆C(s0)∪C(s1)∪C(s2)C(\mathfrak{s})\subseteq C(\mathfrak{s}_{0})\cup C(\mathfrak{s}_{1})\cup C(\mathfrak{s}_{2}).

Proof. If ν∗<h\nu_{*}<h, a family F≠F1\mathcal{F}\ne\mathcal{F}_{1} attaining ν∗\nu_{*} exists, and it is a real character or a pair of conjugate nonreal characters; the configuration lies in C(s1)C(\mathfrak{s}_{1}) or C(s2)C(\mathfrak{s}_{2}), since ν∗≥r\nu_{*}\ge r. Otherwise every family F≠F1\mathcal{F}\ne\mathcal{F}_{1} has ν(F)≥h\nu(\mathcal{F})\ge h, and the configuration lies in C(s0)C(\mathfrak{s}_{0}). □\square

In a reserved child the model uses only the minimality of the reserved family (Lemma 11.3(ii)) and the cap λ2≤h\lambda_{2}\le h (Lemma 11.3(iii)); it never uses that the reserved family is F2\mathcal{F}_{2}. The split of Lemma 11.5 is used with h=min⁡(λ3lo,2)h=\min(\lambda_{3}^{\mathrm{lo}},2) on four inside roots and on five roots that contain supplementary leaves.

Height split for type rc\mathrm{rc}. An inside leaf of type rc\mathrm{rc} is split by Q=μ1Q=\mu_{1} (recall that μ1>0\mu_{1}>0 for this type) into μ1∈[0,1]\mu_{1}\in[0,1] and μ1∈[1,∞)\mu_{1}\in[1,\infty), which changes only the family row (Section 9).

Second-zero height split. An inside leaf of type complex with a finite gap is split by Q=∣y∣Q=|y|, where y=μ1−μρ′=(γ1−γ′)Ly=\mu_{1}-\mu_{\rho'}=(\gamma_{1}-\gamma')L is the normalized height difference between ρ1\rho_{1} and the nondistinguished zero ρ′\rho' of χ1\chi_{1} realizing λ′\lambda' (a zero of χ1\chi_{1}, not of χ‾1\overline{\chi}_{1}; Section 2), into six pieces, according as ∣y∣|y| lies in [0,1][0,1], [1,32][1,\frac{3}{2}], [32,2][\frac{3}{2},2], [2,3][2,3], [3,5][3,5] or [5,∞)[5,\infty). This changes only the family row. A piece may be certified by an exclusion (Definition 10.6(i)) rather than by duals: the near row alone is then infeasible.

Realization and the cover theorem

Proposition 11.6 (Realization). Let s\mathfrak{s} specify a leaf or one of its subcases, with the program constructed in Section 11.5. For every sufficiently large qq whose configuration lies in C(s)C(\mathfrak{s}), there is a feasible pair (x,τ)(x,\tau) such that

W(q)+2η≤V(x).W(q)+2\eta\le V(x).

The column values xx are obtained from the actual character counts and weighted tail masses. The threshold q0q_{0} may be chosen uniformly over the finite set of leaves and subcases.

Proof. Let OO and H\mathcal{H} be as in Section 11.5 (H\mathcal{H} is used only if the configuration is outside). For χ∈O\chi\in O we have νT∗(χ)≤CP\nu_{T^{*}}(\chi)\le C_{P} and νT∗(χ)≥ν(χ)≥r\nu_{T^{*}}(\chi)\ge\nu(\chi)\ge r by (C4). Put xix_{i} equal to the number of χ∈O\chi\in O with νT∗(χ)∈[ξi,ξi+1)\nu_{T^{*}}(\chi)\in[\xi_{i},\xi_{i+1}), xtail=∑w(νT∗(χ))x_{\mathrm{tail}}=\sum w(\nu_{T^{*}}(\chi)) over χ∈O\chi\in O with νT∗(χ)≥R\nu_{T^{*}}(\chi)\ge R, zjz_{j} equal to n2n_{2} for the reserved column containing ν∗\nu_{*} and 0 otherwise, xihx_{i}^{h} equal to the number of χ∈H\chi\in\mathcal{H} whose value tχt_{\chi} (Section 5.7) lies in the hidden bin ii, and xtailh=∑w(tχ)x_{\mathrm{tail}}^{h}=\sum w(t_{\chi}) over the χ∈H\chi\in\mathcal{H} in the hidden tail. In a reserved leaf no χ∈O\chi\in\mathcal{O} lies in an omitted bin, by Lemma 11.3(ii) and λ3≥λ3o\lambda_{3}\geq\lambda_{3}^{\mathrm{o}}.

Far. The characters of O\mathcal{O}, the reserved characters, the first-family characters of H\mathcal{H} and, for an inside configuration, χ1\chi_{1} and χ‾1\overline{\chi}_{1} are distinct, and to each we assign one zero with ∣γ∣≤1|\gamma|\leq1 and parameter at most CPCP: the T∗T^{*}-representative, the height-one representative, the zero realizing tχt_{\chi}, and ρ1\rho_{1} or ρ‾1\overline{\rho}_{1}, which has ∣γ1∣≤CB/L≤1|\gamma_{1}|\leq C_{B}/\mathcal{L}\leq1. By (6.1), and since ww is decreasing, the far constraint holds.

Count. By Lemma 11.3(i), since a character in a counted bin has νT∗(χ)<ξi+1≤λ3o≤λ3\nu_{T^{*}}(\chi)<\xi_{i+1}\leq\lambda_{3}^{\mathrm{o}}\leq\lambda_{3}.

Hidden count and second family. We have ∣H∣≤n|\mathcal{H}|\leq n, a character in the hidden tail contributes w(tχ)/w(end)≤1w(t_{\chi})/w(\mathrm{end})\leq1 to the hidden count, and ∑jzj=n2=nE\sum_{j}z_{j}=n_{2}=n_{E} when the reserved family is charged through columns (otherwise there are no second-family columns and nE=0n_{E}=0).

Near rows. By Proposition 9.7, each near row holds for some threshold, with features at least and diagonals at most those of the columns.

Value. We use (5.2). An ordinary character in a bin costs at most the objective G(ξi)\mathcal{G}(\xi_{i}) at the left end of the bin, a reserved character in a column at most Gϕ2\mathcal{G}_{\phi_{2}} at the left end (with fixed terms, at most Gϕ2(lo2)\mathcal{G}_{\phi_{2}}(\mathrm{lo}_{2})), and the tail characters at most G(R)/w(R)\mathcal{G}(R)/w(R) per unit of far mass, by the monotonicity of G\mathcal{G}, Gϕ2\mathcal{G}_{\phi_{2}} and G/w\mathcal{G}/w (Lemma 5.1(v)). For an outside configuration we apply Proposition 5.9 with the anchor p∘=max⁡(a,p,lo′)p^{\circ}=\max(a,p,\mathrm{lo}^{\prime}), which is admissible because every non-distinguished occurrence of the first family has parameter at least λ′≥p∘\lambda^{\prime}\geq p^{\circ}; then tχ≥p∘t_{\chi}\geq p^{\circ}, a character in a hidden bin [loh,hih)[\mathrm{lo}_{h},\mathrm{hi}_{h}) costs at most the objective e−AlohBϕ1(p∘)e^{-A\mathrm{lo}_{h}}B_{\phi_{1}}(p^{\circ}) since e−Aλe^{-A\lambda} is decreasing, and a character in the hidden tail costs at most e−A⋅endBϕ1(p∘)/w(end)e^{-A\cdot\mathrm{end}}B_{\phi_{1}}(p^{\circ})/w(\mathrm{end}) per unit of far mass, since e−Aλ/w(λ)e^{-A\lambda}/w(\lambda) is decreasing (Lemma 6.2). Including the directions of integer rounding, we obtain

W≤V(x)−finalS+E≤V(x)−5η+2η≤V(x)−2η.W\leq\mathcal{V}(x)-\frac{\mathrm{final}}{S}+E\leq\mathcal{V}(x)-5\eta+2\eta\leq\mathcal{V}(x)-2\eta.

Theorem 11.7 (The zero-sum bound in the middle range). For the case tree and certificates described in this paper, there is a threshold q0q_{0} such that every q≥q0q\geq q_{0} with 0.1≤λ1<1.50.1\leq\lambda_{1}<1.5 satisfies

W(q)<1−2η.W(q)<1-2\eta.

More precisely, its configuration belongs to a leaf or subcase whose program has a feasible pair (x,τ)(x,\tau) with W(q)+2η≤V(x)<1W(q)+2\eta\leq\mathcal{V}(x)<1.

Proof. By Theorem 11.4 the configuration lies in the configuration set of a root. Following the refinement tree and the subdivisions of Section 11.6, each of which is exhaustive, we reach a leaf or subcase whose configuration set contains the configuration; it is not an excluded node, because excluded nodes have no configurations. Proposition 11.6 gives the feasible pair, and Theorem 10.7 and the certificates of Section 14 give V(x)<1\mathcal{V}(x)<1. □\square

The exterior regimes

The case tree covers 0.1≤λ1<1.50.1\leq\lambda_{1}<1.5. We now treat the remaining possibilities: an empty zero rectangle, a first zero with λ1<0.1\lambda_{1}<0.1, and a first zero with λ1≥1.5\lambda_{1}\geq1.5. The small-parameter case is different from the others. Its first zero is real: it is an exceptional zero in the sense of Landau and Page [74, 97]. By the theorems of Siegel and Tatuzawa [118, 123] (see also [102, 103]) it cannot be extremely close to 1, but these bounds do not keep λ1\lambda_{1} away from 0. Although such a zero has a large cost, the Deuring–Heilbronn phenomenon [30, 53, 78] repels the remaining zeros. We use Heath-Brown’s quantitative estimates for this repulsion (see also [8] for explicit versions) and, for a sufficiently close zero, his separate theorem on Siegel zeros [50], by which a very strong exceptional zero is even helpful.

No zero in R(l)R(l)

If no nonprincipal L(s,χ)L(s,\chi) vanishes in R(l)R(l), then for large qq none vanishes in RPR_{P}, because CP≤13log⁡log⁡LC_{P}\le\frac{1}{3}\log\log\mathcal{L} and CP/L≤lC_{P}/\mathcal{L}\le l. So W=0W=0 and Proposition 4.4 gives a prime p≡a(modq)p\equiv a\pmod q with qA<p<qLq^{A}<p<q^{L}. Xylouris notes this case himself [145].

A tiny exceptional zero

Let ηHB(δ)\eta_{\mathrm{HB}}(\delta) be the effectively computable constant of Input 3.10 and put

uexc=min⁡(1ηHB(1/2),0.08).u_{\mathrm{exc}}=\min\left(\frac{1}{\eta_{\mathrm{HB}}(1/2)},0.08\right).

Proposition 12.1 (Primes in the presence of a very close real zero). Let uexcu_{\mathrm{exc}} be the positive cutoff defined above. For all sufficiently large qq, if λ1<uexc\lambda_{1}<u_{\mathrm{exc}}, then P(a,q)≤q3.5P(a,q)\le q^{3.5} for every (a,q)=1(a,q)=1.

Proof. By Proposition 7.1, χ1\chi_{1} and ρ1\rho_{1} are real, so β0=ρ1=1−λ1/L\beta_{0}=\rho_{1}=1-\lambda_{1}/\mathcal{L} is a real zero of the real nonprincipal L(s,χ1)L(s,\chi_{1}). Since λ1<uexc≤13\lambda_{1}<u_{\mathrm{exc}}\le\frac{1}{3} we have β0≥1−1/(3log⁡q)\beta_{0}\ge1-1/(3\log q), and ((1−β0)log⁡q)−1=1/λ1>1/uexc≥ηHB(12)((1-\beta_{0})\log q)^{-1}=1/\lambda_{1}>1/u_{\mathrm{exc}}\ge\eta_{\mathrm{HB}}(\frac{1}{2}). Input 3.10 with δ=12\delta=\frac{1}{2} gives the claim. □\square

A small exceptional zero

For uexc≤λ1≤0.1u_{\mathrm{exc}}\le\lambda_{1}\le0.1 we use a different prime weight, a single triangle, and Heath-Brown’s own estimates. The inputs are the following.

Input 12.2 (Prime detection with a single triangle [51]). Let L′>2K+3L'>2K+3, let f=hL′,Kf=h_{L',K} and F(z)=e−(L′−2K)zF2(z)F(z)=e^{-(L'-2K)z}F_{2}(z) with F2(z)=(1−e−Kzz)2F_{2}(z)=\left(\frac{1-e^{-Kz}}{z}\right)^{2}. For every ε>0\varepsilon>0 there is R=R(ε)R=R(\varepsilon) such that for large qq

∑p≡a (q)log⁡ppf(log⁡pL)≥Lφ(q)[K2−ε−∑χ≠χ0∑ρ′∣F((1−ρ)L)∣],\sum_{p\equiv a\ (q)}\frac{\log p}{p}f\left(\frac{\log p}{\mathcal{L}}\right) \ge \frac{\mathcal{L}}{\varphi(q)} \left[ K^{2}-\varepsilon-\sum_{\chi\ne\chi_{0}}\sum_{\rho}'\left|F\left((1-\rho)\mathcal{L}\right)\right| \right],

where ∑′\sum' runs over the zeros with 1−R/L≤β≤11-R/\mathcal{L}\le\beta\le1, ∣γ∣≤R/L|\gamma|\le R/\mathcal{L}.

Input 12.3 (Per-character bound for the triangular kernel [51]). Let χ≠χ0\chi\ne\chi_{0}, 0<λ≤R0<\lambda\le R, and suppose that L(s,χ)L(s,\chi) has no zero with 1−λ/L<β≤11-\lambda/\mathcal{L}<\beta\le1 and ∣γ∣≤1|\gamma|\le1. Then for every η1>0\eta_{1}>0 and q≥q(η1,λ)q\ge q(\eta_{1},\lambda)

∑ρ′∣F2((1−ρ)L)∣≤Bφ(χ)(λ)+η1,Bϕ(λ)=ϕ21−e−2Kλλ+2Kλ−1+e−2Kλ2λ2.\sum_{\rho}'\left|F_{2}\left((1-\rho)\mathcal{L}\right)\right| \le B_{\varphi(\chi)}(\lambda)+\eta_{1}, \qquad B_{\phi}(\lambda) = \frac{\phi}{2}\frac{1-e^{-2K\lambda}}{\lambda} + \frac{2K\lambda-1+e^{-2K\lambda}}{2\lambda^{2}}.

Input 12.4 (Heath-Brown’s weighted density estimate). In the notation of [51] and [145], the following holds. Let ε,c1,c2>0\varepsilon,c_{1},c_{2}>0 and ϕ=max⁡χφ(χ)\phi=\max_{\chi}\varphi(\chi). For q≥q(ε,c1,c2)q\ge q(\varepsilon,c_{1},c_{2}) and any choice, for each nonprincipal χ\chi with a zero in β≥1−λ0/L\beta\ge1-\lambda_{0}/\mathcal{L}, ∣γ∣≤1|\gamma|\le1 (λ0=13log⁡log⁡L)(\lambda_{0}=\frac{1}{3}\log\log\mathcal{L}), of one such zero with parameter λ(χ)\lambda(\chi),

∑χλ(χ)e(4ϕ+6c1+2c2)λ(χ)−e(2ϕ+4c1)λ(χ)≤2ϕ+2c1+c24c1c2+ε.\sum_{\chi} \frac{\lambda(\chi)} {e^{(4\phi+6c_{1}+2c_{2})\lambda(\chi)}-e^{(2\phi+4c_{1})\lambda(\chi)}} \le \frac{2\phi+2c_{1}+c_{2}}{4c_{1}c_{2}}+\varepsilon.

The left side decreases and the right side increases with ϕ\phi, so the inequality holds with ϕ=13\phi=\frac{1}{3}.

We also use Input 7.2(b),(d) and the following two rows of Heath-Brown’s tables, which hold for all λ1≤0.10\lambda_{1}\le0.10 bounded below by a fixed positive constant [51] (see the remark after Input 7.2; we use them for λ1≥0.08\lambda_{1}\ge0.08): if χ1\chi_{1} and ρ1\rho_{1} are real and λ1≤0.10\lambda_{1}\le0.10, then λ′≥4.96−ε\lambda'\ge4.96-\varepsilon and λ2≥2.83−ε\lambda_{2}\ge2.83-\varepsilon for q≥q(ε)q\ge q(\varepsilon). We have recomputed these rows from Heath-Brown’s printed parameters, obtaining 4.963554.96355 and 2.845772.84577; we use them with ε=0.001\varepsilon=0.001. (Table 5 is stated under λ2≤λ′\lambda_{2}\le\lambda', which is redundant here since λ′≥4.959>2.83\lambda'\ge4.959>2.83.)

Fixed data. L=3.99L=3.99, K=0.1821K=0.1821, As=L−2K=3.6258A_{s}=L-2K=3.6258; c1=0.057c_{1}=0.057, c2=0.1554c_{2}=0.1554, ϕ=13\phi=\frac{1}{3};

ad=43+6c1+2c2=37241875,bd=23+4c1=671750,Vd=2/3+2c1+c24c1c2=1847506993;a_{d}=\frac{4}{3}+6c_{1}+2c_{2}=\frac{3724}{1875}, \qquad b_{d}=\frac{2}{3}+4c_{1}=\frac{671}{750}, \qquad V_{d}=\frac{2/3+2c_{1}+c_{2}}{4c_{1}c_{2}}=\frac{184750}{6993};

α=1.09<1211\alpha=1.09<\frac{12}{11}, B1=K2+K4B_1=K^2+\frac{K}{4}, m∗=αlog⁡12.5=2.7530442…m^*=\alpha\log12.5=2.7530442\ldots, m2=2.829m_2=2.829 and m2′=4.959m'_2=4.959. Let Hs(u)=e−AsuF2(u)H_s(u)=e^{-A_su}F_2(u) and M(u)=(K2−Hs(u))/u\mathcal{M}(u)=(K^2-H_s(u))/u, and

Q(m)=e−(As−ad)m−e−(As−bd)mm=∫As−adAs−bde−tm dt.Q(m)=\frac{e^{-(A_s-a_d)m}-e^{-(A_s-b_d)m}}{m} =\int_{A_s-a_d}^{A_s-b_d}e^{-tm}\,\mathrm{d}t.

Theorem 12.5 (The small exceptional-zero range). For all sufficiently large qq with uexc≤λ1≤0.1u_{\mathrm{exc}}\le\lambda_1\le0.1, and every integer aa with (a,q)=1(a,q)=1, there is a prime pp such that

p≡a(modq),q3.6258<p<q3.99.p\equiv a\pmod q,\qquad q^{3.6258}<p<q^{3.99}.

The threshold is uniform over the fixed interval [uexc,0.1][u_{\mathrm{exc}},0.1].

Proof. Write u=λ1u=\lambda_1. By Proposition 7.1, χ1\chi_1 and ρ1\rho_1 are real. The zero ρ1\rho_1 is simple, since a double zero would give λ′=u\lambda'=u, whereas λ′≥(2−ε)log⁡(1/u)>u\lambda'\ge(2-\varepsilon)\log(1/u)>u (Input 7.2(b)) on [uexc,0.08][u_{\mathrm{exc}},0.08] and λ′≥4.959>u\lambda'\ge4.959>u on [0.08,0.1][0.08,0.1].

The criterion. Apply Input 12.2 with L′=LL'=L, since 3.99>2K+33.99>2K+3; the zero ρ1\rho_1 contributes Hs(u)H_s(u). Enlarging RR only adds nonnegative terms to ∑′\sum', so the inequality of Input 12.2 remains true with RR replaced by max⁡(R,3)\max(R,3); we use it with R≥3R\ge3, so that Input 12.3 applies with the values λ≤m2<3\lambda\le m_2<3 below. It remains to bound the other terms. Put m(u)=αlog⁡(1/u)m(u)=\alpha\log(1/u) on P1=[uexc,0.08]P_1=[u_{\mathrm{exc}},0.08] and m(u)=m2m(u)=m_2 on P2=[0.08,0.1]P_2=[0.08,0.1], and m1(u)=m(u)m_1(u)=m(u) on P1P_1, m1(u)=m2′m_1(u)=m'_2 on P2P_2. By Input 7.2(b),(d) (with ε=1211−α\varepsilon=\frac{12}{11}-\alpha) and the two table rows, every zero ρ≠ρ1\rho\ne\rho_1 of L(s,χ1)L(s,\chi_1) in R(l)R(l) has λρ≥m1(u)\lambda_\rho\ge m_1(u), and every zero of every other nonprincipal L(s,χ)L(s,\chi) in R(l)R(l) has λρ≥m(u)\lambda_\rho\ge m(u).

The first character. L(s,χ1)L(s,\chi_1) has no zero with λ<uexc/2\lambda<u_{\mathrm{exc}}/2 and ∣γ∣≤1|\gamma|\le1, so Input 12.3 with λ=uexc/2\lambda=u_{\mathrm{exc}}/2 and φ(χ1)=14\varphi(\chi_1)=\frac14 gives ∑ρ′∣F2((1−ρ)L)∣≤B1/4(uexc/2)+η1≤B1+η1\sum'_\rho|F_2((1-\rho)\mathcal{L})|\le B_{1/4}(u_{\mathrm{exc}}/2)+\eta_1\le B_1+\eta_1, as BϕB_\phi is decreasing in λ\lambda, with limit ϕK+K2\phi K+K^2 as λ→0\lambda\to0. Hence the zeros of χ1\chi_1 other than ρ1\rho_1 contribute

T1≤e−Asm1(u)(B1+η1).T_1\le e^{-A_sm_1(u)}(B_1+\eta_1).

The other characters. For χ≠χ0,χ1\chi\ne\chi_0,\chi_1 with a zero in the rectangle of Input 12.2, let λ(χ)\lambda(\chi) be the parameter of its zero of least parameter with ∣γ∣≤1|\gamma|\le1; then λ(χ)≥m(u)\lambda(\chi)\ge m(u). Apply Input 12.3 with the fixed value λ=m∗\lambda=m^* on P1P_1 (note m∗≤m(u)m^*\le m(u) there) and λ=m2\lambda=m_2 on P2P_2, and with φ(χ)≤13\varphi(\chi)\le\frac13. Since every zero in the rectangle has parameter at least λ(χ)\lambda(\chi), the zeros of χ\chi contribute at most e−Asλ(χ)(B1/3(mˉ)+η1)e^{-A_s\lambda(\chi)}(B_{1/3}(\bar m)+\eta_1), with mˉ=m∗\bar m=m^* or m2m_2. Since As>adA_s>a_d, the function e−Ast(eadt−ebdt)/t=Q(t)e^{-A_st}(e^{a_dt}-e^{b_dt})/t=Q(t) is decreasing, so e−Asλ(χ)≤Q(m(u))λ(χ)/(eadλ(χ)−ebdλ(χ))e^{-A_s\lambda(\chi)}\le Q(m(u))\lambda(\chi)/(e^{a_d\lambda(\chi)}-e^{b_d\lambda(\chi)}). By Input 12.4, the characters other than χ1\chi_1 contribute

T2≤(B1/3(mˉ)+η1)(Vd+ε2)Q(m(u)).T_2\le(B_{1/3}(\bar m)+\eta_1)(V_d+\varepsilon_2)Q(m(u)).

Monotonicity. We have K2−Hs(u)=∫hL,K(t)(1−e−ut) dtK^2-H_s(u)=\int h_{L,K}(t)(1-e^{-ut})\,\mathrm{d}t, so M(u)=∫hL,K(t)1−e−utu dt\mathcal{M}(u)=\int h_{L,K}(t)\frac{1-e^{-ut}}{u}\,\mathrm{d}t, which decreases in uu because hL,K≥0h_{L,K}\ge0 and (1−e−ut)/u(1-e^{-ut})/u decreases for each t>0t>0. On P1P_1 we have u=e−m/αu=e^{-m/\alpha}, and

Q(m(u))u=∫As−ad−1/αAs−bd−1/αe−tm(u) dt,e−Asm1(u)u=e−(As−1/α)m(u),\frac{Q(m(u))}{u} =\int_{A_s-a_d-1/\alpha}^{A_s-b_d-1/\alpha}e^{-tm(u)}\,\mathrm{d}t, \qquad \frac{e^{-A_sm_1(u)}}{u}=e^{-(A_s-1/\alpha)m(u)},

with As−ad−1/α=236171327000>0A_s-a_d-1/\alpha=\frac{236171}{327000}>0; both increase with uu. On P2P_2 the bounds for T1T_1 and T2T_2 do not depend on uu, and u≥0.08u\ge0.08. Hence on PiP_i

K2−Hs(u)−T1−T2≥ϖiu−E,K^2-H_s(u)-T_1-T_2\ge\varpi_i u-\mathcal{E},

where

ϖ1=M(0.08)−VdB1/3(m∗)Q(m∗)0.08−B1e−Asm∗0.08,ϖ2=M(0.1)−VdB1/3(m2)Q(m2)0.08−B1e−Asm2′0.08,\begin{aligned} \varpi_{1} &= \mathcal{M}(0.08) - V_{d}B_{1/3}(m^{*})\frac{Q(m^{*})}{0.08} - B_{1}\frac{e^{-A_{s}m^{*}}}{0.08}, \\ \varpi_{2} &= \mathcal{M}(0.1) - V_{d}B_{1/3}(m_{2})\frac{Q(m_{2})}{0.08} - B_{1}\frac{e^{-A_{s}m'_{2}}}{0.08}, \end{aligned}

and E=η1(1+(Vd+ε2)(ad−bd))+ε2B1/3(0)(ad−bd)\mathcal{E}=\eta_{1}(1+(V_{d}+\varepsilon_{2})(a_{d}-b_{d}))+\varepsilon_{2}B_{1/3}(0)(a_{d}-b_{d}), using Q≤ad−bdQ\leq a_{d}-b_{d}.

Numerical margins. The interval computations give the following values, displayed here to ten decimal places:

M\mathcal{M}other charactersfirst characterϖi\varpi_{i}
P1\mathcal{P}_{1}0.10884584290.07831167460.00004546550.0304887029
P2\mathcal{P}_{2}0.10500566980.06688746920.00000001530.0381181853

Table 16.

The certified lower bounds give ϖi>0.0304\varpi_{i}>0.0304. Choose η1\eta_{1}, ε2\varepsilon_{2} and the ε\varepsilon of Input 12.2, after uexcu_{\mathrm{exc}}, with E+ε<0.03uexc\mathcal{E}+\varepsilon<0.03u_{\mathrm{exc}}. Then the bracket in Input 12.2 is at least (0.0304−0.03)uexc>0(0.0304-0.03)u_{\mathrm{exc}}>0, and there is a prime p≡a(modq)p\equiv a\pmod{q} with log⁡p/L∈(As,L)\log p/\mathcal{L}\in(A_{s},L). □

At L=3.99L=3.99, the margins are approximately 28% and 36% of the main term. Exploratory evaluations with the same KK, c1c_{1}, and c2c_{2} remain positive at L=3.90L=3.90; the certified exterior result used here is the one at 3.99.

A large first zero.

Theorem 12.6 (The large first-zero range). For all sufficiently large qq with a nonempty zero rectangle and λ1≥1.5\lambda_{1}\geq1.5,

W(q)≤0.9966814−2η.W(q)\leq0.9966814-2\eta.

Consequently, every reduced residue class contains a prime pp with q3.156342<p<q3.99q^{3.156342}<p<q^{3.99}.

Proof. Every nonprincipal character with a zero in RPR_{P} is treated as ordinary, with its height-one representative, of parameter λχ≥λ1≥1.5\lambda_{\chi}\geq\lambda_{1}\geq1.5; by Corollary 5.7 it contributes at most G(λχ)+ε1e−Aλχ\mathcal{G}(\lambda_{\chi})+\varepsilon_{1}e^{-A\lambda_{\chi}}. Consider the leaf program with the bins [1.5+j100,1.5+j+1100)[1.5+\frac{j}{100},1.5+\frac{j+1}{100}), 0≤j<1500\leq j<150, and the tail [3,∞)[3,\infty); the far budget F=(1+η)VF=(1+\eta)V with profile I (Section 6); first =0=0 and final =5η=5\eta; no count, hidden or second-family constraints; and one near row. The near row is Corollary 8.10 with Ξ≡0\Xi\equiv0, s=s1=32s=s_{1}=\frac{3}{2}, detector f^2\hat{f}_{2} (γf=1\gamma_{f}=1), Gram test f^3\hat{f}_{3} (γg=32\gamma_{g}=\frac{3}{2}) and no family entries, normalized as in Section 9.2 with t0=1t_{0}=1 (so IuI_{u} is the upper Riemann sum over 2000 cells of each of [0,1)[0,1) and [1,2)[1,2)): every zero of every nonprincipal LL-function in R(l)\mathbf{R}(l) has parameter at least 1.5=s11.5=s_{1}, so the safe-anchor hypothesis holds, and every entry has no zero to the right of its anchor min⁡(lo⁡,s)=s\min(\operatorname{lo},s)=s. The actual characters give a feasible point as in Proposition 11.6.

The certificate consists of 36 contiguous threshold intervals covering [0,e][0,e]; on each it uses the first-order relaxation and integer duals YF,Z≥0Y_{F},Z\geq0 satisfying (18); there are no exclusions and 5436 column inequalities. The largest value is 0.9966813270344512, on the threshold interval [0.1625,0.165][0.1625,0.165], where YF=5145926349Y_{F}=5145926349 and Z=1593393341176Z=1593393341176. By Theorem 10.7, V(x)<0.9966814\mathcal{V}(x)<0.9966814, and W≤V(x)−2ηW\leq\mathcal{V}(x)-2\eta as in Proposition 11.6. Proposition 4.4 gives the prime. □

In this regime one near row, together with the far-density budget, suffices. All characters are treated by the same per-character bound.

Assembly of the proof

Each preceding argument holds once qq exceeds a threshold depending on fixed choices. We first verify that these choices can be made in a noncircular order. A single threshold then works for every leaf and every exterior case.

The order of the choices

The constants of the argument are chosen in the following order. Each stage depends only on the earlier ones, and there are finitely many choices at each stage.

(S1) The fixed data: the exponent L=3.99L=3.99 and the kernel of Section 4; the two far profiles of Section 6; the test functions, anchors, sieve weights and grids of all near rows (Section 9); the case tree, its leaf programs and their certificates (Sections 10 and 11); the zero-location rows of Section 7; the data of the exterior regimes (Section 12); η=η′=10−6\eta=\eta'=10^{-6}, λ11=0.05\lambda_{11}=0.05 and smax⁡=3s_{\max}=3; and uexc=min⁡(1/ηHB(12),0.08)u_{\mathrm{exc}}=\min(1/\eta_{\mathrm{HB}}(\frac{1}{2}),0.08), with ηHB(δ)\eta_{\mathrm{HB}}(\delta) from Input 3.10; and the envelope tolerance ε1=η/23\varepsilon_{1}=\eta/23 of Section 5.8.

(S2) The prime window: CP≥max⁡(C0(ηH0),3)C_{P}\geq\max(C_{0}(\eta H_{0}),3), with C0C_{0} from Input 4.2.

(S3) The zero count N=N(CP)N=N(C_{P}) of Lemma 5.3; the mesh hAh_{\mathcal{A}} and the finite set of anchors A\mathcal{A} of Section 5.3; then the smoothing parameter θ\theta of Section 5.2, chosen so that (7) holds for ε1=η/23\varepsilon_{1}=\eta/23; then the constant coutc_{\mathrm{out}} of Proposition 5.9.

(S4) The count K0K_{0} of Section 6.3, a bound, uniform in qq, for the number of zeros of all nonprincipal LL-functions modulo qq with λ≤smax⁡\lambda\leq s_{\max} and ∣γ∣≤2|\gamma|\leq2 (Input 3.9). Then the separation MM, so large that for ∣Y∣≥M|Y|\geq M we have ∣Re⁡F(x+iY)∣≤ηmin⁡′′/(2K0+2)|\operatorname{Re}F(x+iY)|\leq\eta_{\min}''/(2K_{0}+2) for every detector and x∈[−smax⁡,0]x\in[-s_{\max},0], where ηmin⁡′′>0\eta_{\min}''>0 is the least of the numbers η′′=η′min⁡(1,IBDu,Du)\eta''=\eta'\min(1,\sqrt{I_{B}D_{u}},D_{u}) of Corollary 8.10 over the finitely many rows, and ∣Re⁡G(−σ+iY)∣≤d1|\operatorname{Re}G(-\sigma+iY)|\leq d_{1} for every Gram test and every σ=s1−δ−δ′\sigma=s_{1}-\delta-\delta' that occurs (Lemma 8.6). Then the buffer CB>max⁡(CP+M,1.5)C_{B}>\max(C_{P}+M,1.5) with 4cout/(H0CB2)≤η4c_{\mathrm{out}}/(H_{0}C_{B}^{2})\leq\eta.

(S5) The tolerances of the published results and of the lemmas proved here, each below the finitely many positive margins it must fit: the tolerances of Inputs 3.3 and 3.4 for the finitely many test functions (those of Section 5, the functions fσθf_{\sigma}^{\theta} with tolerance H0ε1/8H_{0}\varepsilon_{1}/8, and the detectors of the near rows with tolerance ηmin⁡′′/2\eta_{\min}''/2), the disc radius δ0≤13\delta_{0}\leq\frac{1}{3} of Section 5.3, which is at most the radius that Input 3.4 provides for each of these, and the tolerances of (12) and Lemma 7.8, of Input 6.1 (with ε=ηJ2\varepsilon=\eta J^{2}), of Inputs 3.7 and 3.8 and Mertens’ estimate, of Input 4.2 (with ε=ηH0\varepsilon=\eta H_{0}), of the zero tables, and of the small-branch estimates of Theorem 12.5 (with E+ε<0.03uexcE+\varepsilon<0.03u_{\mathrm{exc}}).

(S6) Finally q0q_{0}: the maximum of the finitely many thresholds of the previous stages, together with those of Input 3.5, of the inclusions RP⊆RB⊆the discsR_{P}\subseteq R_{B}\subseteq\text{the discs}, and L>4M(K0+1)L>4M(K_{0}+1) for Lemma 6.3.

The certificates of stage (S1) depend only on fixed data, not on CPC_{P}, MM, CBC_{B} or qq. The count K0K_{0} counts zeros up to height 2, which exceeds 1+δ01+\delta_{0} for every admissible radius; so MM, which depends on K0K_{0}, can be fixed before the radius δ0\delta_{0}, and there is no circular dependence. The strip height T∗=T∗(q)∈[12,1]T^{*}=T^{*}(q)\in[\frac{1}{2},1] of Lemma 6.3 is then determined by qq.

Envelope errors that are summed over unboundedly many characters are weighted by e−Aλχe^{-A\lambda_{\chi}}, and ∑χe−Aλχ<19\sum_{\chi}e^{-A\lambda_{\chi}}<19 by (10); errors in the near rows are uniform per entry or are multiples of (∑jaj)2(\sum_{j}a_{j})^{2} (Theorem 8.4). Thus the error allowances remain uniform as the number of characters grows with qq.

The regimes

Let q≥q0q\geq q_{0}. Exactly one of the following holds.

(R1) No nonprincipal L(s,χ)L(s,\chi) has a zero in R(l)R(l).

(R2) λ1<0.1\lambda_{1}<0.1.

(R3) 0.1≤λ1<1.50.1\leq\lambda_{1}<1.5.

(R4) λ1≥1.5\lambda_{1}\geq1.5.

(The first zero λ1\lambda_{1} is defined when (R1) fails; since R(l)R(l) is compact there are finitely many zeros in it.)

Proof of Theorem 1.1

Let q≥q0q \ge q_{0} and let aa be coprime to qq.

In case (R1), W=0W=0 and Proposition 4.4 gives a prime p≡a(modq)p \equiv a \pmod{q} with qA<p<qLq^{A}<p<q^{L}.

In case (R2), if λ1<uexc\lambda_{1}<u_{\mathrm{exc}} then P(a,q)≤q3.5P(a,q)\le q^{3.5} by Proposition 12.1; if uexc≤λ1<0.1u_{\mathrm{exc}}\le\lambda_{1}<0.1 then Theorem 12.5 gives a prime p≡a(modq)p\equiv a\pmod{q} with q3.6258<p<q3.99q^{3.6258}<p<q^{3.99}.

In case (R3), Theorem 11.7 gives a leaf of the case tree, or a subcase of a leaf, and a feasible point of its program with W≤V(x)−2η<1−2ηW\le\mathcal{V}(x)-2\eta<1-2\eta. In case (R4), Theorem 12.6 gives W<1−2ηW<1-2\eta. In both cases Proposition 4.4 gives a prime p≡a(modq)p\equiv a\pmod{q} with qA<p<qLq^{A}<p<q^{L}.

Thus P(a,q)<q3.99P(a,q)<q^{3.99} for every q≥q0q\ge q_{0} and every (a,q)=1(a,q)=1. There are only finitely many moduli 2≤q<q02\le q<q_{0}, and finitely many reduced residue classes for each such modulus. Dirichlet’s theorem supplies a least prime in each class. Taking the maximum of 1 and the finitely many ratios P(a,q)/q3.99P(a,q)/q^{3.99} gives an absolute CC valid for all q≥2q\ge2. □\square

Remark 13.1 (Effectivity). Heath-Brown’s theorem on Siegel zeros [50], used for the closest exceptional zeros, is effective. The other source estimates used here are also stated with effective constants. However, we have neither audited the effectivity of every step of the assembly nor computed a value of q0q_{0} or CC. Theorem 1.1 asserts their existence and does not supply an explicit numerical bound for either constant.

The computations and their verification

The finite computation has three logical components. The zero-location checks justify the case analysis; enclosures of analytic quantities supply the leaf coefficients; and integer certificates bound the resulting programs. The exterior ranges require separate computations. A check of one component does not by itself verify the others. This section records the evidence for each and the scope of the Lean formalization.

Computer assistance is well established in explicit analytic number theory, for instance in the ternary Goldbach problem [54] and in verifications of the Riemann hypothesis [106, 107], and linear programs with rigorously checked bounds are central to the proof of the Kepler conjecture [96, 119, 46]. In our computations, the real coefficients and margins are enclosed by interval arithmetic [88, 89, 113, 132]. It is implemented with the interval context of mpmath [126], at a working precision of at least 50 digits, and with a vectorized binary64 interval module that widens the result of every operation outward by one unit in the last place and evaluates exp, sin and cos by Taylor polynomials with explicit remainder bounds, not by the platform’s library; the Lean checkers of Section 14.3 use their own verified rational interval arithmetic. Floating-point linear programming (with the HiGHS solver [55] in SciPy [137]) is used only to propose dual multipliers, and every certificate is checked in exact integer arithmetic. Independent numerical audits use multiple-precision arithmetic [126]. The supplementary floating-point checks of certain published table entries are identified separately in Section 7.1. Table 12 summarizes, for each component of the proof, whether it is published or new, how it was obtained and how it was checked.

The data

The certificates are stored in three compressed JSON record files. The inside and outside files are indexed by root case; the supplementary file records the 233 supplementary leaves of 199 inside roots (Section 11.4). Their sizes and SHA-256 hashes are:

FileBytes
inside.jsonl.gz298,536,973
539362a03e7b330dad03b33d6b5037871cfcbbf7b0973e7976fccf12509fe63c
nodes.jsonl.gz159,624,716
b715ac735b6976da17a2671667ecfdf07c3f19b8ab74ae5f0533dbcf6b2a6cfc
outside.jsonl.gz8,626,435
a53d4d5793f71c92647a7f22022110bdc8135cbda1f2575d1dc0c1f487f5ac1d

Table 17.

The first holds the 2768 inside roots with their 2913 other leaves, the second the 233 supplementary leaves, and the third the 1685 outside roots with their 1642 leaves. A leaf record contains the parameters of its near rows and its certificate tree; it does not contain the leaf program itself, which the verifier regenerates from the case tree. Some leaves are split into subcases (Section 11.6), each with its own certificate tree, so the 4788 leaves carry 5591 certificate trees in all: 3949 inside (including the supplementary leaves) and 1642 outside.

Exact verification

All numbers entering a leaf program are enclosed with outward-rounded interval arithmetic and stored as integers scaled by S=1016S=10^{16}, rounded in the safe direction (Section 10). The verification of a certificate uses only integer arithmetic. The following checks have been run.

  1. The replay. A verifier traverses the case tree from its roots. It re-checks every split, location node and exclusion (regenerating the location rows, except that the table of complex polynomial rows and the side conditions [145], eq:4.29, eq:4.34 are checked by separate programs). At every leaf it regenerates the leaf program from the specification and the tree, independently of the stored certificate, and checks the certificate exactly. It passes on all 4453 roots: 4788 leaves, 4,196,879 boxes, and largest certified value 0.9999998223…<10.9999998223\ldots<1 (inside) and 0.9999942478…0.9999942478\ldots (outside).

  2. An independent checker. A second verifier was written from Section 10 alone, without reading or using the checking code of the first, and it uses only exact integer and rational arithmetic. It takes the leaf data from the same generator as the replay (the leaf data are checked separately, in (4) and in Section 14.3). It accepts every certificate of the corpus (29,397,336 relaxation cases in all) and that of Theorem 12.6, and for every certificate tree its value and number of boxes are exactly those of the first verifier. In a negative test on ten leaves, including the tightest, it rejected all 308 certificates made invalid by deliberate forgeries, among them duals lowered, and coefficients raised, by the smallest amount that breaks (10.1); it accepted all 82 controls perturbed by one unit less. It replaces an earlier independent checker, whose code was not preserved and which had also accepted every certificate of the corpus.

  3. A mutation test. On the 19 leaves of three of the tightest inside roots, 133 perturbations of the regenerated leaf data (far budget, objective, first-family charge, correlation term, features), of certificate multipliers and of the allowance 5η5\eta, each in the unfavourable direction, were all rejected by the replay.

  4. Audits of the constants. For 38 leaves of all kinds, the 19,616 objective and 19,616 far coefficients, the tail, hidden and second-family columns, FF, first and all 120 family, shifted and graded near rows were recomputed from their definitions in multiple-precision arithmetic by a program sharing no code with the builders; all lie on the safe side of their true values. (The program of this audit was not preserved; the numeric and row checkers of Section 14.3 now check the same quantities for every leaf, except the two-test rows.) Separately, for 117 roots (100 of them chosen at random), all 530,942 column and budget integers of their 154 leaves (objective, far, count, hidden-count and second-family

coefficients, FF, first and final) were recomputed from their defining formulas in 50-digit arithmetic and agree exactly with the stored ones.

(5) The two-test rows. A separate program replays the 37 roots whose leaves use the row of Section 9.6 (31 with a single test, 2 with a mixture, 4 paired). It compares the row as the certificates use it with a separate implementation of the same inequality, with its own builder and its own certificate checker. The comparison runs on every threshold interval that the 37 certificates use (1079 intervals, from 23,341 boxes) and on 11,211 seeded random intervals, in the first-order relaxation and in both tangent relaxations (for these, the tangent rule applied to the exact inequality of that implementation). The exact costs and budgets are those of that inequality divided by DD, and the integers used are rounded in the safe direction. The row’s integers are regenerated by the interval routines of that implementation and agree with the leaves; every one-unit perturbation of them is detected. These constants rest on those interval routines alone.

The exterior regimes. The margins of Theorem 12.5 were regenerated from their definitions and checked in interval arithmetic by a separate replay of the exterior components, and they have been recomputed independently in multiple-precision interval arithmetic. The program of Theorem 12.6 is a leaf program in the sense of Definition 10.1. Its objective and far integers are regenerated and compared with the stored ones, and its near row is built with the normalization of Section 9.2. The exact checking code of the replay (1) and the independent checker (2) accept its certificate, and so do the verified checkers of Section 14.3, which also accept its numeric data and its near row.

Formal verification

Several parts of the argument have been formalized in the Lean 4 proof assistant [29] with the Mathlib library [125]; see [2] for the role of such formalizations. Earlier formalizations in analytic number theory include the prime number theorem [3, 47, 72] and Dirichlet’s theorem [121]. The formalization uses Lean v4.35.0-rc3 with the corresponding Mathlib release, and contains about 24,000 lines and about 800 theorems and lemmas. The 18 statements of the near lemma and the certificate checker were checked against their proofs with the Comparator tool [75], which replays the proofs in the Lean kernel and in the independent kernel nanoda [5]; the leaf-level and near-row theorems were checked by the Lean kernel. All proofs are complete and use only the three standard axioms (propositional extensionality, the axiom of choice, and the soundness of quotients). The published inputs of Section 3 enter as explicit hypotheses, each stated no more strongly than its printed source. The following results are formally verified.

(a) The graded near-density lemma (Theorem 8.4), from the local explicit formulas, Burgess’s estimate and Graham’s estimate as hypotheses. This includes the response lemma, the threshold form, and the entry bounds of Section 9.

(b) The soundness of the certificate checker of Section 10: an accepted certificate proves that every feasible point of the leaf program has value below 1.

(c) The passage from a certified leaf with valid data to primes. From Xylouris’s criterion and his far-density lemma, stated as printed, and the localized envelope of Proposition 5.5, stated as a hypothesis, a leaf whose certificate and data pass the checks yields a prime qA<p<qLq^{A}<p<q^{L} in every reduced class, for every modulus whose zeros satisfy the leaf’s zero-level hypotheses. The count, hidden-count and second-family integers and final enter as they are (they were recomputed in the audit of (4) above).

(d) Interval arithmetic with proved inclusion, including the exponential function, and enclosures of the functions entering the objective and far data of the leaves. A numeric checker is proved to imply that each leaf’s objective and far coefficients, far budget and first-family charge are valid.

(e) Heath-Brown’s Conditions 1 and 2 for the parabolic test functions of Section 9.

(f) The near rows: each family, shifted and graded near row is a concrete instance of Theorem 8.4. A row checker is proved sound, including the three family terms that keep a second zero and the constant CZC_Z of Section 9. A checked row is a valid row of the leaf program under zero-level hypotheses only, and this is composed with (c). (The formal version of Corollary 8.10 assumes that the numerators D(δj)−d1+ejD(\delta_j)-d_1+e_j themselves are positive, which is stronger than what the proof in Section 8 needs; the row checker verifies this for every checked entry.)

The certificate checker has a Mathlib-free copy that is proved to compute the same value; the numeric checker and the row checker are Mathlib-free themselves and are proved sound directly. Run as compiled programs, the certificate and numeric checkers accept every certificate tree and the objective, far and first-family data of every leaf in the corpus: 5591 trees and 4,196,879 boxes. The compiled row checker accepts every near row of the corpus except the 37 two-test rows of Section 9.6, which it does not cover: 16,842 rows with 11,135,323 entries. The same three checkers accept the leaf program of Theorem 12.6: its certificate (36 boxes), its objective and far data, and its near row (150 entries). So (c) and (f) apply to it as to the leaves of the case tree. In addition, seven certificates are checked by evaluation in the Lean kernel itself: six of the case tree (122 boxes) and that of Theorem 12.6 (36 boxes).

The published inputs (Sections 3, 7 and 12) enter the formal statements as hypotheses. The following parts have conventional proofs in this paper but are not formally verified. Their numerical checks are as described here and in the relevant sections; in particular, a multiple-precision audit is distinct from a formally verified enclosure.

  • the per-character costs of Section 5: the localized envelope, which enters as a hypothesis, and the first-family bounds (the enclosures of JoldJ_{\mathrm{old}} and JnewJ_{\mathrm{new}} are verified, but not the inequalities they bound);

  • the conductor coefficient 14\frac{1}{4} for real characters and the refinement of Section 7.6;

  • the zero-location arguments and the case tree (Sections 7 and 11), that is, the statement that every configuration of zeros falls into a leaf whose zero-level hypotheses it satisfies, including the count, hidden-count and second-family constraints;

  • the exterior regimes (Section 12): the range of Proposition 12.1 and Theorem 12.5, and, for Theorem 12.6, the statement that every configuration with λ1≥1.5\lambda_1 \ge1.5 satisfies the zero-level hypotheses of its program;

  • the two-test row of Section 9.6, which is checked by the replay and compared exactly with a separate implementation in (5), and whose constants were computed only by the interval routines of that implementation;

  • the count-type integers of the leaf programs (checked in the audit of (4));

  • the assembly of Section 13.

A compiled program is trusted to compute what it is proved to compute; its correctness then rests on the Lean compiler and runtime and on the unverified routines that read the data files. The Lean proofs were written by AI agents, like the rest of this work (see the note after the abstract). The Lean kernel checks the proofs, but whether each formal statement says what the corresponding statement of this paper says has so far been reviewed only by AI agents.

Data availability

The development repository contains the certificates, the programs that generate and check them, the verification reports, and the Lean formalization:

https://github.com/enaslund/linniks-constant.

At the date of this revision the repository is private. A public, versioned archive of these materials is needed before submission so that readers can reproduce the computations and identify the precise data used in the proof. This paper and the Lean formalization are publicly available in the repository prepared for the Palomar registry of formalized mathematics:

https://github.com/enaslund/linniks-constant-3.99.

The repository includes instructions for rerunning the replay, mutation tests, numerical audits, and formal checks. On the machine used, the replay took 1.7 hours of wall-clock time with parallel workers, and the compiled Lean checks of all certificates, leaf data and near rows took about 5 hours of wall-clock time. One caveat concerns exact reproducibility: the sieve heights of the graded rows (Section 9.2) are regenerated in binary64 floating point and are not stored in the records. Any nonnegative heights are admissible, so this does not affect soundness, but on a platform whose elementary functions round differently the regenerated rows, and hence the certificates that use them, could differ and would have to be recomputed. Thus the reproducibility requirement concerns the particular stored certificates; admissibility of the analytic sieve weight alone does not guarantee that a different regenerated row will pass those certificates.

Concluding remarks

The difficult configurations. The tightest certified cases have a nonreal first character with λ1∈[0.64,0.80]\lambda_{1} \in[0.64, 0.80] and a nearby second family. The numerical experiments of Section 1.7 suggest that stronger location bounds in this range would be useful. One limitation of a single near row can be seen directly: if NN entries have the same positive feature vv and diagonal DD, then its quadratic inequality permits N≤D/(v2−d)N \le D/(v^{2}-d) when v2>dv^{2}>d (Example 8.9). This is a restriction of that relaxation, not a lower bound on what other tests or arguments can achieve.

Improved far-density constants may also help. Xylouris estimates [145] that halving the constant in his (5.19) would lower the exponent obtained by his method to about 4.6. A corresponding gain for the present method would require new certificates for the complete cover.

Possible improvements. Heath-Brown’s list of possible improvements [51] contains several that we have not used: optimal test functions for the combined inequalities (his item 1), positivity beyond the support of the test functions via Brun–Titchmarsh bounds (item 3; see [129, 86, 60, 81]), and modified sieve coefficients in the density argument (item 6; compare [35]). His item 7, a continuous weight in the density argument, is what Xylouris’s far-density lemma implements, and we use it. For cube-free moduli Burgess’s bounds hold for every kk, and even the Weyl bound is known [100]; Heath-Brown already noted that all the arguments improve in that case. Recent large-value estimates for Dirichlet polynomials [44, 17, 122] and new density theorems [105] improve zero-density estimates away from σ=1\sigma=1; it is not clear whether they help in the range σ=1−O(1/log⁡q)\sigma=1-O(1/\log q) that matters here.

Other approaches. For moduli with bounded cubic part the exponent 4.5 has been claimed [83], and for moduli all of whose prime factors are small much better exponents are known [16]. A different recent approach, through the exceptional set of Goldbach’s problem (compare [104]), claims the exponent 5 in a preprint [148].

The role of the computation. The graded lemma gives quadratic constraints for different test functions and anchors. Their threshold forms can be combined with the far-density budget in linear programs, whose bounds are certified by exact duality. This separates the search for effective test parameters from the verification of the resulting bound. Any improvement of the global exponent must still include the zero-location arguments, a complete case cover, and the exterior regimes.

Appendix A. The complex location rows

The 94 rows below make the implications of Propositions 7.12 and 7.13 explicit. In each row, the hypothesis is that χ1\chi_{1} is nonreal and λ1∈[a,b]\lambda_{1} \in[a,b]. The number hh is the resulting strict lower bound for λ2\lambda_{2} or λ′\lambda', as specified in the caption. The parameters γ\gamma and tt select the test function (11) and the polynomial PtP_{t}.

Each margin is the interval-certified amount, divided by f(0)f(0), by which the necessary inequality fails at error tolerance 0; the displayed margins are rounded for readability. Their positivity allows a common sufficiently large qq. For the λ2\lambda_{2} rows, the real-second-character case uses the test parameter γr\gamma_{r} and t′=0.85t'=0.85; the column labelled “alias” gives the margin in (13). The λ′\lambda' rows use the parent lower bound described after Proposition 7.13 and assume ord⁡χ1≥5\operatorname{ord}\chi_{1}\ge5. For orders 3 and 4, the stronger published bound quoted there applies.

[a,b][a,b]hhγ\gammattγr\gamma_{r}genericrealalias
[.66,.6625][.66,.6625].763.7631.31.3.95.951.151.150.001460.001460.03230.03230.16370.1637
[.6625,.665][.6625,.665].762.7621.31.3.95.951.151.150.000840.000840.03140.03140.16390.1639
[.665,.6675][.665,.6675].76.761.31.3.9.91.151.150.001140.001140.03110.03110.16440.1644
[.6675,.67][.6675,.67].758.7581.31.3.9.91.151.150.001480.001480.03080.03080.16470.1647
[.67,.6725][.67,.6725].757.7571.31.3.9.91.151.150.000910.000910.02990.02990.16490.1649
[.6725,.675][.6725,.675].755.7551.31.3.9.91.151.150.001260.001260.02960.02960.16510.1651
[.675,.6775][.675,.6775].753.7531.31.3.9.91.151.150.001620.001620.02930.02930.16540.1654
[.6775,.68][.6775,.68].752.7521.31.3.9.91.151.150.001060.001060.02840.02840.16550.1655
[.68,.6825][.68,.6825].75.751.31.3.9.91.151.150.001430.001430.02810.02810.16580.1658
[.6825,.685][.6825,.685].749.7491.31.3.9.91.151.150.000880.000880.02720.02720.16580.1658
[.685,.6875][.685,.6875].747.7471.31.3.9.91.151.150.001260.001260.02690.02690.16580.1658
[.6875,.69][.6875,.69].745.7451.31.3.9.91.151.150.001640.001640.02660.02660.16580.1658
[.69,.6925][.69,.6925].744.7441.31.3.9.91.151.150.001100.001100.02580.02580.16580.1658
[.6925,.695][.6925,.695].742.7421.31.3.9.91.151.150.001500.001500.02540.02540.16580.1658
[.695,.6975][.695,.6975].741.7411.31.3.9.91.151.150.000950.000950.02460.02460.16580.1658
[.6975,.7][.6975,.7].739.7391.31.3.9.91.11.10.001360.001360.02430.02430.16570.1657
[.7,.7025][.7,.7025].738.7381.31.3.9.91.11.10.000820.000820.02360.02360.16570.1657
[.7025,.705][.7025,.705].736.7361.31.3.9.91.11.10.001240.001240.02340.02340.16570.1657
[.705,.7075][.705,.7075].734.7341.31.3.9.91.11.10.001670.001670.02320.02320.16570.1657
[.7075,.71][.7075,.71].733.7331.31.3.9.91.11.10.001140.001140.02250.02250.16570.1657
[.71,.7125][.71,.7125].731.7311.31.3.9.91.11.10.001580.001580.02230.02230.16570.1657
[.7125,.715][.7125,.715].73.731.31.3.9.91.11.10.001060.001060.02160.02160.16570.1657
[.715,.7175][.715,.7175].728.7281.251.25.95.951.11.10.001460.001460.02140.02140.16570.1657
[.7175,.72][.7175,.72].727.7271.251.25.95.951.11.10.001050.001050.02070.02070.16570.1657
[.72,.7225][.72,.7225].725.7251.251.25.95.951.11.10.001550.001550.02060.02060.16570.1657

Table 13. Rows of Proposition 7.12: λ1∈[a,b]\lambda_{1}\in[a,b] implies λ2>h\lambda_{2}>h.

[a,b][a,b]hhγ\gammattmargin[a,b][a,b]hhγ\gammattmargin
[.68,.6825][.68,.6825].975.9751.151.15.95.950.001790.00179[.77,.7725][.77,.7725].908.9081.11.1.95.950.001690.00169
[.6825,.685][.6825,.685].973.9731.151.15.95.950.001750.00175[.7725,.775][.7725,.775].906.9061.11.1.95.950.001870.00187
[.685,.6875][.685,.6875].971.9711.151.15.95.950.001710.00171[.775,.7775][.775,.7775].904.9041.11.1.95.950.002050.00205
[.6875,.69][.6875,.69].968.9681.151.15.95.950.002280.00228[.7775,.78][.7775,.78].903.9031.11.1.95.950.001610.00161
[.69,.6925][.69,.6925].966.9661.151.15.95.950.002250.00225[.78,.7825][.78,.7825].901.9011.11.1.95.950.001800.00180
[.6925,.695][.6925,.695].964.9641.151.15.95.950.002220.00222[.7825,.785][.7825,.785].899.8991.11.1.95.950.001990.00199
[.695,.6975][.695,.6975].962.9621.151.15.95.950.002200.00220[.785,.7875][.785,.7875].897.8971.11.1.95.950.002190.00219
[.7,.7025][.7,.7025].958.9581.151.15.95.950.002160.00216[.7875,.79][.7875,.79].896.8961.11.1.95.950.001750.00175
[.7025,.705][.7025,.705].956.9561.151.15.95.950.002150.00215[.79,.7925][.79,.7925].894.8941.11.1.95.950.001960.00196
[.705,.7075][.705,.7075].954.9541.151.15.95.950.002140.00214[.7925,.795][.7925,.795].892.8921.11.1.95.950.002170.00217
[.7075,.71][.7075,.71].952.9521.151.15.95.950.002130.00213[.795,.7975][.795,.7975].891.8911.11.1.95.950.001730.00173
[.71,.7125][.71,.7125].95.951.151.15.9.90.002140.00214[.7975,.8][.7975,.8].889.8891.11.1.95.950.001950.00195
[.7125,.715][.7125,.715].948.9481.151.15.9.90.002180.00218[.8,.8025][.8,.8025].887.8871.11.1.95.950.002170.00217
[.715,.7175][.715,.7175].946.9461.151.15.9.90.002220.00222[.8025,.805][.8025,.805].886.8861.11.1.95.950.001740.00174

Table 14. Rows of Proposition 7.13: λ1∈[a,b]\lambda_{1}\in[a,b] implies λ′>h\lambda'>h (for χ1\chi_{1} of order at least 5).

[a,b][a,b]hhγ\gammattmargin[a,b][a,b]hhγ\gammattmargin
[.7175,.72][.7175,.72].945.9451.11.1.95.950.001700.00170[.805,.8075][.805,.8075].884.8841.11.1.9.90.001970.00197
[.72,.7225][.72,.7225].943.9431.11.1.95.950.001810.00181[.8075,.81][.8075,.81].883.8831.11.1.9.90.001570.00157
[.7225,.725][.7225,.725].941.9411.11.1.95.950.001930.00193[.81,.8125][.81,.8125].881.8811.11.1.9.90.001840.00184
[.725,.7275][.725,.7275].939.9391.11.1.95.950.002050.00205[.8125,.815][.8125,.815].879.8791.11.1.9.90.002120.00212
[.7275,.73][.7275,.73].937.9371.11.1.95.950.002170.00217[.815,.8175][.815,.8175].878.8781.11.1.9.90.001730.00173
[.73,.7325][.73,.7325].936.9361.11.1.95.950.001700.00170[.8175,.82][.8175,.82].876.8761.11.1.9.90.002020.00202
[.7325,.735][.7325,.735].934.9341.11.1.95.950.001830.00183[.82,.8225][.82,.8225].875.8751.11.1.9.90.001630.00163
[.735,.7375][.735,.7375].932.9321.11.1.95.950.001960.00196[.8225,.825][.8225,.825].873.8731.11.1.9.90.001920.00192
[.7375,.74][.7375,.74].93.931.11.1.95.950.002100.00210[.825,.8275][.825,.8275].872.8721.11.1.9.90.001530.00153
[.74,.7425][.74,.7425].928.9281.11.1.95.950.002230.00223[.8275,.83][.8275,.83].87.871.11.1.9.90.001830.00183
[.7425,.745][.7425,.745].927.9271.11.1.95.950.001770.00177[.83,.8325][.83,.8325].868.8681.11.1.9.90.002140.00214
[.745,.7475][.745,.7475].925.9251.11.1.95.950.001920.00192[.8325,.835][.8325,.835].867.8671.11.1.9.90.001750.00175
[.7475,.75][.7475,.75].923.9231.11.1.95.950.002070.00207[.835,.8375][.835,.8375].865.8651.11.1.9.90.002070.00207
[.75,.7525][.75,.7525].921.9211.11.1.95.950.002220.00222[.8375,.84][.8375,.84].864.8641.11.1.9.90.001690.00169
[.7525,.755][.7525,.755].92.921.11.1.95.950.001760.00176[.84,.8425][.84,.8425].862.8621.11.1.9.90.002000.00200
[.755,.7575][.755,.7575].918.9181.11.1.95.950.001920.00192[.8425,.845][.8425,.845].861.8611.11.1.9.90.001630.00163
[.7575,.76][.7575,.76].916.9161.11.1.95.950.002080.00208[.845,.8475][.845,.8475].859.8591.11.1.9.90.001950.00195
[.76,.7625][.76,.7625].915.9151.11.1.95.950.001630.00163[.8475,.85][.8475,.85].858.8581.11.1.9.90.001580.00158
[.7625,.765][.7625,.765].913.9131.11.1.95.950.001790.00179[.85,.8525][.85,.8525].856.8561.051.05.95.950.001930.00193
[.765,.7675][.765,.7675].911.9111.11.1.95.950.001960.00196[.8525,.855][.8525,.855].855.8551.051.05.95.950.001660.00166
[.7675,.77][.7675,.77].909.9091.11.1.95.950.002140.00214

Table 14.

Appendix B. The fixed constants

This appendix collects the constants and parameters of the proof. The error budget that they serve is summarized in Section 2.7.

References

Email address: [email protected]

References

  1. [1]David L. Applegate, William Cook, Sanjeeb Dash, and Daniel G. Espinoza, Exact solutions to linear programming problems, Oper. Res. Lett. 35 (2007), no. 6, 693–699.DOI
  2. [2]Jeremy Avigad, Mathematics and the formal turn, Bull. Amer. Math. Soc. (N.S.) 61 (2024), no. 2, 225–240.arxiv.org/abs/2311.00007
  3. [3]Jeremy Avigad, Kevin Donnelly, David Gray, and Paul Raff, A formally verified proof of the prime number theorem, ACM Trans. Comput. Log. 9 (2007), no. 1, Art. 2.DOI
  4. [4]Eric Bach and Jonathan Sorenson, Explicit bounds for primes in residue classes, Math. Comp. 65 (1996), no. 216, 1717–1735.
  5. [5]Chris Bailey, nanoda lib, GitHub repository ammkrn/nanoda_lib, 2020, an external type checker for Lean 4.
  6. [6]William D. Banks and Igor E. Shparlinski, Bounds on short character sums and L-functions with characters to a powerful modulus, J. Anal. Math. 139 (2019), no. 1, 239–263.DOI
  7. [7]M. B. Barban, Yu. V. Linnik, and N. G. Chudakov, On prime numbers in an arithmetic progression with a prime-power difference, Acta Arith. 9 (1964), no. 4, 375–390.DOI
  8. [8]Kübra Benli, Shivani Goel, Henry Twiss, and Asif Zaman, Explicit Deuring–Heilbronn phenomenon for Dirichlet L-functions, Proc. Amer. Math. Soc. 154 (2026), no. 2, 509–525.DOI
  9. [9]Michael A. Bennett, Greg Martin, Kevin O’Bryant, and Andrew Rechnitzer, Explicit bounds for primes in arithmetic progressions, Illinois J. Math. 62 (2018), no. 1–4, 427–532.arxiv.org/abs/1802.00085
  10. [10]Enrico Bombieri, Le grand crible dans la théorie analytique des nombres, second ed., Astérisque, vol. 18, Société Mathématique de France, Paris, 1987 (French), First edition 1974.
  11. [11]D. A. Burgess, On character sums and L-series, Proc. London Math. Soc. (3) 12 (1962), 193–206.DOI
  12. [12]———, On character sums and primitive roots, Proc. London Math. Soc. (3) 12 (1962), 179–192.DOI
  13. [13]———, On character sums and L-series. II, Proc. London Math. Soc. (3) 13 (1963), 524–536.DOI
  14. [14]———, The character sum estimate with r = 3, J. London Math. Soc. (2) 33 (1986), no. 2, 219–226.DOI
  15. [15]Emanuel Carneiro, Micah B. Milinovich, Emily Quesada-Herrera, and Antonio Pedro Ramos, Fourier optimization, the least quadratic non-residue, and the least prime in an arithmetic progression, Math. Comp. (2025), to appear; published electronically 2 December 2025, doi:10.1090/mcom/4154.DOI
  16. [16]Mei-Chu Chang, Short character sums for composite moduli, J. Anal. Math. 123 (2014), no. 1, 1–33.arxiv.org/abs/1201.0299
  17. [17]Bin Chen, Vishal Gupta, and Yung Chi Li, Large value estimates for Dirichlet polynomials with characters and zero density of Dirichlet L-functions, 2025, arXiv:2507.08296.arxiv.org/abs/2507.08296
  18. [18]Jingrun Chen, On the least prime in an arithmetical progression, Sci. Sinica 14 (1965), 1868–1871. MR 188172
  19. [19]———, On the least prime in an arithmetical progression and two theorems concerning the zeros of Dirichlet’s L-functions, Sci. Sinica 20 (1977), no. 5, 529–562. MR 476668
  20. [20]———, On the least prime in an arithmetical progression and theorems concerning the zeros of Dirichlet’s L-functions. II, Sci. Sinica 22 (1979), no. 8, 859–889. MR 549597
  21. [21]Jingrun Chen and Jianmin Liu, On the least prime in an arithmetical progression. III, Sci. China Ser. A 32 (1989), no. 6, 654–673. MR 1056044
  22. [22]———, On the least prime in an arithmetical progression. IV, Sci. China Ser. A 32 (1989), no. 7, 792–807. MR 1058000
  23. [23]———, On the least prime in an arithmetical progression and theorems concerning the zeros of Dirichlet’s L-functions (V), International Symposium in Memory of Hua Loo Keng, Vol. I: Number Theory (Sheng Gong, Qi-Keng Lu, Yuan Wang, and Lo Yang, eds.), Springer, Berlin, 1991, pp. 19–42.DOI
  24. [24]S. Chowla, On the least prime in an arithmetical progression, J. Indian Math. Soc. (N.S.) 1 (1934), 1–3.DOI
  25. [25]Vašek Chvátal, Linear programming, A Series of Books in the Mathematical Sciences, W. H. Freeman and Company, New York and San Francisco, 1983.
  26. [26]Harold Davenport, Multiplicative number theory, third ed., Graduate Texts in Mathematics, vol. 74, Springer-Verlag, New York, 2000, Revised and with a preface by Hugh L. Montgomery.
  27. [27]Ch.-J. de la Vallée Poussin, Recherches analytiques sur la théorie des nombres premiers, Ann. Soc. Sci. Bruxelles 20 (1896), 183–256, 281–397 (French).
  28. [28]———, Sur la fonction ζ(s) de Riemann et le nombre des nombres premiers inférieurs à une limite donnée, Mém. Couronnés et Autres Mém. Publ. Acad. Roy. Sci. Lett. Beaux-Arts Belg. 59 (1899), 1–74 (French).DOI
  29. [29]Leonardo de Moura and Sebastian Ullrich, The Lean 4 theorem prover and programming language, Automated Deduction – CADE 28 (André Platzer and Geoff Sutcliffe, eds.), Lecture Notes in Comput. Sci., vol. 12699, Springer, 2021, pp. 625–635.DOI
  30. [30]Max Deuring, Imaginäre quadratische Zahlkörper mit der Klassenzahl 1, Math. Z. 37 (1933), no. 1, 405–415.
  31. [31]P. G. Lejeune Dirichlet, Beweis des Satzes, dass jede unbegrenzte arithmetische Progression, deren erstes Glied und Differenz ganze Zahlen ohne gemeinschaftlichen Factor sind, unendlich viele Primzahlen enthält, G. Lejeune Dirichlet’s Werke, Vol. 1 (L. Kronecker, ed.), Reimer, Berlin, 1889, Reissued by Cambridge University Press, 2012, pp. 313–342.
  32. [32]E. Fogels, On the zeros of L-functions, Acta Arith. 11 (1965), no. 1, 67–96.DOI
  33. [33]J. B. Friedlander and H. Iwaniec, Exceptional characters and prime numbers in arithmetic progressions, Int. Math. Res. Not. 2003 (2003), no. 37, 2033–2050.
  34. [34]———, Opera de cribro, Amer. Math. Soc. Colloq. Publ., vol. 57, Amer. Math. Soc., Providence, RI, 2010.DOI
  35. [35]———, Selberg’s sieve of irregular density, Acta Arith. 209 (2023), 385–396.arxiv.org/abs/2206.03479
  36. [36]———, Sifting for small primes from an arithmetic progression, Sci. China Math. 66 (2023), no. 12, 2715–2730.arxiv.org/abs/2303.06122
  37. [37]P. X. Gallagher, A large sieve density estimate near σ = 1, Invent. Math. 11 (1970), no. 4, 329–339.DOI
  38. [38]———, Primes in progressions to prime-power modulus, Invent. Math. 16 (1972), no. 3, 191–201.DOI
  39. [39]S. W. Graham, Applications of sieve methods, Ph.D. thesis, University of Michigan, 1977, 187 pp., ProQuest 7804710. MR 2627480DOI
  40. [40]———, An asymptotic estimate related to Selberg’s sieve, J. Number Theory 10 (1978), no. 1, 83–94.DOI
  41. [41]———, On Linnik’s constant, Acta Arith. 39 (1981), no. 2, 163–179.DOI
  42. [42]Andrew Granville and Carl Pomerance, On the least prime in certain arithmetic progressions, J. London Math. Soc. (2) 41 (1990), no. 2, 193–200.DOI
  43. [43]T. H. Gronwall, Sur les séries de Dirichlet correspondant à des caractères complets, Rend. Circ. Mat. Palermo 35 (1913), 145–159.DOI
  44. [44]Larry Guth and James Maynard, New large value estimates for Dirichlet polynomials, Ann. of Math. (2) 203 (2026), no. 2, 623–675.arxiv.org/abs/2405.20552
  45. [45]H. Halberstam and H.-E. Richert, Sieve methods, London Math. Soc. Monogr., vol. 4, Academic Press, London, 1974.
  46. [46]Thomas Hales, Mark Adams, Gertrud Bauer, Tat Dat Dang, John Harrison, Le Truong Hoang, Cezary Kaliszyk, Victor Magron, Sean McLaughlin, Tat Thang Nguyen, Quang Truong Nguyen, Tobias Nipkow, Steven Obua, Joseph Pleso, Jason Rute, Alexey Solovyev, Thi Hoai An Ta, Nam Trung Tran, Thi Diep Trieu, Josef Urban, Ky Vu, and Roland Zumkeller, A formal proof of the Kepler conjecture, Forum Math. Pi 5 (2017), e2.
  47. [47]John Harrison, Formalizing an analytic proof of the prime number theorem, J. Automat. Reason. 43 (2009), no. 3, 243–261.DOI
  48. [48]D. R. Heath-Brown, Hybrid bounds for Dirichlet L-functions, Invent. Math. 47 (1978), no. 2, 149–170.
  49. [49]———, The density of zeros of Dirichlet’s L-functions, Canad. J. Math. 31 (1979), no. 2, 231–240.
  50. [50]———, Siegel zeros and the least prime in an arithmetic progression, Quart. J. Math. Oxford Ser. (2) 41 (1990), no. 4, 405–418.
  51. [51]———, Zero-free regions for Dirichlet L-functions, and the least prime in an arithmetic progression, Proc. London Math. Soc. (3) 64 (1992), no. 2, 265–338.
  52. [52]———, Burgess’s bounds for character sums, Number Theory and Related Fields: In Memory of Alf van der Poorten (Jonathan M. Borwein, Igor Shparlinski, and Wadim Zudilin, eds.), Springer Proc. Math. Stat., vol. 43, Springer, New York, 2013, pp. 199–213.
  53. [53]Hans Heilbronn, On the class-number in imaginary quadratic fields, Quart. J. Math. Oxford Ser. 5 (1934), no. 1, 150–160.DOI
  54. [54]Harald Andrés Helfgott, The ternary Goldbach problem, 2015, arXiv:1501.05438.arxiv.org/abs/1501.05438
  55. [55]Q. Huangfu and J. A. J. Hall, Parallelizing the dual revised simplex method, Math. Program. Comput. 10 (2018), no. 1, 119–142.arxiv.org/abs/1503.01889
  56. [56]M. N. Huxley, On the difference between consecutive primes, Invent. Math. 15 (1972), no. 2, 164–170.DOI
  57. [57]A. E. Ingham, The distribution of prime numbers, Cambridge Tracts in Mathematics and Mathematical Physics, vol. 30, Cambridge University Press, Cambridge, 1932.
  58. [58]———, On the estimation of N(σ, T), Quart. J. Math. Oxford Ser. 11 (1940), no. 1, 201–202.DOI
  59. [59]H. Iwaniec, On zeros of Dirichlet’s L series, Invent. Math. 23 (1974), no. 2, 97–104.DOI
  60. [60]———, On the Brun–Titchmarsh theorem, J. Math. Soc. Japan 34 (1982), no. 1, 95–123.DOI
  61. [61]———, Conversations on the exceptional character, Analytic Number Theory: Lectures given at the C.I.M.E. Summer School held in Cetraro, Italy, July 11–18, 2002, Lecture Notes in Math., vol. 1891, Springer, Berlin, 2006, pp. 97–132.DOI
  62. [62]H. Iwaniec and Emmanuel Kowalski, Analytic number theory, Amer. Math. Soc. Colloq. Publ., vol. 53, Amer. Math. Soc., Providence, RI, 2004.
  63. [63]Christian Jansson, Rigorous lower and upper bounds in linear programming, SIAM J. Optim. 14 (2004), no. 3, 914–935.DOI
  64. [64]Matti Jutila, A new estimate for Linnik’s constant, Ann. Acad. Sci. Fenn. Ser. A I 471 (1970), 8 pp. MR 271056DOI
  65. [65]———, On a density theorem of H. L. Montgomery for L-functions, Ann. Acad. Sci. Fenn. Ser. A I 520 (1972), 13 pp. MR 327681DOI
  66. [66]———, On Linnik’s constant, Math. Scand. 41 (1977), 45–62.DOI
  67. [67]Habiba Kadiri, Une région explicite sans zéros pour la fonction ζ de Riemann, Acta Arith. 117 (2005), no. 4, 303–339 (French).DOI
  68. [68]———, Explicit zero-free regions for Dirichlet L-functions, Mathematika 64 (2018), no. 2, 445–474.
  69. [69]Anatolij A. Karatsuba, Basic analytic number theory, Springer-Verlag, Berlin, 1993, Translated from the Russian by Melvyn B. Nathanson.DOI
  70. [70]Bryce Kerr, Moments of character sums to composite modulus, 2019, arXiv:1904.04578.arxiv.org/abs/1904.04578
  71. [71]S. Knapowski, On Linnik’s theorem concerning exceptional L-zeros, Publ. Math. Debrecen 9 (1962), 168–178.
  72. [72]Alex Kontorovich, Terence Tao, et al., Prime number theorem and ..., GitHub repository AlexKontorovich/PrimeNumberTheoremAnd blueprint (the PNT+ project), 2024.
  73. [73]Youness Lamzouri, Xiannan Li, and Kannan Soundararajan, Conditional bounds for the least quadratic non-residue and related problems, Math. Comp. 84 (2015), no. 295, 2391–2412, Corrigendum: Math. Comp. 86 (2017), no. 307, 2551–2554, doi:10.1090/mcom/3261.
  74. [74]E. Landau, Über imaginär-quadratische Zahlkörper mit gleicher Klassenzahl, Nachr. Ges. Wiss. Göttingen, Math.-Phys. Kl. (1918), 277–284.
  75. [75]Lean FRO, Comparator, GitHub repository leanprover/comparator, 2025, a trustworthy judge for Lean proofs.
  76. [76]Junxian Li, Kyle Pratt, and George Shakan, A lower bound for the least prime in an arithmetic progression, Quart. J. Math. 68 (2017), no. 3, 729–758.arxiv.org/abs/1607.02543
  77. [77]Yu. V. Linnik, On the least prime in an arithmetic progression. I. The basic theorem, Rec. Math. [Mat. Sbornik] N.S. 15(57) (1944), 139–178. MR 12111
  78. [78]———, On the least prime in an arithmetic progression. II. The Deuring–Heilbronn phenomenon, Rec. Math. [Mat. Sbornik] N.S. 15(57) (1944), 347–368. MR 12112
  79. [79]Ming-Chit Liu and Tianze Wang, A numerical bound for small prime solutions of some ternary linear equations, Acta Arith. 86 (1998), no. 4, 343–383.DOI
  80. [80]Kaisa Matomäki, Jori Merikoski, and Joni Teräväinen, Primes in arithmetic progressions and short intervals without L-functions, 2024, arXiv:2401.17570.arxiv.org/abs/2401.17570
  81. [81]James Maynard, On the Brun–Titchmarsh theorem, Acta Arith. 157 (2013), no. 3, 249–296.arxiv.org/abs/1201.1777
  82. [82]Kevin S. McCurley, Explicit zero-free regions for Dirichlet L-functions, J. Number Theory 19 (1984), no. 1, 7–32.DOI
  83. [83]Zaizhao Meng, A note on the Linnik’s constant, 2010, arXiv:1010.3544.arxiv.org/abs/1010.3544
  84. [84]H. L. Montgomery, Topics in multiplicative number theory, Lecture Notes in Math., vol. 227, Springer-Verlag, Berlin, 1971.
  85. [85]———, Ten lectures on the interface between analytic number theory and harmonic analysis, CBMS Regional Conference Series in Mathematics, vol. 84, Amer. Math. Soc., Providence, RI, 1994.DOI
  86. [86]H. L. Montgomery and R. C. Vaughan, The large sieve, Mathematika 20 (1973), no. 2, 119–134.
  87. [87]———, Multiplicative number theory I. Classical theory, Cambridge Stud. Adv. Math., vol. 97, Cambridge University Press, Cambridge, 2007.
  88. [88]Ramon E. Moore, Interval analysis, Prentice-Hall Series in Automatic Computation, Prentice-Hall, Englewood Cliffs, N. J., 1966.
  89. [89]Ramon E. Moore, R. Baker Kearfott, and Michael J. Cloud, Introduction to interval analysis, Society for Industrial and Applied Mathematics (SIAM), Philadelphia, PA, 2009.DOI
  90. [90]Michael J. Mossinghoff and Timothy S. Trudgian, Nonnegative trigonometric polynomials and a zero-free region for the Riemann zeta-function, J. Number Theory 157 (2015), 329–349.arxiv.org/abs/1410.3926
  91. [91]Y. Motohashi, On a density theorem of Linnik, Proc. Japan Acad. 51 (1975), 815–817.DOI
  92. [92]———, Primes in arithmetic progressions, Invent. Math. 44 (1978), no. 2, 163–178.DOI
  93. [93]———, Lectures on sieve methods and prime number theory, Tata Inst. Fund. Res. Lectures on Math. and Phys., vol. 72, Tata Institute of Fundamental Research and Springer-Verlag, Berlin, 1983.
  94. [94]Władysław Narkiewicz, The development of prime number theory: From Euclid to Hardy and Littlewood, Springer Monographs in Mathematics, Springer-Verlag, Berlin, 2000.
  95. [95]Arnold Neumaier and Oleg Shcherbina, Safe bounds in linear and mixed-integer linear programming, Math. Program. 99 (2004), no. 2, 283–296.DOI
  96. [96]Steven Obua and Tobias Nipkow, Flyspeck II: the basic linear programs, Ann. Math. Artif. Intell. 56 (2009), no. 3–4, 245–272.DOI
  97. [97]A. Page, On the number of primes in an arithmetic progression, Proc. London Math. Soc. (2) 39 (1935), no. 1, 116–141.DOI
  98. [98]Chengdong Pan, On the least prime in an arithmetical progression, Sci. Record (N.S.) 1 (1957), 311–313. MR 105398
  99. [99]———, On the least prime in an arithmetical progression, Acta Sci. Natur. Univ. Pekinensis (1958), no. 1, 3–36 (Chinese).
  100. [100]Ian Petrow and Matthew P. Young, The Weyl bound for Dirichlet L-functions of cube-free conductor, Ann. of Math. (2) 192 (2020), no. 2, 437–486.DOI
  101. [101]E. Phragmén and Ernst Lindelöf, Sur une extension d’un principe classique de l’analyse et sur quelques propriétés des fonctions monogènes dans le voisinage d’un point singulier, Acta Math. 31 (1908), 381–406.
  102. [102]J. Pintz, Elementary methods in the theory of L-functions, V. The theorems of Landau and Page, Acta Arith. 32 (1977), no. 2, 163–171.DOI
  103. [103]———, Elementary methods in the theory of L-functions, VIII. Real zeros of real L-functions, Acta Arith. 33 (1977), no. 1, 89–98.DOI
  104. [104]———, A new explicit formula in the additive theory of primes with applications II. The exceptional set in Goldbach’s problem, 2018, arXiv:1804.09084.arxiv.org/abs/1804.09084
  105. [105]———, Some new density theorems for Dirichlet L-functions, Banach Center Publ. 118 (2019), 231–244.arxiv.org/abs/1804.05552
  106. [106]D. J. Platt, Numerical computations concerning the GRH, Math. Comp. 85 (2016), no. 302, 3009–3027.arxiv.org/abs/1305.3087
  107. [107]D. J. Platt and Tim Trudgian, The Riemann hypothesis is true up to 3 · 10¹², Bull. Lond. Math. Soc. 53 (2021), no. 3, 792–797.DOI
  108. [108]G. Pólya, Über die Verteilung der quadratischen Reste und Nichtreste, Nachr. Ges. Wiss. Göttingen, Math.-Phys. Kl. (1918), 21–29.
  109. [109]Carl Pomerance, A note on the least prime in an arithmetic progression, J. Number Theory 12 (1980), no. 2, 218–223.DOI
  110. [110]Karl Prachar, Primzahlverteilung, Grundlehren der Mathematischen Wissenschaften, vol. 91, Springer-Verlag, Berlin, 1957 (German).
  111. [111]K. A. Rodosskiĭ, On the least prime number in an arithmetic progression, Mat. Sbornik N.S. 34(76) (1954), 331–356. MR 62156
  112. [112]J. Barkley Rosser and Lowell Schoenfeld, Approximate formulas for some functions of prime numbers, Illinois J. Math. 6 (1962), no. 1, 64–94.
  113. [113]Siegfried M. Rump, Verification methods: rigorous results using floating-point arithmetic, Acta Numer. 19 (2010), 287–449.DOI
  114. [114]Stelios Sachpazis, A pretentious proof of Linnik’s estimate for primes in arithmetic progressions, Mathematika 69 (2023), no. 3, 879–902.
  115. [115]Alexander Schrijver, Theory of linear and integer programming, Wiley-Interscience Series in Discrete Mathematics, John Wiley & Sons, Chichester, 1986.
  116. [116]Atle Selberg, On an elementary method in the theory of primes, Norske Vid. Selsk. Forh., Trondhjem 19 (1947), no. 18, 64–67. MR 22871
  117. [117]———, Lectures on sieves, Collected Papers, Vol. II, Springer-Verlag, Berlin, 1991, pp. 65–247.
  118. [118]C. L. Siegel, Über die Classenzahl quadratischer Zahlkörper, Acta Arith. 1 (1935), no. 1, 83–86.
  119. [119]Alexey Solovyev and Thomas C. Hales, Efficient formal verification of bounds of linear programs, Intelligent Computer Mathematics (Calculemus 2011 and MKM 2011, Bertinoro, Italy) (James H. Davenport, William M. Farmer, Josef Urban, and Florian Rabe, eds.), Lecture Notes in Comput. Sci., vol. 6824, Springer, 2011, pp. 123–132.DOI
  120. [120]Kannan Soundararajan and Jesse Thorner, Weak subconvexity without a Ramanujan hypothesis, Duke Math. J. 168 (2019), no. 7, 1231–1268.DOI
  121. [121]Michael Stoll, Dirichlet’s theorem on primes in arithmetic progression, Lean 4 Mathlib source file Mathlib/NumberTheory/LSeries/PrimesInAP.lean, 2024.
  122. [122]Terence Tao, Tim Trudgian, and Andrew Yang, New exponent pairs, zero density estimates, and zero additive energy estimates: a systematic approach, Math. Comp. 95 (2026), no. 362, 2941–2990.arxiv.org/abs/2501.16779
  123. [123]Tikao Tatuzawa, On a theorem of Siegel, Jpn. J. Math. 21 (1951), 163–178. MR 51262DOI
  124. [124]Gérald Tenenbaum, Introduction to analytic and probabilistic number theory, third ed., Graduate Studies in Mathematics, vol. 163, Amer. Math. Soc., Providence, RI, 2015, Translated from the 2008 French edition by Patrick D. F. Ion. MR 3363366DOI
  125. [125]The mathlib Community, The Lean mathematical library, Proceedings of the 9th ACM SIGPLAN International Conference on Certified Programs and Proofs (CPP 2020), ACM, 2020, arXiv:1910.09336, pp. 367–381.DOI
  126. [126]The mpmath development team, mpmath: a Python library for arbitrary-precision floating-point arithmetic (version 1.3.0), 2023, https://mpmath.org/.
  127. [127]Jesse Thorner and Asif Zaman, An explicit version of Bombieri’s log-free density estimate and Sárközy’s theorem for shifted primes, Forum Math. 36 (2024), no. 4, 1059–1080.
  128. [128]———, Refinements to the prime number theorem for arithmetic progressions, Math. Z. 306 (2024), no. 3, Paper No. 54.
  129. [129]E. C. Titchmarsh, A divisor problem, Rend. Circ. Mat. Palermo 54 (1930), 414–429.DOI
  130. [130]———, The theory of functions, second ed., Oxford University Press, Oxford, 1939.
  131. [131]———, The theory of the Riemann zeta-function, second ed., Clarendon Press, Oxford, 1986. Revised by D. R. Heath-Brown.
  132. [132]Warwick Tucker, Validated numerics: A short introduction to rigorous computations, Princeton University Press, Princeton, NJ, 2011.DOI
  133. [133]P. Turán, On a density theorem of Yu. V. Linnik, Magyar Tud. Akad. Mat. Kutató Int. Közl. 6 (1961), 165–179. MR 146155
  134. [134]———, On some recent results in the analytical theory of numbers, 1969 Number Theory Institute (State Univ. New York, Stony Brook, N.Y., 1969), Proc. Sympos. Pure Math., vol. 20, Amer. Math. Soc., Providence, RI, 1971, pp. 359–374. MR 316400DOI
  135. [135]———, On a new method of analysis and its applications, John Wiley & Sons, New York, 1984, a Wiley-Interscience Publication.
  136. [136]I. M. Vinogradov, Sur la distribution des résidus et des nonrésidus des puissances, J. Soc. Phys.-Math. Univ. Perm 1 (1918), 94–98.
  137. [137]Pauli Virtanen, Ralf Gommers, Travis E. Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, et al., SciPy 1.0: fundamental algorithms for scientific computing in Python, Nature Methods 17 (2020), no. 3, 261–272.
  138. [138]Samuel S. Wagstaff, Jr., Greatest of the least primes in arithmetic progressions having a given modulus, Math. Comp. 33 (1979), no. 147, 1073–1080.
  139. [139]Wei Wang, On the least prime in an arithmetic progression, Acta Math. Sinica 29 (1986), no. 6, 826–836. MR 888852
  140. [140]———, On the least prime in an arithmetic progression, Acta Math. Sinica (N.S.) 7 (1991), no. 3, 279–289.DOI
  141. [141]André Weil, Sur les ‘formules explicites’ de la théorie des nombres premiers, Comm. Sém. Math. Univ. Lund [Medd. Lunds Univ. Mat. Sem.] (1952), no. Tome supplémentaire, 252–265 (French). MR 53152
  142. [142]David Vernon Widder, The Laplace transform, Princeton Mathematical Series, vol. 6, Princeton University Press, Princeton, N. J., 1941.
  143. [143]Triantafyllos Xylouris, On Linnik’s constant, 2009, arXiv:0906.2749; Diplomarbeit, Universität Bonn (in German; title page: Über die Linnicksche Konstante).arxiv.org/abs/0906.2749
  144. [144]———, On the least prime in an arithmetic progression and estimates for the zeros of Dirichlet L-functions, Acta Arith. 150 (2011), no. 1, 65–91.DOI
  145. [145]———, Über die Nullstellen der Dirichletschen L-Funktionen und die kleinste Primzahl in einer arithmetischen Progression, Dissertation, Universität Bonn, 2011, Bonner Mathematische Schriften 404.DOI
  146. [146]———, Linniks Konstante ist kleiner als 5, Chebyshevskiĭ Sb. 19 (2018), no. 3, 80–94 (German).
  147. [147]Asif Zaman, On the least prime ideal and Siegel zeros, Int. J. Number Theory 12 (2016), no. 8, 2201–2229.DOI
  148. [148]Genheng Zhao, The exceptional set of Goldbach problem and Linnik’s constant, 2025, arXiv:2511.05631.arxiv.org/abs/2511.05631

Paper details

Contents