Let u(U) count the unordered pairs at distance one in a finite planar
set U. We construct finite sets Uj with ∣Uj∣→∞ and
u(Uj)/∣Uj∣1.04273→∞. The largest exponent previously
claimed, 1.0358, is in the author's unpublished manuscript. The
method is the number-field construction of OpenAI and Sawin: unit
distances come from elements of relative norm one in quadratic
extensions, and the fields come from an infinite pro-2 class field
tower. The new ingredients are quadratic extensions of mixed signature,
with an exact average over their norm-one units, and a tower over the
real quadratic field \Q(\sqrt{241}), in which 2, 3 and 5 split.
Its Golod–Shafarevich function contains two copies of the local
conditions at these primes but only one constant term, and the extra room
lets the tower be ramified only above 2, 3 and 5; the root
discriminant of its fields is about 286. The relative zeta value is bounded through the zeta function of the
degree-512 field generated over \Q(\sqrt{241}) by the square roots of
its {2,3,5}-units, which is the product of the Dedekind zeta function
of \Q(\sqrt{241}) and 255 quadratic Hecke L-functions. Finite facts and numerical
inequalities are certified by exact computation and interval arithmetic. A significant portion of this work was verified in Lean, reducing the result with exponent 1.0427 to an explicit zeta function inequality (Palomar registry, PALOMAR-2026-10-01-000018, version 1).
Introduction
For a finite set U⊂R2, let
u(U)=#{{x,y}⊂U:∣x−y∣=1},u(n)=∣U∣=nmaxu(U),
where the pairs are unordered and the distance is Euclidean. The unit distance problem of Erdős [6] asks for the order of growth of u(n). Our main result is the following.
Theorem 1.1.There are finite sets Uj⊂R2, with ∣Uj∣→∞, such that
∣Uj∣1.04273u(Uj)→∞.
In particular, u(n)≥n1.04273 for arbitrarily large integers n.
The largest exponent previously claimed, 1.0358324, is in the author’s unpublished manuscript [23]. Theorem 1.1 uses the same number-field method with several new ingredients: fields of mixed signature, averaged over their norm-one units; an infinite pro-2 extension of the real quadratic field Q(241), with restricted ramification and a nonabelian local Galois group at the primes above 2; an explicit bound for the relative zeta value (defined in Subsection 1.2) from the quadratic L-functions of a Kummer field of degree 512; and optimized profiles: smooth or locally constant weight functions at the archimedean places and at finitely many primes, whose product determines, through its superlevel sets, the region in which lattice points are counted. Subsection 1.3 describes them. The theorem concerns an unbounded sequence of cardinalities. Its proof uses no unproved hypothesis, such as the generalized Riemann hypothesis, but it is computer-assisted: finite facts about the tower and several numerical inequalities are verified by exact computation and interval arithmetic, as described in Subsection 1.6.
Background
Erdős [6] showed that a suitable section of a scaled integer lattice gives
u(n)≥n1+c0/loglogn
for some c0>0 and arbitrarily large n, and he conjectured that this is essentially optimal [6, 7]. The best upper bound,
u(n)=O(n4/3),
is due to Spencer, Szemerédi and Trotter [34]; Székely [36] gave a proof by crossing numbers.
In 2026 a team at OpenAI [27] disproved Erdős’s conjecture. They constructed sets with u(U)≥∣U∣1+δ for a fixed δ>0 and arbitrarily large ∣U∣, replacing the integer lattice by ideal lattices in number fields of large degree and small root discriminant. Alon, Bloom, Gowers, Litt, Sawin, Shankar, Tsimerman, Wang and Wood [1] simplified the argument and obtained δ≈6⋅10−38. Sawin [32] made every step explicit and sharpened it, and reached the exponent 1.014114.
Improvements followed within days in online discussions. On the Erdős Problems forum, mlewko [20] reached 1.031849 in Sawin’s criterion with finite data found by ChatGPT, and on MathOverflow spiderduckpig [35] adjusted these data to reach 1.031883. In a comment the next day, spiderduckpig proposed a structural change: in the Golod–Shafarevich step of Sawin’s argument [32], Lemma 11, replace the unweighted count of relations by a count weighted by degree in the Zassenhaus filtration (recalled in Section 2). This was reported to give δ≥0.0333487[35]. Weighted counts of this kind underlie the towers of [23] and of the present paper. The author’s MathOverflow answer [21] then combined this count with covolume counting, dyadic ramification and Euler products over a fixed subfield. The manuscript [23] is a formal write-up of that construction, and it records δ=0.0358324. Table 1 lists the explicit exponents, and the repository [28] records further bounds posted online.
Table 1. Explicit lower bounds u(n)≥n1+δ for arbitrarily large n, in increasing order. The exponents are rounded down and constant factors are omitted. The last column gives the kind of source; apart from Erdős’s, none of these bounds had appeared in a refereed journal when this paper was written.
The construction in outline
The construction carries Erdős’s lattice argument to number fields of large degree. Let K be such a field, of degree 2d, with an involution ι whose fixed field is F. We take the points of a translated ideal lattice of K that lie in a bounded region of K⊗R. A complex embedding in which ι acts as complex conjugation maps them injectively to the plane. It sends every difference of relative norm one over F to a unit vector, so two points whose difference has relative norm one are at distance one. The exponent is then set by a balance between four quantities, each normalized as d−1 times the logarithm of a count. First, equal-norm choices: if a prime p of F splits in K as P⋅ιP, the k+1 ideals Pj(ιP)k−j, 0≤j≤k, all have relative norm pk. When the construction uses these k+1 ideals, we say that p is used with exponent k, and we call the primes so used the selected primes. Second, ideal classes: a choice of exponents 0≤jp≤kp, one at each selected prime p (used with exponent kp), gives the product of the ideals Pjp(ιP)kp−jp. For the choices whose ideals ∏pPjp lie in one class of Cl(K) modulo the image of Cl(F), there are elements of relative norm one that generate these products times one fixed fractional ideal, and keeping the most frequent class loses at most a factor equal to the order of that quotient (Lemma 5.6). Through the class-number formula, this loss is governed by the root discriminant λ of K and by the relative zeta valueLF(1), where LF(s)=ζK(s)/ζF(s). Third, the region: its shape decides how many pairs of points with difference of relative norm one lie in it, and its volume fixes the number of points. Fourth, density: using a prime p with exponent k also makes the ideal lattice denser by the factor Npk; since the exponent 1+δ is measured against the number of points, this costs δklogNp before division by d.
Quantitatively, suppose that such fields exist with d→∞ and a fixed root discriminant λ. Let (b,c) be the signature of F, so that d=b+2c, and put θ=c/d; thus 2θ is the proportion of non-real embeddings of F. Suppose also that the selected primes have uniform local types: they are all the primes of F above the primes r of a fixed finite set R, each of them splits in K, and all those above a given r have the same absolute ramification index er and residue degree fr (Definitions 2.41 and 5.4). Then F has exactly d/(erfr) primes above each r, and we use all of them with a common exponent kr. If the points are taken from one translate of a fixed ideal lattice, with no further weighting at the selected primes (unweighted shell profiles, Definition 6.5), the exponent 1+δ is attained when
is the archimedean contribution. In it, JR and JC are the functionals of the profiles at the real and at the complex places of F (Definition 5.22); they compare how much a profile overlaps its translates by elements of relative norm one with its mass. We call the left side of (1) the margin. Proposition 5.26, the geometric transfer, makes this precise for quadratic extensions K/F with K totally imaginary, F with a real place and K/F unramified at every finite place: under conditions on the profiles, if the margin is bounded below by a positive constant along such a family, then there are finite sets Uj⊂R2 with ∣Uj∣→∞ and u(Uj)/∣Uj∣1+δ→∞. Section 6 allows general shell profiles, weighted sums of indicator functions of sets defined by valuations (Definition 6.5 and Proposition 6.21). This replaces each summand of the sum over R in (1) by logFδ,r/(erfr), where Fδ,r is the value of the local functional of Definition 6.1 at the shell profile used above r; we still call the resulting left side the margin (Corollary 6.23). Section 8 proves a positive lower bound for it at δ=0.04273.
The fields must form an infinite family of growing degree with fixed λ and fixed types at the selected primes. Such families come from towers, pro-2 extensions with restricted ramification and prescribed local Galois groups, when these are infinite. Golod–Shafarevich arguments [10] show that a tower is infinite when an explicit function of t∈(0,1), built from the number of generators of its Galois group and from its local conditions, takes a negative value; here that function is the Golod–Shafarevich functionPB of (12), and Proposition 2.39 proves the criterion for it. The use of Frobenius conditions to control splitting in such towers goes back to Hajir, Maire and Ramakrishna [11].
New ingredients
The construction differs from those of Sawin [32] and of [23] in four ways.
Mixed signature. In Sawin’s construction, and in [23], the field K is a CM field: complex conjugation is central in Gal(K/Q), F is totally real, every norm-one element has modulus one at every archimedean place, and the norm-one units are finite. Here K is totally imaginary and Galois over B=Q(241), and F is the fixed field of one complex conjugation ι1, which is not central: its conjugacy class has at least 215 elements, so θ≥θ∗=65535/131072 (Theorem 2.42) and almost every place of F is complex. Above a complex place of F the two coordinates of a norm-one element can have moduli eu and e−u, and the norm-one units, of rank c, move u along a lattice. Averaging over a fundamental domain of that lattice counts the resulting pairs of points exactly, and the class-number formula cancels the regulator of these units together with the capitulation kernel, the kernel of Cl(F)→Cl(K) (Proposition 5.2). Mixed signature therefore costs only a factor π per complex place of F, and a profile that couples the two coordinates at that place more than repays it.
A tower over Q(241). Over a totally real base field in which every prime at which the tower is ramified splits completely, the Golod–Shafarevich function has a local term for every prime of the base above such a prime, but only one constant term (in (12), the terms 2ψD and 4ψC2×C2, and the constant 1). The field B=Q(241) is the real quadratic field of smallest discriminant in which 2, 3 and 5 all split. Over B we prescribe a nonabelian dyadic local Galois group of order 32 at each of the two primes above 2, a tame local group C2×C2 at each of the four primes above 3 and 5, and conditions that make the Frobenius elements have order dividing 4 at the two primes above 29 and at the prime 7, which is inert in B. The relations are counted through their images in the Zassenhaus filtration, and one dyadic relation is implied by the others through Hilbert reciprocity. The resulting tower is infinite, it is ramified only above 2, 3 and 5, and its Galois group has 8 generators. The fields of Theorem 2.42 have root discriminant
λ=29/43615≈286.0,3615=3⋅5⋅241.
Sawin’s fields come from a pro-2 tower of unramified extensions of Q(3⋅5⋯43), so they are ramified over Q at 2 and at the thirteen odd primes up to 43, and their root discriminant is about 1.6⋅108. Theorem 2.42 states the family of fields.
The zeta value. Every field K of Theorem 2.42 contains the Kummer field E=B(V), of degree 512 over Q, where V is the group of {2,3,5}-units of B (the units of OB[1/30]) modulo squares. Its Dedekind zeta function is the product of ζB and 255 quadratic Hecke L-functions of B, which we evaluate rigorously with approximate functional equations (Sections 3 and 4). Together with refinements from the known types of primes in the fields of the tower (Definition 2.41), this gives the bound C=0.04871285 for d−1logLF(1) (Theorem 4.2). It replaces Louboutin’s bound [16], which Sawin uses, and the Euler products over a fixed multiquadratic subfield of [23].
Profiles and shells. Sawin counts points in a region bounded in the sup norm by packing, and [23] counts them by covolume in a Euclidean ball. We use a Gaussian profile at the real places of F; at the complex places, a Student–Bernstein profile, which couples the two coordinates (Definition 7.1); and shell profiles at the selected primes (Definition 6.5). We count points by Poisson summation (Sections 5–7). Table 2 compares Sawin’s example with the present construction.
Table 2. Sawin’s example and the present construction. The pairs (e,f) are absolute ramification indices and residue degrees of the selected primes of F; Sawin’s residue degrees are at most 2.
Why the exponent improves
Write ℓ=logλ,
J=r∈R∑erfrlog(kr+1),H=r∈R∑erkrlogr.
so that the sum in (1) is J−δH. Sawin’s criterion [32], Proposition 10 applies to Galois CM fields K of unbounded degree, with maximal totally real subfield F, whose relative root discriminant (∣ΔK∣/∣ΔF∣)1/d is λ, and counts points in a region bounded in the sup norm, with a parameter R>1. In his example, as for our fields, K/F is unramified at every finite place, so that ∣ΔK∣=∣ΔF∣2 and λ=rd(K) (Definition 2.40). Rearranged, his criterion gives arbitrarily large sets with u(U)≫∣U∣1+δ whenever (2) holds; inequality (3) is (1) with Aδ(θ) written out and regrouped:
In both, the three groups are the gain from equal-norm choices, the loss in selecting a common ideal class, and the count of points together with the shape of the region.
The fields. The dominant change is in ℓ. Sawin’s example has ℓ>18.9; here ℓ<5.657, so the term −21ℓ improves by more than 6.6 per unit of d. The price is paid in the first group: Sawin selects 22 primes with residue degree at most 2, and we select only the primes above 2,3,5,7 and 29, several with residue degree 4 or 8.
Ideal-class selection. Both criteria use the analytic class-number formula, which is exact in either signature; they differ in how they bound the zeta value. At θ=0 the constants agree, and the groups differ only in C against 1+logℓ. In Sawin’s criterion, 1+logℓ is the sum of 1−log2+logℓ, which is Louboutin’s bound [16], Corollary 3 for d−1logLF(1) read through the class-number formula, and log2, for the factor 2d allowed for units in [32], Lemma 6. The manuscript [23] removed the unit factor and sharpened Louboutin’s bound with Euler products over a fixed subfield. Here Proposition 5.2 accounts for the unit index [OF×:NK/FOK×] exactly, and the Kummer field gives C=0.04871285, whereas 1+logℓ would exceed 2.73 even at our value of λ. Mixed signature costs θlogπ in this group, since each place of F carries one factor π in the class-number formula.
Points and region. Besides the term −δH, which both criteria share, Sawin’s packing count of points contributes −2δlog(2R+e−H/2) to (2), which does not involve λ. Counting by covolume, as in [23] and here, contributes δℓ−δlog2 to (3), because a larger discriminant makes the ideal lattice sparser. The term 2log(1−R−1), which measures how much Sawin’s region overlaps its translates by elements of relative norm one, and the volume of that region become the functionals JR and JC of the profiles.
Signature. CM fields have θ=0. With the other data fixed, mixed signature adds θ(JC−2JR−logπ) to the left side of (3). Section 8 proves JC−2JR−logπ>0.6324 (Proposition 8.5(h)) and Lemma 8.10), so with θ≥θ∗ this term exceeds 0.316; at θ=0 the lower bound for the margin proved there would be negative (Remark 8.11).
Table 8 lists, rounded, the terms of the lower bound for the margin that Section 8 proves at δ=0.04273 and θ=θ∗; they sum to about 1.77⋅10−4 (Proposition 8.5). With the same complex-place profile, the real-place profile adjusted to δ as in Definition 8.1, and shell weights re-optimized in the same way, the corresponding lower bound at δ=0.0428 is negative (Remark 8.12).
Term
Value
∑rlogFδ,r/(erfr)
0.7214954821
−(1/2−δ)ℓ
−2.5863212799
−C
−0.0487128500
(1−δ)log2
0.6635290015
(1−θ∗)logπ
0.5723736765
(1−2θ∗)JR
−0.0000032159
θ∗JC∗
0.6778159157
M∗(θ∗)
0.0001767300
Table 8. The terms of M∗(θ∗), rounded to ten decimals.
Organization
Section 2 constructs the tower over B and proves Theorem 2.42, which supplies the fields. Section 3 describes the Kummer field and its quadratic L-functions, and Section 4 proves the upper bound d−1logLF(1)<C for the relative zeta value (Theorem 4.2). Sections 5–7 carry out the geometric construction: the class-number formula for the units of relative norm one and the geometric transfer to planar sets, the shell profiles at the selected primes, and the archimedean profiles. These sections apply to any quadratic extension K/F satisfying the conditions (G1)–(G3) of Section 5. Section 8 proves the final inequality at δ=0.04273 and completes the proof of Theorem 1.1.
Computations and data
The proof uses finite computations: linear algebra over F2 for the tower, exact rational arithmetic for the Golod–Shafarevich function, rigorous evaluations of L-functions in the ball arithmetic of Arb, now part of FLINT [15, 38], and interval arithmetic with directed rounding for the margin [40]. PARI/GP [41] is used for the arithmetic of number fields and for independent checks. Each finite fact is verified by a supplementary program, named in the text where the fact is used. These computations concern exact arithmetic and rigorous inequalities, not the statistical accuracy of floating-point approximations.
The programs and their data are in the directory papers/0.04273/certificates of the public repository [24]. Some routines, for the degree-two approximate functional equations, the archimedean profiles and the shell profiles, are imported unchanged from the supplementary archive of an earlier, unpublished note of the author, superseded by the present paper, in the directory papers/0.0418235/certificates of the same repository, together with a manifest of the SHA-256 hashes of its members. A replay program runs every step and records its results. The computations were run with PARI/GP 2.17.2, python-flint 0.9.0 (FLINT 3.6.0) and mpmath 1.3.0.
A separate Lean 4 development [22], built on Mathlib [2, 39], proves the planar conclusion with the slightly smaller exponent 1.0427 from one explicit numerical inequality for the Dedekind zeta function of E (Remark 4.40). That inequality has not been proved in Lean, so the formalization is conditional, and it is not used as a substitute for the arguments here. The formalization is registered in the Palomar registry of machine-checked Lean proofs as PALOMAR-2026-10-01-000018, version 1 [22].
The Tower over Q(241)
This section constructs the number fields of Theorem 2.42 as finite subextensions of an infinite pro-2 extension of B=Q(241) with prescribed local Galois groups. Subsection 2.1 describes B and its {2,3,5}-units. Subsection 2.2 recalls the local Galois groups involved and constructs the dyadic local field. Subsection 2.3 shows that the Galois group GS of the maximal pro-2 extension of B unramified outside 2,3,5 has a presentation with eight generators and seven local relators (Lemma 2.17). Subsection 2.4 imposes the local conditions and defines the quotient GB of GS (Definition 2.20). Subsection 2.5 computes the first three graded pieces of the Zassenhaus filtration (DnGB)n of GB, recalled below: the prescribed local groups inject into GB/D3GB, and the complex conjugation ι1 (Definition 2.12) has 215 conjugates in GB/D4GB (Lemma 2.26). Subsections 2.6 and 2.7 prove that GB is infinite, by a filtered Fox calculus and the Golod–Shafarevich method [10]: the Golod–Shafarevich function PB of (2.37) takes a negative value (Lemma 2.38). Subsection 2.8 proves Theorem 2.42.
We use the following conventions. Subgroups of profinite groups are closed, generation is topological, and the normal closure of a set is the smallest closed normal subgroup containing it. The commutator is [g,h]=g−1h−1gh. The generator rank and the relation rank of a pro-2 group G are dimH1(G,F2) and dimH2(G,F2). For an extension M/A of local or global fields, DM/A denotes its different and dM/A its relative discriminant. For a finite 2-group Δ with augmentation ideal J⊂F2[Δ], the Zassenhaus filtration is DnΔ={g:g−1∈Jn}; for a pro-2 group it is defined through finite quotients. Then [Dm,Dn]⊂Dm+n, g2∈D2n for g∈Dn, a surjection maps Dn onto Dn, and grΔ=⨁n≥1grnΔ, with graded pieces grnΔ=DnΔ/Dn+1Δ, is a restricted Lie algebra over F2 whose bracket is induced by commutators and whose restricted square v↦v[2] is induced by squaring [13, 31, 8]. By Lazard’s formula [8], Proposition 3.2, Dn=∏i2j≥nγi2j, where γi is the lower central series. In particular D1 is the whole group, D2 is the Frattini subgroup, and D3=D22[D2,D1]. For a free pro-2 group F of rank n, grF is the free restricted Lie algebra on F/D2F≅F2n[31, 8]. We realize it inside the free associative algebra F2⟨X1,…,Xn⟩, with bracket [a,b]=ab+ba and restricted square a[2]=a2; its graded pieces of degrees one, two and three have dimensions n, n(n+1)/2 and (n3−n)/3.
Definition 2.1. A filtered space is a finite-dimensional F2-vector space M with subspaces M=F0M⊇F1M⊇⋯ such that FnM=0 for large n. Its Hilbert polynomial is hM(t)=∑n≥0dim(FnM/Fn+1M)tn. For a finite 2-group Δ we filter F2[Δ] by the powers Jn of its augmentation ideal, so that hF2[Δ](t)=∑n≥0dim(Jn/Jn+1)tn.
Theorem 2.2 (Jennings [13]; see also [31]). Let Δ be a finite 2-group, and let g1,…,gm∈Δ be elements whose classes form a basis of grΔ consisting of homogeneous elements, the class of gi lying in grniΔ. Then the products (g1−1)a1⋯(gm−1)am with ai∈{0,1}, taken in this order, form a basis of F2[Δ], and those with ∑iaini≥n form a basis of Jn. Consequently
hF2[Δ](t)=n≥1∏(1+tn)dimgrnΔ.
The form with the factors in an arbitrary fixed order follows from Quillen’s description of grF2[Δ] as the restricted enveloping algebra of grΔ[31] and the restricted Poincaré–Birkhoff–Witt theorem.
The Base Field
Let B=Q(241), with real places v1(241↦+241) and v2(241↦−241). Put
Their norms NB/Q(αi) are 1,−1,−2,−2,−3,−3,−5,−5, and α2α3=2, α4α5=−3, α6α7=−5. For a nonzero ideal a of OB we write Na=∣OB/a∣ for its absolute norm.
Lemma 2.4. (a) OB=Z[(1+241)/2], the discriminant of B is 241, and B has class number one.
(b) The unit ε has norm NB/Q(ε)=−1, and the classes of −1 and ε form a basis of OB×/OB×2.
(c) In OB we have 2=p1p2, 3=q1q2, 5=r1r2 and 29=t1t2, where
with the sign + for j=1. These eight primes are distinct and of degree one. The prime 7 is inert, and we write t0=7OB, so that Nt0=49. The prime 241 ramifies. At each of the eight primes of degree one, above p∈{2,3,5,29}, the completion of B is Qp, and 241 maps to the square root of 241 in Zp given in Table 3.
place
241↦
α0
α1
α2
α3
α4
α5
α6
α7
p1
7mod32
−1
−5
−10
−5
1
5
−5
1
p2
25mod32
−1
5
5
10
5
1
1
−5
q1
2mod3
−1
1
−1
1
3
−1
−1
−1
q2
1mod3
−1
−1
−1
1
−1
3
−1
−1
r1
1mod5
1
2
2
1
1
2
10
2
r2
4mod5
1
2
1
2
2
1
2
10
t1
3mod29
1
1
2
1
1
2
2
2
t2
26mod29
1
1
1
2
2
1
2
2
v1
+241
−
+
−
−
−
+
−
+
v2
−241
−
−
+
+
+
−
+
−
Table 3. The image of 241 in Zp at each prime of degree one, given by its residue, and in R at the real places; the classes of α0,…,α7 in Qp×/Qp×2 at these primes, and their signs at the real places. At p=2 the classes are written as ±1,±5,±2,±10; at odd p as 1,u,p,up, where u=−1,2,2 is a nonsquare unit for p=3,5,29.
(d) Let S={p1,p2,q1,q2,r1,r2} and V=OB[1/30]×/OB[1/30]×2. The classes of α0,…,α7 form a basis of V, and V is the subgroup of B×/B×2 of classes with even valuation at every prime outside S. Moreover Pic(OB[1/30])=0.
(e) The residue field of t0 is F49, and αi is a square modulo t0 exactly for i=0,4,5,6,7.
Proof. (a) Since 241≡1(mod4), the ring of integers and the discriminant are as stated. The Minkowski bound is 241/2<8, so every ideal class contains an integral ideal of norm at most 7, which is a product of primes of norm at most 7. By (c), these are the six primes above 2,3,5, and they are principal.
(b) A direct computation gives 710110682−241⋅45742252=−1. By Dirichlet’s theorem OB×={±1}×ηZ for a fundamental unit η, and ε=±ηk. Since NB/Q(ε)=−1, the exponent k is odd, so −1 and ε generate OB× modulo squares. Neither −1 nor ±ε is a square, because B is real and NB/Q(±ε)=−1.
(c) We have 241≡1(mod8), 241≡1(mod3), 241≡1(mod5), 241≡9(mod29) and 241≡3(mod7), and 3 is not a square modulo 7. Hence 2,3,5,29 split, 7 is inert and 241 ramifies. The listed generators have norms ±2,±3,±5,29, so they generate primes of degree one, and the displayed products show that the two generators above each p generate the two distinct primes above p; for 29 the product of the two generators is 29. A prime of degree one above p is the preimage of pZp under an embedding B→Qp, which sends 241 to one of the two square roots of 241 in Zp; the completion there is Qp, and the prime contains the stated generator exactly for the square root in Table 3.
(d) Every S-unit u satisfies uOB=∏i≥2(αiOB)ai, so u∏i≥2αi−ai is a unit, and by (b) the classes of α0,…,α7 generate V. If ∏iαiai with ai∈{0,1} is a square, the valuations at S give ai=0 for i≥2, and then (b) gives a0=a1=0. If b∈B× has even valuation outside S, then bOB=a2∏p∈Spbp with a principal by (a), so b is an S-unit times a square. Finally Pic(OB[1/30]) is a quotient of Pic(OB)=0.
(e) By (c), Nt0=49. An element a∈OB prime to t0 is a square modulo t0 exactly when a24≡1, that is, when the reduction of NB/Q(a)≡a8 is a square in F7. By the norms listed after (4), this holds for αi exactly when i=0,4,5,6,7. □
The classes in Table 3 follow from the residues of 241 given there; at p=2 they need its image modulo 32. For example, at p1 we have 4574225≡1 and −71011068≡4 modulo 8, so ε≡4+7≡3≡−5(mod8). The classes modulo the inert prime t0 are given by Lemma 2.4(e).
Local Galois Groups
For a prime p let GQp be the Galois group of the maximal pro-2 extension of Qp, and for a place v of B let Gv be the Galois group of the maximal pro-2 extension of the completion Bv; for a real place, Gv=Gal(C/R). We write (⋅,⋅) for the Hilbert symbol of order two of a local field [19], Chapter III, Section 4. For odd p and p-adic units u,w, and for 2-adic units
[33], Chapter III, Theorem 1, [19], Chapter VIII, Section 5.
Definition 2.6. Let g1,…,gk generate a pro-2 group G, and let Fk be the free pro-2 group on k letters, mapped onto G by sending the letters to g1,…,gk. A defining relator of G with respect to g1,…,gk is an element of Fk whose normal closure is the kernel of Fk→G. If a defining relator r lies in D2Fk, its quadratic initial is its class in gr2Fk.
Lemma 2.7. (a) If ν is real, then Gν=⟨ι⟩≅C2 with the single defining relator ι2, and ι(b)=−b exactly when b<0.
(b) Let p be an odd prime and ϖ a uniformizer of Qp. Then GQp is generated by a generator τ of its inertia group and a Frobenius lift φ with φ(ϖ)=ϖ, with the single defining relator φτφ−1τ−p. For b∈Qp×, τ(b)=−b exactly when ordp(b) is odd; for a unit b, φ(b)=−b exactly when b is not a square modulo p. The relator has quadratic initial [τ,φ]+2p−1τ[2].
(c) Let E2=Q2(−1,2,5), and let x,y,z∈GQ2 multiply the triple (2,−1,5) by the signs (−,+,+), by (+,−,+) and by (+,−,−). Then x,y,z generate GQ2. If b≡(−1)a12a25a3 modulo squares, they multiply b by (−1)a2, (−1)a1 and (−1)a1+a3, that is, by the Hilbert symbols (5,b)2, (−1,b)2 and (−2,b)2. The group GQ2 has a single defining relator r with respect to x,y,z, and every such r has quadratic initial y[2]+[x,y]+[x,z].
Proof. (a) is clear.
(b) By Iwasawa’s theorem [26], Theorem 7.5.3, the Galois group of the maximal tamely ramified extension of Qp is generated by a Frobenius lift φ′ and a generator τ of tame inertia, with the single relation φ′τφ′−1=τp. Since p is odd, every 2-extension of Qp is tamely ramified, so GQp is the pro-2 group with the same presentation. The extension Qp(ϖ) is ramified, so τ moves ϖ; replacing φ′ by φ′τ if necessary gives φ. This substitution does not change the relator, since (φ′τ)τ(φ′τ)−1=φ′τφ′−1. The actions on square roots hold because Qp(b) is ramified exactly when ordp(b) is odd, and because φ acts on unramified extensions as the Frobenius. Modulo D3, the element φτφ−1τ−1 has class [φ,τ]=[τ,φ], and τ1−p has class 2p−1τ[2].
(c) We first show that x,y,z generate GQ2, and then read off the quadratic initial of r from the cup-product pairing on H1(GQ2,F2), which is given by the Hilbert symbol. By Kummer theory the maximal elementary abelian quotient of GQ2 is Gal(E2/Q2), since Q2×/Q2×2 has basis −1,2,5. The images of x,y,z form a basis of it, so x,y,z generate [26]. The action on b is read off from the actions on −1,2,5, and it agrees with the stated Hilbert symbols by (5). By [26], GQ2 is a Demuškin group of rank 3: it has a single defining relator r, and H2(GQ2,F2)≅F2. For b∈Q2× let χb be the character with g(b)=(−1)χb(g)b. The characters dual to x,y,z are χ2,χ−5,χ5. Inflation identifies H2(GQ2,F2) with H2 of the absolute Galois group of Q2[26], and there the invariant of χa∪χb is given by the Hilbert symbol (a,b)2[19]. By (5), the matrix of the indicator of (a,b)2=−1 on 2,−5,5 is
011110100.
The evaluation trr of the transgression at r is an isomorphism H2(GQ2,F2)→F2[26], and so is the invariant, so the two coincide. We now apply the relation–cup identity [26], see also [30]. Write g1,g2,g3 for x,y,z. The identity writes the relator as r=∏jgjqaj∏k<l[gk,gl]aklr′, with r′ in the third term of the descending q-central series, and states that trr of the cup product of the characters dual to gk and gl is −akl for k<l and −(2q)ak for k=l. Here q, the smallest elementary divisor of the abelianization of GQ2, is 2: by local class field theory this abelianization is the pro-2 completion of Q2×, which is Z22×Z/2. The third term of the 2-central series is (D2)2[D2,D1]=D3 by Lazard’s formula, the class of gj2 in gr2 is gj[2], and the signs disappear modulo 2, with (22)=1. Hence the off-diagonal entries of the matrix are the coefficients of [x,y],[x,z],[y,z] in the quadratic initial of r, and the diagonal entries, which are the values of the cup squares χ∪χ, are the coefficients of x[2],y[2],z[2]. For p=2 the cup square of a character is its Bockstein, so the diagonal coefficients are also given by [26]. □
We now construct the dyadic local field. Let U4 be the unramified extension of Q2 of degree four, with Frobenius Fr; let i=−1, k=Q2(i), α=1+2i, and
M8=Q2(i,5,α),M16=M8(2),L2=M16U4.
Since ααˉ=5 with αˉ=1−2i, the field M8 contains αˉ=5/α. Let σy and σz be the automorphisms of M8 given by
σy:i↦−i,5↦5,α↦5/α;σz:i↦−i,5↦−5,α↦5/α.
Lemma 2.8. (a) The extension M8/Q2 is Galois of degree 8, and its Galois group ⟨σy,σz⟩ is dihedral.
(b) The extension L2/Q2 is Galois of degree 32, with ramification index 8 and residue degree 4. There are unique x,y,z∈Gal(L2/Q2) such that x is trivial on M8 and U4 and negates 2; y restricts to σy on M8, fixes 2 and is trivial on U4; and z restricts to σz on M8, fixes 2 and restricts to Fr on U4. They act on E2⊂L2 as in Lemma 2.7(c), and the map sending the generators of
D=⟨x,y,z∣x2,y2,z4,[y,z]2,[x,y],[x,z],[[y,z],z]⟩
to x,y,z is an isomorphism D≅Gal(L2/Q2). The inertia group is ⟨x,y,[y,z]⟩≅C23.
(c) The graded pieces of the Zassenhaus filtration of D are gr1D=⟨xˉ,yˉ,zˉ⟩≅F23 and gr2D=⟨zˉ2,[yˉ,zˉ]⟩≅F22, and D3(D)=1, so the Hilbert polynomial of F2[D] (Definition 2.1) is hF2[D](t)=(1+t)3(1+t2)2.
(d) The different DL2/Q2 has valuation 18 in L2; equivalently, ord2(DL2/Q2)=18/8=9/4, where ord2 is the valuation of L2 normalized by ord2(2)=1.
Proof.Quadratic extensions of k. Let ordk be the normalized valuation of k. Put ϖ=1+i, a uniformizer of k, so that ordk(2)=2, and for b∈k× let c(b) be the exponent of the conductor of k(b)/k, which for a quadratic extension equals the exponent of its discriminant. Then c(5)=0, since Q2(5)/Q2 is unramified. Since α=1+ϖ2, the element ϖα=(α−1−ϖ)/ϖ is a root of the Eisenstein polynomial X2+ϖ2(1+ϖ)X+ϖ2; its derivative at ϖα is the sum of terms of valuations 5 and 2 in k(α), so c(α)=2. Next 2α=ϖ2α′ with α′=2−i, and α′−1 is a root of X2+2X+1−α′, which is Eisenstein since ordk(1−α′)=1; its derivative 2α′ has valuation 4, so c(2α)=4. Similarly 2=−iϖ2, −i≡i modulo squares, and i−1 is a root of X2+2X+1−i, so c(2)=4. Finally c(5b)=c(b) when b∈/k×2∪5k×2: the field k(5,b) is unramified over both k(b) and k(5b), so the two ways of computing its discriminant over k give 2c(b)=2c(5b).
(a) Since k(5)/k is unramified and k(α)/k is ramified, α is not a square in k(5), so [M8:Q2]=8. The field M8 contains the conjugate αˉ of α, so it is Galois over Q2, and an automorphism is determined by the images ±i, ±5 and a square root of the image of α; in particular σy and σz exist. Now σy2=1, σz2 fixes i,5 and negates α, σz has order four, and σyσzσy=σz−1. So Gal(M8/Q2)=⟨σy,σz⟩ is dihedral of order eight, with center ⟨σz2⟩ and maximal abelian subextension Q2(i,5).
(b) The quadratic subfields of M8 are Q2(b) for b=−1,5,−5, so 2∈/M8, and M16 has degree 16 and group Gal(M8/Q2)×C2. This group has elementary abelian abelianization, so the maximal unramified subextension of M16, which is cyclic over Q2 and contains Q2(5), is Q2(5). Hence M16 has ramification index 8, M16∩U4=Q2(5), and [L2:Q2]=16⋅4/2=32. Since L2 contains U4 and has ramification index at least 8, its residue degree is 4 and its ramification index is 8. Restriction identifies Gal(L2/Q2) with the set of triples (σ,ϵ,μ), where σ∈Gal(M8/Q2), ϵ=±1 is the sign by which the element multiplies 2, and μ∈Gal(U4/Q2) agrees with σ on 5; both sets have 32 elements. Since Fr negates 5, the triples
exist and are the unique elements described in the statement. They act on E2 as required. The elements x and w are central involutions, y is an involution, and z has order four. Hence the relations of D hold. Conversely, in D these relations make w=[y,z] central: wy=w−1=w, and [w,z]=1 is imposed. Since zy=yzw, every element of D is xa1ya2za3wa4 with a1,a2,a4∈F2 and a3∈Z/4, and
These 32 normal forms have distinct images in Gal(L2/Q2): the third component determines a3, the second a1, and σya2σza3+2a4 determines a2 and a4. So D≅Gal(L2/Q2). The inertia group consists of the elements trivial on U4, namely ⟨x,y,w⟩.
(c) The Frattini subgroup of D is ⟨z2,w⟩, which is central of exponent two, and D has exponent four. So Lazard’s formula gives D3(D)=[[D,D],D][D,D]2D4=1, and the Hilbert polynomial follows from Theorem 2.2.
(d) The extension M16/k is abelian with group C23, generated by 5,α,2. We apply the conductor–discriminant formula for abelian extensions [19], Chapter V, Theorem 3.27 to the extension Q(i)(5,1+2i,2) of Q(i), whose completion at the unique prime above 1+i is M16/k. Taking (1+i)-adic valuations gives
ordk(dM16/k)=b∑c(b)=0+0+2+2+4+4+4+4=20,
the sum over the classes b=1,5,α,5α,2,10,2α,10α. Hence
Since M16 has residue degree two over Q2, its different has valuation 36/2=18. The extension L2/M16 is unramified, so the different of L2/Q2 also has valuation 18, and 18/8=9/4. □
Remark 2.9. The field E2 is the maximal elementary abelian subextension of L2, so Gal(L2/E2) is the Frattini subgroup ⟨z2,[y,z]⟩, which is central of exponent two. Hence any triple in Gal(L2/Q2) that acts on E2 as x,y,z do differs from x,y,z by central elements of order at most two; it satisfies the relations of D and generates, so it also defines an isomorphism D≅Gal(L2/Q2). The description of the inertia group, however, refers to the triple of Lemma 2.8(b).
The Global Group and Its Complete Presentation
Let BS be the maximal pro-2 extension of B, in a fixed algebraic closure, that is unramified at every finite prime outside S; no condition is imposed at v1,v2. Let GS=Gal(BS/B), and let E=B(α0,…,α7)=B(V) be the Kummer field.
Lemma 2.10.The field E is the maximal elementary abelian extension of B contained in BS, and [E:B]=256. Hence GS/D2GS=Gal(E/B).
Proof. Every prime above 2 lies in S, so a quadratic extension B(b) lies in BS exactly when b has even valuation outside S. By Lemma 2.4(d), E is the maximal elementary abelian extension of B in BS, and [E:B]=256. Since D2GS is the Frattini subgroup of GS, its fixed field is this maximal elementary abelian extension. □
Definition 2.11. Let M⊂BS be a Galois extension of B containing E, for example M=BS. The elementary image of g∈Gal(M/B) is the vector gˉ∈F28 with g(αi)=(−1)gˉiαi for 0≤i≤7. For a∈V, the Kummer characterχa is defined by g(a)=(−1)χa(g)a; its values lie in F2.
By Lemmas 2.4(d) and 2.10, g↦gˉ identifies GS/D2GS=Gal(E/B) with F28, and a↦χa identifies V with the dual of F28, with χαi(g)=gˉi. We write vectors in F28 as strings c0c1⋯c7 of their coordinates.
For a place ν of B, a choice of a place of BS above ν gives a homomorphism Gν→GS, the local map at ν; another choice changes it by an inner automorphism of GS. For a finite place ν∈/S, the local map kills inertia, and its image is generated by a Frobenius element Frobν. Let Σ=S∪{v1,v2}.
Definition 2.12. Fix a place of BS above each place of B, and use it to define the local maps. For ν∈Σ we fix generators of Gν as follows, and use the same names for their images in GS: ιj at vj (Lemma 2.7(a)), so that ι1,ι2 are complex conjugations; τν,φν at ν∈{q1,q2,r1,r2}, with the generator αi of ν as uniformizer ϖ in Lemma 2.7(b); and xj,yj,zj at pj, the images of one fixed triple x,y,z∈GQ2=Gpj lifting the elements x,y,z of Lemma 2.8(b). We also write Frobtk for k=0,1,2. These elements are the local generators: those at ν∈Σ are the generators just fixed, and the local generator at tk is Frobtk.
The triple x,y,z satisfies the hypothesis of Lemma 2.7(c).
Table 4 lists the elementary images. They follow from Lemma 2.7, Table 3 and Lemma 2.4(e). For example, at the places above 3 and 5 the vector of τν marks the odd valuations, and that of φν marks the units that are not squares modulo ν, with a zero at the generator of ν. The columns headed “action” express the same actions as Hilbert symbols, that is, as local Artin symbols [19], Chapter III, Remark 4.5]. The supplementary program kummer241.gp checks the norms of the αi, the generators of the primes in Lemma 2.4(c) and the residues of 241 in Table 3, and recomputes Table 4 from these Hilbert symbols.
element
image
action
element
image
action
ι1
10111010
sign at v1
ι2
11000101
sign at v2
x1
00100000
(5,b)p1
x2
00010000
(5,b)p2
y1
11110010
(−1,b)p1
y2
10000001
(−1,b)p2
z1
10000100
(−2,b)p1
z2
11111000
(−2,b)p2
τq1
00001000
(−1,b)q1
φq1
10100111
(−α4,b)q1
τq2
00000100
(−1,b)q2
φq2
11101011
(−α5,b)q2
τr1
00000010
(2,b)r1
φr1
01100101
(−α6,b)r1
τr2
00000001
(2,b)r2
φr2
01011010
(−α7,b)r2
Frobt1
00100111
(29,b)t1
Frobt2
00011011
(29,b)t2
Frobt0
01110000
(7,b)t0
Table 4. Elementary images in F28 of the local generators. Entry i is 1 exactly when the element negates αi. The third and sixth columns give the action on b for b in the completion.
Fix a minimal presentation 1→R→F→GS→1, where F is a free pro-2 group of rank 8. Minimality means R⊂D2F, so F/D2F=F28 with the coordinates above. Put L=grF, so that
dimL1=8, dimL2=36 and dimL3=168. For v,w∈L1 we have (v+w)[2]=v[2]+w[2]+[v,w], and the square of an element of F with image v has class v[2]. For each local generator g above choose a lift g^∈F, and define the local relators
where p∈{3,5} is the rational prime below ν, and r is a defining relator of GQp with respect to the fixed triple x,y,z, the same for j=1,2. They lie in R, because each local relation holds in Gν. Their classes in L2 are obtained by substituting elementary images into the quadratic initials of Lemma 2.7:
The results below hold for every minimal presentation and every choice of lifts.
The classes (7) have an arithmetic meaning. As in the conventions, we realize L inside the free associative algebra F2⟨X0,…,X7⟩, where X0,…,X7 is the standard basis of L1=F28. The elements Xi[2]=Xi2 and [Xi,Xj]=XiXj+XjXi(i<j) form a basis of L2, so there is a linear isomorphism β from L2 onto the space of symmetric bilinear forms on V with
It is defined by these formulas for basis vectors v,w, which send the basis of L2 to a basis of the symmetric bilinear forms. The formulas then hold for all v,w, because the bracket is bilinear, (v+w)[2]=v[2]+w[2]+[v,w], and squaring is the identity on F2. Equivalently, if ξ∈L2 is written in the free algebra as ∑i,jnijXiXj, then β(ξ)(αi,αj)=nij.
Definition 2.15. For ν∈Σ, the local Hilbert formHν is the symmetric bilinear form on V with Hν(a,b)=1 if (a,b)ν=−1 and Hν(a,b)=0 otherwise.
Lemma 2.16.For every ν∈Σ, β maps the class of ρν to Hν. The eight classes (7), one for each ν∈Σ, sum to zero, and the seven classes with ν=p1 are linearly independent.
Proof. For a local generator g at ν, χa(gˉ) is the value at g of the local Kummer character of the image of a in Bν. Hence β of the class of ρν is the pullback to V of a symmetric bilinear form on Bν×/Bν×2, as is Hν, and it suffices to compare the two on a basis of Bν×/Bν×2. At a real place the only class is −1, and both forms take the value 1 on (−1,−1). At ν above p∈{3,5}, take the basis ϖ,u with ϖ the generator of ν and u a nonsquare unit. By Lemma 2.7(b), χϖ(τ)=χu(φ)=1 and χϖ(φ)=χu(τ)=0, so the form of [τ,φ]+2p−1τ[2] takes the values 2p−1,1,0 on (ϖ,ϖ), (ϖ,u), (u,u); by (5) these are the indicators of (ϖ,ϖ)p=(−1)(p−1)/2, (ϖ,u)p=−1 and (u,u)p=1. At pj the characters χ2,χ−5,χ5 are dual to x,y,z, as in the proof of Lemma 2.7(c). So y[2] contributes the entry at (−5,−5), [x,y] the entries at (2,−5) and (−5,2), and [x,z] those at (2,5) and (5,2): the form of y[2]+[x,y]+[x,z] on the basis 2,−5,5 is the matrix of the Hilbert symbol displayed in that proof, from which this initial was obtained.
For a,b∈V and a place ν′∈/Σ, the place ν′ is finite and not above 2, and a,b are units at ν′, so (a,b)ν′=1[19]. By the product formula for the Hilbert symbol [19], ∑ν∈ΣHν=0, so the eight classes sum to zero.
have Hilbert symbol −1 exactly at {p1,v1}, {p1,v2}, {p2,v2}, {q1,v1}, {q2,v2}, {r1,v1} and {r2,v2}; the supplementary program cup241.gp recomputes these symbols. Evaluating a relation ∑ν=p1nνHν=0, with nν∈F2, at the first two pairs gives nv1=nv2=0, and then at the other five pairs gives np2=nq1=nq2=nr1=nr2=0. □
Lemma 2.17.The group GS has generator rank 8 and relation rank 7. The seven relators ρν with ν∈Σ∖{p1} form a minimal system of defining relations: their normal closure in F is R. In particular the dyadic local relator ρp1 lies in the normal closure of the other seven.
Proof. The generator rank is dimH1(GS,F2)=8, by Lemma 2.10 and [26].
Next we bound H2. Let X=SpecOB[1/30]. Since 2 is invertible on X, the sheaf μ2 is the constant sheaf F2. The finite subextensions of BS/B correspond to the connected finite étale Galois 2-covers Y→X, and the Hochschild–Serre spectral sequence of this system of covers [17] gives an exact sequence
A class in H1(Y,F2) is an étale double cover of Y. Its Galois closure over X is a finite étale Galois cover whose group embeds in the wreath product of Gal(Y/X) with C2, hence a 2-cover in this system, over which the class vanishes. So the direct limit is zero, and H2(GS,F2) injects into H2(X,μ2). The Kummer sequence and Pic(X)=0 identify H2(X,μ2) with Br(X)[2], where Br(X)=H2(X,Gm). By [29], Br(X) is the kernel of the sum of the local invariants ⨁ν∈ΣBr(Bν)→Q/Z. Each of the eight groups Br(Bν)[2] has order two, and the sum of the invariants maps their direct sum onto 21Z/Z, so dimBr(X)[2]=7. We use only the resulting bound dimH2(GS,F2)≤dimBr(X)[2]≤7.
By Lemma 2.16, the classes in L2 of the seven relators ρν, ν=p1, are linearly independent. Since R⊂D2F, we have R2[R,F]⊂D3F, so the classes of these seven relators in R/R2[R,F] are independent. On the other hand dimR/R2[R,F]=dimH2(GS,F2)≤7 by the five-term exact sequence of the presentation [26, 30]. Hence the seven classes form a basis, the relation rank is 7, and by the pro-2 Nakayama lemma [26] the seven relators generate R as a normal subgroup. Finally ρp1∈R. □
Corollary 2.18.Any seven of the eight local relators ρν, ν∈Σ, generate R as a normal subgroup.
Proof. By Lemma 2.16, the only linear relation among the eight classes (7) is that their sum vanishes, and this relation is Hilbert reciprocity. Hence any seven of the eight classes are linearly independent, and the proof of Lemma 2.17 applies to any seven of the eight local relators. □
We omit the dyadic relator at p1.
Remark 2.19. By Lemma 2.17, the bound dimH2(GS,F2)≤dimBr(X)[2]=7 is attained, so H2(GS,F2)→Br(X)[2] is an isomorphism. This can also be seen directly: the cup product χa∪χb maps to the class of the quaternion algebra (a,b), whose invariant at ν is given by (a,b)ν[19], Chapter III, Remark 4.7, and the seven pairs in the proof of Lemma 2.16 give seven classes with independent invariant vectors. The supplementary program cup241.gp computes the invariant vectors at the eight places of Σ of all 36 algebras (αi,αj).
The Prescribed Local Groups
We now cut GS down to prescribed local groups. At the dyadic places we use the restriction Gpj=GQ2→Gal(L2/Q2)=D of Lemma 2.8. At a place ν above p∈{3,5}, the maximal elementary abelian quotient of Gν is Gal(Qp(u,p)/Qp)≅C2×C2 for a nonsquare unit u; this quotient has ramification index and residue degree two. At t1,t2, whose completions are Q29, and at t0, whose residue field is F49, we bound the order of the Frobenius by four.
Definition 2.20. Let N be the normal closure in GS of the images of ker(Gpj→D) for j=1,2, of the Frattini subgroups D2Gν for ν∈{q1,q2,r1,r2}, and of Frobtk4 for k=0,1,2. Put GB=GS/N and GB=GB/D4GB, and let BG⊂BS be the fixed field of N, so that Gal(BG/B)=GB. We call the extension BG/B the tower.
This does not depend on the choice of places of BS above each ν: another choice changes each local map by an inner automorphism of GS, which replaces the images above by conjugates and does not change their normal closure N.
Lemma 2.21. (a) The image of Gpj in GB is a quotient of D, the image of Gν for ν above 3 or 5 is a quotient of C2×C2 in which inertia maps to the image of τν, and Frobtk has order dividing four in GB.
(b) N⊆D2GS. Hence GB/D2GB=Gal(E/B), and E⊆BG is the fixed field of D2GB.
Proof. (a) By Definition 2.20, the maps Gpj→GB and Gν→GB kill ker(Gpj→D) and D2Gν, and Frobtk4∈N. For ν above 3 or 5, Gν/D2Gν≅C2×C2, and the inertia group of Gν is generated by τν (Lemma 2.7(b)).
(b) Since D and GQ2 both have generator rank three, ker(Gpj→D) lies in the Frattini subgroup of Gpj. The images of the Frattini subgroups of the local groups, and the fourth powers Frobtk4, lie in D2GS. Hence N⊆D2GS, and GB/D2GB=GS/D2GS=Gal(E/B) by Lemma 2.10. □
By Lemma 2.26(b) below, the quotients in Lemma 2.21(a) are D and C2×C2, and Frobtk has order exactly four in GB.
Lemma 2.22.The kernel NB of F→GS→GB is the normal closure of the following thirty elements:
(i) the seven relators ρν, ν∈Σ∖{p1};
(ii) for j=1,2: x^j2, [x^j,y^j], [x^j,z^j], [[y^j,z^j],z^j], z^j4, [y^j,z^j]2;
(iii) for ν∈{q1,q2,r1,r2}: τ^ν2 and ϕ^ν2;
(iv) for k=0,1,2: F^k4, where F^k is the chosen lift of Frobtk.
Moreover NB⊂D2F, and ρp1∈NB.
Proof. By Lemma 2.17, it suffices to show that the images in GS of the elements (ii)–(iv) generate N as a normal subgroup. Let F3 be free on x,y,z, and let ND be the normal closure of the seven relators of D in Lemma 2.8, so that F3/ND=D. The relator r maps to 1 in D, so r∈ND, and ker(GQ2→D) is the image of ND. Since ND⊂D2F3, we have ND2[ND,F3]⊂D3F3. The seven relators of D span ND/ND2[ND,F3][26], Corollary 3.9.3; write the class of r as a combination of them. The classes in gr2F3 of x2,y2,[x,y],[x,z] are independent, the relators z4,[y,z]2 and [[y,z],z] lie in D3F3, and r has class y[2]+[x,y]+[x,z]; so y2 has coefficient one. Replacing y2 by r therefore still gives a spanning set, and by [26], Corollary 3.9.3 the elements r,x2,[x,y],[x,z],[[y,z],z],z4,[y,z]2 generate ND as a normal subgroup. Since r=1 in GQ2, the other six generate ker(GQ2→D).
For ν above p∈{3,5}, the quotient of Gν by the normal closure of τ2,ϕ2 is generated by two involutions which commute, because ϕτϕ−1=τp=τ there; so that normal closure is D2Gν. Finally N⊆D2GS (Lemma 2.21(b)) and R⊂D2F give NB⊂D2F, and ρp1∈R⊂NB. □
The Quotient GB/D4GB.
Definition 2.23. For n≥1, the relation spaceRn⊆Ln of GB in degree n is the image of NB∩DnF.
Since NB⊂D2F and DnGB is the image of DnF,
gr1GB=F28,grnGB≅Ln/Rn.
Consider the following 22 elements of L2, written with the vectors of Table 4:
They are the classes of the eight local relators (7) and of the fourteen quadratic elements x^j2, [x^j,y^j], [x^j,z^j], τ^ν2, ϕ^ν2 of Lemma 2.22.
Lemma 2.25.The space R2 is spanned by the 22 elements (8), and dimR2=21; their only linear relation is that the eight classes of local relators sum to zero. Hence dimgr2GB=15. Consequently, if g∈GB has nonzero elementary image v, then g has order at least four in GB/D3GB exactly when v[2]∈/R2; otherwise g has order two there.
Proof. Conjugates of an element of D2F have the same class modulo D3F, classes of products add, and L2 is finite. Hence, by Lemma 2.22, R2 is spanned by the classes of the thirty generators of NB. Those in (i) have the classes (7), and the quadratic elements of (ii) and (iii) have the classes xˉj[2], [xˉj,yˉj], [xˉj,zˉj], τˉν[2], ϕˉν[2]. The remaining generators, [[y^j,z^j],z^j], z^j4, [y^j,z^j]2 and the fourth powers in (iv), lie in D3F. The class of ρp1 also lies in R2, since ρp1∈NB. The supplementary program lie241.py computes the rank of the 22 elements in the 36-dimensional space L2: it is 21. The sum of the eight local classes vanishes by Lemma 2.16, that is, by Hilbert reciprocity, so this is the only relation; none of the fourteen quadratic elements of (ii) and (iii) is involved in it. For the last statement, g2∈D2GB has class v[2] in gr2GB=L2/R2, so g2∈D3GB exactly when v[2]∈R2. □
Lemma 2.26. (a) dimgr3GB=26, so ∣GB∣=28+15+26=249.
(b) For j=1,2, the local map induces an injection D→GB/D3GB, and grnD→grnGB is injective for n=1,2. For ν above 3 or 5, the image of Gν in GB is ⟨τν⟩×⟨ϕν⟩≅C2×C2, and it injects into GB/D2GB. For k=0,1,2, Frobtk has order exactly four in GB and in GB/D3GB. For j=1,2, ιj has order two in GB/D2GB.
(c) The conjugacy class of ι1 in GB has 215 elements.
(d) We have ιˉ1=ιˉ2. For each ν∈{p1,p2,q1,q2,r1,r2,t0,t1,t2}, neither ιˉ1 nor ιˉ2 lies in the span of the elementary images of the local generators at ν.
Proof. (a) Work in F/D4F, in which D2F/D4F is elementary abelian and D3F/D4F is central. Let ρ1,…,ρ21 be the generators of NB listed in (i), the first three elements of (ii) for each j, and (iii) of Lemma 2.22. They lie in D2F, and their classes in L2 are the elements of (8) other than the class of ρp1, which are independent by Lemma 2.25. The generators ωj=[[y^j,z^j],z^j] lie in D3F, and the remaining generators of NB lie in D4F. For ρ∈D2F and g∈F we have g−1ρg=ρ[ρ,g], and modulo D4F the commutator [ρ,g] is central and depends only, and linearly, on the classes of ρ in L2 and of g in L1. Hence the image of NB in F/D4F is the subgroup generated by the ρi, the [ρi,f] for f in a basis of F, and the ωj. An element ∏iρiai times an element of D3F lies in D3F only if all ai=0. Therefore R3 is spanned by the brackets [ρˉi,ξ], for ξ in a basis of L1, and by [[yˉj,zˉj],zˉj]. The supplementary program lie241c.py computes these elements in the free associative algebra, where [a,ξ]=aξ+ξa: they span a space of dimension 142 in the 168-dimensional space L3. So dimgr3GB=26.
(b) The local map Gpj→GB factors through D by Definition 2.20, and the induced map D→GB preserves the Zassenhaus filtrations. It is injective on gr1 because xˉj,yˉj,zˉj are independent, and on gr2 because zˉj[2] and [yˉj,zˉj], the images of z2 and [y,z], are independent modulo R2; the program lie241.py checks both facts. If 1=g∈D, let n≤2 be maximal with g∈Dn(D); then the image of g is nonzero in grnGB, so it is not in D3GB. For ν above 3 or 5, the vectors τˉν and φˉν are independent, and the image of Gν is a quotient of C2×C2. For the Frobenius elements, lie241.py checks that Frobtk[2]∈/R2, so Lemma 2.25 gives order at least four in GB/D3GB, while Definition 2.20 gives order at most four in GB. Finally ιˉj=0.
(c) Write Gˉ=GB and ι=ι1. The map L1→gr2Gˉ, v↦[ιˉ,v], has rank 7 by lie241.py, so its kernel is the line through ιˉ. The map gr2Gˉ→gr3Gˉ, u↦[u,ιˉ], has rank 8 by lie241c.py. Since D2Gˉ is abelian and D3Gˉ is central in Gˉ, the map h↦[h,ι] is a homomorphism D2Gˉ→D3Gˉ; it factors through gr2Gˉ and equals the second map. Its kernel CGˉ(ι)∩D2Gˉ therefore has order 215+26−8=233. If g∈CGˉ(ι), then [ιˉ,gˉ]=0 in gr2Gˉ, so gˉ∈{0,ιˉ}; conversely ι∈CGˉ(ι). Hence ∣CGˉ(ι)∣=234, and the class of ι has 249−34=215 elements.
(d) This is a finite check on Table 4, carried out by lie241.py. □
Filtered Fox Calculus
This subsection and the next prove that GB is infinite (Proposition 2.39) by the Golod–Shafarevich method. Suppose that GB is finite, and let Λ=F2[GB]. The rows (Definition 2.30) of the relators of GB in a presentation with eight generators (Lemma 2.22) span a Λ-module M whose Hilbert polynomial is 1−(1−8t)hΛ (Lemma 2.32). The module M is also the sum of the modules spanned by the rows of the relations of the eleven local groups Γ, each isomorphic to C2, C2×C2, C4 or D; for 0<t<1 the Hilbert polynomial of each of these modules is at most ψΓ(t)hΛ(t), with ψΓ as in Definition 2.33 (Lemma 2.34). The row of the dyadic relator ρp1 of (6) lies both in the module of p1 and in the sum of the other ten, and the module it generates has Hilbert polynomial sDhΛ, with sD as in Lemma 2.35. Comparing these bounds gives 1≤hΛ(t)PB(t) for 0<t<1, with PB as in (12); this fails at t=34/117 (Lemma 2.38).
Let Δ be a finite 2-group, Λ=F2[Δ], and J its augmentation ideal. We use the filtered spaces and Hilbert polynomials of Definition 2.1; in particular FnΛ=Jn. For m≥1, Λm(−1) denotes Λm with Fn=(Jn−1)m for n≥1, so that hΛm(−1)=mthΛ. Subspaces carry the induced filtration unless stated otherwise. For 0<t<1,
hM(t)=(1−t)n≥1∑dim(M/FnM)tn−1.(9)
Lemma 2.28.Let 0<t<1.
(h1) If f:M→M′ is surjective with f(FnM)⊆FnM′, then hM′(t)≤hM(t).
(h2) If U1⊆U2 are subspaces of the same filtered space, then hU1(t)≤hU2(t).
(h3) If f is surjective and strict, f(FnM)=FnM′, then hM=hkerf+hM′.
(h4) For subspaces U1,U2 of a filtered space,
hU1+U2(t)≤hU1(t)+hU2(t)−hU1∩U2(t).(10)
Proof. Statements (h1)–(h3) follow from the definitions and (9). For (h4), apply (h3) to U1⊕U2→U1+U2, (u1,u2)↦u1+u2, whose kernel is U1∩U2, with the image filtration FnU1+FnU2 on U1+U2; this filtration is contained in the induced one, and (h1) applies to the identity map. □
Definition 2.30. Let F be a free pro-2 group with basis f1,…,fn and π:F→Δ a surjection. Since Λn⋊Δ, with Δ acting by left multiplication, is a finite 2-group, there is a unique homomorphism F→Λn⋊Δ sending fi to (ei,π(fi)), where e1,…,en is the standard basis of Λn. We write it as w↦(∂w,π(w)) and call ∂w=(∂1w,…,∂nw)∈Λn the row of w.
For w in the discrete free group on the fi, the row ∂w∈Λn is the image of the vector of free derivatives of w[9].
Lemma 2.31. (F1) ∂(ww′)=∂w+π(w)∂w′. (F2) ∑i∂iw(π(fi)−1)=π(w)−1. (F3) If π(ρ)=1, then ∂(uρu−1)=π(u)∂ρ; and ∂ restricts to a continuous homomorphism on kerπ. (F4) If F′ is free on f1′,…,fm′, u1,…,um∈F, and ∂′ denotes the rows for the map F′→Δ, fl′↦π(ul), then ∂(W(u1,…,um))=∑l∂l′W⋅∂ul for W∈F′.
Proof. (F1) is the multiplication of the semidirect product. For (F2), the pairs (a,g) with ∑ai(π(fi)−1)=g−1 form a subgroup containing the images of the fi. (F3) follows from (F1). For (F4), both sides are the first components of homomorphisms F′→Λn⋊Δ that agree on the basis. □
Lemma 2.32.Let M={a∈Λn:∑iai(π(fi)−1)=0}, with the filtration induced from Λn(−1). Then hM=1−(1−nt)hΛ and M=∂(kerπ). If ρ lies in the normal closure of a set Y⊂kerπ, then ∂ρ lies in the Λ-span of ∂Y. In particular, M is the Λ-span of the rows of any set of normal generators of kerπ.
Proof. Since the π(fi) generate Δ, we have I=∑iΛ(π(fi)−1)⊂∑iF2(π(fi)−1)+I2, hence Im=∑iIm−1(π(fi)−1) for all m≥1, by induction and because I is nilpotent. So Λn(−1)→I, a↦∑ai(π(fi)−1), is strict and surjective with kernel M, and (h3) gives hM=nthΛ−(hΛ−1).
By (F2), ∂(kerπ)⊂M. Conversely, let F∘⊂F be the dense subgroup generated by the fi, which is a free group on them. Then π(F∘)=Δ, and the kernel of F2[F∘]→Λ is F2[F∘](F∘∩kerπ−1). Lift a∈M to a∈F2[F∘]n. Then ∑iai(fi−1)=∑lζl(nl−1) with nl∈F∘∩kerπ and ζl∈F2[F∘]. By Fox’s fundamental formula nl−1=∑i(∂nl/∂fi)(fi−1), where the ∂/∂fi are the free derivatives [9], and the coefficients of the fi−1 are unique, because ∂/∂fj maps ∑iζi(fi−1) to ζj′. So ai=∑lζl∂nl/∂fi. On F∘ the rows are the images of the vectors of free derivatives, so a=∑lζˉl∂nl lies in Λ∂(kerπ), which equals ∂(kerπ) by (F3). For the remaining statements, let MY be the Λ-span of ∂Y. By (F3), the elements n∈kerπ with ∂n∈MY form a closed normal subgroup of F containing Y. □
The next lemma bounds, in terms of the function ψΓ below, the Hilbert polynomial of the Λ-module spanned by the rows of all relations of one local group Γ. The groups Γ below are C2, C2×C2, C4 and D, with standard generators i; τ,φ; a generator; and x,y,z. Let k be the number of standard generators, so k=1,2,1,3. The Hilbert polynomials hF2[Γ] are 1+t, (1+t)2, (1+t)(1+t2) and hF2[D] (Theorem 2.2), and D3(Γ)=1 in each case.
Definition 2.33. For each of these groups Γ, put
ψΓ(t)=kt−1+hF2[Γ](t)−1.
Lemma 2.34.Let Γ≤Δ be a subgroup isomorphic to one of these groups, with g1,…,gk∈Γ corresponding to its standard generators, and suppose that grnΓ→grnΔ is injective for n=1,2. Let ΛΓ=F2[Γ]⊂Λ, with augmentation ideal IΓ, and let MΓ⊂ΛΓk(−1) be the kernel of b↦∑lbl(gl−1).
(a) There are elements ηj∈Λ and integers wj≥0 such that Λ=⨁jηjΛΓ and Im=⨁jηjIΓm−wj for every m, where IΓm′=ΛΓ for m′≤0; moreover ∑jtwj=hΛ/hF2[Γ].
(b) The Λ-module ΛMΓ⊂Λk(−1) has hΛMΓ=ψΓhΛ.
(c) Let F and π be as above, let u1,…,uk∈F with π(ul)=gl, and let Fk be free on k letters, mapping to Γ by the standard generators. Let U⊂Λn(−1) be the Λ-span of the rows ∂W(u1,…,uk) for all W in the kernel of Fk→Γ. Then hU(t)≤ψΓ(t)hΛ(t) for 0<t<1.
Proof. (a) Since D3(Γ)=1 and grΓ→grΔ is injective, Γ∩DmΔ=Dm(Γ) for all m. Choose elements of Γ whose classes form a homogeneous basis of grΓ, and extend them by elements of Δ to a family whose classes form a homogeneous basis of grΔ. By Theorem 2.2, the ordered products of the elements g−1 of this family, each used at most once and with the elements of Γ last, form a basis of Λ, and those of weight at least m form a basis of Im; here the weight of a product is the sum of the degrees of the classes of its factors. The same holds for ΛΓ and the elements of Γ alone. Let the ηj be the ordered products of the other elements, and wj the weight of ηj. The empty product ηj=1 has weight 0, so Im∩ΛΓ=IΓm: the filtration of ΛΓ induced from Λ is its augmentation filtration, with Hilbert polynomial hF2[Γ]. Comparing Hilbert polynomials gives the last assertion.
(b) As in the proof of Lemma 2.32, the map ΛΓk(−1)→IΓ is strict and surjective, so hMΓ=kthF2[Γ]−(hF2[Γ]−1). By (a), Λk(−1)=⨁jηjΛΓk(−1) with Fm=⨁jηjFm−wj, and ΛMΓ=⨁jηjΛMΓ. Hence hΛMΓ=(hΛ/hF2[Γ])hMΓ=ψΓhΛ.
(c) For W in the kernel, the local row ∂′W∈ΛΓk lies in MΓ by (F2), and by (F4) ∂W(u)=∑l∂l′W∂ul. Thus U is contained in the image of ΛMΓ under the map Tu:Λk(−1)→Λn(−1), a↦∑lal∂ul, which preserves filtrations because ∂ul∈Λn=F0. By (h2), (h1) and (b), hU≤hTu(ΛMΓ)≤hΛMΓ=ψΓhΛ. □
In the proof of Proposition 2.39, the row of the dyadic relator ρp1 lies both in the module spanned by the rows of the relations at p1 and in the module spanned by the rows of the relations at the other ten places. The next lemma computes the Hilbert polynomial of the module generated by this row.
Lemma 2.35.In the situation of Lemma 2.34 with Γ=D, suppose that u1,u2,u3 are the basis elements f1,f2,f3 of F. Let r∈F3 map to 1 in D, with class y[2]+[x,y]+[x,z] in gr2F3, and let ρ=r(f1,f2,f3). Then
hΛ∂ρ=sDhΛ,sD(t)=t2(1−hF2[D](t)t7).
Proof. By (F4), ∂ρ=(∂′r,0,…,0), where ∂′r∈ΛD3 is the local row, so Λ∂ρ is the module Λ∂′r⊂Λ3(−1) placed in the first three coordinates. By Lemma 2.34(a), Λ3=⨁jηjΛD3 with Fn(Λ3(−1))=⨁jηjFn−wj(ΛD3(−1)), and b↦ηjb is injective on ΛD3. The submodule Λ∂′r=⨁jηjΛD∂′r decomposes in the same way, so Fn(Λ∂′r)=⨁jηjFn−wj(ΛD∂′r) for the induced filtrations, and hΛ∂ρ=(hΛ/hF2[D])hΛD∂′r. It therefore suffices to show that hΛD∂′r=t2(hF2[D]−t7), for the filtration induced from ΛD3(−1). Write ID for the augmentation ideal of ΛD=F2[D], and X1,X2,X3,X4 for x−1,y−1,z−1,w−1, where w=[y,z].
The linear part of ∂′r. Here π also denotes the map F3→D. For g,h∈F3, (F1) gives ∂′(g2)=(1+π(g))∂′g and ∂′[g,h]=π(g)−1(π(h)−1−1)∂′g+π(g)−1π(h)−1(π(g)−1)∂′h. Moreover ∂′q is congruent modulo ID2 to the image qˉ∈F23 of q in F3/D2F3, and π(q)−1≡∑lqˉlXl modulo ID2 by (F2). Hence, for q∈D2F3, the row ∂′q has entries in ID, and modulo ID2 it is additive in q, because π(q)−1∈ID2. It vanishes modulo ID2 on D3F3=(D2F3)2[D2F3,F3] (Lazard’s formula): for q∈D2F3 the formulas above give ∂′(q2),∂′[q,h]∈ID2, since qˉ=0 and π(q)−1∈ID2. So the row modulo ID2 depends only on the class of q in gr2F3, the classes v[2] and [v,v′] giving the rows with l-th entries vl∑mvmXm and vl∑mvm′Xm+vl′∑mvmXm. For the class y[2]+[x,y]+[x,z] this gives
∂′r=(X2+X3,X2+X1,X1)(modID2).
Three facts about ΛD. Let A=⨁mIDm/IDm+1, and write X,Y,Z and W for the classes of X1,X2,X3 in degree one and of X4 in degree two. By Theorem 2.2 and Lemma 2.8(c), the products X1a1X2a2X3a3X4a4 with a1,a2,a4∈{0,1} and 0≤a3≤3, of weight a1+a2+a3+2a4≥m, form a basis of IDm; here X32=z2−1. Hence the monomials Xa1Ya2Za3Wa4 form a basis of A, and, since the coefficients of hF2[D] are 1,3,5,7,7,5,3,1,
dimJDm=32,31,28,23,16,9,4,1,0(m=0,…,8).
Because x and w are central and x2=y2=w2=z4=1, the elements X and W are central in A and X2=Y2=W2=Z4=0. Moreover ZY=YZ+W, because X3X2+X2X3=zy+yz=zy(w−1)≡X4(modJD3); by induction ZjY=YZj+jZj−1W for j≥1.
First, JD7=F2ΩD, where ΩD=∑g∈Dg: JD7 is one-dimensional and (g−1)JD7⊂JD8=0, so its nonzero element is invariant under left multiplication by D, and hence equal to ΩD.
Second, for m<7 the map Am→Am+13, ξ↦(ξX,ξY,ξZ), is injective. Let ξ=∑μbμμ over the basis monomials μ=Xa1Ya2Za3Wa4 of degree m. Right multiplication by X maps the monomials with a1=0 to distinct basis monomials and kills the others, and right multiplication by Z maps those with a3≤2 to distinct basis monomials and kills those with a3=3. So ξX=ξZ=0 forces ξ=∑a2,a4ba2a4XYa2Z3Wa4. Then
ξY=b00(XYZ3+XZ2W)+b01XYZ3W+b10XYZ2W,
a combination of distinct basis monomials, so ξY=0 forces ξ=b11XYZ3W, which has degree 7. Hence ξ=0 when m<7.
Third, the same holds for ξ↦(ξ(Y+Z),ξ(Y+X),ξX), because this triple is obtained from (X,Y,Z) by the invertible matrix
011110100.
This is the coefficient matrix of y[2]+[x,y]+[x,z], that is, the matrix of the 2-adic Hilbert symbol on 2,−5,5 in the proof of Lemma 2.7(c). So the injectivity used here, and with it the correction term sD, comes from the nondegeneracy of the local Hilbert pairing, that is, from local duality. The supplementary program dyadic241.py confirms these three facts by direct computation in ΛD.
The filtration. If ξ∈JDm∖JDm+1 with m<7, then by the third fact ξ∂′r lies in Fm+2 but not in Fm+3, whereas ΩD∂′r=0 because ΩDJD=0. Now let ξ∈ΛD with ξ∂′r∈Fm+2, and suppose that ξ∈/JDm+F2ΩD. Let m′<m be maximal with ξ∈JDm′+F2ΩD, and write ξ=ξ′+aΩD with ξ′∈JDm′ and a∈F2. Then ξ′∈/JDm′+1, and m′<7 because JD7=F2ΩD; so ξ∂′r=ξ′∂′r is not in Fm′+3⊃Fm+2, a contradiction. Therefore Fm+2(ΛD∂′r)=JDm∂′r for every m≥0, the map ξ↦ξ∂′r has kernel F2ΩD, and hΛD∂′r=∑m<7dim(JDm/JDm+1)tm+2=t2(hF2[D]−t7). □
Infinitude
By Definition 2.33 and the Hilbert polynomials listed before it,
In the proof of Proposition 2.39, the term −8t comes from the eight generators of a free group mapping onto GB; the terms in ψC2,ψC2×C2,ψD and ψC4 come from the local relations at the two real places, the four places above 3 and 5, the two dyadic places and the three places t0,t1,t2; and −sD comes from the row of the dyadic relator ρp1, which lies both in the module of p1 and in the sum of the modules of the other places.
Lemma 2.38.
PB(11734)=−88772225460489675187433948535241<0.
Proof. By (11), (12), the formula for sD in Lemma 2.35 and hF2[D](t)=(1+t)3(1+t2)2 (Lemma 2.8(c)), PB is a rational function with rational coefficients. Exact rational arithmetic, carried out by the supplementary program gs241.py, gives the stated value. □
Proposition 2.39.The group GB is infinite.
Proof. Suppose that GB is finite, and put Δ=GB and Λ=F2[Δ]. By Table 4, x1,y1,z1 are independent; choose g4,…,g8∈GS such that x1,y1,z1,g4,…,g8 is a basis of F28. Let F be free on f1,…,f8, and send f1,f2,f3 to x1,y1,z1 and fl to gl for 4≤l≤8. This is a minimal presentation of GS; we take the lifts x1=f1, y1=f2, z1=f3, and arbitrary lifts of the other local generators. Let π:F→Δ be the composite and M as in Lemma 2.32, so that hM=1−(1−8t)hΛ.
By Lemma 2.26(b), the image in Δ of each local group is the prescribed group Γ, and grnΓ→grnΔ is injective for n=1,2; this is the hypothesis of Lemma 2.34. Indeed, for Γ=D this is stated in Lemma 2.26(b). For Γ=C2 and Γ=C2×C2 we have gr2Γ=0, and gr1Γ=Γ injects into GB/D2GB because ιj=0, respectively because the image of Gν injects into GB/D2GB. For Γ=C4 generated by Frobtk, which has order four in GB/D3GB, we have Frobtk∈/D2GB, since otherwise its square would lie in D4GB⊆D3GB, and Frobtk2∈/D3GB; this is the injectivity on gr1Γ and on gr2Γ=⟨Frobtk2⟩. We use eleven places: the real places v1,v2, with Γ=C2; the four places above 3 and 5, with Γ=C2×C2; the dyadic places p1,p2, with Γ=D; and t0,t1,t2, with Γ=C4 generated by the Frobenius. For each of these places ν, let Uν be the module U of Lemma 2.34(c) for this Γ, formed with the lifts of the local generators at ν. Every generator of NB in Lemma 2.22 is of the form W(u1,…,uk) for one of these places and some W in the kernel of Fk→Γ, and conversely every such element lies in NB. By Lemma 2.32, M=∑νUν. Let U′ be the sum of the Uν over the ten places ν=p1. The row of ρp1 lies in Up1. Since ρp1 lies in the normal closure of the relators (i) of Lemma 2.22, its row also lies in the Λ-span of their rows (Lemma 2.32), which is contained in U′. Hence Λ∂ρp1⊂Up1∩U′. By Lemma 2.28(h4) and (h2), Lemma 2.34(c) and Lemma 2.35, for 0<t<1,
that is, 1≤hΛ(t)PB(t). Since hΛ(t)>0, this contradicts PB(34/117)<0 (Lemma 2.38). □
The Fields
Definition 2.40. The root discriminant of a number field M with discriminant ΔM is rd(M)=∣ΔM∣1/[M:Q].
Definition 2.41. For a number field M and a prime P of M above a rational prime p, the absolute type of P is the pair (e,f) of its ramification index and residue degree over p. For a Galois extension M/A of number fields, all primes of M above a prime p of A have the same ramification index e and residue degree f in M/A; the pair (e,f) is the type of p in M/A.
As in the introduction, let λ=29/43615, where 3615=3⋅5⋅241.
Theorem 2.42.For every power of two m≥249 there is a finite Galois extension K/B of degree m, contained in BG, such that GB→GB factors through Gal(K/B). In particular K contains the fixed field of D4GB, hence the fixed field of D3GB and the Kummer field E. Every such K is totally imaginary. Let ι1 also denote complex conjugation at any place of K above v1; these complex conjugations are conjugate in Gal(K/B), one of them is the image of the element ι1 of Definition 2.12, and the statements below do not depend on the choice. Let F=K⟨ι1⟩ and d=[F:Q]=m, let (b,c) be the signature of F and θ=c/d, and put θ∗=65535/131072=1/2−2−17. Then F has a real place,
rd(K)=rd(F)=λ,θ∗≤θ<21,
and b/d=1/(2Nι), where Nι≥215 is the number of conjugates of ι1 in Gal(K/B). The extension K/F is quadratic and unramified at every finite place. For r∈{2,3,5,7,29}, every prime of F above r splits in K, and the primes of K and of F above r have the absolute type (er,fr) of Table 5; that is, K/F has uniform local types (er,fr)r∈{2,3,5,7,29} in the sense of Definition 5.4. The completions of K at the primes above 2, 3, 5, 7 and 29 are the local extensions listed in Table 5.
r
(er,fr)
local extension over B
factor of rd(K)
2
(8,4)
L2/Q2, group D
29/4
3
(2,2)
Q3(−1,3)/Q3, group C2×C2
31/2
5
(2,2)
Q5(2,5)/Q5, group C2×C2
51/2
7
(1,8)
unramified of degree 4 over Bt0
1
29
(1,4)
unramified of degree 4 over Q29
1
241
e=2
unramified
2411/2
Table 5. Absolute types of the places of K and F above 2, 3, 5, 7 and 29, the corresponding local extensions of the completions of B, and the contributions of these primes to rd(K). The last row is the prime 241, which ramifies in B; K/B is unramified above it.
Proof.Existence. By Proposition 2.39, GB has open normal subgroups of arbitrarily large index, and D4GB is open because GB is finitely generated. Given m, choose an open normal subgroup U′⊂D4GB of GB of index at least m, and put N′=D4GB/U′. A nontrivial normal subgroup of a finite 2-group meets its center, so N′ contains normal subgroups of GB/U′ of every order dividing ∣N′∣. Take one of order [GB:U′]/m, which divides ∣N′∣=[GB:U′]/249. Its preimage U is open and normal in GB, contained in D4GB, and of index m. Let K be its fixed field in BG. Conversely, every K as in the statement is the fixed field of an open normal subgroup U⊂D4GB. We fix such a K and put H=Gal(K/B)=GB/U.
Complex places. Since −1=α0∈E⊂K, the field K is totally imaginary, and [K:F]=2. Let j∈{1,2}. The group H permutes the places of K above vj transitively. If an embedding σ:K→C induces such a place and ισ∈H is its complex conjugation, so that σˉ=σ∘ισ, then for h∈H we have σˉ∘h=(σ∘h)∘(h−1ισh), so the complex conjugation at the place of σ∘h is h−1ισh. By the definition of the local map at vj, the image of ιj in H is the complex conjugation at the place of K below the place of BS fixed in Definition 2.12. Hence every complex conjugation at a place of K above vj is conjugate in H to the image of ιj, and has elementary image ιˉj, since conjugation preserves elementary images. In particular, two choices of ι1 in the statement are conjugate in H, so an element of H maps the field F of one choice onto that of the other, and the statements of the theorem hold for both choices or for neither.
Let σ1:K→C induce the chosen place, so that σˉ1=σ1∘ι1 and σ1(F)⊂R. The embeddings above v1 are σ1∘h with h∈H, and two of them agree on F exactly when they differ by right multiplication by an element of ⟨ι1⟩. The restriction of σ1∘h to F is real exactly when σ1∘ι1h and σ1∘h agree on F, that is, when h−1ι1h∈⟨ι1⟩, that is, when h∈CH(ι1). Similarly, if ι′ is complex conjugation for an embedding σ2 above v2, the restriction of σ2∘h to F is real exactly when h−1ι′h=ι1. This never happens, because conjugation preserves elementary images and ιˉ′=ιˉ2=ιˉ1 (Lemma 2.26(d)). Hence b=∣CH(ι1)∣/2≥1, so F has a real place and θ=(1−b/d)/2<1/2. Moreover d=∣H∣, so b/d=∣CH(ι1)∣/(2∣H∣) is half the reciprocal of the size Nι of the conjugacy class of ι1 in H, that is, b/d=1/(2Nι). This class maps onto the class in GB of the element u1 of Definition 2.12, which has 215 elements by Lemma 2.26(c). Hence Nι≥215, b/d≤2−16 and θ≥(1−2−16)/2=θ∗.
Local structure and splitting. Let P be a prime of K above p∈S∪{t0,t1,t2}. Its decomposition group in H is conjugate to the image of the local map at p. Since U⊂D3GB, Lemma 2.26(b) shows that this image is D, C2×C2 or C4, the last one cyclic on the Frobenius, and that the prescribed local group injects into H. Therefore the completion KP is L2 over Bpj=Q2, the field Qp(u,p) over Bp=Qp for p=3,5, and the unramified extension of degree four of Btk. Since Bt0 is unramified of degree two over Q7 and Bt1=Bt2=Q29, the places of K above 2,3,5,7 and 29 have the absolute types of Table 5. Every element of the decomposition group has elementary image in the span of the images of the local generators at p, whereas every conjugate of ι1 has elementary image ι1. By Lemma 2.26(d), ι1 lies in no decomposition group of a prime above S∪{t0,t1,t2}, so the decomposition group of P in Gal(K/F)=⟨ι1⟩ is trivial. Hence the places of F above 2,3,5,7,29 split in K, and they have the same absolute types.
Ramification and root discriminant. The extension K/B, and hence K/F, is unramified at every prime outside S, in particular at (241). At primes above S the inertia group of K/F is trivial, because it lies in the decomposition group. So K/F is unramified at every finite place, its relative discriminant is trivial, ∣ΔK∣=∣ΔF∣2, and rd(K)=rd(F). To compute rd(K) we use ∣ΔK∣=241[K:B]NdK/B. Since K/B is Galois, for p∈S all primes P above p have the same ramification index e, residue degree and different exponent ordP(DK/B), where ordP and ordp are the normalized valuations at P and p, and ordp(dK/B)=[K:B]ordP(DK/B)/e. For p above 3 or 5 we have e=2 and ordP(DK/B)=1, since over the unramified extension Qp(u) the different of Qp(u,p) is generated by 2p. For p1,p2, Lemma 2.8(d) gives ordP(DK/B)/e=9/4. Since [K:Q]=2[K:B], the base contributes 2411/2, each of the two primes above 3 contributes 31/4, each of the two above 5 contributes 51/4, and each dyadic prime contributes 29/8. Hence rd(K)=29/431/251/22411/2=λ. □
Remark 2.43. The Lean development [22] formalizes a version of this construction; Remark 4.40 states what it proves and the numerical hypothesis it assumes.
The Kummer Field and Its Quadratic L-Functions
The analytic estimate of Section 4 compares every field of Theorem 2.42 with one fixed subfield, the Kummer field
E=B(α0,…,α7)=B(V)
of Subsection 2.3, with α0,…,α7 as in (4) and V as in Lemma 2.4(d). It is the maximal elementary abelian 2-extension of B unramified at the finite primes outside S, that is, the fixed field of the Frattini subgroup of GB (Lemma 3.2). The Lean formalization calls E the genus field; we do not use that name here. This section shows that E lies in every field of Theorem 2.42. It then writes ζE as the product of ζB and 255 quadratic Hecke L-functions of B, and determines their conductors, gamma factors, root numbers and Euler factors.
For v,e∈F28 put ⟨v,e⟩=∑i=07viei∈F2; as in Section 2, vectors are written as strings of their coordinates. For a nonzero ideal a of a number field, Na denotes its absolute norm. The number of positive divisors of an integer n≥1 is τ(n).
The Characters of the Kummer Field
By Lemma 2.10, [E:B]=256, and the elementary image g∈F28 of Definition 2.11, given by g(αi)=(−1)giαi, identifies Gal(E/B) with F28.
Definition 3.1. For e∈F28 put
αe=i=0∏7αiei,Be=B(αe),
and let χe be the character g↦(−1)⟨gˉ,e⟩ of Gal(E/B).
The character χe takes the values ±1; in the additive notation of Definition 2.11, χe(g)=(−1)χαe(g). Since g(αe)=χe(g)αe, the character χe cuts out Be, and the 256 characters χe are all the characters of Gal(E/B). For e=0 the extension Be/B is quadratic, because the classes of α0,…,α7 are independent in B×/B×2 (Lemma 2.4(d)).
Lemma 3.2.The fixed field of the Frattini subgroup D2GB of GB is E. Consequently E is contained in every field K of Theorem 2.42.
Proof. Let GS be the group of Subsection 2.3, so that GS/D2GS=Gal(E/B) (Lemma 2.10), and let GB=GS/N as in Definition 2.20. By Lemma 2.21(b), N⊆D2GS, so GB/D2GB=GS/D2GS=Gal(E/B). A field K of Theorem 2.42 contains the fixed field of D4GB⊆D2GB, and hence contains E. □
Local Data
Definition 3.3. For a prime p∈/S of B, the Frobenius vectorFrobp∈F28 is the elementary image of the Frobenius of p in Gal(E/B), that is, of the Frobenius element Frobp of Subsection 2.3.
For p=tk it is the vector of Frobtk in Table 4. By Euler’s criterion, (Frobp)i=1 exactly when αi is not a square in OB/p. For a place p of B we write (⋅,⋅)p for the quadratic Hilbert symbol of Bp.
Proposition 3.4.Let e∈F28∖{0}. For a prime p of B let cp(e) be the exponent of p in the conductor of χe, and if cp(e)=0, let χe(p)∈{±1} be the value of χe at the Frobenius of p.
If p∈/S, then cp(e)=0 and χe(p)=(−1)⟨Frobp,e⟩. Let p be the rational prime below p. If p splits in B, then p=(p,241−r) with r2≡241(modp), and (Frobp)i=1 exactly when the image of αi under 241↦r is a quadratic nonresidue modulo p. If p is inert, then (Frobp)i=1 exactly when NB/Q(αi) is a quadratic nonresidue modulo p. If p=241, the same holds with 241↦0.
Let p lie above p∈{3,5}, let u=−1 if p=3 and u=2 if p=5, and let αi be the generator of p. If (u,αe)p=−1, then cp(e)=1. Otherwise cp(e)=0 and χe(p)=(−αi,αe)p. The Frobenius generator φp of Definition 2.12, which fixes αi, acts on square roots by (−αi,⋅)p (Table 4).
Let p=pj. If (5,αe)p=−1, then cp(e)=3. If (5,αe)p=1 and (−1,αe)p=−1, then cp(e)=2. Otherwise cp(e)=0 and χe(p)=(−2,αe)p.
The real place vk becomes complex in Be, that is, χe is nontrivial at vk, exactly when αe<0 at vk.
Proof. For (1), p is odd and αe is a unit at p, so p is unramified in Be. By Euler’s criterion the Frobenius multiplies αe by the quadratic residue symbol of αe modulo p, which is ∏i(αi/p)ei. For a split prime, OB/p≅Fp via 241↦r. For an inert prime, OB/p≅Fp2 and the conjugation of B induces the Frobenius x↦xp of the residue field, so the residue of NB/Q(α) is αˉp+1. An element of Fp2× is a square exactly when αˉ(p2−1)/2=(αˉp+1)(p−1)/2 equals 1, that is, when its norm is a square in Fp. The prime above 241 has residue field F241 and 241↦0.
For (2) and (3) we use local class field theory for Bp=Qp. The local Artin map sends b∈Qp× to the automorphism multiplying αe by (b,αe)p. Units map onto inertia, a uniformizer maps to a Frobenius, and the conductor exponent is the least n≥0 such that b↦(b,αe)p is trivial on U(n), where U(0)=Zp× and U(n)=1+pnZp[19].
For p=3,5, the unit group modulo squares is {1,u}, and U(1) consists of squares by Hensel’s lemma. Hence cp(e)≤1, with equality exactly when (u,αe)p=−1. In the unramified case the Frobenius is the image of any uniformizer, such as −αi.
For p=2, the squares of Z2× form U(3). Moreover U(2)=U(3)∪5U(3) and U(0)=U(1)={±1,±5}U(3). The character is thus trivial on U(3). It is trivial on U(2) exactly when (5,αe)=1, and on U(0) exactly when also (−1,αe)=1. This gives the stated exponents; the exponent 1 cannot occur since U(0)=U(1). In the unramified case the Frobenius is the image of the uniformizer −2.
For (4), the real place vk splits in Be exactly when αe>0 at vk. □
Each condition in the proposition is linear in e:
Lemma 3.5.Let e∈F28. Then (5,αe)pj=(−1)⟨xˉj,e⟩, (−1,αe)pj=(−1)⟨yˉj,e⟩ and (−2,αe)pj=(−1)⟨zˉj,e⟩. For a prime p above 3 or 5, with u as in Proposition 3.4(2) and αi the generator of p, (u,αe)p=(−1)⟨τˉp,e⟩ and (−αi,αe)p=(−1)⟨φˉp,e⟩. Finally, αe<0 at vk exactly when ⟨ιˉk,e⟩=1.
Proof. In each case (b,αe)p=(−1)⟨w,e⟩, where w is the elementary image of the local generator that acts on square roots by (b,⋅)p in Table 4: at pj the vectors xˉj,yˉj,zˉj belong to b=5,−1,−2, and above 3 and 5 the images of the inertia and Frobenius generators belong to b=u and b=−αi. Likewise the complex conjugation ιk negates αe exactly when αe<0 at vk (Lemma 2.7(a)), and it multiplies αe by (−1)⟨ιˉk,e⟩. □
We put νk(e)=⟨ιˉk,e⟩∈{0,1} for k=1,2. By Proposition 3.4, the conductor of χe is
fe=p∈S∏pcp(e).
Corollary 3.6.The type (Definition 2.41) of a prime p of B in E/B is (4,2) at p1,p2; (2,2) at q1,q2,r1,r2; and (1,1) or (1,2) at p∈/S, according as Frobp=0 or not. In particular t0,t1,t2 have type (1,2) in E/B.
Proof. For p∈S, local class field theory identifies the decomposition group of p in Gal(E/B)=F28 with the image of Qp× under the local Artin map, and the inertia group with the image of the units. By Lemma 3.5, at pj these are spanned by xˉj,yˉj,zˉj and by xˉj,yˉj; the three vectors are independent (Table 4), so the decomposition group has order 8 and inertia has order 4. Above 3 and 5 the two elementary images are independent, which gives (2,2). For p∈/S the decomposition group is generated by Frobp. The Frobenius vectors of t0,t1,t2 are nonzero (Table 4). □
where L(s,χe)=ζBe(s)/ζB(s) is the Hecke L-function of χe.
Proof. Compare Euler factors at a prime p of B. Let Tp⊆Zp⊆G=Gal(E/B) be its inertia and decomposition groups, and put fp=∣Zp/Tp∣ and gp=∣G/Zp∣. The Euler factor of ζE at p is (1−Xfp)−gp with X=Np−s. A character nontrivial on Tp contributes the factor 1. The characters trivial on Tp are those of G/Tp; they restrict onto the fp characters of the cyclic group Zp/Tp, each with gp extensions. Hence their factors multiply to ∏ζfp=1(1−ζX)−gp=(1−Xfp)−gp. The same identity for Be/B identifies L(s,χe) with ζBe/ζB. The second identity is the case of B/Q. Since 241≡1(mod8), the prime 2 splits in B, and quadratic reciprocity shows that the character of B/Q is χB. □
Put ΓR(s)=π−s/2Γ(s/2) and ΓC(s)=2(2π)−sΓ(s), so that ΓC(s)=ΓR(s)ΓR(s+1) by the duplication formula. For a number field M with discriminant ΔM and signature (r1(M),r2(M)), the completed Dedekind zeta function is
ΛM(s)=∣ΔM∣s/2ΓR(s)r1(M)ΓC(s)r2(M)ζM(s).
Proposition 3.8.Let e=0, write L(s,χe)=∑n≥1an(e)n−s, and put
The coefficients an(e) are integers with ∣an(e)∣≤τ(n). The Euler factor at a rational prime p is ∏p∣p(1−χe(p)Np−s)−1, with χe(p)=0 when cp(e)>0.
The function Λ(s,χe) equals ΛBe(s)/ΛB(s). It is entire of order at most one and satisfies Λ(s,χe)=Λ(1−s,χe); that is, the root number of L(s,χe) is 1.
Exactly 63 of the characters have ν1(e)=ν2(e)=0, exactly 64 have ν1(e)=ν2(e)=1, and 128 have ν1(e)=ν2(e).
Proof. The Euler factor in (1) is Proposition 3.7 applied prime by prime. Above a rational prime there are at most two primes of B. So the local factor is either a product of at most two factors (1−cp−s)−1, or a single factor (1−cp−2s)−1, with c∈{0,±1}. In both cases the coefficient of p−ks has modulus at most k+1=τ(pk). Multiplicativity gives (1).
For (2), the completed Dedekind zeta functions satisfy ΛM(1−s)=ΛM(s), and L(s,χe) extends to an entire function because χe is nontrivial [12, 37]. The discriminants satisfy ∣ΔBe∣=∣ΔB∣2NdBe/B, where dBe/B is the relative discriminant, and the conductor–discriminant formula gives dBe/B=fe for the quadratic extension Be/B. Hence ∣ΔBe∣/∣ΔB∣=241Nfe=Qe. If νk(e)=0, the place vk splits into two real places of Be and contributes ΓR(s)2/ΓR(s)=ΓR(s). If νk(e)=1, it becomes one complex place and contributes ΓC(s)/ΓR(s)=ΓR(s+1). Thus ΛBe/ΛB=Λ(s,χe). It is the completed Hecke L-function of χe, which is entire of order at most one [12, 37].
For (3), the vectors ιˉ1,ιˉ2 are linearly independent. So e↦(ν1(e),ν2(e)) maps F28 onto F22 with fibers of size 64. Removing e=0 from the fiber over (0,0) gives the counts. □
Definition 3.9. Let e=0. The gamma factor of χe is pure if ν1(e)=ν2(e)=ν, and then it equals ΓR(s+ν)2; otherwise it is mixed, and then it equals ΓR(s)ΓR(s+1)=ΓC(s).
Corollary 3.10.For e=0, Qe=241⋅2i⋅m, where 2i=2cp1(e)+cp2(e)∈{1,4,8,16,32,64} and m=3cq1(e)+cq2(e)5cr1(e)+cr2(e) divides 225. The number of characters χe with given 2i and m is given by the table
In particular 723≤Qe≤241⋅14400=3470400.
Proof. Only the primes of S divide the conductors, so Qe=241Nfe with
By Proposition 3.4, cpi(e)∈{0,2,3} and cqj(e),crj(e)∈{0,1}, and by Lemma 3.5 these exponents are determined by the vectors of Table 4; counting the 255 vectors e=0 gives the table, which the supplementary program check255.gp also prints (Remark 3.12). No nontrivial character has Nfe=1, since B has class number one and a unit of norm −1 (Lemma 2.4(a),(b)), so that its narrow class number is one. The range of Qe can be read off from the table. □
Remark 3.12. The supplementary program lfun241.py computes, for each e=0, the integer Qe, the pair (ν1(e),ν2(e)) and the coefficients an(e) for n≤Ne=4⌊Qe⌋+1, the truncation point used in Subsection 4.6, from Propositions 3.4 and 3.8(1); these are the data used in Section 4.
The supplementary program check255.gp compares all 255 of them with PARI/GP’s own number fields and L-functions [41]. For each e it computes the maximal order of Be, checks ∣ΔBe∣=241Qe and the pair (ν1(e),ν2(e)) against the signature of Be, and checks every coefficient an(e) with n≤Ne against the Dirichlet series of ζBe/ζB. All agree. It also prints the Table 3.11, the counts of Proposition 3.8(3) and the range of Qe. These comparisons check the implementation; the identification of the L-functions rests on Propositions 3.4 and 3.8.
2i\m
1
3
5
9
15
25
45
75
225
total
1
0
2
2
1
4
1
2
2
1
15
4
2
4
4
2
8
2
4
4
2
32
8
4
8
8
4
16
4
8
8
4
64
16
1
2
2
1
4
1
2
2
1
16
32
4
8
8
4
16
4
8
8
4
64
64
4
8
8
4
16
4
8
8
4
64
total
15
32
32
16
64
16
32
32
16
255
Table 3.11.
The Relative Zeta Value and an Explicit Upper Bound
We bound the relative zeta value LF(1) needed in the geometric construction. Logarithms of Dedekind zeta functions are normalized by the degree: for a number field A we work with [A:Q]−1logζA(s).
Definition 4.1. Let K be the set of all finite Galois extensions K/B contained in BG, of degree a power of two at least 249, such that GB→GB factors through Gal(K/B) (notation of Definition 2.20). These are the fields K of Theorem 2.42.
For K∈K we use the notation of Theorem 2.42: ι1 is a complex conjugation at a place of K above v1, F=K⟨ι1⟩, d=[F:Q], so that [K:Q]=2d, (b,c) is the signature of F, θ=c/d, and λ=29/43615 is the root discriminant of both K and F. Put
ℓ=logλ=49log2+21log3615=5.65600472355…
Different choices of ι1 are conjugate in Gal(K/B) and give conjugate fields F; hence b,c,θ and the function LF(s)=ζK(s)/ζF(s) depend only on K. The extension K/F is quadratic and unramified at every finite place, K is totally imaginary, and b≥1. Let γ be Euler’s constant. For a prime p of B, let eK(p) and fK(p) be its ramification index and residue degree in the Galois extension K/B, so that (eK(p),fK(p)) is the type of p in K/B (Definition 2.41).
Theorem 4.2.There is an m0 such that every K∈K with [K:B]≥m0 satisfies
d1logLF(1)<C:=0.04871285.
The proof assumes neither the generalized Riemann hypothesis nor a Brauer–Siegel type asymptotic formula, and it does not give an effective value of m0. It has three steps. Monotonicity of the completion of LF reduces the value at s=1 to ζK at a real point σ>1 (Subsection 4.1). The Tsfasman–Vlăduţ inequality bounds the error of this shift (Subsection 4.2, Proposition 4.13). Finally, ζK(σ) is bounded prime by prime by ζE(σ), for the Kummer field E of Section 3, minus corrections at the primes above 2, 7 and 29, whose types in K/B are known, and at the primes of P4 (Definition 4.15), whose residue degree in K/B is at least 4 (Proposition 4.19); ζE(σ) is evaluated rigorously in Subsections 4.4–4.6 (Proposition 4.29). Subsection 4.7 computes these corrections and the lower bounds for the types in K/B that Proposition 4.13 uses, and the last subsection combines the steps. Throughout, σ denotes a real number greater than 1 and ε=σ−1; in this section ε is this positive number, not the unit α1 of (4).
Reduction to s>1.
For a number field A, write hA, RegA, wA and ΔA for its class number, regulator, number of roots of unity and discriminant. The residue of ζA at s=1 is 2r1(A)(2π)r2(A)hARegA/(wA∣ΔA∣1/2). Since K/F is unramified at the finite places, ∣ΔK∣=∣ΔF∣2. Hence
Let s>1 be real. At a prime of F of norm q, the Euler factor of LF is (1−q−s)−1 if the prime splits in K and (1+q−s)−1 if it is inert. Its square is at most the Euler factor of ζK, which is (1−q−s)−2, respectively (1−q−2s)−1. Consequently
dlogLF(s)≤2dlogζK(s)=[K:Q]logζK(s).(14)
The function LF is the Hecke L-function of the quadratic character of K/F. This character is unramified at every finite place and nontrivial at every real place of F, since K is totally imaginary. With ΓR, ΓC and ΛA as in Section 3, put
By Hecke’s theorem [12, 37], LF is entire of order at most one, and it satisfies LF(s)=LF(1−s) because both completed Dedekind zeta functions do.
Lemma 4.6.The function LF(s) is positive and nondecreasing on [1,∞).
Proof. The Euler product shows that LF has no zeros with Res>1, and the functional equation excludes Res<0. So every zero ρ satisfies 0≤Reρ≤1. There is no zero on [1,∞), by the Euler product and by (13). As LF(s)>0 for large real s, it is positive on [1,∞).
Put Ξ(z)=LF(21+z). It is entire of order at most one, even, and real on the real axis. Its zeros z=ρ−21 satisfy ∣Rez∣≤21, and ∑z=0∣z∣−2<∞. Group the nonzero zeros into pairs {z0,−z0}, and let m≥0 be the order of Ξ at 0. By Hadamard’s theorem for functions of order at most one,
ΞΞ′(z)−zm−{z0,−z0}∑z2−z022z
is constant, where the series converges absolutely and locally uniformly away from the zeros. Since Ξ is even, Ξ′/Ξ is odd, so the constant is zero. For real x≥21 the logarithmic derivative is real, so it equals the sum of the real parts. If z0=a+iy, then
Rex2−z022x=(x2−a2+y2)2+4a2y22x(x2−a2+y2)≥0,
because ∣a∣≤21≤x. Since also m/x≥0, the logarithmic derivative is nonnegative, and logΞ is nondecreasing on [21,∞). □
Taking logarithms, dividing by d, and using b/d=1−2θ, c/d=θ and (14), we obtain (17). Since Υ is affine in θ, Υ(ε,θ)→0 as ε↓0, uniformly for 0≤θ≤21. □
The term Υ does not appear in the final estimate. With the value ε=1/300 used in the proof of Theorem 4.2, (17) would add Υ(1/300,θ)≈0.0054 for θ near 21. Instead, the proof of Proposition 4.13 applies (17) with s−1 in place of ε for each fixed s∈(1,2], passes to the limit along a sequence of fields, and then lets s↓1, so that Υ vanishes in the limit; the Tsfasman–Vlăduț inequality then bounds the change from 1+ε to 1 by the term ε(κ∞+D) of that proposition, which is about 0.00215 in the proof of Theorem 4.2. This limit is also the reason why m0 is not effective.
The Tsfasman–Vlăduţ Inequality over B
For q>1 and s≥1 put gs(q)=−log(1−q−s).
Lemma 4.10. Let M be a finite Galois extension of B and s>1. For a prime p of B let eM(p) and fM(p) be its ramification index and residue degree in M/B. Then
where P runs over the primes of M. Each summand on the right does not increase when eM(p) or fM(p) is replaced by a larger real number. If M⊆M′ are both Galois over B, then eM(p)∣eM′(p) and fM(p)∣fM′(p).
Proof. Above p there are [M:B]/(ef) primes of M, each of norm Npf, and [M:Q]=2[M:B]. This gives both formulas. The first equality in (19) follows from −ζM′/ζM(s)=∑PlogNP/(NPs−1). For monotonicity in f, write gs(xf)/f=∑j≥1x−fjs/(fj) for x>1; each term decreases in f. The divisibilities hold because ramification indices and residue degrees are multiplicative in towers. □
Put
κ∞=4ℓ−γ−log(4π),w(q)=2logqm≥1∑qm+11.
Here 2κ∞ is the right side, and w(q) the weight of the prime power q, in the Tsfasman–Vlăduţ inequality (20) below.
Proposition 4.13. Let σ=1+ε with 0<ε≤1, and let Y∗ be a real number with [K:Q]−1logζK(σ)≤Y∗ for every K∈K. For each prime p of B, let ep0,fp0≥1 be integers with eK(p)≥ep0 and fK(p)≥fp0 for every K∈K, that is, lower bounds for the type of p in K/B, and let
D≥p∑2ep0(Np2fp0−1)logNp.
Then every C′>Y∗+ε(κ∞+D) has the following property: there is an m0 such that d−1logLF(1)<C′ for every K∈K with [K:B]≥m0.
Proof.A limiting sequence. Suppose not. Then there are fields K(j)∈K of strictly increasing degree, hence pairwise non-isomorphic, with dj−1logLF(j)(1)≥C′, where [K(j):Q]=2dj. For a prime power q let Nq(K) be the number of primes of K of norm q. If q=pf, then Nq(K)≤[K:Q]/f, since the local degrees above p add up to [K:Q]. By a diagonal argument we may assume that βq=limjNq(K(j))/(2dj) exists for every q.
The Tsfasman–Vlăduţ inequality. Tsfasman and Vlăduţ call a sequence (Ai) of pairwise non-isomorphic number fields asymptotically exact if its genus g(Ai)=log∣ΔAi∣1/2 tends to infinity and the limits ϕq=limNq(Ai)/g(Ai), ϕR=limr1(Ai)/g(Ai) and ϕC=limr2(Ai)/g(Ai) exist. Their unconditional Basic Inequality [42], Section 3.2, Proposition 3.1 states that every such sequence satisfies
For our sequence g(K(j))=djℓ and K(j) is totally imaginary, so ϕq=2βq/ℓ, ϕR=0 and ϕC=1/ℓ, and the inequality becomes
q∑βqw(q)≤2ℓ−γ−log(4π)=2κ∞.(20)
The value at 1. Put Z(s)=∑qβqgs(q) for s≥1. Since g1(q)≤1/(q−1) and w(q)≥2logq/(q+1), we have g1(q)≤3w(q)/(2log2), so Z(1) is finite. For fixed real s>1, (2dj)−1logζK(j)(s)=∑q(Nq(K(j))/2dj)gs(q) tends to Z(s) by dominated convergence: for q=pf the terms are at most gs(q)/f≤2q−s, and ∑qq−s<∞. By Lemma 4.8 with s−1 in place of ε, for 1<s≤2,
Finally Z(1+ε)=limj(2dj)−1logζK(j)(1+ε)≤Y∗. Thus C′≤Z(1)≤Y∗+ε(κ∞+D), a contradiction. □
Comparison with the Kummer Field
We use the set S of the six primes of B above 2, 3 and 5 and the primes t0=7OB and t1,t2 above 29 (Lemma 2.4), the group GB of Definition 2.20 and its Zassenhaus filtration DnGB, the relation space R2 in degree two (Definition 2.23), which Lemma 2.25 computes, the restricted square v[2] of a vector v∈F28 as in Section 2, and the Frobenius vectors Frobp of Definition 3.3.
Definition 4.15. Let Σ2={v∈F28:v[2]∈R2}, and let P4 be the set of primes p∈/S∪{t0,t1,t2} of B with Np≤106 and Frobp∈/Σ2.
Lemma 4.16.Let K∈K, and let p∈/S be a prime of B with Frobp∈/Σ2. Then eK(p)=1 and fK(p)≥4.
Proof. The field K is unramified over B outside S, so eK(p)=1, and a prime P of K above p has a Frobenius FrobP∈Gal(K/B) of order fK(p). Choose g∈GB mapping to FrobP. Its image in GB/D2GB=Gal(E/B) is the Frobenius of p in E (Lemma 3.2), whose elementary image is Frobp; it is nonzero, since 0∈Σ2. By Lemma 2.25, which uses the complete list of generators of NB in Lemma 2.22, and since Frobp[2]∈/R2, g has order at least 4 in GB/D3GB. The projection GB→GB/D3GB factors through Gal(K/B), because K contains the fixed field of D4GB⊆D3GB. So FrobP has order at least 4. □
The next proposition compares ζK(σ) with ζE(σ) and subtracts two corrections. The first concerns the primes p1,p2 above 2 and t0,t1,t2 above 7 and 29, whose types in K/B are known; the second concerns the primes of P4, which are found in Subsection 4.7. Put
Proof. By Lemma 3.2, E⊆K, and both are Galois over B. Apply (18) to M=E and M=K. By Lemma 4.10, each term for K is at most the corresponding term for E. For the primes listed below we bound the term for K from above using what is known about the type in K/B, and subtract the difference between the term for E and this upper bound.
The types in E/B are given by Corollary 3.6: (4,2) at p1,p2, and (1,2) at t0,t1,t2 and at the primes of P4. By Theorem 2.42, the primes of K above 2, 29 and 7 have absolute types (8,4), (1,4) and (1,8). The primes pj, t1 and t2 have residue degree one over Q, and t0 has residue degree two. So the types in K/B are (8,4) at pj and (1,4) at t0,t1,t2. By Lemma 4.16, the primes of P4 have eK=1 and fK≥4. The differences of the terms of E and the upper bounds for the terms of K are exactly the summands of (21) and (22). For example, at pj the term of E is gσ(22)/(2⋅4⋅2) and that of K is gσ(24)/(2⋅8⋅4). □
An Approximate Functional Equation for Two Gamma Factors
We evaluate ζE(σ) through Proposition 3.7. The factors L(σ,χe) are computed from an approximate functional equation. For real μ and x,y>0 define the incomplete gamma and Bessel integrals
Hμ(x)=∫1∞uμ−1e−xudu,Iμ(y)=∫1∞uμK0(yu)du,
where K0 is the modified Bessel function of the second kind of order zero, and their envelopes (upper bounds, by Lemma 4.21(2))
V1(x)=xe−x(1+x1),V2(y)=2πy3/2e−y(1+y1).
Numerical evaluation of L-functions from their functional equations is developed in [3]. Part (1) of the next lemma treats mixed gamma factors, and part (2), for pure gamma factors, uses the Bessel function K0. We prove both.
Lemma 4.20.Let an be real numbers with ∣an∣≤τ(n), let Q>0 and ν1,ν2∈{0,1} (integers, not places), and let L(s)=∑n≥1ann−s. Suppose that Λ(s)=Qs/2ΓR(s+ν1)ΓR(s+ν2)L(s) extends to an entire function of order at most one with Λ(1−s)=Λ(s). Put t=2π/Q and let σ>1 be real.
Proof. In case (1) put k(x)=e−x and ν=0, and in case (2) put k(x)=K0(x). Their Mellin transforms are Γ(w) and 2w−2Γ(w/2)2 for Rew>0[25], 10.43.19. Let Θ(u)=∑nannνk(tnu) for u>0. For Rew>1+ν, termwise integration gives
Θ(w)=∫0∞Θ(u)uw−1du={t−wΓ(w)L(w)=21Λ(w)2w−2t−wΓ(w/2)2L(w−ν)=41Qν/2Λ(w−ν)in case (1),in case (2).
In case (1) we used ΓR(w)ΓR(w+1)=ΓC(w) and Qw/2(2π)−w=t−w. In both cases Θ is entire and Θ(w)=Θ(1+2ν−w).
The function L=Λ/(Qs/2ΓR(s+ν1)ΓR(s+ν2)) of the complex variable s is entire of finite order. Fix η>0. The function L is bounded on Res=1+η, and, by the functional equation and Stirling’s formula, of polynomial growth on Res=−η. By the Phragmén–Lindelöf principle it has polynomial growth in ∣Ims∣ on every vertical strip. With Stirling’s formula this shows that Θ(w) decays exponentially in ∣Imw∣, uniformly on vertical strips. Mellin inversion on a line Rew=c>1+ν and a shift to Rew=1+2ν−c are therefore justified, and no residues occur. With the functional equation of Θ they give
Θ(u)=u−1−2νΘ(1/u)(u>0).
In particular Θ(u) decays rapidly as u↓0. Splitting the Mellin integral at u=1 and substituting u↦1/u below 1 gives
Θ(σ+ν)=∫1∞Θ(u)(uσ+ν−1+uν−σ)du.
Termwise integration, justified by the exponential decay of k, turns the right side into ∑nannν times Hσ(tn)+H1−σ(tn) in case (1), and times Iσ+ν−1(tn)+Iν−σ(tn) in case (2). On the left, Θ(σ)=t−σΓ(σ)L(σ) in case (1), and Θ(σ+ν)=2σ+ν−2t−σ−νΓ((σ+ν)/2)2L(σ) in case (2). □
Lemma 4.21. Let μ∈R.
The functions Hμ and Iμ are Laplace transforms of positive measures on [1,∞). They are completely monotone on (0,∞), extend holomorphically to Rez>0, and satisfy ∣H(z)∣≤H(Rez) there, for H=Hμ or H=Iμ.
and K0 satisfies y2K0′′+yK0′−y2K0=0 and K0′=−K1, where K1 is the modified Bessel function of the second kind of order one.
Proof. The definition exhibits Hμ as the Laplace transform of uμ−1du on [1,∞). Taking the order 0 in [25], 10.32.9 and substituting v=coshr in its integral gives K0(x)=∫1∞e−xv(v2−1)−1/2dv. Substituting this into Iμ, and then v=r/u in the inner integral, shows
Iμ(y)=∫1∞e−yr(∫1ruμ(r2−u2)−1/2du)dr,
with a nonnegative inner integral. Differentiation under the integral gives complete monotonicity. Since the measures are positive and ∣e−zr∣=e−rRez, we get ∣H(z)∣≤H(Rez) for Rez>0, which proves (1).
For (2), uμ−1≤u for u≥1 and μ≤2, and ∫1∞ue−xudu=V1(x). For Iμ we use K0(x)≤K1/2(x)=π/(2x)e−x[25], 10.39.2. It holds because the modified Bessel function of order η≥0 is ∫0∞e−xcoshrcosh(ηr)dr[25], which increases with η. Since uμ−1/2≤u for μ≤3/2,
Iμ(y)≤2yπ∫1∞uμ−1/2e−yudu≤V2(y).
For (3), differentiate under the integral and integrate by parts: xHμ′(x)=−∫1∞uμxe−xudu=−e−x−μHμ(x), and yIμ′(y)=∫1∞uμ+1dudK0(yu)du=−K0(y)−(μ+1)Iμ(y). The last two identities are Bessel’s equation and a standard derivative [25]. □
In the application σ=301/300, so 1<σ≤3/2. The four exponents σ, 1−σ, σ+ν−1 and ν−σ in Lemma 4.20 are then covered by the envelopes: V1 bounds Hσ and H1−σ, and V2 bounds Iσ+ν−1 and Iν−σ.
Lemma 4.22. Let 1<σ≤3/2, let ∣an∣≤τ(n), let t>0, N≥1 and ω>1.
(1) If t(N+1)≥ω−1, then for μ∈{σ,1−σ}
n>N∑∣an∣Hμ(tn)≤ζ(ω)2(N+1)ωV1(t(N+1)).
(2) If ν∈{0,1} and t(N+1)≥ω+ν−23, then for μ∈{σ+ν−1,ν−σ}
n>N∑∣an∣nνIμ(tn)≤ζ(ω)2(N+1)ω+νV2(t(N+1)).
Proof. The logarithmic derivative of xωV1(x)=xω−1e−x(1+1/x) is (ω−1)/x−1−1/(x(x+1))<0 for x≥ω−1. So nωV1(tn)≤(N+1)ωV1(t(N+1)) for n>N. By Lemma 4.21(2),
and the last sum is ζ(ω)2. Part (2) is the same, since yω+νV2(y) is a multiple of yω+ν−3/2e−y(1+1/y), which decreases for y≥ω+ν−23. □
Enclosures of Hμ and Iμ, and Finite Sums
An enclosure of a real number is an interval that contains it. We enclose Hμ and Iμ on a geometric grid. Let H be Hμ or Iμ with μ as in Lemma 4.22, and let V be its envelope, V1 or V2. For J≥0 and ξ>0 put
ρJ(ξ)=V(ξ/2)1−1/1616−J−1.
Lemma 4.23. Let ξ>0 and let hj=H(j)(ξ)/j! be the Taylor coefficients of H at ξ.
(1) (−1)jhj≥0 and ∣hj∣≤V(ξ/2)(2/ξ)j for all j≥0.
(2) For ∣x−ξ∣≤ξ/32, H(x)−∑j=0Jhj(x−ξ)j≤ρJ(ξ). For x=31ξ/32 the difference lies in [0,ρJ(ξ)].
(3) Let kj be the Taylor coefficients at ξ of e−x if H=Hμ, and of K0 if H=Iμ. If H=Hμ, then (j+1)ξhj+1=−kj−(j+μ)hj, and kj=(−1)je−ξ/j!. If H=Iμ, then (j+1)ξhj+1=−kj−(j+μ+1)hj, where k0=K0(ξ), k1=−K1(ξ), and
Proof. Complete monotonicity gives the signs in (1). The disc ∣z−ξ∣≤ξ/2 lies in Rez≥ξ/2, where ∣H(z)∣≤H(ξ/2)≤V(ξ/2) by Lemma 4.21. Cauchy’s estimate gives the bound on hj. For ∣x−ξ∣≤ξ/32, the omitted terms have modulus at most V(ξ/2)16−j, and their sum over j>J is at most ρJ(ξ). For x<ξ every term hj(x−ξ)j is nonnegative by (1), so the remainder is nonnegative. Part (3) follows by comparing coefficients of (x−ξ)j in the differential equations of Lemma 4.21(3), written at x=ξ+(x−ξ). □
The grid points are ξi=64(31/32)i for 0≤i≤i1, where i1=1878 is the least index with ξi1≤2−80. Taylor polynomials are formed at ξ0,…,ξi1−1. The initial values e−ξi, K0(ξi) and K1(ξi) of part (3) of Lemma 4.23 are enclosed with the rigorous exponential and Bessel functions of Arb [15], which is now part of FLINT [38] and is called through its Python interface.
Lemma 4.24. Let J≥0, and suppose that enclosures of e−ξi, K0(ξi) and K1(ξi) are given for 0≤i<i1. Start from the enclosure [0,V(64)] of H(ξ0). For i=0,1,…,i1−1, compute from the enclosure of H(ξi) enclosures of h0,…,hJ at ξi by the recurrences of Lemma 4.23(3), and take
j=0∑Jhj(−ξi/32)j+[0,ρJ(ξi)],
evaluated in interval arithmetic, as the enclosure of H(ξi+1). Then every interval so obtained contains the corresponding true value.
Proof. By Lemma 4.21, H(ξ0)∈[0,V(64)]. Suppose that the enclosure of H(ξi) contains H(ξi). The true Taylor coefficients at ξi are obtained from the true value H(ξi) by the same recurrence, so their enclosures contain them. Since ξi+1=31ξi/32, Lemma 4.23(2) shows that the enclosure of H(ξi+1) contains H(ξi+1). Induction on i proves the lemma. □
Now let t>0, ν∈{0,1}, integers a1,…,aN and J≥0 be given, with tN<ξ0 and t>ξi1. For 0≤i<i1 let Bi={n:1≤n≤N,ξi+1<tn≤ξi}. These sets partition {1,…,N}. For nonempty Bi, let nˉi be the integer part of the average of its least and largest elements, and put
Proof. Let n∈Bi. Then ∣tn−ξi∣<ξi/32, so by Lemma 4.23(2) the value H(tn) differs from ∑m=0Jhi,m(tn−ξi)m by at most ρJ(ξi). Since tn−ξi=t(n−nˉi)+(tnˉi−ξi), the binomial theorem gives ∑m=0Jhi,m(tn−ξi)m=∑j=0JPi,j(n−nˉi)j. Multiply by annν and sum over n∈Bi and over i. □
In the application below t≥2π/3470400>0.0033, because Qe≤3470400 for every e (Corollary 3.10). So every argument tn exceeds 0.0033, and the part of the grid below that point is not used.
The Value of ζE at σ=301/300
From now on σ=301/300. Let e∈F8∖{0}. By Proposition 3.8, L(s,χe) satisfies the hypotheses of Lemma 4.20 with Q=Qe and νk=νk(e). Put t=2π/Qe and
Ne=4⌊Qe⌋+1,
and let Πe be the prefactor of that lemma: Πe=tσ/Γ(σ) if ν1(e)=ν2(e), and Πe=4(t/2)σ+ν/Γ((σ+ν)/2)2 if ν1(e)=ν2(e)=ν; in the first case put ν=0. Let S1 and S2 be the sums ∑i∑j=020Pi,jMi,j of Lemma 4.25, with N=Ne, an=an(e) and J=20, for the two functions of the series of Lemma 4.20: H=Hσ and H=H1−σ in the first case, and H=Iσ+ν−1 and H=Iν−σ in the second; their envelope is V=V1 in the first case and V=V2 in the second. For every e, Ne lies between 105 and 7449, and t(Ne+1) between 24.7 and 25.5 (the supplementary program check255.gp prints both ranges). Hence tNe<ξ0 and t>ξi1, as Lemma 4.25 requires, and the hypotheses of Lemma 4.22 hold for every ω in the list 100101,2020,1011,89,56,45,34,23,2. The value of ω in this list is chosen to minimize the bound
Te=i∑ρ20(ξi)Ai+ζ(ω)2(Ne+1)ω+νV(t(Ne+1)).
Lemma 4.27.For σ=301/300 and every e∈F28∖{0},
∣L(σ,χe)−Πe(S1+S2)∣≤2ΠeTe.(24)
Proof. Split each of the two series of Lemma 4.20 at Ne. By Lemma 4.25, each finite part differs from its approximation S1 or S2 by at most ∑iρ20(ξi)Ai; the same bound serves both series, since their two functions have the same envelope. By Lemma 4.22, each tail is at most ζ(ω)2(Ne+1)ω+νV(t(Ne+1)). Multiplying by Πe gives (24). □
Proposition 4.29.For σ=301/300,
0.08264460177456<512logζE(σ)<0.08264460807138.
Proof (computer-assisted). By Proposition 3.7,
logζE(σ)=logζ(σ)+logL(σ,χB)+e=0∑logL(σ,χe).
The coefficients an(e) for n≤Ne are computed exactly from the Euler factors of Proposition 3.4 by the supplementary program lfun241.py. The supplementary program afe241.py evaluates (24) in Arb midpoint–radius interval arithmetic [15, 38] with 256 bits of precision, rounding every quantity outward. Its routines for Lemmas 4.20–4.25 are taken unchanged from an earlier supplementary archive of the author (Subsection 1.6), and they implement these lemmas as stated here. Every enclosure of L(σ,χe) is positive. The factor ζB(σ)=ζ(σ)L(σ,χB) is evaluated with
L(σ,χB)=241−σa=1∑240(241a)ζ(σ,241a),
where ζ(s,x)=∑n≥0(n+x)−s is the Hurwitz zeta function, using Arb’s rigorous implementation of it [14]. This gives ζB(σ)=725.51409164864… and L(σ,χB)=2.41373420241…. Summing the logarithms gives the enclosure
[0.08264460177456168…,0.08264460807137072…]
of logζE(σ)/512, which the proposition states rounded outward. So the inequalities follow from Lemmas 4.20–4.27, the exact coefficients, and outward-rounded interval arithmetic. □
As a check independent of the approximate functional equation, the supplementary programs yecheck.gp and lvalues255.gp evaluate the same values numerically with PARI/GP [41]: each of the 255 values L(σ,χe) lies in its enclosure, and PARI’s value 0.08264460492293… of logζE(σ)/512 lies in the enclosure above.
Frobenius Vectors and Lower Bounds for the Types.
Lemma 4.30.The set Σ2 of Definition 4.15 has exactly 21 elements: zero and the following twenty vectors, written as strings c0c1⋯c7 of their coordinates, as in Section 2:
In the notation of Table 4, these twenty vectors are the elementary images of the local involutions xj, yj and xjyj for j=1,2, of the nonidentity elements of the four local groups C2×C2 above 3 and 5, and of ι1 and ι2.
Proof. Each of these images lies in Σ2 by Lemma 2.25, since an involution has trivial square. The supplementary program census241.py computes Σ2 by linear algebra over F2 from the relations of Lemma 2.25, and finds no other vectors. □
Proposition 4.31 (Frobenius vectors of the primes of norm at most 106). Among the 78616 primes p∈/S∪{t0,t1,t2} of B with Np≤106, exactly 246 have Frobp=0, exactly 5994 have Frobp∈Σ2∖{0}, and the remaining 72376 form P4.
Proof. The supplementary program census241.py computes the Frobenius vector of each of these primes from Proposition 3.4(1) with exact modular arithmetic, and compares it with the elements of Σ2 (Lemma 4.30). □
Below we use this classification of the primes, not the three counts.
Proof. The supplementary program ceiling241.py uses the classification of Proposition 4.31 and evaluates (21) and (22) in Arb interval arithmetic. It gives Δ2,7,29(σ)=0.03445342991613… and ΔP4(σ)=0.00162984111205…, and the lemma states these values rounded down. □
where q1,q2 and r1,r2 are the primes above 3 and 5 (Lemma 2.4).
Lemma 4.35.For every K∈K and every prime p of B, eK(p)≥ep0 and fK(p)≥fp0. Thus (26) gives lower bounds for the types in K/B, as Proposition 4.13 requires.
Proof. At pj and tk the types in K/B are (8,4) and (1,4) by Theorem 2.42, as in the proof of Proposition 4.19. At qj and rj the type in E/B is (2,2) by Corollary 3.6, and if Frobp=0, then fE(p)=2 by the same corollary. Since E⊆K, the ramification indices and residue degrees in E/B divide those in K/B (Lemma 4.10). On P4 use Lemma 4.16. The bounds (1,1) hold trivially. □
Hence D=0.0085105513 satisfies the hypothesis of Proposition 4.13.
Proof. The primes of norm greater than 106 receive (1,1). At most two primes of B have a given norm, and x↦logx/(x2−1) decreases on [2,∞). So their contribution to the sum is at most
The supplementary program ceiling241.py sums the primes of norm at most 106 in interval arithmetic and adds this bound for the remaining primes; the result, 0.0085105512261…, is less than 0.0085105513. □
Proof of Theorem 4.2.
Corollary 4.38.Let σ=301/300. Every K∈K satisfies [K:Q]−1logζK(σ)≤Y∗, where
Proof. Combine Proposition 4.19, the upper bound of Proposition 4.29 rounded up, and the lower bounds of Lemma 4.32, which are rounded down. □
Proof of Theorem 4.2. Let σ=301/300 and ε=1/300. We apply Proposition 4.13 with Y∗=0.0465613371 (Corollary 4.38), with the lower bounds (26) for the types (Lemma 4.35), and with D=0.0085105513 (Lemma 4.36). Numerically κ∞=(ℓ−γ−log(4π))/4=0.63694120292…<0.6369412030, so
The interval evaluation of the same expression from the unrounded enclosures, by ceiling241.py, gives the upper bound 0.0487128429. Proposition 4.13 with C′=C proves the theorem. □
Remark 4.39. The primes above 2, 3, 5, 7 and 29 alone limit what Proposition 4.13 can give. Let (K(j)), βq and Z(s)=∑qβqgs(q) be as in the proof of that proposition. The absolute types of the primes of K(j) above 2,3,5,7,29 are known exactly (Theorem 2.42), so these primes contribute exactly
to Z(1), a value printed by ceiling241.py; all other contributions are nonnegative. The proof of Proposition 4.13 shows that Z(1)≤Y∗+ε(κ∞+D) for every choice of σ,Y∗,D and lower bounds for the types that satisfies its hypotheses. So no such choice gives an upper bound below this number. The bound C=0.04871285 exceeds it by 0.0070444.
Remark 4.40 (The hypothesis of the Lean formalization). The Lean formalization [22] proves, under one explicit hypothesis that it does not prove, that there are finite planar sets Uj with ∣Uj∣→∞ and u(Uj)/∣Uj∣1.0427→∞. The hypothesis is
In the formalization, E is the subfield of an algebraic closure of Q generated by 241 and square roots of α0,…,α7; it is isomorphic to the Kummer field E, and the formalization calls it the genus field.
The left side of (28) is the bound of this section with all local data taken from the fixed subfield E. By (19) for M=E, −(ζE′/ζE)(2)/512 is the sum defining D in Proposition 4.13 when the lower bounds are the types in E/B. These lower bounds are valid since E⊆K. The choice Y∗=logζE(σ)/512 is valid by Lemma 4.10. So the left side of (28) is the bound given by Proposition 4.13 without the refinements of Proposition 4.19 and with the types in E/B in place of those in K/B. The supplementary program h241_receipt.py bounds it by 0.0848335193<0.0852. It uses the upper endpoint in Proposition 4.29, the value of κ∞, and −(ζE′/ζE)(2)/512<0.0197321539, which it obtains by summing (19) in interval arithmetic over the primes of norm at most 106 and bounding the rest as in the proof of Lemma 4.36.
From (28), the formalization subtracts corrections at the primes above 2, 7 and 29 and at the primes above 41, 47, 53, 59, 61, 67, 79, 83 and 97, which lie in P4. It obtains the weaker bound 0.0495 for d−1logLF(1) for the fields of the tower in the formalization. This suffices for the exponent 1.0427 but not for 1.04273. The bound C=0.04871285 of Theorem 4.2 uses Proposition 4.31 and the lower bounds for the types in K/B; this refinement is not formalized.
Units of Relative Norm One and Planar Point Sets
This section and Sections 6 and 7 turn arithmetic data of a quadratic extension into planar point sets. Let K/F be a quadratic extension of number fields, let ι be the nontrivial automorphism of K over F, and let (b,c) be the signature of F. Throughout these three sections we assume:
(G1) K is totally imaginary, and b≥1;
(G2) K/F is unramified at every finite place;
(G3) rd(K)=rd(F)=λ.
Here the root discriminant rd and the discriminants ΔM are as in Section 2. By (G2), ∣ΔK∣=∣ΔF∣2, so (G3) amounts to rd(F)=λ. Put d=[F:Q]=b+2c, θ=c/d and ℓ=logλ. Then [K:Q]=2d, ∣ΔK∣=λ2d, ∣ΔF∣=λd and 0≤θ<1/2, where ΔK and ΔF are the absolute discriminants. Theorem 2.42 provides such extensions, with λ=29/43615 and θ≥θ∗; there ι is the complex conjugation ι1. The arguments of these three sections use no Galois structure over B or over Q. For a nonzero fractional ideal a of K or of F, Na denotes its absolute norm, so that N(xOK)=∣NK/Qx∣[18]. We take the unit theorem and the covolumes of ideal lattices from [18], and the Herbrand quotient, Hilbert’s Theorem 90 and the analytic class-number formula from [19].
The argument has five steps. Proposition 5.2 is the class-number formula for the units of relative norm one: it expresses their regulator, together with the relative class number and the capitulation kernel, through LF(1). Lemma 5.6 uses the split primes and one ideal class to produce many elements β1,…,βt of relative norm one in a single fractional ideal, with pairwise disjoint cosets βiOK1. Lemmas 5.9 and 5.13 average over the units of relative norm one and over translates of the ideal lattice; this counts exactly, on average, the pairs of lattice points in a window (Definition 5.11) that differ by the image of an element of ⋃iβiOK1. Lemmas 5.16 and 5.18 bound the number of lattice points in a window uniformly, by Poisson summation. Finally, Proposition 5.26, the geometric transfer, takes the window to be a sublevel set of the total energy of a pair of profiles (Definition 5.22 and Lemma 5.25) and combines these steps.
Coordinates
A real place of F extends to K either as two real places or as one complex place, and by (G1) the second case occurs. A complex place of F has two extensions to K, since C has no quadratic extension. Thus K has exactly d complex places. For each real place v of F, let ϕv:K→C be one of the two embeddings inducing the place of K above v. Number the complex places of F from 1 to c. For the j-th one fix an embedding ϕj+:K→C whose restriction τj to F induces it, and put ϕj−=ϕj+∘ι. The embeddings ϕj± induce the two places of K above this place: they are distinct, and they are not complex conjugate, because τj is not real. We identify K⊗QR with Cd through the map x↦x∞, where
x∞=((ϕv(x))v,(ϕj+(x),ϕj−(x))j),
write a∞={x∞:x∈a} for a⊆K, and identify Cd with R2d through real and imaginary parts. Volumes are Lebesgue measure in these coordinates, and ⟨x,ξ⟩=Re∑wxwξw is the real inner product, where w runs over the d coordinates. For a Borel set Ω, ∣Ω∣ denotes its volume; for a finite set, ∣⋅∣ denotes its cardinality.
At a real place v of F, the automorphism ι induces complex conjugation on the completion C of K, so ϕv(ιx)=ϕv(x). Writing v(y)∈R for the image of y∈F under the real embedding v, we obtain for x∈K
If NK/Fβ=1, then by (29)∣ϕv(β)∣=1 for every real place v, and ∣ϕj±(β)∣=e±lj(β) for every j. Accordingly the b coordinates ϕv are called the compact coordinates: there the elements of relative norm one lie on the unit circle. The c pairs (ϕj+,ϕj−) are the pair coordinates: there they have reciprocal moduli. A compact coordinate, or a pair of coordinates, is called a block. The planar sets will be images under one compact coordinate, which is injective on K and sends every element of relative norm one to a unit vector. We count ordered pairs of points at distance one, and divide by two at the end.
The Class-Number Formula for Units of Relative Norm One
Let OK× and OF× be the unit groups, and put
OK1=ker(NK/F:OK×→OF×),IN=[OF×:NK/F(OK×)].
Let wK be the number of roots of unity in K, let hK and hF be the class numbers of K and F, put hrel=hK/hF, let κ be the order of the capitulation kernelker(Cl(F)→Cl(K)), the map being induced by extension of ideals, and let LF(s)=ζK(s)/ζF(s). At s=1, LF(1) denotes the quotient of the residues of ζK and ζF. The regulators RegK and RegF are the ordinary ones, which use twice the logarithm of the absolute value at a complex place [18].
Proposition 5.2.The roots of unity of K lie in OK1, and the restriction of l to OK1 has kernel of order wK and image a full lattice in Rc. Let Reg1 be the covolume of this lattice, with Reg1=1 if c=0. Then
κ=IN/2b−1,RegK/RegF=2c−1INReg1,
and consequently
hrelκReg1wK=λd/2LF(1)2d−1πb+c.(30)
Proof. Signs. Fix a real place v0 of F. By (29), v0(NK/Fx)=∣ϕv0(x)∣2>0 for x∈K×, so −1 is not a norm from K×. If ζ∈K is a root of unity, then NK/Fζ is a root of unity of F, hence ±1 because F has a real place, and it is positive at v0. Thus NK/Fζ=1.
Their images are lattices in the hyperplanes of coordinate sum zero, the kernel of LogF is {±1}[18], and RegK and RegF are the covolumes of the images after one coordinate is deleted; the choice of the deleted coordinate does not matter.
The lattice l(OK1). Let β∈OK1. Then ∣ϕv(β)∣=1 and ∣ϕj±(β)∣=e±lj(β), so
LogK(β)=(0,(2lj(β),−2lj(β))j).
Hence l(β)=0 exactly when every conjugate of β has absolute value one, that is, when β is a root of unity [18]. By the first step the kernel of l on OK1 has order wK. The norm of y∈OF× is y2, so NK/F(OK×) has finite index in OF× and rank b+c−1. By the unit theorem [18], OK× has rank d−1, so OK1 has rank c. As LogK(OK1) is discrete, so is l(OK1), and it is a full lattice in Rc.
The capitulation kernel. The Herbrand quotient h(M)=∣H0(⟨ι⟩,M)∣/∣H1(⟨ι⟩,M)∣[19] of the units satisfies 2h(OK×)=∏v[Kw:Fv], where v runs over the archimedean places of F and w lies above v[19]. Each real place contributes 2 and each complex place 1, so h(OK×)=2b−1. Since H0(⟨ι⟩,OK×)=OF×/NK/F(OK×) has order IN, we get ∣H1(⟨ι⟩,OK×)∣=IN/2b−1. This group is OK1 modulo the units γ/ιγ with γ∈OK×. If a is a fractional ideal of F with aOK=xOK, the class of the unit x/ιx depends neither on the choice of x nor on the choice of a in its ideal class. This defines a homomorphism from ker(Cl(F)→Cl(K)) to H1(⟨ι⟩,OK×). It is injective: if x/ιx=γ/ιγ, then y=x/γ lies in F and aOK=yOK, so a=yOF by unique factorization. It is surjective: by Hilbert’s Theorem 90 [19], Chapter II, Corollary 1.23, an element of OK1 has the form x/ιx with x∈K×, and then the fractional ideal xOK=(ιx)OK is invariant under ι. By (G2), every ι-invariant fractional ideal of K is extended from F: at a split prime PιP invariance forces equal exponents, an inert prime of F stays prime, and no prime ramifies. Hence κ=IN/2b−1.
Regulators. Delete the coordinate of v0 from both logarithmic embeddings, and let UK⊂Rd−1 and UF⊂Rb+c−1 be the resulting lattices, of covolumes RegK and RegF. By (29), LogF(NK/Fx) is obtained from LogK(x) by keeping the coordinates of the real places and adding the two coordinates of each pair. This induces a linear map T:Rd−1→Rb+c−1 such that T(UK) is the image of NK/F(OK×) in UF. The substitution (xj,yj)↦(xj,xj+yj) on each pair has determinant one and turns T into a coordinate projection whose kernel consists of the c coordinates xj. The covolume of a lattice in a product of two coordinate spaces is the covolume of its intersection with the first factor times the covolume of its projection to the second, whenever both are lattices.
A unit x∈OK× lies in the kernel of T exactly when LogF(NK/Fx)=0, because the deleted coordinate is minus the sum of the others. Then NK/Fx=±1, and NK/Fx=1 by the first step. So the intersection of UK with the kernel is the image of OK1, which after the substitution is 2l(OK1), of covolume 2cReg1. Since −1∈/NK/F(OK×), the projection T(UK) has index [OF×:{±1}NK/F(OK×)]=IN/2 in UF. Hence RegK=2cReg1⋅RegFIN/2.
The class-number formula. The analytic class-number formula [19], Chapter V, Theorem 2.4 (compare (13)), applied to K (with r1=0 and r2=d) and to F (with r1=b, r2=c and wF=2), gives
Definition 5.4. Let R be a finite set of rational primes. The extension K/F has uniform local types(er,fr)r∈R if, for every r∈R, every prime of F above r splits in K and has absolute type (er,fr) (Definition 2.41).
If K/F has uniform local types (er,fr)r∈R, then F has exactly d/(erfr) primes above r, each of norm rfr, since the products epfp of the absolute ramification indices and residue degrees of the primes p of F above r add up to [F:Q]=d. Given integers kr≥0, define
H=r∈R∑erkrlogr,J=r∈R∑erfrlog(kr+1).(31)
Lemma 5.6.Assume that K/F has uniform local types (er,fr)r∈R, and fix integers kr≥0. There are a fractional ideal I of K with N(I)=e−dH, an integer t≥edJ/(κhrel), and elements β1,…,βt∈I of relative norm one whose cosets βjOK1 are pairwise disjoint.
Proof. For each prime p of F above a prime r∈R, write pOK=PpιPp and kp=kr. For integer vectors j=(jp) with 0≤jp≤kp put
Aj=p∏Ppjp,Ij=p∏Ppjp(ιPp)kp−jp.
There are ∏p(kp+1)=edJ such vectors. The quotient of Cl(K) by the image of Cl(F) has order hKκ/hF=κhrel, so one class of this quotient contains the classes of Aj for a set of t≥edJ/(κhrel) vectors j. Fix one of them, j0, and put I=Jj0−1. For each j in the set, write AjAj0−1=xaOK with x∈K× and a a fractional ideal of F, and put βj=x/ιx. Then NK/Fβj=1, and since ι fixes aOK,
βjOK=AjAj0−1(ιAj)−1ιAj0=JjI.
As Jj is integral, βj∈I. The ideals JjI are distinct for distinct j, and every element of βjOK1 generates JjI, so the cosets are disjoint. Finally, since each p splits in K, N(I)−1=∏p(Np)kp=∏rrfrkrd/(erfr)=edH. □
Unit Averaging and Translation
In this subsection and the next, I is a fractional ideal of K and β1,…,βt∈I are elements of relative norm one whose cosets βiOK1 are pairwise disjoint; Lemma 5.6 provides such data. For h∈Rc let Wh:Cd→Cd multiply the coordinate ϕj+ by e−hj and ϕj− by ehj, and fix the compact coordinates. It is a real diagonal map of determinant one. Put Λh=WhI∞. For a lattice Λ⊂Cd, covol(Λ) denotes the volume of a fundamental domain. By [18], Proposition 4.26, extended to fractional ideals by scaling and [18], Proposition 4.2,
covol(a∞)=2−d∣ΔK∣1/2Na(32)
for every fractional ideal a of K. Hence covol(Λh)=2−dλdN(I), independently of h; for the ideal of Lemma 5.6,
covol(Λh)=2−dλde−dH.(33)
If NK/Fβ=1, then Whβ∞ has compact coordinates of modulus one and pair coordinates of moduli e±(lj(β)−hj). Let Π1 be a fundamental parallelepiped of the lattice l(OK1); its volume is Reg1. When c=0, integrals over R0 and over Π1 are evaluations at the single point.
Proof. The map l is a homomorphism on K×, so l(βiOK1)=l(βi)+l(OK1), and by Proposition 5.2 each point of this translate of l(OK1) is the image of exactly wK elements of βiOK1. The translates Π1−x, with x∈l(OK1), tile Rc, so by Tonelli’s theorem ∫Π1∑x∈l(OK1)Φ(l(βi)+x−h)dh=∫RcΦ(u)du for each i. □
Definition 5.11. A window is a bounded Borel set Ω⊂Cd of positive volume that is invariant under rotation of each coordinate separately. For u∈Rc let ν(u)∈Cd have compact coordinates 1 and pair coordinates (euj,e−uj), and for a Borel set Ω⊂Cd put
ΦΩ(u)=∣Ω∩(Ω−ν(u))∣.
Lemma 5.12.LetΩbe a window, leth∈Rc, and letβ∈KsatisfyNK/Fβ=1. Then∣Ω∩(Ω−Whβ∞)∣=ΦΩ(l(β)−h).
Proof. If x∈Cd has compact coordinates of modulus one and pair coordinates of moduli e±uj, then a rotation of the coordinates preserves Ω and maps ν(u) to x, so ∣Ω∩(Ω−x)∣=ΦΩ(u). By the remark before Lemma 5.9, x=Whβ∞ is such a point, with u=l(β)−h. □
For a window Ω, the function ΦΩ is continuous, bounded by ∣Ω∣, and vanishes when some e∣uj∣ exceeds the diameter of Ω: if x and x+ν(u) both lie in Ω, then e∣uj∣, the modulus of a coordinate of their difference ν(u), is at most the diameter of Ω.
Let Ω be a window. Fix a Z-basis of I∞, and for y∈[0,1)2d let yI∈Cd be the point with coordinate vector y in this basis. For h∈Rc let X(h,y)=(Λh+WhyI)∩Ω and n(h,y)=∣X(h,y)∣, and let E(h,y) be the number of ordered edges of X(h,y), that is, of pairs (x,β) with x∈X(h,y), β∈⋃iβiOK1 and x+Whβ∞∈Ω. Since β∈I, the point x+Whβ∞ then lies in X(h,y).
Lemma 5.13.Let Ω be a window and suppose n(h,y)≤nmax<∞ for all (h,y). Then there is a pair (h,y) with n(h,y)>0 and
n(h,y)E(h,y)≥Reg1twK∣Ω∣∫RcΦΩ(u)du.(35)
Proof. For fixed x, distinct β give distinct points x+Whβ∞=x, so E≤n2≤nmax2. Both n and E are countable sums of indicator functions of Borel sets in (h,y). For fixed h, the map y↦WhyI has Jacobian covol(Λh) and maps [0,1)2d onto a fundamental domain of Λh, so tiling and Lemma 5.12 give
Average over h∈Π1 with respect to dh/Reg1 and apply (5.10). For the probability measure dhdy/Reg1 on Π1×[0,1)2d this gives EE=ρEn and En=∣Ω∣/covol(Λh)>0, where ρ is the right side of (35). If E<ρn held wherever n>0, then, since E=0 where n=0 and {n>0} has positive measure, we would get EE<ρEn. □
Lemma 5.15.Let Ω be a window, let (h,y) be a pair, and let v0 be a real place of F. For x∈Cd let xv0 be its compact coordinate at v0, and let U={xv0:x∈X(h,y)}⊂C=R2. Then ∣U∣=n(h,y) and u(U)≥E(h,y)/2.
Proof. Two distinct points of X(h,y) differ by Whγ∞ with 0=γ∈I, whose v0-coordinate ϕv0(γ) is nonzero, so ∣U∣=n(h,y). An ordered edge (x,β) gives two points of U at distance ∣ϕv0(β)∣=1, and distinct ordered edges give distinct ordered pairs of points. Hence u(U)≥E(h,y)/2. □
Uniform Lattice Counting
For a lattice Λ⊂Cd let Λ∗={ξ∈Cd:⟨ξ,x⟩∈Z for all x∈Λ} be its dual lattice, and let ∥ξ∥∗=∑w∣ξw∣, a norm on Cd. We use the Fourier transform Ψ(ξ)=∫CdΨ(x)e−2πi⟨x,ξ⟩dx.
Lemma 5.16.Let a be a fractional ideal of K and h∈Rc. Every nonzero ξ in the dual lattice of Wha∞ satisfies
w∏∣ξw∣≥λd(Na)1/22d,∥ξ∥∗≥λ(Na)1/(2d)2d.
In particular, if N(I)=e−dH, as for the ideal of Lemma 5.6, every nonzero dual vector of Λh satisfies
w∏∣ξw∣≥μd,μ=2eH/2−ℓ,∥ξ∥∗≥dμ.(36)
Proof. First let h=0. For x,y∈K we have ⟨x∞,y∞⟩=Re∑wϕw(xy)=21TrK/Q(xy), where the bar is coordinatewise complex conjugation and ϕw is the embedding of the coordinate w, because these embeddings and their complex conjugates are all the embeddings of K. Let a∨={y∈K:TrK/Q(ya)⊂Z}. It is an OK-module, and since the trace form is nondegenerate, the dual basis of a Z-basis of a is a Z-basis of a∨. So a∨ is a fractional ideal, and the dual lattice of a∞ is (2a∨)∞. Dual lattices have reciprocal covolumes, and complex conjugation preserves volume. With (5.7) and N(2a∨)=22dN(a∨), this gives N(a∨)=λ−2d(Na)−1. If 0=y∈a∨, then yOK⊂a∨ and ∣NK/Qy∣=N(yOK)≥N(a∨)[18]. For ξ=(2y)∞,
w∏∣ξw∣=∣NK/Q(2y)∣1/2≥(22dλ−2d(Na)−1)1/2.
For general h, the dual lattice of Wha∞ is the image of the dual lattice of a∞ under W−h, because Wh is real diagonal, and W−h preserves ∏w∣ξw∣. The bound for ∥ξ∥∗ follows from the arithmetic-geometric mean inequality. For (36) take a=I. □
Lemma 5.18.Let Λ⊂Cd be a lattice such that ∏w∣ξw∣≥μd for every nonzero ξ∈Λ∗, where μ>0. Let Ψ:Cd→[0,∞) be smooth, bounded and integrable, with ∫Ψ>0 and all partial derivatives bounded, and suppose that
∣Ψ(ξ)∣≤Mde−σ∥ξ∥∗∫Ψ(ξ∈Cd)
for some M≥1 and σ>0. If the Fourier condition
σμ≥logM+2log5+2(37)
holds, then for every y∈Cd,
x∈Λ+y∑Ψ(x)≤covol(Λ)2∫Ψ.(38)
Proof. A nonzero ξ∈Λ∗ has no zero coordinate, and the arithmetic-geometric mean inequality gives ∥ξ∥∗≥dμ. Hence distinct points of Λ∗ are at ∥⋅∥∗-distance at least dμ, and the open ∥⋅∥∗-balls of radius dμ/2 about them are disjoint. For j≥1, the balls about the points with jdμ≤∥ξ∥∗<(j+1)dμ lie in the ball of radius (j+3/2)dμ about 0. Comparing volumes in R2d bounds the number of these points by (2j+3)2d≤52dj. By the Fourier condition (37) and M≥1,
For ε>0 put Ψε(x)=Ψ(x)e−ε∣x∣2, a Schwartz function. Its Fourier transform is Ψ∗Gε, where Gε(ζ)=(π/ε)de−π2∣ζ∣2/ε is a probability density on R2d. Hence ∣Ψε(ξ)∣≤mεMde−σ∥ξ∥∗∫Ψ, with mε=∫eσ∥ζ∥∗Gε(ζ)dζ. The function y↦∑x∈ΛΨε(x+y) is smooth and Λ-periodic, its Fourier coefficient at ξ∈Λ∗ is Ψε(ξ)/covol(Λ), and these coefficients are absolutely summable. This gives the Poisson summation formula
x∈Λ∑Ψε(x+y)=covol(Λ)1ξ∈Λ∗∑Ψε(ξ)e2πi⟨ξ,y⟩.
Using Ψε(0)=∫Ψε and the bound above, the sum is at most (∫Ψε+mε∫Ψ)/covol(Λ). As ε→0, we have mε→1 by dominated convergence, since Gε is the image of G1 under ζ↦εζ; moreover Ψε increases to Ψ. Monotone convergence gives (38). □
The next bound is used in Section 7 for the Gaussian real-place profile.
Lemma 5.21.If 0<aR≤1 and ΨR(z)=e−aR∣z∣2 on C, then ΨR(ξ)=(π/aR)e−π2∣ξ∣2/aR and ∣ΨR(ξ)∣≤2e−∣ξ∣∫ΨR.
Proof. The formula is the Gaussian integral, and ∫ΨR=π/aR. Since aR≤1, it suffices that π2x2−x+log2≥0 for x≥0, and the minimum of the left side is log2−1/(4π2)>0. □
Profiles at Real and Complex Places.
Definition 5.22. Fix 0<δ<1 and put p=2/(1+δ), the Lebesgue exponent (not a prime), so that 1<p<2 and p(1+δ)=2. A real-place profile is a positive radial function fR on C, used at the compact coordinates, and a complex-place profile is a positive function gC on C2 that depends only on ∣z∣ and ∣w∣, used at the pair coordinates. Their massesAR, AC, overlapsIR, IC and functionalsJR, JC are
AR=∫CfRp,IR=∫CfR(z)fR(z+1)dz,AC=∫C2gCp,
IC=∫R∫C2gC(z,w)gC(z+eu,w+e−u)dzdwdu,
JR=logIR−(1+δ)logAR,JC=logIC−(1+δ)logAC.
When these integrals are finite and positive, the endpoint laws are the probability densities fR(z)fR(z+1)/IR on C and gC(z,w)gC(z+eu,w+e−u)/IC on C2×R. The functions VR=−logfR and VC=−loggC are the energies of the profiles.
The exponent p is chosen so that the terms in the mean energy cancel at the end of the proof of Proposition 5.26.
Definition 5.23. Let M≥1 and σ>0. The pair (fR,gC) is admissible with Fourier constants(M,σ) if:
(P1) fRp and gCp are smooth and bounded, and all their partial derivatives are bounded;
(P2) AR, IR, AC and IC are finite and positive;
(P3) VR and VC are bounded below and coercive, that is, their sublevel sets are bounded;
(P4) VR(z) and VC(z,w) have finite second moments under the endpoint laws;
(P5) ∣fRp(ξ)∣≤Me−σ∣ξ∣AR and ∣gCp(ξ1,ξ2)∣≤M2e−σ(∣ξ1∣+∣ξ2∣)AC for all ξ,ξ1,ξ2∈C.
Section 7 proves admissibility for the profiles we use.
Lemma 5.24.Let (fR,gC) be admissible with Fourier constants (M,σ), and let Ψ=∏vfRp∏jgCp be the product function on Cd, with one factor for each block. Then Ψ satisfies the hypotheses that Lemma 5.18 places on Ψ, other than the Fourier condition (5.19): it is smooth and bounded, with bounded derivatives, ∫Ψ=ARbACc, and
∣Ψ(ξ)∣≤Mde−σ∥ξ∥Ψ∗∫Ψ(ξ∈Cd).
Proof. The first properties follow from (P1) and (P2). The Fourier transform of Ψ is the product of the transforms of the factors, so (P5) gives ∣Ψ(ξ)∣≤Mb+2ce−σ∥ξ∥Ψ∗∫Ψ=Mde−σ∥ξ∥Ψ∗∫Ψ. □
Lemma 5.25.Let fR and gC be profiles satisfying (P1), (P2) and (P4), and let η>0. Let mR and mC be the means of VR(z) and VC(z,w) under the endpoint laws, and let VarR and VarC be their variances. For x∈Cd with compact coordinates xv and pair coordinates (xj+,xj−), let V(x)=∑vVR(xv)+∑jVC(xj+,xj−) be the total energy, put mˉ=(bmR+cmC)/d, and let
Ω={x∈Cd:V(x)≤d(mˉ+η)}.
Then, for all d larger than a bound that depends only on η, VarR and VarC,
∫RcΦΩ(u)du≥21IRbICce2d(mˉ−η).
Proof. By (P1), fR and gC are continuous, so V is continuous and Ω is closed, hence a Borel set. The function (x,u)↦e−V(x)e−V(x+ν(u))/(IRbICc) is the product of the endpoint densities of the blocks, hence a probability density on Cd×Rc. Under it, the energies V(x) and V(x+ν(u)) of the two endpoints are sums of independent terms, one for each block, and V(x) has mean dmˉ. On a compact block the map z↦−z−1, and on a pair block the map (z,w,uj)↦(−z−euj,−w−e−uj,uj), preserves the endpoint density and carries the first endpoint to minus the second. Since the profiles are even, V(x+ν(u)) has the same law as V(x). By Chebyshev’s inequality, each of the two energies differs from dmˉ by more than ηd with probability at most max(VarR,VarC)/(η2d). So the set Y⊂Cd×Rc where both energies lie in [d(mˉ−η),d(mˉ+η)] has probability at least 1/2 for all d larger than a bound that depends only on η, VarR and VarC. On Y both endpoints lie in Ω and e−V(x)e−V(x+ν(u))≤e−2d(mˉ−η). Therefore, with ∣Y∣ the Lebesgue measure of Y,
Proposition 5.26 (Geometric transfer). Let 0<δ<1 and C∈R. Let (Ki/Fi)i≥1 be quadratic extensions satisfying (G1)–(G3) with a common λ and with di=[Fi:Q]→∞, such that logLFi(1)≤diC and such that every Ki/Fi has the same uniform local types (er,fr)r∈R. Fix integers kr≥0, let H and J be as in (31), and let μ=2eH/2−ℓ. Let (fR,gC) be admissible with Fourier constants (M,σ) satisfying the Fourier condition (37). Put
We call M(θ) the margin. If infiM(θi)>0, where θi is the value of θ for Fi, then there are finite sets Ui⊂R2 with ∣Ui∣→∞ and u(Ui)/∣Ui∣1+δ→∞.
Proof.The window. Let mR, mC, VarR and VarC be as in Lemma 5.25; the variances are finite by (P4). Fix η>0 with 4η<infiM(θi), and consider one extension K/F of the sequence, with d large. Take I, t and β1,…,βt from Lemma 5.6, so that Λh is defined; the counts n(h,y) and E(h,y) are those of the window Ω chosen next. With the total energy V and the mean mˉ of Lemma 5.25, let
Ω={x∈Cd:V(x)≤d(mˉ+η)}.
By (P1) and (P3) (see the proof of Lemma 5.25), this is a bounded Borel set, and it is invariant under coordinate rotations because the profiles are radial in each coordinate. By Lemma 5.25,
∫RcΦΩ(u)du≥21IRbICce2d(mˉ−η)
for all large d, uniformly in the signature. In particular ∣Ω∣>0 for large d, so Ω is a window.
Counting points. With the product function Ψ=∏RfRp∏CgCp=e−pV of Lemma 5.24, we have 1Ω≤epd(mˉ+η)Ψ, so
∣Ω∣≤epd(mˉ+η)ARbACc,
and Lemmas 5.16, 5.18 and 5.24 give, for all (h,y),
Combined with the bounds for ∣Ω∣ and for the integral of ΦΩ, (35) shows that the set X=X(h,y) has n=∣X∣>0 points and at least 41edXηn ordered edges, where
Since n≤2edYη and p(1+δ)=2, the terms in mˉ cancel, and
n1+δE(h,y)≥4⋅2δeδdYηedXη=2−2−δed(M(θ)−4η).
The planar set. Finally let v0 be a real place of F, and let U={xv0:x∈X}⊂C=R2. By Lemma 5.15, ∣U∣=n and u(U)≥E(h,y)/2, so
∣U∣1+δu(U)≥2−3−δed(M(θ)−4η),
which tends to infinity along the sequence. Since u(U)≤∣U∣2/2 and δ<1, also ∣U∣→∞. These sets are the Ui. The conclusion concerns a sequence of cardinalities, not every large cardinality. □
Shell Profiles at Split Finite Places
We now place locally constant weights at finitely many finite places of F that split in K; in Section 8 these are the places above 2, 3, 5, 7 and 29. In Lemma 5.6, the count at each prime of F above r∈R corresponds to the indicator function of a product of two balls in the completion, the unweighted shell profile of Definition 6.5; shell profiles replace this indicator by weighted sums of indicators of products of shells (Corollary 6.23). We keep the hypotheses (G1)–(G3) and the notation of Section 5, and put
covol0=covol((OK)∞)=(λ/2)d.
The Local Functional
Let L be a nonarchimedean local field with valuation ring O, uniformizer ϖ and residue field of cardinality Q. We use the additive Haar measure on L with ∣O∣=1, the product measure on the plane L2=L×L, and the multiplicative Haar measure dω on O× of total measure one. For n∈Z and ω∈O× put
The two coordinates of s(n,ω) have product one; they model the finite coordinates of an element of relative norm one at a split place. If g1 and g2 are locally constant and compactly supported on L2, only finitely many n contribute to ∫L2g1Tg2: if g1(z)g2(z+s(n,ω))=0, then both coordinates of s(n,ω) lie in the compact set suppg2−suppg1, which bounds n from above and from below.
Definition 6.1. For locally constant, compactly supported functions g1,g2 on L2 put ⟨g1,g2⟩T=∫L2g1Tg2. For a nonnegative, locally constant, compactly supported function g on L2 with g=0, and 0<δ<1, p=2/(1+δ), put
A(g)=∫L2gp,Z(g)=⟨g,g⟩T,Fδ(g)=A(g)1+δZ(g).(40)
We call Fδ the local functional.
For i∈Z let Bi=ϖ−iO, a ball of measure Qi. Write [x]+=max(x,0).
Proof. Two balls of an ultrametric space are disjoint or nested. Hence Bi1∩(Bi2−x) has measure Qmin(i1,i2) if x∈Bmax(i1,i2), and is empty otherwise. The left side of (41) is therefore Qmin(i1,i2)+min(j1,j2) times the measure of the set of (n,ω) with ϖnω∈Bmax(i1,i2) and ϖ−nω−1∈Bmax(j1,j2). These conditions say −max(i1,i2)≤n≤max(j1,j2) and do not involve ω, and dω has total measure one. □
Definition 6.5. Put S0=O and Si=Bi∖Bi−1 for i≥1. These are the shells, and their measures m0=1 and mi=Qi−Qi−1 are the shell measures. Let k≥0 be an integer, and let w=(wij)i,j≥0 be nonnegative weights, finitely many of them nonzero, with w00>0. The shell profile with parameters (Q,k,w) on L2 is
g=i,j≥0∑wij1Si×ϖ−kSj.
The shell profile with w00=1 and wij=0 otherwise, that is, the indicator function of O×ϖ−kO, is the unweighted shell profile with parameters (Q,k).
Lemma 6.6.A shell profile g with parameters (Q,k,w) is nonnegative, locally constant and compactly supported. It is invariant under (x,y)↦(ωx,ω′y) for all ω,ω′∈O×; in particular it is even. It is also invariant under translation by the group
C=O×ϖ−kO,∣C∣=Qk.(42)
Proof. The first assertion holds because the shells are compact open sets and finitely many of the nonnegative weights are nonzero. Multiplication by a unit preserves valuations and hence each shell. Adding an element of O does not change the shell of a point, whatever the number of shells; this gives the invariance under C, applied to the first coordinate and to ϖk times the second. □
We call C the period of g. Define the symmetric matrix Δ=(Δii′)i,i′≥0 by
All entries of G are nonnegative, and Z(g)≥(k+1)Qkw002>0. In particular Fδ(g) depends only on the parameters, and the unweighted shell profile has Fδ(g)=(k+1)Q−δk.
Proof. The formula for A(g) holds because the sets Si×ϖ−kSj are disjoint of measure Qkmimj. For Z(g), write 1S0=1B0 and 1Si=1Bi−1Bi−1 for i≥1, and note ϖ−kBj=Bj+k. Only balls with nonnegative indices occur, so by (41) and k≥0 the positive part can be dropped, and ⟨1Bi1×Bj1+k,1Bi2×Bj2+k⟩T, for i1,j1,i2,j2≥0, is
For a function φ of two nonnegative integers put (∂φ)(i,i′)=φ(i,i′)−φ(i−1,i′)−φ(i,i′−1)+φ(i−1,i′−1), where terms with an index −1 are omitted; this is the expansion of the shells into balls. Apply ∂ in (i1,i2) and in (j1,j2). A direct computation gives, for i,i′≥0, that ∂Qmin(i1,i2) at (i,i′) is mi1i=i′, and that ∂(max(i1,i2)Qmin(i1,i2)) is Δii′. For example, when 1≤i<i′ the latter is i′Qi−i′Qi−1−(i′−1)Qi+(i′−1)Qi−1=mi, and when i=i′≥1 it is iQi−2iQi−1+(i−1)Qi−1=imi−Qi−1. This gives ⟨1Si×ϖ−kSj,1Si′×ϖ−kSj′⟩T=QkG(i,j),(i′,j′), and hence the formula for Z(g). All entries of G are nonnegative, since Δii=Qi−1(iQ−i−1)≥0 for i≥1; as G(0,0),(0,0)=k+1, this gives Z(g)≥(k+1)Qkw002. For the unweighted shell profile only G(0,0),(0,0)=k+1 contributes. □
By Lemma 6.8, a shell profile g has Z(g)>0, so g(z)g(z+s(n,ω))/Z(g) is a probability density on L2×Z×O×, with counting measure on Z. As at the archimedean places (Definition 5.22), we call it the endpoint law of g; its first endpoint is z and its second endpoint is z+s(n,ω).
Corollary 6.10. Let x=(xi)i≥0 and x′=(xj′)j≥0 be finitely supported vectors of nonnegative numbers with x0x0′>0, put
Remark 6.12. Suppose that only S0 and S1 carry weight, with x=x′=(1,t) and t≥0. Then Sx=1+(Q−1)t2 and Rx=2t+(Q−2)t2, and the difference between logFδ(g) and the value log((k+1)Q−δk) of the unweighted shell profile is
Let P be a finite set of finite places of F that split in K. For p∈P write pOK=PpιPp, where we also write p for the prime ideal of F. The completion of K at Pp is Fp; for x∈K let xp∈Fp be its image under the completion map. The map x↦(xp,(ιx)p) records the completions at Pp and at ιPp, and ι exchanges the two coordinates. If NK/Fβ=1, then βp(ιβ)p=1. Let Op, ϖp and Qp be the valuation ring, a uniformizer and the residue cardinality of Fp.
Definition 6.14. Put
AP=p∈P∏(Fp×Fp),
with the Haar measure that gives ∏p(Op×Op) measure one. For a point (z,ζ) of Cd×AP we call z its archimedean part and ζ its finite part. Let Γ=OK,P be the ring of elements of K that are integral at every prime other than the Pp and ιPp. We embed Γ in Cd×AP by γ↦(γ∞,γf), where γf=(γp,(ιγ)p)p; thus γ∞ and γf are the archimedean part and the finite part of the image of γ. A subgroup of AP of the form
Cf=p∈P∏(ϖpapOp×ϖpbpOp),ap,bp∈Z,
is called rectangular.
A rectangular group has measure ∣Cf∣=∏pQp−ap−bp.
Lemma 6.15. Let Cf be rectangular and put L=∏pPpap(ιPp)bp.
(a) The elements γ∈Γ with γf∈Cf form L, and NL=∣Cf∣−1, so covol(L∞)=covol0/∣Cf∣.
(b) Every coset of Cf in AP contains γf for some γ∈Γ.
(c) Γ is a lattice in Cd×AP, that is, a discrete subgroup with a fundamental domain of finite measure. If PL is a fundamental parallelepiped of L∞, then PL×Cf is a fundamental domain for Γ, of measure covol0.
(d) For h∈Rc, the set Γh={(Whγ∞,γf):γ∈Γ} is again a lattice, with fundamental domain WhPL×Cf of measure covol0. Put Hf=d−1log∣Cf∣. Every nonzero vector ξ of the dual lattice of WhL∞ satisfies ∏v∣ξv∣≥μf, where μf=2eHf/2−ℓ.
Proof. (a) The valuation of γp is the exponent of Pp in γ, and that of (ιγ)p is the exponent of ιPp. Together with integrality at the other primes, γf∈Cf says exactly that γ∈L. As NPp=N(ιPp)=Qp, we get NL=∏pQpap+bp=∣Cf∣−1, and the covolume follows from the covolume formula (5.7).
(b) Choose ap≤ap and bp≤bp such that the given coset lies in Cf=∏p(ϖp−apOp×ϖp−bpOp), and let L⊃L be the corresponding ideal. The map γ↦γf induces a homomorphism L/L→Cf/Cf. It is injective by (a), applied to Cf. Both groups are finite of the same order: ∣Cf/Cf∣=∣Cf∣/∣Cf∣, and ∣L/L∣=NL/NL[18], which is the same number by (a). So the map is bijective, and every coset of Cf in Cf contains γf for some γ∈L⊂Γ. No principality is needed.
(c) Given (z,ζ)∈Cd×AP, by (b) there is γ∈Γ with ζ−γf∈Cf, and then a unique γ′∈L with z−(γ+γ′)∞∈PL. If (z,ζ) had two such representations, the difference of the two elements of Γ would have finite part in Cf, hence lie in L by (a), and then it would vanish because PL is a fundamental domain. Thus PL×Cf is a fundamental domain, of measure (covol0/∣Cf∣)⋅∣Cf∣=covol0. Finally Γ is discrete: for a bounded open set Y∋0 in Cd, its intersection with the neighborhood Y×Cf of 0 consists of the finitely many γ∈L with γ∞∈Y.
(d) The map (z,ζ)↦(Whz,ζ) preserves measure, because Wh is real diagonal of determinant one, and it carries Γ to Γh and PL×Cf to WhPL×Cf; so the first assertion follows from (c). By (a), NL=e−dHf, and Lemma 5.16 gives the bound for the dual vectors. □
Accordingly, the Fourier condition used in this section is
σ⋅2eHf/2−ℓ≥logM+2log5+2.(46)
Lemma 6.17.Let (fR,gC) be admissible with Fourier constants (M,σ), and let Ψ∞=∏vfRp∏jgCp on Cd, with one factor for each real place v and each complex place j of F (Lemma 5.24). Let Ψf≥0 be a function on AP that is invariant under translation by a rectangular group Cf and vanishes outside a finite union of its cosets. If (46) holds, with Hf defined by this Cf, then for every h∈Rc and every z0∈Cd×AP (a point of the whole space, not an archimedean part),
Proof. Group the points by the coset of Cf containing their finite part. If two points of Γh+z0 have finite parts in the same coset, their difference comes from an element of L, by Lemma 6.15(a). So the archimedean parts of the points in one coset form a translate of WhL∞, or the empty set. By Lemma 5.18, with the dual bound of Lemma 6.15(d) and the Fourier bound for Ψ∞ of Lemma 5.24, the sum of Ψ∞ over such a translate is at most 2∣Cf∣∫Ψ∞/covol0. The values of Ψf on its finitely many nonzero cosets sum to ∫Ψf/∣Cf∣. □
Remark 6.19. Let α∈AP have coordinates (ϖpnp,ϖp−np)p. Multiplication by α preserves the Haar measure, and for a rectangular group Cf the group αCf is rectangular, of the same measure. Hence, if Ψf satisfies the hypotheses of Lemma 6.17 with Cf, then Ψf∘α satisfies them with the rectangular group α−1Cf, which gives the same Hf, and ∫Ψf∘α=∫Ψf.
Class Selection
For u∈Rc and n=(np)∈ZP let ν(u,n)∈Cd×AP have archimedean part ν(u) and finite part (ϖpnp,ϖp−np)p.
Lemma 6.20.Let N⊂ZP be finite, and let W:N→[0,∞) have positive sum. There are a subset N′⊂N with
n∈N′∑W(n)≥khrel1n∈N∑W(n),
an element n0∈N′ with W(n0)>0, and elements βn∈Γ, for n∈N′, with NK/Fβn=1 and
βnOK=p∏Ppnp−n0,p(ιPp)−(np−n0,p).
The cosets βnOK1 are pairwise disjoint, and βnOK1 is the set of all elements of relative norm one that generate this ideal. Moreover, let ϖ0∈AP have coordinates (ϖpn0,p,ϖp−n0,p). Then for n∈N′ and β∈βnOK1, the point ϖ0βf has coordinates (ϖpnpωp,ϖp−npωp−1) with units ωp∈Op×.
Proof. Map n to the class of A(n)=∏pPpnp in Cl(K) modulo the image of Cl(F), a group of order κhrel. Let N′ be a fiber on which the sum of W is largest; this sum is positive, so we can choose n0∈N′ with W(n0)>0. For n∈N′ write A(n)A(n0)−1=xnanOK with xn∈K× and an a fractional ideal of F, and put βn=xn/ιxn. As in the proof of Lemma 5.6, NK/Fβn=1 and βn generates the displayed ideal, which is supported on the primes Pp,ιPp; hence βn∈Γ. Two elements of relative norm one generating the same ideal differ by an element of OK1, and distinct n∈N′ give distinct ideals. Finally, if β∈βnOK1, then βp has valuation np−n0,p and (ιβ)p=βp−1, which gives the last assertion. □
Transfer with Shell Profiles.
Proposition 6.21 (Transfer with shell profiles). Let 0<δ<1 and C∈R. Let (Kj/Fj)j≥1 be quadratic extensions satisfying (G1)–(G3) with a common λ, with dj=[Fj:Q]→∞ and logLFj(1)≤djC. For each j let Pj be a finite set of finite places of Fj that split in Kj, and for p∈Pj let gp be a shell profile on Fp2 with parameters (Qp,kp,w(p)), where the parameters range over a fixed finite set independent of j. Let Fδ,p=Fδ(gp), let (fR,gC) be admissible with Fourier constants (M,σ), and suppose that (46) holds for every j, with Cf=∏p∈Pj(Op×ϖp−kpOp) the product of the periods of the gp, that is, with Hf=dj−1∑p∈PjkplogQp. Let θj be the value of θ for Fj, and put
we call Mfin,j the margin of Kj/Fj with these shell profiles. If infjMfin,j>0, then there are finite sets Uj⊂R2 with ∣Uj∣→∞ and u(Uj)/∣Uj∣1+δ→∞.
Compared with (39), the margin (48) has no separate term −δH: by (43), each Fδ,p contains the factor Qp−δkp, and Corollary 6.23 shows that for uniform local types and unweighted shell profiles (48) reduces to (39).
Proof of Proposition 6.21. Fix η>0 with 4η<infjMfin,j, and consider one extension K/F of the sequence, with d large. We drop the sequence index j and write Mfin for the margin of K/F; below, j indexes the complex places of F.
The profile and the probability measure. On Cd×AP let f=∏vfR∏j=1cgC∏p∈Pgp, with one factor for each real place v of F, each complex place j of F and each p∈P, let V=−logf on the set where f>0, and put
A=∫fp=ARbACcp∈P∏A(gp),Z=IRbICcp∈P∏Z(gp).
Besides the points ν(u,n) defined before Lemma 6.20, let ν(u,n,ω), for ω=(ωp)∈∏pOp×, have archimedean part ν(u) and finite part (ϖpnpωp,ϖp−npωp−1)p. Let P be the probability measure on (Cd×AP)×Rc×ZP×∏pOp× with density f(x)f(x+ν(u,n,ω))/Z with respect to dxdudω and counting measure in n, where dω is the Haar probability measure. It is the product of the endpoint laws of the archimedean blocks (Definition 5.22) and of the endpoint laws of the gp. Its endpoints are x and x+ν(u,n,ω). Their energies V(x) and V(x+ν(u,n,ω)) have the same distribution and are sums of independent terms, one for each block and each p∈P: at the archimedean blocks this is shown in the proof of Lemma 5.25, and at p∈P the map (z,n,ω)↦(−z−s(n,ω),n,ω) preserves the endpoint law of gp and carries its first endpoint to minus the second, gp is even, and the energy −loggp of the first endpoint takes finitely many values. Let dmˉ be their common mean; this mˉ includes the places of P and is not the mˉ of Lemma 5.25. The Qp take finitely many values, since the parameters of the shell profiles do, and each is a power of the residue characteristic of p; so these characteristics lie in a fixed finite set, and ∣P∣=O(d). Since the parameters range over a finite set, the variance of either energy is O(d). Let χ be the indicator of the event that both energies lie in [d(mˉ−η),d(mˉ+η)]. By Chebyshev’s inequality, as in the proof of Lemma 5.25, P(χ=1)≥1/2 for large d.
The set Ω. Let Ω be the set of points of Cd×AP where f>0 and V≤d(mˉ+η); it plays the role of the window of Definition 5.11. It is a bounded Borel set: the archimedean energies are coercive and bounded below, and the finite part lies in the compact support of ∏pgp, on which the finite energies are bounded. It is invariant under rotations of the archimedean coordinates, and, by Lemma 6.6, under (xp,yp)↦(ωxp,ω′yp) for units at each p∈P and under translation by Cf. For u∈Rc and n∈ZP let Φ(u,n)=∣Ω∩(Ω−ν(u,n))∣, and let W(n)=∫RcΦ(u,n)du. Only finitely many n have W(n)=0. By these invariances, the measure of the intersection of Ω with its translate by any vector whose archimedean part has the moduli of ν(u) and whose finite part is that of some ν(u,n,ω) equals Φ(u,n). Where χ=1, both endpoints lie in Ω and f(x)f(x+ν(u,n,ω))≤e−2d(mˉ−η). Integrating over x, u and ω and summing over n, we obtain
Class selection and rescaling. Apply Lemma 6.20 to the finitely many n with W(n)>0, obtaining n0, N′ and βn, and let ϖ0 be as in that lemma. Put Ω′=ϖ0−1Ω, where ϖ0 acts on the finite part. Then ∣Ω′∣=∣Ω∣, and Ω′ is invariant under the rectangular group Cf′=ϖ0−1Cf, of measure ∣Cf∣. For n∈N′, β∈βnOK1 and h∈Rc, the element β(h)=(Whβ∞,βf)∈Γh satisfies
by the last assertion of Lemma 6.20 and the invariances of Ω, as in Lemma 5.12.
Points and edges. Let L′ be the ideal attached to Cf′ in Lemma 6.15. Fix a Z-basis of L∞′, let PL′ be the fundamental parallelepiped that it spans, and for y∈[0,1)2d let yL′∈PL′ be the point with coordinate vector y in this basis. For h∈Rc, y∈[0,1)2d and ζ∈Cf′ put z0=(WhyL′,ζ); these points form the fundamental domain WhPL′×Cf′ of Γh of Lemma 6.15(d). Let X(h,z0)=(Γh+z0)∩Ω′ and let N(h,z0)=∣X(h,z0)∣ be the number of its points (the analogue of the count n(h,y) of Section 5), and let E(h,z0) count the pairs (x,β) with x∈X(h,z0), β∈⋃n∈N′βnOK1 and x+β(h)∈Ω′. Averages over (h,z0) are taken with respect to the probability measure dhdydζ/(Regd∣Cf′∣) on Γ1×[0,1)2d×Cf′. Since 1Ω′≤epd(mˉ+η)(f∘ϖ0)p and (f∘ϖ0)p is the product of an archimedean factor and a function invariant under Cf′, Lemma 6.17 and Remark 6.19 give
N(h,z0)≤Nmax=covol02Aepd(mˉ+η)
for all (h,z0). By tiling, for fixed h the average of N(h,z0) over (y,ζ) is ∣Ω′∣/covol0, and the average of E(h,z0) is covol0−1∑n∈N′∑β∈βnOK1Φ(ι(β)−h,n). Averaging over h∈Π1 as in Lemma 5.9, the average of E becomes
By (30) and logLF(1)≤dC, the first factor is at least 2−2−δ2d(1−δ)π(1−θ)dλ−(1/2−δ)de−dC, and Z/A1+δ=exp(bJR+cJC+∑plogFδ,p). Hence E/N1+δ≥2−2−δed(Mfin−4η).
The planar set. Let v0 be a real place of F and let U be the set of v0-coordinates of the points of X(h,z0). Distinct points differ by (Whγ∞,γf) with 0=γ∈Γ, whose v0-coordinate is nonzero, so ∣U∣=N. Each counted pair (x,β) gives two points at distance ∣ϕv0(β)∣=1, and distinct pairs give distinct ordered pairs of points. As in the proof of Proposition 5.26 (Lemma 5.15), u(U)≥E/2, the ratio u(U)/∣U∣1+δ tends to infinity, and the bound u(U)≤∣U∣2/2 forces ∣U∣→∞. □
Uniform Local Types.
Corollary 6.23 (Uniform local types). Let 0<δ<1 and C∈R. Let (Kj/Fj)j≥1 be quadratic extensions satisfying (G1)–(G3) with a common λ, with dj=[Fj:Q]→∞ and logLFj(1)≤djC, and with the same uniform local types (er,fr)r∈R (Definition 5.4). For each r∈R fix an integer kr≥0 and weights w(r) as in Definition 6.5, let Fδ,r be the value of the local functional Fδ at the shell profile with parameters (rfr,kr,w(r)), and let H be as in (31). Let (fR,gC) be admissible with Fourier constants (M,σ) satisfying (37) with μ=2eH/2−ℓ, and put
If infjMR(θj)>0, where θj is the value of θ for Fj, then there are finite sets Uj⊂R2 with ∣Uj∣→∞ and u(Uj)/∣Uj∣1+δ→∞. For unweighted shell profiles, ∑rlogFδ,r/(erfr)=J−δH, so that MR is the margin M of (39).
Proof. Apply Proposition 6.21 with Pj the set of all places of Fj above R, which split in Kj, and with gp the shell profile with parameters (rfr,kr,w(r)) for p above r; these parameters take ∣R∣ values, and Fδ,p=Fδ,r. Since there are d/(erfr) places above r, each with Qp=rfr,
In particular (46) coincides with (37), whatever the shell weights, and the margin (48) of Kj/Fj equals MR(θj). For unweighted shell profiles, Lemma 6.8 gives logFδ,r=log(kr+1)−δkrfrlogr, and the sum in (49) is J−δH. □
Thus Proposition 6.21 contains Proposition 5.26, and the shell weights replace the term J−δH of (39) by ∑rlogFδ,r/(erfr) without changing the Fourier condition.
Archimedean Profiles
We use a Gaussian at the compact coordinates, and at the pair coordinates a Student profile multiplied by a positive polynomial.
Definition 7.1. Let s>0 and a>0, and let P be a real polynomial in two variables that is positive on [0,1]2. For (z,w)∈C2 put
We call gC the Student–Bernstein profile with parameters (s,a,P). For P≡1 it is the Student profiletst′s; for s>1, each factor (1+a∣z∣2)−s is proportional to the density of a bivariate Student t-distribution. If P has degree at most m in each variable, its Bernstein coefficients of degree m are the numbers βij with
The Bernstein polynomials bim are nonnegative on [0,1] and sum to one, so minβ≤P≤maxβ on [0,1]2, where minβ and maxβ are the smallest and largest Bernstein coefficients. From now on, gC denotes a Student–Bernstein profile with parameters (s,a,P). Section 8 uses
s=1.14600585,a=0.0000128633241758,
and the polynomial P of degree m=3 in each variable whose Bernstein coefficients form the symmetric matrix
All finite decimals here are exact rational numbers.
Heuristic. The shape of gC is adapted to the displacements at a pair of coordinates. By (29) they have moduli (eu,e−u), so as u varies they run along a hyperbola. A profile of width a−1/2 in each coordinate stays close to its translate for all u with e∣u∣≲a−1/2, a range of length about log(1/a); this produces the factor log(1/a) in Proposition 7.17, against the term 2δloga in (61), which comes from the exponent 1+δ on the mass AC. For Student–Bernstein profiles the overlap IC reduces to beta integrals and digamma values (Lemmas 7.5–7.10), which makes rigorous bounds possible, and the polynomial P couples the two radii.
In this section we prove that these profiles are admissible in the sense of Definition 5.23, and we derive bounds for the functionals JR and JC (Definition 5.22) that reduce to finite computations with rational data. Section 8 carries out these computations at δ=0.04273. Throughout, 0<δ<1 and p=2/(1+δ). We write ψ=Γ′/Γ, γ for Euler’s constant, (x)n for the rising factorial, Hn=∑k=1n1/k with H0=0, and Beta(α,α′) for the beta distribution on (0,1) with density proportional to tα−1(1−t)α′−1.
The Gaussian at the Compact Coordinates.
Lemma 7.3.Let aR=2pδ and fR(z)=e−aR∣z∣2/p, so that fRp=e−aR∣z∣2. Then
and aR maximizes JR over the Gaussians e−a′∣z∣2/p, a′>0.
Proof. For f=e−a′∣z∣2/p we have AR=π/a′, and, since ∣z∣2+∣z+1∣2=2∣z+1/2∣2+1/2, IR=(πp/(2a′))e−a′/(2p). Hence JR=log(p/2)+δlog(a′/π)−a′/(2p), a concave function of loga′ with its maximum at a′=2pδ. □
By (39), the margin is affine in θ with slope −logπ−2JR+JC. This slope depends on the profiles; for ours it is positive (Section 8).
The Pair Convolution
Lemma 7.5.Let α,α′>0 with A=α+α′−1>0, and let h∈C. Then
Proof. Insert (1+∣z∣2)−α=Γ(α)−1∫0∞rα−1e−r(1+∣z∣2)dr and the analogous formula with a variable r′. Since r∣z∣2+r′∣z+h∣2=(r+r′)∣z+r′h/(r+r′)∣2+rr′∣h∣2/(r+r′), the integral of e−r∣z∣2−r′∣z+h∣2 over z∈C is π(r+r′)−1e−rr′∣h∣2/(r+r′); the factors e−r and e−r′ remain in the Laplace integrals. Substitute r=RT, r′=R(1−T), with Jacobian R, and integrate over R. This gives
(c) Suppose A,B≥1, 0<y<1, ABy2<1 and 2log(1/y)≥H⌈A⌉−1+H⌈B⌉−1. Then every term of the series in (b) is nonnegative, and for every integer n∗≥0 its remainder Remn∗(y) after the terms n≤n∗ satisfies
0≤Remn∗(y)≤log(1/y)1−ABy2(ABy2)n∗+1.
Proof. (a) Split the integral at v=0. On v≥0, first replace (1+ye−2v)−B by 1. Since 0≤1−(1+ye−2v)−B≤Bye−2v, this changes the half-integral by an amount in [0,By/2]. The substitution x=ye2v gives ∫0∞(1+ye2v)−Adv=21∫y∞(1+x)−Adx/x, and
The last integral lies in [0,Ay], because 0≤1−(1+x)−A≤Ax; in particular it tends to 0 with y. On the other hand, the substitution t=1/(1+x) gives
∫y∞x(1+x)Adx=∫01/(1+y)1−ttA−1−1dt+logy1+y.
Letting y→0 in the two expressions yields ϰA=∫01(tA−1−1)(1−t)−1dt=−ψ(A)−γ[25], Equation (5.9.16). Thus the half v≥0 equals 21(log(1/y)−ψ(A)−γ) plus a term in [0,Ay/2] minus a term in [0,By/2]. The half v≤0 is the same with A and B exchanged. Adding the halves proves (a).
(b) The substitutions x=ye2v and then t=x/(x+y2) give
2QAB(y)=∫01tB−1(1−t)A−1(1−(1−y2)t)−Adt.
By Euler’s integral [25], Equation (15.6.1), the right side is Γ(A)Γ(B)F(A,B;A+B;1−y2), where F is Olver’s regularized hypergeometric function. The expansion [25], Equation (15.8.10), in the case where the third parameter is the sum of the first two, gives
for 0<y<1. Dividing by two proves (b). The coefficients grow at most polynomially in n, so the series converges absolutely.
(c) For A≥1 and n≥0, monotonicity of ψ and the recurrence ψ(x+1)=ψ(x)+1/x give ψ(n+1)≤ψ(A+n)≤ψ(⌈A⌉+n)≤ψ(n+1)+H⌈A⌉−1. Hence the bracket in (b) lies between log(1/y)−21(H⌈A⌉−1+H⌈B⌉−1)≥0 and log(1/y). Moreover (A)n/n!=∏1≤k<n(A+k)/(k+1)≤An for A≥1, and similarly for B. Summing the geometric series proves (c). □
Definition 7.9. Let dij be the monomial coefficients of P, so that P(t,t′)=∑i,jdijtit′j, and put η0=2s−1,
where the sum runs over all crosses and, for each cross, T∼Beta(s+i,s+k) and T′∼Beta(s+j,s+l) are independent. Each expectation is finite.
Proof. Expand gC(z,w)=∑i,jdij(1+a∣z∣2)−s−i(1+a∣w∣2)−s−j and multiply out gC(z,w)gC(z+eu,w+e−u). For one cross, the scaling z↦z/a and Lemma 7.5 give
and similarly in w with e−2u, Bjl and T′. By Tonelli’s theorem, the u-integral of the product is E∫R(1+Ye2u)−Aik(1+Y′e−2u)−Bjldu with Y=aT(1−T) and Y′=aT′(1−T′), and the translation u↦u+41log(Y′/Y) turns the inner integral into QAikBjl(YY′). By (53), QAB(y) is at most log(1/y) plus a constant, and log(1/(T(1−T))) is integrable under every beta law; so each term is finite and the finite signed sum is legitimate. □
Admissibility
For the Fourier bound we assume that P has positive Bernstein coefficients βij of some degree m≥1.
Definition 7.11. Let Δ1 and Δ2 be the forward difference operators in the two indices of (βij). For 0<ρ<1, the width of the complex tube in the proof of Lemma 7.12, put ω=(ρ+ρ2)/(1−ρ2),
Lemma 7.12. Let s>0, sp>1 and a>0, let P have positive Bernstein coefficients βij of degree m≥1, let 0<ρ<1, and suppose that the tube constant satisfies q0<1. Then for all ξ1,ξ2∈C,
So the complexified t=D−1 lies within ω of the real t0=(1+∣x∣2)−1∈(0,1], and likewise for t′. By the Bernstein derivative formula, ∂tr∂t′r′P/(r!r′!)=(rm)(r′m)∑i,j(Δ1rΔ2r′β)ijbim−r(t)bjm−r′(t′), which is at most Err′ in absolute value on [0,1]2. Taylor expansion of the polynomial P about (t0,t0′) therefore gives ∣P(t,t′)−P(t0,t0′)∣≤q0minβ≤q0P(t0,t0′). Thus P(t,t′) lies in the right half-plane, and with principal branches,
where x′+iy′ is the variable attached to w. The same bounds hold for ∣y∣,∣y′∣<ρ′ with some ρ′>ρ, since q0 depends continuously on ρ, so the left side is holomorphic in each complex coordinate on a neighborhood of the closed tube.
The function gCp is invariant under rotations of z and of w. Rotate so that ξ1 and ξ2 are positive multiples of the first basis vector, and shift the first real coordinate of x by −iρ, then that of x′ by −iρ. Each shift is justified by Cauchy’s theorem in one variable: the integrand is holomorphic in the strip and bounded there by a constant times (1+∣x∣2)−sp(1+∣x′∣2)−sp, which tends to zero at infinity and is integrable because sp>1. After the shifts the Fourier factor is multiplied by e−2πρ(∣ξ1∣+∣ξ2∣)/a, and the displayed bound gives the lemma. □
Proposition 7.13. Let s≥1, a>0, and let P have positive Bernstein coefficients βij of some degree m≥1 in each variable. Let 0<δ<1 with aR=2pδ≤1, let fR be the Gaussian of Lemma 7.3, and let 0<ρ<1 with q0<1. Then (fR,gC) is admissible with Fourier constants
M=max{2,KC},σ=min{1,σC}.(54)
Proof. (P1) The functions fRp and gCp=(1+a∣z∣2)−sp(1+a∣w∣2)−spP(t,t′)p are smooth, because P≥minβ>0 on [0,1]2, and they and all their derivatives are bounded.
(P2) The integrals at the compact coordinates are computed in Lemma 7.3. The mass AC is finite because sp>1 (Proposition 7.20 below), and 0<IC<∞ because gC≤maxβtst′s and Lemma 7.10 applies to the Student profile tst′s.
(P3) VR=aR∣z∣2/p, and VC=slog(1+a∣z∣2)+slog(1+a∣w∣2)−logP(t,t′) is at least −logmaxβ and tends to infinity with ∣z∣ or ∣w∣.
(P4) The endpoint law at a compact coordinate is Gaussian. For the pair, −logP is bounded, so it suffices to bound the second moments of log(1+a∣z∣2) and log(1+a∣w∣2). Choose 0<τ<s; then τ<2s−1 as well, since s≥1. We have log2(1+x)≤cτ(1+x)τ for x≥0 and some constant cτ, and gC≤maxβtst′s. Hence the second moment of log(1+a∣z∣2) is at most a constant times the integral of (1+a∣z∣2)−(s−τ)(1+a∣z+eu∣2)−s(1+a∣w∣2)−s(1+a∣w+e−u∣2)−s. By the argument of Lemma 7.10, with exponents s−τ>0 and s and A=2s−τ−1>0, this integral is a multiple of EQA,2s−1(y), which is finite by (7.8). The same holds for w.
(P5) By Lemma 5.21 the Gaussian satisfies ∣fRp(ξ)∣≤2e−∣ξ∣AR≤Me−σ∣ξ∣AR, and by Lemma 7.12 the pair profile satisfies the bound with KCe−σC(∣ξ1∣+∣ξ2∣)≤M2e−σ(∣ξ1∣+∣ξ2∣). □
Proof. Start from Lemma 7.10. Since y≤y0=a/4<1, the hypotheses of Lemma 7.7(c) hold at every value of y, with A=Aik≥1 and B=Bjl≥1. Split each QAikBjl(y) into the terms n≤n∗ of Lemma 7.7(b) and the remainder Remn∗(y).
Consider the term of index n for one cross, with α=s+i, α′=s+k, A=Aik=α+α′−1, and similarly for T′. Its factor y2n is a2n(T(1−T))n(T′(1−T′))n, and log(1/y)=log(1/a)−21log(T(1−T))−21log(T′(1−T′)). The beta integral gives E(T(1−T))n=(α)n(α′)n/(A+1)2n, so
Weighting by (T(1−T))n turns Beta(α,α′) into Beta(α+n,α′+n), under which the mean of log(T(1−T)) is ψ(α+n)+ψ(α′+n)−2ψ(A+2n+1). The recurrence ψ(x+j)=ψ(x)+Dx(j) and ψ(1)=−γ give
Summing over crosses gives the displayed main terms.
By Lemma 7.7(c), 0≤Remn∗(y)≤log(1/y)(ABy2)n∗+1/(1−ABy2). The right side increases with y on (0,y0], because the derivative of y2n∗+2log(1/y) is y2n∗+1((2n∗+2)log(1/y)−1)>0 when log(1/y)>1/2. So 0≤Remn∗(y)≤log(4/a)(ABy02)n∗+1/(1−ABy02). For crosses with wijkl>0 we discard the remainder, and for crosses with wijkl<0 we subtract its upper bound. This proves (57). □
With n∗=1 the bound keeps the exact term of order a2; the terms it discards or subtracts are of order a4log(1/a). It is a lower bound for IC only.
An Upper Bound for the Mass AC.
Proposition 7.20.Let gC be a Student–Bernstein profile with parameters (s,a,P) such that q=sp−1>0. Then
AC=a2q2π2AˉC,AˉC=EP(T,T′)p,(58)
where T and T′ are independent with density qtq−1 on (0,1). Partition [0,1]2 into rectangles Q=[l,r]×[l′,r′], and use on Q the coordinates x=(T−l)/(r−l) and x′=(T′−l′)/(r′−l′). Suppose that the Bernstein coefficients, of some degree, of the restriction of P to Q in these coordinates lie in [mQ,MQ] with mQ>0, and put
Proof. The substitution t=(1+a∣z∣2)−1 gives dz=πdt/(at2) on C. Hence
AC=a2π2∫01∫01tsp−2t′sp−2P(t,t′)pdtdt′,
which is (58) since sp−2=q−1. The Bernstein convex-hull property gives ∣WQ∣≤ρQ<1 on Q. For 1<p<2 the binomial coefficients satisfy ∣(k+1p)∣/∣(kp)∣=∣p−k∣/(k+1)<1 for k≥1, so the binomial series of (1+W)p converges uniformly on ∣W∣≤ρQ, and its tail after index N≥1 is at most ∣(N+1p)∣ρQN+1/(1−ρQ). Integrate against the density on Q, whose integral over Q is M0(l,r)M0(l′,r′). The moment formula follows by expanding (v−l)i; the convention 0q=0 is valid since q>0. □
If P is symmetric and both axes are partitioned alike, a rectangle and its reflection contribute equally, so off-diagonal rectangles may be counted twice and diagonal ones once. Subsection 8.5 describes how the quantities in Proposition 7.20 are evaluated.
A Lower Bound for the Functional JC
Corollary 7.24.Under the hypotheses of Propositions 7.17 and 7.20, let Z∗ be the right side of (57) for some n∗, and suppose Z∗>0. If AˉC∗≥AˉC, then
Proof. By Proposition 7.17, IC≥π2a−2Z∗, and by (58), AC≤π2a−2q−2AˉC∗. Substitute into JC=logIC−(1+δ)logAC. □
Interval evaluation of the right side of (61) gives a rigorous lower bound for JC; the upper endpoint of such an enclosure is not an upper bound for JC. For rational s, all digamma values in Proposition 7.17 reduce to rational corrections and to cs=2/η0+ψ(η0)−2ψ(s)+ψ(1), since ψ(2s)=ψ(η0)+1/η0 and ψ(1)=−γ. The digamma function at rational arguments is enclosed as follows.
Lemma 7.26.For x>0 and integers L≥0 and k∗≥1, put z=x+L. Then
and use ∫0∞t2k−1(e2πt−1)−1dt=∣B2k∣/(4k) and (−1)k−1∣B2k∣=B2k. The remaining integral has a fixed sign and is at most ∣B2k∗∣/(2k∗z2k∗) in absolute value, because (t2+z2)−1≤z−2. Finally apply the recurrence ψ(x+1)=ψ(x)+1/xL times. □
Proof of Theorem 1.1
We prove Theorem 1.1 by applying Corollary 6.23 to the fields of Theorem 2.42, with the upper bound for the relative zeta value in Theorem 4.2 and the data of Definition 8.1. Propositions 8.4 and 8.5 are verified by computer, in exact rational arithmetic and in outward-rounded interval arithmetic, as described in Subsection 8.5; the other steps are proved in the text. The inequalities between decimals in Propositions 8.4 and 8.5 and in Remarks 8.11 and 8.12, and the entries of Table 6 and Table 8, are computed and printed by the supplementary program geom241.py. In the inequalities the decimals are rounded outward from interval enclosures, and the table entries are rounded to ten decimals. No rounded decimal is substituted into a later calculation; in particular, AˉC∗ in Definition 8.3 is the exact upper end of a computed enclosure, not its rounded decimal.
r
er
fr
Qr
kr
nr
gr
unweighted
2
8
4
16
7
6
0.0393401103
0.0390666415
3
2
2
9
9
6
0.3675484302
0.3643996093
5
2
2
25
6
6
0.2817099488
0.2801636913
7
1
8
5764801
1
2
0.0034946608
0.0034946569
29
1
4
707281
1
3
0.0294023321
0.0294022443
Table 6. Local data at the primes of R. Here Qr=rfr, the profile gr uses the weights wr,0,…,wr,nr−1, and the last two columns give logFδ,r/(erfr) for gr and for the unweighted shell profile with parameters (Qr,kr), rounded to ten decimals.
Parameters and Local Data
Definition 8.1. The data of this section are the following.
(a) δ=4273/100000=0.04273, p=2/(1+δ)=200000/104273 and C=4871285/108=0.04871285; λ=29/43615 and ℓ=logλ, as in Theorem 2.42.
(b) At the compact coordinates, fR is the Gaussian of Lemma 7.3, with fRp=e−aR∣z∣2 and aR=2pδ=17092/104273<1. (c) At the pair coordinates, gC is the Student–Bernstein profile of Definition 7.1 (the polynomial P below is positive on [0,1]2, since its Bernstein coefficients are positive by Proposition 8.4(b)) with s=1.14600585, a=0.0000128633241758 and the polynomial P of degree m=3 in each variable whose Bernstein coefficients are (50). We put q=sp−1=12492817/10427300.
(d) In Definition 7.11 and Lemma 7.12 we take ρ=10−3, with the Bernstein coefficients of degree m=3 of (c). In Proposition 7.20 we partition [0,1]2 into the 64 squares of side 1/8, use on each square the Bernstein coefficients of degree 3, with mQ and MQ the smallest and largest of these coefficients, and take N=18. In Proposition 7.17 we take n∗=1.
(e) R={2,3,5,7,29}. For r∈R, (er,fr) is the absolute type (Definition 2.41) of the primes of F above r given by Theorem 2.42 (Table 5, repeated in Table 6), the exponent kr is that of Table 6, and Qr=rfr. For a prime p of F above r, gr is the shell profile on Fp2 with parameters (Qr,kr,(wr,iwr,j)i,j≥0), where wr,0=1, the weights wr,i with 1≤i<nr are those of Table 7, and wr,i=0 for i≥nr; these weights are nonnegative by Proposition 8.4(a). By Lemma 6.8, the value Fδ,r=Fδ(gr) depends only on these parameters.
r
i
wr,i
≈wr,i
2
1
2854206002066/59041214756005
4.83426⋅10−2
2
2
198349814036/95709219081499
2.07242⋅10−3
2
3
7762557303/86561984122975
8.96763⋅10−5
2
4
109894931/27828186078256
3.94905⋅10−6
2
5
17033757/96603874664399
1.76326⋅10−7
3
1
5693537714695/63764117760919
8.92906⋅10−2
3
2
567838527095/78237129879259
7.25792⋅10−3
3
3
43356430379/72943614524859
5.94383⋅10−4
3
4
2265261068/45953937834745
4.92942⋅10−5
3
5
74149907/18018553422030
4.11520⋅10−6
5
1
1343081168627/46386563301209
2.89541⋅10−2
5
2
61788974773/82067065665364
7.52908⋅10−4
5
3
1389672153/69896165011649
1.98820⋅10−5
5
4
32644248/60812698834313
5.36800⋅10−7
5
5
511565/34657937034136
1.47604⋅10−8
7
1
2309805/69240611498956
3.33591⋅10−8
29
1
36498535/95956876917664
3.80364⋅10−7
29
2
4/29090285629613
1.37503⋅10−13
Table 7. The exact shell weights wr,i for 1≤i<nr; all wr,0=1, and wr,i=0 for i≥nr. The last column gives the weights rounded to six significant digits, for orientation only; the computations of Section 8 use the exact fractions.
Every parameter written as a finite decimal is exact. We recall θ∗=65535/131072 from Theorem 2.42. Since the weights wr,i are nonnegative (Proposition 8.4(a)) and wr,0=1, Corollary 6.10, with x=x′=(wr,i)i≥0, gives
where Pr=∑imiwr,ip, Sr=∑imiwr,i2 and Rr=∑i,i′Δii′wr,iwr,i′, with the shell measures mi and the matrix Δ of Section 6 for Q=Qr. Table 6 lists the local data and Table 7 the exact weights. At 7 and 29 only S0,S1, respectively S0,S1,S2, carry weight, and these shell profiles improve only slightly on the unweighted shell profiles (Table 6).
Definition 8.3. Let Z∗ be the right side of (57) for the data of Definition 8.1, with n∗=1. Let AˉC∗ be the upper end of the rigorous interval enclosure of AˉC computed by the program of Subsection 8.5 (Proposition 8.5(a)), so that AˉC≤AˉC∗<38.7888697256861071, the last number being AˉC∗ rounded up. Let JC∗ be the right side of (61) for these Z∗ and AˉC∗, and put
Proof. These are finite checks, made by the supplementary program geom241.py as described in Subsection 8.5. The comparisons in (a)–(c), and the first condition in (d), are exact comparisons of rational numbers; the other two conditions in (d) are checked by outward-rounded interval evaluation of log(4/a). □
Proposition 8.5.For the data of Definition 8.1, with Z∗, JC∗ and M∗ as in Definition 8.3, the following hold.
(a) With T and T′ independent with density qtq−1 on (0,1),
Proof. The supplementary program geom241.py evaluates these quantities in outward-rounded interval arithmetic from the exact data, as described in Subsection 8.5: (a) by Proposition 7.20; (b) from (57), with Lemma 7.26 for the digamma values; (c) from (61); (d) from (51); (e) from (62); (f) from the definitions of H and ℓ; (g) from Lemma 7.12 and the definition of μ; and (h) and (i) from Definition 8.3, where AˉC∗ is the upper end of the enclosure computed for (a). □
In (b), the upper endpoint bounds the expression Z∗, not the overlap IC. The proof of Theorem 1.1 uses Proposition 8.4 and, from Proposition 8.5, only the fact, established by the computation for (a), that AˉC≤AˉC∗, the positivity of Z∗ given by (b), and (g), (h) and (i); the other bounds in Proposition 8.5 are not used in the proof. Table 8 lists the terms of M∗(θ∗).
Conclusion of the Proof.
Corollary 8.9.The pair(fR,gC)of Definition 8.1 is admissible with Fourier constantsM=2andσ=1, and these satisfy(37)withμ=2eH/2−ℓ.
Proof. Proposition 7.13 applies, since s≥1, a>0, aR≤1 and 0<ρ<1, the Bernstein coefficients (50) of degree m=3 are positive, and qˉ0<1 (Proposition 8.4(b)). It gives the Fourier constants M=max{2,KC} and σ=min{1,σC}. By Proposition 8.5(g), KC<4 and σC>1, so M=2 and σ=1. The right side of (37) is then log50+2<5.9121, while σμ=μ>17.8683654522 by Proposition 8.5(g). □
Lemma 8.10.For the data of Definition 8.1,JC≥JC∗. Consequently, forθ∗≤θ<1/2, the marginMR(θ)of Corollary 6.23 for these data satisfies
MR(θ)≥M∗(θ)≥M∗(θ∗)>0.000176730033534598.
Proof. Corollary 7.24 applies: the hypotheses of Propositions 7.17 and 7.20 hold by Proposition 8.4(c),(d), Z∗>0 by (64), and AˉC≤AˉC∗ by Proposition 8.5(a) and Definition 8.3. Hence JC≥JC∗. Since MR(θ)−M∗(θ)=θ(JC−JC∗) and θ≥0, we get MR(θ)≥M∗(θ). The function M∗ is affine in θ with slope −logπ−2JR+JC∗, which is positive by Proposition 8.5(h); so M∗(θ)≥M∗(θ∗) for θ≥θ∗, and (65) completes the proof. □
Proof of Theorem 1.1. For each n≥49, let K(n)/B be a Galois extension with [K(n):B]=2n as in Theorem 2.42, let ι1 be a complex conjugation at a place of K(n) above the real place v1 of B, put F(n)=(K(n))⟨ι1⟩ and dn=[F(n):Q], and let θn be the value of θ for F(n). By Theorem 4.2, dn−1logLF(n)(1)<C for all sufficiently large n; we discard the finitely many other fields.
We apply Corollary 6.23 to this sequence, with δ=0.04273, the exponents kr, the weights w(r)=(wr,iwr,j)i,j≥0 and the pair (fR,gC) of Definition 8.1, and check its hypotheses. By Theorem 2.42, each K(n)/F(n) satisfies (G1)–(G3) with the common λ=29/43615, dn=2n→∞, and all these extensions have the uniform local types (er,fr)r∈R of Table 5. By the choice above, logLF(n)(1)≤dnC. The weights are as in Definition 6.5, since wr,0=1, wr,i>0 for i<nr by Proposition 8.4(a), and wr,i=0 for i≥nr. The pair (fR,gC) is admissible, with Fourier constants that satisfy (37) with μ=2eH/2−ℓ, by Corollary 8.9. Finally, θ∗≤θn<1/2 by Theorem 2.42, so Lemma 8.10 gives MR(θn)≥M∗(θ∗)>0 for all n.
Corollary 6.23 therefore gives finite sets Un⊂R2 with ∣Un∣→∞ and u(Un)/∣Un∣1+δ→∞, where 1+δ=1.04273; these are the sets of Theorem 1.1, indexed by n. In particular u(Un)≥∣Un∣1.04273 for all large n, which gives the last assertion of Theorem 1.1 at the cardinalities ∣Un∣. □
Remarks
Remark 8.11. The proof uses the bound θ≥θ∗ only through Lemma 8.10, that is, through the positivity of M∗ on [θ∗,1/2). The program geom241.py also encloses the value at θ=0 and the zero θ0 of the affine function M∗:
M∗(0)<−0.316,θ0<0.49971293.
So M∗(θ)>0 for θ≥0.49971293; these enclosures are not used in the proof. By Theorem 2.42, b/d=1/(2Nι), where Nι is the number of conjugates of ι1 in Gal(K/B), so that θ=1/2−1/(4Nι). Hence Nι≥871 would suffice, whereas that theorem gives Nι≥215. On the other hand, at θ=0 the margin MR(0) does not involve JC and equals M∗(0)<0: with these profiles, mixed signature is essential.
Remark 8.12. The following numerical observations are not used in the proof. For the unweighted shell profiles with the same exponents kr, the sum in Proposition 8.5(e) is replaced by J−δH<0.7165268434. The proof of Proposition 6.21 needs only some η>0 with 4η below the lower bound of the margins, where η is the auxiliary parameter of that proof; for instance, η=10−9 gives
M∗(θ∗)−4η>0.000176726033534598>0.(66)
The value M∗(θ∗) is small compared with the individual terms in Table 8. Since M∗ decreases with slope 1 in C and with slope 1/2−δ in ℓ, it would be absorbed by an increase of 1.77⋅10−4 in C, or of 3.87⋅10−4 in ℓ. For δ=0.0428, with the same pair profile gC (the same s, a and coefficients (7.2)), the Gaussian of Lemma 7.3 for this δ, the same C, and exponents and weights produced by the optimization program shells241.py for this δ (its exponents kr and numbers nr agree with Table 6), the same evaluation gives M∗(θ∗)<−0.0015.
Remark 8.14. The weights wr,i are nonincreasing in i; this is an exact rational comparison. So gr is the nonnegative combination ∑i,j(wr,i−wr,i+1)(wr,j−wr,j+1)1Bi×ϖ−krBj of indicators of products of balls. As a consistency check, the rational number Z(gr)Qr−kr is computed twice for each r: from (44), and from this combination and (41). The two exact values agree.
Computations
The program geom241.py evaluates the formulas of Sections 6 and 7 for the data of Definition 8.1: the local functionals (62) from the exponents and weights of Table 6 and Table 7, and the quantities of Propositions 7.20 and 7.17, of Lemma 7.12 and of Corollary 7.24 from the profile data. For these it uses, unchanged, three modules of an earlier supplementary archive of the author (Subsection 1.6), which implement exactly these formulas. In Proposition 7.20 it uses the symmetry of (7.2): each square off the diagonal is evaluated once and counted twice. The program checks the hypotheses listed in Proposition 8.4 (for (c), on the squares on or above the diagonal; the others follow by the symmetry of (50), since the restriction of P to a reflected square has the transposed Bernstein coefficients), the conditions aR≤1, σC>1, logKC<2log2 and μ>10 (which implies (37) for M=2 and σ=1, since log50+2<10), and the positivity of the slope in Proposition 8.5(h); it also makes the consistency check of Remark 8.14. It then prints the enclosures stated in this section. The exponents and weights were produced by the separate optimization program shells241.py; they enter the proof only as the exact rational data of Table 7.
Rational quantities are computed exactly: q0, the restrictions of P to the squares and their Bernstein coefficients, the coefficients of the powers WQk and the binomial coefficients in Proposition 7.20, Gn and Gn, and the shell sums Sr and Rr; the polynomial arithmetic is done with integers after clearing denominators. The remaining quantities involve the logarithm, the exponential (also through the real powers xy=exp(ylogx) in (59), in cQρ and in wr,iρ), the square root in σC, the constant π, and the digamma function, which enters only through Lemma 7.26; that lemma is applied with exact rational arguments and Bernoulli numbers, L=64 and k∗=16, without rounding the rational parameters of the profile. These quantities are evaluated in the interval arithmetic of mpmath [40], which computes each arithmetic operation and each of these elementary functions with the lower endpoint rounded down and the upper endpoint rounded up, at a working precision of 80 decimal digits; exact rational algebra followed by interval evaluation keeps the bound of Proposition 7.20 valid despite the cancellation in (59). Hence every computed interval contains the exact value, and the decimals stated above, which are coarser than the working precision by more than fifty orders of magnitude, are rigorous bounds.
[1]Noga Alon, Thomas F. Bloom, W. T. Gowers, Daniel Litt, Will Sawin, Arul Shankar, Jacob Tsimerman, Victor Wang, and Melanie Matchett Wood, Remarks on the disproof of the unit distance conjecture, 2026, arXiv:2605.20695v1; https://arxiv.org/abs/2605.20695v1.
[2]Leonardo de Moura and Sebastian Ullrich, The Lean 4 theorem prover and programming language, Automated Deduction—CADE 28 (Cham) (André Platzer and Geoff Sutcliffe, eds.), Lecture Notes in Computer Science, vol. 12699, Springer, 2021, pp. 625–635.
[3]Tim Dokchitser, Computing special values of motivic L-functions, Experimental Mathematics 13 (2004), no. 2, 137–149.
[4]Michael T. M. Emmerich, Optimizing explicit unit-distance lower-bound certificates, 2026, arXiv:2606.03419v5, June 9, 2026 (v1 June 2, 2026); certificate with δ = 0.015263 . . .; https://arxiv.org/abs/2606.03419v5.
[5]Michael T. M. Emmerich and Francesco Cordella, Optimized certificate for the unit distance problem with extended prime number range, Zenodo record, June 5, 2026, Certificate and verification code. https://doi.org/10.5281/zenodo.20551478.
[7]———, Some problems in number theory, combinatorics and combinatorial geometry, Mathematica Pannonica 5 (1994), no. 2, 261–269.
[8]Mikhail Ershov, Golod–Shafarevich groups: a survey, International Journal of Algebra and Computation 22 (2012), no. 5, 1230001, 68 pages.DOI
[9]Ralph H. Fox, Free differential calculus. I: Derivation in the free group ring, Annals of Mathematics 57 (1953), no. 3, 547–560.DOI
[10]E. S. Golod and I. R. Shafarevich, On the class field tower, Izv. Akad. Nauk SSSR Ser. Mat. 28 (1964), no. 2, 261–272 (Russian), English translation: Amer. Math. Soc. Transl. (2) 48 (1965), 91–102; https://www.mathnet.ru/eng/im2955.
[11]Farshid Hajir, Christian Maire, and Ravi Ramakrishna, Cutting towers of number fields, Annales mathématiques du Québec 45 (2021), no. 2, 321–345.arxiv.org/abs/1901.04354
[12]Erich Hecke, Eine neue Art von Zetafunktionen und ihre Beziehungen zur Verteilung der Primzahlen. Zweite Mitteilung, Mathematische Zeitschrift 6 (1920), 11–51, https://doi.org/10.1007/BF01202991.
[13]S. A. Jennings, The structure of the group ring of a p-group over a modular field, Transactions of the American Mathematical Society 50 (1941), no. 1, 175–185, https://doi.org/10.2307/1989916.
[14]Fredrik Johansson, Rigorous high-precision computation of the Hurwitz zeta function and its derivatives, Numerical Algorithms 69 (2015), no. 2, 253–270.arxiv.org/abs/1309.2877
[16]Stéphane Louboutin, Explicit bounds for residues of Dedekind zeta functions, values of L-functions at s = 1, and relative class numbers, Journal of Number Theory 85 (2000), no. 2, 263–282, https://doi.org/10.1006/jnth.2000.2545.
[17]James S. Milne, Lectures on étale cohomology, 2013, Version 2.21, March 22, 2013; https://www.jmilne.org/math/CourseNotes/LEC.pdf.
[18]———, Algebraic number theory, 2020, Version 3.08, July 19, 2020; https://www.jmilne.org/math/CourseNotes/ANT.pdf.
[19]———, Class field theory, 2020, Version 4.03, August 6, 2020; https://www.jmilne.org/math/CourseNotes/CFT.pdf.
[20]mlewko, Post in the discussion thread of Erdős Problem #90, Erdős Problems forum, May 21, 2026, Reports δ = 0.031849981 . . . in Sawin’s criterion, with finite data found by ChatGPT; accessed September 30, 2026; https://www.erdosproblems.com/forum/thread/90.
[21]Eric Naslund, Answer to “What is the unit distance exponent?”, MathOverflow, answer 511576, 2026, Posted May 23, 2026, last edited June 9, 2026; accessed September 30, 2026; https://mathoverflow.net/a/511576.
[22]———, An exponent of 1.04273 for the unit distance problem: Lean formalization at exponent 1.0427, conditional on one zeta inequality, Palomar registry, PALOMAR-2026-10-01-000018, version 1, 2026, https://palomar-registry.org/entry?id=PALOMAR-2026-10-01-000018&version=1; Lean 4 source at https://github.com/enaslund/unit-distance-bound-0.0427/tree/e0ac836176016dbefae28646bc91fc3a5aca30e8/lean.
[23]———, An improved explicit lower bound for the unit distance problem, Unpublished manuscript, 2026, Records exponent 1.0358324. https://github.com/enaslund/unit-distance-bound-0.0427/tree/8ab94da4ba4faf6f6d8fe43c298f04c3b284e796/papers/0.0358324.
[24]———, Supplementary programs and data for “An exponent of 1.04273 for the unit distance problem”, GitHub repository, directory papers/0.04273/certificates, 2026, https://github.com/enaslund/unit-distance-bound-0.0427.
[25]National Institute of Standards and Technology, NIST Digital Library of Mathematical Functions, 2026, Release 1.2.8, September 15, 2026; accessed September 27, 2026; https://dlmf.nist.gov/.
[26]Jürgen Neukirch, Alexander Schmidt, and Kay Wingberg, Cohomology of number fields, 2 ed., Grundlehren der mathematischen Wissenschaften, vol. 323, Springer, Berlin, Heidelberg, 2008.DOI
[27]OpenAI, Planar point sets with many unit distances, Research manuscript, May 20, 2026, https://cdn.openai.com/pdf/74c24085-19b0-4534-9c90-465b8e29ad73/unit-distance-proof.pdf.
[28]Erdős unit distance exponent, Entry 84a of Optimization Constants in Mathematics, a crowdsourced repository started by Terence Tao, 2026, Accessed September 30, 2026; https://teorth.github.io/optimizationproblems/constants/84a.html.
[29]Bjorn Poonen, Rational points on varieties, Graduate Studies in Mathematics, vol. 186, American Mathematical Society, Providence, RI, 2017, https://math.mit.edu/~poonen/papers/Qpoints.pdf.DOI
[30]Claudio Quadrelli, Pro-p groups with few relations and universal Koszulity, Mathematica Scandinavica 127 (2021), no. 1, 28–42, https://doi.org/10.7146/math.scand.a-123644.
[33]Jean-Pierre Serre, A course in arithmetic, Graduate Texts in Mathematics, vol. 7, Springer, New York, 1973.DOI
[34]Joel Spencer, Endre Szemerédi, and William T. Trotter, Jr., Unit distances in the Euclidean plane, Graph Theory and Combinatorics (Béla Bollobás, ed.), Academic Press, London, 1984, https://trotter.math.gatech.edu/papers/44.pdf, pp. 293–303.
[35]spiderduckpig, Answer to “What is the unit distance exponent?”, MathOverflow, answer 511531, 2026, Posted May 21, 2026, with 1 + δ ≥ 1.03188306553964 . . .; comment of May 22, 2026 proposing a weighted count in the Zassenhaus filtration in Sawin’s Golod–Shafarevich step, with δ ≥ 0.0333487156372102; last edited June 9, 2026; accessed September 30, 2026; https://mathoverflow.net/a/511531.
[36]László A. Székely, Crossing numbers and hard Erdős problems in discrete geometry, Combinatorics, Probability and Computing 6 (1997), no. 3, 353–358, https://doi.org/10.1017/S09635483097002976.
[37]John T. Tate, Fourier analysis in number fields and Hecke’s zeta-functions, Algebraic Number Theory (J. W. S. Cassels and A. Fröhlich, eds.), Academic Press, London, 1967, Originally a Princeton University doctoral thesis, 1950; https://www.lms.ac.uk/publications/algebraic-number-theory, pp. 305–347.
[38]The FLINT team, FLINT: Fast Library for Number Theory, 2026, Version 3.6.0, used through python-flint 0.9.0; https://flintlib.org.
[39]The mathlib Community, The Lean mathematical library, Proceedings of the 9th ACM SIGPLAN International Conference on Certified Programs and Proofs, ACM, 2020, pp. 367–381.DOI
[40]The mpmath development team, mpmath: A Python library for arbitrary-precision floating-point arithmetic, 2023, Version 1.3.0; https://github.com/mpmath/mpmath/releases/tag/1.3.0.
[41]The PARI Group, PARI/GP, Université de Bordeaux, 2025, Version 2.17.2; https://pari.math.u-bordeaux.fr/.
[42]M. A. Tsfasman and S. G. Vlăduţ, Infinite global fields and the generalized Brauer–Siegel theorem, Moscow Mathematical Journal 2 (2002), no. 2, 329–402, https://doi.org/10.17323/1609-4514-2002-2-2-329-402.