Introduction

An equation over a group permits coefficients from that group and asks whether the unknowns can be realized in a larger group without identifying distinct coefficients. The exponent sums record the dependence on the unknowns after abelianization. A system whose exponent rows are independent is called nonsingular. We prove that every such finite system is solvable over every coefficient group.

Let GG be a group and let Fn=F(x1,…,xn)F_n=F(x_1,\ldots,x_n) be the free group on nn generators. A system over GG consists of words w1,…,wm∈G∗Fnw_1,\ldots,w_m\in G*F_n, interpreted as equations wi=1w_i=1. A solution over GG is a group HH containing an isomorphic copy of GG, together with elements h1,…,hn∈Hh_1,\ldots,h_n\in H at which all the words evaluate to the identity. Define the exponent-sum matrix A=(aij)∈Mm×n(Z)A=(a_{ij})\in M_{m\times n}(\mathbb{Z}) by the homomorphism

G∗Fn⟶Zn,G⟼0,xj⟼ej:wi⟼(ai1,…,ain).G*F_n\longrightarrow\mathbb{Z}^n,\qquad G\longmapsto0,\quad x_j\longmapsto e_j:\qquad w_i\longmapsto(a_{i1},\ldots,a_{in}).

The system is nonsingular if rank⁡QA=m\operatorname{rank}_{\mathbb{Q}} A=m. Here ⟨⟨w1,…,wm⟩⟩\langle\langle w_1,\ldots,w_m\rangle\rangle denotes normal closure in G∗FnG*F_n.

Theorem 1 (Nonsingular systems). Let GG be any group, let 1≤m≤n1\le m\le n be integers, and let w1,…,wm∈G∗Fnw_1,\ldots,w_m\in G*F_n. If their exponent-sum matrix has row rank mm over Q\mathbb{Q}, then the canonical homomorphism

G⟶(G∗Fn)/⟨⟨w1,…,wm⟩⟩G\longrightarrow(G*F_n)/\langle\langle w_1,\ldots,w_m\rangle\rangle

is injective. Equivalently, the system has a simultaneous solution in an overgroup of GG.

The equivalence follows from the universal property of the displayed quotient. If the canonical map is injective, the quotient itself supplies the overgroup and the images of the free generators supply the solution. Conversely, any solution induces a homomorphism from the quotient whose restriction to GG is injective. Thus solving the equations amounts to showing that their normal closure kills no nonidentity element of GG. Theorem 1 imposes no countability, finiteness, or torsion condition on GG. The same conclusion for arbitrary sets of equations and variables with independent exponent rows follows by a finite-relation argument in Remark 5.1.

History and significance

The general solvability assertion for nonsingular systems is commonly called Howie’s conjecture; see [9], Introduction for this formulation and terminology. Its one-equation, one-variable case is the Kervaire–Laudenbach conjecture: a nonzero exponent sum should guarantee coefficient injectivity. This is stronger than the nontriviality assertion usually called the Kervaire conjecture, which asks whether adjoining one generator and imposing one relation can annihilate a nontrivial group. Chen explains these formulations and their connection with high-dimensional knot groups [3], Section 1. Theorem 1 establishes the finite nonsingular-system assertion for arbitrary coefficient groups.

The foundational system theorem is due to Gerstenhaber and Rothaus [4]. For a square system over a compact connected Lie group, they compute the degree of the word map from the determinant of its exponent-sum matrix. A nonzero determinant therefore forces surjectivity and a simultaneous solution in that Lie group. Their finite-group theorem also produces a finite solution overgroup, using arithmetic specialization and reduction over finite fields. A nonsingular rectangular system reduces to the square case by retaining columns of a nonzero maximal minor and setting the other variables equal to the identity. Their argument established the usefulness of compact Lie groups, cohomology, and degree theory in a problem stated entirely in group-theoretic terms.

A different line of work uses the topology of relative presentations. Howie proved solvability of finite independent systems over locally indicable groups [6], Corollary 4.2; locally indicable means that every nontrivial finitely generated subgroup maps onto Z\mathbb{Z}. Klyachko’s theorem treats a single equation in one variable over a torsion-free group when the exponent sum is ±1\pm1 [8]. This unimodular hypothesis is stronger than a nonzero exponent sum. Chen later recovered that theorem by estimating the complexity of surfaces in HNN extensions [3], Theorem 6.9. These results illustrate two complementary sources of control: algebraic conditions on the coefficient group and topological restrictions on diagrams witnessing a kernel element.

The compact-unitary method also passes to metric ultraproducts. Pestov observed its application to hyperlinear groups [14], Corollary 10.4, and Nitsche and Thom give the nonsingular-system statement explicitly [12], Theorem 1.2 and Lemma 2.1. Here a hyperlinear, or Connes-embeddable, group is one that embeds into a metric ultraproduct of unitary groups with their normalized Hilbert–Schmidt metrics. Their stronger Theorem 1.3 replaces nonsingularity by vanishing second homology of a covering of the presentation complex obtained after deleting the coefficients. For the presentation complex itself, vanishing second homology is precisely independence of the exponent rows. Their covering criterion therefore reaches beyond nonsingular systems while retaining the hypothesis on the coefficient group.

Other developments control the solution overgroup more closely. Klyachko, Mikheenko, and Roman’kov obtain solutions within specified classes of solvable groups with suitable torsion-free abelian factors [9]. Ramirez-Côté and Wise use the Banach fixed-point theorem for groups embedded in Magnus-type power-series groups; their nonsingular-system construction inverts the determinant in the coefficient ring [15], Theorems 9–10. Such conclusions address additional structure beyond existence in an arbitrary overgroup.

The direct antecedent of our proof is the spectral-phase and planar argument in The Kervaire theorem for groups [13], Lemma 2.1, Lemma 3.1, and Theorem 4.1. Its fixed-plane incidence construction treats one unimodular relator. The extension below uses products of unitary groups and rational intersection classes to accommodate several relator types with different multiplicities. The cohomological background is classical [2]; related geometry of unitary eigenvalue strata and their intersection classes is developed by Nicolaescu [11]. In finite dimensions, spectral-phase subadditivity also follows from Thompson’s exponential formula [16]. Antezana, Larotonda, and Varela prove an approximate version, in operator norm, for embeddable II1\mathrm{II}_1 factors [1], Theorem 4.4. Our argument instead uses the faithful trace of the group von Neumann algebra [7], Section 3.3 and proves the required phase inequality directly, so the coefficients need no unitary approximation hypothesis.

Proof and technical contribution

The proof passes from a possible kernel relation to a planar surface. A finite product of conjugates of the relators is represented by disks whose boundaries read the relators or their inverses. Rectangular bands pair occurrences of the same variable with opposite signs. The remaining boundary arcs carry only coefficients from GG. The key assertion is a boundary obstruction: on a connected planar surface of this form, if all but one boundary words are trivial in GG, then the last is trivial as well. Removing the innermost components then proves injectivity.

Two independent estimates establish the obstruction. First, full row rank forces equal numbers of positive and negative disks of each relator type. Suppose these numbers are kik_i for the types present. A topological argument supplies unitary matrices mixing the kik_i copies of each type so that a finite-dimensional unitary matrix has at least 2∑iki2\sum_i k_i fixed directions. Second, a spectral-phase function measures the boundary words through the left regular representation of GG. Its subadditivity turns the fixed-space estimate into an upper bound. A direct calculation around the boundary circles attains that bound plus a nonnegative term measuring the possibly nontrivial boundary word. Euler characteristic makes the two constants agree, and faithfulness of the trace forces that word to be the identity.

The spectral-phase and planar-surface strategy comes from the one-relator unimodular theorem in [13]. We give the required arguments in full. The principal extension is the multi-block fixed-space theorem in Section 2. It concerns block matrices built from independent unitary variables of possibly different sizes, followed by arbitrary fixed unitary matrices. Full row rank of the signed block-multiplicity matrix forces a simultaneous lower bound on the dimensions of their fixed spaces. The proof pairs fixed-subspace incidence cycles with the product of the unitary-variable groups. Only the maximal-length exterior terms in cohomology contribute; their coefficient is a product of nested nonzero minors. This makes rational nonsingularity sufficient and gives an intersection principle that is independent of its application to equations over groups.

Section 3 establishes the phase inequality in the finite operator algebra of the left regular representation and computes the phase of a weighted cyclic shift. Section 4 combines these results to prove the boundary obstruction. Section 5 constructs the planar surfaces from an arbitrary kernel relation and completes the proof. The unitary topology is used only for finite complex matrices on the copy indices; the group coefficients are retained as operators throughout. This separation allows the argument to apply to arbitrary coefficient groups.

Simultaneous fixed spaces

The rank hypothesis enters the proof through the following finite-dimensional statement. A unitary matrix may occur in several blocks, with either exponent 1 or exponent −1-1. Only the signed numbers of its occurrences matter for the lower bound on the total fixed-space dimension. The proof extends the single-matrix incidence argument of [13] by using rational cohomology and nested nonzero minors to accommodate several unitary matrix variables.

Theorem 2 (Simultaneous fixed spaces). Let r≥0r \ge0 and n≥1n \ge1, and let k1,…,krk_1,\ldots,k_r be positive integers. For each j∈{1,…,n}j \in\{1,\ldots,n\}, let BjB_j be a finite ordered list of pairs (i,s)(i,s) with 1≤i≤r1 \le i \le r and s∈{1,−1}s \in\{1,-1\}; repetitions are allowed. Set

dj=∑(i,s)∈Bjki,bij=∑(i′,s)∈Bji′=is,B=(bij).d_j = \sum_{(i,s)\in B_j} k_i,\qquad b_{ij} = \sum_{\substack{(i',s)\in B_j\\ i'=i}} s,\qquad B=(b_{ij}).

Suppose that BB has row rank rr over Q\mathbb{Q}. For arbitrary Pj∈U(dj)P_j \in\mathrm{U}(d_j), define

D=∏i=1rU(ki),Fj(X)=diag⁡(i,s)∈Bj(Xis)Pj(X∈D).D = \prod_{i=1}^{r} \mathrm{U}(k_i), \qquad F_j(X) = \operatorname{diag}_{(i,s)\in B_j}(X_i^s)P_j \quad(X \in D).

Then there is X∈DX \in D such that

∑j=1ndim⁡Cker⁡(Fj(X)−Idj)≥∑i=1rki.(1)\sum_{j=1}^{n} \dim_{\mathbb{C}} \ker(F_j(X)-I_{d_j}) \ge\sum_{i=1}^{r} k_i. \tag*{(1)}

Here U(0)\mathrm{U}(0) and an empty product of groups are points, and the fixed-space dimension in dimension zero is zero.

Proof. Choosing target dimensions. The case r=0r=0 is immediate. Assume r>0r>0, and reorder the rows so that k1≥⋯≥krk_1 \ge\cdots\ge k_r. We first allocate the desired fixed-space dimensions among the maps FjF_j. There are distinct columns j1,…,jrj_1,\ldots,j_r such that

det⁡B[1,…,t∣j1,…,jt]≠0(1≤t≤r).(2)\det B[1,\ldots,t \mid j_1,\ldots,j_t] \ne0 \qquad(1 \le t \le r). \tag*{(2)}

Indeed, the first tt rows have rank tt. The columns chosen at the previous step remain independent on these rows, since their restrictions to the first t−1t-1 rows are independent; they can therefore be extended by one column. Define

ljt=kt(1≤t≤r),lj=0for unchosen columns.l_{j_t}=k_t \qquad(1 \le t \le r), \qquad l_j=0 \quad\text{for unchosen columns}.

These choices satisfy lj≤djl_j \le d_j. To see this for j=jtj=j_t, the nonzero minor in (2) implies that bi,jt≠0b_{i,j_t}\ne0 for some i≤ti\le t. The list BjtB_{j_t} thus contains a block of size ki≥ktk_i\ge k_t. Moreover,

∑jlj=∑iki,∑jlj2=∑iki2=dim⁡RD.(3)\sum_j l_j=\sum_i k_i, \qquad\sum_j l_j^2=\sum_i k_i^2=\dim_{\mathbb{R}}D. \tag*{(3)}

We will obtain dim⁡ker⁡(Fj(X)−Idj)≥lj\dim\ker(F_j(X)-I_{d_j})\ge l_j simultaneously. The equality of dimensions in (3) makes this an intersection problem of complementary dimensions. We develop the cohomology classes that detect the required intersection.

Primitive cohomology classes. All cohomology in this proof has coefficients in Q\mathbb{Q}. The classical computation of unitary-group cohomology [2] gives generators

H∗(U(d);Q)=Λ(e1,d,…,ed,d),∣ea,d∣=2a−1,(4)H^*(\mathrm{U}(d);\mathbb{Q})=\Lambda(e_{1,d},\ldots,e_{d,d}), \qquad|e_{a,d}|=2a-1, \tag*{(4)}

where Λ\Lambda denotes the exterior algebra. Compatibility means that inclusion as a coordinate block, with the identity on its complement, satisfies

ιk,d∗ea,d={ea,k,a≤k,0,a>k.(5)\iota_{k,d}^*e_{a,d}= \begin{cases} e_{a,k}, & a\le k,\\ 0, & a>k. \end{cases} \tag*{(5)}

We will choose the generators to satisfy this compatibility and to be primitive: if μ\mu is group multiplication, then

μ∗ea,d=ea,d⊗1+1⊗ea,d.\mu^*e_{a,d}=e_{a,d}\otimes1+1\otimes e_{a,d}.

Thus these classes add under multiplication, just as degree-one classes do on a torus.

For completeness, these properties follow together from the last-column bundle

U(d−1)⟶U(d)⟶S2d−1.\mathrm{U}(d-1) \longrightarrow\mathrm{U}(d) \longrightarrow S^{2d-1}.

Start with the degree-one generator of U(1)=S1\mathrm{U}(1)=S^1. For d≥2d\ge2, the base is simply connected, so the coefficient system in the multiplicative Serre spectral sequence is constant [10]. Its second page is

E2p,q=Hp(S2d−1;Q)⊗Hq(U(d−1);Q).E_2^{p,q}=H^p(S^{2d-1};\mathbb{Q})\otimes H^q(\mathrm{U}(d-1);\mathbb{Q}).

Its only nonzero columns are p=0,2d−1p=0,2d-1. The only possible differential between them is d2d−1d_{2d-1}, which sends fiber degree qq to degree q−2d+2q-2d+2. Each fiber generator has degree at most 2d−32d-3, so its differential has negative target degree and vanishes. The differential vanishes on all their products by the product rule for the differential. Consequently the spectral sequence collapses. Its edge maps show that restriction to U(d−1)\mathrm{U}(d-1) is an isomorphism in degrees below 2d−12d-1, and that the pullback of the sphere’s top class is nonzero. Lift each earlier generator in its degree, and take this sphere pullback as ed,de_{d,d}. Their exterior products give the basis in (4): they give that basis in the associated graded algebra supplied by the spectral sequence, and their squares vanish because they have odd degree and the coefficients are rational. This also proves compatibility with the standard inclusions.

To check primitivity, subtract the two summands on the right of (2.6) from its left side. The difference restricts to zero on either factor at the identity, so every term in its Künneth decomposition has positive degree in both factors. Each such degree is at most 2d−22d-2. Restriction to U(d−1)×U(d−1)\mathrm{U}(d-1)\times\mathrm{U}(d-1) is injective on all these bidegrees, by the edge-map isomorphism just proved. On this product the difference vanishes: for a<da<d this is the induction hypothesis, and for a=da=d it follows because ed,de_{d,d} is pulled back by the last-column projection, which is constant on the fiber. This proves (2.6); the case d=1d=1 is immediate. Coordinate-block inclusions in other positions are conjugate to the standard one, and conjugation is homotopic to the identity in the connected group U(d)\mathrm{U}(d). Finally, inversion sends ea,de_{a,d} to −ea,d-e_{a,d}, by pulling (2.6) back along X↦(X,X−1)X\mapsto(X,X^{-1}).

The fixed-space incidence class. We next construct a class that detects a fixed space of dimension at least ll, for 0≤l≤d0\le l\le d. Let

R(d,l)={(W,E):W∈U(d), E∈Gr⁡l(Cd), W∣E=IE},p(W,E)=W.R(d,l)=\{(W,E): W\in\mathrm{U}(d),\ E\in\operatorname{Gr}_l(\mathbb{C}^d),\ W|_E=I_E\},\qquad p(W,E)=W.

where Gr⁡l(Cd)\operatorname{Gr}_l(\mathbb{C}^d) is the complex Grassmannian of ll-dimensional subspaces. Over a subspace EE, the possible WW are precisely the unitary transformations of E⊥E^\perp, extended by the identity on EE. Orthonormal frame charts therefore make R(d,l)R(d,l) a smooth bundle over the Grassmannian with fiber U(d−l)\mathrm{U}(d-l). It is compact and without boundary, and

dim⁡RR(d,l)=2l(d−l)+(d−l)2=d2−l2.\dim_{\mathbb{R}}R(d,l)=2l(d-l)+(d-l)^2=d^2-l^2.

The complex orientation of the Grassmannian and an orientation of U(d−l)\mathrm{U}(d-l) orient this bundle: changes of frame act on the fiber by conjugation and preserve orientation. Choose orientations and let

α(d,l)=PD⁡(p∗[R(d,l)])∈Hl2(U(d);Q).\alpha(d,l)=\operatorname{PD}(p_*[R(d,l)])\in H^{l^2}(\mathrm{U}(d);\mathbb{Q}).

Here PD is Poincaré duality [5]. It applies to the homology class pushed forward by pp; the image of pp need not itself be a submanifold. Its image consists exactly of those WW with dim⁡ker⁡(W−Id)≥l\dim\ker(W-I_d)\ge l.

We recall the geometric interpretation of this class. If MM is a closed oriented manifold of dimension l2l^2 and f:M→U(d)f:M\to\mathrm{U}(d), then ⟨f∗α(d,l),[M]⟩\langle f^*\alpha(d,l),[M]\rangle is the intersection number of ff and pp. When these maps are transverse, it is the signed count of pairs (x,(W,E))(x,(W,E)) with f(x)=p(W,E)f(x)=p(W,E). This follows by intersecting f×pf\times p with the diagonal of U(d)×U(d)U(d)\times U(d): transversality makes its inverse image a compact oriented zero-manifold, whose signed count gives the intersection pairing. In particular, disjoint images give intersection number zero: disjoint maps from compact manifolds have disjoint sufficiently small transverse perturbations. The same interpretation holds for products of these maps and classes.

The part of α(d,l)\alpha(d,l) we need is

α(d,l)=cd,le1,d⋯el,d+terms with fewer than l generators,cd,l≠0.(6)\alpha(d,l)=c_{d,l}e_{1,d}\cdots e_{l,d}+\text{terms with fewer than }l\text{ generators},\qquad c_{d,l}\ne0. \tag*{(6)}

For l=0l=0, the map pp is the identity and we take α(d,0)=1\alpha(d,0)=1 and cd,0=1c_{d,0}=1. For l>0l>0, a product of uu distinct generators has degree at least 1+3+⋯+(2u−1)=u21+3+\cdots+(2u-1)=u^2. Since α(d,l)\alpha(d,l) has degree l2l^2, its expansion contains at most ll generators in each term, and the sole possible term of length ll is the one displayed in (6).

To show that its coefficient is nonzero, consider

f:U(l)⟶U(d),f(X)=X⊕(−Id−l).f:U(l)\longrightarrow U(d),\qquad f(X)=X\oplus(-I_{d-l}).

There is exactly one intersection pair for ff and pp: it has X=IlX=I_l, W=W0=Il⊕(−Id−l)W=W_0=I_l\oplus(-I_{d-l}), and E=E0=Cl⊕0E=E_0=\mathbb{C}^l\oplus0. Indeed, the second summand has no fixed vectors, while an ll-dimensional fixed space in the first summand forces X=IlX=I_l. This intersection is transverse. Identify the tangent space at W0W_0 with skew-Hermitian matrices by right multiplication by W0−1W_0^{-1}. Variations of ff fill the upper-left block; variations of pp in the fiber over E0E_0 fill the lower-right block. For a linear map Z:E0→E0⊥Z:E_0\to E_0^\perp, put

T=(0−Z∗Z0).T=\begin{pmatrix}0&-Z^*\\ Z&0\end{pmatrix}.

The path (etTW0e−tT,etTE0)(e^{tT}W_0e^{-tT},e^{tT}E_0) lies in R(d,l)R(d,l) and its image under pp has tangent vector

[T,W0]W0−1=2T.[T,W_0]W_0^{-1}=2T.

These vectors supply every off-diagonal direction. The unique intersection therefore has multiplicity 11 or −1-1. The map ff is homotopic to the coordinate inclusion, so (5) computes its pullback. Terms with an index above ll vanish; any remaining term with fewer than ll generators has degree less than dim⁡U(l)=l2\dim U(l)=l^2. Only the displayed term in (6) can account for the nonzero intersection number, proving cd,l≠0c_{d,l}\ne0. The same argument includes l=dl=d, when the complementary blocks are absent.

Computing the intersection number. We now apply these classes to the maps in the statement. Write ea,ki(i)e_{a,k_i}^{(i)} for the generator pulled back from the iith factor of DD. Primitivity, compatibility with block inclusions, and the inversion formula give

Fj∗ea,dj=∑i:ki≥abijea,ki(i)(1≤a≤dj).(7)F_j^*e_{a,d_j}=\sum_{i:k_i\ge a}b_{ij}e_{a,k_i}^{(i)}\qquad(1\le a\le d_j). \tag*{(7)}

To justify the use of multiplication here, write a block diagonal map as the product of maps that act on one block and are the identity on all others. Each positive occurrence contributes one generator and each negative occurrence its negative. Right multiplication by PjP_j is homotopic to the identity and has no effect on cohomology.

Consider the top-degree class

Ω=∏j=1nFj∗α(dj,lj)∈H∑iki2(D;Q).\Omega=\prod_{j=1}^{n}F_j^*\alpha(d_j,l_j)\in H^{\sum_i k_i^2}(D;\mathbb{Q}).

The exterior algebra H∗(D;Q)H^*(D;\mathbb{Q}) has altogether ∑iki\sum_i k_i generators, all of positive degree. A nonzero monomial of top degree must contain every one of them. By (7), pullback preserves the number of generators in each term, unless that term vanishes. It follows that any use of a shorter term from (6) contributes zero to Ω\Omega. Thus we need only compute the product of the displayed leading terms.

For each a∈{1,…,k1}a \in\{1,\ldots,k_1\}, put ta=∣{i:ki≥a}∣t_a = |\{i:k_i \ge a\}|. The rows with ki≥ak_i \ge a are 1,…,ta1,\ldots,t_a, and the selected columns with lj≥al_j \ge a are j1,…,jtaj_1,\ldots,j_{t_a}. Grouping the exterior factors of degree 2a−12a-1 together gives

∏t=1ta(∑i=1tabi,jtea,ki(i))=det⁡B[1,…,ta∣j1,…,jta]∏i=1taea,ki(i).\prod_{t=1}^{t_a}\left(\sum_{i=1}^{t_a}b_{i,j_t}e_{a,k_i}^{(i)}\right)=\det B[1,\ldots,t_a\mid j_1,\ldots,j_{t_a}]\prod_{i=1}^{t_a}e_{a,k_i}^{(i)}.

Consequently the coefficient of the top monomial in Ω\Omega, up to an overall sign, is

(∏j=1ncdj,lj)∏a=1k1det⁡B[1,…,ta∣j1,…,jta].(8)\left(\prod_{j=1}^{n}c_{d_j,l_j}\right)\prod_{a=1}^{k_1}\det B[1,\ldots,t_a\mid j_1,\ldots,j_{t_a}]. \tag*{(8)}

Every factor is nonzero, by (2) and (6). Hence ⟨Ω,[D]⟩≠0\langle\Omega,[D]\rangle\ne0.

Finally, let F=(F1,…,Fn)F=(F_1,\ldots,F_n) and let ptot=∏jpp_{\mathrm{tot}}=\prod_j p be the map from ∏jR(dj,lj)\prod_j R(d_j,l_j) to ∏jU(dj)\prod_j U(d_j). Up to the harmless orientation sign, Ω\Omega is the pullback by FF of the Poincaré dual of the cycle defined by ptotp_{\mathrm{tot}}. Its nonzero evaluation implies that these two maps have intersecting images. At an intersection, Fj(X)F_j(X) fixes an ljl_j-plane for every jj, which gives (1) by (3). □\square

Spectral phase

The planar argument will associate unitary operators to the boundary circles of a surface. We need a nonnegative numerical invariant that detects the identity, is subadditive under multiplication, and can be calculated on a weighted cyclic shift. We develop these properties in the operator algebra of an arbitrary group, following the spectral-phase argument of [13].

The trace and the choice of phase

Let GG be any group, and let {δx:x∈G}\{\delta_x:x\in G\} be the standard orthonormal basis of ℓ2(G)\ell^2(G). Define left and right translations by

L(a)δx=δax,Raδx=δxa−1.L(a)\delta_x=\delta_{ax},\qquad R_a\delta_x=\delta_{xa^{-1}}.

They commute. We work in the von Neumann algebra

M={T∈B(ℓ2(G)):TRa=RaT for every a∈G},\mathcal{M}=\{T\in B(\ell^2(G)):TR_a=R_aT\text{ for every }a\in G\},

which contains every L(a)L(a). For a positive integer dd, regard Md(M)M_d(\mathcal{M}) as operators on Cd⊗ℓ2(G)\mathbb{C}^d\otimes\ell^2(G) and set

τd(T)=∑j=1d(Tjj)1,1,τ=τ1.(9)\tau_d(T)=\sum_{j=1}^{d}(T_{jj})_{1,1},\qquad\tau=\tau_1. \tag*{(9)}

Here (Tjj)1,1(T_{jj})_{1,1} is the coefficient of δ1\delta_1 in Tjjδ1T_{jj}\delta_1. The trace is unnormalized: τd(I)=d\tau_d(I)=d. For the group von Neumann algebra and this trace construction, see [7].

We record why this is a faithful normal trace, including when GG is uncountable. If T∈MT\in\mathcal{M} and

Tδ1=∑xtxδx,T\delta_1=\sum_x t_x\delta_x,

commutation with right translations gives

Tx,y=txy−1.T_{x,y}=t_{xy^{-1}}.

@ Bibliography keys For A,B∈MA, B \in\mathcal{M} with corresponding coefficient families a,ba,b,

τ(AB)=∑x∈Gax−1bx=∑x∈Gbx−1ax=τ(BA).\tau(AB) = \sum_{x \in G} a_{x^{-1}}b_x = \sum_{x \in G} b_{x^{-1}}a_x = \tau(BA).

These sums converge absolutely by Cauchy–Schwarz; each square-summable family has countable support. Summing over the matrix indices proves traciality of τd\tau_d. Positivity and normality follow from its expression as a finite sum of vector functionals. If T≥0T \geq0 and τd(T)=0\tau_d(T) = 0, then

0=∑j=1d∥T1/2(ej⊗δ1)∥2.0 = \sum_{j=1}^{d}\left\|T^{1/2}(e_j \otimes\delta_1)\right\|^2.

The square root commutes with all simultaneous right translations, so it kills ej⊗δxe_j \otimes\delta_x for every j,xj,x. Thus T=0T = 0, proving faithfulness.

The bounded Borel functional calculus of a normal operator in Md(M)M_d(\mathcal{M}) remains in this algebra. For a unitary UU, its spectral projections define a finite measure μU(E)=τd(1E(U))\mu_U(E) = \tau_d(1_E(U)) on the unit circle T\mathbb{T}, and

τd(f(U))=∫Tf dμU,μU(T)=d.(10)\tau_d(f(U)) = \int_{\mathbb{T}} f \,d\mu_U,\qquad\mu_U(\mathbb{T}) = d. \tag*{(10)}

In particular, uniformly bounded pointwise convergence of Borel functions implies convergence of these trace evaluations.

Choose the phase with its cut at 11 by

h(e2πit)=t−⌊t⌋,ℓ(U)=τd(h(U)).h(e^{2\pi i t}) = t - \lfloor t \rfloor,\qquad\ell(U) = \tau_d(h(U)).

Thus h(1)=0h(1) = 0, and hh is right-continuous in the angle. Since h(U)≥0h(U) \geq0 and U=exp⁡(2πih(U))U = \exp(2\pi i h(U)), faithfulness gives

ℓ(U)≥0,ℓ(U)=0⟺U=I.(11)\ell(U) \geq0,\qquad\ell(U) = 0 \Longleftrightarrow U = I. \tag*{(11)}

The phase is invariant under unitary conjugation and additive on direct sums. On scalar matrices, viewed as matrices over M\mathcal{M}, it is the sum of the phases of the eigenvalues, counted with multiplicity.

Subadditivity

Lemma 3.1 (Spectral-phase subadditivity). For every group GG, every positive integer dd, and all unitaries U,V∈Md(M)U,V \in M_d(\mathcal{M}),

ℓ(UV)≤ℓ(U)+ℓ(V).\ell(UV) \leq\ell(U) + \ell(V).

Proof. Put D=h(V)D = h(V) and join UU to UVUV by the unitary path

Wu=Ue2πiuD,0≤u≤1.W_u = Ue^{2\pi i uD},\qquad0 \leq u \leq1.

We first calculate the change of a smooth function along this path, and then approximate the discontinuous function hh from the correct side of its cut. For f∈C∞(T)f \in C^\infty(\mathbb{T}) write f′(e2πit)=ddtf(e2πit)f'(e^{2\pi i t}) = \frac{d}{dt}f(e^{2\pi i t}). We claim that

dduτd(f(Wu))=τd(f′(Wu)D).\frac{d}{du}\tau_d(f(W_u)) = \tau_d(f'(W_u)D).

Indeed, Wu′=2πiWuDW'_u = 2\pi i W_uD. The product and inverse rules, followed by cyclic permutation inside the trace, give for every integer kk

dduτd(Wuk)=2πikτd(WukD).\frac{d}{du}\tau_d(W_u^k) = 2\pi i k\tau_d(W_u^kD).

This proves (3.5) for Laurent polynomials. For smooth ff, its Fourier coefficients satisfy ∑k∈Z(1+∣k∣)∣f^(k)∣<∞\sum_{k\in\mathbb{Z}}(1+|k|)|\widehat{f}(k)|<\infty, while ∥dduWuk∥≤2π∣k∣∥D∥\left\|\frac{d}{du}W_u^k\right\|\le2\pi|k|\|D\|. Thus the Fourier series and its differentiated series converge uniformly in operator norm, proving (3.5) in general.

If ff is real-valued and f′≤1f'\le1, then I−f′(Wu)≥0I-f'(W_u)\ge0. Although DD need not commute with WuW_u, traciality gives the required inequality:

τd(D)−τd(f′(Wu)D)=τd(D1/2(I−f′(Wu))D1/2)≥0.\tau_d(D)-\tau_d(f'(W_u)D)=\tau_d\left(D^{1/2}(I-f'(W_u))D^{1/2}\right)\ge0.

Integration of (3.5) therefore yields

τd(f(UV))−τd(f(U))≤τd(D).(12)\tau_d(f(UV))-\tau_d(f(U))\le\tau_d(D). \tag*{(12)}

For 0<ε<10<\varepsilon<1, choose a nonnegative smooth function ρε\rho_\varepsilon supported in (0,ε)(0,\varepsilon) with integral one, and define

fε(e2πit)=∫0ερε(s)h(e2πi(t+s)) ds.f_\varepsilon(e^{2\pi i t})=\int_0^\varepsilon\rho_\varepsilon(s)h(e^{2\pi i(t+s)})\,ds.

This is a smooth function on the circle, with 0≤fε≤10\le f_\varepsilon\le1. For δ>0\delta>0 the fractional-part function satisfies

h(e2πi(t+δ))−h(e2πit)≤δ.h(e^{2\pi i(t+\delta)})-h(e^{2\pi i t})\le\delta.

Averaging this inequality and taking the difference quotient shows that fε′≤1f_\varepsilon'\le1. Right-continuity gives fε(λ)→h(λ)f_\varepsilon(\lambda)\to h(\lambda) for every λ∈T\lambda\in\mathbb{T}. At the cut this follows explicitly from 0≤fε(1)≤ε0\le f_\varepsilon(1)\le\varepsilon; symmetric smoothing would not have this property. Apply (12) to fεf_\varepsilon and pass to the limit by (10). Since τd(D)=ℓ(V)\tau_d(D)=\ell(V), the result is the stated inequality.

The same right-continuity also gives, for any fixed c≥0c\ge0,

lim⁡ε↓0ℓ(e2πicεU)=ℓ(U).\lim_{\varepsilon\downarrow0}\ell(e^{2\pi i c\varepsilon}U)=\ell(U).

In the opposite direction the identity has a jump: ℓ(e−2πiεId)=d(1−ε)\ell(e^{-2\pi i\varepsilon I_d})=d(1-\varepsilon) for 0<ε<10<\varepsilon<1. These two behaviors at the cut will distinguish the possibly nontrivial boundary circle from the trivial ones in the planar argument.

The phase of a weighted cycle

Subadditivity will give an upper bound on the operator associated to a surface. The following exact calculation will express its phase in terms of the boundary words.

Lemma 3.2 (Weighted-cycle formula). Let aa be a positive integer, let Q1,…,Qa∈MQ_1,\ldots,Q_a\in\mathcal{M} be unitary, and define T∈Ma(M)T\in M_a(\mathcal{M}) by

T(et⊗v)=et+1⊗Qtv(1≤t≤a),ea+1=e1.T(e_t\otimes v)=e_{t+1}\otimes Q_t v\quad(1\le t\le a),\qquad e_{a+1}=e_1.

If C=QaQa−1⋯Q1C=Q_aQ_{a-1}\cdots Q_1 is the product around the cycle starting at the first coordinate, then

ℓ(T)=a−12+τ(h(C)).\ell(T)=\frac{a-1}{2}+\tau(h(C)).

Proof. Put ζ=e2πi/a\zeta=e^{2\pi i/a}. Conjugating TT by the scalar diagonal matrix with entries 1,ζ,…,ζa−11,\zeta,\ldots,\zeta^{a-1} gives ζT\zeta T. Hence ℓ(ζjT)=ℓ(T)\ell(\zeta^jT)=\ell(T) for every integer jj. The scalar identity

∑j=0a−1h(ζjλ)=a−12+h(λa),λ∈T,\sum_{j=0}^{a-1}h(\zeta^j\lambda)=\frac{a-1}{2}+h(\lambda^a),\qquad\lambda\in\mathbb{T},

follows by listing the fractional parts of the aa equally spaced angles; it holds also when one of them is zero. Applying bounded Borel functional calculus and taking the unnormalized trace gives

aℓ(T)=a(a−1)2+τa(h(Ta)).a\ell(T)=\frac{a(a-1)}{2}+\tau_a(h(T^a)).

The operator TaT^a is diagonal. Its first diagonal entry is CC, and the other entries are the products around the same cycle with different starting coordinates. Consecutive such products are conjugate by the corresponding QtQ_t, so every diagonal entry has phase trace τ(h(C))\tau(h(C)). Thus τa(h(Ta))=aτ(h(C))\tau_a(h(T^a))=a\tau(h(C)), proving (3.8).

The planar boundary obstruction

We now combine the fixed-space theorem with the phase inequality. The result is a statement about a planar surface assembled from copies of the relators: if every boundary component but one reads the identity in GG, then the remaining component does too. This extends the one-relator boundary argument of [13], Theorem 4.1 and Section 4; the new fixed-space input allows the different relator types to be treated simultaneously.

Cyclically conjugating the words does not change their normal closure or their exponent sums. We may therefore write

wi=xνi,1si,1gi,1⋯xνi,Nisi,Nigi,Ni,gi,t∈G,si,t∈{1,−1},1≤νi,t≤n.w_i=x_{\nu_{i,1}}^{s_{i,1}}g_{i,1}\cdots x_{\nu_{i,N_i}}^{s_{i,N_i}}g_{i,N_i},\qquad g_{i,t}\in G,\quad s_{i,t}\in\{1,-1\},\quad1\leq\nu_{i,t}\leq n.

Here each occurrence is a single variable letter, coefficients may be the identity, and Ni≥1N_i\geq1 because the iith row of AA is nonzero.

A word disk of type ii and sign ++ or −- is an oriented disk whose boundary reads wiw_i or wi−1w_i^{-1}, respectively. Its variable letters occupy disjoint closed intervals called slots, separated by nonempty corner arcs carrying the coefficients. On both signs of disk, index the slots by the positions tt in (4.1). Indices are cyclic modulo NiN_i. The boundary data are then

disk signsuccessor of ttvariable at ttcorner after tt
++t+1t+1xνi,tsi,tx_{\nu_{i,t}}^{s_{i,t}}gi,tg_{i,t}
−-t−1t-1xνi,t−si,tx_{\nu_{i,t}}^{-s_{i,t}}gi,t−1−1g_{i,t-1}^{-1}

Table 4.2.

The exponent si,ts_{i,t} or −si,t-s_{i,t} in this table is the slot’s traversal sign. Attach rectangular bands to pairs of slots so that the orientations extend across the bands. Every slot is used once, and paired slots must have the same variable index and opposite traversal signs. The remaining boundary consists of corner arcs and band sides; its word in GG is obtained by reading the corner labels, with no label contributed by a band side. Changing the starting corner conjugates this word, so its triviality is well defined.

Proposition 4.1 (Planar boundary obstruction). Let w1,…,wm∈G∗Fnw_1,\ldots,w_m\in G*F_n have exponent-sum matrix of row rank mm over Q\mathbb{Q}. Let Σ\Sigma be a connected oriented surface of genus zero obtained from a nonempty finite collection of signed word disks for these words by attaching bands that pair every slot exactly once, always with the same variable index and opposite traversal signs. If the words on all boundary components except possibly one are trivial in GG, then every boundary word is trivial in GG.

Proof. The proof estimates one unitary operator in two ways. Mixing disks of the same type gives an upper bound for its phase. Reading the operator around the boundary gives an exact phase formula. A choice of scalar phases at the end makes the two expressions differ only by the phase of the possibly nontrivial boundary word.

Disk counts and boundary cycles. Let pip_i and mim_i be the numbers of positive and negative disks of type ii. Each band pairs opposite occurrences of the same variable, whence

∑i(pi−mi)aij=0(1≤j≤n).\sum_i (p_i-m_i)a_{ij}=0 \qquad(1\le j\le n).

Row independence gives pi=mip_i=m_i. In the rest of this proof, let II be the set of types that occur and write ki=pi=mi≥1k_i=p_i=m_i\ge1 for i∈Ii\in I. Set

K=∑i∈Iki,b=∑i∈INiki.K=\sum_{i\in I} k_i,\qquad b=\sum_{i\in I}N_i k_i.

Thus Σ\Sigma has 2K2K disks, 2b2b slots, and bb bands.

Let D\mathcal{D} be the set of slots. The successor permutation σ\sigma follows the oriented boundary of each disk; the fixed-point-free involution θ\theta pairs the ends of each band. For d∈Dd\in\mathcal{D}, write cd∈Gc_d\in G for the label on the corner from the end of dd to the start of σd\sigma d, as specified in eq:4.2. Starting at the start of slot dd, a boundary path crosses a band side to the end of θd\theta d, then follows its corner to the start of σθd\sigma\theta d; see Figure 1. Consequently the cycles zz of σθ\sigma\theta are exactly the boundary components. If z=(d1,…,da)z=(d_1,\ldots,d_a) is written in this order, put

One boundary step: a band side connects d to θd, followed by corner cθd to the start of σθd.

Figure 1. One boundary step, with the attached slots drawn as thick gray intervals. The arrow crosses a band side and then follows the corner after the partner slot. This gives the permutation σθ\sigma\theta and the label cθdc_{\theta d}. The drawing is local: the disk signs and types are unspecified, and the two band ends may lie on the same disk.

a(z)=a,rz=cθd1⋯cθda.a(z)=a,\qquad r_z=c_{\theta d_1}\cdots c_{\theta d_a}.

The nonempty corner arcs ensure that there is at least one boundary component. If qq denotes their number, genus zero gives

2−q=χ(Σ)=2K−b,∑za(z)=2b.2-q=\chi(\Sigma)=2K-b,\qquad\sum_z a(z)=2b.

Distributing phases among boundary components. For any real numbers (ηz)z(\eta_z)_z with ∑zηz=0\sum_z\eta_z=0, we can choose real numbers (γd)d∈D(\gamma_d)_{d\in\mathcal{D}} such that

γθd=−γd,∑d∈zγd=ηz.\gamma_{\theta d}=-\gamma_d,\qquad\sum_{d\in z}\gamma_d=\eta_z.

To see this, form a graph whose vertices are the boundary cycles and whose edges are the pairs {d,θd}\{d,\theta d\}, with an end incident to the cycle containing that slot. Loops and multiple edges are allowed. Connectedness of Σ\Sigma means that σ\sigma and θ\theta act transitively on D\mathcal{D}. Since σθ\sigma\theta and θ\theta generate the same permutation group, this graph is connected. Choose a spanning tree and set the values on all other edges to zero. At a leaf, set the value at its end of the remaining edge equal to its prescribed sum and the value at the other end to its negative; remove the leaf and adjust the prescribed sum at its neighbor. The zero total ensures the last vertex is satisfied. This proves (4.6), including the one-vertex case.

Use the algebra M2b(M)M_{2b}(\mathcal{M}) of Section 3, acting on CD⊗ℓ2(G)\mathbb{C}^{D} \otimes\ell^2(G), and define

S(ed⊗v)=eσd⊗L(cd−1)v,Θ(ed⊗v)=e2πiγdeθd⊗v.\begin{aligned} S(e_d \otimes v) &= e_{\sigma d} \otimes L(c_d^{-1})v,\\ \Theta(e_d \otimes v) &= e^{2\pi i\gamma_d}e_{\theta d} \otimes v. \end{aligned}

Both operators are unitary, and (4.6) gives Θ2=I\Theta^2=I. We identify scalar matrices with their tensor products with the identity on ℓ2(G)\ell^2(G); their trace τ2b\tau_{2b} is therefore the ordinary matrix trace.

The upper phase bound. We will choose a scalar unitary involution JJ and use the factorization SΘ=(SJ)(JΘ)S\Theta=(SJ)(J\Theta). We arrange that SJSJ is an involution with known phase, and use the fixed-space theorem to give JΘJ\Theta many fixed vectors. For each i∈Ii\in I, order the kik_i disks of each sign. At each position tt, their slots form two coordinate spaces, each identified with Cki\mathbb{C}^{k_i}. Given Xi∈U(ki)X_i\in U(k_i), define a scalar unitary involution JJ by sending the positive array at (i,t)(i,t) to the negative array by XiX_i, and the negative array back by Xi−1X_i^{-1}. The same XiX_i is used at every position of type ii.

The corner conventions give

JSJ=S−1.JSJ=S^{-1}.

Indeed, on positive arrays SS moves from tt to t+1t+1 with coefficient L(gi,t−1)L(g_{i,t}^{-1}), whereas on negative arrays it moves from tt to t−1t-1 with coefficient L(gi,t−1)L(g_{i,t-1}). Applying JJ on either side interchanges these rules; its factors XiX_i and Xi−1X_i^{-1} cancel because the coefficients are constant across each array and commute with scalar matrices. Thus SJSJ is a unitary involution. It exchanges the two disk signs, so its diagonal blocks, and hence its trace, are zero. Its −1-1 spectral projection (I−SJ)/2(I-SJ)/2 has trace bb, giving

ℓ(SJ)=b2.(13)\ell(SJ)=\frac{b}{2}. \tag*{(13)}

We now choose the XiX_i so that JΘJ\Theta has many fixed vectors. Regroup the scalar coordinates by variable index jj and traversal sign into spaces Vj,+V_{j,+} and Vj,−V_{j,-}. For each occurrence (i,t)(i,t) with νi,t=j\nu_{i,t}=j, each space has one block of dimension kik_i: the positive or negative disk array is selected according to its traversal sign. In particular both spaces have dimension

dj=∑i∈I, 1≤t≤Niνi,t=jki,∑jdj=b.d_j=\sum_{\substack{i\in I,\ 1\leq t\leq N_i\\ \nu_{i,t}=j}} k_i,\qquad\sum_j d_j=b.

Both JJ and Θ\Theta exchange Vj,+V_{j,+} and Vj,−V_{j,-}. Under the array coordinates, let Pj∈U(dj)P_j\in U(d_j) represent Θ:Vj,+→Vj,−\Theta:V_{j,+}\to V_{j,-}. On the return map J:Vj,−→Vj,+J:V_{j,-}\to V_{j,+}, the block at (i,t)(i,t) is Xi−si,tX_i^{-s_{i,t}}. Hence

(JΘ)∣Vj,+=diag⁡(i,t): νi,t=j(Xi−si,t)Pj.(J\Theta)|_{V_{j,+}}=\operatorname{diag}_{(i,t):\,\nu_{i,t}=j}\left(X_i^{-s_{i,t}}\right)P_j.

The signed block counts here are ∑t:νi,t=j(−si,t)=−aij\sum_{t:\nu_{i,t}=j}(-s_{i,t})=-a_{ij}. The matrix formed by the rows of −A-A indexed by II has full row rank, so Theorem 2 supplies XiX_i for which

dim⁡Cker⁡((JΘ−I)∣⨁jVj,+)≥K.\dim_{\mathbb{C}}\ker\left((J\Theta-I)\big|_{\bigoplus_j V_{j,+}}\right)\geq K.

There are equally many fixed vectors on ⨁jVj,−\bigoplus_j V_{j,-}: JJ interchanges the two sums and satisfies J(JΘ)J=(JΘ)−1J(J\Theta)J=(J\Theta)^{-1}. Thus the full fixed-space dimension ff of the scalar matrix JΘJ\Theta is at least 2K2K.

The same conjugacy makes its eigenvalue multiset invariant under λ↦λ−1\lambda\mapsto\lambda^{-1}. Since h(λ)+h(λ−1)=1h(\lambda)+h(\lambda^{-1})=1 for λ≠1\lambda\ne1 and is zero at 11,

2ℓ(JΘ)=2b−f≤2b−2K.2\ell(J\Theta)=2b-f\le2b-2K.

Using SΘ=(SJ)(JΘ)S\Theta=(SJ)(J\Theta), (13) and Lemma 3.1 now give

ℓ(SΘ)≤32b−K.(14)\ell(S\Theta)\le\frac{3}{2}b-K. \tag*{(14)}

This holds for every phase distribution in (4.6); the auxiliary choice of JJ may depend on that distribution.

The boundary formula and the cut limit. It remains to read the left side of (14) from the boundary. On the coordinates of a cycle z=(d1,…,da)z=(d_1,\ldots,d_a) of σθ\sigma_\theta, the operator SΘS\Theta is a weighted cyclic shift TzT_z whose step from dtd_t to dt+1d_{t+1} has weight e2πiγdtL(cθdt−1)e^{2\pi i\gamma_{d_t}L(c_{\theta d_t}^{-1})}. The circuit starting at d1d_1 therefore has weight

e2πi∑d∈zγdL(cθda−1⋯cθd1−1)=e2πiηzL(rz−1).e^{2\pi i\sum_{d\in z}\gamma_dL(c_{\theta d_a}^{-1}\cdots c_{\theta d_1}^{-1})}=e^{2\pi i\eta_zL(r_z^{-1})}.

Lemma 3.2 yields the exact formula

ℓ(Tz)=a(z)−12+τ(h(e2πiηzL(rz−1))).\ell(T_z)=\frac{a(z)-1}{2}+\tau\bigl(h(e^{2\pi i\eta_z}L(r_z^{-1}))\bigr).

Choose the possibly exceptional component z0z_0, and for ϵ>0\epsilon>0 set

ηz=−ϵ(z≠z0),ηz0=(q−1)ϵ.\eta_z=-\epsilon\quad(z\ne z_0),\qquad\eta_{z_0}=(q-1)\epsilon.

These phases sum to zero, so (14) applies. For z≠z0z\ne z_0, the assumption rz=1r_z=1 makes the last term of (4.12) equal to 1−ϵ1-\epsilon when 0<ϵ<10<\epsilon<1. For z0z_0, the scalar phase approaches zero from the nonnegative side. Thus h(e2πi(q−1)ϵλ)→h(λ)h(e^{2\pi i(q-1)\epsilon}\lambda)\to h(\lambda) at every point of the unit circle, including λ=1\lambda=1. Bounded convergence in the spectral measure gives

τ(h(e2πi(q−1)ϵL(rz0−1)))⟶τ(h(L(rz0−1))).\tau\bigl(h(e^{2\pi i(q-1)\epsilon}L(r_{z_0}^{-1}))\bigr)\longrightarrow\tau\bigl(h(L(r_{z_0}^{-1}))\bigr).

This also covers q=1q=1, when the scalar phase is identically zero. Summing (4.12) over all cycles and using (4.5), we obtain

lim⁡ϵ↓0ℓ(SΘ)=2b−q2+q−1+τ(h(L(rz0−1)))=32b−K+τ(h(L(rz0−1))).\begin{aligned} \lim_{\epsilon\downarrow0}\ell(S\Theta)&=\frac{2b-q}{2}+q-1+\tau\bigl(h(L(r_{z_0}^{-1}))\bigr)\\ &=\frac{3}{2}b-K+\tau\bigl(h(L(r_{z_0}^{-1}))\bigr). \end{aligned}

The upper bound (14) forces the last trace to be zero. Faithfulness of τ\tau and nonnegativity of hh imply L(rz0−1)=IL(r_{z_0}^{-1})=I. Applying this operator to δ1\delta_1 gives rz0=1r_{z_0}=1, as required.

From a kernel relation to a solution

We now apply Proposition 4.1 to prove Theorem 1. The required surfaces must be embedded in a disk: an arbitrary pairing of variable occurrences would not guarantee planarity. Following the construction in [13], Section 5, we obtain the bands as level arcs of the variable-circle maps on a punctured disk. This also ensures that the coefficient map is constant along every band.

A punctured disk for the relation

Let g∈Gg \in G lie in the normal closure of the relators. By definition, there is a finite expression in G∗FnG * F_n

g=∏α=1ubαwiαϵαbα−1,bα∈G∗Fn,ϵα∈{1,−1}.g = \prod_{\alpha=1}^{u} b_\alpha w_{i_\alpha}^{\epsilon_\alpha} b_\alpha^{-1}, \qquad b_\alpha\in G * F_n,\qquad\epsilon_\alpha\in\{1,-1\}.

If u=0u = 0, then g=1g = 1. Assume u>0u > 0. Use the cyclic representatives of the relators chosen in Section 4; conjugating a relator only changes the corresponding conjugators in (5.1).

Choose a based CW presentation space YY with π1(Y)=G\pi_1(Y) = G, and put

Y′=Y∨S11∨⋯∨Sn1.Y' = Y \vee S_1^1 \vee\cdots\vee S_n^1.

The added circles represent the variables, so π1(Y′)=G∗Fn\pi_1(Y') = G * F_n. Choose based loops in YY for the coefficients and for gg. No asphericity or finiteness property of YY is needed.

Let Δ\Delta be an oriented polygonal disk and remove the interiors of uu disjoint polygonal disks DαD_\alpha in its interior. On the resulting punctured disk

P=Δ∖⋃α=1uint⁡DαP = \Delta\setminus\bigcup_{\alpha=1}^{u} \operatorname{int} D_\alpha

there is a map f:P→Y′f:P \to Y' with these boundary values: the outer boundary reads gg in YY, and ∂Dα\partial D_\alpha, oriented as the boundary of the missing disk, reads wiαϵαw_{i_\alpha}^{\epsilon_\alpha}. Each variable letter is represented by one monotone traversal of its circle. The intervening coefficient intervals are mapped to the chosen loops in YY; retain an interval even when its label is the identity.

Here is an explicit reason that these values extend over PP. Cut PP along disjoint arcs from the outer boundary to the holes, and map those arcs to the loops for bαb_\alpha. Place their outer endpoints in a constant portion of the outer loop, with the holes encountered in reverse order after the outer traversal. The boundary of the cut disk reads

g(buwiuϵubu−1)−1⋯(b1wi1ϵ1b1−1)−1.g(b_u w_{i_u}^{\epsilon_u} b_u^{-1})^{-1}\cdots(b_1 w_{i_1}^{\epsilon_1} b_1^{-1})^{-1}.

It is nullhomotopic by (5.1). A nullhomotopy extends the map across the cut disk. The two copies of each cut arc have the same map with opposite traversal, so the extension descends to PP. Notice that the induced boundary orientation of PP on a hole is opposite to the missing-disk orientation; this explains the inverses in the displayed word.

Disjoint bands from circle levels

Let fj:P→Sj1=R/Zf_j:P \to S_j^1 = \mathbb{R}/\mathbb{Z} be the projection of ff that collapses all the other wedge summands, and let v:P→Yv:P \to Y be the coefficient projection. We construct bands on which vv is exactly constant, without changing the boundary values of any fjf_j.

Choose a common sufficiently fine finite triangulation of PP and piecewise affine circle maps f~j\widetilde f_j satisfying

f~j∣∂P=fj∣∂P,dist⁡(f~j(p),fj(p))<112(p∈P).\widetilde f_j|_{\partial P} = f_j|_{\partial P}, \qquad\operatorname{dist}(\widetilde f_j(p),f_j(p)) < \frac{1}{12} \quad(p \in P).

These approximations can be constructed directly. The boundary maps are already piecewise affine in the angular coordinate. Subdivide to respect their breakpoints and so finely that the image of each triangle under each fjf_j lies in a short arc of the circle. Lift that arc to R\mathbb{R} and interpolate its vertex values. The lifts on a shared edge differ by an integer, so the interpolants agree as circle maps. Uniform continuity gives the stated error bound.

For each jj, choose tj∈(1/3,2/3)t_j \in(1/3,2/3) that is not the image of any vertex under f~j\widetilde{f}_j. The level set f~j−1(tj)\widetilde{f}_j^{-1}(t_j) is a finite disjoint union of polygonal circles and properly embedded polygonal arcs. On each triangle it consists of line segments, and avoidance of vertices makes the segments join without branching. Since the approximation is fixed on the boundary, there is exactly one endpoint for each occurrence of xj±1x_j^{\pm1} on a hole and no endpoint on the outer boundary or on a coefficient interval.

These level sets are disjoint for different indices. Indeed, at a point of f~j−1(tj)\widetilde{f}_j^{-1}(t_j) the original fjf_j is at distance more than 1/41/4 from the circle basepoint. Thus ff lands in the jjth variable circle away from the wedge point, and all its other circle projections vanish. The same remains true in a neighborhood of the level set. In particular, vv is constant at the basepoint on that neighborhood. This argument uses the original wedge-valued map ff; the separately approximated maps need not themselves combine to a map into the wedge.

Discard the closed level curves and choose thin, mutually disjoint rectangular neighborhoods of the proper arcs inside these neighborhoods. They are the bands. Each meets the holes in small intervals inside the corresponding variable traversals, and every occurrence has exactly one such interval. Coorientation of a circle level makes the two endpoint crossing signs of an arc opposite with respect to the boundary orientation of PP. Reversing both orientations to the missing-disk orientations leaves the signs opposite. Hence each band pairs the same variable with opposite traversal signs, as required in Proposition 4.1.

Use the smaller attachment intervals as the slots. The pieces of a variable traversal outside its slot project constantly to YY, so shrinking the slots does not change any coefficient label between them. Moreover, vv is constant on every band. Reinsert the disks DαD_\alpha as the signed word disks. Their union with the bands is a compact surface embedded in the interior of Δ\Delta, with every slot paired. Every connected component is planar, and its boundary words are exactly the loops supplied by vv on those boundary circles.

Filling from the inside out

It remains to account for the possibility of several components nested inside one another. Each connected component has one exterior boundary circle whose Jordan disk contains the component; its other boundary circles enclose the inner complementary disks. The exterior Jordan disks of distinct components are disjoint or nested. Choose a component whose exterior disk is minimal under inclusion. Its inner complementary disks contain no other component, and hence contain no missing word disk. The coefficient map vv is already defined on each of them. Their boundary words are therefore trivial in GG.

Proposition 4.1 now implies that the exterior boundary word of the chosen component is also trivial. Its exact boundary loop extends to a map of its Jordan disk into YY. Replace vv inside that disk by this extension. The two maps agree on the boundary, so they paste continuously. No other component lies in the disk, and the coefficient labels and band constants on all remaining components are unchanged.

Repeating this finite procedure removes all components. The result is a map Δ→Y\Delta\to Y extending the original outer loop for gg. Consequently g=1g=1 in GG. This proves the injectivity in Theorem 1; its quotient supplies the simultaneous solution.

Remark 5.1. The finite statement also implies the version with arbitrary sets of variables and equations. Assume the exponent rows, each of finite support, are linearly independent over Q\mathbb{Q}. Any element of GG killed in the quotient has a normal-closure expression involving only finitely many relators and finitely many variables, including those in its conjugators. The corresponding finite exponent matrix still has independent rows, so Theorem 1 makes that element trivial.

References

  1. [1]J. Antezana, G. Larotonda, and A. Varela, Thompson-type formulae, J. Funct. Anal. 262 (2012), no. 4, 1515–1528. doi:10.1016/j.jfa.2011.11.011.DOI
  2. [2]A. Borel, Sur la cohomologie des espaces fibrés principaux et des espaces homogènes de groupes de Lie compacts, Ann. of Math. (2) 57 (1953), no. 1, 115–207. doi:10.2307/1969728.DOI
  3. [3]L. Chen, The Kervaire conjecture and the minimal complexity of surfaces, Trans. Amer. Math. Soc. 379 (2026), 587–626. arXiv:2302.09811.arxiv.org/abs/2302.09811
  4. [4]M. Gerstenhaber and O. S. Rothaus, The solution of sets of equations in groups, Proc. Nat. Acad. Sci. U.S.A. 48 (1962), no. 9, 1531–1533. doi:10.1073/pnas.48.9.1531.DOI
  5. [5]A. Hatcher, Algebraic Topology, Cambridge University Press, Cambridge, 2002.
  6. [6]J. Howie, On pairs of 2-complexes and systems of equations over groups, J. Reine Angew. Math. 324 (1981), 165–174. doi:10.1515/crll.1981.324.165.DOI
  7. [7]V. F. R. Jones, Von Neumann Algebras, lecture notes, 1 October 2009. https://math.berkeley.edu/~vfr/VonNeumann2009.pdf.
  8. [8]A. A. Klyachko, A funny property of sphere and equations over groups, Comm. Algebra 21 (1993), no. 7, 2555–2575. doi:10.1080/00927879308824692.DOI
  9. [9]A. A. Klyachko, M. A. Mikheenko, and V. A. Roman’kov, Equations over solvable groups, J. Algebra 638 (2024), 739–750. doi:10.1016/j.jalgebra.2023.10.004.
  10. [10]J. McCleary, A User’s Guide to Spectral Sequences, second ed., Cambridge Studies in Advanced Mathematics, vol. 58, Cambridge University Press, Cambridge, 2001.DOI
  11. [11]L. I. Nicolaescu, Schubert calculus on the Grassmannian of hermitian lagrangian spaces, Adv. Math. 224 (2010), no. 6, 2361–2434. doi:10.1016/j.aim.2010.02.003.DOI
  12. [12]M. Nitsche and A. Thom, Universal solvability of group equations, J. Group Theory 25 (2022), no. 1, 1–10. doi:10.1515/jgth-2019-0167.DOI
  13. [13]OpenAI, The Kervaire theorem for groups, OpenAI Math Release preprint OAI:The-Kervaire-Theorem-for-Groups-September-24-2026, 2026.
  14. [14]V. G. Pestov, Hyperlinear and sofic groups: a brief guide, Bull. Symbolic Logic 14 (2008), no. 4, 449–480. arXiv:0804.3968.arxiv.org/abs/0804.3968
  15. [15]A. Ramirez-Côté and D. T. Wise, The Kervaire conjecture via Banach fixed points in power series groups, Math. Z. 311 (2025), article 6. doi:10.1007/s00209-025-03753-3.
  16. [16]R. C. Thompson, Proof of a conjectured exponential formula, Linear Multilinear Algebra 19 (1986), no. 2, 187–197. doi:10.1080/03081088608817715.DOI

Paper details

Contents