2026-07-15 –, Executive Conference Room
What would count as evidence that large language models use syntactic structure during generation rather than producing output that merely looks well-formed? The current debate contrasts stochastic parrots with emergent human-like syntax; I argue this dichotomy does not exhaust the possibilities, distinguishing human-theoretic syntax, alien syntax, and fragmentary versions of each. I then present a behavioral test built on doubly center-embedded sentences:
[The mayorₙ₁ [the
reportersₙ₂ [the investigatorₙ₃ questionsᵥ₃]interview/sᵥ₂] is late.]
whose rarity in corpora limits memorization and whose structure puts hierarchical agreement rules and nearest-noun heuristics into direct competition: at the middle verb (V₂), the nearest noun (N₃) is not the controller (N₂). Using 6,000 minimal pairs from 2,000 sentence families, I evaluate eleven open-weight autoregressive base models (124M–72B parameters) on agreement at three sites per sentence (N₁, N₂, N₃). Larger models show substantial accuracy drops at V₂ under attractor interference, and mismatch at N₁, which is neither the controller nor the nearest noun, interferes at least as strongly as N₃ mismatch in several models. Isolated agreement success is weak evidence of syntax use; interference profiles discriminate more finely, and structured failure may indicate syntactic rules unlike, though functionally analogous to, human-theoretic syntax.
What would count as evidence that large language models use syntactic structure during generation, rather than producing output that merely looks well-formed? I argue that the current debate, organized around a contrast between stochastic parrots and emergent human-like syntax, does not exhaust the possibilities. After distinguishing human-theoretic syntax, alien syntax, and fragmentary versions of each, I present diagnostic tests that separate structure-sensitive behavior from surface heuristics.
Doubly center-embedded sentences are a useful stimulus class: their nested dependencies are rare in ordinary corpora, and correct agreement requires maintaining non-local dependencies across embedded clauses:
[The mayorₙ₁ [the
reportersₙ₂ [the investigatorₙ₃ questionsᵥ₃]interview/sᵥ₂] is late.]
Language models are known to make agreement-attraction errors consistent with a nearest-noun heuristic:
"The
keysₙ₁ in the cabinetₙ₂is/areon the desk"
matching the verb to the nearest noun rather than to its syntactic controller. Center-embedding puts the hierarchical rule and the heuristic into direct competition: at V₂, the nearest noun (N₃) is not the controller (N₂).
I introduce a dataset of 2,000 sentence families yielding 6,000 minimal pairs, testing agreement at three sites within the same sentence (N₁–copula, N₂–V₂, N₃–V₃) while manipulating whether N₁ and N₃ match N₂ in number. I evaluate eleven open-weight autoregressive base models (GPT-2, Pythia, Llama, and Qwen families) from 124 million to 72 billion parameters.
The results complicate two views. Against the inference from scale to syntactic competence, larger models show substantial accuracy drops at V₂ under attractor interference. Against the nearest-noun account, mismatch at N₁, the noun farthest from V₂ and irrelevant to its agreement, interferes at V₂ at least as strongly as mismatch at the adjacent N₃ in several models, and the two mismatches interact.
The upshot is both methodological and philosophical. Success at isolated agreement tasks is weak evidence of syntax use; interference profiles under center-embedding discriminate more finely among hypotheses. And structured failure on human-syntactic tasks may indicate syntactic rules that differ from human-theoretic syntax while playing a functionally similar role in generation.
