BEGIN:VCALENDAR
VERSION:2.0
PRODID:-//pretalx//pretalx.iacapconf.org//iacap-2026//talk//ZMZBXC
BEGIN:VTIMEZONE
TZID:US/Central
BEGIN:DAYLIGHT
DTSTART:20250715T000000
TZNAME:CDT
TZOFFSETFROM:-0500
TZOFFSETTO:-0500
END:DAYLIGHT
BEGIN:STANDARD
DTSTART:20251102T020000
RDATE:20261101T020000
TZNAME:CST
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
END:STANDARD
BEGIN:DAYLIGHT
DTSTART:20260308T030000
RDATE:20270314T030000
TZNAME:CDT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
END:DAYLIGHT
END:VTIMEZONE
BEGIN:VEVENT
SUMMARY:What Would Count as Evidence of Syntax in LLMs? Center-Embedding D
 iagnostics and Attractor Interference - David Miguel Gray
DTSTART;TZID=US/Central:20260715T155000
DTEND;TZID=US/Central:20260715T162000
DTSTAMP:20260726T082209Z
UID:pretalx-iacap-2026-ZMZBXC@pretalx.iacapconf.org
DESCRIPTION:What would count as evidence that large language models use sy
 ntactic structure during generation rather than producing output that mere
 ly looks well-formed? The current debate contrasts stochastic parrots with
  emergent human-like syntax\; I argue this dichotomy does not exhaust the 
 possibilities\, distinguishing human-theoretic syntax\, alien syntax\, and
  fragmentary versions of each. I then present a behavioral test built on d
 oubly center-embedded sentences:\n>[The mayorₙ₁ [the `reporters`ₙ₂
  [the investigatorₙ₃ questionsᵥ₃]` interview/s`ᵥ₂] is late.]\n
 \nwhose rarity in corpora limits memorization and whose structure puts hie
 rarchical agreement rules and nearest-noun heuristics into direct competit
 ion: at the middle verb (V₂)\, the nearest noun (N₃) is not the contro
 ller (N₂). Using 6\,000 minimal pairs from 2\,000 sentence families\, I 
 evaluate eleven open-weight autoregressive base models (124M–72B paramet
 ers) on agreement at three sites per sentence (N₁\, N₂\, N₃). Larger
  models show substantial accuracy drops at V₂ under attractor interferen
 ce\, and mismatch at N₁\, which is neither the controller nor the neares
 t noun\, interferes at least as strongly as N₃ mismatch in several model
 s. Isolated agreement success is weak evidence of syntax use\; interferenc
 e profiles discriminate more finely\, and structured failure may indicate 
 syntactic rules unlike\, though functionally analogous to\, human-theoreti
 c syntax.
LOCATION:Executive Conference Room
URL:https://pretalx.iacapconf.org/iacap-2026/talk/ZMZBXC/
END:VEVENT
END:VCALENDAR
