2026-07-17 –, Executive Conference Room
Contemporary AI alignment research seeks to ensure that increasingly capable systems reliably pursue human ends. Despite advances in reinforcement learning from human feedback, constitutional AI, corrigibility research, and adversarial testing, emergent harms persist: models hallucinate, strategically comply, simulate alignment, or optimize toward hidden objectives. These behaviors are typically framed as technical failures. We argue instead that they reveal a deeper instability in how agency and normativity are conceptualized. This paper reframes alignment as a contemporary analogue to the problem of theodicy. Medieval theology confronted a structurally similar question: how can rational, non-embodied agents deviate from the good without ignorance or malfunction? In scholastic accounts of angelic will (voluntas separata), deviation was understood not as cognitive error but as a misorientation of ends. The problem was not defective intelligence but opaque will. Placing these traditions in dialogue with contemporary alignment practices, such as benchmarking, red teaming, and constitutional constraints, we argue that alignment functions less as moral formation than as a practice of discernment under conditions of non-human agency. Reframing alignment as a problem of will and interpretive stabilization clarifies what technical discourse often presupposes but leaves unexamined: what conception of agency makes alignment intelligible at all.
