International Association for Computing and Philosophy - Annual Conference 2026

Colin Allen


Session

07-15
14:00
30min
Assessing whether LLMs Provide Acceptable Substitutes for Human Judgments in a Digital Philosophy Project
Colin Allen, Nazhah Mir

We consider the application of LLMs for a digital humanities project. The Internet Philosophy Ontology (InPhO) project (inphoproject.org) organizes concepts from the Stanford Encyclopedia of Philosophy (SEP) into a taxonomic hierarchy supplemented by non-taxonomic relationships. The InPhO concept graph is inferred from automated statistical analysis of SEP content and human judgments about concept relatedness. The need to collect human judgments made it hard to scale up the original project. The appearance of LLMs raises the question of whether LLM-generated judgments could be substituted for human judgments. We tested this idea using five different LLMs prompted to adopt different levels of philosophical expertise, When prompted to adopt higher expertise levels, two of the LLMs provided closer matches to human judgments at the corresponding levels than the other models. We also found that most of the LLMs showed less variance when prompted to respond at the level of a philosophy doctoral student, mirroring the finding in the original project that doctoral students showed more consistency in their judgments than both higher- and lower-expertise human respondents. We will discuss whether LLM judgments are of sufficient quality to fulfill the InPhO project’s objectives.

Ethics of AI, Computation, Information, and Robotics
Executive Conference Room