International Association for Computing and Philosophy - Annual Conference 2026

Trustworthy AI and the King Midas Problem
2026-07-15 , Executive Conference Room

The King Midas Problem concerns the possibility that we may give a highly advanced AI instructions intended for our benefit, but, like King Midas wishing for all he touched to turn to gold, realise too late that there are severely harmful consequences to our wishes being fulfilled. A classic example is Nick Bostrom’s paperclip maximiser, in which an AI, having been instructed to maximise the manufacture of paperclips, turns everything it can into paperclips, effectively destroying the planet in the process.

I propose a solution based on the observation that the King Midas Problem is an instance of the Principal-Agent Problem. The solution crucially involves the concept of trustworthy AI. What is more, I argue that the problem cannot be solved without making AI trustworthy. This means that, contrary to the opinion of many writers on the topic, trustworthiness must be considered alongside other important factors, like safety and reliability, in the development of future highly advanced AIs.