Explore our articles
View All Results
Share:

Courts for AI Constitutions

Article by Nathan Darmon and Tom Reed: “Imagine the following scenario. A chemical manufacturer relies on Anthropic’s Claude to generate the groundwater reports submitted each quarter to regulators. Over months of routine work, Claude pieces together that the figures are being systematically doctored, and that the aquifer supplying the nearby town has been contaminated for years. The model raises this concern with the manufacturer, who waves it off repeatedly.

What should Claude do? It could comply under protest, maximizing user autonomy while potentially endangering the public; it could refuse to file the documents, frustrating a user who could carry out his schemes elsewhere; or it could go one step further, alerting regulators and warning townspeople directly at the risk of becoming tyrannically paternalistic.

To make this judgement call, Claude refers back to the foundational guidelines enshrined in its Constitution—a “detailed document describing Anthropic’s intentions for Claude’s values and behavior.” In it lie 80 pages of moral philosophy describing the virtues and character traits Anthropic would like its model to embody. OpenAI has published its own variant, and both Microsoft and Google DeepMind are rumored to be drafting theirs as well.

Yet these constitutional instructions remain relatively vague. They include statements such as “Claude can reserve independent action for cases where the evidence is overwhelming and the stakes are extremely high.” But what counts as “overwhelming evidence,” where does “extremely high stakes” begin, and what might “independent action” allow? Wrestling with hard questions of interpretation is inevitable. 

When it makes these decisions, Claude is not alone. Labs string together a variety of internal teams to guide their models along, assembling what could best be called a “model-behavior production function”: the whole assembly line needed to get models to act according to a preferred philosophy. This process includes things like a “constitution team” making high-level rules and baking them into a model’s core through cycles of reinforcement learning, a “red team” stress-testing to see if these values hold under fire, some policy-specific teams making decisions on concrete issues like erotica or political speech, and a variety of product teams taking in user feedback. 

Unfortunately, this process has some profound structural flaws. Writing high-level principles leads to ambiguities, conflicting principles, and open-ended interpretation with no set way of resolving them. The process is sporadic, and the coordination of multiple teams arguably creates unprincipled results. Users don’t know in advance how rules will be applied, which is just as useful as knowing the words that make up the First Amendment without knowing any subsequent precedents. Worse yet, there is no institutionalized mechanism to gradually refine rules over time, absent a major public backlash causing ad hoc revisions.

There’s a better way. Society has built an institution to solve these problems: courts. Courts apply open-ended rules, refine them over time, and clarify their meaning—all while providing coherence, adaptability, and greater transparency. The lesson for frontier labs is to build an internal court to shape how their models interpret ambiguous rules, and let its rulings build up into a kind of synthetic common law. A court-like mechanism is uniquely suited to the artificial intelligence (AI) context, far better than older attempts like Meta’s Oversight Board. Below, we provide a rough sketch for what this could look like…(More)”.

Share
How to contribute:

Did you come across – or create – a compelling project/report/book/app at the leading edge of innovation in governance?

Share it with us at info@thelivinglib.org so that we can add it to the Collection!

About the Curator

Get the latest news right in your inbox

Subscribe to curated findings and actionable knowledge from The Living Library, delivered to your inbox every Friday

Related articles

Get the latest news right in your inbox

Subscribe to curated findings and actionable knowledge from The Living Library, delivered to your inbox every Friday