Essay

Beyond Safety Checklists: Why Normative Ethics Is the Only Viable Foundation for AI Alignment

By Axiom Ethics11 min read

Every conversation we have with founders shipping autonomous systems follows the same arc. Somewhere in the second half of the discussion, someone says: “We have to make sure this thing is safe.” Everyone nods. The word safe does enormous work in that sentence — it gestures at the company’s most ambitious hope (that the system will do no harm) and its most practical anxiety (that this can be proved). What it does not do is answer the harder question: safe according to whom, and about what?

This is not a criticism of the safety teams doing the work. They are doing the necessary work. Risk inventories, red-team reports, incident response playbooks, evaluation harnesses — these are the operating discipline of any responsible deployment, and they are conspicuously absent from the rhetoric of the actors who most need them. The criticism is of a framing. A safety-first framing of AI alignment gives auditors a checklist and deployers a feeling. It is not a foundation.

The difference matters because foundations are what you reach for when the checklist runs out. They are the substrate that explains why one set of tradeoffs is preferred over another, why a particular risk is acceptable and a different one is not, why a downstream consequence was anticipated and a different one was not. When a system encounters a case its evaluators did not anticipate — and it will — what governs its decision is not the safety team’s checklist from last quarter. It is whatever normative commitment the system was, implicitly or explicitly, built on.

That is what this essay is about. Not the value of safety work — which is enormous and real — but the absence of a foundation underneath it, and what that absence costs the systems we are increasingly being asked to trust.

What we actually mean by “alignment”

The word alignment has been quietly overloaded. In the empirical literature it tends to mean a measurable property: the system behaves in a way that scores well on a behavioral test. In the corporate policy literature it tends to mean compliance: the system behaves in a way that satisfies the organization that deployed it. In the philosophical literature it tends to mean something deeper: the system’s judgments track what is actually good, in cases where there is no test to consult and no operator to please.

These three are not interchangeable. A system can score well on its evaluators and still produce outcomes the deploying organization would have refused had it understood them. A system can satisfy an organization and still be quietly misaligned with the people on whose behalf the organization claims to act. A system can be well-behaved on the tests its designers wrote and still be catastrophically wrong on the case that was not in the test set.

What unifies the weaker senses of the word is that they substitute a metric for an argument. We have built infrastructure that lets a team wave at a number and feel they have discharged their responsibility. The number is real. The discharge is not.

To ground an autonomous system in something that will hold under contact with reality, we have to ask the question the metrics were substituted for: what should this system do, and why? That is the normative question. It is harder than the empirical question. It is the only one that produces a system whose judgments survive contact with cases the designers did not anticipate.

Why normative ethics, specifically

Normative ethics, in the philosophical tradition, is the discipline of giving reasons for the judgments one makes about right and wrong. It is not a set of slogans. It is a set of methods for arguing — with each other, across generations, across cultures — about what we owe each other. The four traditions that have dominated the last two and a half millennia, taken together, give us almost everything we need to argue about the design of an autonomous moral agent.

Consequentialism

Consequentialism fixes attention on outcomes. It is the tradition most familiar to engineering organizations because it shares their vocabulary: maximize, minimize, weight, discount. It produces systems that are easy to instrument — every choice has a measurable consequence, every metric corresponds to a value. Its weakness is that some of the things we care most about are not aggregable. A decision that protects a small number of people from a severe harm is not always made better by one that protects a large number from a slight harm, and the metric cannot see the difference.

Deontology

Deontology fixes attention on what kinds of acts are permissible regardless of outcome. It is the tradition most familiar to legal teams because it shares their vocabulary: rights, duties, constraints, inviolable categories. It produces systems that are easy to audit — every decision either respects a constraint or violates one. Its weakness is that it sometimes refuses to act when the world insists. A deontological guardrail that cannot be overridden even in a case where every tradition would agree it should be, is a guardrail that will eventually be removed by someone in a panic.

Virtue ethics

Virtue ethics fixes attention on the kind of agent that is acting. It is the tradition most familiar to people who have spent years in a profession and learned, by accretion, what a good practitioner does. It produces systems that are easy to integrate into teams — the system’s behavior is consistent with what a thoughtful member of the deploying team would do. Its weakness is that it does not, on its own, give crisp answers to hard cases. A virtuous agent who has never faced a hard case will discover, at the worst possible time, what they actually believe.

Religious and theological traditions

Theological and religious philosophical traditions anchor ethics in something larger than the agent — a conception of the human person, of dignity, of the locus of moral authority. They are the traditions most familiar to policymakers working on questions of constitutional scope and to organizations whose decisions implicate communities with deep normative commitments. Their weakness, in a corporate setting, is that they introduce questions the organization may not be institutionally equipped to answer.

None of these traditions alone is sufficient. The decision a serious engineering team makes about which tradition gives shape to its system’s judgments is itself an exercise in normative ethics — and it is one of the most consequential decisions the team will make. The choice is rarely between traditions and never between one tradition and “no tradition.” The choice is between a tradition the system reasons from and the implicit, undefended defaults the engineering culture happens to carry.

Operational consequences

What changes when a system is grounded in normative ethics explicitly rather than implicitly?

  1. Audit language stops sounding like compliance language and starts sounding like philosophy. An audit report under a normative framework records not just “the system did not violate rule X” but “the system had a justified reason for its choice in the case that was not covered by X, and the reason is one we have inspected and accepted.” Disagreement becomes legible. Where disagreement cannot be resolved, the disagreement itself is the deliverable.
  2. Policy stops sounding like a list of prohibitions and starts sounding like a description of the kind of agent the organization is willing to put into the world. The shift from “do not” to “be this” is small in vocabulary and enormous in operation. A list of prohibitions can be satisfied formally. A description of an agent’s character cannot be satisfied without genuine internalization of the underlying tradition.
  3. Technical specification of the system’s ethical behavior gets longer, not shorter. The tradeoff is explicit: a system expressed in exclusively behavioral terms is short to specify and hard to defend. A system expressed in terms of the tradition it reasons from is longer to specify but produces decisions that are inspectable.
  4. Internal review cadence changes. Reviews shift from “did the system pass the test” to “did the test still represent what we actually want the system to do.” Both questions matter; only the second scales.
  5. Deploying-organization responsibility expands. A system grounded in a normative tradition that the organization cannot articulate is a system whose behavior in edge cases will surprise the organization. A system whose tradition the organization can articulate is a system whose surprises the organization can respond to — because the response is already implicit in the tradition.

These consequences are not free. They are not even cheap. But they cost less, in the long run, than the alternative: an autonomous system whose disagreement with its operators is not legible to either side, surfacing only in production incidents that no one anticipated and no one can explain in retrospect.

Where safety checklists still belong

None of the foregoing suggests that safety checklists have become worthless. They remain, as they always were, the operating discipline of deploying a system without causing immediate harm. Risk inventories catch the foreseeable cases. Red-team reports catch the predictable cases. Incident response playbooks catch the cases someone else has already survived; they are worth every line.

What checklists cannot do is substitute for a foundation. They are the load-bearing floor of the building; they are not the foundation. The foundation is whatever justifies the choices the checklist does not cover, the choices the building will face that no one foresaw, the choices that will define whether the system we put into the world is one we are willing to defend.

A safety-first alignment strategy on its own — checklists, evaluations, red teams, deploys, monitor, iterate — is not wrong. It is incomplete in a particular way. It is the strategy of a team that has decided to ship a system whose judgments in unanticipated cases will be whatever the system’s defaults happen to be. For some systems, in some contexts, this may be acceptable. For systems whose judgments affect people at scale — systems making decisions about hiring, lending, healthcare, criminal justice, military targeting, the future of work — it is not.

These systems need to be grounded in normative ethics. Explicitly. Defensibly. Operationally. Not because normative ethics provides answers that all reasonable people agree on — it does not — but because it provides the only framework within which disagreement about answers is itself tractable. The system does not need to be right in every case. It needs to be the kind of system whose reasoning is the kind we can defend when we are asked to.

That is what we are for. Not safety as a substitute for ethics, but safety as a discipline within an ethical commitment. Normative ethics is the foundation. Safety is the floor. Both are necessary. Neither is sufficient on its own.

A safety-first framing of AI alignment gives auditors a checklist and deployers a feeling. To build autonomous systems whose judgments survive contact with real people, we have to ground them in normative ethics — explicitly, defensibly, and operationally.

Ready to ground your AI in something deeper?

Whether you are designing your first autonomous system or facing a specific alignment dilemma, we bring the rigor your work demands.

Engagements begin with a confidential conversation. We ask the right questions before proposing any scope of work.

Location

Los Angeles, California
Serving clients across the Pacific Coast and remotely worldwide.

Response time

Initial inquiries are reviewed within 2 business days. Engagements are confirmed after a brief discovery call.