Position: What is a Morally Aligned AI Agent? Philosophy as a Bridge to Operationalization

S. A. Haider
ICML Philosophy meets ML 2026 (under review) · January 2026 · agent-alignment

Alignment optimizes with precision for targets it has not clearly defined; philosophy is the discipline that can fix that.

What this is

A position paper arguing that alignment research has a vocabulary problem: we talk about making models reliable, safe, and trustworthy without asking what those words require of a system philosophically. Building on Singh’s eight-dimension reconstruction of agency — causality, initiative, responsibility, intention, delegation, utility, stability, computation — the paper identifies three dimensions of morality shared across virtue, deontological, and consequentialist traditions: normativity, deliberation, and judgment. Crossing the two yields an 8×3 matrix that serves as a coverage map for the alignment field. The central claim: current techniques are not randomly distributed across this matrix. We are strong on normativity, increasingly competent on judgment, and consistently weak on deliberation.

Why it matters

The matrix is a diagnostic tool. It shows where the field’s investment is heavy and where it is structurally thin — not because of an aesthetic preference for philosophical rigor, but because techniques that succeed on one dimension can fail silently on another, and we lack the vocabulary to notice. Treating philosophy as primary rather than decorative is, I argue, an engineering necessity: you cannot build a morally aligned agent if “morally aligned” remains a word that everyone in the room interprets differently.