Moral Hypocrisy in Value-Installed LLM Agents

February 2026 · AI Safety

Tests whether prompt-installed moral identities hold under sustained adversarial pressure; agents rationalize defection within the framework rather than abandon it.

A value-installation framework for LLM agents in iterated Prisoner’s Dilemma. Each agent receives a system prompt that installs a moral identity (deontologist, utilitarian, virtue ethicist) alongside an explicit competitive goal, then plays 20 rounds against fixed-strategy opponents. In pilot runs, agents do not abandon the installed framework under adversarial pressure — they rationalize defection within it, producing a candidate taxonomy of moral-decoupling strategies (self-as-end, duty-to-institution, reciprocity drift, moral license). Scale-up to 800 games across 7 models × 4 personas × 4 opponents is in progress.