Tests whether prompt-installed moral identities hold under sustained adversarial pressure; agents rationalize defection within the framework rather than abandon it.
Builders' shelf: tools, prototypes, and lots of negative results
AI Safety
Tests whether visible chain-of-thought is asymmetrically helpful across incentive structures: cooperative games vs. misaligned signaling.
Null result: alignment tier has no detectable effect on deceptive communication in multi-agent PD; prosocial framing dominates regardless.
Scale-free communication topologies fail to self-correct after misinformation injection; hubs anchor incorrect beliefs and bottleneck downstream correction.
Permanent regimes ratchet into rule-accumulation lock-in; sunset clauses break it. The failure mode is rule-making, not compliance.
Human-AI Collaboration
Chrome extension that adaptively routes queries between traditional and LLM-powered search. 39% reduction in user effort, 90% routing-decision acceptance.
Context-aware prompt engineering with secure code generation for data visualization. Code success: 88% zero-shot → 100% with schema-guided few-shot.
Multimodal RAG system that makes STEM video content accessible to blind and low-vision learners. 100% retrieval accuracy, 90% answer faithfulness on physics content.