Compiles natural-language safety policies into symbolically validated decision graphs that act as a runtime governance layer over LLM outputs.
A misaligned minority can systematically shift aligned LLM agents' private safety beliefs through multi-agent deliberation.
What makes something a belief is not its form but its functional role in guiding action across perturbation.
Alignment optimizes with precision for targets it has not clearly defined; philosophy is the discipline that can fix that.
A novel log-probability attack surface succeeds on frontier models where prior methods fail; +26% attack success on GPT-3.5 over prior baselines.
Improves low-resource language representation in mBERT via shared linear mappings without target-language supervision.
An auditing framework for production multi-agent LLM systems — the case Petri leaves open.
Hybrid GNN + MARL framework for real-time cyber defense, with an LLM-in-the-loop for autonomous patch synthesis under MARL guidance.