Log-Probability Guided Adversarial Attacks on Multi-Agent LLM Debate
A novel log-probability attack surface succeeds on frontier models where prior methods fail; +26% attack success on GPT-3.5 over prior baselines.
A novel log-probability attack surface succeeds on frontier models where prior methods fail; +26% attack success on GPT-3.5 over prior baselines.