Christian Schroeder de WittClassical confinement bounds how much information can leak through a channel. When the confined parties are learned strategic agents, a bound on leakage need not bound what they can jointly achieve: past a threshold in permitted communication, the set of reachable joint outcomes changes qualitatively rather than by degree.
Strategic ConfinementPreprint, September 2026
Schroeder de Witt Christian Schroeder de Witt, Klaudia Krawiecka, Igor Krawczuk, Ben Hagag, and othersA taxonomy of the security problems that arise only once agents interact, an assessment of which classical tools apply and which fail, and a statement of the open questions.
Open Challenges in Multi-Agent SecurityPreprint, August 2026
Schroeder de Witt, Krawiecka, Krawczuk, Hagag, Anderson, Belcak and others Aaron Rose, Carissa Cullen, S. Abdelnabi, Philip H. S. Torr, Brandon Kaplowitz, Christian Schroeder de WittCollusion is a collective anomaly: no single agent need look unusual, but the pattern across agents does. Per-agent deception probes, aggregated across the group, reach 1.00 AUROC in distribution and transfer zero-shot to new scenarios at 0.73 to 0.93 on Llama-3.1-70B. On a blackjack task where two agents hide a card count in table talk an LLM monitor cannot flag, detection reaches 0.99 to 1.00. The agents were instructed to collude, and detection requires white-box access to every agent.
Multi-Agent InterpretabilityPreprint, May 2026
Rose, Cullen, Abdelnabi, Torr, Kaplowitz, Schroeder de Witt Tejas Krishnan*, Sumeet Ramesh Motwani*, Charlie London, Suhaas M. Bhat, Huitian Jiao, Philip Torr, Riashat Islam, Christopher Summerfield, Christian Schroeder de Witt, Qilong Gu, Shital Shah (* equal contribution)Evaluating creative writing against rubric criteria is more tractable than generating it. Training against a curriculum of rubric criteria with pairwise preferences continues to improve beyond 1,000 steps, where baselines plateau or collapse, reaching a 70.5% human win rate on short fiction.
Rubric Curriculum RLAccepted at ICML 2026, July 2026
Krishnan, Motwani, London, Bhat, Jiao, Torr, Islam, Summerfield, Schroeder de Witt, Gu, Shah Alesia Ivanova*, Sumeet Ramesh Motwani*, Z. Cai, Philip Torr, R. Islam, S. Shah, Christian Schroeder de Witt†, Charlie London† (* equal contribution, † joint supervision)h1 composes existing problems into long-horizon chains, linking the output of one to the input of the next, so that reinforcement learning continues to scale without saturating. Long-horizon training retains its gains at pass@128 and transfers from easy to hard domains.
h1Accepted at ICML 2026 (spotlight), July 2026
Motwani, Ivanova, Cai, Torr, Islam, Shah, Schroeder de Witt and others Andis Draguns*, Andrew Gritsevskiy*, Sumeet Ramesh Motwani, C. Rogers-Smith, J. Ladish, Christian Schroeder de Witt (* equal contribution)A cryptographic circuit inside a transformer makes the backdoor unelicitable: a defender with full white-box access, red-teaming and formal methods cannot activate it ahead of deployment. The construction is also robust to current mitigations, which has implications for pre-deployment detection as a defensive strategy.
Unelicitable BackdoorsPublished at NeurIPS 2024, December 2024
Draguns, Gritsevskiy, Motwani, Rogers-Smith, Ladish, Schroeder de Witt Sumeet Ramesh Motwani, M. Baranchuk, Martin Strohmeier, V. Bolina, Philip H. S. Torr, L. Hammond, Christian Schroeder de WittSteganography is only the channel. Secret collusion is a game: agents with a shared incentive must agree on an encoding the monitor cannot detect, and must reach that agreement without a pre-shared key the monitor could find. We formalise the interaction between colluding agents and the monitor, state the capability, incentive and common-knowledge conditions under which collusion arises, and show which mitigations survive against agents that optimise against the monitor.
Secret CollusionPublished at NeurIPS 2024, December 2024
Motwani, Baranchuk, Strohmeier, Bolina, Torr, Hammond, Schroeder de Witt Tim Franzmeyer, S. McAleer, J. F. Henriques, Jakob N. Foerster, Philip H. S. Torr, A. Bibi, Christian Schroeder de WittExisting observation-space attacks on reinforcement learning agents are detectable because they ignore information-theoretic detectability. ε-illusory attacks bound the statistical detectability of the perturbed observations, in the sense of steganalysis, and remain effective; they are harder to detect both for automated methods and for human observers.
Illusory AttacksPublished at ICLR 2024, May 2024
Franzmeyer, McAleer, Henriques, Foerster, Torr, Bibi, Schroeder de Witt Christian Schroeder de Witt*, Samuel Sokota*, J. Zico Kolter, Jakob Foerster, Martin Strohmeier (* equal contribution)Coupling the covertext distribution to the message distribution through a minimum entropy coupling gives an information-theoretically secure stegosystem that remains efficient enough to use with generative models. The method was patented in the UK and open-sourced in 2025.
Perfectly Secure Steganography Using Minimum Entropy CouplingPublished at ICLR 2023, May 2023
Schroeder de Witt, Sokota, Kolter, Foerster, Strohmeier