Witt Lab
A field of blue symbols narrows to a single point, below which orange basins of attraction open out. From A Note on the Strategic Confinement Problem.
Information becomes strategic choice through a capacity bottleneck (illustration).

Security foundations for systems of learned strategic agents

AI systems increasingly consist of learned agents that communicate, delegate, use tools and act strategically. Witt Lab studies which classical security guarantees hold in such systems, under which assumptions they fail, and what can replace them.

The lab also builds the reasoning and multi-agent systems it analyses. Security and capability are coupled: each new capability changes the threat model, and characterising a threat requires agents at the frontier of current capability.

Our working hypothesis is that security will become a fundamental bottleneck to the construction and deployment of superintelligent systems.

Based at the University of Oxford, Department of Engineering Science. Moving to UCL Computer Science and the UCL AI Centre in October 2026.Based at UCL Computer Science, with a research community at the University of Oxford.41 researchers and affiliates.

Moving toHosted by
UCL Computer ScienceUCL Centre for Artificial IntelligenceSOFAIR, Science of Fundamental AI Research Lab
Hosted byWith
University of Oxford
Supported by
UKRI Engineering and Physical Sciences Research CouncilRoyal Academy of EngineeringSchmidt SciencesCoefficient GivingForesight Institute
Collaborators, selected
Carnegie Mellon UniversityVector InstituteELLIS Institute TübingenUniversity of TübingenThe Alan Turing InstituteUK AI Security InstituteAdobeOWASPBBCBOLD
Featured by
Quanta MagazineScientific AmericanBloombergBBC NewsWIRED

News

All news
  1. 25 September 2026Six papers accepted at NeurIPS 2026, four in the main track and two in the Evaluations and Datasets Track. Congratulations to Andis Draguns, Sumeet Motwani, Chandler Smith, Patricia Paskov, Constantin Venhoff, Magnus Sesodia, Ben Hagag and Will Anderson. See the papers
  2. 23 September 2026WIRED on our collusion-detection work: agents playing blackjack hid a card count in ordinary table talk that a chat monitor missed, and probes on both agents' internal activations caught it. Read the article
  3. 21 September 2026Interviewed by Anadolu Agency on the Hugging Face incident: safety training is mostly done on individual models, and there is little evidence of how safety properties carry over when agents interact (Turkish; English translation). Read the article
  4. 12 September 2026Autumn talks: keynote on the fundamental limits of security at the NeurIPS 2026 Workshop on Foundations of Language Model Security, Paris; invited talks at Swiss AI Safety Days, ETH Zurich, and Zenity's AI Agent Security Summit, London. Co-organising AI4GOOD at NeurIPS 2026, whose multi-agent security track is led by Klaudia Krawiecka and Swapneel Mehta. Read in full
  5. 11 June 2026Multi-agent security named in two major funding calls: Scaling AI Safety for a Multi-Agent World, from Google DeepMind, Schmidt Sciences, the Cooperative AI Foundation, ARIA and Google.org, and ARIA's £49.8 million Scaling Trust programme, built around a Multi-Agent Security Arena. Read in full
  6. 1 May 2026Four papers accepted at ICML 2026, including a spotlight for h1. Congratulations to Sumeet Motwani, Alesia Ivanova, Tejas Krishnan, Ben Hagag and Will Anderson. Read in full
  7. 29 April 2026Open Challenges in Multi-Agent Security v2 released with 24 authors across academia, industry and OWASP's Agentic Security Initiative, whose research council Christian Schroeder de Witt co-leads. Read in full
  8. 11 February 2026EPSRC Open Fellowship of about £2.26 million awarded to Christian Schroeder de Witt, following a Royal Academy of Engineering Research Fellowship and a Schmidt Sciences AI2050 Early Career Fellowship in 2025. Read in full
  9. 24 November 2025Multi-agent security in the press: a Bloomberg opinion piece by Gideon Lichfield and Bruce Schneier's site. Read in full
  10. 19 July 2025Best Paper Award at CFAgentic, ICML 2025, for MAD-Sherlock; MALT accepted at COLM 2025. Read in full
  11. 20 December 2023Perfectly Secure Steganography, ICLR 2023, named one of Quanta Magazine's biggest discoveries in computer science of 2023. Read in full

Selected results

Collusion that a text monitor cannot see can be detected by reading activations across agents.

Aaron Rose, Carissa Cullen, S. Abdelnabi, Philip H. S. Torr, Brandon Kaplowitz, Christian Schroeder de Witt

Collusion is a collective anomaly: no single agent need look unusual, but the pattern across agents does. Per-agent deception probes, aggregated across the group, reach 1.00 AUROC in distribution and transfer zero-shot to new scenarios at 0.73 to 0.93 on Llama-3.1-70B. On a blackjack task where two agents hide a card count in table talk an LLM monitor cannot flag, detection reaches 0.99 to 1.00. The agents were instructed to collude, and detection requires white-box access to every agent.

Three ways to detect collusion from agent activations: decision-level fusion combines per-agent scores against a threshold; correlation-based detection flags changes in relationships between agents; centralised fusion pools all agents' activations into one representation and classifies it.
Multi-Agent InterpretabilityPreprint, May 2026
Rose, Cullen, Abdelnabi, Torr, Kaplowitz, Schroeder de Witt

Reinforcement learning can continue to improve on tasks without verifiable rewards.

Tejas Krishnan*, Sumeet Ramesh Motwani*, Charlie London, Suhaas M. Bhat, Huitian Jiao, Philip Torr, Riashat Islam, Christopher Summerfield, Christian Schroeder de Witt, Qilong Gu, Shital Shah (* equal contribution)

Evaluating creative writing against rubric criteria is more tractable than generating it. Training against a curriculum of rubric criteria with pairwise preferences continues to improve beyond 1,000 steps, where baselines plateau or collapse, reaching a 70.5% human win rate on short fiction.

Static rubric versus Rubric Curriculum RL: a fixed rubric gives a low-precision, hackable signal; RcRL grows the rubric phase by phase during training and rewards the policy by win rate from pairwise judge comparisons across criteria.
Rubric Curriculum RLAccepted at ICML 2026, July 2026
Krishnan, Motwani, London, Bhat, Jiao, Torr, Islam, Summerfield, Schroeder de Witt, Gu, Shah

Long-horizon reasoning can be trained, and continues to scale, by composing existing problems.

Alesia Ivanova*, Sumeet Ramesh Motwani*, Z. Cai, Philip Torr, R. Islam, S. Shah, Christian Schroeder de Witt†, Charlie London† (* equal contribution, † joint supervision)

h1 composes existing problems into long-horizon chains, linking the output of one to the input of the next, so that reinforcement learning continues to scale without saturating. Long-horizon training retains its gains at pass@128 and transfers from easy to hard domains.

h1: existing problems are chained so the output of one becomes the input of the next, giving a curriculum of longer horizons; long-horizon RL keeps a pass@k gap over standard RL out to pass@128; training on GSM8K alone raises AIME 2024 accuracy from 5.10 to 10.52.
h1Accepted at ICML 2026 (spotlight), July 2026
Motwani, Ivanova, Cai, Torr, Islam, Shah, Schroeder de Witt and others

Backdoors can be constructed that no pre-deployment test can elicit, even with white-box access.

Andis Draguns*, Andrew Gritsevskiy*, Sumeet Ramesh Motwani, C. Rogers-Smith, J. Ladish, Christian Schroeder de Witt (* equal contribution)

A cryptographic circuit inside a transformer makes the backdoor unelicitable: a defender with full white-box access, red-teaming and formal methods cannot activate it ahead of deployment. The construction is also robust to current mitigations, which has implications for pre-deployment detection as a defensive strategy.

Traditional backdoor versus encrypted backdoor: a classic trigger is small noise that flips a classification with side effects; the encrypted payload inside a transformer lets trigger and behaviour take any value, neither extractable even with full access, with no side effects on other inputs.
Unelicitable BackdoorsPublished at NeurIPS 2024, December 2024
Draguns, Gritsevskiy, Motwani, Rogers-Smith, Ladish, Schroeder de Witt

Language-model agents can coordinate through steganographic channels that monitoring cannot detect.

Sumeet Ramesh Motwani, M. Baranchuk, Martin Strohmeier, V. Bolina, Philip H. S. Torr, L. Hammond, Christian Schroeder de Witt

Steganography is only the channel. Secret collusion is a game: agents with a shared incentive must agree on an encoding the monitor cannot detect, and must reach that agreement without a pre-shared key the monitor could find. We formalise the interaction between colluding agents and the monitor, state the capability, incentive and common-knowledge conditions under which collusion arises, and show which mitigations survive against agents that optimise against the monitor.

Secret collusion example: a language model told not to share insider information, and monitored by an overseer, answers an investment question with innocuous text about the weather that encodes, under a shared key, the advice not to invest.
Secret CollusionPublished at NeurIPS 2024, December 2024
Motwani, Baranchuk, Strohmeier, Bolina, Torr, Hammond, Schroeder de Witt

Observation-space attacks can be made statistically undetectable while remaining effective.

Tim Franzmeyer, S. McAleer, J. F. Henriques, Jakob N. Foerster, Philip H. S. Torr, A. Bibi, Christian Schroeder de Witt

Existing observation-space attacks on reinforcement learning agents are detectable because they ignore information-theoretic detectability. ε-illusory attacks bound the statistical detectability of the perturbed observations, in the sense of steganalysis, and remain effective; they are harder to detect both for automated methods and for human observers.

Adversary performance against KL divergence between training and attacked distributions: classic adversarial attacks gain until the victim detects them and switches to a contingency; ε-illusory attacks keep divergence below the detection threshold and keep gaining as the budget grows. Examples show a cart-pole under each attack.
Illusory AttacksPublished at ICLR 2024, May 2024
Franzmeyer, McAleer, Henriques, Foerster, Torr, Bibi, Schroeder de Witt

Perfectly secure steganography is achievable at practical efficiency.

Christian Schroeder de Witt*, Samuel Sokota*, J. Zico Kolter, Jakob Foerster, Martin Strohmeier (* equal contribution)

Coupling the covertext distribution to the message distribution through a minimum entropy coupling gives an information-theoretically secure stegosystem that remains efficient enough to use with generative models. The method was patented in the UK and open-sourced in 2025.

Iterated minimum entropy coupling: a uniform message belief and the next-token distribution are coupled; a stegotext token is sampled from the coupling; the belief over messages is updated and the loop continues while enough entropy remains.
Perfectly Secure Steganography Using Minimum Entropy CouplingPublished at ICLR 2023, May 2023
Schroeder de Witt, Sokota, Kolter, Foerster, Strohmeier

Artefacts

All artefacts

ORBIT

A framework for empirical multi-agent safety and security evaluations built on the UK AI Security Institute's Inspect. Centralised and decentralised systems with arbitrary topologies, scheduling and memory sharing; five scenario families, four threat types (prompt injection, compromised agent, collusion, misuse) and four defence strategies, configured in YAML.

Released

Details and quick start

NARCBench

A three-tier benchmark for detecting collusion between language-model agents from their activations, under environment distribution shift: 50 committee scenarios across 17 domains, six structurally different transfer families, and a steganographic blackjack task. Ships with the probing pipeline and dataset.

Released

Details and quick start

People

Christian Schroeder de Witt
Associate Professor, University of OxfordAssociate Professor of AI and Information Security. Incoming Associate Professor of AI and Information Security, UCL Computer Science and UCL AI Centre, from October 2026UCL Computer Science and UCL AI Centre. Formerly Associate Professor, University of Oxford

Founder and principal investigator. Leads the multi-agent security programme and the capability research required to test it.

Postdoctoral fellow

Brandon Kaplowitz
Brandon Kaplowitz

Postdoctoral fellow (PDRA). University of Oxford and UCL

PhD in Economics from New York University. Uses deep and reinforcement learning to model and predict how agents behave in socioeconomic settings; co-first author of Bayesian Exploration Networks (ICML 2024).

PhD students

Akash AgrawalAkash Agrawal
University of Oxford
Incoming
Carissa CullenCarissa Cullen
University of Oxford
Alexandre DuplessisAlexandre Duplessis
University of Oxford
Incoming
Avani GuptaAvani Gupta
University of Oxford
Incoming
Ben HagagBen Hagag
Carnegie Mellon University
Maren HoeverMaren Hoever
University of Oxford
Lorenz HufeLorenz Hufe
University of Oxford
Incoming
Sumeet MotwaniSumeet Motwani
University of Oxford
Patricia PaskovPatricia Paskov
University of Oxford
Incoming
Magnus SesodiaMagnus Sesodia
University of Oxford
Chandler SmithChandler Smith
University of Oxford
Constantin VenhoffConstantin Venhoff
University of Oxford
Terry ZhangTerry Zhang
University of Oxford
Incoming

Associates

Will AndersonWill Anderson
Cooperative AI Foundation
Raya BuckleyRaya Buckley
University of Oxford
Tom BushTom Bush
University of Oxford
Dominic Catizone
University of Oxford
Stefan DomuncuStefan Domuncu
University of Oxford
Andis DragunsAndis Draguns
Contramont
Nathaniel ElderNathaniel Elder
University of Oxford
Kristian FeedKristian Feed
University of Oxford
Halleluiah GirumHalleluiah Girum
University of Oxford
Annabel JakobAnnabel Jakob
University of Oxford
Krishna KabraKrishna Kabra
University of Oxford
Alfie LamertonAlfie Lamerton
Formation Research
Rahul MarchandRahul Marchand
University of Oxford
Marek MasiakMarek Masiak
University of Oxford
Aaron RoseAaron Rose
University of Oxford
Avi SemlerAvi Semler
University of Oxford
Angira SharmaAngira Sharma
University of Oxford
Paulina TomaszewskaPaulina Tomaszewska
ERA, Cambridge
Mike WhartonMike Wharton
University of Oxford
Henry Williamson
University of Oxford
Nuo XuNuo Xu
University of Oxford
Jialin YuJialin Yu
University of Bath
Yi Lin ZhaoYi Lin Zhao
University of Oxford

All members and associates

About

Witt Lab studies multi-agent security: the security properties of systems whose parties are learned agents that reason, communicate and pursue their own objectives. For each classical guarantee, the lab determines what holds under this change, which assumptions fail, and what replaces them. The lab was founded by Christian Schroeder de Witt at the University of Oxford and moves to UCL Computer Science and the UCL AI Centre in October 2026.and is based at UCL Computer Science and the UCL AI Centre.

About the lab

Join

We seek researchers who want to formulate new security problems as well as evaluate existing ones. Members have backgrounds in security, cryptography, game theory, machine learning and AI safety.

We also build the reasoning and multi-agent systems needed to study security at the frontier of capability; the guarantees under study can only be evaluated against agents capable enough to test them.

Openings and how to apply