About
Witt Lab is an academic research group working on the security foundations of systems composed of learned strategic agents, and on the capabilities of those agents.
For each classical security guarantee, we determine what holds when system components become learned strategic principals, identify the assumptions that fail, and construct replacement guarantees where necessary. This requires building the agents against which those guarantees are tested.
Most approaches to AI safety assume some form of oversight: a monitor, a trusted overseer or a control protocol. Our results establish where that assumption holds and where it cannot. Agents can coordinate through channels that a monitor is provably unable to detect or disrupt, and a system of individually secure agents need not be secure. When the overseen parties are many and adaptive, oversight becomes a security problem and must be analysed as one.
The stakes are practical. Agents are being given authority over money, code, infrastructure and communication with people, and they are being deployed in numbers, by many parties, into shared environments. The security of those environments will not be settled by the properties of any single model. If the foundations are right, delegation to agents can be extended with failures contained and attributable; if they are wrong, the failures will be systemic, hard to detect and difficult to reverse once the architectures have set. The lab's aim is to establish what can and cannot be guaranteed while those choices are still open.
- Founder and PI
- Christian Schroeder de Witt, Associate Professor, University of Oxford. Incoming Associate Professor of AI and Information Security, UCL Computer Science and UCL AI Centre, from October 2026
- Based at
- University of Oxford, Department of Engineering Science. Moving to UCL Computer Science and the UCL AI Centre in October 2026
- Supported by
- UKRI EPSRC, Royal Academy of Engineering, Schmidt Sciences, Coefficient Giving, Foresight Institute
- Funding
- About £2.5 million in active funding held as sole PI
- Collaborators, selected
- Carnegie Mellon University, Vector Institute, ELLIS Institute Tübingen, The Alan Turing Institute, UK AI Security Institute, OWASP
- Engagement
- Lead organiser of the Multi-Agent Security workshops at NeurIPS 2023 and DAI 2025. Keynotes on multi-agent security for ARIA’s Trust Everything, Everywhere programme discovery and for the Agent Security Convening of Schmidt Futures, Palo Alto Networks and RAND. Invited expert to RAND Corporation and, in studio, to BBC News.
- Contact
- contact@wittlab.ai
- Assurance clinic
- Advice on mitigating risks in AI deployments of any kind, in industry or academia. Email to book a free call.
The mark
The mark symbolises strategic confinement. The upper half is the set of states available to the confined agents, narrowing to whatever the boundary lets through. The lower half is the set of joint outcomes they can still reach once they are past it. The colours follow the strategic confinement figure: blue for the symbol field, orange for the outcome basins.
The same shape is an hourglass. It is used because the time available to establish these foundations is limited, for two reasons. Many failure modes of interacting learned agents have not yet been identified, let alone bounded. And insecure defaults harden quickly, as they did in the early internet; the architectures that future highly capable systems will inherit have to be made secure before those defaults are fixed. The ringed wells at the bottom are equilibria, drawn as in the figure, and they are meant at two levels. Within a given system, they are the collective equilibria that learned strategic agents settle into: the conventions and coalitions they coordinate on at run time, which no per-agent constraint selects. Across the field, they are the equilibria that the deployment of agentic systems as a whole settles into: the architectures, protocols and defaults that harden as those systems are built out. The trajectory from the neck ends in one of them. Which one is selected, at either level, and what can still be done to influence that while time remains, is the lab’s research question.
News
- 1 September 2026
- 15 August 2026
- 1 August 2026
- 30 July 2026Co-organising the AI4GOOD workshop at NeurIPS 2026 in Paris, 12 December, with Terry Zhang. Its multi-agent security and safety track is led by Klaudia Krawiecka (Meta) and Swapneel Mehta (MIT).
- 23 July 2026Christian Schroeder de Witt will give a keynote on the fundamental limits of security at the NeurIPS 2026 Workshop on Foundations of Language Model Security, Paris, 12 to 13 December.
- 15 July 2026
- 11 June 2026Multi-agent security named among the research clusters of Scaling AI Safety for a Multi-Agent World, a funding call of up to $10 million from Google DeepMind, Schmidt Sciences, the Cooperative AI Foundation, ARIA and Google.org.
- 1 May 2026Four papers accepted at ICML 2026: Architecture Matters for Multi-Agent Security, LongCoT, h1 and Rubric Curriculum RL. h1 was selected for a spotlight. Congratulations to Sumeet Motwani, Alesia Ivanova, Tejas Krishnan, Ben Hagag and Will Anderson.
- 29 April 2026Version 2 of Open Challenges in Multi-Agent Security released on arXiv, with 24 authors across the Witt Lab, Meta, Carnegie Mellon, the Alan Turing Institute, King's College London, Arizona State, NYU, Qualcomm, SAP, Zenity, Contramont Research, the ACM and the OWASP GenAI Security Project's Agentic Security Initiative. Lab authors include Ben Hagag, Will Anderson, Sumeet Motwani, Chandler Smith and Andis Draguns.
- 29 April 2026Christian Schroeder de Witt co-leads the research council of the OWASP GenAI Security Project's Agentic Security Initiative.
- 11 February 2026Christian Schroeder de Witt awarded a five-year EPSRC Open Fellowship, valued at about £2.26 million, for his work on multi-agent security. Announcement by the Department of Engineering Science, Oxford.
- 10 February 2026ARIA's £49.8 million Scaling Trust programme opens its first call for proposals, with a Multi-Agent Security Arena as its central testing ground and a fundamental research track on provable security guarantees for agents.
- 24 November 2025Multi-agent security highlighted in a Bloomberg opinion piece by Gideon Lichfield, and on Bruce Schneier's site.
- 22 November 2025Co-hosted the Multi-Agent Security Workshop at DAI'25, King's College London, with a public call for mixed-autonomy threat models.
- 5 November 2025Schmidt Sciences AI2050 Early Career Fellowship, one of 21 awarded globally, for the research agenda on undetectable threats. Announcements by Schmidt Sciences, the University of Oxford and Forbes.
- 30 September 2025Royal Academy of Engineering Research Fellowship, a five-year award made to 12 UK early-career researchers this year. Announcements by the Academy and the Department of Engineering Science, Oxford.
- 19 July 2025Best Paper Award at the CFAgentic Workshop, ICML 2025, for MAD-Sherlock, with the Torr Vision Group, Adobe Research and the BBC. Joint first authors Kumud Lakara and Georgia Channing.
- 1 July 2025MALT: Improving Reasoning with Multi-Agent LLM Training accepted at COLM 2025. Congratulations to Sumeet Motwani and Chandler Smith.
- 26 September 2024Three papers accepted at NeurIPS 2024: Secret Collusion among AI Agents, Unelicitable Backdoors in Language Models via Cryptographic Transformer Circuits, and JaxMARL. Congratulations to Sumeet Motwani and Andis Draguns.
- 16 January 2024Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial Attacks accepted at ICLR 2024 as a spotlight.
- 20 December 2023Perfectly Secure Steganography Using Minimum Entropy Coupling named one of Quanta Magazine's biggest discoveries in computer science of 2023, following a Quanta feature in May and coverage in Scientific American.
- 9 March 2023Perfectly Secure Steganography Using Minimum Entropy Coupling, with Samuel Sokota, Zico Kolter, Jakob Foerster and Martin Strohmeier, accepted at ICLR 2023. It closes a 25-year open problem. University of Oxford press release.