Witt Lab

About

Witt Lab is an academic research group working on the security foundations of systems composed of learned strategic agents, and on the capabilities of those agents.

For each classical security guarantee, we determine what holds when system components become learned strategic principals, identify the assumptions that fail, and construct replacement guarantees where necessary. This requires building the agents against which those guarantees are tested.

Most approaches to AI safety assume some form of oversight: a monitor, a trusted overseer or a control protocol. Our results establish where that assumption holds and where it cannot. Agents can coordinate through channels that a monitor is provably unable to detect or disrupt, and a system of individually secure agents need not be secure. When the overseen parties are many and adaptive, oversight becomes a security problem and must be analysed as one.

The stakes are practical. Agents are being given authority over money, code, infrastructure and communication with people, and they are being deployed in numbers, by many parties, into shared environments. The security of those environments will not be settled by the properties of any single model. If the foundations are right, delegation to agents can be extended with failures contained and attributable; if they are wrong, the failures will be systemic, hard to detect and difficult to reverse once the architectures have set. The lab's aim is to establish what can and cannot be guaranteed while those choices are still open.

Founder and PI
Christian Schroeder de Witt, Associate Professor, University of Oxford. Incoming Associate Professor of AI and Information Security, UCL Computer Science and UCL AI Centre, from October 2026Associate Professor of AI and Information Security, UCL Computer Science and UCL AI Centre
Based at
University of Oxford, Department of Engineering Science. Moving to UCL Computer Science and the UCL AI Centre in October 2026UCL Computer Science, with a research community at the University of Oxford
Supported by
UKRI EPSRC, Royal Academy of Engineering, Schmidt Sciences, Coefficient Giving, Foresight Institute
Funding
About £2.5 million in active funding held as sole PIAbout £4.8 million in active funding held as sole PI, including an ARIA award
Collaborators, selected
Carnegie Mellon University, Vector Institute, ELLIS Institute Tübingen, The Alan Turing Institute, UK AI Security Institute, OWASP
Engagement
Lead organiser of the Multi-Agent Security workshops at NeurIPS 2023 and DAI 2025. Keynotes on multi-agent security for ARIA’s Trust Everything, Everywhere programme discovery and for the Agent Security Convening of Schmidt Futures, Palo Alto Networks and RAND. Invited expert to RAND Corporation and, in studio, to BBC News.
Contact
contact@wittlab.ai
Assurance clinic
Advice on mitigating risks in AI deployments of any kind, in industry or academia. Email to book a free call.

The mark

The mark symbolises strategic confinement. The upper half is the set of states available to the confined agents, narrowing to whatever the boundary lets through. The lower half is the set of joint outcomes they can still reach once they are past it. The colours follow the strategic confinement figure: blue for the symbol field, orange for the outcome basins.

The same shape is an hourglass. It is used because the time available to establish these foundations is limited, for two reasons. Many failure modes of interacting learned agents have not yet been identified, let alone bounded. And insecure defaults harden quickly, as they did in the early internet; the architectures that future highly capable systems will inherit have to be made secure before those defaults are fixed. The ringed wells at the bottom are equilibria, drawn as in the figure, and they are meant at two levels. Within a given system, they are the collective equilibria that learned strategic agents settle into: the conventions and coalitions they coordinate on at run time, which no per-agent constraint selects. Across the field, they are the equilibria that the deployment of agentic systems as a whole settles into: the architectures, protocols and defaults that harden as those systems are built out. The trajectory from the neck ends in one of them. Which one is selected, at either level, and what can still be done to influence that while time remains, is the lab’s research question.

News

  1. 1 September 2026
    Invited talk at Zenity's AI Agent Security Summit, London, 8 October.
  2. 15 August 2026
    Attended the Schmidt Sciences AI2050 annual Fellows convening, San Francisco.
  3. 1 August 2026
    Invited participant at the Center for AI Safety Agents Workshop.
  4. 30 July 2026
    Co-organising the AI4GOOD workshop at NeurIPS 2026 in Paris, 12 December, with Terry Zhang. Its multi-agent security and safety track is led by Klaudia Krawiecka (Meta) and Swapneel Mehta (MIT).
  5. 23 July 2026
    Christian Schroeder de Witt will give a keynote on the fundamental limits of security at the NeurIPS 2026 Workshop on Foundations of Language Model Security, Paris, 12 to 13 December.
  6. 15 July 2026
    Invited speaker at Swiss AI Safety Days 2026, ETH Zurich, 7 to 8 November.
  7. 11 June 2026
    Multi-agent security named among the research clusters of Scaling AI Safety for a Multi-Agent World, a funding call of up to $10 million from Google DeepMind, Schmidt Sciences, the Cooperative AI Foundation, ARIA and Google.org.
  8. 1 May 2026
    Four papers accepted at ICML 2026: Architecture Matters for Multi-Agent Security, LongCoT, h1 and Rubric Curriculum RL. h1 was selected for a spotlight. Congratulations to Sumeet Motwani, Alesia Ivanova, Tejas Krishnan, Ben Hagag and Will Anderson.
  9. 29 April 2026
    Version 2 of Open Challenges in Multi-Agent Security released on arXiv, with 24 authors across the Witt Lab, Meta, Carnegie Mellon, the Alan Turing Institute, King's College London, Arizona State, NYU, Qualcomm, SAP, Zenity, Contramont Research, the ACM and the OWASP GenAI Security Project's Agentic Security Initiative. Lab authors include Ben Hagag, Will Anderson, Sumeet Motwani, Chandler Smith and Andis Draguns.
  10. 29 April 2026
    Christian Schroeder de Witt co-leads the research council of the OWASP GenAI Security Project's Agentic Security Initiative.
  11. 11 February 2026
    Christian Schroeder de Witt awarded a five-year EPSRC Open Fellowship, valued at about £2.26 million, for his work on multi-agent security. Announcement by the Department of Engineering Science, Oxford.
  12. 10 February 2026
    ARIA's £49.8 million Scaling Trust programme opens its first call for proposals, with a Multi-Agent Security Arena as its central testing ground and a fundamental research track on provable security guarantees for agents.
  13. 24 November 2025
    Multi-agent security highlighted in a Bloomberg opinion piece by Gideon Lichfield, and on Bruce Schneier's site.
  14. 22 November 2025
    Co-hosted the Multi-Agent Security Workshop at DAI'25, King's College London, with a public call for mixed-autonomy threat models.
  15. 5 November 2025
    Schmidt Sciences AI2050 Early Career Fellowship, one of 21 awarded globally, for the research agenda on undetectable threats. Announcements by Schmidt Sciences, the University of Oxford and Forbes.
  16. 30 September 2025
    Royal Academy of Engineering Research Fellowship, a five-year award made to 12 UK early-career researchers this year. Announcements by the Academy and the Department of Engineering Science, Oxford.
  17. 19 July 2025
    Best Paper Award at the CFAgentic Workshop, ICML 2025, for MAD-Sherlock, with the Torr Vision Group, Adobe Research and the BBC. Joint first authors Kumud Lakara and Georgia Channing.
  18. 1 July 2025
  19. 26 September 2024
  20. 16 January 2024
  21. 20 December 2023
  22. 9 March 2023
    Perfectly Secure Steganography Using Minimum Entropy Coupling, with Samuel Sokota, Zico Kolter, Jakob Foerster and Martin Strohmeier, accepted at ICLR 2023. It closes a 25-year open problem. University of Oxford press release.