Witt Lab

Publications

Strand
Year
Download BibTeX
  1. 2026
    Humanity's Last Exam: A benchmark of expert-level academic questions to assess AI capabilitiesL. Phan, A. Gatti, Z. Han, N. Li, J. Hu, H. Zhang, C. B. C. Zhang, M. Shaaban, and others
    Nature 649, 1139–1146Paper
  2. 2026
    Architecture Matters for Multi-Agent SecurityB. Hagag, W. L. Anderson, C. Schroeder de Witt, S. Scheffler
    ICML 2026Paper
  3. 2026
    Tool Use Enables Undetectable Steganography in Multi-Agent LLM SystemsJ. L. Rippin, S. C. Marshall, D. D. Africa, C. Schroeder de Witt
    arXiv preprint 2606.28425Paper
  4. 2026
    A Decision-Theoretic Formalisation of Steganography With Applications to LLM MonitoringU. Anwar, J. Piskorz, D. D. Baek, D. Africa, J. Weatherall, M. Tegmark, and others
    arXiv preprint 2602.23163Paper
  5. 2026
    Detecting Multi-Agent Collusion Through Multi-Agent InterpretabilityA. Rose, C. Cullen, S. Abdelnabi, P. Torr, B. G. Kaplowitz, C. Schroeder de Witt
    arXiv preprint 2604.01151Paper
  6. 2026
    LongCoT: Benchmarking Long-Horizon Chain-of-Thought ReasoningS. R. Motwani, D. Nichols, Charlie London, P. Li, and others
    ICML 2026Paper
  7. 2026
    A Note on the Strategic Confinement ProblemC. Schroeder de Witt
    arXiv preprint 2606.09931Paper
  8. 2026
    Chronos: The AI Co-HistorianL. Hufe, N. Griesshaber, G. Greif, S. O. Eck, P. Francois, W. Samek, C. Schroeder de Witt, and others
    arXiv preprint 2604.03553Paper
  9. 2026
    Spatial Generalization Tests for Machine Learning-based Weather Models to Assess Physical ConsistencyM. Höver, M. Klöwer, C. Schroeder de Witt, H. M. Christensen
    arXiv preprint; abstract at EGU 2026Paper
  10. 2026
    Exploring the Cryptographic Limits of Transformer NetworksS. Domunco, A. Draguns, P. Torr, I. Robinson, C. Schroeder de Witt
    arXiv preprint 2606.29389Paper
  11. 2026
    A Low-Rank Subspace Analysis of LLM InterventionsA. Sharma, C. Schroeder de Witt, P. Torr, A. Calinescu, J. Yu
    arXiv preprint 2606.14388Paper
  12. 2026
    When Language Representations Interact: Separability and Cross-Lingual Effects in LLMsB. Marinov, A. Sharma, C. Schroeder de Witt, P. Torr, A. Calinescu, J. Yu
    arXiv preprint; workshop version at CoLoRAI, ICML 2026Paper
  13. 2026
    Multimodal Model Diffing for Feature Discovery and ControlH. Batra, L. Naghashyar, A. Khakzar, P. Torr, R. Clark, C. Schroeder de Witt, C. Venhoff
    arXiv preprint; workshop version at AI4GOOD, ICML 2026Paper
  14. 2026
    Towards Understanding Multimodal Fine-Tuning: Spatial FeaturesL. Naghashyar, H. Batra, A. Khakzar, P. Torr, R. Clark, C. Schroeder de Witt, C. Venhoff
    arXiv preprint 2602.08713Paper
  15. 2026
    Entangled Representations Amplify Collateral Damage in UnlearningE. Wybitul, T. G. J. Rudner, C. Schroeder de Witt
    arXiv preprint 2609.02285Paper
  16. 2026
    OpenSanctions Pairs: Large-Scale Entity Matching with LLMsC. Smith, M. Sesodia, F. Lindenberg, C. Schroeder de Witt, and others
    arXiv preprint 2603.11051Paper
  17. 2026
    PerturbAgent: An Agentic AI System for Analysis and Prediction of Genetic PerturbationsK. Pei, S. Qu, P. Torr, J. G. Hedley, C. Schroeder de Witt
    Second Workshop on XAI4Science, 2026
  18. 2026
    LLM-guided Acquisition for Pathway-specific Perturb-seq Design under Experimental BudgetsM. Aiyar, K. Pei, S. Qu, P. Torr, C. Schroeder de Witt, W. J. Bolton, J. G. Hedley
    Workshop on Generative and Agentic AI for Biology, 2026
  19. 2026
    AI Models Can Provably Hide Arbitrary CapabilitiesA. Draguns, S. R. Motwani, R. Douglas, C. Schroeder de Witt
    Preprint
  20. 2026
    h1: Bootstrapping LLMs to Reason over Longer Horizons via Reinforcement LearningA. Ivanova*, S. R. Motwani*, Z. Cai, P. Torr, R. Islam, S. Shah, C. Schroeder de Witt, Charlie London (* equal contribution, † joint supervision)
    ICML 2026Paper
  21. 2026
    Rubric Curriculum RL: Exploiting the Generation-Verification Gap in Non-Verifiable DomainsT. Krishnan*, S. R. Motwani*, C. London, S. M. Bhat, H. Jiao, P. Torr, R. Islam, C. Summerfield, C. Schroeder de Witt, Q. Gu, S. Shah (* equal contribution)
    ICML 2026
  22. 2026
    Vet Your Agent: Towards Host-independent Autonomy via Verifiable Execution TracesA. Grigor, C. Schroeder de Witt, S. Birnbach, I. Martinovic
    ACM ASIA CCS 2026
  23. 2026
    Delta-Influence: Unlearning Poisons via Influence FunctionsW. Li, J. Li, P. Zeng, C. Schroeder de Witt, A. Prabhu, A. Sanyal
    Transactions on Machine Learning ResearchPaper
  24. 2026
    Fact-checking with Contextual Narratives: Leveraging Retrieval-augmented LLMs for Social Media AnalysisA. U. Dey, M. J. Awan, G. Channing, C. Schroeder de Witt, J. Collomosse
    IEEE Transactions on Computational Social Systems
  25. 2025
    PSyDUCK: Hiding Information in the Denoising Process of Latent Diffusion ModelsA. Mahfuz, G. Channing, M. van der Wilk, P. H. S. Torr, F. Pizzati, C. Schroeder de Witt
    IEEE WIFS 2025Paper
  26. 2025
    Multi-Agent Security Tax: Trading Off Security and Collaboration Capabilities in Multi-Agent SystemsP. Peigné, M. Kniejski, F. Sondej, M. David, J. Hoelscher-Obermaier, C. Schroeder de Witt, E. Kran
    AAAI 2025Paper
  27. 2025
    Multi-Agent Risks from Advanced AIL. Hammond, A. Chan, J. Clifton, J. Hoelscher-Obermaier, A. Khan, and others
    arXiv preprint 2502.14143Paper
  28. 2025
    Open Challenges in Multi-Agent Security: Towards Secure Systems of Interacting AI AgentsC. Schroeder de Witt, K. Krawiecka, I. Krawczuk, B. Hagag, and others
    arXiv preprint 2505.02077Paper
  29. 2025
    MALT: Improving Reasoning with Multi-Agent LLM TrainingS. R. Motwani, C. Smith, R. J. Das, R. Rafailov, I. Laptev, P. H. S. Torr, F. Pizzati, R. Clark, C. Schroeder de Witt
    COLM 2025Paper
  30. 2025
    Fundamental Limitations in Pointwise Defences of LLM Finetuning APIsX. Davies, E. Winsor, A. Souly, T. Korbak, R. Kirk, C. Schroeder de Witt, Y. Gal
    NeurIPS 2025Paper
  31. 2025
    Mitigating Goal Misgeneralization via Minimax RegretK. A. Sadek, M. Farrugia-Roberts, U. Anwar, and others
    arXiv preprint 2507.03068Paper
  32. 2025
    AnnoCaseLaw: A Richly-annotated Dataset for Benchmarking Explainable Legal Judgment PredictionM. Sesodia, A. Petrova, J. Armour, T. Lukasiewicz, C. Schroeder de Witt, and others
    arXiv preprint 2503.00128Paper
  33. 2025
    Mixture of Experts Made Intrinsically InterpretableX. Yang, C. Venhoff, A. Khakzar, C. Schroeder de Witt, P. K. Dokania, A. Bibi, P. Torr
    arXiv preprint 2503.07639Paper
  34. 2025
    REAL: Benchmarking Autonomous Agents on Deterministic Simulations of Real WebsitesD. Garg, D. Caples, A. Draguns, N. Ravi, P. Putta, N. Garg, P. Hebbar, and others, C. Schroeder de Witt, S. Motwani
    NeurIPS 2025Paper
  35. 2025
    Architecting Resilient LLM Agents: A Guide to Secure Plan-then-Execute ImplementationsR. F. Del Rosario, K. Krawiecka, C. Schroeder de Witt
    arXiv preprint 2509.08646Paper
  36. 2025
    Extending the OWASP Multi-Agentic System Threat Modeling Guide: Insights from Multi-Agent Security ResearchK. Krawiecka, C. Schroeder de Witt
    arXiv preprint 2508.09815Paper
  37. 2025
    Predicting Weak-to-Strong Generalization from Latent RepresentationsB. Wilop, C. Schroeder de Witt, Y. Gal, P. Torr, C. Venhoff
    Preprint
  38. 2025
    DEEDEE: Fast and Scalable Out-of-Distribution Dynamics DetectionT. Aljaafari, V. Kanade, P. Torr, C. Schroeder de Witt
    arXiv preprint 2510.21638Paper
  39. 2025
    Testing the Limits of the World's Largest Control Task: Solar Geoengineering as a Deep Reinforcement Learning ProblemE. Agrawal, C. Schroeder de Witt
    Geoengineering and Climate Change: Methods, Risks, and Governance (Wiley, 2025)
  40. 2025
    Hidden in Plain Text: Emergence and Mitigation of Steganographic Collusion in LLMsY. Mathew, O. Matthews, R. McCarthy, J. Velja, C. Schroeder de Witt, D. Cope, and others
    IJCNLP 2025. Social Impact AwardPaper
  41. 2025
    LLM-Consensus (formerly MAD-Sherlock): Multi-Agent Debate for Visual Misinformation DetectionK. Lakara, G. Channing, J. Sock, C. Rupprecht, P. Torr, J. Collomosse, C. Schroeder de Witt
    CFAgentic Workshop, ICML 2025. Oral and Best Paper AwardPaper
  42. 2025
    Efficient Dictionary Learning with Switch Sparse AutoencodersA. Mudide, J. Engels, E. J. Michaud, M. Tegmark, C. Schroeder de Witt
    ICLR 2025Paper
  43. 2024
    Illusory Attacks: Information-Theoretic Detectability Matters in Adversarial AttacksT. Franzmeyer, S. M. McAleer, J. F. Henriques, J. N. Foerster, P. Torr, A. Bibi, C. Schroeder de Witt
    ICLR 2024, SpotlightPaper
  44. 2024
    Secret Collusion among AI Agents: Multi-Agent Deception via SteganographyS. R. Motwani, M. Baranchuk, M. Strohmeier, V. Bolina, P. H. S. Torr, L. Hammond, C. Schroeder de Witt
    NeurIPS 2024Paper
  45. 2024
    Unelicitable Backdoors in Language Models via Cryptographic Transformer CircuitsA. Draguns*, A. Gritsevskiy*, S. R. Motwani, C. Rogers-Smith, J. Ladish, C. Schroeder de Witt (* equal contribution)
    NeurIPS 2024Paper
  46. 2024
    Computing Low-Entropy Couplings for Large-Support DistributionsS. Sokota, D. Sam, C. Schroeder de Witt, S. Compton, J. Foerster, J. Z. Kolter
    UAI 2024Paper
  47. 2024
    JaxMARL: Multi-Agent RL Environments and Algorithms in JAXA. Rutherford, B. Ellis, M. Gallici, J. Cook, A. Lupu, G. Ingvarsson, T. Willi, and others
    NeurIPS 2024, Datasets and BenchmarksPaper
  48. 2024
    Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and DetectionL. Nasvytis, K. Sandbrink, J. Foerster, T. Franzmeyer, C. Schroeder de Witt
    AAMAS 2024, OralPaper
  49. 2024
    Foundational Challenges in Assuring Alignment and Safety of Large Language ModelsU. Anwar, A. Saparov, J. Rando, D. Paleka, M. Turpin, P. Hase, E. S. Lubana, and others
    Transactions on Machine Learning ResearchPaper
  50. 2024
    Position: Near to Mid-term Risks and Opportunities of Open-Source Generative AIF. Eiras, A. Petrov, B. Vidgen, C. Schroeder de Witt, F. Pizzati, K. Elkins, and others
    ICML 2024Paper
  51. 2024
    Risks and Opportunities of Open-Source Generative AIF. Eiras, A. Petrov, B. Vidgen, C. Schroeder de Witt, F. Pizzati, K. Elkins, and others
    arXiv preprint 2405.08597Paper
  52. 2024
    IDs for AI SystemsA. Chan, N. Kolt, P. Wills, U. Anwar, C. Schroeder de Witt, N. Rajkumar, L. Hammond, D. Krueger, L. Heim, M. Anderljung
    arXiv preprint 2406.12137Paper
  53. 2024
    Comparative Global AI Regulation: Policy Perspectives from the EU, China, and the USJ. Chun, C. Schroeder de Witt, K. Elkins
    arXiv preprint 2410.21279Paper
  54. 2024
    Toward Robust Real-World Audio Deepfake Detection: Closing the Explainability GapG. Channing, J. Sock, R. Clark, P. Torr, C. Schroeder de Witt
    arXiv preprint 2410.07436Paper
  55. 2024
    SAGE: Scalable Ground Truth Evaluations for Large Sparse AutoencodersC. Venhoff, A. Calinescu, P. Torr, C. Schroeder de Witt
    arXiv preprint 2410.07456Paper
  56. 2024
    The Danger of Arrogance: Welfare Equilibria as a Solution to Stackelberg Self-Play in Non-Coincidental GamesJ. Levi, C. Lu, T. Willi, C. Schroeder de Witt, J. Foerster
    arXiv preprint 2402.01088Paper
  57. 2024
    Can Reinforcement Learning Model Learning across Development? Online Lifelong Learning through Adaptive Intrinsic MotivationK. J. Sandbrink, B. Christian, L. M. Nasvytis, C. Schroeder de Witt, P. Butlin
    Proceedings of the Annual Meeting of the Cognitive Science Society 46
  58. 2024
    Using Adaptive Intrinsic Motivation in RL to Model Learning across DevelopmentK. J. Sandbrink, B. Christian, L. Nasvytis, C. Schroeder de Witt, P. Butlin
    Intrinsically Motivated and Open-Ended Learning Workshop, NeurIPS 2024
  59. 2024
    Proofs of Autonomy: Scalable and Practical Verification of AI AutonomyA. Grigor, C. Schroeder de Witt, I. Martinovic
    ICML Workshop on Technical AI Governance, 2024
  60. 2024
    Bayesian Exploration NetworksM. Fellows, B. Kaplowitz, C. Schroeder de Witt, S. Whiteson
    ICML 2024Paper
  61. 2023
    Perfectly Secure Steganography Using Minimum Entropy CouplingC. Schroeder de Witt*, S. Sokota*, J. Z. Kolter, J. N. Foerster, M. Strohmeier (* equal contribution)
    ICLR 2023. Covered by Scientific American, Quanta Magazine and Bruce SchneierPaper
  62. 2023
    Cheap Talk Discovery and Utilization in Multi-Agent Reinforcement LearningY. L. Lo, C. Schroeder de Witt, S. Sokota, J. N. Foerster, S. Whiteson
    ICLR 2023
  63. 2022
    Communicating via Markov Decision ProcessesS. Sokota, C. Schroeder de Witt, M. Igl, L. M. Zintgraf, P. Torr, M. Strohmeier, Z. Kolter, S. Whiteson, J. Foerster
    ICML 2022Paper
  64. 2022
    Equivariant Networks for Zero-Shot CoordinationD. Muglich, C. Schroeder de Witt, E. van der Pol, S. Whiteson, J. Foerster
    NeurIPS 2022Paper
  65. 2022
    Discovered Policy OptimisationC. Lu, J. Kuba, A. Letcher, L. Metz, C. Schroeder de Witt, J. Foerster
    NeurIPS 2022Paper
  66. 2022
    Model-Free Opponent ShapingC. Lu, T. Willi, C. Schroeder de Witt, J. Foerster
    ICML 2022Paper
  67. 2022
    Mirror Learning: A Unifying Framework of Policy OptimisationJ. G. Kuba, C. Schroeder de Witt, J. Foerster
    ICML 2022Paper
  68. 2022
    Generalized Beliefs for Cooperative AID. Muglich, L. M. Zintgraf, C. Schroeder de Witt, S. Whiteson, J. Foerster
    ICML 2022Paper
  69. 2022
    Amortized Rejection Sampling in Universal Probabilistic ProgrammingS. Naderiparizi, A. Ścibior, A. Munk, M. Ghadiri, A. G. Baydin, B. Gram-Hansen, C. Schroeder de Witt, R. Zinkov, P. Torr, T. Rainforth, Y. W. Teh, F. Wood
    AISTATS 2022
  70. 2022
    Revealing Robust Oil and Gas Company Macro-Strategies Using Deep Multi-Agent Reinforcement LearningD. Radovic, L. Kruitwagen, C. Schroeder de Witt, B. Caldecott, S. Tomlinson, M. Wolf
    arXiv preprint 2211.11043Paper
  71. 2021
    FACMAC: Factored Multi-Agent Centralised Policy GradientsB. Peng, T. Rashid, C. Schroeder de Witt, P. A. Kamienny, P. Torr, W. Böhmer, S. Whiteson
    NeurIPS 2021Paper
  72. 2021
    Randomized Entity-wise Factorization for Multi-Agent Reinforcement LearningS. Iqbal, C. Schroeder de Witt, B. Peng, W. Böhmer, S. Whiteson, F. Sha
    ICML 2021Paper
  73. 2021
    RainBench: Towards Data-Driven Global Precipitation Forecasting from Satellite ImageryC. Schroeder de Witt, C. Tong, V. Zantedeschi, D. De Martini, A. Kalaitzis, M. Chantry, D. Watson-Parris, P. Bilinski
    AAAI 2021Paper
  74. 2021
    Fixed Points in Cyber Space: Rethinking Optimal Evasion Attacks in the Age of AI-NIDSY. Huang, C. Schroeder de Witt, P. H. S. Torr, M. Strohmeier
    arXiv preprint 2111.12197Paper
  75. 2021
    Coordination and Communication in Deep Multi-Agent Reinforcement LearningC. Schroeder de Witt
    DPhil thesis, University of Oxford
  76. 2020
    Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningT. Rashid, M. Samvelyan, C. Schroeder de Witt, G. Farquhar, J. Foerster, S. Whiteson
    Journal of Machine Learning Research 21Paper
  77. 2020
    Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?C. Schroeder de Witt, T. Gupta, D. Makoviichuk, V. Makoviychuk, P. H. S. Torr, M. Sun, S. Whiteson
    arXiv preprint 2011.09533Paper
  78. 2020
    Deep Multi-Agent Reinforcement Learning for Decentralized Continuous Cooperative ControlC. Schroeder de Witt, B. Peng, P. A. Kamienny, P. Torr, W. Böhmer, S. Whiteson
    arXiv preprint 2003.06709Paper
  79. 2020
    Towards Data-Driven Physics-Informed Global Precipitation Forecasting from Satellite ImageryV. Zantedeschi, D. De Martini, C. Tong, C. Schroeder de Witt, A. Kalaitzis, M. Chantry, D. Watson-Parris
    AI for Earth Sciences Workshop, NeurIPS 2020
  80. 2020
    Artificial Intelligence and Climate Change: Supplementary Impact ReportT. Walsh, A. Evatt, C. Schroeder de Witt
    Report
  81. 2019
    The StarCraft Multi-Agent ChallengeM. Samvelyan, T. Rashid, C. Schroeder de Witt, G. Farquhar, N. Nardelli, T. G. J. Rudner, C. M. Hung, P. H. S. Torr, J. Foerster, S. Whiteson
    AAMAS 2019Paper
  82. 2019
    Multi-Agent Common Knowledge Reinforcement LearningC. Schroeder de Witt, J. Foerster, G. Farquhar, P. Torr, W. Böhmer, S. Whiteson
    NeurIPS 2019Paper
  83. 2019
    Stratospheric Aerosol Injection as a Deep Reinforcement Learning ProblemC. Schroeder de Witt, T. Hornigold
    Tackling Climate Change with Machine Learning Workshop, ICML 2019. Best Idea AwardPaper
  84. 2019
    Hijacking Malaria Simulators with Probabilistic ProgrammingB. Gram-Hansen, C. Schroeder de Witt, T. Rainforth, P. H. S. Torr, Y. W. Teh, A. G. Baydin
    AI for Social Good Workshop, ICML 2019Paper
  85. 2015
    Safe Screening for Support Vector MachinesJ. Zimmert, C. Schroeder de Witt, G. Kerg, M. Kloft
    OPT Workshop, NIPS 2015
  86. 2014
    The ZX-Calculus is Incomplete for Quantum MechanicsC. Schroeder de Witt, V. Zamdzhiev
    Quantum Physics and Logic (QPL) 2014Paper

A selection of the lab's publications. The full record is on Google Scholar.