mAI alignment lab

Junior Research Group at University of Bonn focusing on AI alignment and safety issues

About

Welcome to the mAI alignment lab, a Junior Research Group at the University of Bonn led by Dr. Florian Mai.

Our research focuses on AI alignment and safety issues, exploring how to ensure that current and future advanced AI systems are acting reliably in accordance with human values.

Current Projects

  • Scalable Oversight by Learning to Decompose Tasks: Exploring how AI systems can learn to break complex tasks into manageable subtasks for reliable human oversight, advancing the frontier of superalignment research.

  • Emergent Misalignment: Investigating how narrow finetuning can produce broadly misaligned language models and developing methods to prevent such misalignment.

  • Backdoor Detection: Detecting hidden behaviors in language models — such as backdoors, sleeper agents, sandbagging, and censorship — without prior knowledge of the trigger or the target behavior.

  • Value Alignment: Researching methods to ensure AI systems align with human values and preferences.

Join Us

We currently have no open positions available. However, if you are interested in collaborating with our research group, please feel free to send an email to Dr. Florian Mai at fmai@uni-bonn.de.

news

Aug 15, 2026 Lukasz Karwacki joined us as a research assistant! Lukasz is funded by our grant from Coefficient Giving, which supports our work on Scalable Oversight by Learning to Decompose Tasks. Welcome, Lukasz! ⚛️
Aug 11, 2026 New preprint! Data Attribution of Emergent Misalignment with Persona Features traces the persona features behind emergent misalignment back to pre-training data, and finds that human-written documents alone do not reliably induce it — synthetic instruction-response pairs from the same content do. Great work, Clemens and David! 🧬
Aug 1, 2026 Mohammad Faiz joined us as a HiWi! Mohammad is a master’s student at the University of Bonn and works on Emergent Misalignment. Welcome, Mohammad! 🌱
Jul 23, 2026 Our paper A Unified Moral-Value Dataset for Instruction Tuning has been accepted at the 4th International Workshop on Value Engineering in AI (VALE 2026), co-located with IJCAI-ECAI 2026 in Bremen. It merges existing moral-value datasets into a single instruction-tuning corpus, available on Hugging Face. Congratulations to Zhaohui! ⚖️
Jul 15, 2026 Our paper All too perfect: bias and aspiration in persona generation with LLMs has been published in Artificial Intelligence Review! We generate 40,000 synthetic personas across four languages and find that LLMs overwhelmingly produce middle-aged, aspirational, upbeat characters — a narrative sanitization that flattens representational diversity. Congratulations to Nicholas, Rafaela, David, Ana, and Julia! 🎭

selected publications

  1. ICML
    In-Training Defenses against Emergent Misalignment in Language Models
    David Kaczér, Magnus Jørgenvåg, Clemens Vetter, and 4 more authors
    In Proceedings of the 43rd International Conference on Machine Learning (ICML), Jul 2026
  2. IASEAI
    AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
    Leonard Dung, and Florian Mai
    In IASEAI’26: International Association for Safe and Ethical AI Conference, Feb 2026
  3. COLM
    Learning to Plan for Language Modeling from Unlabeled Data
    Nathan Cornille, Marie-Francine Moens, and Florian Mai
    In First Conference on Language Modeling, Oct 2024