mAI alignment lab
Junior Research Group at University of Bonn focusing on AI alignment and safety issues
About
Welcome to the mAI alignment lab, a Junior Research Group at the University of Bonn led by Dr. Florian Mai.
Our research focuses on AI alignment and safety issues, exploring how to ensure that current and future advanced AI systems are acting reliably in accordance with human values.
Current Projects
-
Scalable Oversight by Learning to Decompose Tasks: Exploring how AI systems can learn to break complex tasks into manageable subtasks for reliable human oversight, advancing the frontier of superalignment research.
-
Emergent Misalignment: Investigating how narrow finetuning can produce broadly misaligned language models and developing methods to prevent such misalignment.
-
Backdoor Detection: Detecting hidden behaviors in language models — such as backdoors, sleeper agents, sandbagging, and censorship — without prior knowledge of the trigger or the target behavior.
-
Value Alignment: Researching methods to ensure AI systems align with human values and preferences.
Join Us
We currently have no open positions available. However, if you are interested in collaborating with our research group, please feel free to send an email to Dr. Florian Mai at fmai@uni-bonn.de.
news
| Aug 15, 2026 | Lukasz Karwacki joined us as a research assistant! Lukasz is funded by our grant from Coefficient Giving, which supports our work on Scalable Oversight by Learning to Decompose Tasks. Welcome, Lukasz! ⚛️ |
|---|---|
| Aug 11, 2026 | New preprint! Data Attribution of Emergent Misalignment with Persona Features traces the persona features behind emergent misalignment back to pre-training data, and finds that human-written documents alone do not reliably induce it — synthetic instruction-response pairs from the same content do. Great work, Clemens and David! 🧬 |
| Aug 1, 2026 | Mohammad Faiz joined us as a HiWi! Mohammad is a master’s student at the University of Bonn and works on Emergent Misalignment. Welcome, Mohammad! 🌱 |
| Jul 23, 2026 | Our paper A Unified Moral-Value Dataset for Instruction Tuning has been accepted at the 4th International Workshop on Value Engineering in AI (VALE 2026), co-located with IJCAI-ECAI 2026 in Bremen. It merges existing moral-value datasets into a single instruction-tuning corpus, available on Hugging Face. Congratulations to Zhaohui! ⚖️ |
| Jul 15, 2026 | Our paper All too perfect: bias and aspiration in persona generation with LLMs has been published in Artificial Intelligence Review! We generate 40,000 synthetic personas across four languages and find that LLMs overwhelmingly produce middle-aged, aspirational, upbeat characters — a narrative sanitization that flattens representational diversity. Congratulations to Nicholas, Rafaela, David, Ana, and Julia! 🎭 |