publications

publications by categories in reversed chronological order. generated by jekyll-scholar.

selected

  1. ICML
    In-Training Defenses against Emergent Misalignment in Language Models
    David Kaczér, Magnus Jørgenvåg, Clemens Vetter, and 4 more authors
    In Proceedings of the 43rd International Conference on Machine Learning (ICML), Jul 2026
  2. IASEAI
    AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
    Leonard Dung, and Florian Mai
    In IASEAI’26: International Association for Safe and Ethical AI Conference, Feb 2026
  3. COLM
    Learning to Plan for Language Modeling from Unlabeled Data
    Nathan Cornille, Marie-Francine Moens, and Florian Mai
    In First Conference on Language Modeling, Oct 2024

2026

  1. Data Attribution of Emergent Misalignment with Persona Features
    Clemens Vetter, David Kaczér, Lucie Flek, and 1 more author
    Aug 2026
  2. VALE
    A Unified Moral-Value Dataset for Instruction Tuning
    Zhaohui Zeng, and Florian Mai
    In Fourth International Workshop on Value Engineering in AI (VALE), co-located with IJCAI-ECAI 2026, Aug 2026
  3. AIR
    All Too Perfect: Bias and Aspiration in Persona Generation with LLMs
    Nicholas Kluge Corrêa, Rafaela Weber Mallmann, David Kaczér, and 3 more authors
    Artificial Intelligence Review, Jul 2026
  4. ICML
    In-Training Defenses against Emergent Misalignment in Language Models
    David Kaczér, Magnus Jørgenvåg, Clemens Vetter, and 4 more authors
    In Proceedings of the 43rd International Conference on Machine Learning (ICML), Jul 2026
  5. Preprint
    Detecting Hidden Behaviors in LLMs via Activation-matched Finetuning
    Robin Haselhorst, Lucie Flek, and Florian Mai
    Jun 2026
    Preprint hosted on this website
  6. Beyond Liars’ Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs
    Amr Moustafa, Max Feser, and Florian Mai
    AI Transparency Journal, Jun 2026
    Presented at the AI Transparency Conference; forthcoming in the first edition of the AI Transparency Journal.
  7. Reinforcement Learning Amplifies Emergent Misalignment from Harmless Rewards
    Magnus Jørgenvåg, David Kaczér, Lasse Ruttert, and 3 more authors
    May 2026
  8. Reasoning Primitives in Hybrid and Non-Hybrid LLMs: Do Architectural Differences Yield Advantages in State-Tracking and Recall?
    Shivam Rawat, Lucie Flek, Florian Mai, and 1 more author
    Apr 2026
  9. Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
    Shiza Fatimah, Aniket Sen, Sophia Falk, and 3 more authors
    Mar 2026
  10. Understanding Artificial Theory of Mind: Perturbed Tasks and Reasoning in Large Language Models
    Christian Nickel, Laura Schrewe, Florian Mai, and 1 more author
    Feb 2026
  11. IASEAI
    AI Alignment Strategies from a Risk Perspective: Independent Safety Mechanisms or Shared Failures?
    Leonard Dung, and Florian Mai
    In IASEAI’26: International Association for Safe and Ethical AI Conference, Feb 2026
  12. LM4UC
    Pluralistic AI Alignment: A Cross-Cultural Pilot Survey
    Khashayar Alavi, Lucie Flek, and Florian Mai
    In Second Workshop on Language Models for Underserved Communities (LM4UC), Jan 2026

2025

  1. EMNLP
    Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models
    Mehdi Ali, Manuel Brack, Max Lübbering, and 16 more authors
    In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP), Nov 2025
  2. Survey-to-Behavior: Downstream Alignment of Human Values in LLMs via Survey Questions
    Shangrui Nie, Florian Mai, David Kaczér, and 3 more authors
    Aug 2025
  3. BiAlign
    Superalignment with Dynamic Human Values
    Florian Mai, David Kaczér, Nicholas Kluge Corrêa, and 1 more author
    In ICLR 2025 Workshop on Bidirectional Human-AI Alignment, Apr 2025

2024

  1. WiNLP
    Improving Language Modeling by Increasing Test-time Planning Compute
    Florian Mai, Nathan Cornille, and Marie-Francine Moens
    In Eighth Widening NLP Workshop (WiNLP 2024) Phase II, Nov 2024
  2. COLM
    Learning to Plan for Language Modeling from Unlabeled Data
    Nathan Cornille, Marie-Francine Moens, and Florian Mai
    In First Conference on Language Modeling, Oct 2024