• Hi! I’m Constanza 👋🏼

    I’m a DDSA Postdoc and MATS fellow. I received my PhD from the University of Copenhagen, supervised by Anders Søgaard. During my PhD I was lucky to collaborate with Anthropic Alignment Science team, and do internships at Google DeepMind in the language understanding team, and at Apple in the Siri Understanding team. Before starting my PhD, I worked as a Software Engineer at Google, scaling ML classifiers development at Youtube (based in Paris). Previous to that, I completed my engineering degree in Computer Science at the University of Chile.

    I’m generally interested in knowledge representations and reasoning, multilinguality, and interpretability to build more trustworthy models. I also enjoy learning about psychology, philosophy and cognitive sciences. Non work related, I really enjoy running, hiking, cooking, taking photos, and exploring new places.

Selected publications

Steering Language Models with Weight Arithmetic

Constanza Fierro, Fabien Roger
The Fourteenth International Conference on Learning Representations (ICLR), 2026.
paper code

How Do Multilingual Language Models Remember Facts?

Constanza Fierro, Negar Foroutan, Desmond Elliott, Anders Søgaard
Findings of the Association for Computational Linguistics (ACL), 2025.
paper code

Defining Knowledge: Bridging Epistemology and Large Language Models

Constanza Fierro, Ruchira Dhar, Filippos Stamatiou, Nicolas Garneau, Anders Søgaard
Empirical Methods in Natural Language Processing (EMNLP), 2024.
paper

Learning to Plan and Generate Text with Citations

Constanza Fierro, Reinald Kim Amplayo, Fantine Huot, Nicola De Cao, Joshua Maynez, Shashi Narayan, Mirella Lapata
Association for Computational Linguistics (ACL), 2024.
paper

MuLan: A Study of Fact Mutability in Language Models

Constanza Fierro, Nicolas Garneau, Emanuele Bugliarello, Yova Kementchedjhieva, Anders Søgaard
North American Chapter of the Association for Computational Linguistics (NAACL), 2024.
paper code