My research has moved across several areas of computational linguistics, from lexical semantics and discourse to authorship, framing, language models, evaluation, and responsible NLP. What’s always stayed at the centre is language: how people use it, what we can learn from their choices, and what happens when these choices are modelled or mediated by machines.
There are frequently situations where I have to summarise in just one line what I’m interested in research-wise. Observing what I’ve resorted to say during the years, paired with what I do in practice, I’d say now that what I’m most interested in is how people write. This does not cover everything I do, but it definitely captures a recurring thread across much of my work.
Meaning, framing, and perspective
I have long worked on the relation between linguistic choices and interpretation. My earlier research addressed lexical semantics, reference, metonymy, information structure, and discourse; more recent work focuses on framing, perspective, narratives, and event representation.
In particular, I am interested in how alternative linguistic descriptions of the same event affect interpretation and attribution of responsibility, and in computational representations that preserve multiple perspectives.
Individual and social variation in language
I have studied linguistic variation across individuals, groups, and genres. A substantial part of this work has focused on authorship profiling and attribution, including shared tasks and international evaluation campaigns, and on the relation between writing style and social or demographic variables.
Language models: making, representation, and control
I’ve worked on the development and analysis of neural language models, particularly beyond English. This includes BERTje for Dutch, GePpeTto and IT5 for Italian, as well as work on multilingual and under-resourced settings.
My more recent work includes stylistic and personalised generation, persona, interpretability, and model steering, with a particular interest in how model behaviour can be better understood, and controlled.
Evaluation
I’ve worked extensively on evaluation, from multiple viewpoints. At first, especially through the organisation of (and participation to) EVALITA campaigns, evaluation mostly focused on shared tasks and providing the field with reference benchmarks (notably DUMB for Dutch, and multiple ones for Italian) aimed at stimulating the development of systems to solve given tasks. In more recent times, I’ve put a lot of effort into creating a solid native Italian benchmark which could test the abilities of LLMs, rather than focusing on creating datasets for model development. I am also interested in human evaluation and reproducibility, see e.g. work at GEM and as part of ReproHum.
Language games
I love games, in particular language games, and one reason I love my job is that I can use language games as controlled settings for studying language models. They are a great testbed for testing linguistic competence and reasoning in models, and are great fun to work with. They are also an excellent tool to have students improve their technical skills.
Recent work includes rebuses, Italian crosswords, Taboo, and Dr Denker puzzles. Some of these have also become shared or community evaluation tasks, including EurekaRebus within CALAMITA and Cruciverb-IT at EVALITA.
Language technology in society
I also work on the assumptions and societal consequences of NLP systems, including bias and fairness, responsible NLP, AI literacy, and human-centred language technology. Broadly, I’m concerned with how language technology affects people, education, communication, and access to information. Multiple outreach events stem from such concerns. See for example the puzzle project.
Ongoing and recent projects
PAGINA (2024-2027)
Promoting the Application of Generative AI for Improved News Accessibility
Through a collaboration with the Dagblad van het Noorden, the main regional newspaper in the north of the Netherlands, PAGINA investigates how language technology can make Dutch news more accessible, with work on readability, text simplification, reading behaviour, perspective, and human-centred evaluation.
All Sides of the Story (2023-2027)
Language Technology for Multifaceted Event Understanding
In collaboration with Prof Hedderik van Rijn, this project studies computational representations of events that preserve multiple perspectives, with a focus on framing, perspective, and event understanding, and education as target application.
HAICu (2023-2029)
Digital Humanities - Artificial Intelligence - Cultural Heritage
HAICu is a NWA project which develops AI methods for cultural-heritage collections. Our work focuses on language technology for limited-data settings, contextual interpretation, and the representation of multiple perspectives.
CALAMITA
Challenging the Abilities of Large Language Models in Italian
CALAMITA is a community initiative for the systematic evaluation of large language models in Italian, involving more than 80 researchers from the Italian NLP community.
Framing Situations in the Dutch Language (2019-2024)
This NWO project in collaboration with VU Amsterdam studied alternative linguistic descriptions of the same event and developed computational methods and resources for analysing framing. We studied in particular news reporting of femicides and responsibility framing in Italian and Dutch. Otje Minnema’s thesis, which contains a lot of this project’s work, won the AVT/Anela Best Dissertation award.
Evaluation and Adaptation of Neural Language Models for Under-Resourced Languages (2019-2023)
This project, based on my NWO Aspasia grant, investigated neural language modelling beyond English and contributed to the development and evaluation of dedicated models for Dutch and Italian. BERTje and DUMB are associated with this project, as well as the Make the best of cross-lingual transfer paper.
Past things
See here a collection of topics that I worked on in the past but are not currently the focus of my research.
For a complete list of papers, see Publications.