INTEGRATOR
INTEGRATOR logo
Dirk Hovy, scientific director of DMI and Professor of computer science, has won an ERC starting grant of 1.5mln euros. His project INTEGRATOR, funded under grant agreement 949944, introduces demographic factors into language processing systems, which will improve algorithmic performance, avoid racism, sexism, and ageism, and open up new applications. What if I wrote that “winning an ERC Grant, Dirk Hovy got a sick result?”. Those familiar with the use of “sick” as a synonym for “great” or “awesome” among teenagers would think that Bocconi Knowledge hired a very young writer (or someone posing as such). The rest would think I went crazy. Current artificial intelligence-based language systems wouldn’t have a clue. “Natural language processing (NLP) technologies,” Prof. Hovy says, “fail to account for demographics both in understanding language and in generating it. And this failure prevents us from reaching human-like performance. It limits possible future applications and it introduces systematic bias against underrepresented demographic groups”.
🗞️🗞️ Related articles featured in Corriere Innovazione and Bocconi News.
Related
Publications
Consistency is Key: Disentangling Label Variation in Natural Language Processing with Intra-Annotator Agreement
November, 2025
No for Some, Yes for Others: Persona Prompts and Other Sources of False Refusal in Language Models
November, 2025
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
September, 2025
The AI Gap: How Socioeconomic Status Affects Language Technology Interactions
July, 2025
Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions
July, 2025
Educators' Perceptions of Large Language Models as Tutors: Comparing Human and AI Tutors in a Blind Text-only Setting
July, 2025
Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps
November, 2024
Countering Hateful and Offensive Speech Online - Open Challenges
September, 2024
Divine LLaMAs: Bias, Stereotypes, Stigmatization, and Emotion Representation of Religion in Large Language Models
September, 2024
Compromesso! Italian Many-Shot Jailbreaks Undermine the Safety of Large Language Models
August, 2024
My Answer is C: First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models
August, 2024
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviors in Large Language Models
July, 2024
Beyond Flesch-Kincaid: Prompt-based Metrics Improve Difficulty Classification of Educational Texts
May, 2024
Wisdom of Instruction-Tuned Language Model Crowds. Exploring Model Label Variation
May, 2024
Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
May, 2024
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
April, 2024
Angry Men, Sad Women: Large Language Models Reflect Gendered Stereotypes in Emotion Attribution
March, 2024
Conversations as a Source for Teaching Scientific Concepts at Different Education Levels
March, 2024
Emotion Analysis in NLP: Trends, Gaps and Roadmap for Future Directions
March, 2024
Explaining Speech Classification Models via Word-Level Audio Segments and Paralinguistic Features
March, 2024
Classist Tools: Social Class Correlates with Performance in NLP
March, 2024
Impoverished Language Technology: The Lack of (Social) Class in NLP
March, 2024
Subjective isms? On the Danger of Conflating Hate and Offence in Abusive Language Detection
March, 2024
A Tale of Pronouns: Interpretability Informs Gender Bias Mitigation for Fairer Instruction-Tuned Machine Translation
December, 2023
Mirages. On Anthropomorphism in Dialogue Systems
December, 2023
The Empty Signifier Problem: Towards Clearer Paradigms for Operationalising 'Alignment' in Large Language Models
November, 2023
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
November, 2023
The Past, Present and Better Future of Feedback Learning in Large Language Models for Subjective Human Preferences and Values
October, 2023
Wisdom of Instruction-Tuned Language Model Crowds: Exploring Model Label Variation
July, 2023
MilaNLP at SemEval-2023 Task 10: Ensembling Domain-Adapted and Regularized Pretrained Language Models for Robust Sexism Detection
July, 2023
Respectful or Toxic? Using Zero-Shot Learning with Language Models to Detect Hate Speech
July, 2023
Temporal and Second Language Influence on Intra-Annotator Agreement and Stability in Hate Speech Labelling
July, 2023
The Ecological Fallacy in Annotation: Modeling Human Label Variation goes beyond Sociodemographics
July, 2023
The State of Profanity Obfuscation in Natural Language Processing Scientific Publications
July, 2023
What about ''em''? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns
July, 2023
What about ''em''? How Commercial Machine Translation Fails to Handle (Neo-)Pronouns
July, 2023
Computer says “No”: The Case Against Empathetic Conversational AI
June, 2023
Can Demographic Factors Improve Text Classification? Revisiting Demographic Adaptation in the Age of Transformers
May, 2023
Know Your Audience: Do LLMs Adapt to Different Age and Education Levels?
Viewpoint: Artificial Intelligence Accidents Waiting to Happen?
January, 2023
Twitter-Demographer: A Flow-based Tool to Enrich Twitter Data
December, 2022
Bridging Fairness and Environmental Sustainability in Natural Language Processing
December, 2022
SocioProbe: What, When, and Where Language Models Learn about Sociodemographics
December, 2022
Data-Efficient Strategies for Expanding Hate Speech Detection into Under-Resourced Languages
October, 2022
Is It Worth the (Environmental) Cost? Limited Evidence for the Benefits of Diachronic Continuous Training
October, 2022
Welcome to the Modern World of Pronouns: Identity-Inclusive Natural Language Processing beyond Gender
October, 2022
Guiding the Release of Safer E2E Conversational AI through Value Sensitive Design
September, 2022
Hard and Soft Evaluation of NLP models with BOOtSTrap SAmpling - BooStSa
May, 2022
Benchmarking Post-Hoc Interpretability Approaches for Transformer-based Misogyny Detection
April, 2022
Measuring Harmful Sentence Completion in Language Models for LGBTQIA+ Individuals
April, 2022
Pipelines for Social Bias Testing of Large Language Models
April, 2022
Two Contrasting Data Annotation Paradigms for Subjective NLP Tasks
April, 2022
XLM-EMO: Multilingual Emotion Prediction in Social Media Text
April, 2022
Fair and Argumentative Language Modeling for Computational Argumentation
April, 2022
Entropy-based Attention Regularization Frees Unintended Bias Mitigation from Lists
March, 2022
SAFETYKIT: First Aid for Measuring Safety in Open-domain Conversational Systems
March, 2022
Five sources of bias in natural language processing
August, 2021
Exposing the limits of Zero-shot Cross-lingual Hate Speech Detection
August, 2021
HONEST: Measuring Hurtful Sentence Completion in Language Models
June, 2021
The Importance of Modeling Social Factors of Language: Theory and Practice
June, 2021