<?xml version="1.0" encoding="utf-8" standalone="yes" ?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>current | MilaNLP Lab @ Bocconi University</title>
    <link>https://milanlproc.github.io/tags/current/</link>
      <atom:link href="https://milanlproc.github.io/tags/current/index.xml" rel="self" type="application/rss+xml" />
    <description>current</description>
    <generator>Source Themes Academic (https://sourcethemes.com/academic/)</generator><language>en-us</language><copyright>© MilaNLP, 2026</copyright><lastBuildDate>Fri, 12 Dec 2025 00:00:00 +0000</lastBuildDate>
    <image>
      <url>img/map[gravatar:%!s(bool=false) shape:circle]</url>
      <title>current</title>
      <link>https://milanlproc.github.io/tags/current/</link>
    </image>
    
    <item>
      <title>SALMON</title>
      <link>https://milanlproc.github.io/project/salmon/</link>
      <pubDate>Fri, 12 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://milanlproc.github.io/project/salmon/</guid>
      <description>&lt;p&gt;&lt;strong&gt;SALMON – Social Awareness for better Large Language Model Learning&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Funded by: 
&lt;a href=&#34;https://tef.tech/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Tech Europe Foundation (TEF)&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;Despite consuming massive amounts of textual data, current large language models (LLMs) still struggle with tasks that require social awareness, such as understanding moral values and making decisions in sensitive situations. As a result, models are ill-aligned with human values and lack actionable social knowledge (i.e., &amp;ldquo;reading the room&amp;rdquo;). This gap is critical, as LLMs are increasingly used in sensitive areas, such as intercultural business negotiations, educational applications, and mental health support, that require social awareness.&lt;/p&gt;
&lt;p&gt;However, it is impossible to learn true social awareness through passive consumption of text. Current training approaches do not explicitly address critical social elements or explicitly represent social awareness as an objective, but approximate its effects through post-hoc instruction fine-tuning. Continuing a text-only training paradigm and hoping that social awareness will somehow emerge limits performance, potentially harming users by catastrophically misreading situations. Instead, models need to acquire social awareness as an integral part of their pre-training to match the areas they are already expected to cover. To develop genuine social understanding&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>TOLD</title>
      <link>https://milanlproc.github.io/project/told/</link>
      <pubDate>Fri, 12 Dec 2025 00:00:00 +0000</pubDate>
      <guid>https://milanlproc.github.io/project/told/</guid>
      <description>&lt;p&gt;&lt;strong&gt;TOLD – Thinking Out Loud: A Speech-Based Data Collection Framework&lt;/strong&gt;&lt;br&gt;
&lt;strong&gt;Funded by: 
&lt;a href=&#34;https://tef.tech/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Tech Europe Foundation (TEF)&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;This project is in collaboration with 
&lt;a href=&#34;https://gattanasio.cc/&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Dr. Giuseppe Attanasio&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;Recent advances in language modeling show that the quality of training data matters more than its quantity. Yet collecting meaningful and representative language data remains costly and slow because annotation still relies almost entirely on written text. Voice-based feedback offers a powerful alternative: it elicits richer and more natural descriptions, reflects personal experiences and subjective perspectives, and conveys paralinguistic cues such as prosody and timing that written text cannot capture. Despite this potential, voice is still largely underused in annotation.&lt;/p&gt;
&lt;p&gt;TOLD aims to show that a voice-based annotation paradigm can outperform traditional written feedback in NLP. Speaking rather than typing produces more informative and expressive data, enables faster and more efficient collection, and can lead to models that learn more effectively from the resulting annotations.&lt;/p&gt;
&lt;p&gt;By shifting data collection from text to voice, TOLD introduces a new way of capturing how people think, react, and interpret information. The project seeks to establish voice as a natural, scalable, and more powerful medium for annotating language data.&lt;/p&gt;
</description>
    </item>
    
    <item>
      <title>PERSONAE</title>
      <link>https://milanlproc.github.io/project/personae/</link>
      <pubDate>Mon, 27 Feb 2023 00:00:00 +0000</pubDate>
      <guid>https://milanlproc.github.io/project/personae/</guid>
      <description>&lt;p&gt;Debora Nozza have been recently awarded a €1.5m ERC Starting Grant project 2023 for my project PERSONAE.&lt;/p&gt;
&lt;p&gt;PERSONAE will make language technology (LT) accessible and valuable to everyone. I will revolutionize research in subjective tasks in NLP such as abusive language detection and sentiment and emotion analysis by developing a new field called personal NLP, yielding new datasets, tasks, and algorithms. This new research area will explore subjective tasks from the perspective of the individual as information receiver, making users active actors in the creation of LTs instead of mere recipients. This will allow for a more tailored, effective approach to NLP model design, resulting in better models overall.&lt;/p&gt;
&lt;p&gt;Each person has their own interests and preferences based on their background and experience. These factors impact their views of what makes them happy, angry, or depressed over time. Language technologies (LTs) can consider individual preferences. However, current research presumes a static view of subjectivity: that a single ground truth underlies subjective tasks such as abusive language detection, an assumption that lacks human variability and prevents universal access to LTs.&lt;/p&gt;
&lt;p&gt;Language-based AI such as virtual assistants is widely available. But despite significant scientific advances, most LT applications are inaccessible to individuals and their public&amp;rsquo;s opinion has become increasingly negative. GPT-3&amp;rsquo;s 2020 release boosted business-oriented applications such as copywriting and chatbots, yet few that let people improve their lives—for example, by controlling what they see on social media. This gap becomes more pronounced for subjective tasks.&lt;/p&gt;
&lt;p&gt;PERSONAE will help design subjective LTs that can be adapted by individuals at will over time. Based on an ambitious meta approach able to generalize from existing, disconnected work, PERSONAE will rely on fully personalizable privacy-aware algorithms that can be used by anyone. It will reveal benefits of LT far beyond those of existing systems, paving the way for future applications.&lt;/p&gt;
&lt;p&gt;🌏🌏 Check out the 
&lt;a href=&#34;https://www.knowledge.unibocconi.eu/notizia.php?idArt=25755&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;&lt;strong&gt;web article&lt;/strong&gt;&lt;/a&gt; on my project!&lt;/p&gt;
&lt;p&gt;🎙️🎙️ Check out my latest &lt;strong&gt;interview&lt;/strong&gt; on 
&lt;a href=&#34;https://www.radio24.ilsole24ore.com/programmi/smart-city/puntata/trasmissione-7-dicembre-2023-6500-2400735627083863&#34; target=&#34;_blank&#34; rel=&#34;noopener&#34;&gt;Radio 24&lt;/a&gt;  in Italian!&lt;/p&gt;
</description>
    </item>
    
  </channel>
</rss>
