Training Data

What Is The Political Content in LLMs' Pre- and Post-Training Data?

Large language models (LLMs) are known to generate politically biased text. Yet, it remains unclear how such biases arise, making it difficult to design effective mitigation strategies. We hypothesize that these biases are rooted in the composition …