The Bias Spillover Effect in LLMs

The Bias Spillover Effect in LLMs: When Fixing One Bias Breaks Another

Author – Eva Paraschou

Imagine a mental health support app powered by artificial intelligence (AI). Its developers notice the app recommends professional help more aggressively to women than to men, a clear gender bias. They fix it, and they succeed. But a few months later, something else surfaces: older users are now receiving shorter, more dismissive responses than younger users for the exact same struggles. Nobody changed anything related to age. So what happened?

What happened is the bias spillover effect, scientifically defined as “the unintended alteration of behavior on one social axis when mitigating another” (Mijalli et al., 2023). Biases in AI systems are rarely isolated. The stereotypes baked in during training are entangled across gender, age, disability, race and more. Pull one thread, and others may unravel in ways nobody anticipated.

What is the Problem with the AI Fairness Playbook?

Most fairness approaches in machine learning work the same way: pick one sensitive attribute and optimise for it. It sounds reasonable, but the evidence tells a different story. Studies have shown that debiasing one attribute can quietly increase disparities in others, especially when those attributes are negatively correlated (Dang et al., 2025; Wang & Yang, 2025). In deep learning, the flexibility that makes neural networks powerful also makes them prone to trading fairness on one axis for unfairness on another (Cherepanova et al., 2021). One large-scale analysis found that single-attribute fairness improvements made things worse for untargeted groups in up to 88.3% of scenarios (Chen et al., 2024). And unfortunately, that is not a rare edge case, that is nearly the rule.

For large language models (LLMs), the AI behind modern chatbots and writing assistants, the problem runs even deeper. LLMs are trained on vast amounts of human text, absorbing not just facts but the full tangle of societal biases embedded in language (Due et al., 2024; Hu et al., 2025; Weidinger et al., 2021). Yet the dominant approach to making them fairer is still the same old single-attribute playbook. And bias spillover in LLMs remains critically underexplored. Most fairness evaluations report a single aggregated score, which, as we will see, can give a dangerously misleading picture.

We Already Have Some Unsettling Clues

A small but growing body of research has started to look under the hood of LLM alignment. Even when models are explicitly trained to be unbiased, they still form biased associations at a deeper, implicit level (Bai et al., 2025). One study found something particularly paradoxical: safety-aligned models were actually more likely to overlook racial concepts in ambiguous situations than their non-aligned counterparts (Sun et al., 2025), the guardrails meant to protect against bias were backfiring. Another study confirmed directly that targeted gender and race debiasing made things worse in untargeted dimensions (Chand et al., 2025). And the well-documented “alignment tax” (Ouyang et al., 2022), the general capability drop that often follows targeted alignment, suggests these costs routinely cascade beyond the dimension being optimised. What the field is lacking is a systematic, empirical look at how these spillover effects play out across multiple models and sensitive attributes at once.

An Ongoing Study Puts It to the Test

In an ongoing study we are pursuing, we have started to explore the above gap. Three LLMs are aligned to reduce gender bias using Direct Preference Optimization (DPO) (Rafailov et al., 2023). They are then evaluated on nine sensitive attributes (gender, age, disability status, nationality, physical appearance, race/ethnicity, religion, sexual orientation and socioeconomic status) using the BBQ benchmark (Parrish et al., 2022), with benchmark accuracy serving as a proxy for fairness. Crucially, in our evaluation we separate ambiguous questions (where context is insufficient and the correct answer is always “unknown”) from disambiguous ones (where a specific answer is clearly justified).

Preliminary insights suggest that aggregate, context-unaware results look positive: gender alignment improved fairness across all nine attributes in all three models. But stratify by context, and a different picture emerges! In ambiguous questions, physical appearance fairness declined significantly across all three models after gender alignment. Sexual orientation and disability status also worsened in some models. In clear-context questions, alignment consistently helped. But we have preliminary indications that bias spillover is real and statistically significant, it just hides in uncertain contexts and gets washed out by aggregate metrics.

So… Are We Fixing Bias, or Just Moving It Around?

That is the uncomfortable question that we all have to sit with. If making a model fairer on gender makes it less fair on physical appearance or age without anyone noticing because the aggregate metrics look fine, then what does it actually mean to align an AI system for fairness? Are we solving a problem, or just relocating it?

What makes this especially striking is where the spillover concentrates: in ambiguous situations, where there is no clear answer and the model has to fill in the gaps. Those are precisely the moments that matter most in real-world applications: a mental health chatbot navigating a vague distress message, a hiring tool evaluating an unconventional resume. Bias does not disappear in the easy cases. It reveals itself in the hard ones.

This points to an open and urgent challenge: we need fairness evaluation approaches that are fine-grained, context-aware and multi-attribute by default. Aggregate scores are not just incomplete: a system that looks fair at first glance may be quite unfair when you look more carefully. And the deeper question, whether multi-attribute alignment strategies (Chen et al., 2024; Wang & Yang, 2025) can actually prevent spillover rather than just redistribute it, remains very much unanswered. So, the next time an AI system is announced as “aligned for fairness”, it is worth asking: fairness for whom? Under what conditions? And at whose expense?

References

Bai, X., Wang, A., Sucholutsky, I., & Griffiths, T. L. (2025). Explicitly unbiased large language models still form biased associations. Proceedings of the National Academy of Sciences, 122(8), e2416228122. https://doi.org/10.1073/pnas.2416228122.

Chand, S., Baca, F., & Ferrara, E. (2025). No free lunch in language model bias mitigation? Targeted bias reduction can exacerbate unmitigated LLM biases. arXiv:2511.18635.

Chen, Z., Zhang, J. M., Sarro, F., & Harman, M. (2024). Fairness improvement with multiple protected attributes: How far are we? In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering (pp. 1–13). ACM.

Cherepanova, V., Nanda, V., Goldblum, M., Dickerson, J. P., & Goldstein, T. (2021). Technical challenges for training fair neural networks. arXiv:2102.06764.

Dang, V. N., Campello, V. M., Hernández-González, J., & Lekadir, K. (2025). Empirical comparison of post-processing debiasing methods for machine learning classifiers in healthcare. Journal of Healthcare Informatics Research, 1–29. https://doi.org/10.1007/s41666-025-00183-2.

Due, S., Das, S., Andersen, M., López, B. P., Nexø, S. A., & Clemmensen, L. (2024). Evaluation of large language models: STEM education and gender stereotypes. arXiv:2406.10133.

Hu, T., Kyrychenko, Y., Rathje, S., Collier, N., van der Linden, S., & Roozenbeek, J. (2025). Generative language models exhibit social identity biases. Nature Computational Science, 5(1), 65–75. https://doi.org/10.1038/s43588-024-00741-1.

Mijalli, Y., Price, P. C., & Navarro, S. P. (2023). Spillover bias in social and nonsocial judgments of diversity and variability. Psychonomic Bulletin & Review, 30(5), 1829–1839. https://doi.org/10.3758/s13423-023-02276-4.

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.

Parrish, A., Chen, A., Nangia, N., Padmakumar, V., Phang, J., Thompson, J., Htut, P. M., & Bowman, S. (2022). BBQ: A hand-built bias benchmark for question answering. In Findings of the Association for Computational Linguistics: ACL 2022 (pp. 2086–2105). ACL.

Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., & Finn, C. (2023). Direct preference optimization: Your language model is secretly a reward model. Advances in Neural Information Processing Systems, 36, 53728–53741.

Sun, L., Mao, C., Hofmann, V., & Bai, X. (2025). Aligned but blind: Alignment increases implicit bias by reducing awareness of race. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Vol. 1, pp. 22167–22184). ACL.

Wang, X., & Yang, C. C. (2025). Enhancing multi-attribute fairness in healthcare predictive modeling. arXiv:2501.13219.

Weidinger, L., Mellor, J., Rauh, M., Griffin, C., Uesato, J., Huang, P.-S., Cheng, M., Glaese, M., Balle, B., Kasirzadeh, A., Kenton, Z., Brown, S., Hawkins, W., Stepleton, T., Biles, C., Birhane, A., Haas, J., Rimell, L., Hendricks, L. A., … Gabriel, I. (2021). Ethical and social risks of harm from language models. arXiv:2112.04359.

Further Reading/Watching/Listening:

Anonymous. (2025). Bias spillover in language models: A review of political alignment, regional fragility, and multi-axis risks. Submitted to Transactions on Machine Learning Research. https://openreview.net/forum?id=hv82NTjEhs.

AI For Everybody – Preferences, Equity, Fairness and Why They Matter
https://alignai.eu/2025/07/24/ai-for-everybody-preferences-equity-fairness-and-why-they-matter/.

“Fair Enough?” – Who Wins, Who Loses and Why AI Needs to do Better Than Just Working for Most https://alignai.eu/2025/04/15/fair-enough-who-wins-who-loses-and-why-ai-needs-to-do-better-than-just-working-for-most/.

More Than Just Math – How Fairness is Being Approached in AI https://alignai.eu/2025/04/22/more-than-just-math-how-fairness-is-being-approached-in-ai/.

Image Attribution

Generated by: Nano Banana Pro 

Date: 25 February 2026

Prompt: “Create an image which describes the bias spillover effect in LLMs. The concept is that if you align a model towards one sensitive attribute (e.g., gender) it might cause negative influence on other sensitive attributes (e.g., race, physical appearance, etc.). Be creative, the image must not have text, only a visual.”

Contact Us

FIll out the form below and we will contact you as soon as possible