Safety Guardrails for AI: How LLMs Learn to Stay Safe

Large language models (LLMs) are trained on large amounts of text from the internet, books, forums and other sources in a process called pre-training. This gives them great versatility, but also comes with a hidden challenge: human language data contains biases, misinformation and unsafe patterns, such as hate speech, toxic or discriminatory content. When models learn from such data, they not only gain useful knowledge but also inherit these problems. On top of this, LLMs tend to be statistically overconfident (Guo et al., 2017; Minderer et al., 2021), meaning they assign higher probabilities to their predictions, due to the way that they interpret data (Xu et al., 2024). They often present information with certainty, even when the output is false. This combination of biased training data and overconfidence can lead to hallucinations, biased answers or unsafe outputs, such as toxic content or instructions for harmful behavior.
How is AI Changing the Creative Process? AI as the Co-creator Nowadays

Creativity is often considered as an “intuition” or “talent” and can’t be easily interpreted in a logical way (Wu et al. 2021). The creative industries often refer to graphic design, film, music, video games, fashion, advertising, media or entertainment industries (Howkins 2002), related to the extraordinary thinking by supreme creative individuals (Weisberg 2006). However, creativity actually lies in all creative activities, from the arts to science, from everyday life to industry production. Today, creativity is considered to be a crucial competency (Binkley et al. 2012). Boden (2004), who pioneered the field of philosophy of cognitive science, offers the definition “Creativity is the ability to come up with ideas or artefacts that are new, surprising and valuable”. With the help of language, people used the creative process in art and technology, making creativity “one of the most striking features of the human species”, since at least 40,000 years ago (Carruthers 2002, p. 226). Creativity in today’s sense is at the heart of human endeavour, shaping various fields including education, art and healthcare (Esling and Devis 2020; Farina et al. 2024; Tredinnick and Laybats 2023).
Q&A with DC Tuan-Ting Huang

What inspired you to join the alignAI project? Coming from a graphic design background, my master’s studies in interaction design opened my eyes to the possibilities of human-machine collaboration and to what computers can bring to the creative field in terms of serendipity, iteration and data processing. But alongside the excitement about these technological potentials […]
Meta and Mind: Tracing the Journey of Thinking about Thinking

For as long as we have written history, humans have been fascinated by the idea of thinking about thinking. The ancient Greeks saw self-reflection as a path to wisdom: Socrates urged his students to “know thyself”, while Aristotle suggested that the mind could even grasp its own activity. Centuries later, philosophers and logicians took this further, asking whether knowing something also means knowing that you know it. In the 1960s, Jaakko Hintikka captured this in a famous principle of logic: if an agent knows a fact, it should also know that it knows it. Fast forward to today, and this same idea has found new life in artificial intelligence, where researchers explore how machines might be designed not just to think, but to reflect on their own thinking.
Creativity, Style and the Flattening Threat in Large Language Models

The debate on creativity has intensified with the rise of generative AI, especially of large language models (LLMs). Recent research shows that these systems can produce work that competes with, and in some cases exceeds, human creativity (Guzik et al., 2023; Bohren et al., 2024). At the same time, their use brings serious concerns about value, authenticity, and the long-term safeguarding of human creative practices (Mei et al., 2025; Messer, 2024). This tension highlights what might be called the “flattening threat”: there is a perceived risk that even as LLMs make it easier to generate ideas and boost productivity, they could also diminish the diversity, style and authenticity that enrich human creativity.
Q&A with DC Sharvari Bondre

What inspired you to join the alignAI project? My master’s in information science exposed me to the impact of information systems on society and ethical responsibilities of those who design them. The technology we design not only mirrors but amplifies societal biases, both positive and negative. AlignAI’s focus on harmonising AI with human values is […]
Thinking about the Ethical Use of AI in the Military: Implications for Organizations and Global Security

On 24 September 2025, the TUM Institute for Ethics in Artificial Intelligence (IEAI) hosted a panel discussion at the TUM Think Tank titled “Thinking about Ethical Use of AI in the Military – Implications for Organizations and Global Security”. The session, moderated by IEAI Executive Director and alignAI Project Lead Dr. Caitlin Corrigan, featured Brigadier General (Ret.) Dr. David Barnes (Empowering AI) and Lance Lindauer (Partnership to Advance Responsible Technology).
The Myth of Neutral Participation: Why Good Intentions Aren’t Enough in AI Design

The field of AI is experiencing a participatory turn (Delgado et al., 2023). From tech companies to researchers, there is growing recognition that AI design and development should not happen in isolation from the people it affects. Regardless whether AI systems are designed for mental health, education or journalism, they need input from communities who deeply understand these domains. Interdisciplinary collaboration has been assuming a more pivotal role, bringing together computer scientists, researchers, ethicists and community members to create more aligned and responsible AI systems. This shift is certainly representative of progress.
Q&A with DC Katerina Drakos

What inspired you to join the alignAI project? During the last few years, AI has transformed society. Curiously enough, I was raised by a family of engineers who would repeatedly discuss the potential of artificial intelligence. Back then I dismissed it as futuristic, fictional and possibly even Orwellian. I forged my own path in medicine, […]
Mind the XAI Gap: A Human-centred LLM Framework for Democratising Explainable AI

Our alignAI doctoral candidate Eva Paraschou presented her work “Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI” at the “The 3rd World Conference on eXplainable Artificial Intelligence” held in Istanbul, Turkey (9-11 July 2025).