AI
Ethics

Flemish language models: where innovation meets conscience

Vlaamse taalmodellen

As language models evolve in speed, intelligence and scale, a joint event hosted by VRT, VAIA and the Flanders AI Research Program emphasized a key nuance: AI’s future is shaped not only by size, but by empathy, context and responsibility. 

Under the title "Flemish Language Models: Big and Small", academics, media makers and technology experts gathered at VRT for an afternoon of reflection, discussion and insight. While the core focus was on technical innovation, the standout topics of the day were ethics, usability, trust and empathy in an AI-driven future. 

What are language models? 

Language models essentially predict the next token (i.e. words, punctuation marks, ...) in a sentence. This so-called next token ‘prediction principle’, enriched with clever algorithms, allows language models to generate meaningful and useful responses. Building such a model involves two main phases: pre-training and post-training. 

Pieter Delobelle - Vlaamse (L)LMs: Vlaamse taalmodellen, groot en klein - 2025
Pieter Delobelle - Vlaamse (L)LMs: Vlaamse taalmodellen, groot en klein - 2025
Pieter Delobelle - Vlaamse (L)LMs: Vlaamse taalmodellen, groot en klein - 2025

In the pre-training phase, a model is trained on large amounts of text from the Internet. It learns to recognize statistical patterns and meaningful relationships purely through self-supervised learning, without manually labelled data. Instead, it discovers which words and phrases frequently appear together and how they relate to each other. The goal of this phase is to give the model a general understanding of language so it can perform a wide range of tasks flexibly. 

This phase is followed by post-training, where the model is fine-tuned using techniques such as fine-tuning or Reinforcement Learning from Human Feedback (RLHF). These make the model better at specific tasks such as translation, summarization or question answering.  

Another technique, in-context learning, allows the model to learn new tasks from examples in the prompt without adjusting the underlying parameters. 

Open or closed models? 

Language models come in both open and commercial versions. However, what 'open' means is not straightforward: although open models are often available for free, it is not always clear what data they were trained on. These 'open' models pose many risks to users, such as unknowingly using biased or one-sided ethical viewpoint models. Additionally, users are unaware of the model's capabilities and limitations because the training data is unknown. 

Language and culture: Flemish models 

Large language models work well in Dutch because they include Dutch content such as VRT NWS articles. However, they often run into cultural barriers: typical Flemish proverbs, cultural heritage or questions about local culture. They were not sufficiently trained on this due to limited publicly available data online. 

Pieter Delobelle - Vlaamse (L)LMs: Vlaamse taalmodellen, groot en klein - 2025

However, there is currently a lack of good benchmarks for evaluating Flemishlanguage models (and its knowledge of Dutch heritage, culture, ...). This absence makes it difficult to systematically compare models or track their progress over time. Benchmarks are essential because they provide a common framework for assessing performance on key tasks such as information extraction, grammar, and logical reasoning. Without them, it's nearly impossible to identify specific strengths and weaknesses or to measure improvements in a reliable way. There are differences in how well models perform in tasks such as information extraction, grammar or logical reasoning. Therefore, developing dedicated, high-quality evaluation benchmarks tailored to Flemish is a critical step toward advancing the performance and reliability of Flemish language models in this context. 

Empathy and personal AI 

Language models produce speech, but human communication involves much more: intonation, body language and facial expressions. Empathy is central to trust, collaboration and conflict resolution. In human-machine interactions, empathy is the difference between a dry chatbot and a meaningful conversation. 

But empathic language models are still in their infancy. Again, the bottleneck is data: there are hardly any high-quality datasets in which emotions are accurately labeled or that contain authentic conversations. Labeling emotional data is particularly challenging because emotion is subjective and context dependent. What one person interprets as sadness, another might see as frustration. Moreover, emotional expression varies greatly across cultures, making consistent labeling even more difficult. 

Can language models think? 

The first language models struggled with reasoning. 'Regular LLMs' see a question and immediately jump to providing an answer, simply predicting the next word without verifying the logical consistency of their responses. The breakthrough came with the introduction of Chain-of-Thought (CoT), where a model internally processes the problem step by step, exploring different perspectives and considering evidence before delivering its final response. This allows it to detect errors, test alternatives and break down complex problems into smaller parts. 

Vincent Ginis - Hoe nieuwe schalingswetten het redeneren van grote taalmodellen vormgeven - 2025

These types of reasoning models (such as OpenAI's o-series) perform particularly well in tasks such as maths or coding which require logic thinking.  

Bias and recourse  

An incorrect decision by an AI system is difficult to detect and even harder to  correct. This is because AI systems often make decisions based on complex algorithms that are not transparent. When something goes wrong, it’s not immediately clear why the system made a particular choice. Additionally, human intervention or correction is often difficult, as the model may be highly confident in its decision, even when it’s incorrect. Human intervention is often overshadowed by the model. Bias testing should focus on what you really want to measure, not just the benchmark used. It's important to focus on what you really want to measure, rather than relying solely on traditional metrics that may overlook these critical factors. 

An additional problem with biases in language models is that even when a person is 'in the loop', they often unconsciously follow the model, especially in cases of doubt or 'grey areas'. This means that errors or biases in the model are not actively corrected but simply reinforced. 

How VRT is betting on it 

VRT sees AI as the third major wave of innovation in media - after digital TV and social media. Generative AI has the potential to fundamentally change the way media is consumed and produced. Think of AI assistants that communicate with users in natural language, and dynamic media that adapts to the user in real time. 

At the same time, AI also poses a threat: users are reading news via summaries from tools like ChatGPT, rather than clicking through to the original articles. This undermines the visibility and revenue model of Flemish media companies. In addition, recent BBC research shows that AI summaries of news are often inaccurate, damaging trust in the media.  

That's why VRT is working on AI strategies that can be used in the short and long term. At this moment, this means integrating AI into existing workflows. For example, the Smart News Assistant allows news articles to be automatically reformatted for different channels. Experiments are also being conducted with conversational AI, mainly for professional use, with a view to future applications for end users. 

An LLM tailored to Flemish language and media use can make a difference: deeper knowledge of language, culture and context results in authentic content and better user interaction. VRT is actively involved in research projects such as Narrate (AI storytelling), Solid4Media (personal and secure AI agents) and Elliot (multimodal AI combining text, audio and video).

Flemish and European initiatives 

The AI revolution is moving at lightning speed. The developments around and the impact of AI are visible throughout society and are becoming more entrenched. Flanders can't afford to miss this technological (high-speed) train. That is why the Flanders Artificial Intelligence (AI) Policy Plan aims to expand the existing AI knowledge base, grow the AI-expertise in Flemish industry and foster the rollout of AI in Flanders. The Flemish Policy Plan AI invests 36 million euros per year in research, implementation and digitisation. 

Additionally, several initiatives have been launched at the European level to support non-English language projects, addressing the lack of data and improving cultural and non-English language models in Europe. For example, the European Language Data Space (LDS) is creating a European marketplace for multilingual and multimodal language data. ALT-EDIC aims to build a European language technology infrastructure with a focus on cultural diversity and domain-specific applications. 

Article written by Fleur Maselis, with the input from Luk Overmeire (VRT) and Kevin D'hooghe (IMEC)