Large Language Models (LLMs) that power AI assistants present a fundamental challenge: they tend to reflect the values and knowledge of culturally dominant regions, leading to biased interpretations that perpetuate prejudices and a lack of diversity. This limitation is not merely technical, it has profound implications for the development of global digital technologies.
The Challenge of Cultural Bias in AI
The proliferation of large-scale Language Models has highlighted a critical issue: the unintentional transfer of cultural biases from the Global North, particularly the United States and China, to the rest of the world. According to the United Nations Educational, Scientific and Cultural Organization (UNESCO), through its Recommendation on the Ethics of Artificial Intelligence, AI systems risk perpetuating historical inequalities and stereotypes if governance frameworks that ensure diversity and inclusion from the design stage are not implemented.
In Latin America, the consequences go beyond limiting the technical accuracy of AI—they also lead to the erasure of linguistic nuances and socioeconomic realities across the continent.
In this context, research institutions have identified that imbalances in training data produce models that tend to reflect only the linguistic and social perspectives of regions with greater digital representation. In Latin America, the impact is even greater because algorithms often fail to capture the cultural differences and language variants of the region.
Franco-Chilean collaboration for an innovative methodology
In November 2025, the LLACA initiative (LLM Assimilation of Cultural Aspects) was launched—a binational research project addressing the challenge of cultural bias in AI by developing methodologies to measure and improve the assimilation of cultural aspects in language models, with a special focus on Latin America.
The project consolidates as a high-level Franco-Chilean scientific collaboration ecosystem, integrating the strategic capabilities of researchers from Inria Chile, the Universidad de Chile, and the ALMAnaCH project-team at Inria Paris Centre in France. It was conceived under the Franco-Chilean Binational Center for Artificial Intelligence, established by Chile’s Ministry of Science, Technology, Knowledge, and Innovation and Inria at the end of 2024, and operated by Inria Chile.
LLACA’s goal is to develop and validate systematic methodologies for evaluating and integrating cultural aspects into large language models, emphasizing the representation of Latin American knowledge and values. This lays the foundation for:
- Reducing cultural biases in globally used language models;
- Improving the representation of Latin American perspectives in AI technologies;
- Developing cultural evaluation standards for AI systems;
- Promoting AI technologies that respect cultural diversity.
Lack of knowledge of Latin American cultures in large language models: the project’s initial findings
The LLACA team of scientists aims to establish reproducible standards for measuring cultural assimilation in AI systems and to develop scalable techniques for integrating cultural knowledge.
In its first months of execution, the project has achieved results in the following areas:
- Knowledge base construction: Selection of regions of interest and large-scale collection of articles through category analysis and digital encyclopedias.
- Specialized curation and filtering: Manual validation by experts and the design of an automated model.
- Controlled data generation: Creation of question-and-answer datasets using the selected articles as exclusive context.
- Integration of theoretical approaches: Experimentation with various definitions of culture (anthropological, sociological, symbolic) to guide information generation.
- Hallucination mitigation: Implementation of technical instructions and strict criteria to prevent data fabrication.
Within the LLACA framework, the team has published its first findings in the paper “Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America”, authored by Yannis Karmim (Inria, Inria Chile) as the lead author, along with researchers from Inria, Inria Chile, the Universidad de Chile, the Pontificia Universidad Católica and CENIA. The paper was presented by Valentin Barriere of the University of Chile at the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL 2026), held from March 24 to 29, 2026, in Rabat, Morocco.
The study reveals that large language models (LLMs) exhibit cultural bias, leading to a lack of understanding of Latin American cultures compared to, for example, those of Spain. To measure this gap, the authors created LatamQA, a dataset of 26,000 Wikipedia and Wikidata-based questions about various countries in the region, demonstrating that AI struggles to comprehend local Latin American contexts.
The LLACA team managed to quantify the level of knowledge of several large language models and discovered:
- A performance discrepancy among Latin American countries, with some being easier for most models than others;
- Models perform better in their original language; and
- Peninsular Hispanic culture is more widely recognized than Latin American cultures.
Toward an AI with identity
The LLACA initiative establishes itself as an international collaboration ecosystem, integrating expertise in technology, language, and social sciences. Its results will drive the development of more equitable artificial intelligence systems, laying the groundwork for improving the representation of regional perspectives in global technologies and developing cultural evaluation standards for AI systems.
Beyond immediate technical outcomes, the project seeks to establish replicable standards and a structured knowledge base that empowers underrepresented regions. By fostering technologies that respect and integrate diversity, the initiative aims to move toward innovation with its own identity, one that effectively addresses the demands of our societies.
To know more:
Yannis Karmim, Renato Pino, Hernán Contreras, Hernán Lira, Sebastián Cifuentes, et al.. Leveraging Wikidata for Geographically Informed Sociocultural Bias Dataset Creation: Application to Latin America. Workshop on Multilingual and Multicultural Evaluation (MME) of the 19th Conference of the European Chapter of the Association for Computational Linguistics (EACL), Pinzhen Chen; Vilém Zouhar; Hanxu Hu; Simran Khanuja; Wenhao Zhu; Barry Haddow; Alexandra Birch; Alham Fikri Aji; Rico Sennrich; Sara Hooker, Mar 2026, Rabat, Morocco. ⟨hal-05510068v3⟩ https://inria.hal.science/hal-05510068/document