Current advances in Artificial Intelligence (AI) mainly come from the development and training of neural based architectures in machine learning. In particular, Pre-trained Language Models (PLMs) and, among them, Large Language Models (LLMs), Visual Language Models (VLMs) and Multi-modal LLMs have largely impacted all the fields of AI.
Pre-trained Language Models are generally deep neural architectures made of billions of parameters and require huge amounts of data to build efficient and relevant representations, either for texts or images. Furthermore, they generally need additional and annotated data to learn specific tasks (e.g. translation, question answering, cancer detection...). They are also highly demanding in terms of computational power: training and deploying such models typically requires massive GPU clusters. This greed for data and computation contrasts with many realistic scenarios: domains and languages where only small, noisy, or highly specialized datasets are available, limited number of digital documents on a specific domain that should remain confidential (e.g. document from a company, in a low resource language) or that can be annotated by just a limited number of experts (e.g. medical domain), embedded systems with tight energy or memory budgets. In addition, the current development of AI models poses important issues in terms of their environmental cost, and raises many ethical questions.
Many challenges remain: Pre-trained Language Models, and thus applications, only exist for a handful of languages – around 100 while more that 5, 000 different languages with a written system exist –, and performance drops for low-resource, specialized domains. Large language or vision models are computationally too expensive for most practitioners, either academic or professional, with issues related to data privacy, leading to an increased interest in smaller, specialized models. Both the theoretical foundations of these methods and their joint behavior in concrete truly low-resource environments (simultaneously low compute and low data) remain only partially understood despite the importance of the topic both for society and IT companies.
The goal of this trimester is to develop a deeper understanding of these models and related architectures under severe data or computational constraints and to study learning paradigms tailored to low-data regimes, with a particular emphasis on low-resource languages and specialized domains.
This semester has received the support of the International Center of Mathematics and Computer Science in Toulouse (CIMI), and the CNRS through the GT RAG of the GDR TAL.
Program
Global Trimester event link
Workshop: Frugality & explainability
October 6 to 9, Campus Toulouse Rangueil
> Event link
One-day event: Generation and use of low-resource data in the One Health field
October, 15th, École ISIS Castres
> Event link
One-day event: Content-Based Multimedia Indexing
October, 21th, Université Toulouse Capitole
> Event link
Workshop: Low-resource languages and domains
November 16 to 20, Campus Toulouse Rangueil
> Event link
One-day event: Engagement with Users and Stakeholders – The Case of Occitan
December, 2nd INSPE, UT2J, Toulouse
One-day event: Law, Health, and AI Systems: An Equation to Solve?
December, 10th, École ISIS Castres
Organizers
Josiane Mothe, IRIT - UT
José Moreno, IRIT - UT
Chloé Braud, IRIT - CNRS
Local committee
Julien Aligon, IRIT - UT
Hugo Boisaubert, IRIT - IUT Castres
Hugo Carlesso, IRIT - UT
Lotfi Chaari, IRIT
Yohann Chasseray, IRIT - Castres)
Moncef Garouani, IRIT-UTC
Philippe Muller, IRIT - UT)
Thomas Pellegrini, IRIT - UT
Antonin Poché, IRIT - IRT
Mathieu Serrurier, IRIT - UT2J
Ronan Sicre, IRIT - UT
Ludovic Tanguy, CLLE - UT2J
Rufin VanRullen, CERCO - CNRS
Marianne Vergez-Couret, CLLE - Univ. Montpellier
