This project addresses an urgent scientific and social need: the lack of massive and open corpora that allow for the objective analysis of language use and information dissemination on social media. At the academic level, the availability of a corpus of this nature will allow progress on key sociolinguistic issues, such as the study of varieties, languages in contact, multilingualism and the relationship between standard and vernacular languages, as well as in communication technologies and cultural geography. At the social level, the existence of verifiable and accessible data will be crucial for analyzing the formation of digital communities and the differences in their speech patterns.
The objective is twofold: (1) to construct a global corpus of word frequencies from billions of geolocated and anonymized tweets; (2) to apply this resource to specific applications, such as cultural mapping, the coexistence of languages and the quantification of linguistic variation processes. Our methodology is based on artificial intelligence techniques and complex systems (clusters, PCA, network theory, statistical models) applied to lexical counts enriched with linguistic, temporal and geographical metadata. The main deliverable will be an open-access database with a query interface, accompanied by scientific publications describing both the corpus and its applications. The expected impact is both scientific and social. On the one hand, it will provide a resource of unprecedented size and spatial resolution, guaranteeing reproducibility and reliability of results. On the other hand, it will offer new tools for understanding digital culture and the dynamics of languages in the online public sphere. Exploitation of geolocation and resident verification will transform this corpus into a unique infrastructure with great potential for interdisciplinary research and for addressing the global challenges of disinformation in a context of cultural diversity.