Harmony through alignment: unifying multimodal representations of sound and text across languages
Carregando...
Data
Autores
Título da Revista
ISSN da Revista
Título de Volume
Editor
Universidade Federal de Goiás
Resumo
The alignment of data modalities has achieved significant success, particularly in the text-image domain. However, extending these principles to create robust, multilingual representations of language and audio introduces distinct challenges that remain partially unaddressed in the literature. These include the scarcity of linguistically diverse text-audio datasets and the prohibitively high computing cost of training large-scale models. To address the high computational and data costs
associated with multimodal expansion, this dissertation first introduces CACARA, a framework designed to integrate new modalities efficiently. CACARA employs an emergent alignment learning strategy. In our case, a new modality, audio, is aligned with a pre-trained and frozen multilingual image-text model. The model inherits the text encoder's rich multilingual capabilities by training our framework using only English data. This approach enables zero-shot cross-lingual and cross-modal retrieval, reducing training time and energy consumption. As a complementary contribution, this research presents PALMA (Pre-trained Audio-Language Multilingual Alignment via Mixture-of-Layers), a framework for the systematic design and optimization of bimodal language-audio models. PALMA goes beyond conventional methods by comprehensively evaluating state-of-the-art (SOTA) encoders and advanced training paradigms. Its central innovation is the “Mixture-of-Layers” (MoL) method. This learnable aggregation strategy fuses features from multiple encoder layers to create a richer, more nuanced audio representation. Empirical evaluations validate the proposed approaches, which achieve competitive results. Our findings give strong evidence that emergent alignment is a viable, scalable, and low-cost multimodal expansion strategy. This work contributes to the development of more accessible, efficient, and linguistically inclusive multimodal systems.
Descrição
Citação
FERREIRA, Alef Iury Siqueira. Harmony through alignment: unifying multimodal representations of sound and text across languages. 2026. 239 f. Dissertação (Mestrado em Ciência da Computação) - Instituto de Informática (INF)Universidade Federal de Goiás, Goiânia, 2026.