Abstract
The urgency and linguistic diversity inherent in emergency response scenarios necessitate robust and rapid communication tools. Traditional speech translation (ST) systems, however, falter in low-resource settings due to their heavy reliance on extensive parallel corpora, a scarcity for many languages crucial in crisis situations. This paper introduces a novel approach utilizing Self-Supervised Cross-Modal Transformers (SS-CMT) designed to overcome these limitations for multilingual speech translation in emergency contexts. Our SS-CMT model leverages vast amounts of readily available monolingual speech and text data through a multi-task self-supervision framework, effectively learning rich, language-agnostic representations. By integrating speech and text encoders within a unified transformer architecture and employing contrastive learning across modalities, the system builds robust inter-modal mappings. Subsequent fine-tuning on minimal parallel data demonstrates significant performance gains over purely supervised baselines, particularly for critically low-resource languages. Experimental results on simulated emergency dialogues, augmented with environmental noise, validate the SS-CMT's superior translation quality, enhanced robustness, and reduced data dependency, paving the way for more inclusive and effective global emergency communication.