Research Article

Self-Supervised Cross-Modal Transformers for Low-Resource Multilingual Speech Translation in Emergency Response Scenarios

23 reads
J Ong Artific Int Innov, 2026, 1 (1), 41-47, doi: , ISSN

Abstract

The urgency and linguistic diversity inherent in emergency response scenarios necessitate robust and rapid communication tools. Traditional speech translation (ST) systems, however, falter in low-resource settings due to their heavy reliance on extensive parallel corpora, a scarcity for many languages crucial in crisis situations. This paper introduces a novel approach utilizing Self-Supervised Cross-Modal Transformers (SS-CMT) designed to overcome these limitations for multilingual speech translation in emergency contexts. Our SS-CMT model leverages vast amounts of readily available monolingual speech and text data through a multi-task self-supervision framework, effectively learning rich, language-agnostic representations. By integrating speech and text encoders within a unified transformer architecture and employing contrastive learning across modalities, the system builds robust inter-modal mappings. Subsequent fine-tuning on minimal parallel data demonstrates significant performance gains over purely supervised baselines, particularly for critically low-resource languages. Experimental results on simulated emergency dialogues, augmented with environmental noise, validate the SS-CMT's superior translation quality, enhanced robustness, and reduced data dependency, paving the way for more inclusive and effective global emergency communication.

Keywords emergency response self-supervised learning Speech Translation Low-Resource Languages Cross-Modal Transformers
Authors 3

The team behind this paper

3 authors, 3 institutions.

This paper Technical University of Munich — Germany Technical University of… 1 author Kyoto University — Japan Kyoto University 1 author African Institute of Mathematical Sciences — South Africa African Institute of Ma… 1 author Prof. Elena Rostova — corresponding author ER Prof. Elena Rostova ✉ Dr. Hiroshi Tanaka HT Dr. Hiroshi Tanaka Dr. Amara Okonkwo AO Dr. Amara Okonkwo

Readership

23 reads over 1 month.

#7 most read in this journal this month
23
August 2026

Blockchain Confirmation

Loading...
If you want to upload this article to SciMatic Hybrid Blockchain, install MetaMask extension to your web browser, create a wallet and buy SCI coins at SciMatic using credit or contact your country coordinator.
One article costs 10 SCI coins to be in the Blockchain. Buy SCI Coins

Bibliographic Information

Prof. Elena Rostova, Dr. Hiroshi Tanaka, Dr. Amara Okonkwo, (2026). Self-Supervised Cross-Modal Transformers for Low-Resource Multilingual Speech Translation in Emergency Response Scenarios, Journal of Ongoing Artificial Intelligence Innovations, 1(1): 41-47
Bibtex Citation
@article{prof._elena_rostova2026joaii,
author = {Prof. Elena Rostova and Dr. Hiroshi Tanaka and Dr. Amara Okonkwo},
title = {Self-Supervised Cross-Modal Transformers for Low-Resource Multilingual Speech Translation in Emergency Response Scenarios},
journal = {Journal of Ongoing Artificial Intelligence Innovations},
year = {2026},
volume = {1},
number = {1},
pages = {41-47},
doi = {},
url = {https://scimatic.org/index.php/show_manuscript/9116}
}
APA Citation
Rostova, P.E., Tanaka, D.H., Okonkwo, D.A., (2026). Self-Supervised Cross-Modal Transformers for Low-Resource Multilingual Speech Translation in Emergency Response Scenarios. Journal of Ongoing Artificial Intelligence Innovations, 1(1), 41-47. https://doi.org/

Author Information

  • To change your profile photo, login to scimatic.org, go to your profile and change the photo.
  • Provide a face photo, and not full body.
  • It is better to remove the background from your photo. Go to Remove Background and then upload to profile
  • If you are unable to login, go to Reset My Password provide your email registered with the article and get new password.
  • In case of any other problem, contact your editor directly or write to us at info @ scimatic.org