Vivek Yadav,
Utkarsh,
Sudhakar Verma,
- Student, Department of Computer Science and Engineering (Data Science), ABES Engineering College, Uttar Pradesh, India
- Student, Department of Computer Science and Engineering (Data Science), ABES Engineering College, Uttar Pradesh, India
- Student, Department of Computer Science and Engineering (Data Science), ABES Engineering College, Uttar Pradesh, India
Abstract
Different languages have been proved a great obstacle to global communication despite the internet’s role in allowing information sharing all over the world. While presenting an extensive number of current solutions, traditional Machine Translation (MT) systems are unable to convey complex contextual information and dialects including “Hinglish”. Above all, the voice of the interlocutor is lost and is replaced with an artificial one, programmed to mimic the voice of the interlocutor, creating a good distance between interlocutors in terms of personality. However, these multimodal limitations have been tackled through the introduction of a novel Generative AI-based system for real-time identity-preserving speech-to-speech (S2S) translation, the “Speech-Text-Speech Translator. These multimodal limitations, however, have been addressed here by introducing a novel Generative AI-based system for real-time identity-preserving speech-to-speech (S2S) translation, the ‘Speech-Text-Speech Translator’. Its extensive support of major Indic languages includes Hindi, Bengali, Urdu, and Punjabi; it is based on a three-stage cascaded pipeline and a Python Flask back-end. To begin with, the initial spoken prompt is correctly converted to textual content by an Automatic Speech Recognition (ASR) module. Second, it’s a Large Language Model (the Google Gemini API) that is doing semantic translation with context. The strict prompt engineering eliminates the LLM from producing conversational filler in order to assure stability in the computational pipeline. Last but not least, a Zero-Shot Text-to-Speech (ZS-TTS) module is a quick analysis of a short text embedding of the source audio can instantly translate and synthesize the translated text in the speaker’s natural voice timbre. The system is implemented as a serverless client-server system, with asynchronous processing to tackle cumulative API latency effectively. Considering critical ethical and privacy issues surrounding voice spoofing, the framework is completely stateless with audio profiles being only kept in the volatile temporary memory. This new study pushes beyond merely translating text to producing a more natural and authentic dialogue using state-of-the-art ASR, LLM, and ZS-TTS tools. Quantitative evaluations demonstrate a high translation accuracy (BLEU = 78.19) and strong speaker-identity preservation (Cosine Similarity = 0.87), alongside a measured end-to-end processing latency of approximately 30 seconds. This architecture is a secure base for future applications such as emotion and prosody preservation, real-time streaming and video cross-lingual dubbing.
Keywords: Generative AI, Speech-to-Speech (S2S), Voice Cloning, Zero-Shot Text-to-Speech (ZS-TTS), Large Language Model (LLM), Context-Aware Translation, Flask, Google Gemini, Indic Languages
[This article belongs to Journal of Mechatronics and Automation ]
References
- Team G, Anil R, Borgeaud S, Alayrac JB, Yu J, Soricut R, Schalkwyk J, Dai AM, Hauth A, Millican K, Silver D. Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:2312.11805. 2023 Dec 19.
- Kamenskaya A.S. Adaptation of Google Cloud Speech-to-text API for automatic transcription of web conferences in real time. Automation and Software Engineering. 2019(2 (28)):19–23.
- Zheng Z, Peng P, Diwan A, Huynh CP, Sun X, Liu Z, Bhat V, Harwath D. Voicecraft-x: Unifying multilingual, voice-cloning speech synthesis and speech editing. InProceedings of the 2025 Conference on Empirical Methods in Natural Language Processing 2025 Nov (pp. 2737-2756).
- Ashish V. Attention is all you need. Advances in neural information processing systems. 2017.
- Devlin J, Chang MW, Lee K, Toutanova K. Bert: Pre-training of deep bidirectional transformers for language understanding. InProceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers) 2019 Jun (pp. 4171-4186).
- Radford A, Kim JW, Xu T, Brockman G, McLeavey C, Sutskever I. Robust speech recognition via large-scale weak supervision. InInternational conference on machine learning 2023 Jul 3 (pp. 28492-28518). PMLR.
- Baevski A, Zhou Y, Mohamed A, Auli M. wav2vec 2.0: A framework for self-supervised learning of speech representations. Advances in neural information processing systems. 2020;33:12449-60.
- Jia Y, Ramanovich MT, Remez T, Pomerantz R. Translatotron 2: High-quality direct speech-to-speech translation with voice preservation. InInternational conference on machine learning 2022 Jun 28 (pp. 10120-10134). PMLR.
- Sperber M, Paulik M. Speech translation and the end-to-end promise: Taking stock of where we are. InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics 2020 Jul (pp. 7409-7421).
- Lakhotia K, Kharitonov E, Hsu WN, Adi Y, Polyak A, Bolte B, Nguyen TA, Copet J, Baevski A, Mohamed A, Dupoux E. On generative spoken language modeling from raw audio. Transactions of the Association for Computational Linguistics. 2021 Dec 6;9:1336-54.
- Shen J, Pang R, Weiss RJ, Schuster M, Jaitly N, Yang Z, Chen Z, Zhang Y, Wang Y, Skerrv-Ryan R, Saurous RA. Natural tts synthesis by conditioning wavenet on mel spectrogram predictions. In2018 IEEE international conference on acoustics, speech and signal processing (ICASSP) 2018 Apr 15 (pp. 4779-4783). IEEE.
- Kim J, Kong J, Son J. Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech. InInternational conference on machine learning 2021 Jul 1 (pp. 5530-5540). PMLR..
- Wang C, Chen S, Wu Y, Zhang Z, Zhou L, Liu S, Chen Z, Liu Y, Wang H, Li J, He L. Neural codec language models are zero-shot text to speech synthesizers. arXiv preprint arXiv:2301.02111. 2023 Jan 5.
- Arık SÖ, Chrzanowski M, Coates A, Diamos G, Gibiansky A, Kang Y, Li X, Miller J, Ng A, Raiman J, Sengupta S. Deep voice: Real-time neural text-to-speech. InInternational conference on machine learning 2017 Jul 17 (pp. 195-204). PMLR.
- Bezzaoucha I. Generative AI or the Doom of Translation as we Know it?. Cahiers de Traduction. 2024 Jun 2;30(1):07-15.

Journal of Mechatronics and Automation
| Volume | 13 | |
| Issue | 02 | |
| Received | 26/06/2026 | |
| Accepted | 27/08/2026 | |
| Published | 31/08/2026 | |
| Publication Time | 66 Days |