Skip to main content

Showing 1–7 of 7 results for author: Federico, M

Searching in archive eess. Search in all archives.
.
  1. arXiv:2404.07336  [pdf, other

    cs.CV cs.MM eess.AS

    PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores

    Authors: Lucas Goncalves, Prashant Mathur, Chandrashekhar Lavania, Metehan Cekic, Marcello Federico, Kyu J. Han

    Abstract: Recent advancements in audio-visual generative modeling have been propelled by progress in deep learning and the availability of data-rich benchmarks. However, the growth is not attributed solely to models and benchmarks. Universally accepted evaluation metrics also play an important role in advancing the field. While there are many metrics available to evaluate audio and visual content separately… ▽ More

    Submitted 10 April, 2024; originally announced April 2024.

    Comments: 24 pages

  2. arXiv:2311.00697  [pdf, other

    cs.CL eess.AS

    End-to-End Single-Channel Speaker-Turn Aware Conversational Speech Translation

    Authors: Juan Zuluaga-Gomez, Zhaocheng Huang, Xing Niu, Rohit Paturi, Sundararajan Srinivasan, Prashant Mathur, Brian Thompson, Marcello Federico

    Abstract: Conventional speech-to-text translation (ST) systems are trained on single-speaker utterances, and they may not generalize to real-life scenarios where the audio contains conversations by multiple speakers. In this paper, we tackle single-channel multi-speaker conversational ST with an end-to-end and multi-task training model, named Speaker-Turn Aware Conversational Speech Translation, that combin… ▽ More

    Submitted 1 November, 2023; originally announced November 2023.

    Comments: Accepted at EMNLP 2023. Code: https://github.com/amazon-science/stac-speech-translation

  3. arXiv:2305.13204  [pdf, other

    cs.CL cs.SD eess.AS

    Improving Isochronous Machine Translation with Target Factors and Auxiliary Counters

    Authors: Proyag Pal, Brian Thompson, Yogesh Virkar, Prashant Mathur, Alexandra Chronopoulou, Marcello Federico

    Abstract: To translate speech for automatic dubbing, machine translation needs to be isochronous, i.e. translated speech needs to be aligned with the source in terms of speech durations. We introduce target factors in a transformer model to predict durations jointly with target language phoneme sequences. We also introduce auxiliary counters to help the decoder to keep track of the timing information while… ▽ More

    Submitted 22 May, 2023; originally announced May 2023.

    Comments: Accepted at INTERSPEECH 2023

  4. arXiv:2302.12979  [pdf, other

    cs.CL cs.SD eess.AS

    Jointly Optimizing Translations and Speech Timing to Improve Isochrony in Automatic Dubbing

    Authors: Alexandra Chronopoulou, Brian Thompson, Prashant Mathur, Yogesh Virkar, Surafel M. Lakew, Marcello Federico

    Abstract: Automatic dubbing (AD) is the task of translating the original speech in a video into target language speech. The new target language speech should satisfy isochrony; that is, the new speech should be time aligned with the original video, including mouth movements, pauses, hand gestures, etc. In this paper, we propose training a model that directly optimizes both the translation as well as the spe… ▽ More

    Submitted 24 February, 2023; originally announced February 2023.

    Comments: 5 pages

  5. arXiv:2204.02530  [pdf, other

    cs.CL cs.LG cs.SD eess.AS

    Prosodic Alignment for off-screen automatic dubbing

    Authors: Yogesh Virkar, Marcello Federico, Robert Enyedi, Roberto Barra-Chicote

    Abstract: The goal of automatic dubbing is to perform speech-to-speech translation while achieving audiovisual coherence. This entails isochrony, i.e., translating the original speech by also matching its prosodic structure into phrases and pauses, especially when the speaker's mouth is visible. In previous work, we introduced a prosodic alignment model to address isochrone or on-screen dubbing. In this wor… ▽ More

    Submitted 5 April, 2022; originally announced April 2022.

    Comments: 5 pages, 2 figures, 3 tables, Submitted to Interspeech 2022

  6. arXiv:2110.03847  [pdf, other

    cs.CL cs.SD eess.AS

    Machine Translation Verbosity Control for Automatic Dubbing

    Authors: Surafel M. Lakew, Marcello Federico, Yue Wang, Cuong Hoang, Yogesh Virkar, Roberto Barra-Chicote, Robert Enyedi

    Abstract: Automatic dubbing aims at seamlessly replacing the speech in a video document with synthetic speech in a different language. The task implies many challenges, one of which is generating translations that not only convey the original content, but also match the duration of the corresponding utterances. In this paper, we focus on the problem of controlling the verbosity of machine translation output… ▽ More

    Submitted 7 October, 2021; originally announced October 2021.

    Comments: Accepted at IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 2021

  7. arXiv:2001.06785  [pdf, other

    cs.CL cs.SD eess.AS

    From Speech-to-Speech Translation to Automatic Dubbing

    Authors: Marcello Federico, Robert Enyedi, Roberto Barra-Chicote, Ritwik Giri, Umut Isik, Arvindh Krishnaswamy, Hassan Sawaf

    Abstract: We present enhancements to a speech-to-speech translation pipeline in order to perform automatic dubbing. Our architecture features neural machine translation generating output of preferred length, prosodic alignment of the translation with the original speech segments, neural text-to-speech with fine tuning of the duration of each utterance, and, finally, audio rendering to enriches text-to-speec… ▽ More

    Submitted 2 February, 2020; v1 submitted 19 January, 2020; originally announced January 2020.

    Comments: 5 pages, 4 figures