Skip to main content

Showing 1–11 of 11 results for author: McGraw, I

Searching in archive eess. Search in all archives.
.
  1. arXiv:2303.08343  [pdf, ps, other

    eess.AS cs.AI cs.LG cs.SD

    Sharing Low Rank Conformer Weights for Tiny Always-On Ambient Speech Recognition Models

    Authors: Steven M. Hernandez, Ding Zhao, Shao** Ding, Antoine Bruguier, Rohit Prabhavalkar, Tara N. Sainath, Yanzhang He, Ian McGraw

    Abstract: Continued improvements in machine learning techniques offer exciting new opportunities through the use of larger models and larger training datasets. However, there is a growing need to offer these new capabilities on-board low-powered devices such as smartphones, wearables and other embedded environments where only low memory is available. Towards this, we consider methods to reduce the model siz… ▽ More

    Submitted 14 March, 2023; originally announced March 2023.

    Comments: Accepted to IEEE ICASSP 2023

  2. arXiv:2204.06164  [pdf, other

    eess.AS cs.LG cs.SD

    A Unified Cascaded Encoder ASR Model for Dynamic Model Sizes

    Authors: Shao** Ding, Weiran Wang, Ding Zhao, Tara N. Sainath, Yanzhang He, Robert David, Rami Botros, Xin Wang, Rina Panigrahy, Qiao Liang, Dongseong Hwang, Ian McGraw, Rohit Prabhavalkar, Trevor Strohman

    Abstract: In this paper, we propose a dynamic cascaded encoder Automatic Speech Recognition (ASR) model, which unifies models for different deployment scenarios. Moreover, the model can significantly reduce model size and power consumption without loss of quality. Namely, with the dynamic cascaded encoder model, we explore three techniques to maximally boost the performance of each model size: 1) Use separa… ▽ More

    Submitted 24 June, 2022; v1 submitted 13 April, 2022; originally announced April 2022.

    Comments: Accepted by INTERSPEECH 2022

  3. arXiv:2204.03793  [pdf, other

    eess.AS cs.LG cs.SD

    Personal VAD 2.0: Optimizing Personal Voice Activity Detection for On-Device Speech Recognition

    Authors: Shao** Ding, Rajeev Rikhye, Qiao Liang, Yanzhang He, Quan Wang, Arun Narayanan, Tom O'Malley, Ian McGraw

    Abstract: Personalization of on-device speech recognition (ASR) has seen explosive growth in recent years, largely due to the increasing popularity of personal assistant features on mobile devices and smart home speakers. In this work, we present Personal VAD 2.0, a personalized voice activity detector that detects the voice activity of a target speaker, as part of a streaming on-device ASR system. Although… ▽ More

    Submitted 24 June, 2022; v1 submitted 7 April, 2022; originally announced April 2022.

    Comments: Accepted by INTERSPEECH 2022

  4. arXiv:2202.12169  [pdf, other

    eess.AS cs.LG stat.ML

    Closing the Gap between Single-User and Multi-User VoiceFilter-Lite

    Authors: Rajeev Rikhye, Quan Wang, Qiao Liang, Yanzhang He, Ian McGraw

    Abstract: VoiceFilter-Lite is a speaker-conditioned voice separation model that plays a crucial role in improving speech recognition and speaker verification by suppressing overlap** speech from non-target speakers. However, one limitation of VoiceFilter-Lite, and other speaker-conditioned speech models in general, is that these models are usually limited to a single target speaker. This is undesirable as… ▽ More

    Submitted 26 April, 2022; v1 submitted 24 February, 2022; originally announced February 2022.

  5. arXiv:2107.01201  [pdf, other

    eess.AS cs.LG cs.SD

    Multi-user VoiceFilter-Lite via Attentive Speaker Embedding

    Authors: Rajeev Rikhye, Quan Wang, Qiao Liang, Yanzhang He, Ian McGraw

    Abstract: In this paper, we propose a solution to allow speaker conditioned speech models, such as VoiceFilter-Lite, to support an arbitrary number of enrolled users in a single pass. This is achieved by using an attention mechanism on multiple speaker embeddings to compute a single attentive embedding, which is then used as a side input to the model. We implemented multi-user VoiceFilter-Lite and evaluated… ▽ More

    Submitted 8 November, 2021; v1 submitted 2 July, 2021; originally announced July 2021.

  6. arXiv:2104.13970  [pdf, other

    eess.AS cs.LG cs.SD

    Personalized Keyphrase Detection using Speaker and Environment Information

    Authors: Rajeev Rikhye, Quan Wang, Qiao Liang, Yanzhang He, Ding Zhao, Yiteng, Huang, Arun Narayanan, Ian McGraw

    Abstract: In this paper, we introduce a streaming keyphrase detection system that can be easily customized to accurately detect any phrase composed of words from a large vocabulary. The system is implemented with an end-to-end trained automatic speech recognition (ASR) model and a text-independent speaker verification model. To address the challenge of detecting these keyphrases under various noisy conditio… ▽ More

    Submitted 15 June, 2021; v1 submitted 28 April, 2021; originally announced April 2021.

  7. arXiv:2104.12870  [pdf, other

    eess.AS cs.CL cs.LG cs.SD

    Multi-Task Learning for End-to-End ASR Word and Utterance Confidence with Deletion Prediction

    Authors: David Qiu, Yanzhang He, Qiujia Li, Yu Zhang, Liangliang Cao, Ian McGraw

    Abstract: Confidence scores are very useful for downstream applications of automatic speech recognition (ASR) systems. Recent works have proposed using neural networks to learn word or utterance confidence scores for end-to-end ASR. In those studies, word confidence by itself does not model deletions, and utterance confidence does not take advantage of word-level training signals. This paper proposes to joi… ▽ More

    Submitted 26 April, 2021; originally announced April 2021.

    Comments: Submitted to Interspeech 2021

  8. arXiv:2103.06716  [pdf, other

    eess.AS cs.CL cs.LG

    Learning Word-Level Confidence For Subword End-to-End ASR

    Authors: David Qiu, Qiujia Li, Yanzhang He, Yu Zhang, Bo Li, Liangliang Cao, Rohit Prabhavalkar, Deepti Bhatia, Wei Li, Ke Hu, Tara N. Sainath, Ian McGraw

    Abstract: We study the problem of word-level confidence estimation in subword-based end-to-end (E2E) models for automatic speech recognition (ASR). Although prior works have proposed training auxiliary confidence models for ASR systems, they do not extend naturally to systems that operate on word-pieces (WP) as their vocabulary. In particular, ground truth WP correctness labels are needed for training confi… ▽ More

    Submitted 11 March, 2021; originally announced March 2021.

    Comments: To appear in ICASSP 2021

  9. arXiv:2006.01416  [pdf, other

    cs.CL cs.SD eess.AS

    Analyzing the Quality and Stability of a Streaming End-to-End On-Device Speech Recognizer

    Authors: Yuan Shangguan, Kate Knister, Yanzhang He, Ian McGraw, Francoise Beaufays

    Abstract: The demand for fast and accurate incremental speech recognition increases as the applications of automatic speech recognition (ASR) proliferate. Incremental speech recognizers output chunks of partially recognized words while the user is still talking. Partial results can be revised before the ASR finalizes its hypothesis, causing instability issues. We analyze the quality and stability of on-devi… ▽ More

    Submitted 14 August, 2020; v1 submitted 2 June, 2020; originally announced June 2020.

    Comments: Accepted at Interspeech 2020

  10. arXiv:1909.12408  [pdf, other

    cs.CL cs.LG eess.AS

    Optimizing Speech Recognition For The Edge

    Authors: Yuan Shangguan, Jian Li, Qiao Liang, Raziel Alvarez, Ian McGraw

    Abstract: While most deployed speech recognition systems today still run on servers, we are in the midst of a transition towards deployments on edge devices. This leap to the edge is powered by the progression from traditional speech recognition pipelines to end-to-end (E2E) neural architectures, and the parallel development of more efficient neural network topologies and optimization techniques. Thus, we a… ▽ More

    Submitted 6 February, 2020; v1 submitted 26 September, 2019; originally announced September 2019.

  11. arXiv:1908.10992  [pdf, other

    cs.CL cs.SD eess.AS

    Two-Pass End-to-End Speech Recognition

    Authors: Tara N. Sainath, Ruoming Pang, David Rybach, Yanzhang He, Rohit Prabhavalkar, Wei Li, Mirkó Visontai, Qiao Liang, Trevor Strohman, Yonghui Wu, Ian McGraw, Chung-Cheng Chiu

    Abstract: The requirements for many applications of state-of-the-art speech recognition systems include not only low word error rate (WER) but also low latency. Specifically, for many use-cases, the system must be able to decode utterances in a streaming fashion and faster than real-time. Recently, a streaming recurrent neural network transducer (RNN-T) end-to-end (E2E) model has shown to be a good candidat… ▽ More

    Submitted 28 August, 2019; originally announced August 2019.