Papers
arxiv:1804.00015

ESPnet: End-to-End Speech Processing Toolkit

Published on Mar 30, 2018
Authors:
,
,
,
,
,
,
,
,
,
,
,

Abstract

This paper introduces a new open source platform for end-to-end speech processing named ESPnet. ESPnet mainly focuses on end-to-end automatic speech recognition (ASR), and adopts widely-used dynamic neural network toolkits, Chainer and PyTorch, as a main deep learning engine. ESPnet also follows the Kaldi ASR toolkit style for data processing, feature extraction/format, and recipes to provide a complete setup for speech recognition and other speech processing experiments. This paper explains a major architecture of this software platform, several important functionalities, which differentiate ESPnet from other open source ASR toolkits, and experimental results with major ASR benchmarks.

Community

Sign up or log in to comment

Models citing this paper 655

Browse 655 models citing this paper

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/1804.00015 in a dataset README.md to link it from this page.

Spaces citing this paper 188

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.