arxiv:2402.08093

BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Published on Feb 12

· Submitted by

akhaliq on Feb 14

#1 Paper of the day

Upvote

Authors:

Guillermo Cámbara ,

Yang Li ,

Fatih Beyhan ,

Arent van Korlaar ,

Adam Michalski ,

Haohan Guo ,

Abstract

We introduce a text-to-speech (TTS) model called BASE TTS, which stands for Big Adaptive Streamable TTS with Emergent abilities. BASE TTS is the largest TTS model to-date, trained on 100K hours of public domain speech data, achieving a new state-of-the-art in speech naturalness. It deploys a 1-billion-parameter autoregressive Transformer that converts raw texts into discrete codes ("speechcodes") followed by a convolution-based decoder which converts these speechcodes into waveforms in an incremental, streamable manner. Further, our speechcodes are built using a novel speech tokenization technique that features speaker ID disentanglement and compression with byte-pair encoding. Echoing the widely-reported "emergent abilities" of large language models when trained on increasing volume of data, we show that BASE TTS variants built with 10K+ hours and 500M+ parameters begin to demonstrate natural prosody on textually complex sentences. We design and share a specialized dataset to measure these emergent abilities for text-to-speech. We showcase state-of-the-art naturalness of BASE TTS by evaluating against baselines that include publicly available large-scale text-to-speech systems: YourTTS, Bark and TortoiseTTS. Audio samples generated by the model can be heard at https://amazon-ltts-paper.com/.

View arXiv page View PDF Add to collection

Community

mrfakename

Feb 14

This is incredible! It’s really great with emotions. We need an open sourced implementation!

Alignment-Lab-AI

Feb 14

thi is a really cool idea for an implementation, its really awesome, almost like a hash

jnemecek

Feb 14

"However, due to the potential misuse of this capability, we have decided against open-sourcing this model as a precautionary measure." I'm really tired of seeing this excuse in speech generation models.

mrfakename

Feb 14

cc: @lucidrains

MichaelBarryUK

Feb 14

"However, due to the potential misuse of this capability, we have decided against open-sourcing this model as a precautionary measure." I'm really tired of seeing this excuse in speech generation models.

It's virtue signalling, and can be translated as "F you, this belongs to me", but of course they are saints, and saints can't say such things.