Krisshvamsi
/

TTS

@@ -18,7 +18,7 @@ pipeline_tag: text-to-speech
 # Text-to-Speech (TTS) with Transformer trained on LJSpeech
-This repository provides all the necessary tools for Text-to-Speech (TTS)  with SpeechBrain using a [Tacotron2](https://arxiv.org/abs/1712.05884) pretrained on [LJSpeech](https://keithito.com/LJ-Speech-Dataset/).
 The pre-trained model takes in input a short text and produces a spectrogram in output. One can get the final waveform by applying a vocoder (e.g., HiFIGAN) on top of the generated spectrogram.
@@ -32,15 +32,16 @@ pip install speechbrain
 ```python
 import torchaudio
-from speechbrain.inference.TTS import Tacotron2
 from speechbrain.inference.vocoders import HIFIGAN
-# Intialize TTS (tacotron2) and Vocoder (HiFIGAN)
-tacotron2 = Tacotron2.from_hparams(source="speechbrain/tts-tacotron2-ljspeech", savedir="tmpdir_tts")
 hifi_gan = HIFIGAN.from_hparams(source="speechbrain/tts-hifigan-ljspeech", savedir="tmpdir_vocoder")
 # Running the TTS
-mel_output, mel_length, alignment = tacotron2.encode_text("Mary had a little lamb")
 # Running Vocoder (spectrogram-to-waveform)
 waveforms = hifi_gan.decode_batch(mel_output)
@@ -49,19 +50,8 @@ waveforms = hifi_gan.decode_batch(mel_output)
 torchaudio.save('example_TTS.wav',waveforms.squeeze(1), 22050)
 ```
-If you want to generate multiple sentences in one-shot, you can do in this way:
-```
-from speechbrain.pretrained import Tacotron2
-tacotron2 = Tacotron2.from_hparams(source="speechbrain/TTS_Tacotron2", savedir="tmpdir")
-items = [
-       "A quick brown fox jumped over the lazy dog",
-       "How much wood would a woodchuck chuck?",
-       "Never odd or even"
-     ]
-mel_outputs, mel_lengths, alignments = tacotron2.encode_batch(items)
-```
 ### Inference on GPU
 To perform inference on the GPU, add  `run_opts={"device":"cuda"}`  when calling the `from_hparams` method.

 # Text-to-Speech (TTS) with Transformer trained on LJSpeech
+This repository provides all the necessary tools for Text-to-Speech (TTS)  with SpeechBrain using a [Transformer](https://arxiv.org/pdf/1809.08895.pdf) pretrained on [LJSpeech](https://keithito.com/LJ-Speech-Dataset/).
 The pre-trained model takes in input a short text and produces a spectrogram in output. One can get the final waveform by applying a vocoder (e.g., HiFIGAN) on top of the generated spectrogram.
 ```python
 import torchaudio
 from speechbrain.inference.vocoders import HIFIGAN
+texts = ["This is a sample text for synthesis."]
+# Intialize TTS (Transformer) and Vocoder (HiFIGAN)
+my_tts_model = TTSModel.from_hparams(source="/content/")
 hifi_gan = HIFIGAN.from_hparams(source="speechbrain/tts-hifigan-ljspeech", savedir="tmpdir_vocoder")
 # Running the TTS
+mel_output, mel_length = my_tts_model.encode_text(texts)
 # Running Vocoder (spectrogram-to-waveform)
 waveforms = hifi_gan.decode_batch(mel_output)
 torchaudio.save('example_TTS.wav',waveforms.squeeze(1), 22050)
 ```
+If you want to generate multiple sentences in one-shot, pass the sentences as items in a list.
 ### Inference on GPU
 To perform inference on the GPU, add  `run_opts={"device":"cuda"}`  when calling the `from_hparams` method.