Search Results for author: Penny Karanasou

Found 10 papers, 0 papers with code

A Comparative Analysis of Pretrained Language Models for Text-to-Speech

no code implementations • 4 Sep 2023 • Marcel Granero-Moya, Penny Karanasou, Sri Karlapati, Bastian Schnell, Nicole Peinelt, Alexis Moinet, Thomas Drugman

In this study, we aim to address this gap by conducting a comparative analysis of different PLMs for two TTS tasks: prosody prediction and pause prediction.

Natural Language Understanding Prosody Prediction

Paper
Add Code

Controllable Emphasis with zero data for text-to-speech

no code implementations • 13 Jul 2023 • Arnaud Joly, Marco Nicolis, Ekaterina Peterova, Alessandro Lombardi, Ammar Abbas, Arent van Korlaar, Aman Hussain, Parul Sharma, Alexis Moinet, Mateusz Lajszczak, Penny Karanasou, Antonio Bonafonte, Thomas Drugman, Elena Sokolova

We show that this technique significantly closes the gap to methods that require explicit recordings.

Sentence

Paper
Add Code

eCat: An End-to-End Model for Multi-Speaker TTS & Many-to-Many Fine-Grained Prosody Transfer

no code implementations • 20 Jun 2023 • Ammar Abbas, Sri Karlapati, Bastian Schnell, Penny Karanasou, Marcel Granero Moya, Amith Nagaraj, Ayman Boustati, Nicole Peinelt, Alexis Moinet, Thomas Drugman

We show that eCat statistically significantly reduces the gap in naturalness between CopyCat2 and human recordings by an average of 46. 7% across 2 languages, 3 locales, and 7 speakers, along with better target-speaker similarity in FPT.

Paper
Add Code

Simple and Effective Multi-sentence TTS with Expressive and Coherent Prosody

no code implementations • 29 Jun 2022 • Peter Makarov, Ammar Abbas, Mateusz Łajszczak, Arnaud Joly, Sri Karlapati, Alexis Moinet, Thomas Drugman, Penny Karanasou

In this paper, we examine simple extensions to a Transformer-based FastSpeech-like system, with the goal of improving prosody for multi-sentence TTS.

Language Modelling Sentence

Paper
Add Code

CopyCat2: A Single Model for Multi-Speaker TTS and Many-to-Many Fine-Grained Prosody Transfer

no code implementations • 27 Jun 2022 • Sri Karlapati, Penny Karanasou, Mateusz Lajszczak, Ammar Abbas, Alexis Moinet, Peter Makarov, Ray Li, Arent van Korlaar, Simon Slangen, Thomas Drugman

In this paper, we present CopyCat2 (CC2), a novel model capable of: a) synthesizing speech with different speaker identities, b) generating speech with expressive and contextually appropriate prosody, and c) transferring prosody at fine-grained level between any pair of seen speakers.

Paper
Add Code

Multi-Scale Spectrogram Modelling for Neural Text-to-Speech

no code implementations • 29 Jun 2021 • Ammar Abbas, Bajibabu Bollepalli, Alexis Moinet, Arnaud Joly, Penny Karanasou, Peter Makarov, Simon Slangens, Sri Karlapati, Thomas Drugman

We propose a novel Multi-Scale Spectrogram (MSS) modelling approach to synthesise speech with an improved coarse and fine-grained prosody.

Sentence

Paper
Add Code

A learned conditional prior for the VAE acoustic space of a TTS system

no code implementations • 14 Jun 2021 • Penny Karanasou, Sri Karlapati, Alexis Moinet, Arnaud Joly, Ammar Abbas, Simon Slangen, Jaime Lorenzo Trueba, Thomas Drugman

Many factors influence speech yielding different renditions of a given sentence.

Sentence

Paper
Add Code

Prosodic Representation Learning and Contextual Sampling for Neural Text-to-Speech

no code implementations • 4 Nov 2020 • Sri Karlapati, Ammar Abbas, Zack Hodari, Alexis Moinet, Arnaud Joly, Penny Karanasou, Thomas Drugman

In Stage II, we propose a novel method to sample from this learnt prosodic distribution using the contextual information available in text.

Graph Attention Representation Learning +2

Paper
Add Code

Cross-lingual Transfer Learning for Japanese Named Entity Recognition

no code implementations • NAACL 2019 • Andrew Johnson, Penny Karanasou, Judith Gaspers, Dietrich Klakow

This work explores cross-lingual transfer learning (TL) for named entity recognition, focusing on bootstrapping Japanese from English.

Cross-Lingual Transfer named-entity-recognition +3

Paper
Add Code

Selecting Machine-Translated Data for Quick Bootstrapping of a Natural Language Understanding System

no code implementations • NAACL 2018 • Judith Gaspers, Penny Karanasou, Rajen Chatterjee

The goal is to decrease the cost and time needed to get an annotated corpus for the new language, while still having a large enough coverage of user requests.

Machine Translation Natural Language Understanding +1

Paper
Add Code

Cannot find the paper you are looking for? You can Submit a new open access paper.