Problem with Utterance-level Prosody extractor of DelightfulTTS

I've recently been experimenting with your implementation of DelightfulTTS and the voice quality is awesome. However I found out that the embedding vector output of Utterance-level Prosody extractor is very small, making the that of Utterance-level Prosody predictor small as well (L2 is roughly 12 and each element in the vector is roughly 0.2 to 0.3). Vectors with element close to zero means this layer mostly doesn't add any information at all. Have you find any solution to this?

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Problem with Utterance-level Prosody extractor of DelightfulTTS #7

Metadata

Assignees

Labels

Projects

Milestone

Relationships

Development

Problem with Utterance-level Prosody extractor of DelightfulTTS #7

Description

Metadata

Metadata

Assignees

Labels

Projects

Milestone

Relationships

Development

Issue actions