Rectifying the Emotional Flow: Aligning Priors and Dynamic Guidance for High-Arousal Text-to-Speech

Anonymous Authors

Method overview: Emotion-Rectified Noise Prior and Likelihood-Inverse Guidance
Overview of the proposed inference framework: Emotion-Rectified Noise Prior (ERNP) and Likelihood-Inverse Guidance (LIG).

Abstract

While diffusion and flow-matching models have advanced TTS, generating high-arousal emotions remains a persistent challenge due to the trade-off between stability and expressiveness. Existing systems often suffer from linguistic collapse when pursuing high intensity or fail to meet target emotional levels under stable settings. In this work, we identify that standard Gaussian initialization introduces a neutral-prosody bias, while uniform Classifier-Free Guidance can distort the acoustic manifold and lead to artifacts. To address this, we propose an inference framework that rectifies the emotional trajectory. An Emotion-Rectified Noise Prior injects a semantic gradient at initialization to align sampling with the target emotional manifold, and Likelihood-Inverse Guidance adaptively schedules guidance via a conditional/unconditional likelihood ratio, strengthening guidance only when the trajectory drifts toward a neutral fallback. Experiments show improved stability and expressiveness without retraining.

Where our method helps

  • More accurate high-arousal emotion: ERNP reduces neutral-prosody drift early in sampling.
  • Fewer linguistic failures: LIG increases guidance only when needed, avoiding over-guidance that can cause broken phrasing or unnatural pauses.
  • Cleaner acoustics: adaptive guidance better respects the acoustic manifold, reducing audible artifacts.
  • More stable phonation: improved continuity and fewer sudden energy spikes under intense expressions.

Audio Demos