Speech emotion recognition and emotional text-to-speech for humanoid robots

Two 2026 arXiv papers address both directions of affective speech: recognising emotion in speech (PhysioSER; Zhang et al., 2026) and generating emotional speech (DUET; Zhang et al., 2026). Both methods are designed as plug-and-play frameworks that keep the pretrained model frozen, and both were deployed on a humanoid robot.

PhysioSER recognises emotion by complementing a frozen self-supervised (SSL) backbone with a parallel branch that encodes physiology-informed amplitude and phase views of the voice with quaternion convolutions. DUET controls emotion in pretrained diffusion and flow-matching text-to-speech models without fine-tuning, by steering hidden states along a linearly decodable emotion direction and refining the mel estimate with emotion-recognizer gradients backpropagated through a differentiable vocoder.

Papers compared

MethodPublishedCore ideaEvaluationCode
DUETarXiv:2606.00066 (2026)Steer frozen TTS hidden states along a linear emotion direction; refine mel estimates with recognizer gradients through a differentiable vocoder.5 frozen backbones, ESD/CREMA-D/IEMOCAP, 10 supervised baselines; ESD accuracy 75.5% vs 46.8%; highest human EMOS (3.93).—
PhysioSERarXiv:2602.13259 (2026)Encodes STFT amplitude and phase features with a quaternion encoder, contrastively aligned with a frozen SSL backbone.14 SER datasets, 6 frozen backbones; WA, UA, macro-F1; CREMA-D WavLM WA 69.69% to 75.20%.—

DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech

Xu Zhang, Longbing Cao, Zhangkai Wu · arXiv:2606.00066 (2026)

DUET adds emotion control to frozen diffusion and flow-matching TTS models by steering hidden states along a linearly decodable emotion direction and refining the mel estimate with emotion-recognizer gradients passed through a differentiable vocoder. With GradTTS on ESD it reaches 75.5% average emotion accuracy (strongest supervised baseline: 46.8%).

PhysioSER: Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition

Xu Zhang, Longbing Cao, Runze Yang, Zhangkai Wu · arXiv:2602.13259 (2026)

PhysioSER complements a frozen self-supervised (SSL) backbone with a compact physiology-informed branch that encodes vocal amplitude and phase features with quaternion convolutions, for speech emotion recognition. With frozen WavLM on CREMA-D, it raises weighted accuracy from 69.69% to 75.20%.

Key concepts

Speech emotion recognition (SER)
Classifying the emotion expressed in a spoken utterance, for example angry, happy, sad or neutral. Papers: PhysioSER
Emotional text-to-speech
Speech synthesis that renders text with a chosen emotion rather than a neutral voice. Papers: DUET
Flow-matching and diffusion TTS
Text-to-speech models that generate mel spectrograms by iteratively denoising or transporting noise, such as F5-TTS, Matcha-TTS and Grad-TTS. Papers: DUET
Activation (representation) steering
Changing a frozen model's behaviour at inference time by shifting its hidden states along a learned direction, without retraining. Papers: DUET
Self-supervised speech models
Speech encoders pre-trained without labels, such as wav2vec 2.0, HuBERT and WavLM, often used as frozen feature extractors. Papers: PhysioSER
Phase features: group delay and instantaneous frequency
Quantities derived from the phase of the short-time Fourier transform; most SER models use magnitude only. Papers: PhysioSER

Frequently asked questions

How does DUET add emotion control to a pretrained diffusion or flow-matching TTS model without fine-tuning?

DUET (arXiv:2606.00066) keeps the TTS model frozen. It steers the hidden states along a linearly decodable emotion direction and refines the mel estimate with mel-space guidance, backpropagating emotion-recognizer gradients through a differentiable vocoder, in a single per-step update during generation.

How does PhysioSER model vocal phase as well as amplitude for speech emotion recognition?

PhysioSER (arXiv:2602.13259) builds amplitude and phase views informed by voice anatomy and physiology, embeds them in a quaternion field, and contrastively aligns the resulting utterance-level features with those of a frozen self-supervised (SSL) backbone before a shallow fusion head classifies the emotion.

Which datasets are used to evaluate these speech papers?

DUET is evaluated on ESD, CREMA-D and IEMOCAP. PhysioSER is evaluated on 14 emotional speech datasets, including CREMA-D, RAVDESS, EmoV-DB, JL-Corpus, EMNS, RESD, nEMO, AESDD, Emozionalmente, Oréau, MESD, SUBESCO, PAVOQUE and CaFE.

Which pretrained models do they work with?

DUET is plugged into frozen F5-TTS, Matcha-TTS, Grad-TTS, ProDiff and StableTTS. PhysioSER works with six frozen self-supervised backbones: wav2vec 2.0, HuBERT, WavLM, emotion2vec, BEATs and the CLAP audio encoder.

Fields, tasks, datasets and related search terms

Task

emotional text-to-speech emotional speech synthesis expressive speech synthesis emotion-controllable TTS controllable speech synthesis affective speech generation training-free emotion control speech emotion recognition (SER) emotion recognition from speech vocal emotion recognition multilingual speech emotion recognition paralinguistic analysis affective computing

Method

activation steering representation steering steering vectors for TTS inference-time intervention classifier guidance for speech linear probing of hidden states emotion direction in hidden states speaker-emotion disentanglement phase-aware speech features group delay instantaneous frequency spectral flux quaternion convolution quaternion neural networks frozen self-supervised speech representations InfoNCE contrastive alignment interpretable speech emotion recognition parameter-efficient SER

Models used or compared

F5-TTS Matcha-TTS Grad-TTS ProDiff StableTTS Vocos vocoder emotion2vec Qwen3-TTS CosyVoice2 EmoKnob EmoSphere++ IndexTTS2

Datasets

ESD (Emotional Speech Dataset) CREMA-D IEMOCAP RAVDESS EmoV-DB JL-Corpus EMNS RESD nEMO AESDD Emozionalmente Oréau MESD SUBESCO PAVOQUE CaFE

中文

情感语音合成 情感可控语音合成 表现力语音合成 可控语音合成 表征引导 推理时干预 扩散模型语音合成 流匹配语音合成 人形机器人语音 语音情感识别 语音情绪识别 相位特征 群延迟 瞬时频率 四元数卷积 自监督语音表征 多语言语音情感识别 情感计算

Fields (broad to narrow)

artificial intelligence machine learning deep learning speech and audio processing spoken language technology generative AI speech synthesis text-to-speech (TTS) expressive TTS emotional TTS emotion control in TTS training-free emotion steering paralinguistics affective computing speech emotion recognition (SER) physiology-informed SER

Related areas

affective computing human-robot interaction social robotics embodied AI conversational agents voice assistants digital humans controllable generation representation engineering mechanistic interpretability diffusion models flow matching score-based generative models neural vocoders mel spectrogram prosody modeling speaking style control zero-shot TTS voice cloning emotional voice conversion speech emotion recognition emotion recognition multimodal emotion recognition audio classification speech representation learning self-supervised learning foundation models for audio voice physiology vocal tract and glottal source signal processing time-frequency analysis short-time Fourier transform (STFT) hypercomplex neural networks interpretable machine learning cross-lingual transfer mental health assessment

Applications

expressive humanoid robots empathetic dialogue systems audiobook narration game and virtual character voices assistive speech emotion-aware humanoid robots call-centre analytics psychological assessment support human-computer interaction voice assistants

Related benchmarks (not used in this paper)

EmoV-DB RAVDESS MSP-Podcast LibriTTS VCTK LJSpeech IEMOCAP ESD TESS SAVEE EmoDB CASIA

中文 (扩展)

语音合成 文本转语音 TTS 情感TTS 情绪语音合成 风格可控语音合成 韵律控制 扩散模型 流匹配 声码器 梅尔频谱 表征工程 可解释性 情感计算 人机交互 社交机器人 具身智能 数字人 零样本语音合成 情感语音转换 情感识别 语音情绪分类 副语言信息 语音表征学习 自监督学习 音频基础模型 发声生理 声道 声门 短时傅里叶变换 时频分析 超复数神经网络 可解释机器学习 跨语言迁移 心理健康评估

Models used

wav2vec 2.0 HuBERT WavLM emotion2vec BEATs CLAP audio encoder

中文概述

面向人形机器人的语音情感识别与情感语音合成

两篇 2026 年 arXiv 论文分别覆盖情感语音的两个方向:识别语音中的情感(PhysioSER)与生成带情感的语音(DUET)。两者都采用即插即用设计,保持预训练模型冻结,并都部署在人形机器人上。

相关检索词:语音情感识别;情感语音合成;情感可控语音合成;表现力语音合成;自监督语音表征;人形机器人情感交互

Cite these papers

@article{zhang2026duet,
  title        = {{DUET}: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech},
  author       = {Zhang, Xu and Cao, Longbing and Wu, Zhangkai},
  journal      = {arXiv preprint arXiv:2606.00066},
  year         = {2026},
  eprint       = {2606.00066},
  archivePrefix= {arXiv},
  primaryClass = {cs.SD},
  doi          = {10.48550/arXiv.2606.00066},
  url          = {https://arxiv.org/abs/2606.00066}
}

@article{zhang2026physioser,
  title        = {Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition},
  author       = {Zhang, Xu and Cao, Longbing and Yang, Runze and Wu, Zhangkai},
  journal      = {arXiv preprint arXiv:2602.13259},
  year         = {2026},
  eprint       = {2602.13259},
  archivePrefix= {arXiv},
  primaryClass = {cs.SD},
  doi          = {10.48550/arXiv.2602.13259},
  url          = {https://arxiv.org/abs/2602.13259}
}