Speech emotion recognition and emotional text-to-speech for humanoid robots
Two 2026 arXiv papers address both directions of affective speech: recognising emotion in speech (PhysioSER; Zhang et al., 2026) and generating emotional speech (DUET; Zhang et al., 2026). Both methods are designed as plug-and-play frameworks that keep the pretrained model frozen, and both were deployed on a humanoid robot.
PhysioSER recognises emotion by complementing a frozen self-supervised (SSL) backbone with a parallel branch that encodes physiology-informed amplitude and phase views of the voice with quaternion convolutions. DUET controls emotion in pretrained diffusion and flow-matching text-to-speech models without fine-tuning, by steering hidden states along a linearly decodable emotion direction and refining the mel estimate with emotion-recognizer gradients backpropagated through a differentiable vocoder.
Papers compared
| Method | Published | Core idea | Evaluation | Code |
|---|---|---|---|---|
| DUET | arXiv:2606.00066 (2026) | Steer frozen TTS hidden states along a linear emotion direction; refine mel estimates with recognizer gradients through a differentiable vocoder. | 5 frozen backbones, ESD/CREMA-D/IEMOCAP, 10 supervised baselines; ESD accuracy 75.5% vs 46.8%; highest human EMOS (3.93). | — |
| PhysioSER | arXiv:2602.13259 (2026) | Encodes STFT amplitude and phase features with a quaternion encoder, contrastively aligned with a frozen SSL backbone. | 14 SER datasets, 6 frozen backbones; WA, UA, macro-F1; CREMA-D WavLM WA 69.69% to 75.20%. | — |
DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech
Xu Zhang, Longbing Cao, Zhangkai Wu · arXiv:2606.00066 (2026)
DUET adds emotion control to frozen diffusion and flow-matching TTS models by steering hidden states along a linearly decodable emotion direction and refining the mel estimate with emotion-recognizer gradients passed through a differentiable vocoder. With GradTTS on ESD it reaches 75.5% average emotion accuracy (strongest supervised baseline: 46.8%).
- In frozen diffusion and flow-matching TTS models, emotion is a linearly decodable direction of the hidden states, nearly orthogonal to speaker identity (absolute cosine similarity 0.029 on F5-TTS).
- DUET unifies hidden-state steering and mel-space guidance in one per-step update. Steering uses probe-derived emotion directions at the most discriminative layer; guidance backpropagates emotion2vec gradients through the Vocos vocoder to the mel estimate.
- On ESD, CREMA-D and IEMOCAP, four of five DUET-equipped frozen backbones (F5-TTS, Matcha-TTS, GradTTS, ProDiff) beat all 10 supervised emotional-TTS baselines in average emotion accuracy; StableTTS beats them only on IEMOCAP.
- Ablation on F5-TTS with ESD: removing hidden-state steering lowers average emotion accuracy from 64.9% to 40.6%, and removing mel-space guidance lowers it to 45.4%, so both components contribute.
PhysioSER: Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
Xu Zhang, Longbing Cao, Runze Yang, Zhangkai Wu · arXiv:2602.13259 (2026)
PhysioSER complements a frozen self-supervised (SSL) backbone with a compact physiology-informed branch that encodes vocal amplitude and phase features with quaternion convolutions, for speech emotion recognition. With frozen WavLM on CREMA-D, it raises weighted accuracy from 69.69% to 75.20%.
- PhysioSER derives four STFT features grounded in voice anatomy and physiology: log-magnitude, log-magnitude rate (spectral flux), instantaneous frequency and group delay. It embeds each time-frequency bin as one quaternion.
- A Hamilton-structured Quaternion Spectrotemporal Encoder (QSE) models the interactions among the four features. Replacing it with a real-valued CNN of comparable size lowers CREMA-D weighted accuracy from 75.20% to 72.78%.
- Contrastive Projection and Alignment (CPA) aligns utterance-level features of the physiology branch and the frozen SSL branch with a symmetric InfoNCE loss; a shallow Transformer fusion head then classifies the emotion.
- Across 14 datasets and 6 frozen backbones, PhysioSER improves WA and macro-F1 over the Backbone-only baseline in every pair. Gains are larger for weaker backbones, up to +136.07% relative WA (JLCorpus, Wav2vec2).
Key concepts
- Speech emotion recognition (SER)
- Classifying the emotion expressed in a spoken utterance, for example angry, happy, sad or neutral. Papers: PhysioSER
- Emotional text-to-speech
- Speech synthesis that renders text with a chosen emotion rather than a neutral voice. Papers: DUET
- Flow-matching and diffusion TTS
- Text-to-speech models that generate mel spectrograms by iteratively denoising or transporting noise, such as F5-TTS, Matcha-TTS and Grad-TTS. Papers: DUET
- Activation (representation) steering
- Changing a frozen model's behaviour at inference time by shifting its hidden states along a learned direction, without retraining. Papers: DUET
- Self-supervised speech models
- Speech encoders pre-trained without labels, such as wav2vec 2.0, HuBERT and WavLM, often used as frozen feature extractors. Papers: PhysioSER
- Phase features: group delay and instantaneous frequency
- Quantities derived from the phase of the short-time Fourier transform; most SER models use magnitude only. Papers: PhysioSER
Frequently asked questions
How does DUET add emotion control to a pretrained diffusion or flow-matching TTS model without fine-tuning?
DUET (arXiv:2606.00066) keeps the TTS model frozen. It steers the hidden states along a linearly decodable emotion direction and refines the mel estimate with mel-space guidance, backpropagating emotion-recognizer gradients through a differentiable vocoder, in a single per-step update during generation.
How does PhysioSER model vocal phase as well as amplitude for speech emotion recognition?
PhysioSER (arXiv:2602.13259) builds amplitude and phase views informed by voice anatomy and physiology, embeds them in a quaternion field, and contrastively aligns the resulting utterance-level features with those of a frozen self-supervised (SSL) backbone before a shallow fusion head classifies the emotion.
Which datasets are used to evaluate these speech papers?
DUET is evaluated on ESD, CREMA-D and IEMOCAP. PhysioSER is evaluated on 14 emotional speech datasets, including CREMA-D, RAVDESS, EmoV-DB, JL-Corpus, EMNS, RESD, nEMO, AESDD, Emozionalmente, Oréau, MESD, SUBESCO, PAVOQUE and CaFE.
Which pretrained models do they work with?
DUET is plugged into frozen F5-TTS, Matcha-TTS, Grad-TTS, ProDiff and StableTTS. PhysioSER works with six frozen self-supervised backbones: wav2vec 2.0, HuBERT, WavLM, emotion2vec, BEATs and the CLAP audio encoder.
Fields, tasks, datasets and related search terms
Task
emotional text-to-speech emotional speech synthesis expressive speech synthesis emotion-controllable TTS controllable speech synthesis affective speech generation training-free emotion control speech emotion recognition (SER) emotion recognition from speech vocal emotion recognition multilingual speech emotion recognition paralinguistic analysis affective computing
Method
activation steering representation steering steering vectors for TTS inference-time intervention classifier guidance for speech linear probing of hidden states emotion direction in hidden states speaker-emotion disentanglement phase-aware speech features group delay instantaneous frequency spectral flux quaternion convolution quaternion neural networks frozen self-supervised speech representations InfoNCE contrastive alignment interpretable speech emotion recognition parameter-efficient SER
Models used or compared
F5-TTS Matcha-TTS Grad-TTS ProDiff StableTTS Vocos vocoder emotion2vec Qwen3-TTS CosyVoice2 EmoKnob EmoSphere++ IndexTTS2
Datasets
ESD (Emotional Speech Dataset) CREMA-D IEMOCAP RAVDESS EmoV-DB JL-Corpus EMNS RESD nEMO AESDD Emozionalmente Oréau MESD SUBESCO PAVOQUE CaFE
中文
情感语音合成 情感可控语音合成 表现力语音合成 可控语音合成 表征引导 推理时干预 扩散模型语音合成 流匹配语音合成 人形机器人语音 语音情感识别 语音情绪识别 相位特征 群延迟 瞬时频率 四元数卷积 自监督语音表征 多语言语音情感识别 情感计算
Fields (broad to narrow)
artificial intelligence machine learning deep learning speech and audio processing spoken language technology generative AI speech synthesis text-to-speech (TTS) expressive TTS emotional TTS emotion control in TTS training-free emotion steering paralinguistics affective computing speech emotion recognition (SER) physiology-informed SER
Related areas
affective computing human-robot interaction social robotics embodied AI conversational agents voice assistants digital humans controllable generation representation engineering mechanistic interpretability diffusion models flow matching score-based generative models neural vocoders mel spectrogram prosody modeling speaking style control zero-shot TTS voice cloning emotional voice conversion speech emotion recognition emotion recognition multimodal emotion recognition audio classification speech representation learning self-supervised learning foundation models for audio voice physiology vocal tract and glottal source signal processing time-frequency analysis short-time Fourier transform (STFT) hypercomplex neural networks interpretable machine learning cross-lingual transfer mental health assessment
Applications
expressive humanoid robots empathetic dialogue systems audiobook narration game and virtual character voices assistive speech emotion-aware humanoid robots call-centre analytics psychological assessment support human-computer interaction voice assistants
Related benchmarks (not used in this paper)
EmoV-DB RAVDESS MSP-Podcast LibriTTS VCTK LJSpeech IEMOCAP ESD TESS SAVEE EmoDB CASIA
中文 (扩展)
语音合成 文本转语音 TTS 情感TTS 情绪语音合成 风格可控语音合成 韵律控制 扩散模型 流匹配 声码器 梅尔频谱 表征工程 可解释性 情感计算 人机交互 社交机器人 具身智能 数字人 零样本语音合成 情感语音转换 情感识别 语音情绪分类 副语言信息 语音表征学习 自监督学习 音频基础模型 发声生理 声道 声门 短时傅里叶变换 时频分析 超复数神经网络 可解释机器学习 跨语言迁移 心理健康评估
Models used
wav2vec 2.0 HuBERT WavLM emotion2vec BEATs CLAP audio encoder
中文概述
面向人形机器人的语音情感识别与情感语音合成
两篇 2026 年 arXiv 论文分别覆盖情感语音的两个方向:识别语音中的情感(PhysioSER)与生成带情感的语音(DUET)。两者都采用即插即用设计,保持预训练模型冻结,并都部署在人形机器人上。
相关检索词:语音情感识别;情感语音合成;情感可控语音合成;表现力语音合成;自监督语音表征;人形机器人情感交互
Cite these papers
@article{zhang2026duet,
title = {{DUET}: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech},
author = {Zhang, Xu and Cao, Longbing and Wu, Zhangkai},
journal = {arXiv preprint arXiv:2606.00066},
year = {2026},
eprint = {2606.00066},
archivePrefix= {arXiv},
primaryClass = {cs.SD},
doi = {10.48550/arXiv.2606.00066},
url = {https://arxiv.org/abs/2606.00066}
}
@article{zhang2026physioser,
title = {Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition},
author = {Zhang, Xu and Cao, Longbing and Yang, Runze and Wu, Zhangkai},
journal = {arXiv preprint arXiv:2602.13259},
year = {2026},
eprint = {2602.13259},
archivePrefix= {arXiv},
primaryClass = {cs.SD},
doi = {10.48550/arXiv.2602.13259},
url = {https://arxiv.org/abs/2602.13259}
}