PhysioSER: Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition
Xu Zhang, Longbing Cao, Runze Yang, Zhangkai Wu
arXiv:2602.13259 (2026) · Free to read on arXiv (arXiv non-exclusive distribution license)
Also indexed in: Semantic Scholar · OpenAlex · dblp · Hugging Face · alphaXiv

Quick facts
| Method name | PhysioSER |
|---|---|
| Authors | Xu Zhang, Longbing Cao, Runze Yang, Zhangkai Wu |
| Published in | arXiv:2602.13259 (2026) |
| arXiv | arXiv:2602.13259 (DOI 10.48550/arXiv.2602.13259) |
| Task | Utterance-level speech emotion recognition (SER), motivated by humanoid-robot uses such as social interaction and robotic psychological diagnosis |
| Architecture | Compact plug-and-play physiology-informed branch running in parallel with a frozen self-supervised (SSL) backbone; CPA aligns the two, and a shallow attention fusion head classifies |
| Inputs | Raw waveform for the SSL branch; for the physiology branch, an STFT-derived quartet: log-magnitude M, log-magnitude rate ρ (spectral flux), instantaneous frequency f_inst, group delay τ_g |
| Encoder | Quaternion Spectrotemporal Encoder (QSE): each bin embedded as Q = M + iρ + j f_inst + k τ_g, then Hamilton-product quaternion convolutions, quaternion batch normalization and radial qReLU |
| Alignment and fusion | Contrastive Projection and Alignment (CPA): projection heads and a symmetric InfoNCE loss, with the SSL model frozen; then a shallow Transformer encoder and compact classifier |
| SSL backbones | Six, all frozen: Wav2vec2, HuBERT, WavLM, Emotion2Vec, BEATs and the CLAP audio encoder, covering four pre-training paradigms |
| Datasets | 14 public acted or simulated emotional speech datasets: 5 English (CREMA-D, EMNS, JLCorpus, EmoV-DB, RAVDESS) and 9 non-English (RESD, nEMO, AESDD, Emozionalmente, Oréau, MESD, SUBESCO, PAVOQUE, CaFE) |
| Metrics and protocol | Weighted accuracy (WA), unweighted accuracy (UA), macro-F1; 7:1:2 train/validation/test split; mean of five random seeds; every Δ is a relative % change |
| Baselines | Backbone-only (same frozen backbone, masked attention pooling and classifier, no physiology branch); on CREMA-D also fully fine-tuned WavLM and a real-valued CNN variant |
| Headline result | CREMA-D with frozen WavLM: WA 69.69% → 75.20% (+7.91% relative) with 2.03M trainable parameters; fully fine-tuned WavLM reaches 73.66% with 94.61M |
| Deployment | Real-time on the Ameca humanoid robot: vocal emotion is mapped to Ameca's facial expressions without automatic speech recognition; the paper links a demo video |
| BibTeX key | zhang2026physioser |
Main results
| Setting | Result | Compared with |
|---|---|---|
| CREMA-D, frozen WavLM backbone | WA = 75.20% (UA 75.62%, macro-F1 75.63%)“On dataset CREMA-D with a frozen WavLM [8], WA increases from 69.69% to 75.20%.” | Backbone-only frozen WavLM: WA 69.69% (+7.91% relative) |
| CREMA-D, parameter efficiency | WA = 75.20% with 2.03M trainable parameters (frozen WavLM)“PhysioSER achieves a superior 75.20% WA with only 2.03M trainable parameters (2.1% of the baseline)” | Fully fine-tuned WavLM: WA 73.66% with 94.61M parameters |
| CREMA-D ablation: quaternion QSE vs real-valued CNN of comparable size | WA = 75.20% (quaternion) vs 72.78% (Normal CNN)“While Normal CNN improves over WavLM Only, it trails our quaternion model by 3.22% WA.” | Real-valued CNN is 3.22% lower (relative) |
| CREMA-D ablation: feature subsets | Magnitude-only WA = 74.60% (−0.80% relative); Phase-only WA = 70.50% (−6.25% relative)“removing phase information (Magnitude-only) causes a minor 0.80% drop, whereas discarding magnitude (Phase-only) leads to a substantial 6.25% decline.” | Full quartet: WA 75.20% |
| JLCorpus (10 classes), Wav2vec2 backbone | WA = 66.88%“the model yields large relative improvements for weaker backbones (e.g., +136.07% WA on JLCorpus with Wav2vec2)” | Backbone-only: WA 28.33% (+136.07% relative, the largest WA gain in Tables I-II) |
| CaFE (French, 7 classes), Wav2Vec2 backbone | UA = 35.45% (−6.64% relative); WA = 35.83% (+11.65%); F1 = 35.86% (+12.77%)“Wav2Vec2 32.09 35.83 +11.65 37.97 35.45 −6.64 31.80 35.86 +12.77” | Backbone-only: UA 37.97%, WA 32.09%, F1 31.80%; the only negative Δ among the 252 reported in Tables I-II |
Quoted text is verbatim from the paper.
Key points
- PhysioSER derives four STFT features grounded in voice anatomy and physiology: log-magnitude, log-magnitude rate (spectral flux), instantaneous frequency and group delay. It embeds each time-frequency bin as one quaternion.
- A Hamilton-structured Quaternion Spectrotemporal Encoder (QSE) models the interactions among the four features. Replacing it with a real-valued CNN of comparable size lowers CREMA-D weighted accuracy from 75.20% to 72.78%.
- Contrastive Projection and Alignment (CPA) aligns utterance-level features of the physiology branch and the frozen SSL branch with a symmetric InfoNCE loss; a shallow Transformer fusion head then classifies the emotion.
- Across 14 datasets and 6 frozen backbones, PhysioSER improves WA and macro-F1 over the Backbone-only baseline in every pair. Gains are larger for weaker backbones, up to +136.07% relative WA (JLCorpus, Wav2vec2).
- PhysioSER runs in real time on the Ameca humanoid robot, mapping vocal emotion directly to Ameca's facial expressions without automatic speech recognition.
Abstract
Speech emotion recognition (SER) is essential for humanoid robot tasks such as social robotic interactions and robotic psychological diagnosis, where interpretable and efficient models are critical for safety and performance. Existing deep models trained on large datasets remain largely uninterpretable, often insufficiently modeling underlying emotional acoustic signals and failing to capture and analyze the core physiology of emotional vocal behaviors. Physiological research on human voices shows that the dynamics of vocal amplitude and phase correlate with emotions through the vocal tract filter and the glottal source. However, most existing deep models solely involve amplitude but fail to couple the physiological features of and between amplitude and phase. Here, we propose PhysioSER, a physiology-informed vocal spectrotemporal representation learning method, to address these issues with a compact, plug-and-play design. PhysioSER constructs amplitude and phase views informed by voice anatomy and physiology (VAP) to complement SSL models for SER. This VAP-informed framework incorporates two parallel workflows: a vocal feature representation branch to decompose vocal signals based on VAP, embed them into a quaternion field, and use Hamilton-structured quaternion convolutions for modeling their dynamic interactions; and a latent representation branch based on a frozen SSL backbone. Then, utterance-level features from both workflows are aligned by a Contrastive Projection and Alignment framework, followed by a shallow attention fusion head for SER classification. PhysioSER is shown to be interpretable and efficient for SER through extensive evaluations across 14 datasets, 10 languages, and 6 backbones, and its practical efficacy is validated by real-time deployment on a humanoid robotic platform.
Frequently asked questions
What is PhysioSER?
PhysioSER is a physiology-informed method for speech emotion recognition. It adds a compact branch that encodes vocal amplitude and phase features with Hamilton-structured quaternion convolutions, aligns it with a frozen self-supervised (SSL) backbone through Contrastive Projection and Alignment (CPA), and classifies the emotion with a shallow attention fusion head.
Why model phase as well as amplitude?
Physiological research shows that the dynamics of vocal amplitude and phase correlate with emotion through the vocal tract filter and the glottal source. Yet most existing deep SER models use amplitude only. Methods that do use phase often concatenate it with amplitude as a separate feature or treat the two as independent channels, instead of modelling the physiology-informed coupling between them.
Which acoustic features does PhysioSER extract?
It extracts four features from the short-time Fourier transform: log-magnitude M, log-magnitude rate ρ (spectral flux), instantaneous frequency f_inst and group delay τ_g. The authors describe them as spanning vocal-tract timbre, temporal energy modulation, laryngeal oscillation rate and phase dispersion. Each time-frequency bin is embedded as the quaternion M + iρ + j f_inst + k τ_g.
How much does PhysioSER improve a frozen WavLM, and at what cost?
On CREMA-D, weighted accuracy rises from 69.69% (frozen WavLM plus classifier) to 75.20%, a 7.91% relative gain, with 2.03M trainable parameters. Fully fine-tuning WavLM uses 94.61M parameters and reaches 73.66% WA.
Which backbones and datasets was PhysioSER evaluated on?
It was tested with six frozen backbones (Wav2vec2, HuBERT, WavLM, Emotion2Vec, BEATs and the CLAP audio encoder) on 14 public emotional speech datasets, 5 English and 9 non-English. WA and macro-F1 improve over the Backbone-only baseline for every backbone-dataset pair; the only drop in the main result tables (Tables I-II) is UA on CaFE with Wav2Vec2 (−6.64% relative).
Which robot was PhysioSER deployed on?
PhysioSER was deployed on the Ameca humanoid robot. It processes continuous audio in Ameca's perception-action loop and maps vocal emotion directly to Ameca's facial expressions, without automatic speech recognition or linguistic analysis. The paper links a demo video.
How do I cite PhysioSER?
Cite it as: Zhang, X., Cao, L., Yang, R., & Wu, Z. (2026). Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition. arXiv. https://doi.org/10.48550/arXiv.2602.13259 BibTeX is available at https://codezx6.github.io/papers/physioser.html#cite and in https://codezx6.github.io/publications.bib.
中文摘要
学习生理信息引导的嗓音时频表示用于语音情感识别(PhysioSER)
PhysioSER 是生理信息引导的语音情感识别方法:基于嗓音解剖与生理(VAP)提取对数幅度、谱通量、瞬时频率与群时延,嵌入四元数域并用 Hamilton 结构四元数卷积建模交互,经对比投影与对齐(CPA)与冻结自监督骨干的特征对齐后融合分类。CREMA-D 上以冻结 WavLM 为骨干时,加权准确率由 69.69% 升至 75.20%;已在 Ameca 人形机器人上实时部署。
关键词:语音情感识别;生理信息引导的嗓音表示;四元数神经网络;自监督学习;幅度与相位;人形机器人
Keywords
speech emotion recognition physiology-informed vocal representation quaternion neural networks self-supervised learning amplitude and phase group delay and instantaneous frequency humanoid robot
Also referred to as: Physiology-Informed Vocal Spectrotemporal Representation Learning for Speech Emotion Recognition; physiology-informed speech emotion recognition; voice anatomy and physiology (VAP)-informed representation; Quaternion Spectrotemporal Encoder (QSE); Contrastive Projection and Alignment (CPA); PhysioSER-Full.
Research area map and related search terms
Field path (broad to narrow): artificial intelligence › machine learning › deep learning › speech and audio processing › spoken language technology › paralinguistics › affective computing › speech emotion recognition (SER) › physiology-informed SER › PhysioSER
Task
speech emotion recognition (SER) emotion recognition from speech vocal emotion recognition multilingual speech emotion recognition paralinguistic analysis affective computing
Method
phase-aware speech features group delay instantaneous frequency spectral flux quaternion convolution quaternion neural networks frozen self-supervised speech representations InfoNCE contrastive alignment interpretable speech emotion recognition parameter-efficient SER
Models used
wav2vec 2.0 HuBERT WavLM emotion2vec BEATs CLAP audio encoder
Datasets
CREMA-D RAVDESS EmoV-DB JL-Corpus EMNS RESD nEMO AESDD Emozionalmente Oréau MESD SUBESCO PAVOQUE CaFE
中文
语音情感识别 语音情绪识别 相位特征 群延迟 瞬时频率 四元数卷积 自监督语音表征 多语言语音情感识别 情感计算
Related areas
emotion recognition multimodal emotion recognition audio classification speech representation learning self-supervised learning foundation models for audio voice physiology vocal tract and glottal source signal processing time-frequency analysis short-time Fourier transform (STFT) hypercomplex neural networks interpretable machine learning cross-lingual transfer human-robot interaction social robotics mental health assessment
Applications
emotion-aware humanoid robots call-centre analytics psychological assessment support human-computer interaction voice assistants
Related benchmarks (not used in this paper)
IEMOCAP MSP-Podcast ESD TESS SAVEE EmoDB CASIA
中文 (扩展)
情感识别 语音情绪分类 副语言信息 语音表征学习 自监督学习 音频基础模型 发声生理 声道 声门 短时傅里叶变换 时频分析 超复数神经网络 可解释机器学习 跨语言迁移 人机交互 心理健康评估
Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.
Related papers
Affective speech
- DUET: Unified Dual-Space Emotion Control for Diffusion and Flow-Matching Driven Text-to-Speech (arXiv:2606.00066, 2026)
Other topics
- DSTCN: Exploiting dynamic spatio-temporal correlations for origin-destination demand prediction (2026)
- URMDet-SimFire: Trustworthy multimodal detection with uncertainty and reliability modeling for fire and smoke analytics (2026)
- S2CMEN: A Mutually Enhancement Network for Superpixel Segmentation and Classification of Hyperspectral Image (2025)
- MR-UFP: Enhancing urban flow prediction via mutual reinforcement with multi-scale regional information (2025)
- BiST-IF: Enhancing origin–destination flow prediction via bi-directional spatio-temporal inference and interconnected feature evolution (2025)
- Automatic visual recognition for leaf disease based on enhanced attention mechanism (2024)
- S3CFSL: Spatial-Spectral–Semantic Cross-Domain Few-Shot Learning for Hyperspectral Image Classification (2024)
- ST-FCL: Spatio-temporal fusion and contrastive learning for urban flow prediction (2023)
- MC-STL: Mask- and Contrast-Enhanced Spatio-Temporal Learning for Urban Flow Prediction (2023)
Cite this paper
@article{zhang2026physioser,
title = {Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition},
author = {Zhang, Xu and Cao, Longbing and Yang, Runze and Wu, Zhangkai},
journal = {arXiv preprint arXiv:2602.13259},
year = {2026},
eprint = {2602.13259},
archivePrefix= {arXiv},
primaryClass = {cs.SD},
doi = {10.48550/arXiv.2602.13259},
url = {https://arxiv.org/abs/2602.13259}
}Zhang, X., Cao, L., Yang, R., & Wu, Z. (2026). Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition. arXiv. https://doi.org/10.48550/arXiv.2602.13259