PhysioSER: Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition

Xu Zhang, Longbing Cao, Runze Yang, Zhangkai Wu

arXiv:2602.13259 (2026) · Free to read on arXiv (arXiv non-exclusive distribution license)

Also indexed in: Semantic Scholar · OpenAlex · dblp · Hugging Face · alphaXiv

TL;DR. PhysioSER complements a frozen self-supervised (SSL) backbone with a compact physiology-informed branch that encodes vocal amplitude and phase features with quaternion convolutions, for speech emotion recognition. With frozen WavLM on CREMA-D, it raises weighted accuracy from 69.69% to 75.20%.
Architecture diagram of PhysioSER: a general-purpose self-supervised (SSL) speech backbone branch and a voice anatomy and physiology (VAP)-informed vocal feature branch with quaternion spectrotemporal encoding, aligned by contrastive projection and fused by a shallow Transformer encoder for speech emotion classification.
Structure of the PhysioSER. The model consists of two parallel workflows: (a) an upper latent speech representation workflow to extract general features from the raw waveform using a frozen SSL backbone, followed by a latent representation transform; and (b) a lower vocal feature representation workflow to decompose vocal signals based on voice anatomy and physiology (VAP)-informed knowledge. In (b), the Hamilton structured Quaternion Spectrotemporal Encoder (QSE) embeds a physiology-aligned quartet—log-magnitude (M), log-magnitude rate (ρ) (spectral flux), instantaneous frequency (f_inst), and group delay (τ_g)—and models their structured, dynamic interactions. The two branches are separately summarized into utterance-level embeddings and aligned by the Contrastive Projection and Alignment (CPA) framework via projection heads with an InfoNCE objective. Finally, a shallow Transformer encoder fuses the aligned latent and vocal representations for SER. Figure 1 of arXiv:2602.13259 (Zhang et al., 2026).

Quick facts

Method namePhysioSER
AuthorsXu Zhang, Longbing Cao, Runze Yang, Zhangkai Wu
Published inarXiv:2602.13259 (2026)
arXivarXiv:2602.13259 (DOI 10.48550/arXiv.2602.13259)
TaskUtterance-level speech emotion recognition (SER), motivated by humanoid-robot uses such as social interaction and robotic psychological diagnosis
ArchitectureCompact plug-and-play physiology-informed branch running in parallel with a frozen self-supervised (SSL) backbone; CPA aligns the two, and a shallow attention fusion head classifies
InputsRaw waveform for the SSL branch; for the physiology branch, an STFT-derived quartet: log-magnitude M, log-magnitude rate ρ (spectral flux), instantaneous frequency f_inst, group delay τ_g
EncoderQuaternion Spectrotemporal Encoder (QSE): each bin embedded as Q = M + iρ + j f_inst + k τ_g, then Hamilton-product quaternion convolutions, quaternion batch normalization and radial qReLU
Alignment and fusionContrastive Projection and Alignment (CPA): projection heads and a symmetric InfoNCE loss, with the SSL model frozen; then a shallow Transformer encoder and compact classifier
SSL backbonesSix, all frozen: Wav2vec2, HuBERT, WavLM, Emotion2Vec, BEATs and the CLAP audio encoder, covering four pre-training paradigms
Datasets14 public acted or simulated emotional speech datasets: 5 English (CREMA-D, EMNS, JLCorpus, EmoV-DB, RAVDESS) and 9 non-English (RESD, nEMO, AESDD, Emozionalmente, Oréau, MESD, SUBESCO, PAVOQUE, CaFE)
Metrics and protocolWeighted accuracy (WA), unweighted accuracy (UA), macro-F1; 7:1:2 train/validation/test split; mean of five random seeds; every Δ is a relative % change
BaselinesBackbone-only (same frozen backbone, masked attention pooling and classifier, no physiology branch); on CREMA-D also fully fine-tuned WavLM and a real-valued CNN variant
Headline resultCREMA-D with frozen WavLM: WA 69.69% → 75.20% (+7.91% relative) with 2.03M trainable parameters; fully fine-tuned WavLM reaches 73.66% with 94.61M
DeploymentReal-time on the Ameca humanoid robot: vocal emotion is mapped to Ameca's facial expressions without automatic speech recognition; the paper links a demo video
BibTeX keyzhang2026physioser

Main results

SettingResultCompared with
CREMA-D, frozen WavLM backboneWA = 75.20% (UA 75.62%, macro-F1 75.63%)
“On dataset CREMA-D with a frozen WavLM [8], WA increases from 69.69% to 75.20%.”
Backbone-only frozen WavLM: WA 69.69% (+7.91% relative)
CREMA-D, parameter efficiencyWA = 75.20% with 2.03M trainable parameters (frozen WavLM)
“PhysioSER achieves a superior 75.20% WA with only 2.03M trainable parameters (2.1% of the baseline)”
Fully fine-tuned WavLM: WA 73.66% with 94.61M parameters
CREMA-D ablation: quaternion QSE vs real-valued CNN of comparable sizeWA = 75.20% (quaternion) vs 72.78% (Normal CNN)
“While Normal CNN improves over WavLM Only, it trails our quaternion model by 3.22% WA.”
Real-valued CNN is 3.22% lower (relative)
CREMA-D ablation: feature subsetsMagnitude-only WA = 74.60% (−0.80% relative); Phase-only WA = 70.50% (−6.25% relative)
“removing phase information (Magnitude-only) causes a minor 0.80% drop, whereas discarding magnitude (Phase-only) leads to a substantial 6.25% decline.”
Full quartet: WA 75.20%
JLCorpus (10 classes), Wav2vec2 backboneWA = 66.88%
“the model yields large relative improvements for weaker backbones (e.g., +136.07% WA on JLCorpus with Wav2vec2)”
Backbone-only: WA 28.33% (+136.07% relative, the largest WA gain in Tables I-II)
CaFE (French, 7 classes), Wav2Vec2 backboneUA = 35.45% (−6.64% relative); WA = 35.83% (+11.65%); F1 = 35.86% (+12.77%)
“Wav2Vec2 32.09 35.83 +11.65 37.97 35.45 −6.64 31.80 35.86 +12.77”
Backbone-only: UA 37.97%, WA 32.09%, F1 31.80%; the only negative Δ among the 252 reported in Tables I-II

Quoted text is verbatim from the paper.

Key points

Abstract

Speech emotion recognition (SER) is essential for humanoid robot tasks such as social robotic interactions and robotic psychological diagnosis, where interpretable and efficient models are critical for safety and performance. Existing deep models trained on large datasets remain largely uninterpretable, often insufficiently modeling underlying emotional acoustic signals and failing to capture and analyze the core physiology of emotional vocal behaviors. Physiological research on human voices shows that the dynamics of vocal amplitude and phase correlate with emotions through the vocal tract filter and the glottal source. However, most existing deep models solely involve amplitude but fail to couple the physiological features of and between amplitude and phase. Here, we propose PhysioSER, a physiology-informed vocal spectrotemporal representation learning method, to address these issues with a compact, plug-and-play design. PhysioSER constructs amplitude and phase views informed by voice anatomy and physiology (VAP) to complement SSL models for SER. This VAP-informed framework incorporates two parallel workflows: a vocal feature representation branch to decompose vocal signals based on VAP, embed them into a quaternion field, and use Hamilton-structured quaternion convolutions for modeling their dynamic interactions; and a latent representation branch based on a frozen SSL backbone. Then, utterance-level features from both workflows are aligned by a Contrastive Projection and Alignment framework, followed by a shallow attention fusion head for SER classification. PhysioSER is shown to be interpretable and efficient for SER through extensive evaluations across 14 datasets, 10 languages, and 6 backbones, and its practical efficacy is validated by real-time deployment on a humanoid robotic platform.

Frequently asked questions

What is PhysioSER?

PhysioSER is a physiology-informed method for speech emotion recognition. It adds a compact branch that encodes vocal amplitude and phase features with Hamilton-structured quaternion convolutions, aligns it with a frozen self-supervised (SSL) backbone through Contrastive Projection and Alignment (CPA), and classifies the emotion with a shallow attention fusion head.

Why model phase as well as amplitude?

Physiological research shows that the dynamics of vocal amplitude and phase correlate with emotion through the vocal tract filter and the glottal source. Yet most existing deep SER models use amplitude only. Methods that do use phase often concatenate it with amplitude as a separate feature or treat the two as independent channels, instead of modelling the physiology-informed coupling between them.

Which acoustic features does PhysioSER extract?

It extracts four features from the short-time Fourier transform: log-magnitude M, log-magnitude rate ρ (spectral flux), instantaneous frequency f_inst and group delay τ_g. The authors describe them as spanning vocal-tract timbre, temporal energy modulation, laryngeal oscillation rate and phase dispersion. Each time-frequency bin is embedded as the quaternion M + iρ + j f_inst + k τ_g.

How much does PhysioSER improve a frozen WavLM, and at what cost?

On CREMA-D, weighted accuracy rises from 69.69% (frozen WavLM plus classifier) to 75.20%, a 7.91% relative gain, with 2.03M trainable parameters. Fully fine-tuning WavLM uses 94.61M parameters and reaches 73.66% WA.

Which backbones and datasets was PhysioSER evaluated on?

It was tested with six frozen backbones (Wav2vec2, HuBERT, WavLM, Emotion2Vec, BEATs and the CLAP audio encoder) on 14 public emotional speech datasets, 5 English and 9 non-English. WA and macro-F1 improve over the Backbone-only baseline for every backbone-dataset pair; the only drop in the main result tables (Tables I-II) is UA on CaFE with Wav2Vec2 (−6.64% relative).

Which robot was PhysioSER deployed on?

PhysioSER was deployed on the Ameca humanoid robot. It processes continuous audio in Ameca's perception-action loop and maps vocal emotion directly to Ameca's facial expressions, without automatic speech recognition or linguistic analysis. The paper links a demo video.

How do I cite PhysioSER?

Cite it as: Zhang, X., Cao, L., Yang, R., & Wu, Z. (2026). Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition. arXiv. https://doi.org/10.48550/arXiv.2602.13259 BibTeX is available at https://codezx6.github.io/papers/physioser.html#cite and in https://codezx6.github.io/publications.bib.

中文摘要

学习生理信息引导的嗓音时频表示用于语音情感识别(PhysioSER)

PhysioSER 是生理信息引导的语音情感识别方法:基于嗓音解剖与生理(VAP)提取对数幅度、谱通量、瞬时频率与群时延,嵌入四元数域并用 Hamilton 结构四元数卷积建模交互,经对比投影与对齐(CPA)与冻结自监督骨干的特征对齐后融合分类。CREMA-D 上以冻结 WavLM 为骨干时,加权准确率由 69.69% 升至 75.20%;已在 Ameca 人形机器人上实时部署。

关键词:语音情感识别;生理信息引导的嗓音表示;四元数神经网络;自监督学习;幅度与相位;人形机器人

Keywords

speech emotion recognition physiology-informed vocal representation quaternion neural networks self-supervised learning amplitude and phase group delay and instantaneous frequency humanoid robot

Also referred to as: Physiology-Informed Vocal Spectrotemporal Representation Learning for Speech Emotion Recognition; physiology-informed speech emotion recognition; voice anatomy and physiology (VAP)-informed representation; Quaternion Spectrotemporal Encoder (QSE); Contrastive Projection and Alignment (CPA); PhysioSER-Full.

Research area map and related search terms

Field path (broad to narrow): artificial intelligence › machine learning › deep learning › speech and audio processing › spoken language technology › paralinguistics › affective computing › speech emotion recognition (SER) › physiology-informed SER › PhysioSER

Task

speech emotion recognition (SER) emotion recognition from speech vocal emotion recognition multilingual speech emotion recognition paralinguistic analysis affective computing

Method

phase-aware speech features group delay instantaneous frequency spectral flux quaternion convolution quaternion neural networks frozen self-supervised speech representations InfoNCE contrastive alignment interpretable speech emotion recognition parameter-efficient SER

Models used

wav2vec 2.0 HuBERT WavLM emotion2vec BEATs CLAP audio encoder

Datasets

CREMA-D RAVDESS EmoV-DB JL-Corpus EMNS RESD nEMO AESDD Emozionalmente Oréau MESD SUBESCO PAVOQUE CaFE

中文

语音情感识别 语音情绪识别 相位特征 群延迟 瞬时频率 四元数卷积 自监督语音表征 多语言语音情感识别 情感计算

Related areas

emotion recognition multimodal emotion recognition audio classification speech representation learning self-supervised learning foundation models for audio voice physiology vocal tract and glottal source signal processing time-frequency analysis short-time Fourier transform (STFT) hypercomplex neural networks interpretable machine learning cross-lingual transfer human-robot interaction social robotics mental health assessment

Applications

emotion-aware humanoid robots call-centre analytics psychological assessment support human-computer interaction voice assistants

Related benchmarks (not used in this paper)

IEMOCAP MSP-Podcast ESD TESS SAVEE EmoDB CASIA

中文 (扩展)

情感识别 语音情绪分类 副语言信息 语音表征学习 自监督学习 音频基础模型 发声生理 声道 声门 短时傅里叶变换 时频分析 超复数神经网络 可解释机器学习 跨语言迁移 人机交互 心理健康评估

Terms are grouped by role. "Related areas", "Applications" and "Related benchmarks (not used in this paper)" describe the surrounding field, not results of this paper. See the site-wide research area map.

Related papers

Affective speech

Other topics

Cite this paper

@article{zhang2026physioser,
  title        = {Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition},
  author       = {Zhang, Xu and Cao, Longbing and Yang, Runze and Wu, Zhangkai},
  journal      = {arXiv preprint arXiv:2602.13259},
  year         = {2026},
  eprint       = {2602.13259},
  archivePrefix= {arXiv},
  primaryClass = {cs.SD},
  doi          = {10.48550/arXiv.2602.13259},
  url          = {https://arxiv.org/abs/2602.13259}
}
Zhang, X., Cao, L., Yang, R., & Wu, Z. (2026). Learning Physiology-Informed Vocal Spectrotemporal Representations for Speech Emotion Recognition. arXiv. https://doi.org/10.48550/arXiv.2602.13259