MARATTO

article

S-PHiNe: Physics-Informed Multichannel Speech Enhancement Using Spectro-Spatial Fusion for Low-SNR Conditions

Abstract

Multichannel speech enhancement (MCSE) under adverse noise conditions remains a major challenge. This study proposes S-PHiNe, a neural physics-informed framework integrating deep graph convolutional network (DGCN)-based spatial embedding for steering vector estimation, a complex beamforming neural network (CBNN) for spectral mask estimation, and physics-informed MVDR beamforming within end-to-end physics-constrained training. The model outperforms state-of-the-art (SOTA) neural postfilter and beamforming methods. Evaluations demonstrate how masking, spatial embeddings, and physics-informed constraints jointly improve speech intelligibility, perceptual quality, and multichannel generalization, establishing S-PHiNe as a robust solution for highly noisy multichannel environments.

Research topics

  • Speech and Audio Processing
  • Emotion and Mood Recognition
  • Speech Recognition and Synthesis

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/icassp55912.2026.11464982

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.