Robust speech denoising using a parallel target-field non-causal WaveNet under stationary and non-stationary noise
Abstract
The presence of stationary or non-stationary noise substantially impairs speech intelligibility and its perceptual quality in present-day communication systems. Many existing spectral-domain methods like Wiener filtering and spectral subtraction rely on strong statistical assumptions about complex-valued data, while state-of-the-art neural models such as Conv-TasNet, DCCRN, and MetricGAN+ achieve remarkable performance gains for speech enhancement, but tend to suffer from heavy computational complexity and inference latency due to their complex architecture. In this paper, we present a new supervised non-causal WaveNet with parallel target-field prediction for efficient end-to-end speech denoising. The proposed model captures both past and future temporal context by enabling symmetric receptive fields and removing autoregressive dependencies. Traditional autoregressive WaveNet models generate samples one at a time, which require considerable redundant convolution operations due to the repetitive nature of their structures and tasks, while the presented method is capable of predicting a target field of samples in a single forward pass, allowing for parallel inference. We validate the proposed model on the NSDTSEA dataset for input SNR levels of 2.5 dB to 17.5 dB and under stationary as well as non-stationary noise types. We showcase experimental results that prove that the proposed framework consistently outperforms strong classical baselines, with over 1.22 dB gain in output SNR and 46% reduction in Mel Cepstral Distortion (MCD) under non-stationary noise. Moreover, when compared to state-of-the-art deep learning models such as Conv-TasNet, DCCRN, and MetricGAN+, the proposed method achieves a viable trade-off between performance, robustness, and computational efficiency.
Copyright (c) 2026 Jella Sandhya, Shafiq Ur Rehman, Shaik Mazhar Hussain

This work is licensed under a Creative Commons Attribution 4.0 International License.
References
[1]Hu Y, Loizou PC. Evaluation of Objective Quality Measures for Speech Enhancement. IEEE Transactions on Audio, Speech, and Language Processing. 2008; 16(1): 229–238. doi: 10.1109/TASL.2007.911054
[2]Boll S. Suppression of acoustic noise in speech using spectral subtraction. IEEE Transactions on Acoustics, Speech, and Signal Processing. 1979; 27(2): 113–120. doi: 10.1109/TASSP.1979.1163209
[3]Wang Y, Wang D. Towards Scaling Up Classification-Based Speech Separation. IEEE Transactions on Audio, Speech, and Language Processing. 2013; 21(7): 1381–1390. doi: 10.1109/TASL.2013.2250961
[4]Pascual S, Bonafonte A, Serrà J. SEGAN: Speech Enhancement Generative Adversarial Network. In: Proceedings of the Interspeech 2017; 20–24 August 2017; Stockholm, Sweden. pp. 3642–3646. doi: 10.21437/Interspeech.2017-1428
[5]Luo Y, Mesgarani N. Conv-TasNet: Surpassing Ideal Time–Frequency Magnitude Masking for Speech Separation. IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2019; 27(8): 1256–1266. doi: 10.1109/TASLP.2019.2915167
[6]Hu Y, Liu Y, Lv S, et al. DCCRN: Deep Complex Convolution Recurrent Network for Phase-Aware Speech Enhancement. In: Proceedings of the Interspeech 2020; 25–29 October 2020; Shanghai, China. pp. 2472–2476. doi: 10.21437/Interspeech.2020-2537
[7]Pandey A, Wang D. A New Framework for CNN-Based Speech Enhancement in the Time Domain. IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2019; 27(7): 1179–1188. doi: 10.1109/TASLP.2019.2913512
[8]van den Oord A, Dieleman S, Zen H, et al. WaveNet: A Generative Model for Raw Audio. arXiv preprint. 2016. doi: 10.48550/arXiv.1609.03499
[9]Qian K, Zhang Y, Chang S, et al. Speech Enhancement Using Bayesian Wavenet. In: Proceedings of the Interspeech 2017; 20–24 August 2017; Stockholm, Sweden. pp. 2013–2017. doi: 10.21437/Interspeech.2017-1672
[10]Rethage D, Pons J, Serra X. A Wavenet for Speech Denoising. In: Proceedings of the 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 15–20 April 2018; Calgary, AB, Canada. pp. 5069–5073. doi: 10.1109/ICASSP.2018.8462417
[11]Fu SW, Liao CF, Tsao Y, et al. MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement. arXiv preprint. 2019. doi: 10.48550/ARXIV.1905.04874
[12]Oppenheim AV, Schafer RW. Digital Signal Processing. Prentice-Hall; 1975.
[13]Liu C, Wang L, Dang J. Masking based Spectral Feature Enhancement for Robust Automatic Speech Recognition. In: Proceedings of the 2020 IEEE International Conference on Artificial Intelligence and Computer Applications (ICAICA); 27–29 June 2020; Dalian, China. pp. 287–291. doi: 10.1109/ICAICA50127.2020.9181915
[14]Défossez A, Usunier N, Bottou L, et al. Music Source Separation in the Waveform Domain. arXiv preprint. 2019. doi: 10.48550/ARXIV.1911.13254
[15]Lin J, Van Wijngaarden AJDL, Wang KC, et al. Speech Enhancement Using Multi-Stage Self-Attentive Temporal Convolutional Networks. IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2021; 29: 3440–3450. doi: 10.1109/TASLP.2021.3125143
[16]Choi HS, Kim JH, Huh J, et al. Phase-aware Speech Enhancement with Deep Complex U-Net. arXiv preprint. 2019. doi: 10.48550/ARXIV.1903.03107
[17]Lv S, Hu Y, Zhang S, et al. DCCRN+: Channel-wise Subband DCCRN with SNR Estimation for Speech Enhancement. arXiv preprint. 2021. doi: 10.48550/arXiv.2106.08672
[18]Bishop CM. Pattern Recognition and Machine Learning. Springer; 2006.
[19]Fu SW, Yu C, Hsieh TA, et al. MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement. In: Proceedings of the Interspeech 2021; 30 August 30–3 September 2021; Brno, Czech Republic. pp. 201–205. doi: 10.21437/Interspeech.2021-599
[20]Chen J, Mao Q, Liu D. Dual-Path Transformer Network: Direct Context-Aware Modeling for End-to-End Monaural Speech Separation. In: Proceedings of the Interspeech 2020; 25–29 October 2020; Shanghai, China. pp. 2642–2646. doi: 10.21437/Interspeech.2020-2205
[21]Gulati A, Qin J, Chiu CC, et al. Conformer: Convolution-augmented Transformer for Speech Recognition. In: Proceedings of the Interspeech 2020; 25–29 October 2020; Shanghai, China. pp. 5036–5040. doi: 10.21437/Interspeech.2020-3015
[22]Luo Y, Chen Z, Yoshioka T. Dual-Path RNN: Efficient Long Sequence Modeling for Time-Domain Single-Channel Speech Separation. In: Proceedings of the ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP); 4–8 May 2020; Barcelona, Spain. pp. 46–50. doi: 10.1109/ICASSP40776.2020.9054266
[23]Nachmani E, Adi Y, Wolf L. Voice Separation with an Unknown Number of Multiple Speakers. arXiv preprint. 2020. doi: 10.48550/arXiv.2003.01531
[24]Kong Z, Ping W, Huang J, et al. DiffWave: A Versatile Diffusion Model for Audio Synthesis. arXiv preprint. 2020. doi: 10.48550/ARXIV.2009.09761
[25]Richter J, Welker S, Lemercier JM, et al. Speech Enhancement and Dereverberation with Diffusion-Based Generative Models. IEEE/ACM Transactions on Audio, Speech, and Language Processing. 2023; 31: 2351–2364. doi: 10.1109/TASLP.2023.3285241
[26]Welker S, Richter J, Gerkmann T. Speech Enhancement with Score-Based Generative Models in the Complex STFT Domain. In: Proceedings of the Interspeech 2022; 18–22 September 2022; Incheon, South Korea. pp. 2928–2932. doi: 10.21437/Interspeech.2022-10653
[27]Emura S. Estimation of Output SI-SDR Solely from Enhanced Speech Signals in Diffusion-Based Generative Speech Enhancement Method. In: Proceedings of the 2024 32nd European Signal Processing Conference (EUSIPCO); 26-30 August 2024; Lyon, France. pp. 236–240. doi: 10.23919/EUSIPCO63174.2024.10714970
[28]van den Oord A, Li Y, Babuschkin I, et al. Parallel WaveNet: Fast High-Fidelity Speech Synthesis. arXiv preprint. 2017. doi: 10.48550/ARXIV.1711.10433
[29]Braun S, Tashev I. Data augmentation and loss normalization for deep noise suppression. arXiv preprint. 2020. doi: 10.48550/ARXIV.2008.06412
[30]Défossez A. Hybrid Spectrogram and Waveform Source Separation. arXiv preprint. 2022. doi: 10.48550/arXiv.2111.03600
[31]Kubichek R. Mel-cepstral distance measure for objective speech quality assessment. In: Proceedings of IEEE Pacific Rim Conference on Communications Computers and Signal Processing; 19–21 May 1993; Victoria, BC, Canada. pp. 125–128. doi: 10.1109/PACRIM.1993.407206




