WSEAS Transactions on Computer Research
Print ISSN: 1991-8755, E-ISSN: 2415-1521
Volume 13, 2025
Frequency-based Pre-processing for Reducing the Effect of Sex in Machine-Learning Voice Analysis for Alcohol Intoxication Detection
Authors: ,
Search Articles
Abstract: In machine learning-based voice analysis the sex of the subjects is the most important covariate, as vocal differences between sexes lead to significant discrepancies in acoustic features and model generalization. We propose a signal processing pipeline that normalizes audio recordings of female subjects to exhibit characteristics more similar to those of males, or vice versa. We computed ratios of fundamental frequency, formants, breathiness, and pitch variability from a baseline dataset and used them to build a pipeline based on the decomposition of a signal into fundamental frequency contour, spectral envelope and aperiodicity. We tested on a classification task for the detection of intoxicated versus sober subjects and observed a systematic increase in accuracy when using our processing on the dataset. Accuracies increase up to 5.5% across all three considered algorithms: Support Vector Machine, Random Forest, and Neural Network, with the latter model applied on processed data achieving a state-of-the-art accuracy of 81.79%. Summary: This paper presents a frequency-based audio preprocessing method to reduce gender-related acoustic variability in machine learning-based voice analysis, particularly for detecting alcohol intoxication. The authors developed a pipeline to normalize female voice characteristics toward male standards, addressing differences in fundamental frequency, formants, breathiness, and pitch variability. Applied to a dataset of sober and intoxicated speech samples, this preprocessing significantly improved classification accuracy for intoxication detection, achieving up to a 5.5% accuracy increase, with a neural network reaching a state-of-the-art accuracy of 81.79%. The pipeline effectively minimized sex-based acoustic variability without introducing notable perceptual artifacts, thus enhancing generalization and model reliability across genders in voice-based machine learning tasks.
Pages: 660-668
DOI: 10.37394/232018.2025.13.59