<doi_batch xmlns="http://www.crossref.org/schema/4.4.0" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" version="4.4.0"><head><doi_batch_id>06ad22e8-1511-4897-a1e7-a8976f180323</doi_batch_id><timestamp>20251110115604612</timestamp><depositor><depositor_name>wseas:wseas</depositor_name><email_address>mdt@crossref.org</email_address></depositor><registrant>MDT Deposit</registrant></head><body><journal><journal_metadata language="en"><full_title>WSEAS TRANSACTIONS ON COMPUTER RESEARCH</full_title><issn media_type="electronic">2415-1521</issn><issn media_type="print">1991-8755</issn><archive_locations><archive name="Portico"/></archive_locations><doi_data><doi>10.37394/232018</doi><resource>http://wseas.org/wseas/cms.action?id=13372</resource></doi_data></journal_metadata><journal_issue><publication_date media_type="online"><month>1</month><day>10</day><year>2025</year></publication_date><publication_date media_type="print"><month>1</month><day>10</day><year>2025</year></publication_date><journal_volume><volume>13</volume><doi_data><doi>10.37394/232018.2025.13</doi><resource>https://wseas.com/journals/cr/2025.php</resource></doi_data></journal_volume></journal_issue><journal_article language="en"><titles><title>Frequency-based Pre-processing for Reducing the Effect of Sex in Machine-Learning Voice Analysis for Alcohol Intoxication Detection</title></titles><contributors><person_name sequence="first" contributor_role="author"><given_name>Cesarini</given_name><surname>Valerio</surname><affiliation>Department of Electronic Engineering, University of Rome Tor Vergata, Via del Politecnico, 1, 00100 Rome, ITALY</affiliation></person_name><person_name sequence="additional" contributor_role="author"><given_name>Costantini</given_name><surname>Giovanni</surname><affiliation>Department of Electronic Engineering, University of Rome Tor Vergata, Via del Politecnico, 1, 00100 Rome, ITALY</affiliation></person_name></contributors><jats:abstract xmlns:jats="http://www.ncbi.nlm.nih.gov/JATS1"><jats:p>In machine learning-based voice analysis the sex of the subjects is the most important covariate, as vocal differences between sexes lead to significant discrepancies in acoustic features and model generalization. We propose a signal processing pipeline that normalizes audio recordings of female subjects to exhibit characteristics more similar to those of males, or vice versa. We computed ratios of fundamental frequency, formants, breathiness, and pitch variability from a baseline dataset and used them to build a pipeline based on the decomposition of a signal into fundamental frequency contour, spectral envelope and aperiodicity. We tested on a classification task for the detection of intoxicated versus sober subjects and observed a systematic increase in accuracy when using our processing on the dataset. Accuracies increase up to 5.5% across all three considered algorithms: Support Vector Machine, Random Forest, and Neural Network, with the latter model applied on processed data achieving a state-of-the-art accuracy of 81.79%. Summary: This paper presents a frequency-based audio preprocessing method to reduce gender-related acoustic variability in machine learning-based voice analysis, particularly for detecting alcohol intoxication. The authors developed a pipeline to normalize female voice characteristics toward male standards, addressing differences in fundamental frequency, formants, breathiness, and pitch variability. Applied to a dataset of sober and intoxicated speech samples, this preprocessing significantly improved classification accuracy for intoxication detection, achieving up to a 5.5% accuracy increase, with a neural network reaching a state-of-the-art accuracy of 81.79%. The pipeline effectively minimized sex-based acoustic variability without introducing notable perceptual artifacts, thus enhancing generalization and model reliability across genders in voice-based machine learning tasks.</jats:p></jats:abstract><publication_date media_type="online"><month>11</month><day>10</day><year>2025</year></publication_date><publication_date media_type="print"><month>11</month><day>10</day><year>2025</year></publication_date><pages><first_page>660</first_page><last_page>668</last_page></pages><publisher_item><item_number item_number_type="article_number">59</item_number></publisher_item><ai:program xmlns:ai="http://www.crossref.org/AccessIndicators.xsd" name="AccessIndicators"><ai:free_to_read start_date="2025-11-10"/><ai:license_ref applies_to="am" start_date="2025-11-10">https://wseas.com/journals/cr/2025/b205118-014(2025).pdf</ai:license_ref></ai:program><archive_locations><archive name="Portico"/></archive_locations><doi_data><doi>10.37394/232018.2025.13.59</doi><resource>https://wseas.com/journals/cr/2025/b205118-014(2025).pdf</resource></doi_data><citation_list><citation key="ref0"><doi>10.3390/electronics11233935</doi><unstructured_citation>J. L. Bautista, Y. K. Lee, H. S. Shin, «Speech Emotion Recognition Based on Parallel CNNAttention Networks with Multi-Fold Data Augmentation», Electronics, vol. 11, fasc. 23, Art. fasc. 23, Jan 2022, doi: 10.3390/electronics11233935. </unstructured_citation></citation><citation key="ref1"><doi>10.3390/s23073461</doi><unstructured_citation>G. Costantini, V. Cesarini, E. Brenna, «HighLevel CNN and Machine Learning Methods for Speaker Recognition», Sensors, vol. 23, fasc. 7, Art. fasc. 7, Jan 2023, doi: 10.3390/s23073461. </unstructured_citation></citation><citation key="ref2"><doi>10.3390/app142311446</doi><unstructured_citation>Cesarini, V.; Costantini, G. Reverb and Noise as Real-World Effects in Speech Recognition Models: A Study and a Proposal of a Feature Set. Appl. Sci. 2024, 14, 11446. https://doi.org/10.3390/app142311446. </unstructured_citation></citation><citation key="ref3"><doi>10.1016/j.jvoice.2021.01.018</doi><unstructured_citation>M. Alves, G. Silva, B. C. Bispo, M. E. Dajer, P. M. Rodrigues, «Voice Disorders Detection Through Multiband Cepstral Features of Sustained Vowel», J. Voice, vol. 37, fasc. 3, pp. 322–331, May 2023, doi: 10.1016/j.jvoice.2021.01.018. </unstructured_citation></citation><citation key="ref4"><doi>10.1007/bf00994018</doi><unstructured_citation>C. Cortes, V. Vapnik, «Support-vector networks», Mach. Learn., vol. 20, fasc. 3, pp. 273–297, Sept. 1995, doi: 10.1007/BF00994018. </unstructured_citation></citation><citation key="ref5"><unstructured_citation>G. Fant, “Acoustic Theory of Speech Production”. Book, Ed. Walter de Gruyter, 1970. </unstructured_citation></citation><citation key="ref6"><doi>10.1016/j.protcy.2013.12.124</doi><unstructured_citation>J. P. Teixeira, C. Oliveira, C. Lopes, «Vocal Acoustic Analysis – Jitter, Shimmer and HNR Parameters», Procedia Technol., vol. 9, pp. 1112–1122, 2013, doi: 10.1016/j.protcy.2013.12.124. </unstructured_citation></citation><citation key="ref7"><unstructured_citation>B. P. Bogert, «The quefrency alanysis of time series for echoes ; Cepstrum, pseudo autocovariance, cross-cepstrum and saphe cracking», Time Ser. Anal., pp. 209–243, 1963. </unstructured_citation></citation><citation key="ref8"><doi>10.1007/978-981-15-9647-6_39</doi><unstructured_citation>J. Kaur e A. Kumar, «Speech Emotion Recognition Using CNN, k-NN, MLP and Random Forest», 2021, pp. 499–509. doi: 10.1007/978-981-15-9647-6_39. </unstructured_citation></citation><citation key="ref9"><doi>10.3390/app13158562</doi><unstructured_citation>Cesarini, V.; Saggio, G.; Suppa, A.; Asci, F.; Pisani, A.; Calculli, A.; Fayad, R.; HajjHassan, M.; Costantini, G. Voice Disorder Multi-Class Classification for the Distinction of Parkinson’s Disease and Adductor Spasmodic Dysphonia. Appl. Sci. 2023, 13, 8562. https://doi.org/10.3390/app13158562. </unstructured_citation></citation><citation key="ref10"><doi>10.1109/icassp.2002.1005790</doi><unstructured_citation>S. Umesh, S. V. Bharath Kumar, M. K. Vinay, R. Sharma and R. Sinha, "A simple approach to non-uniform vowel normalization," 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing, Orlando, FL, USA, 2002, pp. I517-I-520, doi: 10.1109/ICASSP.2002.5743768. </unstructured_citation></citation><citation key="ref11"><doi>10.1044/jshr.3704.769</doi><unstructured_citation>J. Hillenbrand, R. A. Cleveland, e R. L. Erickson, «Acoustic Correlates of Breathy Vocal Quality», J. Speech Lang. Hear. Res., vol. 37, fasc. 4, pp. 769–778, Aug. 1994, doi: 10.1044/jshr.3704.769. </unstructured_citation></citation><citation key="ref12"><doi>10.1016/0271-5309(94)00011-z</doi><unstructured_citation>C. Henton, «Pitch dynamism in female and male speech», Lang. Commun., vol. 15, fasc. 1, pp. 43–61, Jan 1995, doi: 10.1016/0271- 5309(94)00011-Z. </unstructured_citation></citation><citation key="ref13"><doi>10.1371/journal.pone.0156870</doi><unstructured_citation>G. Pino Escobar, J. Terry, B. P. Kriengwatana, P. Escudero, «Speech normalization across speaker, sex and accent variation is handled similarly by listeners of different language backgrounds: Australasian International Conference on Speech Science and Technology», Proc. Sixt. Australas. Int. Conf. Speech Sci. Technol. 6-9 Dec. 2016 Parramatta Aust., pp. 161–164, 2016. </unstructured_citation></citation><citation key="ref14"><doi>10.1002/9781119184096.ch6</doi><unstructured_citation>K. Johnson, M. J. Sjerps, «Speaker Normalization in Speech Perception», in The Handbook of Speech Perception, John Wiley &amp; Sons, Ltd, 2021, pp. 145–176. doi: 10.1002/9781119184096.ch6. </unstructured_citation></citation><citation key="ref15"><unstructured_citation>J. D. Miller, «Auditory-perceptual interpretation of the vowel», J. Acoust. Soc. Am., vol. 85, fasc. 5, pp. 2114–2134, 1989, doi: 10.1121/1.397862. </unstructured_citation></citation><citation key="ref16"><doi>10.1017/s0954394509990160</doi><unstructured_citation>A. H. Fabricius, D. Watt, e D. E. Johnson, «A comparison of three speaker-intrinsic vowel formant frequency normalization algorithms for sociophonetics», Lang. Var. Change, vol. 21, fasc. 3, pp. 413–435, Oct. 2009, doi: 10.1017/S0954394509990160. </unstructured_citation></citation><citation key="ref17"><doi>10.21437/interspeech.2010-117</doi><unstructured_citation>T. Schaaf, F. Metze, «Analysis of gender normalization using MLP and VTLN features», presented at Proc. Interspeech 2010, pp. 306–309. doi: 10.21437/Interspeech.2010- 117. </unstructured_citation></citation><citation key="ref18"><doi>10.21437/interspeech.2008-414</doi><unstructured_citation>P. Fousek, L. Lamel, J.-L. Gauvain, «Transcribing broadcast data using MLP features», presented at Proc. Interspeech 2008, pp. 1433–1436. doi: 10.21437/Interspeech.2008-414. </unstructured_citation></citation><citation key="ref19"><unstructured_citation>L.-M. Zhang, Y. Li, Y.-T. Zhang, G. W. Ng, Y.-B. Leau, H. Yan, «A Deep Learning Method Using Gender-Specific Features for Emotion Recognition», Sensors, vol. 23, fasc. 3, Art. fasc. 3, Jan 2023, doi: 10.3390/s23031355. </unstructured_citation></citation><citation key="ref20"><doi>10.3390/electronics11101594</doi><unstructured_citation>D. Rizhinashvili, A. H. Sham, G. Anbarjafari, «Gender Neutralisation for Unbiased Speech Synthesising», Electronics, vol. 11, fasc. 10, Art. fasc. 10, Jan 2022, doi: 10.3390/electronics11101594. </unstructured_citation></citation><citation key="ref21"><unstructured_citation>P. Boersma, D. Weenink, «PRAAT, a system for doing phonetics by computer», Glot Int., vol. 5, pp. 341–345, gen. 2001. </unstructured_citation></citation><citation key="ref22"><unstructured_citation>«Gender Recognition by Voice(original)». Accessed on: March 15, 2025. [Online]. https://www.kaggle.com/datasets/murtadhanaj im/gender-recognition-by-voiceoriginal </unstructured_citation></citation><citation key="ref23"><doi>10.1007/s10579-011-9139-y</doi><unstructured_citation>F. Schiel, C. Heinrich, S. Barfüßer, T. Gilg, «ALC: Alcohol Language Corpus», in Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC’08), Marrakech, Morocco: European Language Resources Association (ELRA), May 2008. </unstructured_citation></citation><citation key="ref24"><unstructured_citation>A. de Cheveigne, H. Kawahara, «YIN, a fundamental frequency estimator for speech and musica)», J Acoust Soc Am, vol. 111, fasc. 4, 2002. </unstructured_citation></citation><citation key="ref25"><doi>10.1109/tassp.1980.1163489</doi><unstructured_citation>A. Gray, D. Wong, «The Burg algorithm for LPC speech analysis/Synthesis», IEEE Trans. Acoust. Speech Signal Process., vol. 28, fasc. 6, pp. 609–615, dic. 1980, doi: 10.1109/TASSP.1980.1163489. </unstructured_citation></citation><citation key="ref26"><doi>10.1587/transinf.2015edp7457</doi><unstructured_citation>M. Morise, F. Yokomori, K. Ozawa, «WORLD: A Vocoder-Based High-Quality Speech Synthesis System for Real-Time Applications», IEICE Trans. Inf. Syst., vol. E99.D, fasc. 7, pp. 1877–1884, 2016, doi: 10.1587/transinf.2015EDP7457. </unstructured_citation></citation><citation key="ref27"><doi>10.21437/interspeech.2017-68</doi><unstructured_citation>M. Morise, «Harvest: A High-Performance Fundamental Frequency Estimator from Speech Signals», in Interspeech 2017, ISCA, ago. 2017, pp. 2321–2325. doi: 10.21437/Interspeech.2017-68. </unstructured_citation></citation><citation key="ref28"><doi>10.1016/j.specom.2014.09.003</doi><unstructured_citation>«CheapTrick, a spectral envelope estimator for high-quality speech synthesis, Speech Communication, Vol. 67, March 2015, pp. 1- 7. https://doi.org/10.1016/j.specom.2014.09.003. </unstructured_citation></citation><citation key="ref29"><doi>10.1016/j.specom.2016.09.001</doi><unstructured_citation>M. Morise, «D4C, a band-aperiodicity estimator for high-quality speech synthesis», Speech Commun., vol. 84, pp. 57–65, nov. 2016, doi: 10.1016/j.specom.2016.09.001. </unstructured_citation></citation><citation key="ref30"><doi>10.1145/2729095.2729097</doi><unstructured_citation>F. Eyben, B. Schuller, «openSMILE:): the Munich open-source large-scale multimedia feature extractor», ACM SIGMultimedia Rec., vol. 6, fasc. 4, pp. 4–13, gen. 2015, doi: 10.1145/2729095.2729097. </unstructured_citation></citation><citation key="ref31"><doi>10.21437/interspeech.2016-129</doi><unstructured_citation>B. Schuller et al., «The INTERSPEECH 2016 computational paralinguistics challenge: 17th Annual Conference of the International Speech Communication Association, INTERSPEECH 2016», Proc. Annu. Conf. Int. Speech Commun. Assoc. INTERSPEECH, vol. 08-12-September-2016, pp. 2001–2005, 2016, doi: 10.21437/Interspeech.2016-129. </unstructured_citation></citation><citation key="ref32"><doi>10.1109/89.326616</doi><unstructured_citation>H. Hermansky, N. Morgan, «RASTA processing of speech», IEEE Trans. Speech Audio Process., vol. 2, fasc. 4, pp. 578–589, ott. 1994, doi: 10.1109/89.326616. </unstructured_citation></citation><citation key="ref33"><unstructured_citation>M. A. Hall, «Correlation-based Feature Selection for Machine Learning», PhD Thesis, University of Waikato, New Zealand, 1999. </unstructured_citation></citation><citation key="ref34"><doi>10.1201/9781315108230-10</doi><unstructured_citation>M. K. and K. Johnson, “Stepwise Selection”, from Feature Engineering and Selection: A Practical Approach for Predictive Models. Accessed on: February 19, 2023, [Online]. https://bookdown.org/max/FES/greedystepwise-selection.html (Accessed Date: October 10, 2024). </unstructured_citation></citation><citation key="ref35"><doi>10.1007/bf00116251</doi><unstructured_citation>J. R. Quinlan, «Induction of decision trees», Mach. Learn., vol. 1, fasc. 1, pp. 81–106, mar. 1986, doi: 10.1007/BF00116251. </unstructured_citation></citation><citation key="ref36"><unstructured_citation>D. P. Kingma e J. Ba, «Adam: A Method for Stochastic Optimization», January 30, 2017, arXiv: arXiv:1412.6980. doi: 10.48550/arXiv.1412.6980. </unstructured_citation></citation><citation key="ref37"><doi>10.1016/j.eswa.2024.125656</doi><unstructured_citation>F. Amato, V. Cesarini, G. Olmo, G. Saggio, G. Costantini, «Beyond breathalyzers: AIpowered speech analysis for alcohol intoxication detection», Expert Syst. Appl., vol. 262, p. 125656, March 2025, doi: 10.1016/j.eswa.2024.125656. </unstructured_citation></citation><citation key="ref38"><doi>10.21437/interspeech.2011-805</doi><unstructured_citation>D. Bone, M. Black, M. Li, A. Metallinou, S. Lee, e S. Narayanan, “Intoxicated Speech Detection by Fusion of Speaker Normalized Hierarchical Features and GMM Supervectors.”, Interspeech 2011, p. 3220. doi: 10.21437/Interspeech.2011-805. </unstructured_citation></citation><citation key="ref39"><doi>10.3389/fpsyg.2024.1412372</doi><unstructured_citation>VL. Holmes, G. Rieger, e S. Paulmann, «The effect of sexual orientation on voice acoustic properties», Front. Psychol., vol. 15, ago. 2024, doi: 10.3389/fpsyg.2024.1412372.</unstructured_citation></citation></citation_list></journal_article></journal></body></doi_batch>