|
Deepfake speech detection using perceptual pathological features |
|---|---|
| รหัสดีโอไอ | |
| Title | Deepfake speech detection using perceptual pathological features |
| Creator | Anuwat Chaiwongyen |
| Contributor | Waree Kongprawechnon, Advisor |
| Publisher | Thammasat University |
| Publication Year | 2568 |
| Keyword | Spoof detection, Anti-spoofing, Speech features, Speaker verification, Speaker recognition, Speech synthesis, Voice conversion, Speech recognition, Sound detection, Speech signal processing, Spoofing countermeasure, Data fusion, Deep learning, การตรวจจับการปลอมแปลง, การป้องกันการปลอมแปลง, ลักษณะเด่นของเสียงพูด, การยืนยันตัวตนผู้พูด, การรู้จำผู้พูด, การสังเคราะห์เสียงพูด, การแปลงเสียงพูด, การรู้จำเสียงพูด, การตรวจจับเสียง, การประมวลผลสัญญาณเสียงพูด, มาตรการรับมือการปลอมแปลง, การหลอมรวมข้อมูล, การเรียนรู้เชิงลึก |
| Abstract | The exponential growth of generative artificial intelligence in speech synthesis has established deepfake speech detection as a critical research priority. Synthetic vocal technologies introduce severe vulnerabilities to security infrastructures, particularly concerning biometric authentication, voice activated interfaces, and automatic speaker verification (ASV) systems. Consequently, fortifying the defensive mechanisms of these platforms is imperative to neutralize emerging spoofing threats. To address this, this research explores the diagnostic utility of perceptual speech-pathological features, traditionally applied in clinical settings to assess vocal abnormalities, as novel indicators for synthetic speech detection. Specifically, the study evaluates a comprehensive suite of timbral characteristics, including hardness, depth, brightness, roughness, sharpness, warmth, boominess, and reverberation. Analytical findings confirm that these perceptual parameters effectively differentiate bona fide human speech from artificial generation. Furthermore, extending the dimensional analysis of these timbral features directly enhances the overall robustness of the detection framework. To improve the identification of synthetic speech, this research proposes a dual-branch architecture. The first branch employs a Deep Neural Network (DNN) to evaluate multidimensional speech-pathological features, while the second branch utilizes a Gammatone filterbank and ResNet-18 framework to simulate human cochlear processing. Experimental validation on the ASVspoof 2019 dataset indicates that this combined method successfully outperforms standard baseline models, attaining a competitive Equal Error Rate (EER) of 5.93%. |