Automated stress identification and speaker diarisation for speech-language pathology
Loading...
Files
Date
Authors
Researcher ID
Supervisors
Journal Title
Journal ISSN
Volume Title
Publisher
North-West University
Record Identifier
Abstract
This study presents an engineering-based approach to assist speech-language pathologists (SLPs) in their analysis of parent-child speech interactions. Currently, SLPs spend numerous hours evaluating video recordings of contact sessions to identify each speaker's specific speech and associated behaviours. This study o↵ers an autonomous solution that utilises ECG and accelerometer sensors to identify stress and movement, respectively, and speaker diarisation to identify who spoke when. The algorithms were trained and evaluated using open-source datasets such as CommonVoice [1], LibriSpeech [2], MIT-BIH [3], WESAD [4], and HAPT [5]. Obtaining respectable results with a diarisation error rate (DER) of 5.207% for speaker diarisation, a validation accuracy of 93.939% (32-true positive, 2-false positive, 8-false negative, 165-total size) for stress identification with a beat detection error rate of 0.877% for normal beats, and a validation accuracy of 99.038% (6355-true positive, 76-false positive, 45-false negative, 12575-total size) for movement identification. These algorithms are integrated into a single analyser application that provides audio and visual results to the SLP. Additionally, a second recorder application was created to capture a single-channel audio recording and multiple ECG and accelerometer signals from patients wearing Polar H10 devices. These two applications automate the previous manual analysis procedures for speech, stress, and movement assessments. The study's findings o↵er a
novel, automated tool for speech interaction analyses that can potentially benefit SLPs in
their clinical assessments.
Sustainable Development Goals
Quality Education
Description
Dissertation, Master of Engineering in Computer and Electronic Engineering -- North-West University
