Logo image
Modelling Filled Particles and Prolongation Using End-to-end Automatic Speech Recognition Systems: A Quantitative and Qualitative Analysis
Conference proceeding   Open access   Peer reviewed

Modelling Filled Particles and Prolongation Using End-to-end Automatic Speech Recognition Systems: A Quantitative and Qualitative Analysis

Vincenzo Norman Vitale, Loredana Schettino and Francesco Cutugno
Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024), Vol.3878, pp.1-7
3878
Tenth Italian Conference on Computational Linguistics (Clic-it 2024) (Pisa, 04/12/2024–06/12/2024)
2024
Handle:
https://hdl.handle.net/10863/53060

Abstract

Disfluencies Speech recognition Probing Interpretability Explainability
State-of-the-art automatic speech recognition systems based on End-to-End models (E2E-ASRs) achieve remarkable performances. However, phenomena that characterize spoken language such as fillers ( ) or segmental prolongations (the) are still mostly considered as disrupting objects that should not be included to obtain optimal transcriptions, despite their acknowledged regularity and communicative value. A recent study showed that two types of pre-trained systems with the same Conformer-based encoding architecture but different decoders – a Connectionist Temporal Classification (CTC) decoder and a Transducer decoder – tend to model some speech features that are functional for the identification of filled pauses and prolongation in speech. This work builds upon these findings by investigating which of the two systems is better at fillers and prolongations detection tasks and by conducting an error analysis to deepen our understanding of how these systems work.
pdf
107_main_long1.69 MBDownloadView
Open Access
url
https://ceur-ws.org/Vol-3878/107_main_long.pdfView

Details

Metrics

1 Record Views
Logo image