article
An effective conversion of recognised phonemes to words has a significant impact on the overall performance of an automatic lip-reading system. In this paper, transformer-based models have been examined for phoneme-to-text translation. The benchmark BBC Lip-reading Sentence 2 (LRS2) data is used in this study, and the T5 model (Text-To-Text Transfer Transformer) and the GPT-2 model are considered the conversion models. A detailed discussion is provided on the analysis procedure, data pre-processing, model implementation, and the metrics employed for model evaluation. The experimental results have demonstrated that the T5 model can achieve a substantial improvement over the GPT-2 model with a superior accuracy and fluency in the phoneme-to-text translation, offering a potentially promising solution for enhancing the precision of automatic translation and speech processing systems. This work contributes to the development of more robust applications in the field of speech recognition and communication technologies.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iccsi62669.2024.10799296
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.