Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition
Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.
Computer Science > Computation and Language
Title:Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition
Abstract:Phoneme-centric visual speech recognition reconstructs sentences from intermediate phoneme predictions, making overall recognition performance highly dependent on the robustness of the phoneme-to-text reconstruction model. Existing reconstruction approaches are commonly trained on clean phoneme sequences or synthetically corrupted inputs, leading to a mismatch between training conditions and the realistic phoneme prediction errors encountered during inference. To address this limitation, this paper proposes progressive error curriculum training (PECT). This curriculum learning framework progressively adapts a No Language Left Behind (NLLB)-based phoneme-to-text reconstruction model using synthetic phoneme perturbations, multi-domain pseudo-labels, and target-domain pseudo-labels generated by a visual speech recognizer. By gradually exposing the reconstruction model to increasingly realistic phoneme prediction errors, the proposed framework improves robustness while preserving sentence-reconstruction accuracy. Experiments on the LRS2 and LRS3 benchmarks demonstrate that PECT consistently improves reconstruction performance across multiple phoneme-based visual speech recognition frontends, including visual automatic speech recognition (V-ASR), point visual automatic speech recognition (PV-ASR), and head-pose-aware visual speech recognition (HP-VSR) variants. In particular, PECT reduces the word error rate (WER) of HP-VSR-FiLMFuse (L4) from 23.3% to 22.2% on LRS2 and reduces the WER of HP-VSR-ResFiLM from 30.3% to 29.7% on LRS3. Comprehensive ablation studies and qualitative analyses further demonstrate the effectiveness of progressively adapting the reconstruction model to realistic phoneme prediction errors. These results show that PECT provides an effective and generalizable curriculum learning strategy for phoneme-to-text reconstruction in phoneme-centric visual speech recognition.
| Comments: | Submitted for journal publication and currently under consideration |
| Subjects: | Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Sound (cs.SD) |
| Cite as: | arXiv:2609.20839 [cs.CL] |
| (or arXiv:2609.20839v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2609.20839
arXiv-issued DOI via DataCite
|
Submission history
From: Matthew Kit Khinn Teng Mr [view email][v1] Thu, 30 Jul 2026 04:56:46 UTC (276 KB)
Access Paper:
- View PDF
- HTML (experimental)
- TeX Source
Current browse context:
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — NLP / Computation & Language
-
A Mechanistic Study of AI-Text Detection Neurons in Frozen BERT: Sparse Probing and Activation Patching on RAID
Sep 28
-
Manifold Projection and Iterative Autoencoder Refinement for Masked Language Modeling
Sep 28
-
Not All Memories Are Equal: Hierarchical Collaborative Memory for Validity-Aware Retrieval in LLM Agents
Sep 28
-
Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.