Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
This paper improves the accuracy of Automatic Speech Recognition (ASR) models by controlling the style of the transcription, allowing for more reliable word-level timing and disfluency detection. Practitioners can benefit from this work to improve the quality of ASR systems in real-world applications.