Reddit - r/MachineLearning

Best current methods for finetuning whisper on domain specific vocabulary? [P]

Best current methods for finetuning whisper on domain specific vocabulary?

Hey everyone, I’m wondering whether there are any newer or more effective methods for fine tuning whisper on domain specific speech.

I’m working on a project where the model needs to reliably detect certain specific words and technical terms. The vocabulary and context are mostly in Spanish.

Does anyone have experience with a similar use case? Roughly how many hours of labeled audio would be needed before seeing the model converged?

I know about LoRA, QLoRA, and Spectrum, but I’m curious if there are any newer or better ways to adapt whisper to specific vocabulary.

Any help is welcome!

Comments

No comments yet. Start the discussion.