← all repositories

YuanGongND/whisper-at

A joint audio tagging and speech recognition model that extends OpenAI Whisper with audio event detection capability at minimal additional computational cost.

whisper-at
Not currently ranked — collecting fresh signals.
star history

Whisper-AT provides pretrained models that perform both speech recognition (with identical performance to original Whisper) and general audio event tagging across 527 AudioSet classes. The model can output audio event labels at various temporal resolutions alongside transcription. It offers a Python package, HuggingFace Space demo, and Google Colab notebook for easy experimentation.

Frequently asked

What is YuanGongND/whisper-at?
A joint audio tagging and speech recognition model that extends OpenAI Whisper with audio event detection capability at minimal additional computational cost.
Is whisper-at open source?
Yes — YuanGongND/whisper-at is open source, released under the BSD-2-Clause license.
What language is whisper-at written in?
YuanGongND/whisper-at is primarily written in Python.
How popular is whisper-at?
YuanGongND/whisper-at has 421 stars on GitHub.
Where can I find whisper-at?
YuanGongND/whisper-at is on GitHub at https://github.com/YuanGongND/whisper-at.

heatdrop uses Google Analytics to see which pages get read — nothing else. Your call. How we handle data.