Home / Articles / AI Courses / Audio (transofrmers)

Audio (transofrmers)

🔷 Key Takeaways

Course nameHugging Face Audio Course
PlatformHugging Face
PriceFree
DurationSelf-paced, ~7 units
LevelIntermediate
PrerequisitesDeep learning & transformers familiarity
SkillsAudio preprocessing, speech recognition, text-to-speech, audio classification, transformer-based modeling

About

The Hugging Face Audio Course is a free, self-paced program focused on applying transformer models to audio data. Designed for those with a grounding in deep learning, this course walks learners through using state-of-the-art models for tasks like speech recognition, audio classification, and speech generation—no prior audio experience required.

Who is teaching?

The course is developed by Hugging Face’s open-source audio team, including:

  • Sanchit Gandhi, ML Research Engineer (focus: speech recognition/translation)
  • Matthijs Hollemans, ML Engineer and audio plugin creator
  • Maria Khalusova, educational content creator
  • Vaibhav Srivastav, ML Developer Advocate (focus: low-resource TTS)

What is covered?

You’ll explore:

  • Audio data processing and preparation
  • Using transformer pipelines for audio tasks
  • Architecture insights into audio transformers
  • Building real-world apps like genre classifiers and speech transcribers
  • Generating speech from text with TTS models

Each unit combines theory, hands-on exercises, and quizzes to reinforce learning.

Skills you’ll develop

  • Audio preprocessing & feature extraction
  • Fine-tuning audio transformers
  • Speech-to-text & text-to-speech systems
  • Deploying real-world audio AI applications

Level

Intermediate – A solid understanding of deep learning and transformer models is expected. No audio domain expertise needed.