Mel Spectrogram
المخطط الطيفي ميل
تمثيل بصري لترددات الصوت عبر الزمن مُرجَّح بمقياس ميل الذي يقارب إدراك السمع البشري، يستخدمه Whisper بـ80 قناة.
A visual representation of audio frequencies over time, weighted by the Mel scale to approximate human hearing perception. Whisper uses 80-channel log-magnitude Mel spectrograms.
Also translated asطيف ميل، مخطط ميل الطيفي اللوغاريتمي
First appears in this corpus in: Deep Speech 2: End-to-End Speech Recognition in English and Mandarin (2015)
Appears in these papers
- Deep Speech 2: End-to-End Speech Recognition in English and Mandarin2015in the sky ✦
- ImageBind: One Embedding Space to Bind Them All2023in the sky ✦
- ImageBind: One Embedding Space to Bind Them All2023in the sky ✦
- Jukebox: A Generative Model for Music2020in the sky ✦
- Jukebox: A Generative Model for Music2020in the sky ✦
- MusicLM: Generating Music From Text2023in the sky ✦
- MusicLM: Generating Music From Text2023in the sky ✦
- Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions2018in the sky ✦
- Natural TTS Synthesis by Conditioning WaveNet on Mel Spectrogram Predictions2018in the sky ✦
- WaveNet: A Generative Model for Raw Audio2016in the sky ✦
- Robust Speech Recognition via Large-Scale Weak Supervision2022in the sky ✦