Scribe is ElevenLabs' first Speech to Text (STT) model.
Think of it as a tool that listens to audio and converts it into written text โ like a highly accurate digital transcriptionist.
๐ฏ Key claim: It is described as the world's most accurate transcription model.
Before understanding why Scribe is impressive, you need to understand Word Error Rate (WER).
| Language | Accuracy |
|---|---|
| Italian | 98.7% accurate |
| English | 96.7% accurate |
| Serbian, Cantonese, Malayalam | Competing models exceed 40% WER; Scribe dramatically reduces this |
Scribe was tested against industry-leading models:
โ Scribe outperformed all competitors on both benchmarks across 99 languages
Scribe isn't just about accuracy. It comes packed with three core features:
[laughter], [applause], [music]One of Scribe's most important contributions is reducing the gap for underserved languages.
Many STT models are trained heavily on English and a few major languages. This leaves speakers of languages like:
...with very poor transcription quality (40%+ error rates).
๐ Scribe supports 99 languages and dramatically improves accuracy for these traditionally underserved communities.
There are two ways to access Scribe:
| Concept | Key Takeaway |
|---|---|
| What is Scribe? | ElevenLabs' Speech-to-Text model |
| Word Error Rate | Lower = better; Scribe achieves industry-low WER |
| Benchmarks | Outperforms Gemini, Whisper, Deepgram |
| Core Features | Timestamps, Speaker ID, Audio-event tagging |
| Language Access | Supports 99 languages including underserved ones |
| How to Use | API (developers) or Dashboard (creators/businesses) |
๐ก Bottom Line: Scribe is a highly accurate, feature-rich, and inclusive speech-to-text model designed to work reliably across languages, speakers, and real-world audio conditions.