README.md
15.3 KB · 366 lines · markdown Raw
1 ---
2 license: cc-by-4.0
3 base_model: nvidia/parakeet-unified-en-0.6b
4 base_model_relation: quantized
5 library_name: transcribe.cpp
6 pipeline_tag: automatic-speech-recognition
7 language:
8 - en
9 tags:
10 - gguf
11 - transcribe.cpp
12 - asr
13 - speech-to-text
14 - parakeet
15 - conformer
16 - rnnt
17 transcribe_cpp:
18 schema_version: 2
19 wer_fleurs_en:
20 q8_0: 3.99
21 wer_librispeech_test_clean:
22 f32: 1.59
23 f16: 1.59
24 q8_0: 1.6
25 q6_k: 1.61
26 q5_k_m: 1.58
27 q4_k_m: 1.62
28 rtf_m4_max:
29 cpu: 37.32
30 metal: 208.06
31 rtf_ryzen_4750u:
32 cpu: 13.96
33 vulkan: 25.28
34 streaming: true
35 translate: false
36 lang_detect: false
37 timestamps: token
38 ---
39
40 # parakeet-unified-en-0.6b: transcribe.cpp GGUF
41
42 GGUF conversions of [nvidia/parakeet-unified-en-0.6b](https://huggingface.co/nvidia/parakeet-unified-en-0.6b) for use
43 with [transcribe.cpp](https://github.com/handy-computer/transcribe.cpp).
44
45 Ported from upstream commit
46 [d4ac992](https://huggingface.co/nvidia/parakeet-unified-en-0.6b/commit/d4ac992),
47 pinned 2026-05-10.
48 Validated against the NeMo reference at transcribe.cpp commit
49 [42528dd](https://github.com/handy-computer/transcribe.cpp/tree/42528dd)
50 on 2026-05-10.
51
52 English speech-to-text with punctuation and capitalization. A 0.6B-parameter FastConformer encoder with an RNN-T transducer decoder, trained as a 'unified' streaming/offline model. This port runs the model in both offline and buffered streaming modes.
53
54
55 ## Downloads
56
57 | Quantization | Download | Size | WER (LibriSpeech test-clean, offline) |
58 | --- | --- | ---: | ---: |
59 | F32 | [parakeet-unified-en-0.6b-F32.gguf](https://huggingface.co/handy-computer/parakeet-unified-en-0.6b-gguf/resolve/main/parakeet-unified-en-0.6b-F32.gguf) | 2.47 GB | 1.59% |
60 | F16 | [parakeet-unified-en-0.6b-F16.gguf](https://huggingface.co/handy-computer/parakeet-unified-en-0.6b-gguf/resolve/main/parakeet-unified-en-0.6b-F16.gguf) | 1.24 GB | 1.59% |
61 | Q8_0 | [parakeet-unified-en-0.6b-Q8_0.gguf](https://huggingface.co/handy-computer/parakeet-unified-en-0.6b-gguf/resolve/main/parakeet-unified-en-0.6b-Q8_0.gguf) | 731 MB | 1.60% |
62 | Q6_K | [parakeet-unified-en-0.6b-Q6_K.gguf](https://huggingface.co/handy-computer/parakeet-unified-en-0.6b-gguf/resolve/main/parakeet-unified-en-0.6b-Q6_K.gguf) | 602 MB | 1.61% |
63 | Q5_K_M | [parakeet-unified-en-0.6b-Q5_K_M.gguf](https://huggingface.co/handy-computer/parakeet-unified-en-0.6b-gguf/resolve/main/parakeet-unified-en-0.6b-Q5_K_M.gguf) | 541 MB | 1.58% |
64 | Q4_K_M | [parakeet-unified-en-0.6b-Q4_K_M.gguf](https://huggingface.co/handy-computer/parakeet-unified-en-0.6b-gguf/resolve/main/parakeet-unified-en-0.6b-Q4_K_M.gguf) | 477 MB | 1.62% |
65
66 WER on the full LibriSpeech test-clean split (2,620 utterances), batch size 1, timestamps none. Figures without a commit were published before provenance was recorded.
67
68 Greedy RNN-T decoding, no external LM. F32 reference baseline: 1.59%. NVIDIA's
69 self-reported number on the same split is 1.63%.
70
71
72 ## Usage
73
74 Build transcribe.cpp from source:
75
76 ```bash
77 git clone git@github.com:handy-computer/transcribe.cpp.git
78 cd transcribe.cpp
79 cmake -B build && cmake --build build
80 ```
81
82 Run on a 16 kHz mono WAV:
83
84 ```bash
85 build/bin/transcribe-cli \
86 -m parakeet-unified-en-0.6b-Q8_0.gguf \
87 input.wav
88 ```
89
90 If your audio isn't already 16 kHz mono WAV, convert it first:
91
92 ```bash
93 ffmpeg -i input.mp3 -ar 16000 -ac 1 output.wav
94 ```
95
96 See the [transcribe.cpp model page](https://github.com/handy-computer/transcribe.cpp/blob/main/docs/models/parakeet-unified-en-0.6b.md) for performance
97 numbers, numerical validation, and reproduction steps.
98
99 ## License
100
101 Inherited from the base model: **CC-BY-4.0**. See the
102 [upstream model card](https://huggingface.co/nvidia/parakeet-unified-en-0.6b) for full terms.
103
104 ---
105
106 ## Original Model Card
107
108 > The section below is reproduced from
109 > [nvidia/parakeet-unified-en-0.6b](https://huggingface.co/nvidia/parakeet-unified-en-0.6b) at commit
110 > `d4ac992` for offline reference. The upstream card is the
111 > authoritative source.
112
113 # 🦜Parakeet-unified-en-0.6b: Unified ASR model for offline and streaming inference
114
115 | [Model architecture](#model-architecture) | [Model size](#model-architecture) | [Language](#datasets) |
116 |---|---|---|
117
118 Parakeet-unified-en-0.6b is an English automatic speech recognition (ASR) model based on transducer architecture (RNN-T) combining both offline and streaming inference (with a minimum latency of 160ms) in one model [1]. It is trained mostly on the English part of the Granary dataset [4], which contains approximately 250,000 hours of US English (en-US) speech across diverse acoustic conditions. The model transcribes speech to English alphabet, spaces, and apostrophes with punctuation and captalization support.
119
120 <figure align="center">
121 <img src="figures/wer_comparison.png" width="1250" />
122 <figcaption>
123 Average WER comparison on the HF ASR Leaderboard datasets including offline and streaming inference with different latency values.
124 </figcaption>
125 </figure>
126
127 Why Choose nvidia/parakeet-unified-en-0.6b?
128
129 - **One model for both tasks:** You need to utilize only one unified model for both offline and streaming inference with a minimum latency of 160ms.
130 - **Better accuracy performance:** The unified model achieves better accuracy performance on the HF ASR Leaderboard datasets compared to the previous transducer-based offline and streaming only models.
131 - **Streaming chunk size flexibilty:** Enables you to choose the optimal streaming latency (chunk + right context) from 2080ms to 160ms with step of 80ms.
132 - **Punctuation & Capitalization:** Built-in support for punctuation and capitalization in output text
133
134 This model consists of a 🦜 Parakeet (FastConformer) encoder (jointly trained in offline and streaming modes) with an RNN-T decoder. It is designed for offline and streaming speech-to-text applications where latency can be as low as 160ms, such as voice assistants, live captioning, and conversational AI systems. The current inference pipeline supports only buffered streaming (left context is recomputed for each chunk) that can be longer than cache-aware streaming.
135
136 This model is ready for commercial/non-commercial use.
137
138 ## License/Terms of Use:
139
140 Governing Terms: Use of the model is governed by the [NVIDIA Open Model License Agreement](https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license/).
141
142 ## Deployment Geography:
143
144 Global
145
146
147 ## Use Case:
148
149 This model is for transcription of English audio in offline and streaming modes.
150
151
152 ## Release Date:
153
154 - Hugging Face [04/07/2026] via [https://huggingface.co/nvidia/parakeet-unified-en-0.6b](https://huggingface.co/nvidia/parakeet-unified-en-0.6b)
155
156
157 ## Model Architecture
158
159 **Architecture Type:** Unified-FastConformer-RNNT
160
161 The unified model architecture is presented in [1]. The model is based on the FastConformer encoder architecture [2] with 24 encoder layers and an RNNT (Recurrent Neural Network Transducer) decoder. The model was trained jointly in offline and streaming modes. In the offline mode we used standard offline training with full-context self-attention and non-causal convolutions. In the streaming mode we applied chunked self-attention masks (incluing left, middle/chunk and right context) together with Dynamic Chunked Convolutions inside each FastConformer layer [3] to adapt the model to both decoding scenarios. We also introduced a novel mode-consistency regularization loss to further reduce the gap between offline and streaming performance. All the model parameters are shared between offline and streaming modes (encoder, predictor, and joint networks), including initial x8 subsampling with non-causal convolutions.
162
163 **Network Architecture:**
164
165 - Encoder: Unified FastConformer with 24 layers
166 - Decoder: RNNT (Recurrent Neural Network Transducer)
167 - Parameters: 600M
168
169 ## NVIDIA NeMo
170
171 ## How to Use this Model
172
173 For now, we provide only inference support for the unified model. We will release the unified training pipeline soon.
174
175 ### Loading the Model
176
177 ```python
178 import nemo.collections.asr as nemo_asr
179 asr_model = nemo_asr.models.ASRModel.from_pretrained(model_name="nvidia/parakeet-unified-en-0.6b")
180 ```
181
182 ### Offline Inference
183
184 ```python
185 output = asr_model.transcribe([wav_file_path])
186 print(output[0].text)
187 ```
188
189 ### Streaming Inference
190
191 For streaming inference you can use statfull chunked RNN-T decoding script from NeMo - [/NeMo/blob/main/examples/asr/asr_chunked_inference/rnnt/speech_to_text_streaming_infer_rnnt.py](https://github.com/NVIDIA-NeMo/NeMo/blob/main/examples/asr/asr_chunked_inference/rnnt/speech_to_text_streaming_infer_rnnt.py)
192
193 ```bash
194 cd NeMo
195 python examples/asr/asr_chunked_inference/rnnt/speech_to_text_streaming_infer_rnnt.py \
196 model_path=<model_path> \
197 dataset_manifest=<dataset_manifest> \
198 output_filename=<output_json_file> \
199 left_context_secs=<left_context_secs> \ # left context in seconds, 5.6s by default
200 chunk_secs=<chunk_secs> \ # chunk size in seconds, 0.56s by default
201 right_context_secs=<right_cintext_secs> \ # right context in seconds, 0.56s by default
202 att_context_size_as_chunk=true \ # set to true to use chunked self-attention masks
203 batch_size=<batch_size>
204 ```
205
206 You can also run streaming inference through the pipeline method, which uses [NeMo/examples/asr/conf/asr_streaming_inference/buffered_rnnt.yaml](https://github.com/NVIDIA-NeMo/NeMo/blob/main/examples/asr/conf/asr_streaming_inference/buffered_rnnt.yaml) configuration file to build end‑to‑end workflows with punctuation and capitalization (PnC), inverse text normalization (ITN), and translation support.
207
208 ```python
209 from nemo.collections.asr.inference.factory.pipeline_builder import PipelineBuilder
210 from omegaconf import OmegaConf
211
212 # Path to the buffered rnnt config file downloaded from above link
213 cfg_path = 'buffered_rnnt.yaml'
214 cfg = OmegaConf.load(cfg_path)
215
216 # Pass the paths of all the audio files for inferencing
217 audios = ['/path/to/your/audio.wav']
218
219 # Create the pipeline object and run inference
220 pipeline = PipelineBuilder.build_pipeline(cfg)
221 output = pipeline.run(audios)
222
223 # Print the output
224 for entry in output:
225 print(entry['text'])
226 ```
227
228 ---
229
230 ### Setting up Streaming Configuration
231
232 Latency is defined as the sum of the chunk size (middle part) and the right context.
233 For the left context we use 5.6s by default (5.6s was used during the model training), but you can try to find the optimal value for better accuracy/speed trade-off.
234
235 We would recommend to use the following context parameters for different latencies:
236
237 | Left, s | Chunk, s | Right, s | Latency (C+R), s |
238 | :---: | :---: | :---: | :---: |
239 | 5.6 | 1.04 | 1.04 | 2.08 |
240 | 5.6 | 0.56 | 0.56 | 1.12 |
241 | 5.6 | 0.16 | 0.40 | 0.56 |
242 | 5.6 | 0.08 | 0.24 | 0.32 |
243 | 5.6 | 0.08 | 0.16 | 0.24 |
244 | 5.6 | 0.08 | 0.08 | 0.16 |
245
246 ### Input
247
248 - Input Type(s): Audio
249 - Input Format(s): wav
250 - Input Parameters: One-Dimensional (1D)
251 - Other Properties Related to Input: Maximum Length in seconds specific to GPU Memory, No Pre-Processing Needed, Mono channel is required. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
252
253
254 ### Output
255
256 - Output Type(s): Text String in English
257 - Output Format(s): String
258 - Output Parameters: One-Dimensional (1D)
259 - Other Properties Related to Output: No Maximum Character Length, transcribe punctuation and capitalization. By leveraging NVIDIA’s hardware (e.g. GPU cores) and software frameworks (e.g., CUDA libraries), the model achieves faster training and inference times compared to CPU-only solutions.
260
261
262 ## Datasets
263
264 ### Training Datasets
265
266 The majority of the training data comes from the English portion of the Granary dataset [4]:
267
268 - YouTube-Commons (YTC) (109.5k hours)
269 - YODAS2 (102k hours)
270 - Mosel (14k hours)
271 - LibriLight (49.5k hours)
272
273 In addition, the following datasets were used:
274
275 - Librispeech 960 hours
276 - Fisher Corpus
277 - Switchboard-1 Dataset
278 - WSJ-0 and WSJ-1
279 - National Speech Corpus (Part 1, Part 6)
280 - VCTK
281 - VoxPopuli (EN)
282 - Europarl-ASR (EN)
283 - Multilingual Librispeech (MLS EN)
284 - Mozilla Common Voice (v11.0)
285 - Mozilla Common Voice (v7.0)
286 - Mozilla Common Voice (v4.0)
287 - People Speech
288 - AMI
289
290 **Data Modality:** Audio and text
291
292 **Audio Training Data Size:** 530k hours
293
294 **Data Collection Method:** Human - All audios are human recorded
295
296 **Labeling Method:** Hybrid (Human, Synthetic) - Some transcripts are generated by ASR models, while some are manually labeled
297
298 ### Evaluation Datasets
299
300 The model was evaluated on the HuggingFace ASR Leaderboard datasets:
301
302 - AMI
303 - Earnings22
304 - Gigaspeech
305 - LibriSpeech test-clean
306 - LibriSpeech test-other
307 - SPGI Speech
308 - TEDLIUM
309 - VoxPopuli
310
311 ## Performance
312
313 ## ASR Performance (w/o PnC)
314
315 ASR performance is measured using the Word Error Rate (WER). Both ground-truth and predicted texts are processed using [whisper-normalizer](https://pypi.org/project/whisper-normalizer/) version 0.1.12. The obtained results for other models can be slightly different from the official HF model cards because of the different evaluation machines.
316
317 The following table show the WER on the [HuggingFace OpenASR leaderboard](https://huggingface.co/spaces/hf-audio/open_asr_leaderboard) datasets including offline and streaming inference with different latency values:
318
319
320 | Model setup | Offline | 2.08s | 1.12s | 0.56s | 0.40s | 0.32s | 0.24s | 0.16s | 0.08s |
321 | :--- | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: | :---: |
322 | nvidia/parakeet-tdt-0.6b-v2 | 6.04 | 7.99 | 22.83 | 69.55 | 95.12 | — | — | — | — |
323 | nvidia/nemotron-speech-streaming-en-0.6b | 6.92 | 7.46 | 6.92 | 7.09 | 9.52 | 7.64 | 8.01 | **7.84** | **8.70** |
324 | nvidia/parakeet-unified-en-0.6b | **5.91** | **6.14** | **6.29** | **6.52** | **6.70** | **6.92** | **7.35** | 8.44 | 15.63 |
325
326
327 Parakeet-unified-en-0.6b model outperforms previous NVIDIA transducer-based models in offline and streaming (up to 240ms latency) inference modes. At 160ms latency, the unified model start to degrade because of the ansence of enough right context, yielding slightly to the strong streaming baseline. For 80ms latency we would recommend to use nemotron-speech-streaming-en-0.6b model instead.
328
329 ## Software Integration
330
331 **Runtime Engine:** NeMo 2.7.3
332
333 **Supported Hardware Microarchitecture Compatibility:**
334
335 - NVIDIA Ampere
336 - NVIDIA Blackwell
337 - NVIDIA Hopper
338 - NVIDIA Volta
339
340 **Test Hardware:**
341
342 - NVIDIA V100
343 - NVIDIA A100
344 - NVIDIA A6000
345 - DGX Spark
346
347 **Preferred/Supported Operating System(s):** Linux
348
349 ## Ethical Considerations
350
351 NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
352
353 Please report model quality, risk, security vulnerabilities or NVIDIA AI Concerns [here](https://www.nvidia.com/en-us/support/submit-security-vulnerability/).
354
355 ## References
356
357 [1] [Reducing the Offline-Streaming Gap for Unified ASR Transducer with Consistency Regularization](https://arxiv.org/abs/2604.19079)
358
359 [2] [Fast Conformer with Linearly Scalable Attention for Efficient Speech Recognition](https://arxiv.org/abs/2305.05084)
360
361 [3] [Dynamic Chunk Convolution for Unified Streaming and Non-Streaming Conformer ASR](https://arxiv.org/abs/2304.09325)
362
363 [4] [NVIDIA Granary](https://huggingface.co/datasets/nvidia/Granary)
364
365 [5] [NVIDIA NeMo Framework](https://github.com/NVIDIA/NeMo)
366