Run Audio8 ASR Infinite for 24/7 Chinese and English Transcription
A 3B-decoder streaming ASR model with 240–560 ms delay. With the authors' adapted vLLM build, a rolling KV cache lets it transcribe audio of any length at constant memory.
The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.
Audio8 ASR Infinite is a streaming speech recognition model from Edge0 for Chinese and English. It combines the causal audio tower from Voxtral Realtime 4B with a decoder initialized from Qwen2.5-3B-Instruct, and it emits text while audio is still arriving. Its main feature is length: the checkpoint has a native context of 30 seconds, but with the authors’ adapted vLLM build a rolling KV cache keeps memory and latency constant. The card says this lets it transcribe 24/7 without drifting. This is a preview release of the transcription base.
Key specs
| Languages | Chinese, English |
| Audio clock | 80 / 120 / 160 ms (12.5 / 8.3 / 6.25 text decisions per second) |
| Transcription delay | 240–560 ms, configurable via target_delay_ms |
| Decoder | 36 layers, hidden 2048, from Qwen2.5-3B-Instruct |
| Audio tower | 32 layers, hidden 1280, 128 mel bins, from Voxtral Realtime 4B |
| Weights | 8.17 GB model.safetensors, bfloat16, plus semantic_vad_heads.safetensors |
| Rolling KV window | 30 s with exact RoPE re-basing (vLLM path) |
| Extras | Semantic VAD heads (8 classes, horizons 0.5 / 1.0 / 2.0 / 3.0 s) |
The card reports these error rates at 480 ms delay with the 80 ms clock (lower is better):
| test set | metric | Audio8 ASR Infinite | Voxtral-Mini-4B-Realtime-2602 | nemotron-3.5-asr-streaming-0.6b (@560 ms) |
|---|---|---|---|---|
| aishell1/test | CER | 1.750 | 16.795 | 12.927 |
| aishell4/test | CER | 2.893 | 16.456 | 14.677 |
| librispeech test.clean | WER | 3.042 | 2.210 | 3.353 |
| librispeech test.other | WER | 6.808 | 5.552 | 7.140 |
On these Chinese test sets the gap is large. On English LibriSpeech, Voxtral Realtime has lower error than Audio8.
Post-trained operating points:
| audio clock | frame_len |
streaming_n_left_pad_tokens |
target_delay_ms |
|---|---|---|---|
| 80 ms | 4 | 18 | 240 / 320 / 480 / 560 |
| 120 ms | 6 | 12 | 240 / 480 |
| 160 ms | 8 | 9 | 320 / 480 |
Install
The card doesn’t give a pip install line. Here is what the code needs:
transformers,torchandnumpy.- The
audio8_asr_infinitePython package, which provides the model class, the streaming decode helper and the example scripts. The card calls this “the embedded remote code” but gives no install step for it. The card also links a GitHub repository (github.com/Edge0-AI/Audio8-ASR-Infinite), but it doesn’t say whether the package or thedocker/directory used for vLLM serving comes from there. - A CUDA GPU, because the card’s example calls
.cuda()and loads in bfloat16.
The weights load from the Hub with trust_remote_code=True.
Run it
This is the card’s simulated-streaming decode for one 16 kHz mono clip stored as a NumPy array with values in [-1, 1]:
import numpy as npimport torchfrom transformers import AutoFeatureExtractor, AutoTokenizer
from audio8_asr_infinite.modeling.modeling_audio8_asr_infinite import ( Audio8ASRInfiniteForConditionalGeneration, resolve_qwen_language_token_id, resolve_qwen_streaming_special_token_ids,)from audio8_asr_infinite.streaming_inference import simulated_streaming_greedy_decode_batch
checkpoint = "Edge0/Audio8-ASR-Infinite"tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True)feature_extractor = AutoFeatureExtractor.from_pretrained(checkpoint, trust_remote_code=True)model = Audio8ASRInfiniteForConditionalGeneration.from_pretrained( checkpoint, trust_remote_code=True, torch_dtype=torch.bfloat16).eval().cuda()
class AudioConfig: raw_audio_samples_per_token = 1280 # 80 ms @ 16 kHz streaming_n_left_pad_tokens = 18 sampling_rate = 16000
waveform = np.load("sample.npy", allow_pickle=False).astype(np.float32)results = simulated_streaming_greedy_decode_batch( model=model, tokenizer=tokenizer, feature_extractor=feature_extractor, waveforms=[waveform], language_token_ids=[resolve_qwen_language_token_id(tokenizer, "zh")], special_ids=resolve_qwen_streaming_special_token_ids(tokenizer), audio_config=AudioConfig(), num_delay_tokens=[480 // 80], right_pad_text_tokens=10, dtype=torch.bfloat16, device=next(model.parameters()).device, max_new_tokens=512,)print(results[0]["final_text"])If you have a WAV file instead, use the card’s CLI wrapper:
python -m audio8_asr_infinite.examples.torch_streaming_decode \ --checkpoint /path/to/checkpoint \ --audio sample.wav --language zh --transcription-delay-ms 480Batch-transcribe a mixed Chinese and English folder
This is an untested extension of the card’s single-clip example. The decode helper’s arguments are lists (waveforms, language_token_ids, num_delay_tokens), but the card only ever passes one item. It doesn’t show that one call can take clips of different lengths or mix zh and en language tokens. It also never shows resolve_qwen_language_token_id(tokenizer, "en"), only "zh". Treat the script below as a starting point to check, not documented behavior:
from pathlib import Path
import numpy as npimport torchfrom transformers import AutoFeatureExtractor, AutoTokenizer
from audio8_asr_infinite.modeling.modeling_audio8_asr_infinite import ( Audio8ASRInfiniteForConditionalGeneration, resolve_qwen_language_token_id, resolve_qwen_streaming_special_token_ids,)from audio8_asr_infinite.streaming_inference import simulated_streaming_greedy_decode_batch
checkpoint = "Edge0/Audio8-ASR-Infinite"tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True)feature_extractor = AutoFeatureExtractor.from_pretrained(checkpoint, trust_remote_code=True)model = Audio8ASRInfiniteForConditionalGeneration.from_pretrained( checkpoint, trust_remote_code=True, torch_dtype=torch.bfloat16).eval().cuda()
class AudioConfig: raw_audio_samples_per_token = 1280 streaming_n_left_pad_tokens = 18 sampling_rate = 16000
# Files named like "meeting01.zh.npy" or "call07.en.npy"files = sorted(Path("clips").glob("*.npy"))waveforms = [np.load(f, allow_pickle=False).astype(np.float32) for f in files]languages = [f.suffixes[-2].lstrip(".") for f in files]
results = simulated_streaming_greedy_decode_batch( model=model, tokenizer=tokenizer, feature_extractor=feature_extractor, waveforms=waveforms, language_token_ids=[resolve_qwen_language_token_id(tokenizer, lang) for lang in languages], special_ids=resolve_qwen_streaming_special_token_ids(tokenizer), audio_config=AudioConfig(), num_delay_tokens=[480 // 80] * len(files), right_pad_text_tokens=10, dtype=torch.bfloat16, device=next(model.parameters()).device, max_new_tokens=512,)
for f, r in zip(files, results): print(f"{f.name}\t{r['final_text']}")If mixed batches don’t work, call the helper once per file with single-item lists, as the card does.
You may want to keep the clips short. The checkpoint’s native context is 30 seconds, and the card only credits the vLLM build with unlimited length. The card doesn’t say what the torch path does with longer audio, so this is our inference, not a documented limit.
Live captions around the clock with vLLM
For a continuous stream such as a meeting room microphone, a broadcast feed or a support line, use the card’s canonical deployment: Docker Compose with the adapted vLLM build. The service uses the 30 s rolling KV window with RoPE re-basing to keep memory flat.
cd dockerAUDIO8_MODEL_DIR=/path/to/checkpoint docker compose up -dThe stack also serves a web demo at http://localhost:8080/ and a TLS version at https://localhost:8443/, which uses a self-signed certificate. To stream a file into the realtime WebSocket from a terminal, use the card’s client. The card passes a --pace flag but doesn’t explain what it does.
python -m audio8_asr_infinite.examples.vllm_realtime_client \ --ws-url ws://127.0.0.1:18191/v1/realtime \ --audio sample.wav --language zh --target-delay-ms 480 --paceThe card doesn’t document the WebSocket message schema. To write your own client, start from the source of vllm_realtime_client.
Tune latency against accuracy
The delay is a direct trade-off: more delay gives the model more right context before it commits a token. On the 80 ms clock, all four delays (240, 320, 480 and 560 ms) are post-trained. This untested script runs the same clip at each one so you can pick a delay for your product. It makes one single-clip call per delay, which matches the card’s example, instead of mixing delays in one batch:
import numpy as npimport torchfrom transformers import AutoFeatureExtractor, AutoTokenizer
from audio8_asr_infinite.modeling.modeling_audio8_asr_infinite import ( Audio8ASRInfiniteForConditionalGeneration, resolve_qwen_language_token_id, resolve_qwen_streaming_special_token_ids,)from audio8_asr_infinite.streaming_inference import simulated_streaming_greedy_decode_batch
checkpoint = "Edge0/Audio8-ASR-Infinite"tokenizer = AutoTokenizer.from_pretrained(checkpoint, trust_remote_code=True)feature_extractor = AutoFeatureExtractor.from_pretrained(checkpoint, trust_remote_code=True)model = Audio8ASRInfiniteForConditionalGeneration.from_pretrained( checkpoint, trust_remote_code=True, torch_dtype=torch.bfloat16).eval().cuda()
class AudioConfig: raw_audio_samples_per_token = 1280 streaming_n_left_pad_tokens = 18 sampling_rate = 16000
waveform = np.load("sample.npy", allow_pickle=False).astype(np.float32)delays_ms = [240, 320, 480, 560]
for d in delays_ms: results = simulated_streaming_greedy_decode_batch( model=model, tokenizer=tokenizer, feature_extractor=feature_extractor, waveforms=[waveform], language_token_ids=[resolve_qwen_language_token_id(tokenizer, "zh")], special_ids=resolve_qwen_streaming_special_token_ids(tokenizer), audio_config=AudioConfig(), num_delay_tokens=[d // 80], right_pad_text_tokens=10, dtype=torch.bfloat16, device=next(model.parameters()).device, max_new_tokens=512, ) print(f"{d} ms: {results[0]['final_text']}")To test this against the vLLM server instead, change --target-delay-ms on the client.
Gotchas
- Remote code is required. Tokenizer, feature extractor and model all load with
trust_remote_code=True, and the model class is imported from theaudio8_asr_infinitepackage. - Use full weights only. The card supports only the full merged weight directory. Adapter-style or partially converted weights won’t work.
- Match the Python example’s input format. In the card’s Python example, the
.npyinput is 16 kHz mono float32 in [-1, 1]. The vLLM client and torch CLI take a.wavfile, and the card doesn’t say what format they need. The model loads in bfloat16, and the card’s decode call also passesdtype=torch.bfloat16. - Delays must line up with the clock.
target_delay_msmust be an integer multiple of the clock. Combinations outside the table above are allowed but weren’t post-trained. The card’s Python example sets up only the 80 ms clock (1280samples,18left-pad tokens). - Only the vLLM build is shown running unbounded audio. The 30 s rolling window and 24/7 claim apply to the Docker/vLLM path.
- Two different ports. The compose stack publishes
18191on the host. The service listens on18190inside the compose network. - Semantic VAD has no documented API yet. The heads ship as
semantic_vad_heads.safetensors, but the card shows no API for reading their output. - This is a preview. Frame-level semantic perception beyond transcription is still listed as in progress. The arXiv paper is marked “coming soon”.
When to pick it
Pick Audio8 ASR Infinite if you need low-latency streaming transcription of Mandarin, especially for long-running or always-on audio. Its aishell CERs (1.750 and 2.893 at 480 ms) are far below the two streaming baselines the card compares against. The Docker/vLLM setup is aimed at 24/7 captioning with bounded memory.
Don’t pick it for English-only workloads where accuracy matters most: Voxtral-Mini-4B-Realtime-2602 has lower WER on both LibriSpeech splits in the card’s own table. Skip it for languages other than Chinese and English, and for anything that needs a stable, documented semantic VAD API today. If you need to run without a GPU, note that the card only shows GPU usage.
Related
Run Confucius4-R2T2 for Append-Only Real-Time Speech Recognition
NetEase Youdao's streaming ASR model built on Qwen3-ASR: transcript text never gets rewritten, with 80 ms to 2 s chunks. Setup, code and caveats.
netease-youdao/Confucius4-R2T2
AuK: Clone, Edit, and Clean Up Speech Through One Instruction Interface
Tencent's MIT-licensed AuK does zero-shot TTS, speech editing, enhancement, and separation from natural-language instructions.
tencent/AuK
Run Breeze TTS 2 for Voice Cloning, Design and Streaming
Breeze TTS 2 is an open-weight English and Chinese TTS model with voice cloning, text-described voices and a streaming API. Weights are non-commercial.
BreezeBlue/Breeze-TTS-2