Use AIUnderstand AIBuild with AI
Build with AI·Model deep dive·· 8 min read

Run Confucius4-R2T2 for Append-Only Real-Time Speech Recognition

NetEase Youdao's streaming ASR model built on Qwen3-ASR: transcript text never gets rewritten, with 80 ms to 2 s chunks. Setup, code and caveats.

The code in this post comes from the model's docs and hasn't been run in our CI yet. If something breaks, let us know.

Confucius4-R2T2 is a streaming speech recognition model from NetEase Youdao. It is built on Qwen3-ASR, and the card lists Qwen3-ASR-1.7B as the base model. R2T2 stands for “Real Real-Time Transcription”. Its main feature is append-only output: once the model emits text, it doesn’t revise it. Many streaming ASR systems rewrite earlier words as more audio arrives. That makes captions flicker and gives downstream code a moving target. R2T2 is trained to emit only a “longest stable prefix” and to wait for more audio when it isn’t sure. That makes it a good fit for live captions, simultaneous translation, and LLM agents that need to act on text as soon as it arrives.

Key specs

  • Base model: Qwen/Qwen3-ASR-1.7B
  • Decoding chunk size: configurable from 80 ms to 2 s
  • Average latency: 200 to 600 ms, according to the card
  • Output mode: append-only. Emitted text stays committed.
  • Backends: vLLM (shown in the card) and a Hugging Face transformers backend (mentioned but not documented on the card)
  • Languages: optimized for Chinese and English. The card says it “retains useful cross-lingual streaming capability” on French, German, Italian, Japanese, Korean, Portuguese, Russian, Spanish, Arabic and others.
  • Extras: context and hotword prompts, batched offline inference, and a WebSocket server with VAD
  • License: code is Apache 2.0. Weights use the NetEase Model Use License Agreement.

The card reports these results at a 160 ms chunk size. The English rows are WER % and the Chinese row (CN-RealSI) is CER %. Lower is better. All numbers are the vendor’s own:

Dataset R2T2 (160 ms) Qwen3-ASR base (160 ms) X-ASR (160 ms) Nemotron (160 ms) Voxtral (160 ms) Qwen3-ASR (pseudo-streaming, 2 s)
LS-clean (English, WER) 2.13 22.30 3.86 3.71 2.49 1.67
LS-other (English, WER) 4.88 25.74 9.64 8.27 7.15 3.54
Earnings22 (English, WER) 9.36 29.72 15.95 17.22 11.66 6.68
AMI (English, WER) 11.37 24.79 14.41 18.11 15.94 9.25
CN-RealSI (Chinese, CER) 3.48 39.72 4.92 11.52 8.74 3.34

At the same 160 ms chunk size, R2T2 has the lowest error among the open-source true-streaming models in these tables. Its lead over X-ASR and Voxtral is sometimes small, though. On TED-LIUM it scores 3.34 against X-ASR’s 3.75, and on SPGI it scores 3.00 against Voxtral’s 3.06. R2T2 still trails pseudo-streaming Qwen3-ASR with 2 s chunks. Some proprietary systems in the card also do better on several datasets: Commercial B gets 2.48 on LS-other and 8.44 on AMI. The card marks those proprietary systems as pseudo-streaming, which means they may revise text they have already emitted.

Install

The GitHub repo installs the package that provides qwen_asr, which the card’s code imports. Python 3.10+ is supported, and 3.12 is the version the authors test against.

Terminal window
git clone https://github.com/netease-youdao/Confucius4-R2T2.git
cd Confucius4-R2T2
uv venv --python 3.12
source .venv/bin/activate
uv pip install -e .

If the vLLM install fails to resolve, the card says to check the version matrix on the vLLM website and pin a combination that matches your CUDA runtime.

The card doesn’t show a download command for the R2T2 weights. It says MODEL_PATH / --model_path accepts a local path or an HF repo id, so you can pass netease-youdao/Confucius4-R2T2 directly. We prefer a local copy, so we fetch the weights ourselves with the hf CLI. The examples below assume this layout:

Terminal window
uv pip install -U "huggingface_hub[cli]"
hf download netease-youdao/Confucius4-R2T2 --local-dir checkpoints/Confucius4-R2T2

The card also offers Docker as the fastest way to get a preconfigured CUDA and runtime environment. R2T2 runs on the official qwenllm/qwen3-asr image, which already ships the runtime libraries it needs. You need the NVIDIA Container Toolkit installed first.

Run it

The quickest test is the bundled script. Audio can be mono or stereo at any sample rate, because it is resampled to 16 kHz internally.

Terminal window
./run_example.sh /path/to/audio.wav \
--model_path checkpoints/Confucius4-R2T2 \
--infer_mode stream_vllm \
--language English \
--chunk_size_ms 160

Use --infer_mode onetime_vllm for offline decoding. Logs go to run_example.log.

Live captions from an audio stream

This example follows the card’s streaming API. It feeds 160 ms chunks, the same way you would forward microphone frames. The card says to keep max_new_tokens small for low-latency streaming.

live_captions.py
import librosa
from qwen_asr import Qwen3ASRModel
CHUNK_SEC = 0.16
SR = 16000
if __name__ == "__main__":
asr = Qwen3ASRModel.LLM(
model="checkpoints/Confucius4-R2T2",
gpu_memory_utilization=0.4,
max_new_tokens=4,
)
wav, sr = librosa.load("meeting.wav", sr=SR, mono=True)
state = asr.init_streaming_state(
context="",
language="English",
unfixed_chunk_num=0,
unfixed_token_num=1,
chunk_size_sec=CHUNK_SEC,
)
step = int(CHUNK_SEC * SR)
for pos in range(0, len(wav), step):
seg = wav[pos : pos + step]
_, text = asr.streaming_transcribe(seg, state, max_new_tokens=2)
print("text:", text)
asr.finish_streaming_transcribe(state)
print("final:", state.text)

The model supports chunks from 80 ms to 2 s. The per-step max_new_tokens=2 above is the value the card’s example uses with 160 ms chunks. The card doesn’t say how to set it for other chunk sizes. The repo’s example.py uses adaptive max_new_tokens and handles lookahead for the first chunk, so start there if you change CHUNK_SEC or write production code.

Domain terms with hotword context

Product names, people’s names and jargon are where ASR usually fails. R2T2 accepts a context string that is prepended to the prompt. The card does not specify a format for this string. A plain list of terms is a reasonable first try.

hotwords.py
import librosa
from qwen_asr import Qwen3ASRModel
if __name__ == "__main__":
asr = Qwen3ASRModel.LLM(
model="checkpoints/Confucius4-R2T2",
gpu_memory_utilization=0.4,
max_new_tokens=4,
)
wav, sr = librosa.load("earnings_call.wav", sr=16000, mono=True)
state = asr.init_streaming_state(
context="Youdao, Confucius4, R2T2, vLLM, EBITDA",
language="English",
unfixed_chunk_num=0,
unfixed_token_num=1,
chunk_size_sec=0.16,
)
step = int(0.16 * 16000)
for pos in range(0, len(wav), step):
asr.streaming_transcribe(wav[pos : pos + step], state, max_new_tokens=2)
asr.finish_streaming_transcribe(state)
print(state.text)

From the command line, set the CONTEXT environment variable for run_example.sh, or pass --context when calling example.py directly.

Batch transcription of recorded files

You don’t need streaming for archived audio. The offline path takes lists, and batched inference is supported. Inputs can be local paths, URLs, base64 data, or (np.ndarray, sr) tuples. The card says adding streaming support did not reduce offline accuracy.

batch_transcribe.py
import librosa
from qwen_asr import Qwen3ASRModel
FILES = ["call_01.wav", "call_02.wav", "call_03.wav"]
if __name__ == "__main__":
asr = Qwen3ASRModel.LLM(
model="checkpoints/Confucius4-R2T2",
gpu_memory_utilization=0.5,
max_inference_batch_size=32,
max_new_tokens=4096,
)
audio = [(librosa.load(f, sr=16000, mono=True)[0], 16000) for f in FILES]
results = asr.transcribe(
audio=audio,
language=[None] * len(audio), # card allows None; it doesn't say what None does
return_time_stamps=False,
)
for path, r in zip(FILES, results):
print(f"{path}\t{r.language}\t{r.text}")

The card shows [None] as an option for language, and each result includes a language field. That suggests the model reports a detected language when you don’t give one, but the card doesn’t say this directly. If you know the language, pass it.

Serving many clients over WebSocket

The repo includes ws_server.py, which serves multi-client streaming. It needs the FireRedVAD streaming detector. Run these commands from the repository root. The download puts the detector in checkpoints/vad/Stream-VAD, which is the default for --vad_model_path, so the server finds it without the flag. The card doesn’t say how the launcher resolves a relative --model_path, so we pass an absolute path:

Terminal window
# from the Confucius4-R2T2 repository root
hf download FireRedTeam/FireRedVAD \
--include "Stream-VAD/*" \
--local-dir checkpoints/vad
./run_start_server.sh start \
--model_path "$(pwd)/checkpoints/Confucius4-R2T2" \
--port 8272 \
--gpu 0
python ws_client.py --uri ws://localhost:8272/asr_stream_api_v1 --audio resources/test.wav

If you put the VAD model somewhere else, pass --vad_model_path explicitly.

Clients send raw 16 kHz mono int16 PCM as binary frames. The reference client sends about 160 ms per frame, which is 2560 samples. To end a session, the client sends the string "YOUDAO_ONETIME_ASR_STREAM_EOS". The server replies with JSON messages whose msg.text field contains only the new text since the previous message. Concatenate these on the client side. Each message also has a reset field, which the card doesn’t explain. Log it, and check the repo’s ws_server.py to see what it means before you rely on plain concatenation.

Gotchas

  • Wrap vLLM code in if __name__ == "__main__":. Without it you hit vLLM’s multiprocessing spawn error.
  • vLLM is picky about versions. It has strict CUDA/PyTorch compatibility requirements. If the install fails to resolve, pin versions using the vLLM compatibility matrix, or use the Docker image, which comes preconfigured.
  • No VRAM numbers. The card’s examples use gpu_memory_utilization=0.4 for streaming and 0.5 for offline. Tune this yourself, especially when sharing a GPU.
  • The default language is Chinese. run_example.sh defaults LANGUAGE to Chinese. Set it explicitly for English audio. The Python API also accepts None.
  • max_new_tokens for other chunk sizes isn’t documented. The card’s example pairs 160 ms chunks with max_new_tokens=2 per step. See example.py for its adaptive max_new_tokens before changing the chunk size.
  • Accuracy beyond Chinese and English is only described as “useful”. The card gives no benchmark numbers for other languages. Evaluate on your own data.
  • The transformers backend is undocumented. The card mentions it but shows no code, so expect to read the repo.
  • Undocumented reset field. WebSocket responses include reset, and the card doesn’t explain it.
  • Docker networking. Services inside the container must bind to 0.0.0.0, not 127.0.0.1, for port mapping to work.
  • Weight license. The weights are not under an open-source license. Read the NetEase Model Use License Agreement before commercial deployment. Apache 2.0 covers only the code.
  • Self-reported evaluations. The comparisons come from NetEase, and the technical report has not been published yet.

When to pick it

Pick R2T2 if:

  • You need live transcription where emitted text must not change, such as captions, simultaneous translation, or voice agents that act on partial transcripts.
  • Your audio is mostly Chinese or English.
  • You can run vLLM on an NVIDIA GPU.
  • You want a ready-made WebSocket server with VAD instead of building streaming infrastructure yourself.
  • In the card’s 160 ms results, it is the most accurate of the open true-streaming models compared. It beats plain Qwen3-ASR at the same chunk size by a wide margin, and its lead over X-ASR and Voxtral ranges from small on some datasets to large on others.

Skip it if:

  • You need CPU or Apple Silicon inference and don’t want to dig into undocumented code. The documented path is CUDA plus vLLM. The card mentions a transformers backend but doesn’t document it.
  • Your main languages aren’t Chinese or English.
  • Your legal team needs an OSI-approved license for the weights.

If you only transcribe recorded files, note that R2T2’s offline accuracy is unverified on the card. The card says adding streaming support didn’t reduce offline accuracy, but it publishes no offline numbers. The tables compare R2T2 in 160 ms streaming mode, not offline mode. Benchmark it against your current offline model on your own audio before switching.

Related