Model deep dive·
Run Confucius4-R2T2 for Append-Only Real-Time Speech Recognition
NetEase Youdao's streaming ASR model built on Qwen3-ASR: transcript text never gets rewritten, with 80 ms to 2 s chunks. Setup, code and caveats.
netease-youdao/Confucius4-R2T2