Instructions to use appautomaton/confucius4-r2t2-bf16-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use appautomaton/confucius4-r2t2-bf16-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir confucius4-r2t2-bf16-mlx appautomaton/confucius4-r2t2-bf16-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Confucius4-R2T2 — MLX (bf16)
MLX-native bf16 conversion of NetEase Youdao's Confucius4-R2T2, a streaming speech recognition model finetuned from Qwen3-ASR-1.7B. It runs through the mlx-speech runtime on Apple Silicon with no PyTorch, vLLM, or cloud API at inference time.
Any modifications made to the original model in this Derivative Work are not endorsed, warranted, or guaranteed by the original right-holder of the original model, and the original right-holder disclaims all liability related to this Derivative Work.
Model Details
- Upstream developer: NetEase Youdao (
netease-youdao/Confucius4-R2T2) - MLX conversion and runtime: App Automaton
- Architecture: Qwen3-ASR graph (audio encoder + 1.7B text decoder)
- Precision: bf16. The weights are not quantized; keys are remapped to the MLX module tree and audio Conv2D weights transposed to MLX layout.
- Input: 16 kHz mono audio
- Languages: 30, as declared by upstream
How to Get Started
pip install "mlx-speech>=0.5.3"
Streaming — feed microphone PCM as it arrives:
import mlx_speech
asr = mlx_speech.asr.load("confucius4-r2t2")
session = asr.stream_session(language="English", chunk_ms=160, lookahead_ms=160)
for pcm_chunk in microphone: # float32, 16 kHz mono, any length
update = session.feed(pcm_chunk)
print(update.committed) # append-only committed text
print(session.finalize().text)
Offline:
result = asr.generate("speech.wav")
print(result.text)
Or download once and load by path:
hf download appautomaton/confucius4-r2t2-bf16-mlx \
--local-dir models/netease/confucius4_r2t2/mlx-bf16
Streaming
mlx-speech implements NetEase's chunk loop: prefix rollback, per-window token budget, and append-only commit. Each window reuses the previous window's mel frames, closed audio-encoder blocks, and decoder KV prefix, and recomputes only what changed.
Notes
- This repo contains the MLX runtime artifact only (no PyTorch checkpoint).
License
The model weights are licensed under the NetEase Youdao Model Use License Agreement; a copy is included in this repository as LICENSE (original: MODEL_LICENSE). By using these weights you agree to its terms, including the restrictions on large-scale commercial use (section 2.2) and prohibited high-risk use (section 4.2). Downstream redistribution must retain the license and this notice.
The mlx-speech runtime is licensed separately under its own terms.
- Downloads last month
- 25
Quantized