Text Generation
Transformers
Safetensors
English
code
qwen3
coderion
coding
reasoning
small-language-model
0.6b
chronological-reasoning
high-reasoning
compact-model
conversational
text-generation-inference
Instructions to use OrionLLM/NanoCoder-0.6b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OrionLLM/NanoCoder-0.6b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="OrionLLM/NanoCoder-0.6b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("OrionLLM/NanoCoder-0.6b") model = AutoModelForCausalLM.from_pretrained("OrionLLM/NanoCoder-0.6b", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use OrionLLM/NanoCoder-0.6b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OrionLLM/NanoCoder-0.6b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OrionLLM/NanoCoder-0.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/OrionLLM/NanoCoder-0.6b
- SGLang
How to use OrionLLM/NanoCoder-0.6b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "OrionLLM/NanoCoder-0.6b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OrionLLM/NanoCoder-0.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "OrionLLM/NanoCoder-0.6b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OrionLLM/NanoCoder-0.6b", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use OrionLLM/NanoCoder-0.6b with Docker Model Runner:
docker model run hf.co/OrionLLM/NanoCoder-0.6b
File size: 1,776 Bytes
c729bfb 123017a c729bfb 123017a c729bfb 123017a c729bfb 123017a c729bfb 123017a 67cd2de 123017a c729bfb 123017a c729bfb 123017a c729bfb c628203 c729bfb 123017a c729bfb c628203 c729bfb 123017a c729bfb 123017a c729bfb 123017a c729bfb 123017a c729bfb 123017a c729bfb a970605 c729bfb 123017a c729bfb 123017a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 | ---
language:
- en
- code
pipeline_tag: text-generation
license: apache-2.0
tags:
- coderion
- code
- coding
- reasoning
- small-language-model
- 0.6b
- chronological-reasoning
- high-reasoning
- compact-model
library_name: transformers
datasets:
- nvidia/OpenCodeReasoning
base_model:
- Qwen/Qwen3-0.6B
---
<p align="center">
<img src="https://huggingface.co/proxy/cdn-uploads.huggingface.co/production/uploads/685ea8ff7b4139b6845ce395/YF0kEDYMGJhcM3Lbl2EOD.png" alt="logo" width="250">
</p>
<p align="center"><b>A compact 0.6B coding model built for strong reasoning efficiency.</b></p>
---
**NanoCoder** is a **small 0.6B parameter coding-focused language model** designed for **high and xhigh chronological reasoning** in programming tasks.
It is built to deliver **surprisingly strong structured reasoning and coding performance for its size**, focusing on consistency, logical step progression, and efficient problem solving.
While **NanoCoder is not intended to be a general everyday assistant**, it is a **small but capable specialist model** that performs well within its class and remains **reliable for compact code reasoning workloads**.
---
## Key Characteristics
- **0.6B parameters**
- **Dedicated to code**
- **Optimized for high reasoning intensity**
- **Chronological reasoning style**
- **Strong consistency for a compact model**
- **Designed for efficient performance despite its small size**
---
## Limitations
NanoCoder is a **small specialized model**.
Because of that:
- It may not match larger models on broad real-world assistant tasks
- It is not primarily designed for daily casual use
- It performs best when used for **focused coding and reasoning workloads**
- Its main strength is **efficiency, consistency, and reasoning quality relative to size** |