Text Generation
Transformers
Safetensors
deepseek_v3
conversational
custom_code
text-generation-inference
fp8
Instructions to use tngtech/DeepSeek-TNG-R1T2-Chimera with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use tngtech/DeepSeek-TNG-R1T2-Chimera with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="tngtech/DeepSeek-TNG-R1T2-Chimera", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("tngtech/DeepSeek-TNG-R1T2-Chimera", trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained("tngtech/DeepSeek-TNG-R1T2-Chimera", trust_remote_code=True, device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use tngtech/DeepSeek-TNG-R1T2-Chimera with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "tngtech/DeepSeek-TNG-R1T2-Chimera" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tngtech/DeepSeek-TNG-R1T2-Chimera", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/tngtech/DeepSeek-TNG-R1T2-Chimera
- SGLang
How to use tngtech/DeepSeek-TNG-R1T2-Chimera with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "tngtech/DeepSeek-TNG-R1T2-Chimera" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tngtech/DeepSeek-TNG-R1T2-Chimera", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "tngtech/DeepSeek-TNG-R1T2-Chimera" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "tngtech/DeepSeek-TNG-R1T2-Chimera", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use tngtech/DeepSeek-TNG-R1T2-Chimera with Docker Model Runner:
docker model run hf.co/tngtech/DeepSeek-TNG-R1T2-Chimera
File size: 5,217 Bytes
5577f86 934649a 5577f86 934649a e40ca83 934649a 7531624 934649a 18ac100 3a9f1f8 3b51f47 3a9f1f8 3b51f47 f60e279 f6ba7e0 975dbeb 3b51f47 f6ba7e0 3b51f47 f6ba7e0 3b51f47 213b8a6 3a9f1f8 a5907ba 934649a a5907ba 934649a 6e71df0 3ad73cf 7531624 6e71df0 934649a 975dbeb f6ba7e0 3b51f47 934649a 975dbeb 162fed2 975dbeb 3c3d48d 934649a 3c3d48d 934649a 3a9f1f8 934649a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 | ---
license: mit
library_name: transformers
base_model:
- deepseek-ai/DeepSeek-V3-0324
- deepseek-ai/DeepSeek-R1
- deepseek-ai/DeepSeek-R1-0528
pipeline_tag: text-generation
---
# DeepSeek-TNG-R1T2-Chimera
<div align="center">
<img src="https://354918363417-runtime-assets.s3.eu-central-1.amazonaws.com/company_logo_light.svg"
alt="TNG Logo"
width="400"
style="display: inline-block; vertical-align: middle;"/>
</div>
<br>
<div align="center">
<a href="LICENSE" style="margin: 2px;">
<img alt="License" src="https://img.shields.io/badge/License-MIT-f5de53?&color=f5de53" style="display: inline-block; vertical-align: middle;"/>
</a>
</div>
<br>
<div align="center">
<img alt="Intelligence Score" src="intelligence_score_vs_output_tokens.png" style="display: inline-block; vertical-align: middle;" width="750"/>
</div>
**Assembly of Experts Chimera model constructed with the DeepSeek [R1-0528](https://huggingface.co/deepseek-ai/DeepSeek-R1-0528), [R1](https://huggingface.co/deepseek-ai/DeepSeek-R1) and [V3-0324](https://huggingface.co/deepseek-ai/DeepSeek-V3-0324) parent models**
We present our new **DeepSeek-TNG R1T2 Chimera** 671B model, the first successor to our original [*DeepSeek R1T Chimera*](https://huggingface.co/tngtech/DeepSeek-R1T-Chimera) that was released on April 26th. Unlike the original Chimera, which was based on the *two parent models* V3-0324 and R1, the new Chimera is a **Trimind** *with three parents*, namely additionally R1-0528. It is constructed using the Assembly of Experts-method with relatively fine-granular direct brain edits. This more refined assembly allowed, among other improvements, the fixing of the <think> token consistency issue, which was a weakness of R1T and is now solved for R1T2.
**Sweet spot**
R1T2 operates at a new sweet spot in intelligence vs. output token length. It appears to be...
- about **20% faster than** the regular **R1**, and more than **twice as fast as R1-0528**
- significantly **more intelligent than** the regular **R1** in benchmarks such as **GPQA** and **AIME-24**
- much **more intelligent** and also **think-token consistent** compared to the first **R1T Chimera** 0426
- and generally well-behaved and a **nice persona** to talk to, even without any system prompt.
**Recommendations for your model decision**
*R1T2* compared...
- *vs R1:* We hope that R1T2 is a very desirable, almost universal **better and drop-in replacement for R1**
- *vs R1-0528:* R1T2 is a much **cheaper alternative to full R1-0528**, if the fullest 0528-level intelligence is not required
- *vs R1T:* R1T2 is usually **recommended over R1T**, unless the specific personality of R1T was optimal, the think-token issue not important, or R1T's higher speed crucial
- *vs V3-0324:* V3 is so much faster that if you can live with the **lower intelligence, take V3**, however, if you **need reasoning, R1T2** is the go-to model
**Limitations**
- **R1-0528** is thinking much longer, but also is achieving **better hard benchmark results** than R1T2
- As measured by SpeechMap.ai (courtesy of xlr8harder), **R1T2** is significantly **more reserved** than R1T, but not as much as R1-0528
- Due to the influence of its R1 parent, which does not support function calling, **R1T2 is not yet recommended for function-calling** intensive applications at this stage (this may be fixed at a later stage)
**Technological background**
For details on the AoE construction process, you can read our [Paper on arXiV](https://arxiv.org/abs/2506.14794).
## Model Details
- **Architecture**: DeepSeek-MoE transformer-based language model
- **Combination Method**: Assembly of Experts from the three DeepSeek parent models R1-0528, R1 and V3-0324
- **Release Date**: 2025-07-02
- **Design Team**: Robert Dahlke, Henrik Klagges, Benjamin Merkel, Fabian Klemm and David Reiss, Munich, Germany
- **Extra Thanks**: Big thanks to DeepSeek for their great models and open-source generosity, and to the other researchers that have published on model merging methodologies.
## Use, Out-of-scope Use, Other Limitations, Risks, Recommendations et al.
Regarding the R1T/R1T2-Chimeras, we ask you to follow the careful guidelines that Microsoft has created for their "MAI-DS-R1" DeepSeek-based model.
These professional guidelines are available [here on Hugging Face](https://huggingface.co/microsoft/MAI-DS-R1).
## EU AI Act
Due to the strict new guidelines of the EU AI Act that take effect on August 2nd 2025, we recommend that each R1T/R1T2 user in the EU either familiarizes themselves with these requirements and assess their compliance, or ceases using the model in the EU after August 1st, 2025.
## Contact, especially for your user feedback
Please give us your feedback, especially if you find deficiencies in the model:
- Email: research@tngtech.com
- X.com: @tngtech
## Citation
```
@misc{tng_technology_consulting_gmbh_2025_07_0x,
author = { TNG Technology Consulting GmbH },
title = { DeepSeek-TNG-R1T2-Chimera },
year = 2025,
month = { July },
url = { https://huggingface.co/tngtech/DeepSeek-TNG-R1T2-Chimera },
doi = { xxx },
publisher = { Hugging Face }
}
``` |