TMax

💻 Code · 🤗 Models & Data · 📜 Paper · 📓 Blog

TMax v1.1 27B

This terminal-agent model was trained using DPPO on Qwen/Qwen3.6-27B with the TMax hardened prefilter task set.

This model is part of the TMax v1.1 collection. The original TMax recipe is described in our paper.

Checkpoints

This branch contains step 400. The main branch contains step 400, the highest-scoring available checkpoint in the completed TB-Lite sweep. Checkpoint selection uses TB-Lite, not Terminal-Bench 2.1.

Checkpoints 100, 200, 300, 400, 500 are provided as branches named step-0100, step-0200, etc.

Evaluation Results

Training step TB-Lite (%) ± SE (pp) Selection
100 73.05 ± 1.56
200 75.18 ± 1.67
300 72.81 ± 1.60
400 76.93 ± 1.55 main
500 71.87 ± 1.64

The selected checkpoint scores 46.21 ± 1.89% on Terminal-Bench 2.1 (88 tasks, three trials per task; configure-git-webserver excluded).

These v1.1 scores use the fixed Vanillux0.2.3 harness with Sandfleet through Harbor, three trials per task (96 TB-Lite tasks). Scores and standard errors are taken from the completed evaluation sweep updated on October 5, 2026. The harness differs from the original release; use the original model cards for v1.0 results. The per-checkpoint results and pinned source revisions are included in release_manifest.json.

Model Details

Model Description

Use

Serve the model with vLLM and use the terminal-agent harness in our codebase:

uvx vllm==0.19.1 serve TMaxxx/qwen-27b-dppo-hardened-prefilter-sc3 \
  --served-model-name tmax-v1.1 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --port 8008 \
  --max-model-len 65536 \
  --tensor-parallel-size 8

The Qwen3.5/3.6 checkpoints contain the text-only causal language model; use the Qwen3 XML tool parser.

Training Details

  • Released training step: 400
  • Main checkpoint: 400, selected on TB-Lite
  • Available milestone checkpoints: 100, 200, 300, 400, 500

The source checkpoint is TMaxxx/qwen-27b-dppo-hardened-prefilter-sc3 at step 400. Original checkpoint files are preserved, including campaign_provenance.json where supplied. See release_manifest.json for the immutable source commits and evaluation results. See the training launch scripts and original provenance for run-specific settings.

License

This model is licensed under Apache 2.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.

Citation

If you use our model or data, please cite our paper:

@misc{ivison2026tmaxsimplerecipeterminal,
      title={Tmax: A simple recipe for terminal agents}, 
      author={Hamish Ivison and Junjie Oscar Yin and Rulin Shao and Teng Xiao and Nathan Lambert and Hannaneh Hajishirzi},
      year={2026},
      eprint={2606.23321},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2606.23321}, 
}
Downloads last month
33
Safetensors
Model size
2.65M params
Tensor type
BF16
·
Video Preview
loading

Model tree for TMaxxx/qwen-27b-dppo-hardened-prefilter-sc3

Base model

Qwen/Qwen3.6-27B
Finetuned
(411)
this model

Dataset used to train TMaxxx/qwen-27b-dppo-hardened-prefilter-sc3

Paper for TMaxxx/qwen-27b-dppo-hardened-prefilter-sc3