Introduction
If you use AI to write code, you have probably used the same model through more than one tool: Claude models in Claude Code, Cursor, Pi, or Cline. And you may have noticed it doesn’t behave the same way in each. It plans differently, reaches for different tools, and finishes tasks in one that it gets stuck on in another.
The difference can be measured, and it comes from the program wrapped around the model, called an agent harness. It runs the loop, decides which tools the model gets, writes the context the model reads, makes sense of what the model sends back, and decides when to stop. Change the harness, and you change what the model sees and what it is allowed to do, so the results change too. On SWE-bench Pro, Joel Niklaus measured the same model, GLM-5.2, at 23% in one harness and 52% in another. The ranking doesn’t carry over between models either: Codex was the second-best of ten harnesses for GLM-5.2 and the ninth for Gemma 4 26B-A4B.
The harness matters most for the open-weight models you run yourself. A model that was never trained in your harness can struggle there. It calls tools the harness doesn’t have, or writes output the harness can’t read. Training it in one harness doesn’t fix this either, because it then learns that harness’s habits and struggles in the others.
The fix is to train the model inside the harnesses themselves, with reinforcement learning on multi-step tool use, which is called agentic reinforcement learning (agentic RL). The question behind this work is whether you can take a small open model, train it that way inside the harnesses people actually use, such as Claude Code, Codex and OpenCode, and see it improve across all of them. The harnesses run exactly as they ship: none of them was written with training in mind, and we don’t change their code. This article introduces an open framework that makes it possible, built on OpenEnv, with Harbor supplying the tasks and sandboxes and TRL as the trainer.
The article has two parts. The first is a guide: the basics of harnesses, agentic RL and RL environments, then why models need training across harnesses, with what recent papers and model reports show. The second is our own work: how the framework works, and what happened when we used it to train a small open model, LFM2.5-2.6B, across four harnesses at once. Trained across all four, it went from solving 42% of held-out tasks to 54%, averaged over the four harnesses, and used 31% fewer tool calls on the tasks it already solved. The same model trained in OpenCode alone improved mostly in OpenCode, and much less in the three harnesses it never trained in. We also compare RL with supervised fine-tuning (SFT) on a larger model’s successful rollouts, which helped less. Our earlier runs with Qwen3.5-2B are in a collapsible section of the training chapter.
The basics
What is a harness?
A 2026 position paper on benchmark disclosure defines an agent harness as “the software layer between the model and the task that constructs the context the model sees, mediates its tool calls, validates its outputs, and decides when to retry, escalate, or stop” (Zhang et al., 2026). To the model, the harness is its interface to the task. To a trainer, it is a program someone else wrote, which the trainer cannot change.
A few neighbouring terms are easy to mix up with it, so this is how the article uses them:
- The policy is the model being trained: the weights that turn a prompt into a completion, with no loop and no memory between calls. We also call it the model.
- The agent is the policy running inside a harness, the whole program that plans, calls tools and decides when to stop.
- The sandbox is the isolated place where actions actually run, such as a container, a VM or a tmux session.
- A benchmark is a fixed task set plus a protocol for comparing results on it.
Hugging Face’s agent glossary covers the same terms in more detail.
The model chooses the tool and its arguments. The harness executes the call and adds the result back into the model’s context before asking it what to do next.
What is agentic RL?
Agentic RL is reinforcement learning for a model that acts over many steps. The actions are tool calls, the observations are tool results, and the reward comes from a grader that looks at the finished trajectory. A single rollout has this shape: write, run, read the result, decide what to try next, and eventually submit something a grader can score. The trainer runs a group of these rollouts for each task and learns from how their rewards compare.
What are RL environments?
An RL environment is the stateful system a policy acts on. The environment takes an action, updates its internal state, returns an observation, and at some point returns a reward. So the environment owns the state, the transitions and the reward, while a harness runs the loop.
The figure below puts the training loop, which samples, scores and updates, next to the environment it acts on.
The environment box holds a task set, a prompt template, tools, observations, an execution backend, world state, a reward rule and a termination condition. Frameworks do not always give each piece a name. In the box labelled “Tools / Harness”, harness means the tool surface, not the whole agent program. This figure and the rollout above come from The ultimate guide to RL environments, which covers each component with an example if you want to go deeper.
White-box vs black-box RL environments
Who actually drives the rollout?
In a white-box environment, the trainer owns the loop. The trainer samples an action, calls
env.step(), reads the observation, and then samples again. The environment waits to be called.
Every token the policy produced is already in the trainer’s hands, because the trainer is the one
that sampled it. All six frameworks compared in the earlier guide work this way, and it is what TRL’s
GRPOTrainer expects when you hand it an environment factory.
In a black-box environment, the harness owns the loop. It starts inside a sandbox, runs its own loop, calls its own tools, compacts its own context, and stops when it decides to. It is an agent program that was never written with training in mind. The trainer is outside that box. All it sees is a sequence of calls arriving at a model endpoint.
Microsoft’s Agent Lightning team named the two modes (He et al., 2026):
“In traditional agentic RL, the training engine owns the environment interaction loop. In harnessed agentic RL, the harness owns this loop, while the training engine observes only a sequence of LLM request-response pairs.”
Harnesses differ wildly on the inside, but every one of them has to call a model, and the model API is the one interface guaranteed to exist outside the harness. Everything else, the harness’s state and the environment’s, is hidden from the trainer.
The KAT-Coder report (KwaiKAT Team, 2026) uses these terms to describe how a harness organises its trajectories. It calls Mini-SWE-Agent white-box because it has a simple loop without trajectory compression, and Claude Code, Codex, OpenClaw and OpenHands black-box because they compress and reorganise their context. In this article, the distinction is who owns the rollout loop: the trainer or the harness. Context handling and ownership of the loop are separate choices.
Moving the loop into the harness changes two things. The first is who makes each call, and in what order.
The second is what the trainer ends up recording. A harness could report events the trainer did not cause, but nothing obliges it to. The figure below shows one rollout three times: as it happened, and as each architecture records it.
Why rewards are not enough
An on-policy policy gradient (Williams, 1992) needs two things per token: which token was sampled, and the probability it was sampled with. A harness that hands back text and a score gives you neither.
Different token sequences can produce the same text. If you save only the text and tokenize it again, you can get different token IDs from those the model sampled. TRL’s write-up puts it directly: “in RL, you optimize on the exact tokens the model produced” (Gallouédec & Rasul, 2026).
When a harness builds the next prompt, it may add role markers or change whitespace in a tool call. Some harnesses also repair malformed JSON. If the trainer learns from this edited text, it may update the model using a response it never generated.
The probabilities need to be recorded during generation too. In asynchronous RL, the model may have been updated since a rollout was sampled. Recomputing log probabilities with the current weights does not recover the probabilities used during that rollout. Importance sampling compares the current probability with the one used during generation; using the current value for both would make their ratio 1 even when the policies differ.
To avoid both errors, the sampler must return token ids and per-token log probabilities, and the trainer must keep them. Never re-encode text you decoded. Keep the sampled ids and append the template’s suffix by id concatenation. This works because most chat templates extend token for token when a tool message is appended, and eighteen of the nineteen models tested in TRL’s write-up do so unmodified. It is also why a system that wants the exact tokens has to intercept at the model endpoint rather than at the text boundary.
Other teams keep exact tokens from the trainer’s side instead. Prime Intellect’s renderers library turns a model’s chat template into Python code that renders messages to token ids and extends a multi-turn rollout without re-rendering what the model sampled. It needs the trainer to own the loop, and its announcement warns that harnesses which repair tool calls or compact their history break it, which is the case the capture proxy is built for.
Not everyone agrees the token-level machinery is necessary at all. OpenForgeRL builds the same proxy architecture and reconstructs training samples from the prompt and response pairs the proxy collects. The paper never mentions token ids or log probabilities, and it still reports the strongest multi-harness results published so far (Yu et al., 2026). This article takes the token-faithful side, and the question is not settled.
- Gallouédec, Q., & Rasul, K. (2026). Agentic RL: Token-In, Token-Out Done Right. https://huggingface.co/blog/huggingface/tito
- He, Z., Zhang, S., Zhou, Z., Yang, Y., Kang, Y., Zhang, Y., Qiu, L. K., Tsui, T. Y., Xu, J., & Luo, C. (2026). Agent Lightning v1.0: Towards Harnessed Agentic RL. arXiv Preprint arXiv:2608.17528. https://huggingface.co/papers/2608.17528
- KwaiKAT Team. (2026). KAT-Coder-V2.5 Technical Report. arXiv Preprint arXiv:2607.05471. https://huggingface.co/papers/2607.05471
- Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8(3), 229–256. 10.1007/BF00992696
- Yu, X., Peng, B., Xu, R., Zou, H., Wu, Q., Cheng, H., Yao, W., Singh, N., Yu, Z., & Gao, J. (2026). OpenForgeRL: Train Harness-native Agents in Any Environment. arXiv Preprint arXiv:2607.21557. https://huggingface.co/papers/2607.21557
- Zhang, Y., Wang, J., Ge, Y., Xu, W., Hamm, J., & Reddy, C. K. (2026). Stop Comparing LLM Agents Without Disclosing the Harness. arXiv Preprint arXiv:2605.23950. https://huggingface.co/papers/2605.23950
Why train across harnesses
Benchmark scores now come with a harness attached
Model cards have started saying which harness a score came from. GLM-4.7 promises “significant improvements on complex tasks in mainstream agent frameworks such as Claude Code, Kilo Code, Cline, and Roo Code” (Z.ai, 2025). Kimi K2 reports Terminal-Bench (Merrill et al., 2026) twice: 25.0 under Terminus and 30.0 under Moonshot’s own framework (Kimi Team, 2025). MiniMax M2 names a harness for almost every benchmark it lists (MiniMax, 2025). DeepSeek-V3.2’s thinking mode wouldn’t run under Terminus at all, so its Terminal-Bench number had to come from a different harness (DeepSeek-AI, 2025).
The harness can even change which model comes out ahead. Zhang et al. (2026) took three models that sit within three points of each other on a public coding leaderboard: GLM-5.1, GPT-5.4, and Kimi K2.6. They ran all three on the same 100 SWE-bench Verified (Jimenez et al., 2024) tasks under three harness configurations, with everything else held fixed. The configurations build on each other: Minimal has verbose tools, no context compression and no retries, Improved trims the tools and adds compression and retries, and Full adds self-checks, drift checks and rollback.
Each model comes out on top under a different harness. Swapping the harness moved GLM-5.1 by 13 points, while swapping the model inside the same harness moved the score by only 2.5 to 5. Joel Niklaus found the same across ten harnesses on SWE-bench Pro.
You can see the same thing outside a controlled test. Claude Opus 4.5 scores 45.9% on SWE-bench Pro (Deng et al., 2025) on Scale’s SEAL leaderboard and 55.4% inside Claude Code, with the same weights (Zhang et al., 2026).
When a model leaves the harness it was trained in
A model can perform well in the harness it was trained in and struggle when moved to another. The Orchard paper measured this transfer across harnesses (Peng et al., 2026). It took two open coding models trained with one or two harnesses and ran them under three, including Kimi-CLI, which none of the systems compared had used in training. Orchard’s own model ran alongside them. It was fine-tuned on trajectories from OpenHands and Mini-SWE-Agent, then trained with RL in Mini-SWE-Agent.
Moving OpenSWE-32B from OpenHands, the harness the paper treats as its native one, to Mini-SWE-Agent costs 7.5 points. Under Kimi-CLI, it drops 58.8 points, to 3.6% on SWE-bench Verified, and scores zero on Terminal-Bench 2.0. Scale-SWE stops producing valid tool calls anywhere but its home harness. Orchard’s model stays between 45.0 and 64.3 on SWE-bench Verified across all three.
The paper calls these a degraded resolve rate, where the model still works but solves less, and a catastrophic format failure, where its output stops being usable at all. Each row of that figure is one fixed model, and each benchmark uses the same tasks under every harness, so the drop comes from the harness.
Why models overfit to a single harness
The KwaiKAT team names the cause (KwaiKAT Team, 2026):
“In agentic RL, if training relies solely on a single fixed harness, the model often learns not ‘how to solve the task’ but ‘how to solve the task under that particular harness’s interface conventions.’”
They split it into three kinds of overfitting: to the harness’s action format, to its context structure, and to its control flow, the retries and stop conditions the model learns to lean on.
Harnesses differ along exactly those lines. They speak different API dialects, meaning the request and response format they use to talk to a model server: OpenAI chat-completions, OpenAI Responses, Anthropic Messages, Google Gemini. They differ in whether they use tool calling at all, how they compact context, and how much of the control flow they take away from the model.
Mini-SWE-Agent’s README lists avoiding scaffold overfitting as a reason for its simple design. Aider uses a different interface: it asks the model to write edit blocks in prose instead of structured tool calls. A model trained on tool calls has to learn that output format too.
In practice, the failure often shows up as an error rather than a lower score. A model trained in one harness calls a tool by that harness’s name for it, with that harness’s argument names, and a harness that spells them differently rejects the call before the tool ever runs.
The list of harnesses keeps growing. Harbor 0.22.0 ships adapters for more than 40 of them, and OpenEnv has validated ten end-to-end, so a model trained on a few will keep meeting ones it never saw.
Frontier labs now train across harnesses on purpose
KwaiKAT’s answer, which they call harness scaling, is to vary the harness during training. Poolside’s Laguna report does this (Poolside, 2026):
“To encourage generalization across diverse agent harnesses, our training data includes trajectories from external frameworks such as OpenHands, OpenCode, and Mini-SWE-Agent.”
They added 1.3 billion tokens of these trajectories to their supervised training and kept each
harness’s own habits in on purpose. Reinforcement learning then runs only in the harness they ship in
production, pool, and the report doesn’t say why the other harnesses stop there.
Kimi K3’s RL environment builds harnesses such as Claude Code and Codex from composable modules, so the model trains on many harness configurations instead of one (Kimi Team et al., 2026). Qwen3-Coder-Next generates its agentic coding data in six different harnesses (Qwen Team, 2026), and MiMo-V2.6 trains on a pool of task-adapted mini-harnesses (LLM-Core Xiaomi, 2026). Liquid AI does the same at a much smaller size. For LFM2.5-2.6B, the model we train later, Liquid picks a harness at random for each training task. Their release post gives the reason: “training directly inside Hermes Agent, OpenClaw, and other harnesses exposes the model to their tools, system prompts, and interaction patterns, helping it work reliably across agent environments.”
OpenForgeRL checked whether that variety costs peak performance. It trained the same model on one harness and on three, and the three-harness model scored higher even on the single harness’s home turf, 48.5 to 46.0, and almost twice as high under Codex (Yu et al., 2026).
Model reports have only recently started saying which harness a number came from. In 2025, most named none or one per benchmark. By 2026, every report on the timeline below trains across several.
What is missing
What is still missing is an open, reliable way to do this with any harness. It has to capture everything RL training uses from harnesses you don’t control, and plug into whichever trainer you already use.
It also doesn’t take a frontier-sized budget. Agent Lightning takes Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified with roughly 6,000 training examples (He et al., 2026). The same paper names the hard part. The harness owns the loop, so the trainer only sees requests and responses going past, and turning those back into training samples is still an open problem.
We built this on OpenEnv (Meta PyTorch & Hugging Face, 2025), around a capture proxy that sits between the harness and the model. Every call the agent makes passes through the proxy, which records the prompt and completion tokens as the model saw and produced them, with their log probabilities. It speaks the API formats coding agents use, so Claude Code, Codex, and OpenCode run unmodified, and a custom harness only has to point at it. Because it all sits behind OpenEnv’s open environment interface, the same setup can feed different trainers, and ours was TRL (von Werra et al., 2020).
On top of that, the Harbor integration (Kolavi, 2026) serves Harbor’s containerized tasks as OpenEnv environments, so the harness and the sandbox become settings you choose per rollout.
- DeepSeek-AI. (2025). DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models. arXiv Preprint arXiv:2512.02556. https://huggingface.co/papers/2512.02556
- Deng, X., Da, J., Pan, E., He, Y. Y., & others. (2025). SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? arXiv Preprint arXiv:2509.16941. https://huggingface.co/papers/2509.16941
- He, Z., Zhang, S., Zhou, Z., Yang, Y., Kang, Y., Zhang, Y., Qiu, L. K., Tsui, T. Y., Xu, J., & Luo, C. (2026). Agent Lightning v1.0: Towards Harnessed Agentic RL. arXiv Preprint arXiv:2608.17528. https://huggingface.co/papers/2608.17528
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? The Twelfth International Conference on Learning Representations. https://huggingface.co/papers/2310.06770
- Kimi Team. (2025). Kimi K2: Open Agentic Intelligence. arXiv Preprint arXiv:2507.20534. https://huggingface.co/papers/2507.20534
- Kimi Team, Bai, T., Bai, Y., Bao, Y., & others. (2026). Kimi K3: Open Frontier Intelligence. arXiv Preprint arXiv:2607.24653. https://huggingface.co/papers/2607.24653
- Kolavi, A. S. (2026). Harbor integration: serve Harbor task datasets through OpenEnv as trainable environments. https://github.com/huggingface/OpenEnv/pull/1036
- KwaiKAT Team. (2026). KAT-Coder-V2.5 Technical Report. arXiv Preprint arXiv:2607.05471. https://huggingface.co/papers/2607.05471
- LLM-Core Xiaomi. (2026). MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement. https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL/blob/main/MiMo_V2_6_technical_report.pdf
- Merrill, M. A., Shaw, A. G., Carlini, N., Li, B., & others. (2026). Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces. arXiv Preprint arXiv:2601.11868. https://huggingface.co/papers/2601.11868
- Meta PyTorch, & Hugging Face. (2025). OpenEnv: Agentic Execution Environments. https://github.com/huggingface/OpenEnv
- MiniMax. (2025). MiniMax-M2 Model Card. https://huggingface.co/MiniMaxAI/MiniMax-M2
- Peng, B., Yao, W., Wu, Q., Cheng, H., Yu, X., Yang, R., Ge, T., Sordoni, A., Yuan, X., Shen, Y., He, P., Zhang, T., Yu, Z., & Gao, J. (2026). Orchard: An Open-Source Agentic Modeling Framework. arXiv Preprint arXiv:2605.15040. https://huggingface.co/papers/2605.15040
- Poolside. (2026). Laguna M.1/XS.2 Technical Report. arXiv Preprint arXiv:2605.27605. https://huggingface.co/papers/2605.27605
- Qwen Team. (2026). Qwen3-Coder-Next Technical Report. arXiv Preprint arXiv:2603.00729. https://huggingface.co/papers/2603.00729
- von Werra, L., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., Huang, S., Rasul, K., & Gallouédec, Q. (2020). TRL: Transformers Reinforcement Learning. https://github.com/huggingface/trl
- Yu, X., Peng, B., Xu, R., Zou, H., Wu, Q., Cheng, H., Yao, W., Singh, N., Yu, Z., & Gao, J. (2026). OpenForgeRL: Train Harness-native Agents in Any Environment. arXiv Preprint arXiv:2607.21557. https://huggingface.co/papers/2607.21557
- Z.ai. (2025). GLM-4.7. https://huggingface.co/zai-org/GLM-4.7
- Zhang, Y., Wang, J., Ge, Y., Xu, W., Hamm, J., & Reddy, C. K. (2026). Stop Comparing LLM Agents Without Disclosing the Harness. arXiv Preprint arXiv:2605.23950. https://huggingface.co/papers/2605.23950 back: 1, 2
How the framework works
The framework has three parts. OpenEnv is the interface the harness, the environment and the trainer all connect to, and its capture proxy records the tokens. Harbor supplies the tasks and the sandboxes. TRL trains on what the proxy recorded. This chapter follows one rollout through all of them, then covers each part in turn.
OpenEnv
OpenEnv is a standard interface for RL environments (Meta PyTorch & Hugging Face, 2025). It offers Gymnasium-style reset, step, and
state over a client and server transport, with typed actions and observations, packaged as Docker
and publishable to the Hub. It began as a Meta PyTorch and Hugging Face collaboration and now lives
at huggingface/OpenEnv, steered by a committee of twelve organisations, under BSD-3-Clause.
In this article, OpenEnv is the shared interface that the harness, the environment, and the trainer all connect to. It does not train anything or define rewards. The capture proxy records what RL training needs from any agent that calls a model API. The Harbor integration (Kolavi, 2026) serves Harbor’s containerised tasks as OpenEnv environments, with the harness and sandbox chosen per rollout.
Training through the typed TrainingTrace API needs OpenEnv 0.7.0
or later and Python 3.12 for the Harbor extra. On Python 3.10 or 3.11, openenv[harbor] does not
install the Harbor dependencies. The harbor_env package also comes from the repository’s envs/
directory, not the OpenEnv wheel. For a checkout from main:
# Run in a Python 3.12 virtual environment.
git clone --depth 1 --branch main https://github.com/huggingface/OpenEnv.git
python -m pip install -e './OpenEnv[harbor]'
export PYTHONPATH="$PWD/OpenEnv/envs${PYTHONPATH:+:$PYTHONPATH}"
See TRL’s Harbor example for the trainer and inference-server dependencies and the two-GPU launch commands.
One rollout, end to end
Before looking at each part, here is the whole path. The figure follows one rollout, three turns long, of a task from SmolDataEnvs, our suite of data-analysis tasks built from real Kaggle notebooks, through every part of the system. Switching the harness changes only how Harbor wires it to the proxy, the format of its model calls and replies, and the names of its tools.
The trainer and the inference server, vLLM (Kwon et al., 2023), each get their own GPU. The OpenEnv server can run next to them or on a Hugging Face Space, and every rollout gets its own sandbox. The hardware for our runs is in the training setup. The harness inside the sandbox never talks to vLLM or to the trainer. It reaches the model only through the capture proxy, and the only credential it holds for that is a session id.
The capture proxy
A harness finds the proxy the way it would find any model provider, through a base URL and an API key. The API key is a capture session id minted for that rollout, which is how one proxy on one port serves every concurrent rollout. A call carrying a key the proxy never registered gets a 401.
Harnesses speak four different model APIs, as the harness table shows. The proxy works out which one a request uses from its path, then its headers, then the shape of its body. It converts the request to chat completions with converters vendored from NVIDIA’s Polar gateway (Xu et al., 2026), and replays the answer in the caller’s own format.
The call to the engine never streams. The proxy waits for the whole completion, stores it, and replays it to the harness as a stream if the harness asked for one. Apart from a later first token, the harness cannot tell the difference, and the proxy never has to parse partial deltas.
What gets recorded
For each call the proxy asks the engine for the prompt’s token ids, the sampled token ids and a
logprob for every sampled token. It also samples from the full distribution, with no top_p or
top_k. Truncating the distribution biases what the policy samples away from what it would sample
on its own, which in RL is a known route to entropy collapse, and vLLM computes its processed
logprobs after the truncation, so they would not match the policy either. In our runs, moving top_p
to 1.0 took the importance ratio from 0.985–0.993 to 0.9984–0.9999. On vLLM this needs two server
flags:
vllm serve <model> --return-tokens-as-token-ids --logprobs-mode processed_logprobs
The proxy does not take any of this on trust. The first time it sees an engine it sends a probe and
grades the result, and it also checks whether the logprobs are raw or processed. An engine below the
tokens level still serves rollouts, but they are marked as evaluation only, and asking for training
data from one raises an error. Hosted APIs such as OpenAI, Anthropic and Hugging Face Inference
Providers land there, so they can evaluate a harness but not train through it.
From calls to training sequences
Harnesses retry, spawn subagents and compact their context, so a rollout is rarely one clean conversation. The proxy stores each model call as a node and links it to the earlier call whose prompt and completion form the longest exact token prefix of its own prompt. A retry shows up as a sibling that never continued. A subagent has its own system prompt, so its first call extends nothing and starts a new root, and a compacted context does the same.
Each path from a root to a leaf becomes one training sequence. Context tokens get a loss mask of 0 and sampled tokens get 1. A turn whose logprobs are missing or misaligned stays in the sequence as context and is never used as a target. These checks run as each call arrives rather than at export, because a misaligned turn is easiest to diagnose while the proxy still knows which turn it was.
Capturing your own harness
Nothing in the proxy is specific to Harbor, so any program that calls one of the four APIs can be
captured. Start the proxy, mint a session with POST /sessions, give your agent the proxy’s URL as
its base URL and the session id as its API key, and read the rollout back from
GET /sessions/{id}/rollout.
python -m openenv.core.harness.capture.server --llm-url http://127.0.0.1:8000 --model <model>
The proxy binds to localhost by default. On any other interface it needs an admin key, which guards the routes that mint and read sessions.
The proxy is a single process. In our runs its health check started to starve at around 200 concurrent sessions, and the process crashed at 320. A rollout that hits its model-call budget is also stopped by the proxy itself, so that final reply is never part of the training data.
Harbor
Harbor, from the team behind Terminal-Bench (Merrill et al., 2026), runs agents against containerised tasks and keeps the task, the harness and the sandbox independent of each other. Version 0.22.0 ships adapters for more than 40 agent harnesses and 26 execution backends, from local Docker to Daytona, Modal and E2B. About 80 task datasets already use its format, including Terminal-Bench and SWE-bench (Jimenez et al., 2024). Any harness can attempt any task on any backend, so one task set can be trained against many harnesses, and a run is one command that names the dataset, the agent and the backend. Adithya wrote about why it is the right abstraction for coding RL in RL Coding Environments 101: Why Harbor Exists.
A Harbor task is a directory with an instruction, an environment and a test script. The ones below come from SmolDataEnvs, the task suite used for every run in this article. Its questions come from Kaggle notebooks in the jupyter-agent dataset, each with an automatic grader. Pick a task and a file to read it as the agent and the verifier see it.
The task names no harness, sandbox backend or trainer, because those are chosen when it runs. The Harbor visualiser opens the whole training set the same way, all 5,000 tasks.
Note: Harbor calls its sandbox backends environments. In this article, sandbox means the box the agent runs in, and environment means the OpenEnv server.
Serving Harbor through OpenEnv
Harbor already runs an agent against a task and returns a score. A trainer also needs the tokens the policy generated and their probabilities. Harbor’s RL documentation lists two ways to get them: intercept the tokens at the inference server, or have the harness return them with its result. The second needs a harness that puts its tokens in the result metadata, and the docs say that support is still being added to Terminus 2, Harbor’s own harness. The integration takes the first, through the capture proxy, so it works with any harness that calls a model API.
The trainer talks to the environment over HTTP instead of running rollouts in its own process. An
earlier in-process version let an exception from inside someone else’s agent reach the training loop,
where one crashed rank left the others waiting at the NCCL barrier indefinitely. Now nothing raises
across that boundary. A failed rollout comes back with ok=False and reward=None, and the trainer
handles a value instead of catching an exception.
The four commands
| Command | What it does |
|---|---|
openenv harbor info | Reports what this machine can run: whether the engine returns token ids, which sandbox backends have working credentials, which datasets resolve and how many tasks each holds, and which harnesses are validated. Starts nothing. |
openenv harbor rollout | Runs rollouts without an environment server. If a rollout works here and fails under serve, the fault is in the serving layer. |
openenv harbor serve | The environment server: a Task API, a run_rollout tool over MCP, a web UI and the capture proxy. |
openenv harbor push | Deploys the same server to a Hugging Face Space, with configuration as Space variables and credentials as secrets. |
openenv harbor info --llm-url $LLM --dataset org/train,org/eval
openenv harbor rollout --llm-url $LLM --dataset org/train --task-index 0 -n 5 --harness opencode --sandbox e2b
openenv harbor serve --dataset org/train,org/eval
openenv harbor push --llm-url $LLM --dataset org/train,org/eval --repo-id you/harbor-env
rollout requires --llm-url and has no default, because a stale endpoint produces rollouts that
look fine and carry no token ids. serve can start without one, because each rollout names its own
engine. That let one deployment serve both training, pointed at the trainer’s vLLM, and evaluation.
Here -n 5 (or --n-tasks 5) runs tasks 0 through 4 once each, starting at --task-index 0.
It does not sample five attempts at one task. The E2B example matches the reported experiments;
the companion tutorial uses Daytona. A rollout prints one line when it finishes:
[opencode / e2b] task 0: 0000_369_369503_qa_1 ...
ok reward=1.00 turns=9 roots=2 multi-turn tokens=1043 atif=match 48s
The next two parts cover how each harness is wired to the proxy and how rewards and captures are checked. The Harbor page of the OpenEnv docs has the full reference.
Wiring each harness, picking the reward, and checking the capture
Wiring each harness
Harbor already knows how to install and launch each harness. OpenEnv adds only the place each one reads its model URL and API key from, kept as one table entry per harness, so supporting a new harness means adding an entry rather than writing a new environment. Ten harnesses have passed the capture contract and the trace check described below, across 16,000 rollouts on the SmolDataEnvs test set.
Most harnesses read environment variables, and OpenEnv sets those for the agent’s run only. The
session id Codex receives as OPENAI_API_KEY therefore never reaches the verifier, which may need a
real key of its own. Claude Code and Gemini CLI also read the variables in the server process, so
they get them there too. OpenCode and Pi take a config file written into the sandbox, and Terminus 2
runs on the host and is handed the URL and key directly.
Picking the reward, and checking the capture
Harbor’s verifier can return several named scores, and GRPO (Shao et al., 2024) needs one number per rollout. The
integration uses the only score if there is one, or the one named reward, and otherwise asks the
caller to choose. It never combines scores itself, because whatever weighting it picked would become
the training objective. An earlier run had a +0.2 for submitting anything term. The policy learned
to submit immediately, and held-out accuracy fell from 0.740 to 0.178 while training reward still
looked healthy. When the verifier never runs, the reward stays None all the way to the trainer, so
a sandbox that died is not scored as a wrong answer.
The proxy’s own checks can only look at data it captured. As an independent check, each rollout is also compared with the trajectory file Harbor writes itself, in its ATIF format, call by call: the number of turns, the completion tokens per call, and which calls count as agent steps. Any mismatch fails the rollout.
ATIF completion_tokens : [37, 36, 104, 264, 255, 119, 32, 27] total 874
intercept turn_lengths : [37, 36, 104, 264, 255, 119, 32, 27] total 874
ATIF step-1 prompt_tokens 7990 == intercept prompt_len 7990
This caught a harness that sent an empty tools array, got a 400 from the inference server and was cut short, leaving a rollout graph that looked well formed. The proxy now drops empty tools arrays before forwarding.
Deploying it and training against it
On one machine, serve listens on two ports, one for trainers and the web UI and one for the capture
proxy. Sandboxes usually run somewhere else, so the proxy reaches them through a tunnel, set with
--expose and Gradio by default. When the server itself runs on a Hugging Face Space, it has only
one port, so the proxy is mounted at /capture. That Space has to be public, because the agent in
the sandbox cannot send the auth header a private Space requires. It stays safe because the proxy
only answers registered session ids, and minting one needs an admin key. The sandboxes themselves
stay on E2B or another Harbor backend, one per rollout.
Before training, raise the server’s concurrency limit, max_concurrent_envs. It caps how many
rollouts one server runs at once, and the default of four is fewer than the eight rollouts in each
of our GRPO groups. Past the limit, serve rejects new rollouts with CAPACITY_REACHED.
OpenEnv’s HarborSessionFactory produces sessions in the shape TRL’s HarnessRolloutWorker
expects when the agent owns the loop. The worker was merged in
TRL #6947 and lives under trl.experimental.
As of 1 October 2026, it is on main but not in the latest release, v1.14.1. Install TRL from main
until the next release:
python -m pip install 'trl @ git+https://github.com/huggingface/trl.git@main'
Each session requests one rollout, waits for it, and returns a typed TrainingTrace through
fetch_training_trace(), with token IDs, behavior logprobs and loss masks. TRL consumes that object
directly. Given a prepared task dataset and the model’s tokenizer, the worker wiring is:
from functools import partial
from harbor_env.harness import HarborSessionFactory
from trl.experimental.async_grpo.openenv_harness import HarnessRolloutWorker
def correctness_reward(outcome):
return outcome.env_reward # None keeps a failed verifier unscorable.
def make_rollout_worker(dataset, tokenizer, server_url, vllm_url, model):
factory = partial(
HarborSessionFactory,
server_url=server_url,
split="FineEnvs/SmolDataEnvs-harbor-train",
llm_url=vllm_url,
model=model,
harness="opencode",
sandbox="daytona",
reward_key="correctness,reward",
)
return HarnessRolloutWorker(
harness_session_factory=factory,
harness_adapter=None,
rollout_reward_fn=correctness_reward,
model_name=model,
dataset=dataset,
processing_class=tokenizer,
reward_funcs=[],
num_generations=8,
max_inflight_tasks=8,
vllm_server_url=vllm_url,
max_tokens=4096,
temperature=1.0,
top_p=1.0,
top_k=-1,
chat_template_kwargs={"enable_thinking": False},
)
Pass this worker as rollout_worker to AsyncGRPOTrainer. The worker calls the factory with
sampling= before generation, so the capture policy matches training. harness_adapter=None
lets the harness run its own loop. This minimal callback forwards correctness; the
training setup adds the tool-efficiency bonus. The
full TRL example
also shows task discovery, AsyncGRPOConfig, the trainer and checkpoint saving.
In our runs every rollout in a GRPO group used the same harness, so the advantage compared actions rather than harnesses. The sampling temperature sent with each rollout also matched the one the trainer used to recompute logprobs.
Try the SmolDataEnvs Harbor Space. Choose a task, connect a model and run a harness to inspect its rollout and traces. The training tutorial includes the same environment source and commands for HF Jobs or a local cluster. The link uses a fixed revision so the instructions remain available as the repository evolves.
Footnotes
- Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2024). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? The Twelfth International Conference on Learning Representations. https://huggingface.co/papers/2310.06770
- Kolavi, A. S. (2026). Harbor integration: serve Harbor task datasets through OpenEnv as trainable environments. https://github.com/huggingface/OpenEnv/pull/1036
- Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. Proceedings of the 29th Symposium on Operating Systems Principles. https://huggingface.co/papers/2309.06180
- Merrill, M. A., Shaw, A. G., Carlini, N., Li, B., & others. (2026). Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces. arXiv Preprint arXiv:2601.11868. https://huggingface.co/papers/2601.11868
- Meta PyTorch, & Hugging Face. (2025). OpenEnv: Agentic Execution Environments. https://github.com/huggingface/OpenEnv
- Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y. K., Wu, Y., & Guo, D. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv Preprint arXiv:2402.03300. https://huggingface.co/papers/2402.03300
- Xu, B., Zhang, H., Zhang, S., Han, S., Liu, M., Hu, J., Diao, S., Jin, Z., Zou, Y., Demoret, M., Kautz, J., & Dong, Y. (2026). Polar: Agentic RL on Any Harness at Scale. arXiv Preprint arXiv:2605.24220. https://huggingface.co/papers/2605.24220
Training small models across harnesses
We trained LFM2.5-2.6B with the setup from the previous chapter, once in OpenCode alone and once across four harnesses, and then compared both runs with supervised fine-tuning on a larger model’s rollouts. Our first runs used Qwen3.5-2B. They shaped the LFM setup, and they are in a collapsible section at the end.
Setup
- Tasks. 1,000 SmolDataEnvs training tasks, 400 medium and 600 hard, in one fixed order shared by both runs. The test set is 250 held-out tasks with no overlap in notebook, question, or instruction.
- Two runs. One trains in OpenCode only. The other is multi-harness (OpenCode, Claude Code, Codex, and Mini-SWE-Agent), where each GRPO (Shao et al., 2024) group of eight rollouts uses one of the four harnesses.
- Evaluation. Both runs are evaluated the same way, including under the three harnesses the OpenCode-only run never trained in. Every 100 steps, each of the 250 test tasks runs once under each of the four harnesses, 1,000 cells in all. A cell counts as solved when its first graded attempt is correct.
- Recipe. Async GRPO in TRL for 1,000 steps. Async means generation and training run apart. One H100 trains while a second serves the policy with vLLM (Kwon et al., 2023) and keeps producing rollouts, so training never waits for the slowest agent. The price is that rollouts come from weights usually two optimizer steps old, and never more than four.
- Hardware and time. Two H100s per run on the Hugging Face cluster, and one E2B sandbox per rollout with 1 CPU and 4 GB of memory. Training took about 32 hours for the OpenCode-only run and 46 hours for the multi-harness run, not counting evaluation, queueing and restarts. We did not track the full cost closely enough to quote it.
The companion tutorial has readable multi-harness and native OpenCode training scripts, plus HF Jobs and Slurm instructions. The reported LFM runs used Harbor for both policies and E2B sandboxes. The current tutorial uses Daytona and includes native OpenCode as a separate comparison, so it is not an exact replay of those historical runs.
Each cell runs once, so a point or two between neighbouring checkpoints is noise. Trust the trend across checkpoints over any single one.
Where the base model starts
The baseline evaluation already shows the problem from the earlier chapters. The same weights solve 62% of tasks under Mini-SWE-Agent, but only 33% under Claude Code. LFM2.5-2.6B was not trained in our four harnesses, as far as its release says: Liquid AI names Hermes Agent and OpenClaw as its training harnesses.
What we rewarded
We rewarded LFM for two things:
- Correctness: 1 for a correct answer and 0 for a wrong one.
- Tool efficiency: a small bonus, at most 0.1, for a correct answer that takes fewer calls. A wrong answer still scores 0, so the model can’t earn the bonus by giving up early.
The bonus comes from the Qwen runs, which were rewarded for correctness alone. That gives the model no reason to stop exploring once it has the answer, and in an early Qwen run the tool calls per rollout crept from 13 to 41.
The bonus is at most 0.1, but GRPO learns from differences inside a group, so even a small bonus counts. GRPO runs eight rollouts of the same task and trains on how each one compares with the group’s average. When all eight are correct, correctness alone makes them identical, and the group teaches nothing. The bonus is then the only thing that tells them apart, and it favours the shortest solutions.
Here are two real groups from the multi-harness run, scored both ways.
In 97 of 555 OpenCode-only groups (17.5%) and 155 of 686 multi-harness groups (22.6%), every scorable rollout was correct but tool counts differed. The efficiency bonus supplied their only reward contrast. These audited counts cover groups that contributed to the retained training updates; they exclude work discarded on resume. Tool counts come from the harness’s own trace. The reward audit found no formula mismatches and no bonus on wrong answers.
In both LFM runs, tool calls per rollout fell during training and under every harness, as the curves below show. There was no LFM run without the bonus, so we cannot say the bonus caused this on its own.
Training curves
The share of correct rollouts climbs over the first few hundred steps in both LFM runs. After that it rises and falls with the difficulty of each stretch of tasks, which is the same for both runs because they follow the same task order. Tool calls fall throughout. On the test set, the OpenCode-only model performs better than the multi-harness model under OpenCode (58% vs 50%), while the multi-harness model performs better under Claude Code (49% vs 42%) and Codex (54% vs 43%).
The tool-call curves look reversed between training and test. During training, the multi-harness run makes more calls per rollout, because its average includes Claude Code, Codex and Mini-SWE-Agent rollouts, and those harnesses make more calls than OpenCode even before any training. On the test set, both models run under all four harnesses, and there the multi-harness model makes fewer calls under every one of them. The OpenCode-only model learned to use fewer calls in OpenCode, and that carried over only in part, mainly to Codex.
Accuracy
Both runs climb in accuracy, by 10 points for OpenCode-only and 12 for multi-harness, most of it in the first 200 steps, and hold it to step 1,000. They end at 52.3% for OpenCode-only and 54.2% for multi-harness. Both gains are well outside the noise. The 1.9-point gap between the runs is within the noise, so overall accuracy does not separate them.
The per-harness split is the clearest result. The OpenCode-only model is best under OpenCode, the harness it trained in. The multi-harness model is better under Claude Code and Codex, and the two tie under Mini-SWE-Agent. Most of the OpenCode-only model’s gain is in OpenCode (34 → 58%), while the multi-harness model gained under all four.
Tool calls and tokens
Both models solve more tasks with fewer tool calls, and the multi-harness model cuts more, 31% fewer calls than the base model on the tasks both solved, against 11% for OpenCode-only. Fewer calls also means less history resent on every call, so input tokens fall with them. The biggest saving is under Codex, where the multi-harness model uses about half the calls. The clear exception is OpenCode-only under Claude Code, a harness it never trained in, where it used more calls and more tokens than the base model.
The multi-harness model’s Codex savings grow at almost every checkpoint, and by step 1,000 every harness is blue on all three measures. The OpenCode-only model saves steadily in OpenCode and more and more in Codex. Under Claude Code, it generated more tokens than the base model at nearly every checkpoint, up to 62% more at step 500.
How much each run saw
Both runs took 1,000 optimizer steps, but they did not see the same amount of data.
| OpenCode only | Multi-harness | |
|---|---|---|
| Distinct training tasks | 555 | 626 |
| Supervised tokens | 12.5M | 18.7M |
| Tokens processed in training | 162M | 451M |
Claude Code rewrites its own history as it goes, so each of its rollouts became about eight training rows, each carrying its own copy of the context, while the other harnesses averaged about one. With one seed per run and unequal exposure, the comparison between the two runs is observational.
SFT vs RL
RL is not the only way to learn from a harness. The capture that gives RL its exact tokens also records whole trajectories, so a larger model’s successful runs can be used for supervised fine-tuning (SFT). We tried this on LFM2.5-2.6B to see how far imitation gets next to the RL runs above.
The experiment
We ran Qwen3.8-27B as a teacher on SmolDataEnvs training tasks in all four harnesses, through the same OpenEnv and Harbor setup, with sandboxes on Daytona. Each task got up to three attempts per harness, and we kept the first one the verifier graded correct. That gave 3,189 successful rollouts from 888 tasks, 377 medium and 511 hard.
Every assistant turn in a rollout becomes one training example. The conversation up to that turn, tool outputs included, is the context, and the loss is on the next assistant response only. The examples are tokenized with LFM’s own chat template, not the teacher’s.
We trained two SFT models from the same LFM2.5-2.6B base as the RL runs:
- OpenCode SFT on the 801 OpenCode rollouts, 4,825 examples.
- Multi-harness SFT on all 3,189 rollouts, 17,929 examples, shuffled together with no balancing between harnesses.
Both are full fine-tunes with TRL’s SFT trainer for two epochs, at a learning rate of 3e-6 and an effective batch size of 8. Each epoch is evaluated on the same 1,000 test cells as the RL runs.
The data is public as SmolDataEnvs-multiharness-sft.
It has one config per harness, an all config with the four together, and configs already tokenized
for LFM2.5-2.6B and Qwen3.5-2B, next to a training script.
The raw teacher rollouts are in qwen38-27b-harbor-rollouts.
The two methods learn from different things:
| SFT | RL | |
|---|---|---|
| Learns from | the teacher’s successful rollouts | its own rollouts |
| Training signal | the teacher’s next response | correctness, plus the tool-call bonus |
| Rewards fewer tool calls | no | yes |
| Training tasks (OpenCode / multi-harness) | 801 / 888 | 555 / 626 |
| Supervised tokens, whole run | 1.2M / 8.2M | 12.5M / 18.7M |
| Training time | 1.5 h / 6.5 h | 32 h / 46 h |
The SFT times are the trainer’s own loop and leave out collecting the teacher’s rollouts, which took a 27B model through all four harnesses.
Results
For RL we use each run’s best checkpoint on the test set: step 900 for OpenCode only and step 700 for multi-harness. Choosing on the test set flatters RL slightly, but both are within half a point of their step-1,000 scores.
RL ends ahead of both SFT runs. Multi-harness RL reaches 54.6%. The better SFT model, OpenCode SFT, reaches 47.5%, and multi-harness SFT 43.1%, 11.5 points below the RL model trained in the same four harnesses.
OpenCode SFT looks like a smaller OpenCode RL. Almost all of its gain is in OpenCode, from 33.6% to 51.6%, close to OpenCode RL’s 56.0%. Under Claude Code it stays where the base model was, 33.2%, and under Mini-SWE-Agent it drops a little, from 62.1% to 58.8%. Training in one harness mostly helped that harness, by imitation or by RL.
Multi-harness SFT spreads the gain, then loses it in Mini-SWE-Agent. It improves under Claude Code (33.2 → 42.4%), Codex (40.0 → 46.8%) and OpenCode (33.6 → 38.0%). Under Mini-SWE-Agent it falls from 62.1% to 45.2%, which cancels the rest, so overall it ends at 43.1%, within the noise of the base model’s 42.2%. After the first epoch it was below the base model, at 38.3%, with OpenCode down to 22.4%. Multi-harness RL improved under all four, Mini-SWE-Agent included. We have not yet worked out why SFT hurt Mini-SWE-Agent.
SFT cuts tool calls without being rewarded for it. SFT has no efficiency term, yet on the tasks both it and the base model solved, multi-harness SFT makes 24.2% fewer tool calls, with fewer calls under every harness. OpenCode SFT makes 8.5% fewer overall, but more than the base model under Claude Code (10.5% more) and Mini-SWE-Agent (4.8% more). RL shows the same pattern: training across the four harnesses cut calls everywhere, 26.7% for multi-harness RL at step 700 and 31.1% by step 1,000, and training in OpenCode alone cut them unevenly.
What this comparison can’t tell us. The two SFT runs differ in more than the harness mix. The multi-harness run has about seven times the supervised tokens and nearly four times the steps. SFT and RL also differ in objective, data source, task pool and compute, and each setup ran once. These are observations, not a ranking of methods. The runs we want next are SFT followed by RL against RL from the base model, harness mixes compared on the 692 tasks the teacher solved in all four harnesses, and SFT mixes with equal supervised tokens per harness.
The Qwen3.5-2B runs, and what they taught us
The runs
Qwen came first, trained with the correctness-only reward on an easier pool of 1,000 tasks (150 easy, 600 medium, 250 hard). It starts almost three times lower than LFM on the same test set, and the Qwen3.5 model card names no training harness.
There were three Qwen runs, all evaluated on the same 1,000 test cells:
- Harbor multi-harness, rotating the four harnesses across GRPO groups.
- Harbor OpenCode-only, the same pipeline with one harness.
- Standalone OpenCode, OpenCode in its own environment outside Harbor, running in Daytona.
Up, then down
All three improved at first, which showed that an unmodified harness can be trained through once you have the exact tokens. Multi-harness more than doubled, from 14.6% to 37.0% at step 500. OpenCode-only peaked at 39.5% at step 700. Then both fell back to about 26% by step 1,000. The standalone run rose more slowly and ended at its best, 29.8%.
The standalone run trained only in OpenCode, yet it gained about twice as much under Claude Code, Codex and Mini-SWE-Agent as under OpenCode.
Why the two Harbor runs declined
The two declines look different in the traces.
- Multi-harness ran out of output budget. Late in training, answers grew until they hit the 4,096-token output limit used in evaluation. Cells cut off at that limit rose from 9 to 556 of 1,000, often a long explanation where the calculation should have been. Training allowed responses up to 16,384 tokens, so the model learned a habit that evaluation cut short. Under OpenCode, calls per task collapsed from 16 to 3.
- OpenCode-only kept working without finishing. Calls per task rose from 17 to 21, and in the cells where submission was tracked, the share that submitted an answer fell from 69% to 41%.
- A resume bug replayed old tasks. After the multi-harness run restarted at step 684, almost every new rollout came from a task it had already seen. That changed its late training, but the decline started at step 500, before the restart, so the bug does not explain its start.
Tool calls without a bonus
With correctness as the only reward, tool calls per rollout in the Qwen OpenCode-only run doubled from about 10 to 20 during training, and the Qwen multi-harness run stayed between 11 and 16. In both LFM runs, which had the bonus, they fell.
Other things the Qwen runs showed
- Many steps taught nothing. Between 35% and 58% of optimizer steps had no reward contrast in any group, because all eight rollouts were right or all were wrong.
- Claude Code dominated the training data. Each harness got about a quarter of the rollouts, but Claude Code produced 77% of the training rows and 35% of the supervised tokens.
- SETA, a synchronous run in a plain bash environment with its own evaluator, went from 18.8% to 38.0% in 150 steps before we stopped it on purpose.
- Harder data did not rescue the decline. Continuing the multi-harness checkpoint from step 500 on 500 new hard tasks, on H200s through Hugging Face Jobs, did not recover the peak. Nearly two thirds of its steps had no reward contrast, and even counting every ungraded cell as correct, its final score could not reach its starting 37.0%.
What changed for LFM
The LFM setup changed three things. The reward gained the tool-call bonus, so all-correct groups still carry a signal and long, unfinished loops cost something. The training pool moved to medium and hard tasks. And the output cap was 4,096 tokens in both training and evaluation, so the model could not learn answers that evaluation would cut off.
Reproducing it
- SmolDataEnvs Multi-harness RL collection: everything below in one place, plus the base models.
- Trained models: LFM multi-harness RL and OpenCode-only RL at step 1,000, with the best checkpoints (steps 700 and 900) on
step-700andstep-900branches; LFM multi-harness SFT and OpenCode SFT after epoch 2; and the best Qwen checkpoints, multi-harness (step 500), OpenCode-only (step 700) and standalone OpenCode (step 1,000). - Training comparison dashboard: every training curve and checkpoint evaluation.
- Harbor environment and standalone OpenCode environment: the current companion environment servers. The standalone
opencode_envis deprecated since OpenEnv 0.7.0 and is retained for the historical comparison; new runs should use Harbor withharness="opencode". - SmolDataEnvs-multiharness-sft: the SFT data, per harness and pre-tokenized for LFM2.5-2.6B and Qwen3.5-2B, with a training script.
- Merged integrations: OpenEnv #1036 for the capture proxy and Harbor, OpenEnv #1280 for the typed
TrainingTraceAPI, and TRL #6947 for the worker. TRL’s Harbor example shows the full trainer wiring; install TRL from main until that worker is released.
Every run in this article uses a small model, 2 to 2.6 billion parameters, trained for 1,000 steps on a single seed. We are now setting up larger runs with bigger models on the same stack, and we will add the results here when they are in. Stay tuned.
- Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. Proceedings of the 29th Symposium on Operating Systems Principles. https://huggingface.co/papers/2309.06180
- Shao, Z., Wang, P., Zhu, Q., Xu, R., Song, J., Bi, X., Zhang, H., Zhang, M., Li, Y. K., Wu, Y., & Guo, D. (2024). DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models. arXiv Preprint arXiv:2402.03300. https://huggingface.co/papers/2402.03300
Conclusions
Results compiled through 29 September 2026. Article updated 1 October 2026.
We set out to answer one question: can a small open model be trained inside the harnesses people actually use, such as Claude Code, Codex and OpenCode, and improve across all of them? It can. The harnesses run as they ship, with no changes to their code. The capture proxy sits between the harness and the model and records the exact tokens the model produced. Harbor supplies the tasks and the sandboxes, and TRL trains on the result.
What we found
The harness changes the score before any training. With the same weights, LFM2.5-2.6B solved 62% of our test tasks under Mini-SWE-Agent and 33% under Claude Code. A score measured in one harness describes the model in that harness, not the model in general.
A model gets better where it trains. This held for both methods we tried. OpenCode-only RL raised LFM’s OpenCode score from 34% to 58%, and OpenCode SFT raised it to 52%. Both did much less in the other three harnesses. Training across all four spread the gains: multi-harness RL improved under every harness and ended at 54% overall, against 52% for OpenCode-only RL. That overall gap is within the noise. The difference is where the gains landed: higher under Claude Code and Codex, lower in OpenCode.
The reward decides how long the agent keeps working. In our earlier Qwen runs, with correctness as the only reward, nothing told the model to stop, and between 35% and 58% of optimizer steps taught nothing because every rollout in each group scored the same. A bonus of at most 0.1 for fewer tool calls gave LFM a signal even when a whole group was correct. The multi-harness LFM model made 31% fewer calls than the base model on the tasks both solved. There was no LFM run without the bonus, so we can’t separate its effect from the other changes between the Qwen and LFM setups.
Imitation helped less than RL. Fine-tuning LFM on a 27B teacher’s successful rollouts reached 47.5% with OpenCode data and 43.1% with all four harnesses, both below the two RL runs. Multi-harness SFT still made 24% fewer tool calls, and fewer in every harness, with nothing in its objective asking for it. Its overall score, within the noise of the base model, hid gains of 4 to 9 points in three harnesses and a 17-point drop under Mini-SWE-Agent.
Exact tokens are only the starting point. They make the update match what the model generated. The earlier Qwen runs had exact tokens and still peaked and fell back. One run’s answers grew past the output limit that evaluation allowed, and the other kept calling tools without submitting. The LFM runs used the same output limit in training and evaluation and a reward that values finishing, and they did not decline, though they also changed the model and the task pool.
Steps are a poor unit of comparison. The two LFM RL runs took the same 1,000 steps, but the multi-harness run saw 626 distinct tasks against 555 and processed 451M training tokens against 162M. Much of the difference comes from Claude Code, which rewrites its history as it goes, so each of its rollouts became about eight training rows.