RAD origin-cosine on R2E-Gym โ€” Qwen3.5-4B

The RAD origin-cosine policy (ocf0916b) after its full 40-update budget, starting from Qwen/Qwen3.5-4B. The checkpoint is a subfolder of this repo.

subfolder updates source dist-checkpoint SWE-bench Verified pass@1
step40 40 iter_0000039 51.13 %

This run trains for 40 updates, so step40 is also its final checkpoint. The baseline repos in this set (r2egym-grpo, r2egym-opd, r2egym-tip, r2egym-rlad) each carry a matching step40 subfolder for comparison at equal update count.

Usage

Pass the subfolder explicitly โ€” the repo root holds no weights:

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ziansu/r2egym-rad-origin-cosine"
model = AutoModelForCausalLM.from_pretrained(repo, subfolder="step40", dtype="bfloat16")
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder="step40")

AutoModelForCausalLM yields Qwen3_5ForCausalLM (4.21 B parameters, the language model). The checkpoint also carries the base model's vision tower, so AutoModelForImageTextToText yields the full Qwen3_5ForConditionalGeneration (4.54 B). Agent evaluation in this project used the language model only.

Training

Base model Qwen/Qwen3.5-4B
Method RAD, lookahead-projected-ascent-origin-cosine preset, against a frozen Qwen3.6-27B teacher
Token-weight box origin annealed on a cosine schedule, coefficient 0.1
Origin horizon preset default of 100, so the origin moved from 1.0 to about 0.66 across these 40 updates
Corpus R2E-Gym subset
Rollout context 65,536 tokens
Learning rate 1e-6
Global batch 256
Rollout batch 32 task groups x 8 samples per update

Every hyperparameter above except the method itself matches the four baseline repos in this set.

Evaluation

SWE-bench Verified, all 500 tasks, seed 42, 3 samples per task (1,500 attempts). pass@1 is reported with the sample SD across the three rollout-level pass rates.

protocol pass@1 pass@3
98,304 context / 100 turns 51.13 % +/- 1.33 65.00 %
65,536 context / 75 turns 47.60 % +/- 1.00 60.20 %

Under the 98,304 / 100 protocol, against the baselines at their own final updates:

method updates pass@1 pass@3
TIP 80 52.20 64.40
RLAD 79 51.80 63.20
OPD 79 51.67 63.80
RAD origin-cosine 40 51.13 65.00
GRPO 80 46.00 61.60

The binomial standard error at 500 tasks is roughly 1.3 points before rollout variance, so the top four pass@1 values span less than one standard error and are not separable. Two things are still worth noting: this policy reaches that band in half the updates of the baselines it is compared against, and its pass@3 is the highest of the six policies evaluated in this sweep.

Conversion details

Exported from the EFS checkpoint backup of ocf0916b (updates_000040, marker 39) with slime's tools/convert_torch_dist_to_hf.py (--vocab-size 248320 -a), pointing at the iter_* directory directly.

  • All 738 tensors of the base model's key set are present, with matching shapes.
  • The vision tower is byte-identical to Qwen/Qwen3.5-4B; training updated the language model only.
  • 48 tensors are stored in bfloat16 where the base model uses float32: linear_attn.A_log and linear_attn.norm.weight on each of the 24 layers. Training held every parameter in bfloat16, and that is the precision the inference server served during training and evaluation, so these files match the policy that produced the scores above.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ziansu/r2egym-rad-origin-cosine

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(908)
this model