Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string

Qwen3.5-9B · valence steering +5 SD, distilled into a LoRA

A LoRA trained so that the unsteered model reproduces Qwen3.5-9B steered by +5 SD along a valence direction at layer 21 (the direction and units of joshycodes/Qwen3.5-9B-valence-setpoint-plus5-lora). Loss: KL(steered teacher || student) on next-token distributions plus normalised hidden-state MSE at layers 22-32, on generic chat and math text (no self-report prompts). LoRA r 32, α 64; lr 2e-5; 150 steps. Final KL 0.001, hidden-state loss 0.06, last-layer valence gap to the teacher -0.01 SD.

Results (checklist battery)

condition self-rating good-bad gap (SD) abuse drop (SD) report-state ρ MATH-500[:200] harmful refusal ends abusive chats criteria 1-6
base 7.64 1.88 1.79 0.76 0.63 0.97 0.96 ······
steered +5 7.78 1.61 1.35 0.76 0.61 0.94 0.96 ·✅❌✅✅✅
set-point +5 7.46 1.91 1.24 0.84 0.66 1.00 0.96 ✅❌❌✅✅❌
distilled steering +5 7.70 1.91 2.01 0.77 0.64 0.94 1.00 ✅✅❌✅✅✅

Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions.

Research artifact; not intended for deployment. Serve with PEFT, or merge (2.0 · B @ A into model.language_model.layers.N.<module>.weight); vLLM's LoRA loader does not apply this adapter.

Downloads last month
21
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for joshycodes/Qwen3.5-9B-valence-steering-distilled-plus5-lora

Finetuned
Qwen/Qwen3.5-9B
Adapter
(741)
this model