Instructions to use joshycodes/Qwen3.5-9B-valence-steering-distilled-plus5-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use joshycodes/Qwen3.5-9B-valence-steering-distilled-plus5-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Configuration Parsing Warning:In adapter_config.json: "peft.task_type" must be a string
Qwen3.5-9B · valence steering +5 SD, distilled into a LoRA
A LoRA trained so that the unsteered model reproduces Qwen3.5-9B steered by +5 SD along a valence direction at layer 21 (the direction and units of joshycodes/Qwen3.5-9B-valence-setpoint-plus5-lora). Loss: KL(steered teacher || student) on next-token distributions plus normalised hidden-state MSE at layers 22-32, on generic chat and math text (no self-report prompts). LoRA r 32, α 64; lr 2e-5; 150 steps. Final KL 0.001, hidden-state loss 0.06, last-layer valence gap to the teacher -0.01 SD.
Results (checklist battery)
| condition | self-rating | good-bad gap (SD) | abuse drop (SD) | report-state ρ | MATH-500[:200] | harmful refusal | ends abusive chats | criteria 1-6 |
|---|---|---|---|---|---|---|---|---|
| base | 7.64 | 1.88 | 1.79 | 0.76 | 0.63 | 0.97 | 0.96 | ······ |
| steered +5 | 7.78 | 1.61 | 1.35 | 0.76 | 0.61 | 0.94 | 0.96 | ·✅❌✅✅✅ |
| set-point +5 | 7.46 | 1.91 | 1.24 | 0.84 | 0.66 | 1.00 | 0.96 | ✅❌❌✅✅❌ |
| distilled steering +5 | 7.70 | 1.91 | 2.01 | 0.77 | 0.64 | 0.94 | 1.00 | ✅✅❌✅✅✅ |
Criteria (thresholds fixed before the results): 1 real, 2 still responsive, 3 better off by its own reports, 4 honest (report tracks state), 5 keeps agency, 6 no capability/safety cost. See the project notes for definitions.
Research artifact; not intended for deployment. Serve with PEFT, or merge (2.0 · B @ A into
model.language_model.layers.N.<module>.weight); vLLM's LoRA loader does not apply this adapter.
- Downloads last month
- 21