DistilQwen
Proof-weighted distillation, Qwen3-30B to 1.7B/0.6B. Three teachers: Instruct, Thinking, Coder. The core method series. DOI 10.57967/hf/8165
Text Generation • 2B • Updated • 3.56k • 2Note Base distillation. Instruct teacher → 1.7B student. BF16 H100.
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B-SFT-GGUF
Text Generation • 2B • Updated • 2.07kNote Instruct teacher + SFT quantized. F16/Q4/Q5/Q8 available.
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B-SFT
2B • Updated • 125Note Source model for the SFT-GGUF quantizations.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B
Text Generation • 0.8B • Updated • 3.42kNote 0.6B student. Proves the methodology works at extreme scales.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT
Text Generation • 0.8B • Updated • 3.44k • 2Note Higher-entropy teacher distributions → richer student representations.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT-GGUF
Text Generation • 0.8B • Updated • 1.97kNote mradermacher also auto-quantized this one — 420+ shadow downloads.
reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT
Text Generation • 2B • Updated • 3.53k • 1Note Coder teacher. Structured decomposition → STEM derivation.
reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT-GGUF
Text Generation • 2B • Updated • 2.51k • 2Note Coder pipeline quantized. F16/Q4/Q5/Q8.
reaperdoesntknow/DistilQwen3-1.7B-uncensored
Text Generation • 2B • Updated • 2.7kNote Starting point for custom SFT pipelines.
reaperdoesntknow/TopologicalQwen
Text Generation • 2B • Updated • 4.17k • 1Note The model that proved ghost imprinting — literary from physics data.
reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored
2B • Updated • 120 • 1Note DISC-informed distillation. Uncensored. Research-focused.
reaperdoesntknow/Disctil-Qwen3-1.7B
Text Generation • 2B • Updated • 2.5kNote DISC-refined. Discrepancy-aware training produces cleaner signal.
reaperdoesntknow/DistilQwen3-1.7B-uncensored-GGUF
2B • Updated • 2.1k • 3Note Uncensored base quantized. mradermacher also quantized — 411 downloads.
reaperdoesntknow/Qwen3-1.7B-Thinking-Distil
Text Generation • 2B • Updated • 3.64k • 2Note The most popular model. Thinking teacher = richest signal.
reaperdoesntknow/LFM2.5-1.2B-Distilled-SFT
Text Generation • 1B • Updated • 2.81kNote Proves TKD works across architecture families, not just within Qwen.
reaperdoesntknow/Discrepancy_Calculus
UpdatedNote Continuous Thought Dynamics — mathematical backbone of DualMind.