view article Article Decoding LLM Alignment: A Concise Guide for GRPO and its variants Sneha7 • 4 days ago • 5
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Paper • 2609.13141 • Published 28 days ago • 69
view article Article LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation LiquidAI • Aug 19 • 37
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • Sep 3 • 147
Agora: Git as Shared Memory for Collective AutoResearch Paper • 2609.18094 • Published 23 days ago • 58
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 22 days ago • 57
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 24 days ago • 78
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published about 1 month ago • 82
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 23 days ago • 84
AgentGrad: Intervention-guided Prompt Optimization for Multi Agent Systems Paper • 2609.08572 • Published about 1 month ago • 108
RSIAgent: Autonomous Exploration for Recursive Self-improvement in New Environments Paper • 2609.15364 • Published 25 days ago • 84
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 22 days ago • 93
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published Sep 4 • 119
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104