-
RL + Transformer = A General-Purpose Problem Solver
Paper • 2501.14176 • Published • 28 -
Towards General-Purpose Model-Free Reinforcement Learning
Paper • 2501.16142 • Published • 31 -
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Paper • 2501.17161 • Published • 127 -
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
Paper • 2412.12098 • Published • 4
Collections
Discover the best community collections!
Collections including paper arxiv:2508.20722
-
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Paper • 2406.04151 • Published • 23 -
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Paper • 2510.16872 • Published • 113 -
Scaling Generalist Data-Analytic Agents
Paper • 2509.25084 • Published • 22 -
Scaling Agents via Continual Pre-training
Paper • 2509.13310 • Published • 117
-
SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
Paper • 2504.08600 • Published • 33 -
Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval
Paper • 2509.21710 • Published • 19 -
TTRL: Test-Time Reinforcement Learning
Paper • 2504.16084 • Published • 123 -
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Paper • 2508.03680 • Published • 142
-
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
Paper • 2506.21506 • Published • 52 -
Distilling LLM Agent into Small Models with Retrieval and Code Tools
Paper • 2505.17612 • Published • 82 -
Efficient Agent Training for Computer Use
Paper • 2505.13909 • Published • 44 -
Scaling Agents via Continual Pre-training
Paper • 2509.13310 • Published • 117
-
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Paper • 2509.02547 • Published • 237 -
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
Paper • 2508.16279 • Published • 70 -
rStar2-Agent: Agentic Reasoning Technical Report
Paper • 2508.20722 • Published • 121
-
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
Paper • 2404.10155 • Published -
rStar2-Agent: Agentic Reasoning Technical Report
Paper • 2508.20722 • Published • 121 -
Testing LLMs on Code Generation with Varying Levels of Prompt Specificity
Paper • 2311.07599 • Published -
Nerdsking/nerdsking-python-coder-3B-i
Text Generation • 3B • Updated • 398 • 15
-
ExGRPO: Learning to Reason from Experience
Paper • 2510.02245 • Published • 83 -
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
Paper • 2508.07407 • Published • 100 -
rStar2-Agent: Agentic Reasoning Technical Report
Paper • 2508.20722 • Published • 121 -
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
Paper • 2508.19828 • Published • 8
-
FastVLM: Efficient Vision Encoding for Vision Language Models
Paper • 2412.13303 • Published • 77 -
rStar2-Agent: Agentic Reasoning Technical Report
Paper • 2508.20722 • Published • 121 -
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
Paper • 2508.16279 • Published • 70 -
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
Paper • 2509.12201 • Published • 107
-
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Paper • 2402.17764 • Published • 631 -
MiniMax-01: Scaling Foundation Models with Lightning Attention
Paper • 2501.08313 • Published • 304 -
Group Sequence Policy Optimization
Paper • 2507.18071 • Published • 323 -
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
Paper • 2509.03867 • Published • 213
-
RL + Transformer = A General-Purpose Problem Solver
Paper • 2501.14176 • Published • 28 -
Towards General-Purpose Model-Free Reinforcement Learning
Paper • 2501.16142 • Published • 31 -
SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Paper • 2501.17161 • Published • 127 -
MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximization
Paper • 2412.12098 • Published • 4
-
The Fault in our Stars: Quality Assessment of Code Generation Benchmarks
Paper • 2404.10155 • Published -
rStar2-Agent: Agentic Reasoning Technical Report
Paper • 2508.20722 • Published • 121 -
Testing LLMs on Code Generation with Varying Levels of Prompt Specificity
Paper • 2311.07599 • Published -
Nerdsking/nerdsking-python-coder-3B-i
Text Generation • 3B • Updated • 398 • 15
-
AgentGym: Evolving Large Language Model-based Agents across Diverse Environments
Paper • 2406.04151 • Published • 23 -
DeepAnalyze: Agentic Large Language Models for Autonomous Data Science
Paper • 2510.16872 • Published • 113 -
Scaling Generalist Data-Analytic Agents
Paper • 2509.25084 • Published • 22 -
Scaling Agents via Continual Pre-training
Paper • 2509.13310 • Published • 117
-
ExGRPO: Learning to Reason from Experience
Paper • 2510.02245 • Published • 83 -
A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems
Paper • 2508.07407 • Published • 100 -
rStar2-Agent: Agentic Reasoning Technical Report
Paper • 2508.20722 • Published • 121 -
Memory-R1: Enhancing Large Language Model Agents to Manage and Utilize Memories via Reinforcement Learning
Paper • 2508.19828 • Published • 8
-
SQL-R1: Training Natural Language to SQL Reasoning Model By Reinforcement Learning
Paper • 2504.08600 • Published • 33 -
Think-on-Graph 3.0: Efficient and Adaptive LLM Reasoning on Heterogeneous Graphs via Multi-Agent Dual-Evolving Context Retrieval
Paper • 2509.21710 • Published • 19 -
TTRL: Test-Time Reinforcement Learning
Paper • 2504.16084 • Published • 123 -
Agent Lightning: Train ANY AI Agents with Reinforcement Learning
Paper • 2508.03680 • Published • 142
-
FastVLM: Efficient Vision Encoding for Vision Language Models
Paper • 2412.13303 • Published • 77 -
rStar2-Agent: Agentic Reasoning Technical Report
Paper • 2508.20722 • Published • 121 -
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
Paper • 2508.16279 • Published • 70 -
OmniWorld: A Multi-Domain and Multi-Modal Dataset for 4D World Modeling
Paper • 2509.12201 • Published • 107
-
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
Paper • 2506.21506 • Published • 52 -
Distilling LLM Agent into Small Models with Retrieval and Code Tools
Paper • 2505.17612 • Published • 82 -
Efficient Agent Training for Computer Use
Paper • 2505.13909 • Published • 44 -
Scaling Agents via Continual Pre-training
Paper • 2509.13310 • Published • 117
-
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey
Paper • 2509.02547 • Published • 237 -
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications
Paper • 2508.16279 • Published • 70 -
rStar2-Agent: Agentic Reasoning Technical Report
Paper • 2508.20722 • Published • 121
-
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Paper • 2402.17764 • Published • 631 -
MiniMax-01: Scaling Foundation Models with Lightning Attention
Paper • 2501.08313 • Published • 304 -
Group Sequence Policy Optimization
Paper • 2507.18071 • Published • 323 -
Drivel-ology: Challenging LLMs with Interpreting Nonsense with Depth
Paper • 2509.03867 • Published • 213