Microsoft Open-Sources Lightweight Agentic RL Framework
Microsoft Research Asia open-sourced Agent Lightning v1.0 on October 7, a lightweight framework of approximately 3,500 lines of code designed to allow the same agent harness used in deployment to participate directly in reinforcement learning.
Traditional agentic RL often requires reimplementing agent logic within the training framework, leading to discrepancies between training and production environments and high costs. Agent Lightning inserts an LLM proxy between the agent and the model, allowing existing harness code to remain unchanged while connecting to RL training.
The framework natively supports running agents as Kubernetes jobs, eliminating dependency on paid commercial sandbox services. In coding agent experiments, using about 6,000 training samples, it improved Qwen3.5-9B's Pass@1 on SWE-bench Verified from 41.8% to 56.4%, an absolute gain of 14.6 percentage points. These results are self-reported by Microsoft and await independent reproduction.
ソース:microsoft.com