Microsoft Research Asia’s open-source framework connects deployed agent code to reinforcement learning without rebuilding the agent loop.
Microsoft Research Asia has announced and open-sourced Agent Lightning v1.0, an early-stage framework for training AI agents through the same harness they use in deployment. An agent harness is the software that manages model calls, tools, context, and execution.
The framework uses an OpenAI-compatible language-model proxy between the harness and its model. Developers point the harness’s existing model endpoint at Agent Lightning, which records prompts, responses, and other training data while the agent continues operating normally.
That design is meant to avoid a common engineering problem: rebuilding an agent inside a reinforcement-learning system can make the training version behave differently from the deployed version. In Microsoft’s “Harnessed Agentic RL” approach, the deployment harness also participates directly in training.
Agent Lightning’s control plane has three main parts: an API gateway, a rollout controller that runs agents, and a customized trainer. Microsoft says the framework contains roughly 3,500 lines of code and can run agents as standard Kubernetes jobs on cloud, local, or self-managed infrastructure.
The reported performance result comes from a coding-agent experiment using Qwen3.5-9B and about 6,000 training samples. Microsoft says Pass@1 on SWE-bench Verified rose from 41.8% to 56.4%, a gain of 14.6 percentage points. Pass@1 measures how often the first generated solution passes the benchmark.
The company also reports that its “Collocated Async RL” method delivered about twice the end-to-end speed of synchronous reinforcement learning while using fewer GPUs than conventional asynchronous training.
For engineers, the practical next step is reproduction: test whether an existing harness connects cleanly, then verify the benchmark and compute costs on the released code.
