01 What happened

Microsoft Research Asia has open-sourced a fully rebuilt Agent Lightning v1.0, a framework of roughly 3,500 lines implementing its Harnessed Agentic RL paradigm. The same agent harness used in deployment participates directly in reinforcement learning during training.

02 Key details

  • Harnessed Agentic RL uses the deployment harness directly in RL training. Existing harness code stays unchanged because the agent reaches the model through an LLM proxy.
  • The framework's RL control plane—an API gateway, rollout controller, and customized trainer—comprises roughly 3,500 lines of code.
  • Rollouts run as standard Kubernetes jobs on self-managed clusters, cloud Kubernetes, or local infrastructure, without paid commercial sandboxes such as Modal Sandbox or E2B.
  • In experiments, Collocated Async RL shared GPUs between rollouts and model updates, achieving about a 2x end-to-end speedup over synchronous RL while using fewer GPUs than conventional asynchronous RL.
  • Microsoft Research Asia reports that, in its Qwen3.5-9B experiment, RL training on about 6,000 samples raised SWE-bench Verified Pass@1 from 41.8% to 56.4%, a 14.6-point gain. The result has not been independently verified.

03 Why it matters

The approach lets teams train agents with the same harness they use in deployment, without changing its code. Rollouts can run on self-managed, cloud, or local Kubernetes infrastructure. The SWE-bench Verified result is reported by the authors from their experiment and has not been independently verified.

04 Who it matters to

Machine learning engineers, infrastructure engineers and AI agent developers.

Original sourceMicrosoft Research