01 What happened
Microsoft plans to have GitHub Copilot decide automatically whether a task runs on a local model or on cloud-scale models. The company says the change will arrive by the end of October 2026. Microsoft AI also describes a local, quantized build of MAI Code 1.1 Flash, a coding model it puts at 53GB, 80% smaller than the Bfloat16 cloud variant.
02 Key details
- Microsoft says automatic routing between on-device and cloud models is scheduled to arrive by the end of October 2026.
- Microsoft puts the quantized local MAI Code 1.1 Flash at 53GB, with peak memory of 75.5GB at 256k context on Surface Laptop Ultra.
- The SWE-Bench Verified scores (70.80% on device, 72.6% in the cloud) are Microsoft's own and not independently verified.
- Local inference does not make a Copilot session offline. Built-in file tools are checked against policy by the agent harness, not isolated at the OS level.
03 Why it matters
For developers on supported Windows PCs, some tasks may run locally within the memory budget and others in the cloud, without manual infrastructure choices. The size and performance figures, however, remain the vendor's own claims.
04 Who it matters to
Software developers and teams evaluating local AI programming tools.