At its Oct 7 Windows event Microsoft shipped an on-device MAI-Code-1.1-Flash (137B total / 6.8B active MoE, 256K context, ~3.3 bits per weight) on RTX Spark PCs, scoring 70.80% SWE-Bench Verified and 66.29% Terminal-Bench 2.1 locally. GitHub HydraFusion will route Copilot tasks between local and cloud models in experimental preview later in October, with no inference charge for local calls; MXC agent containment went GA the same day.
Key Takeaways
- ✓On-device: 70.80% SWE-Bench Verified (cloud 72.6%) and 66.29% Terminal-Bench 2.1 (cloud 62.9%); GPT-OSS-120B scored 32.0% / 23.6% in the same test (Microsoft Command Line)
- ✓Footprint: ~3.3 bits/weight mixed precision plus DFlash2 speculative decoding; 75.5GB peak memory at 256K; 923.5 / 769.8 tok/s prompt processing at 64K / 128K
- ✓Routing: HydraFusion local+cloud routing reaches the Copilot app, Copilot CLI and VS Code in experimental preview later in October, with no inference charge for local calls (Windows Blog)
- ✓Available now: Copilot CLI 1.0.94-0+ discovers local Ollama models via
/model(tool calling + streaming required) (GitHub Changelog) - ✓Security: MXC agent containment is GA on Windows 11, already supported by Codex, Copilot, OpenClaw and LM Studio, with Claude Code and Hermes Agent among those to follow

Key Decision Metrics at a Glance
Turn your technical choice into a development budget
Compare 40 dev plans & simulate token costs vs $20/mo subscriptions
Project Links & Resources
Direct AccessIn-Depth Technical Analysis
At its Oct 7 Windows event Microsoft introduced 'hybrid intelligence': run agent work locally when it makes sense and reach the cloud for frontier capability, with no inference charge for local calls. MAI-Code-1.1-Flash, a 137B-total / 6.8B-active MoE coding model, now ships in an on-device build using ~3.3 bits-per-weight mixed-precision quantization (nearly 80% smaller) and DFlash2 sliding-window speculative decoding on a Windows ARM64 llama.cpp CUDA runtime, keeping a 256K context. In Microsoft's Oct 5 tests on Surface Laptop Ultra it scored 70.80% on SWE-Bench Verified (cloud 72.6%) and 66.29% on Terminal-Bench 2.1 (cloud 62.9%), versus 32.0% / 23.6% for GPT-OSS-120B, with 75.5GB peak memory at 256K and 923.5 / 769.8 tok/s prompt processing at 64K / 128K; these are vendor numbers without independent replication yet. GitHub's HydraFusion orchestrator will route Copilot tasks between local and cloud models in experimental preview later in October across the Copilot app, Copilot CLI and VS Code. Available today: Copilot CLI 1.0.94-0 discovers local Ollama models via /model (tool calling and streaming required; offline mode stays explicit via COPILOT_OFFLINE=true). Microsoft Execution Containers (MXC) went GA on Windows 11 for policy-enforced agent sandboxing, already supported by Codex, GitHub Copilot, OpenClaw, Replit, LM Studio and others, and llama.cpp support was added to Windows ML.
Benchmark side-by-side against alternatives, or calculate monthly token cost vs subscription break-even.
Discussion & Comments
0Sign in to join the discussion
Connect with AI developers to exchange benchmark insights.