Microsoft rolled out MAI Code 1.1 Flash as the next-generation coding model for GitHub Copilot, but early benchmarks and pricing tables put it behind DeepSeek’s open-weight offering on both performance and cost.

Benchmark reality versus the press release

Microsoft’s launch notes promise a 25 % jump in token efficiency and a 75 % cut in cost compared with the June version of its MAI model. The company also touts a 4 % rise in “code survival” after training the model in hundreds of thousands of reinforcement-learning environments inside Copilot.

Independent testing tells a different story. On the SWE-bench Verified suite, MAI Code 1.1 Flash records a 72.6 % pass rate, nudging ahead of Claude Haiku 4.5 (69.8 %) and GPT-5.4 mini (69.2 %). The advantage is modest and limited to a single metric.

The gap widens on Terminal Bench 2.1, which measures a model’s ability to complete real-world command-line tasks. Here MAI Code 1.1 Flash scores 62.9 %, while DeepSeek-V4-Flash-0731 posts an 82.7 % result. That difference is large enough to affect developers who rely on terminal-centric code generation.

Pricing pressure points

Microsoft markets the new model as a “budget-friendly” option, but token pricing says otherwise. DeepSeek-V4-Flash charges $0.14 per input token and $0.28 per output token. MAI Code 1.1 Flash lists $0.20 for input and $1.20 for output. Claude Haiku 4.5 sits at $1.00 and $5.00 respectively, confirming its premium status.

The steep jump in output-token cost means any workflow that generates large blocks of code will see Microsoft’s model become the more expensive choice, even after the promised reductions from its predecessor. For high-scale developers and enterprises, the economics tilt sharply toward DeepSeek.

How MAI Code 1.1 Flash arrived

The model is part of Microsoft’s broader push to replace third-party services in the Copilot stack with internally built alternatives. By training on extensive reinforcement-learning data drawn from Copilot usage, Microsoft hopes to tighten the feedback loop between user behavior and model improvement. The launch also follows a pattern of incremental upgrades: the June version was positioned as a cost-saver, and the Flash iteration is framed as the next step in that trajectory.

Winners, losers and the broader stakes

Enterprises looking to control AI spend may find DeepSeek’s pricing more attractive, especially when output volume is high. Microsoft, meanwhile, gains tighter control over the Copilot ecosystem and the ability to steer revenue away from external providers such as OpenAI or Anthropic.

Microsoft’s possible defense

Microsoft’s own data emphasizes token-efficiency gains and a modest edge on SWE-bench.

What to watch next

  • Adoption metrics: Copilot’s usage statistics will reveal whether developers switch to the new default or import alternative models via the API.
  • Pricing adjustments: If Microsoft’s output-token price stays high, market pressure could force a revision.
  • Benchmark updates: New versions of terminal-focused tests may shift the performance balance, especially as Microsoft refines its reinforcement-learning pipeline.
  • Open-weight competition: Additional open-weight models entering the coding space could intensify the price-performance race.

Takeaway

MAI Code 1.1 Flash brings modest gains on a narrow benchmark but falls short of DeepSeek’s open-weight model on both terminal performance and token cost, leaving developers to weigh integration convenience against raw efficiency.