Skip to content
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Apps
  • Discover
  • Models
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • SDK
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Coding

Best AI Models for Coding

Model rankings updated July 2026 based on real usage data.

Compare the best AI models for coding, ranked by real usage from developers on OpenRouter. Whether you're generating code, debugging, refactoring or building an AI coding assistant, these LLMs deliver strong performance across popular languages and frameworks.

This collection features top coding models from Anthropic, Google, xAI, OpenAI and more, all accessible through a single API. From agentic coding workflows to one-off code generation, find the right model for your engineering needs.

Browse All ModelsCompare Models

LLM Leaderboard for Programming Models

1.
Favicon for xiaomi
Mimo V2.5
by xiaomi
2.21T
46.6%
2.
Favicon for nvidia
Nemotron 3 Ultra 550B A55B (free)
by nvidia
338B
7.1%
3.
Favicon for deepseek
Deepseek V4 Pro
by deepseek
281B
5.9%
4.
Favicon for z-ai
GLM 5.2
by z-ai
268B
5.6%
5.
Favicon for deepseek
Deepseek V4 Flash
by deepseek
219B
4.6%
6.
Favicon for tencent
Hy3
by tencent
174B
3.7%
7.
Favicon for inclusionai
Ling 3.0 Flash (free)
by inclusionai
163B
3.4%
8.
Favicon for minimax
Minimax M3
by minimax
141B
3.0%
9.
Favicon for moonshotai
Kimi K3
by moonshotai
116B
2.4%
10.
Favicon for unknown
Others
837B
17.6%

Top Coding Models on OpenRouter

Based on top weekly usage data from millions of users accessing AI models for coding through OpenRouter.

Favicon for xiaomi

Xiaomi: MiMo-V2.5

11.4T tokens

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding tasks. Its 1M context window supports complete documents, extended conversations, and complex task contexts in a single pass, making it ideal for integration with agent frameworks where strong reasoning, rich perception, and cost efficiency all matter.

by xiaomi1.05M context$0.112/M input tokens$0.224/M output tokens20% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Flash

7.12T tokens

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and high-throughput workloads, while maintaining strong reasoning and coding performance.

The model includes hybrid attention for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for applications such as coding assistants, chat systems, and agent workflows where responsiveness and cost efficiency are important.

by deepseek1.05M context$0.09/M input tokens$0.18/M output tokens
Favicon for tencent

Tencent: Hy3

5.21T tokens

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort: a direct no-think mode by default, plus low and high chain-of-thought modes for complex math, coding, and multi-step problems. With a 256K context window, Hy3 targets long-horizon tasks, including improved coreference resolution, multi-turn constraint tracking, and stable tool-calling that generalizes across agent scaffoldings.

Tencent positions it as a reliable, cost-effective option across coding, document processing, financial analysis, game development, and frontend design, with a strong emphasis on grounded, anti-hallucination behavior that answers when grounded and flags when evidence is missing rather than fabricating.

by tencent262K context$0.1288/M input tokens$0.5336/M output tokens8% off
Favicon for deepseek

DeepSeek: DeepSeek V4 Pro

3.54T tokens

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding, and long-horizon agent workflows, with strong performance across knowledge, math, and software engineering benchmarks.

Built on the same architecture as DeepSeek V4 Flash, it introduces a hybrid attention system for efficient long-context processing. Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is well suited for complex workloads such as full-codebase analysis, multi-step automation, and large-scale information synthesis, where both capability and efficiency are critical.

by deepseek1.05M context$0.435/M input tokens$0.87/M output tokens
Favicon for z-ai

Z.ai: GLM 5.2

3.49T tokens

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering, and complex multi-step automation.

Reasoning efforts high and xhigh are supported; xhigh maps to max reasoning. It is particularly strong at coding and tool use across long-running tasks, able to maintain engineering context and follow standards consistently through a full development workflow, from requirements to multi-platform deployment, in a single task.

by z-ai1.05M context$0.76/M input tokens$2.42/M output tokens
Favicon for nvidia

NVIDIA: Nemotron 3 Ultra (free)

2.61T tokens

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it supports text input and output with a context window of up to 1M tokens. It is suited for long-running agentic workflows, including agent orchestration, coding agents, deep research, and complex enterprise tasks.

It is particularly strong at multi-step reasoning and planning, with high-throughput inference designed for high-volume agent pipelines. It is part of the NVIDIA Nemotron family of open models for agentic AI.

by nvidia1M context$0/M input tokens$0/M output tokens
Favicon for minimax

MiniMax: MiniMax M3

2.2T tokens

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding, and tool use. It is built on MiniMax Sparse Attention (MSA), which replaces full attention with KV-block selection to cut per-token compute at long context — roughly 1/20 the cost of the previous generation at 1M tokens, with substantially faster prefill and decode while retaining quality across most tasks.

Trained as a native multimodal model on interleaved data and tuned for multi-turn, production-like collaboration via an interactive user-simulator framework, the model is oriented toward sustained, multi-step tasks rather than single-turn execution.

by minimax1.05M context$0.24/M input tokens$0.96/M output tokens60% off
Favicon for stepfun

StepFun: Step 3.7 Flash

2.11T tokens

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters per token. The model supports a 256K context window and exposes selectable reasoning levels (high/medium/low), letting callers trade off speed, cost, and depth of reasoning.

Designed for coding, agentic workflows, structured outputs, and long-context productivity tasks.

by stepfun262K context$0.20/M input tokens$1.15/M output tokens
Favicon for anthropic

Anthropic: Claude Opus 4.8

1.36T tokens

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token context window. It is suited for highly autonomous agents, long-horizon agentic work, knowledge work, and memory-driven tasks where coherence over extended sessions matters.

It is particularly strong on multi-step reasoning, complex coding, and end-to-end project orchestration - large codebases, multi-stage debugging, and long-running asynchronous agent pipelines. Beyond coding, it handles knowledge work such as drafting documents, building presentations, and analyzing data, maintaining quality across very long outputs.

by anthropic1M context$5/M input tokens$25/M output tokens
Favicon for moonshotai

MoonshotAI: Kimi K3

1.32T tokens

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback. Its architecture uses KDA and Attention Residuals for computational efficiency.

by moonshotai1.05M context$3/M input tokens$15/M output tokens