Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Models

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai
  • Favicon for openai
    OpenAI GPT Astra LatestOpenAI GPT Astra Latest

    This model always redirects to the latest model in the OpenAI GPT Astra family.

    by openaiSep 11, 20261.05M context$10/M input tokens$50/M output tokens
  • Favicon for openai
    OpenAI GPT Sol LatestOpenAI GPT Sol Latest
    50% off

    This model always redirects to the latest model in the OpenAI GPT Sol family.

    by openaiSep 11, 20261.05M context$2/M input tokens$10/M output tokens
  • Favicon for openai
    OpenAI GPT Terra LatestOpenAI GPT Terra Latest

    This model always redirects to the latest model in the OpenAI GPT Terra family.

    by openaiSep 11, 20261.05M context$2/M input tokens$12/M output tokens
  • Favicon for openai
    OpenAI GPT Luna LatestOpenAI GPT Luna Latest

    This model always redirects to the latest model in the OpenAI GPT Luna family.

    by openaiSep 11, 20261.05M context$0.20/M input tokens$1.20/M output tokens
  • Favicon for sakana
    Sakana: Fugu Ultra v2Fugu Ultra v2
    89.7M tokens

    Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route tasks across a fixed pool of open and specialized models and to recursively call instances of itself. Fugu Ultra v2 prioritizes answer quality on complex multi-step reasoning, autonomous research, and full-stack software development, and does not rely on individual proprietary frontier models in its pool. It supports configurable reasoning effort (high, xhigh, max), function calling, structured outputs, image and PDF input, and built-in web search. Prompts above 272K tokens are billed at a higher rate. Orchestration tokens consumed by the system are billed as standard input/output tokens.

    by sakanaSep 11, 20261M context$5/M input tokens$30/M output tokens
  • Favicon for sakana
    Sakana: Fugu MaxFugu Max
    32M tokens

    Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route tasks across a fixed pool of open-weights and specialized models, including NVIDIA's Nemotron family, and to recursively call instances of itself. Fugu Max dynamically selects efficient combinations of expert agents to improve quality and cost together, and is priced flat regardless of context length. It supports configurable reasoning effort (high, xhigh, max), function calling, structured outputs, image and PDF input, and built-in web search and web fetch. Orchestration tokens consumed by the system are billed as standard input/output tokens.

    by sakanaSep 11, 20261M context$2/M input tokens$6/M output tokens
  • Favicon for black-forest-labs
    Black Forest Labs: FLUX Video EditFLUX Video Edit
    <1 hour

    FLUX Video Edit [fast] takes a source video and an edit prompt and returns a precisely edited video. Add, remove, or replace objects and characters, rebuild the setting, edit on-screen text, change colors, materials and visual effects, change or translate dialogue with lip sync, or change the events of a clip. Output preserves the source duration, aspect ratio, and audio; inputs above 720p are downscaled to 720p. Source clips up to 15 seconds and 50 MiB.

    by black-forest-labsSep 10, 2026$0.03/second
  • Favicon for inclusionai
    inclusionAI: Ling 3.0 Flash VL (free)Ling 3.0 Flash VL (free)Free variant
    21.8B tokens

    Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual agent capabilities. Hybrid instant/reasoning model with tool calling.

    by inclusionaiSep 10, 2026262K context$0/M input tokens$0/M output tokens
  • Favicon for inclusionai
    inclusionAI: Ling 3.0 Flash VLLing 3.0 Flash VL
    6.73M tokens

    Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual agent capabilities. Hybrid instant/reasoning model with tool calling.

    by inclusionaiSep 10, 2026131K context$0.06/M input tokens$0.18/M output tokens
  • Favicon for deepseek
    DeepSeek: DeepSeek V4.1 FlashDeepSeek V4.1 Flash
    1.4T tokens
    Finance (#12)
    Marketing (#44)
    Programming (#22)
    Roleplay (#47)
    Science (#28)

    DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on output from a 552B-parameter backbone, an asymmetric split that keeps per-token compute low relative to the model's total size. Image understanding is native to the architecture, with visual and text embeddings trained jointly from the start of pre-training rather than added afterward as in the earlier experimental V4 Flash Vision Exp. It is suited for coding, terminal, and computer-use agents, along with long-horizon tasks that must run to completion across many steps and long-context analysis. Compressed KV caching cuts cache memory to roughly a quarter of the previous Flash generation, significantly reducing costs on agentic workloads. DeepSeek positions it as the cost-efficient tier of the V4.1 family and reports that it exceeds V4 Pro on performance, speed, and task completion time.

    by deepseekSep 10, 20261.05M context$0.15/M input tokens$0.60/M output tokens
  • Favicon for openai
    OpenAI: GPT Image 2.5 SunburstGPT Image 2.5 Sunburst
    509M tokens

    GPT Image 2.5 Sunburst is an image generation and editing model from OpenAI, positioned as the precision-oriented tier of the GPT Image 2.5 series. It is suited to detailed creative work where editing accuracy matters more than generation speed, via the dedicated Images API.

    by openaiSep 9, 2026400K context$8/M input tokens$30/M output tokens
  • Favicon for openai
    OpenAI: GPT Image 2.5 FlareGPT Image 2.5 Flare
    161M tokens

    GPT Image 2.5 Flare is an image generation and editing model from OpenAI, positioned as the speed-oriented tier of the GPT Image 2.5 series. It is suited to high-volume everyday generation, creator content, and rapid prototyping via the dedicated Images API.

    by openaiSep 9, 2026400K context$8/M input tokens$30/M output tokens
  • Favicon for inception
    Inception: Mercury 2.5Mercury 2.5
    80% off
    20.6B tokens

    Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving 1,107 tokens/sec on standard GPUs. It delivers a 10+ point jump in intelligence over Mercury 2, comparable quality to cost-optimized frontier models like GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Mercury 2.5 supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output. It's built for production workloads where latency compounds: search agents, voice pipelines, and coding subagents. Read more in the blog post.

    by inceptionSep 8, 2026260K context$0.04/M input tokens$0.15/M output tokens
  • Favicon for nex-agi
    Nex AGI: Nex-N2.5-Mini (free)Nex-N2.5-Mini (free)Free variant
    62.3B tokens

    Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file changes, run commands, launch applications, interact with browser and desktop interfaces, and test software from the user's perspective. When observed behavior does not match the intended result, Nex-N2.5 can diagnose the issue, revise its implementation, and test again. This makes it especially effective for autonomous software engineering, GUI-based QA, computer-use automation, deep research, and scientific workflows where success must be demonstrated in the environment—not merely inferred from generated code.

    by nex-agiSep 8, 2026262K context$0/M input tokens$0/M output tokens
  • Favicon for nex-agi
    Nex AGI: Nex-N2.5-Pro (free)Nex-N2.5-Pro (free)Free variant
    167B tokens

    Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback loop: it can explore codebases, implement multi-file changes, run commands, launch applications, interact with browser and desktop interfaces, and test software from the user's perspective. When observed behavior does not match the intended result, Nex-N2.5 can diagnose the issue, revise its implementation, and test again. This makes it especially effective for autonomous software engineering, GUI-based QA, computer-use automation, deep research, and scientific workflows where success must be demonstrated in the environment—not merely inferred from generated code.

    by nex-agiSep 8, 2026262K context$0/M input tokens$0/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Astra (batch)GPT-6 Astra (batch)Batch variant
    144M tokens

    GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use.

    by openaiSep 4, 20261.05M context$5/M input tokens$25/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 AstraGPT-6 Astra
    666B tokens
    Academia (#50)
    Finance (#48)
    Marketing (#47)
    Programming (#23)
    Science (#22)

    GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon agentic tasks that involve computer and browser use.

    by openaiSep 4, 20261.05M context$10/M input tokens$50/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Astra Pro (batch)GPT-6 Astra Pro (batch)Batch variant
    86.6M tokens

    GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

    by openaiSep 4, 20261.05M context$5/M input tokens$25/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Astra ProGPT-6 Astra Pro
    118B tokens

    GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

    by openaiSep 4, 20261.05M context$10/M input tokens$50/M output tokens
  • Favicon for microsoft
    Microsoft AI: MAI-Image-2.6MAI-Image-2.6
    47.6M tokens

    MAI-Image-2.6 is an image generation and editing model from Microsoft AI, the precision tier of the MAI-Image-2.6 family alongside the faster MAI-Image-2.6 Flash. It is suited for design-ready visuals and controlled, iterative editing, and is particularly strong at multi-reference editing that combines people, products, styles, and scenes from up to five input images while preserving composition, style, and detail. Aspect ratio can be fixed or left to the model to choose for the composition.

    by microsoftSep 4, 20264K context$8/M input tokens$38/M output tokens