Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Models

Models

CompareDiscover Models
Favicon for anthropic
Favicon for openai
  • Favicon for openai
    OpenAI: GPT-6 Luna ProGPT-6 Luna Pro
    1.35B tokens

    GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

    by openaiSep 22, 20261.05M context$0.10/M input tokens$0.50/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Luna Pro (batch)GPT-6 Luna Pro (batch)Batch variant
    56.4M tokens

    GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

    by openaiSep 22, 20261.05M context$0.05/M input tokens$0.25/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 LunaGPT-6 Luna
    5.99B tokens

    GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic tasks, and at higher reasoning effort it can take on complex software engineering and computer-use tasks that previously called for a Sol-tier model. It shares the GPT-6 family's gains in factual reliability and its clearer, more concise communication style.

    by openaiSep 22, 20261.05M context$0.10/M input tokens$0.50/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Luna (batch)GPT-6 Luna (batch)Batch variant
    229K tokens

    GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and latency-sensitive workloads such as chat, classification, and lightweight agentic tasks, and at higher reasoning effort it can take on complex software engineering and computer-use tasks that previously called for a Sol-tier model. It shares the GPT-6 family's gains in factual reliability and its clearer, more concise communication style.

    by openaiSep 22, 20261.05M context$0.05/M input tokens$0.25/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Sol ProGPT-6 Sol Pro
    148M tokens

    GPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

    by openaiSep 22, 20261.05M context$2/M input tokens$10/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Sol Pro (batch)GPT-6 Sol Pro (batch)Batch variant

    GPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

    by openaiSep 22, 20261.05M context$1/M input tokens$5/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 SolGPT-6 Sol
    3.17B tokens

    GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional work, agentic coding, business workflow automation, and computer use, and is particularly strong at long-horizon software engineering tasks in real codebases. It approaches Astra-level factual reliability at a much lower cost and shares Astra's clearer, more concise communication style in technical and coding conversations.

    by openaiSep 22, 20261.05M context$2/M input tokens$10/M output tokens
  • Favicon for openai
    OpenAI: GPT-6 Sol (batch)GPT-6 Sol (batch)Batch variant
    41K tokens

    GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna tier. It is suited for demanding professional work, agentic coding, business workflow automation, and computer use, and is particularly strong at long-horizon software engineering tasks in real codebases. It approaches Astra-level factual reliability at a much lower cost and shares Astra's clearer, more concise communication style in technical and coding conversations.

    by openaiSep 22, 20261.05M context$1/M input tokens$5/M output tokens
  • Favicon for inclusionai
    inclusionAI: Ming Image 0.1 DesignMing Image 0.1 Design
    4.3M tokens

    Ming Image 0.1 Design is a text-to-image model from inclusionAI aimed at graphic-design output, with an emphasis on legible text rendering inside the generated image. It generates from a prompt only and does not accept reference images. Output format can be requested as PNG, JPEG, or WebP. Image dimensions are chosen by the model rather than by the request, so explicit sizes and aspect ratios are rejected instead of silently reshaped.

    by inclusionaiSep 22, 2026$0/M input tokens$0/M output tokens
  • Favicon for anthropic
    Anthropic: Claude Opus 5.5 (batch)Claude Opus 5.5 (batch)Batch variant
    855K tokens

    Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code review and bug finding, financial and scientific analysis, and reading dense charts, diagrams, and screenshots, and it is more careful than its predecessor about only stating figures and citing sources it can back up. The model completes comparable tasks in fewer steps and with fewer tokens than Opus 5, and reports on its work in plainer language, with clear updates on what it did, what it found, and what it needs from the user. Thinking is always adaptive, so effort is the main lever for trading off depth, latency, and cost, and lower effort settings remain effective for latency-sensitive workloads.

    by anthropicSep 22, 20261M context$2/M input tokens$10/M output tokens
  • Favicon for anthropic
    Anthropic: Claude Opus 5.5Claude Opus 5.5
    8.72B tokens

    Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is particularly strong at multi-step changes in large codebases, code review and bug finding, financial and scientific analysis, and reading dense charts, diagrams, and screenshots, and it is more careful than its predecessor about only stating figures and citing sources it can back up. The model completes comparable tasks in fewer steps and with fewer tokens than Opus 5, and reports on its work in plainer language, with clear updates on what it did, what it found, and what it needs from the user. Thinking is always adaptive, so effort is the main lever for trading off depth, latency, and cost, and lower effort settings remain effective for latency-sensitive workloads.

    by anthropicSep 22, 20261M context$4/M input tokens$20/M output tokens
  • Favicon for assemblyai
    AssemblyAI: Universal-3.5 ProUniversal-3.5 Pro
    50% off
    115K characters

    Universal-3.5 Pro is AssemblyAI's speech-to-text model served through its Sync API, returning a complete transcript with word-level timestamps in a single synchronous response for audio clips up to 120 seconds. It accepts 16-bit WAV input and supports free-text prompting, keyterms, and conversation context to steer transcription toward domain vocabulary.

    by assemblyaiSep 22, 2026from $0.000063/second
  • Favicon for xiaomi
    Xiaomi: MiMo-V2.6-Pro-UltraSpeedMiMo-V2.6-Pro-UltraSpeed
    7.64B tokens

    MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly 10x the output speed. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers top-tier performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.

    by xiaomiSep 21, 20261.05M context$4.35/M input tokens$8.70/M output tokens
  • Favicon for xiaomi
    Xiaomi: MiMo-V2.6-FlashMiMo-V2.6-Flash
    115B tokens

    MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and 15B activated per token, it employs a hybrid attention mechanism for greater computational efficiency. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers strong performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.

    by xiaomiSep 21, 20261.05M context$0.14/M input tokens$0.28/M output tokens
  • Favicon for xiaomi
    Xiaomi: MiMo-V2.6-ProMiMo-V2.6-Pro
    157B tokens

    MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of capability for the most demanding workloads. The model features a 1M-token context window and native multimodal capabilities. Optimized for agentic workflows, it delivers top-tier performance across coding, visual, general, and research scenarios, excelling at complex, long-horizon tasks with robust generalization across a diverse range of agent harnesses.

    by xiaomiSep 21, 20261.05M context$0.435/M input tokens$0.87/M output tokens
  • Favicon for x-ai
    SpaceXAI: Grok 4.7Grok 4.7
    62.3B tokens

    Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running software engineering tasks, verifying its own work, and managing long context, and it improves on its predecessor at professional knowledge work such as drafting documents and presentations. The model was trained with a longer reinforcement learning run weighted toward problems that take many hours to complete, and natively understands the Grok Bot harness for conversational tasks. It ships with a new safeguard stack that pairs strong jailbreak resistance with low refusal rates for legitimate cybersecurity and biology work. SpaceXAI's reported benchmark results use the xhigh reasoning effort.

    by x-aiSep 21, 2026500K context$1.60/M input tokens$4.80/M output tokens
  • Favicon for qwen
    Qwen: Qwen3.8 Omni FlashQwen3.8 Omni Flash
    61K tokens

    Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video understanding. It is suited for audio-video analysis and summarization, video editing and production workflows, audio-video dialogue, coding, knowledge work, and GUI interaction, and it is particularly strong at long-form multimedia tasks that combine speech, sound, and visual context. It also supports two-channel and four-channel spatial audio understanding.

    by qwenSep 21, 20261M context$0.15/M input tokens$0.47/M output tokens
  • Favicon for prism-ml
    PrismML: Ternary Bonsai 2 27BTernary Bonsai 2 27B
    222M tokens

    Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image understanding with a 262K-token context window. Ternary compression shrinks the language-model weights to roughly 8.5 GB while retaining 98.2% of the base model's average score across PrismML's 14 thinking-mode benchmarks, enabling efficient inference on consumer hardware. The model thinks by default and defaults to xhigh reasoning effort.

    by prism-mlSep 18, 2026262K context$0.075/M input tokens$0.50/M output tokens
  • Favicon for z-ai
    Z.ai: GLM 5.3 FlashXGLM 5.3 FlashX
    112B tokens
    Programming (#43)

    GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture (320B total parameters, 18B active), it is suited for efficient coding, visual understanding, and long-horizon agent tasks with a 1M-token context window.

    by z-aiSep 18, 20261.05M context$0.37/M input tokens$1.25/M output tokens
  • Favicon for typesafe
    TypeSafe: Jev LatestJev Latest

    This model always redirects to the latest model in the Jev family.

    by typesafeSep 18, 202632K context$0.042/M input tokens$0/M output tokens