AI Models Ranking
The most capable AI models ranked and compared
| Rank | Model | Developer | Modality | Context | Score | Actions |
|---|---|---|---|---|---|---|
| #1 | Anthropic: Claude Sonnet 5 Claude Sonnet 5 is Anthropic's frontier model featuring a 1,000,000-token context window and selectable reasoning effort levels. It supports up to 128,000 output tokens per request at lower pricing than Claude 3.5 Sonnet and GPT-4o. The model is optimized for complex coding, agentic automation, and long-context analysis. | Anthropic | text | 1000K | 9.7 | |
| #2 | GPT-4o OpenAI's flagship multimodal model with best-in-class reasoning and vision. | OpenAI | text | 128K | 9.7 | |
| #3 | Claude Opus 5 (Fast) Claude Opus 5 (Fast) is Anthropic's high-throughput variant of Opus 5, featuring a 1M token context window and 128K max output tokens. It provides frontier-tier reasoning and coding performance at significantly higher output speeds. Priced at $10.00/1M input and $50.00/1M output tokens, it targets latency-critical enterprise workflows. | Anthropic | text | 1000K | 9.6 | |
| #4 | Claude Opus 5 Claude Opus 5 is Anthropic's flagship model engineered for complex reasoning, long-horizon agentic workflows, and end-to-end software engineering. It features a 1,000,000-token context window and a 128,000-token output limit, surpassing Claude 3.5 Sonnet and GPT-4o in generation depth and context retention. Priced at $5.00 input and $25.00 output per million tokens, it targets enterprise workloads requiring extreme context handling. | Anthropic | text | 1000K | 9.6 | |
| #5 | OpenAI: GPT-5.6 Luna Pro GPT-5.6 Luna Pro is OpenAI's reasoning-focused variant of the Luna architecture, configured with extended pro-mode inference for complex technical tasks. It features a 1.05M token context window and 128K maximum output limit at $0.20 per million input tokens and $1.20 per million output tokens. It accepts text, image, and file inputs, undercutting GPT-4o and Claude 3.5 Sonnet on both cost and context capacity. | Openai | text | 1050K | 9.6 | |
| #6 | OpenAI: GPT-5.6 Terra Pro GPT-5.6 Terra Pro is OpenAI's extended-reasoning variant of the GPT-5.6 Terra model, utilizing Pro reasoning mode for high-complexity analytical and technical tasks. It pairs a 1.05M token context window with a 128K max output token capacity, competing directly against Claude 3.5 Sonnet and GPT-4o for complex multi-step reasoning. | Openai | text | 1050K | 9.6 | |
| #7 | OpenAI: GPT-5.6 Terra Pro (batch) OpenAI's GPT-5.6 Terra Pro (batch) pairs extended pro reasoning depth with a 1,050,000 token context window for asynchronous workloads. Priced at $1.00 input and $6.00 output per million tokens, it offers a cost-effective option for large-scale analysis. It supports up to 128,000 output tokens per request. | Openai | text | 1050K | 9.6 | |
| #8 | OpenAI: GPT-5.6 Sol Pro Developed by OpenAI, GPT-5.6 Sol Pro runs on an extended reasoning engine designed for complex logic, math, and code synthesis. It features a massive 1.05M token context window and 128K max output limit, outperforming GPT-4o and Claude 3.5 Sonnet on deep analytical tasks at $2.00/$10.00 per million tokens. | Openai | text | 1050K | 9.6 | |
| #9 | OpenAI: GPT-5.6 Sol Pro (batch) OpenAI's GPT-5.6 Sol Pro (batch) runs the GPT-5.6 architecture with high-compute 'pro' reasoning enabled over asynchronous batch APIs. Featuring a 1.05M token context window and 128K output limit at $1.00/$5.00 per million tokens, it significantly surpasses GPT-4o and Claude 3.5 Sonnet in large-scale technical analysis and complex reasoning workloads. | Openai | text | 1050K | 9.6 | |
| #10 | OpenAI: GPT-5.6 Sol GPT-5.6 Sol is OpenAI's flagship model tailored for complex reasoning, multi-step coding, and agentic workflows. It offers a 1.05M token context window alongside a massive 128K max output token capacity. Priced at $2.00 input and $10.00 output per million tokens, it undercuts GPT-4o and Claude 3.5 Sonnet on input costs while significantly expanding context capabilities. | Openai | text | 1050K | 9.6 | |
| #11 | OpenAI: GPT-5.6 Sol (batch) GPT-5.6 Sol (Batch) is OpenAI's flagship reasoning and multi-step coding model accessible via asynchronous batch processing. It provides a 1,050,000-token context window with up to 128,000 output tokens at a discounted price point. The model targets high-volume agentic workflows, long-context code synthesis, and complex background analysis. | Openai | text | 1050K | 9.6 | |
| #12 | Anthropic: Claude Sonnet 5 (batch) Claude Sonnet 5 (Batch) delivers Anthropic's flagship coding, reasoning, and agentic capabilities via asynchronous batch processing. It features a 1,000,000-token context window and a 128,000-token maximum output limit at a discounted rate of $1.00/1M input tokens and $5.00/1M output tokens. | Anthropic | text | 1000K | 9.6 | |
| #13 | Claude 3.5 Sonnet Anthropic's most capable model — excels at analysis, coding, and nuanced writing. | Anthropic | text | 200K | 9.6 | |
| #14 | Claude Opus 5 (batch) Claude Opus 5 (batch) delivers Anthropic's flagship reasoning and long-horizon coding capabilities via discounted asynchronous batch processing. Featuring a 1,000,000 token context window and a massive 128,000 output token limit, it handles complex codebase analysis and agentic tasks far beyond GPT-4o context boundaries. | Anthropic | text | 1000K | 9.5 | |
| #15 | OpenAI: GPT-5.6 Luna Pro (batch) GPT-5.6 Luna Pro (batch) pairs OpenAI's Luna model in high-reasoning 'pro' mode with batch execution. It features a 1.05M token context window and 128K output capacity at ultra-low pricing ($0.10/1M input). It targets non-real-time asynchronous workloads requiring deep analytical reasoning and massive context ingestion. | Openai | text | 1050K | 9.5 | |
| #16 | Anthropic: Claude Opus 4.8 Anthropic Claude Opus 4.8 is a flagship frontier model built for complex reasoning, large-scale codebases, and extensive document analysis. It features a 1M-token context window alongside a 128K output token capacity, outperforming GPT-4o and Claude 3.5 Sonnet on long-context processing. It is priced at $5.00/1M input tokens and $25.00/1M output tokens. | Anthropic | text | 1000K | 9.5 | |
| #17 | Midjourney v6 Industry-leading image generation with photorealistic quality and artistic range. | Midjourney | image | — | 9.5 | |
| #18 | Anthropic: Claude Fable 5 Anthropic's Claude Fable 5 is a high-capability model optimized for complex coding and autonomous knowledge work. It features a 1M-token context window and a 128K output token limit, making it ideal for deep reasoning over large codebases and complex documents. | Anthropic | text | 1000K | 9.4 | |
| #19 | Anthropic: Claude Fable 5 (batch) Claude Fable 5 (batch) is Anthropic's high-capacity model optimized for asynchronous autonomous knowledge work and repository-scale coding. It provides a 1,000,000-token input context window and a 128,000-token max output limit for text, image, and file inputs. Priced at $5.00 input and $25.00 output per million tokens, it targets large-scale automated reasoning. | Anthropic | text | 1000K | 9.4 | |
| #20 | Anthropic: Claude Opus 4.8 (Fast) Claude Opus 4.8 (Fast) is Anthropic's flagship model variant optimized for higher output generation speed. It features a 1,000,000-token context window and a 128,000-token max output limit for complex analysis and large codebases. It provides identical intelligence to standard Opus 4.8 at double the API cost. | Anthropic | text | 1000K | 9.4 | |
| #21 | Anthropic: Claude Opus 4.8 (batch) Anthropic's Claude Opus 4.8 (batch) delivers top-tier reasoning and multimodal processing at a 50% cost reduction via asynchronous execution. It features a 1M-token input context window and 128K max output tokens for extensive document and codebase analysis. It is designed for latency-tolerant, high-volume enterprise workloads. | Anthropic | text | 1000K | 9.4 | |
| #22 | Gemini 1.5 Pro Google's longest-context model with 1M token window and strong multimodal support. | text | 1000K | 9.4 | ||
| #23 | Google: Gemini 3.7 Flash (batch) Gemini 3.7 Flash (batch) is Google's high-efficiency multimodal model designed for asynchronous workloads and complex agentic tasks. It combines a 1-million-token context window with native reasoning and a 64K output limit at half the cost of standard API endpoints. It targets bulk data processing, long-context code analysis, and high-volume media ingestion. | text | 1049K | 9.3 | ||
| #24 | Qwen: Qwen3.8 Max Qwen3.8 Max is Alibaba's flagship multimodal reasoning model featuring a 1,000,000-token context window and support for text, image, and video inputs. It offers competitive performance across complex reasoning and long-context tasks at lower price points ($2.00/1M input, $6.00/1M output) than GPT-4o and Claude 3.5 Sonnet. Its 131,072-token max output capability makes it uniquely suited for large-scale code and document generation. | Qwen | text | 1000K | 9.3 | |
| #25 | OpenAI: GPT-5.6 Terra (batch) GPT-5.6 Terra (batch) is OpenAI's mid-tier model designed for asynchronous, high-volume workloads involving text, images, and files. Featuring a 1.05M-token context window and a 128K output limit, it drastically exceeds GPT-4o and Claude 3.5 Sonnet in sequence capacity. At $1.00/1M input tokens via batch API, it provides a cost-efficient solution for offline agentic tasks and bulk document processing. | Openai | text | 1050K | 9.3 | |
| #26 | Google: Gemini 3.7 Flash Gemini 3.7 Flash is a multimodal model from Google built for fast agentic execution, complex coding, and multi-step reasoning. It features a 1,048,576 token input context window, a 65,536 token max output window, and native support for audio, video, image, and text inputs. Priced at $0.38 per million input tokens, it offers low-cost, high-throughput execution compared to GPT-4o and Claude 3.5 Sonnet. | text | 1049K | 9.2 | ||
| #27 | MoonshotAI: Kimi K3 Kimi K3 by Moonshot AI is a 2.8-trillion parameter open-weight multimodal reasoning model built for complex coding and agentic workflows. It supports text, image, and video inputs paired with a 1,048,576-token context window and a 943K output limit. At $3.00/1M input and $15.00/1M output tokens, it matches Claude 3.5 Sonnet's pricing while offering significantly larger context capacity. | Moonshotai | text | 1049K | 9.2 | |
| #28 | OpenAI: GPT-5.6 Terra OpenAI's GPT-5.6 Terra is a mid-tier model designed for coding, reasoning, and agentic workflows. It features a 1,050,000-token context window and a 128,000 max output limit at $2.00 input and $12.00 output per million tokens. It significantly expands context capacity and output length compared to GPT-4o and Claude 3.5 Sonnet. | Openai | text | 1050K | 9.2 | |
| #29 | Llama 3.1 405B The largest open-weight model — competitive with proprietary leaders. | Meta | text | 128K | 9.2 | |
| #30 | DeepSeek: DeepSeek V4 Flash Vision Exp DeepSeek V4 Flash Vision Exp is an experimental multimodal model by DeepSeek. It extends the DeepSeek V4 Flash 0731 text model with image understanding capabilities, maintaining its existing text performance, and features an industry-leading 1M token context window at highly competitive pricing. | Deepseek | text | 1049K | 9.1 | |
| #31 | Z.ai: GLM 5.3 Flash GLM-5.3-Flash from Z.ai is a native multimodal model supporting text, image, and video inputs with a 1.3M-token context window. Built on a hybrid sparse-linear attention architecture, it is designed for low-latency coding and agent execution. It offers exceptionally low operational costs at $0.07 per million input tokens. | Z AI | text | 1311K | 9.1 | |
| #32 | Qwen: Qwen3.8 2.4T A95B Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model created by Alibaba's Qwen team, utilizing 95 billion active parameters out of 2.4 trillion total. It provides a 1,048,576 token context window and a 131,072 max output length at pricing lower than GPT-4o and Claude 3.5 Sonnet. | Qwen | text | 1049K | 9.1 | |
| #33 | DeepSeek: DeepSeek V4 Pro 0813 DeepSeek V4 Pro 0813 is a large-scale Mixture-of-Experts text model featuring a 1M-token context window and up to 943K output tokens. Priced at $1.12/1M input and $3.37/1M output, it offers extensive long-context processing at a fraction of the cost of GPT-4o and Claude 3.5 Sonnet. | Deepseek | text | 1049K | 9.1 | |
| #34 | SpaceXAI: Grok 4.6 Developed by xAI, Grok 4.6 is a high-performance language model featuring a 500,000-token context window and a 450,000-token maximum output limit. Priced at $2.00 per million input tokens and $6.00 per million output tokens, it offers lower API costs and significantly larger output capacity compared to GPT-4o and Claude 3.5 Sonnet. | X AI | text | 500K | 9.1 | |
| #35 | Meta: Muse Spark 1.2 Meta's Muse Spark 1.2 is a reasoning model optimized for complex agentic workflows and long-context processing. It supports multimodal inputs—including text, audio, video, and documents—across a 1-million-token context window at a fraction of GPT-4o's API cost. | Meta | text | 1049K | 9.1 | |
| #36 | Google: Gemini 3.6 Flash Gemini 3.6 Flash is Google's high-efficiency model optimized for web development, agentic workflows, and fast multimodal processing. It combines a 1-million-token context window with a large 65,536 max output limit at a fraction of GPT-4o's API costs. The model trades top-tier complex reasoning for speed, massive context volume, and cost efficiency. | text | 1049K | 9.1 | ||
| #37 | Thinking Machines: Inkling Developed by Thinking Machines Lab, Inkling is an open-weight mixture-of-experts model with 41B active parameters out of 975B total. It features a 1M-token context window and handles text, image, and audio inputs for reasoning, coding, and agentic tasks. At $0.95/1M input tokens and a 262K output limit, it offers deep context processing at a fraction of GPT-4o costs. | Thinkingmachines | text | 1049K | 9.1 | |
| #38 | Thinking Machines: Inkling (batch) Inkling (batch) is an open-weight, 975B-parameter multimodal Mixture-of-Experts model (41B active) developed by Thinking Machines Lab. It accepts text, image, and audio inputs across a 524K context window, offering high-volume reasoning and coding at a fraction of GPT-4o and Claude 3.5 Sonnet API costs. | Thinkingmachines | text | 524K | 9.1 | |
| #39 | Meta: Muse Spark 1.1 Meta's Muse Spark 1.1 is a 1M-token context multimodal reasoning model engineered for agentic workflows across text, image, audio, and video inputs. At $1.25/1M input and $4.25/1M output tokens, it significantly undercuts GPT-4o and Claude 3.5 Sonnet on cost while providing a 943K max output token limit. It is targeted at heavy context ingestion, video/audio processing, and agent orchestration. | Meta | text | 1049K | 9.1 | |
| #40 | OpenAI: GPT-5.6 Luna Developed by OpenAI, GPT-5.6 Luna is a high-speed, cost-optimized model tailored for low-latency and high-volume operations. It provides a 1,050,000-token context window and a 128,000 max output token limit at $0.20 per million input tokens. It targets high-throughput classification, routing, and lightweight agentic tasks. | Openai | text | 1050K | 9.1 | |
| #41 | OpenAI: GPT-5.6 Luna (batch) OpenAI's GPT-5.6 Luna (batch) is a cost-optimized model designed for high-volume text and multimodal processing. It combines a 1.05M-token context window with a 128K output limit at $0.10 per million input tokens. It targets bulk asynchronous tasks like document extraction, classification, and large-scale data processing. | Openai | text | 1050K | 9.1 | |
| #42 | SpaceXAI: Grok 4.5 Grok 4.5 by xAI delivers high-tier performance in coding, STEM, and analytical tasks using a 500,000-token context window. At $2.00/1M input and $6.00/1M output tokens, it undercuts GPT-4o and Claude 3.5 Sonnet while offering unmatched 450,000 max output token limits. | X AI | text | 500K | 9.1 | |
| #43 | xAI: Grok Latest xAI Grok Latest routes API requests to xAI's flagship model, offering top-tier reasoning and vision capabilities. It features an industry-leading 500,000 token context window and a 450,000 max output limit. At $2.00/1M input and $6.00/1M output, it undercuts both GPT-4o and Claude 3.5 Sonnet on pricing. | X AI | text | 500K | 9.1 | |
| #44 | Google: Nano Banana Pro (Gemini 3 Pro Image) Nano Banana Pro (Gemini 3 Pro Image) is Google's multimodal model built for interleaved text and image generation and editing. It offers a 131,072 token context window with up to 32,768 max output tokens. Priced at $2.00/1M input and $12.00/1M output tokens, it combines visual generation with extended context reasoning. | image | 131K | 9.1 | ||
| #45 | Z.ai: GLM 5.2 GLM 5.2 is a large-scale reasoning model from Z.ai optimized for project-level software engineering and long-horizon agent workflows. It features a 1M-token context window and a 256K output token limit at pricing substantially lower than GPT-4o and Claude 3.5 Sonnet. | Z AI | text | 1049K | 9.1 | |
| #46 | NVIDIA: Nemotron 3 Ultra NVIDIA Nemotron 3 Ultra is a 550B parameter (55B active) hybrid Transformer-Mamba mixture-of-experts model built for complex reasoning and agent orchestration. It offers a 512K context window and a 461K max output token limit at low inference costs. Its architecture makes it a cost-effective alternative to GPT-4o and Claude 3.5 Sonnet for long-context workloads. | Nvidia | text | 512K | 9.1 | |
| #47 | Qwen: Qwen3.7 Plus Qwen3.7 Plus is Alibaba's high-efficiency multimodal model featuring a 1M token input context window and a 131K token output limit. At $0.32/1M input tokens, it processes text and image inputs at a fraction of the cost of GPT-4o and Claude 3.5 Sonnet. It is designed for large-scale document analysis, high-volume batch processing, and long-form text generation. | Qwen | text | 1000K | 9.1 | |
| #48 | Google: Gemini 3.5 Flash Google's Gemini 3.5 Flash is a multimodal model featuring a 1M-token context window and a 65,536-token max output ceiling. Designed for high-speed coding and parallel agentic execution, it provides lower latency than Pro-tier equivalents. At $1.50/1M input tokens, it offers lower API costs than GPT-4o and Claude 3.5 Sonnet. | text | 1049K | 9.1 | ||
| #49 | Whisper v3 State-of-the-art speech recognition with multilingual support. | OpenAI | audio | — | 9.1 | |
| #50 | DeepSeek V4 Flash Latest DeepSeek V4 Flash Latest is an ultra-low-cost, text-only model alias featuring a 1,310,720 token context window and 131,072 max output capacity. Developed by DeepSeek, it targets high-throughput tasks at $0.03 per million input tokens, drastically undercutting GPT-4o and Claude 3.5 Sonnet on price and context capacity. | Deepseek | text | 1311K | 9.0 | |
| #51 | Qwen: Qwen3.7 Flash Qwen3.7 Flash is a high-speed multimodal vision-language model by Alibaba featuring a 1,000,000-token context window with text, image, and video input support. Optimized for visual coding, spatial reasoning, and agentic workflows, it operates at a fraction of the cost of GPT-4o and Claude 3.5 Sonnet. | Qwen | text | 1000K | 9.0 | |
| #52 | Tencent: Hy3 Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent featuring 21B active parameters and top-8 routing. Built for reasoning and agentic workloads, it offers a 262K context window and an exceptionally high 128K max output limit at a fraction of baseline API costs. | Tencent | text | 262K | 9.0 | |
| #53 | Anthropic: Claude Fable Latest Claude Fable Latest is an Anthropic pointer model that routes requests to the newest iteration in the Claude Fable family. It provides a 1,000,000-token context window and up to 128,000 output tokens for massive document processing and long-form generation. At $10/1M input and $50/1M output tokens, it operates at a premium price point for specialized long-context tasks. | Anthropic | text | 1000K | 9.0 | |
| #54 | Nex AGI: Nex-N2-Pro Nex-N2-Pro is a 397B parameter mixture-of-experts model (17B active) built on the Qwen3.5 architecture. It features a 262K context window with an exceptionally high 235K max output limit for text and vision inputs. Priced at $0.25/$1.00 per million tokens, it provides high-throughput agentic capabilities at a fraction of baseline API costs. | Nex Agi | text | 262K | 9.0 | |
| #55 | NVIDIA: Nemotron 3 Ultra (batch) NVIDIA Nemotron 3 Ultra (Batch) is a 550B total parameter (55B active) hybrid Transformer-Mamba MoE model designed for complex reasoning and agent orchestration. It supports a 512K token context window and up to 461K output tokens per call via batch processing. Priced at $0.60 per million input tokens, it offers a cost-effective alternative to GPT-4o and Claude 3.5 Sonnet for bulk, long-context text processing. | Nvidia | text | 512K | 9.0 | |
| #56 | MiniMax: MiniMax M3 MiniMax M3 is a multimodal foundation model supporting text, image, and video inputs with text output. It features a 1M-token context window alongside an exceptionally high 512K output token limit. Priced at $0.30/$1.20 per million tokens, it provides high-throughput processing for long-horizon agentic workflows and code analysis at a fraction of baseline API costs. | Minimax | text | 1049K | 9.0 | |
| #57 | Qwen: Qwen3.7 Max Qwen3.7 Max is Alibaba's flagship text model featuring a 1-million-token context window and 131K output limit. It is optimized for agentic workloads, coding, and complex reasoning at pricing well below GPT-4o and Claude 3.5 Sonnet. | Qwen | text | 1000K | 9.0 | |
| #58 | SpaceXAI: Grok Build 0.1 x-ai's Grok Build 0.1 is a fast coding model designed specifically for agentic software engineering tasks. It combines a 256K context window with an exceptionally high 230,400 max output token capacity and multimodal input capabilities. At $1.00 per 1M input and $2.00 per 1M output tokens, it offers highly aggressive pricing for code-generation pipelines. | X AI | text | 256K | 9.0 | |
| #59 | Mistral Large European AI leader with strong multilingual and coding capabilities. | Mistral AI | text | 128K | 9.0 | |
| #60 | DeepSeek: DeepSeek V4 Flash 0731 DeepSeek V4 Flash 0731 is a 284B total (13B active) sparse Mixture-of-Experts model designed for high-throughput coding, reasoning, and agentic workflows. It offers an expansive 1.3M-token context window and output generation up to 943K tokens at roughly 1/50th the input cost of GPT-4o and Claude 3.5 Sonnet. While ideal for massive document processing and rapid agent execution, its lightweight active parameter count limits high-level reasoning depth on complex edge cases. | Deepseek | text | 1311K | 8.9 | |
| #61 | Google: Gemini 3.6 Flash (batch) Google's Gemini 3.6 Flash (Batch) provides asynchronous inference with a 1,048,576 token context window and 65,536 max output tokens. Designed for high-volume background processing, it handles coding, agentic workflows, and long-document analysis at lower batch pricing. It offers extreme context capacity at a fraction of GPT-4o and Claude 3.5 Sonnet API costs. | text | 1049K | 8.9 | ||
| #62 | Google: Nano Banana 2 (Gemini 3.1 Flash Image) Developed by Google, Gemini 3.1 Flash Image (Nano Banana 2) provides integrated text and image generation with editing capabilities within a 131K token context window. Operating at $0.50 per million input tokens and $3.00 per million output tokens, it delivers fast multimodal visual synthesis. It serves as a cost-effective solution for workflows requiring combined visual creation and prompt processing. | image | 131K | 8.9 | ||
| #63 | Qwen: Qwen3.8 Flash Qwen3.8 Flash is a multimodal reasoning model from Alibaba supporting text, image, and video inputs. It features a 1,000,000 token context window and a 131,072 max output token limit at $0.15 per million input tokens. It is designed for codebase analysis, high-throughput agentic tasks, and video processing. | Qwen | text | 1000K | 8.8 | |
| #64 | Z.ai: GLM 5.3 GLM 5.3 is a reasoning model developed by Z.ai built for software engineering and long-horizon agentic workloads. It features a 1M-token context window alongside an exceptionally high 131,072 max output token ceiling. Priced at $1.40/1M input and $4.40/1M output, it offers a cost-effective alternative to GPT-4o and Claude 3.5 Sonnet for long-context text processing. | Z AI | text | 1049K | 8.8 | |
| #65 | ByteDance Seed: Seed 2.1 Turbo ByteDance Seed 2.1 Turbo is a multimodal AI model designed for long-horizon agent execution and coding tasks. It features a 262K context window alongside a massive 235K token output limit at a fraction of the cost of GPT-4o. The model accepts text, image, and video inputs to assist with end-to-end software engineering workflows. | Bytedance Seed | text | 262K | 8.8 | |
| #66 | Thinking Machines: Inkling Small Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines with 12B active out of 276B total parameters. It features a 1M token context window and supports text, image, and audio inputs at a fraction of standard API costs. | Thinkingmachines | text | 1049K | 8.8 | |
| #67 | Poolside: Laguna S 2.1 Laguna S 2.1 is an agentic coding model developed by Poolside using a 118B total and 8B active parameter MoE architecture. It features a 1-million-token context window and 131K max output limit, optimized for terminal-based software engineering tasks. | Poolside | text | 1049K | 8.8 | |
| #68 | Meituan: LongCat 2.0 LongCat 2.0 is a sparse Mixture-of-Experts language model from Meituan utilizing 48B active parameters out of 1.6T total. Designed for repository-level software development and long-horizon agentic workflows, it features a 1,048,576 token context window and a 262,144 token max output. At $0.30/$1.20 per million tokens, it offers low-cost long-context processing relative to major benchmarks. | Meituan | text | 1049K | 8.8 | |
| #69 | Kwaipilot: KAT-Coder-Pro V2.5 Developed by Kwaipilot, KAT-Coder-Pro V2.5 is an agentic coding model designed for autonomous issue resolution and software workflow execution. It features a 256K context window and an 80K output token limit at a fraction of the cost of GPT-4o and Claude 3.5 Sonnet. | Kwaipilot | text | 256K | 8.8 | |
| #70 | Poolside: Laguna XS 2.1 Poolside's Laguna XS 2.1 is a 33B-parameter specialized coding model designed for autonomous software agent workflows. Featuring a 262K token context window and a 32K max output limit, it delivers extreme cost efficiency for high-volume code generation compared to GPT-4o and Claude 3.5 Sonnet. | Poolside | text | 262K | 8.8 | |
| #71 | MoonshotAI: Kimi K2.7 Code MoonshotAI Kimi K2.7 Code is a long-context, coding-optimized Mixture-of-Experts (MoE) model with native vision support. It features a 262K token context window and an unusually high 235K token output limit designed for multi-file codebase generation. It operates at roughly a quarter of the API cost of GPT-4o and Claude 3.5 Sonnet. | Moonshotai | text | 262K | 8.8 | |
| #72 | MoonshotAI: Kimi K2.7 Code (batch) MoonshotAI Kimi K2.7 Code (batch) is a multimodal coding model designed for asynchronous processing of large codebases. It features a 262K token context window and supports outputs up to 235K tokens at discounted batch rates. It targets automated repository refactoring, multi-file code generation, and visual architecture analysis. | Moonshotai | text | 262K | 8.8 | |
| #73 | Sora Text-to-video generation with realistic physics and long-form coherence. | OpenAI | video | — | 8.8 | |
| #74 | MiniMax: MiniMax M3 (batch) MiniMax M3 (batch) is a multimodal model from MiniMax supporting text, image, and video inputs with text output. It features a 524K-token context window alongside a massive 471,859 maximum output token limit. Offered at $0.30/1M input tokens in batch mode, it targets high-volume, cost-sensitive long-context workflows. | Minimax | text | 524K | 8.7 | |
| #75 | StepFun: Step 3.7 Flash Step 3.7 Flash is a 196B-parameter Mixture-of-Experts multimodal model from StepFun activating roughly 11B parameters per token. It features native image and video understanding across a 262,144-token context window. At $0.20 per million input tokens, it delivers high-throughput multimodal processing at a fraction of GPT-4o's cost. | Stepfun | text | 262K | 8.7 | |
| #76 | Z.ai: GLM Latest Z.ai: GLM Latest is a text-to-text model developed by ~z-ai, offering a 1,048,576 token context window and a 131,072 token max output, making it suitable for extensive long-context processing. | Z AI | text | 1049K | 8.6 | |
| #77 | ByteDance Seed: Seed-2.0-Code ByteDance Seed-2.0-Code is a specialized model designed for agentic programming workflows, frontend development, and multilingual coding. It features a 262K context window, an unusually large 131K max output token limit, and multimodal input capabilities. Priced at $0.50 input and $3.00 output per million tokens, it provides a low-cost alternative to Claude 3.5 Sonnet and GPT-4o for coding tools. | Bytedance Seed | text | 262K | 8.6 | |
| #78 | Sakana: Fugu Ultra Sakana AI's Fugu Ultra is a multi-agent orchestration system that dynamically routes complex queries across specialized sub-models. It provides a 1M-token context window alongside an exceptionally high 128K output token limit. The architecture targets complex reasoning and long-form task execution. | Sakana | text | 1000K | 8.6 | |
| #79 | Meta: Muse Spark 1.2 Contributor Meta: Muse Spark 1.2 Contributor is a reasoning model from Meta, positioned as a cost-effective option for developers. It offers an exceptionally large context window and broad multimodal input capabilities at a significantly lower price point compared to current market leaders. | Meta | text | 1049K | 8.5 | |
| #80 | Qwen: Qwen3.8 27B Qwen3.8 27B is an open-weight, dense vision-language model from Qwen that processes text, image, and video inputs to generate text. It features a 1,000,000-token context window and competitive API pricing, positioning it for long-context multimodal tasks. | Qwen | text | 1000K | 8.5 | |
| #81 | Google: Gemini 3.5 Flash Lite Developed by Google, Gemini 3.5 Flash Lite is a high-efficiency model designed for focused task execution in multi-agent workflows. It pairs a 1-million-token context window with low pricing ($0.30/1M input) and a 65,536 max output limit. While less capable in complex reasoning than GPT-4o or Claude 3.5 Sonnet, it excels at high-throughput, low-cost operations. | text | 1049K | 8.5 | ||
| #82 | Upstage: Solar Pro 4 Solar Pro 4 is a text LLM from Upstage featuring a 524K context window and a 131K output token limit. At $0.03 per million input tokens, it is significantly cheaper than GPT-4o and Claude 3.5 Sonnet. It targets high-volume document processing, long-horizon agentic workflows, and enterprise text tasks. | Upstage | text | 524K | 8.4 | |
| #83 | Meta: Muse Glimmer 30B Meta's Muse Glimmer 30B is an open-weight, 30-billion parameter multimodal model distilled for long-horizon autonomous agents. It features an exceptionally high output limit of 117,964 tokens and low API pricing at $0.35/$1.50 per million tokens. While ideal for cost-sensitive edge deployments, its smaller parameter count limits peak reasoning depth compared to GPT-4o or Claude 3.5 Sonnet. | Meta | text | 131K | 8.4 | |
| #84 | Ling-3.0-flash Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) text model developed by inclusionai, activating 5.1B parameters per token. It features a 262K token context window, a 32K token max output, and extremely low API costs ($0.02/$0.06 per 1M tokens). The architecture targets high-throughput agentic workflows and budget-focused enterprise text tasks. | Inclusionai | text | 262K | 8.4 | |
| #85 | Kwaipilot: KAT-Coder-Air V2.5 Kwaipilot KAT-Coder-Air V2.5 is an agentic coding model built for autonomous software issue resolution and workflow automation. It provides a 256K token context window with an exceptionally high 80K max output limit at a fraction of the price of Claude 3.5 Sonnet and GPT-4o ($0.15/1M input, $0.60/1M output). It targets automated long-context code generation and codebase refactoring. | Kwaipilot | text | 256K | 8.4 | |
| #86 | Tencent: Hy-MT2-1.8B Tencent's Hy-MT2-1.8B is a compact 1.8B-parameter model specialized for machine translation. It supports 33 language pairs and 5 Chinese dialect/minority language pairs, offering advanced workflows like glossary-based and style-guided translation for specific enterprise needs. | Tencent | text | 8K | 8.3 | |
| #87 | NVIDIA: Nemotron 3.5 Lightning NVIDIA Nemotron 3.5 Lightning is an open Mixture-of-Experts (MoE) model featuring 3B active out of 30B total parameters, optimized for high-throughput agentic execution. It delivers a massive 262K context window and an unprecedented 131K max output limit at a tiny fraction of frontier model pricing. It trades peak reasoning depth for extreme inference speed, high token volume, and low operational cost. | Nvidia | text | 262K | 8.3 | |
| #88 | Sakana: Sakana Namazu Sakana Namazu is a Japanese-specialized reasoning model developed by Sakana AI using Moonshot AI's Kimi K2.6 architecture. It combines a 262K token context window with fine-tuning targeted specifically at Japanese instruction following and business domain tasks. | Sakana | text | 262K | 8.3 | |
| #89 | Google: Gemini 3.5 Flash Lite (batch) Google Gemini 3.5 Flash Lite (batch) is a lightweight, multimodal model engineered for high-volume asynchronous tasks and subagent workflows. It offers a 1-million-token context window and a 65,536-token output limit at a fraction of the cost of standard frontier models. It sacrifices real-time response latency and deep reasoning capabilities to achieve high cost-efficiency. | text | 1049K | 8.3 | ||
| #90 | Auto Router (Beta) Auto Router (Beta) is a task-aware model router developed by OpenRouter that dynamically directs prompts to top-performing models based on real-time task spend and usage data. It features a 2,000,000 token context window and handles multimodal inputs, simplifying multi-model orchestration through a single API endpoint. | Openrouter | text | 2000K | 8.3 | |
| #91 | Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) Google Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is a high-speed, cost-efficient model optimized for developer image generation and editing pipelines. It combines a 65K context window with low-latency multimodal processing at a fraction of standard API costs. It serves high-volume workflows where speed and cost supersede peak visual fidelity. | image | 66K | 8.3 | ||
| #92 | OpenRouter: Fusion OpenRouter Fusion is a multi-model orchestration system that executes prompts across a panel of expert LLMs in parallel with active web retrieval. It synthesizes consensus answers from multiple models across a 1,000,000 token context window. The system is designed for high-accuracy research and multi-perspective deliberation rather than raw speed. | Openrouter | text | 1000K | 8.3 | |
| #93 | Nex AGI: Nex-N2-Mini Nex-N2-Mini is an open-source, agentic Mixture-of-Experts (MoE) model developed by Nex AGI for coding, tool interaction, and multimodal processing. It offers a 262,144-token context window with a massive max output capacity at ultra-low inference costs. | Nex Agi | text | 262K | 8.2 | |
| #94 | Tencent: Hy-MT2-30B-A3B Tencent's Hy-MT2-30B-A3B is a specialized translation model supporting 33 language pairs alongside 5 Chinese regional dialects. It incorporates workflows for glossary control, contextual translation, and delimiter-based formatting at a fraction of general-purpose API costs. It is designed strictly for localization pipelines rather than general reasoning. | Tencent | text | 8K | 7.8 | |
| #95 | AionLabs: Aion-3.0 Developed by AionLabs on the GLM base architecture, Aion-3.0 is a multi-model system optimized for roleplay and narrative generation. It pairs a 131K context window with an unusually high 32K max output token limit at $3.00/$6.00 per million tokens. Unlike general-purpose benchmarks like GPT-4o and Claude 3.5 Sonnet, it focuses primarily on collaborative creative writing rather than technical reasoning or coding. | Aion Labs | text | 131K | 7.8 | |
| #96 | AionLabs: Aion-3.0-Mini Developed by AionLabs, Aion-3.0-Mini is a narrative and roleplay model built on DeepSeek architectures. It features a 131K context window and an unusually high 32K max output token limit at significantly lower costs than GPT-4o. However, it is specialized for creative narrative tasks rather than general-purpose software development or enterprise reasoning. | Aion Labs | text | 131K | 7.6 | |
| #97 | Tencent: Hy-MT2-7B Tencent: Hy-MT2-7B is a 7B-parameter model developed by Tencent, specialized in machine translation across 33 language pairs and specific Chinese dialects, offering structured and contextual translation workflows. | Tencent | text | 8K | 7.5 |
No models match your search or filter.