Last updated October 5, 2026.
Our research team spent the first half of 2026 evaluating 42 AI models using data compiled from the Artificial Analysis Intelligence Index, LM Council’s independently run benchmark leaderboard, vals.ai’s standardized SWE-bench harness, Humanity’s Last Exam, and official developer documentation published by each model’s creator. Our full sources list is available at the end of the article. We scored each model on seven weighted criteria:
- AA Intelligence Index (25%): Artificial Analysis composite measure of model capability, aggregating performance across coding, reasoning, math, and knowledge tasks into a single 0–100 scale.
- Agentic Performance (20%): AA-Briefcase Elo, which evaluates models on realistic tasks involving research, analysis, spreadsheets, and other professional deliverables. The benchmark uses rubric-based grading plus pairwise assessments of analytical quality and presentation.
- Coding and Terminal Performance (20%): Terminal-Bench 4.0 score, measuring a model’s ability to complete complex tasks in a real terminal environment across software, machine learning, science, operations, security, hardware, and media.
- Context Window (10%): Maximum number of tokens the model can process in a single inference call. Relevant for long documents, large codebases, and extended agent sessions.
- Output Speed (10%): Median tokens per second across providers, measured by Artificial Analysis under standardized conditions.
- Cost Per Task (10%): Artificial Analysis’s cost per Intelligence Index task, which accounts for input, output, reasoning, and cache-related costs and provides a more realistic value than list-price API rates alone.
- Multimodal Capabilities (5%): Breadth of supported input and output modalities, including text, images, audio, and video.
Models without publicly available data for a criterion receive a conservative below-average score for that factor.
The Top AI Models of 2026
| # | Model | Developer | AA Intel. | Agentic Performance | Coding and Terminal Performance | Context Window | Output Speed (tok/s) | Cost Per Task | Modalities |
| 1 | Claude Fable 5.1 | Anthropic | 53 | 1,662 | 52.0% | 1M | 67 | $7.63 | Text, vision |
| 2 | GPT-6 Astra | OpenAI | 53 | 1,562 | 59.1% | 1M | 53 | $3.26 | Text, vision |
| 3 | Claude Opus 5 | Anthropic | 51 | 1,645 | 55.0% | 1M | 50 | $5.86 | Text, vision |
| 4 | Claude Fable 5 | Anthropic | 50 | 1,530 | 44.5% | 1M | 63 | $4.36 | Text, vision |
| 5 | Muse Spark 1.3 | Meta | 48 | 1,589 | 33.0% | 1M | 206 | $1.60 | Text, vision |
| 6 | GPT-5.6 Sol | OpenAI | 47 | 1,528 | 39.9% | 1M | 57 | $1.99 | Text, vision, audio, images |
| 7 | Qwen3.8 Max | Alibaba | 45 | 1,622 | 39.0% | 984K | 40 | $5.41 | Text, vision |
| 8 | GLM-5.3 | Z AI | 45 | 1,511 | 42.0% | 1M | 66 | $2.01 | Text, vision |
| 9 | Grok 4.6 | SpaceXAI | 44 | 1,545 | 21.0% | 500K | 54 | $1.86 | Text, vision |
| 10 | Kimi K3 | Moonshot AI | 44 | 1,488 | 13.0% | 1.05M | 37 | $2.00 | Text, vision |
| 11 | GLM-5.3 Flash | Z AI | 42 | 1,655 | 33% | 1M | 117 | $0.25 | Text, vision |
| 12 | GPT-5.6 Terra | OpenAI | 42 | 1,416 | 21.5% | 1M | 108 | $1.40 | Text, vision, audio, images |
| 13 | Gemini 3.8 Flash | 41 | 1,426 | 19.1% | 1M | 341 | $1.24 | Text, vision, audio, video | |
| 14 | Qwen3.8 2.4T A95B | Alibaba | 40 | 1,483 | 30.3% | 984K | 40 | $2.16 | Text, vision |
| 15 | DeepSeek V4.1 Flash | DeepSeek | 40 | 1,424 | 27.0% | 1M | 216 | $0.27 | Text, vision |
All benchmark data sourced from Artificial Analysis, vals.ai, benchlm.ai, and official developer documentation as of September 2026.
Here is what each model on this list is best suited for, who it’s designed for, and what separates it from others in our evaluation.
| Model | Best For | What Sets It Apart |
| Claude Fable 5.1 | Research teams, enterprises, and developers handling complex work and software projects | Tied for the highest Intelligence Index score and leads the AA-Briefcase agentic knowledge-work benchmark, but the most expensive option on the list. |
| GPT-6 Astra | Teams that want a general-purpose model for coding, research, and professional workflows | Tied for the highest Intelligence Index score, with broad capabilities spanning software engineering, browsing, computer use, science, at a more affordable price point than Claude’s Fable 5.1. |
| Claude Opus 5 | Engineering, research, and professional teams working on difficult multi-step problems | Combines a 51 Intelligence Index score with the third-highest AA-Briefcase score, giving it particularly strong performance on complex knowledge work and reasoning tasks. |
| Claude Fable 5 | Organizations that need a high-capability model for demanding coding tasks | Remains a frontier model despite being superseded by Fable 5.1, with a 50 Intelligence Index score and strong performance across all ranking factors. |
| Muse Spark 1.3 | High-volume applications that need strong model performance without sacrificing response speed | Combines a 48 Intelligence Index score with very high output speed, making it one of the more interesting options for applications where both capability and throughput matter. |
| GPT-5.6 Sol | General-purpose business applications that need multimodal capabilities | Offers a broad combination of text, vision, audio, and image capabilities while retaining a lower cost profile than the newest frontier models. |
| Qwen3.8 Max | Research, coding, and agentic applications | Scores 1,622 on AA-Briefcase, placing it among the strongest models, while Qwen’s broader model family also provides open-weight deployment options. |
| GLM-5.3 | Developers and enterprises looking for a strong open-weight alternative to proprietary frontier models | Part of the leading open-weight group tracked by Artificial Analysis, combining high capability with greater deployment flexibility. |
| Grok 4.6 | Teams working with current information and research that benefit from web-connected AI | Combines a 44 Intelligence Index score with a large 500K-token context window and access to xAI’s web-connected ecosystem, making it suited to information-heavy workflows. |
| Kimi K3 | Developers building long-context applications and AI agents | Offers a 1.05M-token context window and a 44 Intelligence Index score, giving it a combination of long-context capacity and frontier-level general performance. |
| GLM-5.3 Flash | High-volume AI applications where throughput and cost are major constraints | Combines a 42 Intelligence Index score with very high output speed and a low cost per task, making it one of the strongest cost-and-throughput-oriented models in the current dataset. |
| GPT-5.6 Terra | Teams building fast, general-purpose AI applications that need multimodal capabilities | Provides a lower-cost, high-throughput alternative within OpenAI’s GPT-5.6 family, with support for text, vision, audio, and image inputs and outputs. |
| Gemini 3.8 Flash | High-volume applications involving text, images, audio, and video | Its combination of multimodal support and very high output speed makes it particularly suitable for interactive applications and workloads that process large amounts of content. |
| Qwen3.8 2.4T A95B | Organizations that want a high-capability open-weight model | Reaches a 40 Intelligence Index score while remaining an open-weight model, giving developers greater control over deployment, customization, and infrastructure. |
| DeepSeek V4.1 Flash | Cost-sensitive AI applications that still require strong reasoning and agentic performance | High output speed, a low cost profile, and a 1M-token context window make it particularly well suited to high-volume workloads. |
In-Depth Model Reviews
1. Claude Fable 5.1, for Peak Reasoning and Complex Knowledge Work
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Claude Fable 5.1 | High scores across the board; the most capable AI model available | The most expensive model evaluated, making it harder to justify for high-volume workloads | Complex research, strategic analysis, software engineering, and long-running professional workflows |
Claude Fable 5.1 is Anthropic’s latest flagship model and one of two models tied for the highest score on the Artificial Analysis Intelligence Index, with a score of 53. Out of all models evaluated, it has the strongest performance on complex, multi-step knowledge work, and also performs strongly on software engineering and terminal-based tasks.
On Terminal-Bench 4.0, it achieves a 52.0% score under the configuration used for the current comparison, while GPT-6 Astra scores higher at 59.1%. Astra has an advantage on the current terminal-coding benchmark, while Fable 5.1 leads the pair on the broader AA-Briefcase evaluation.
The tradeoff is cost. Artificial Analysis lists Fable 5.1 at approximately $7.63 per Intelligence Index task, putting it toward the expensive end of the current frontier-model market. That makes it difficult to justify for simple summarization, routine content generation, or other high-volume tasks where a faster, cheaper model can produce an adequate result. For complex research synthesis, difficult software projects, strategic analysis, and other workflows where errors or additional human review are expensive, however, its combination of general intelligence and strong agentic performance makes the premium easier to justify.
2. GPT-6 Astra, for Fast, High-Level Reasoning and Coding
| Model | Biggest Pro | Biggest Con | Best Use Case |
| GPT-6 Astra | Combines top-tier intelligence with leading Terminal-Bench performance | Performance varies noticeably by reasoning effort | Complex coding, agentic workflows, research, and general-purpose reasoning |
GPT-6 Astra is OpenAI’s flagship model and ties Claude Fable 5.1 for the highest score on Artificial Analysis’s Intelligence Index, at 53. Rather than dominating only one category, Astra performs particularly well across agentic work, coding, and general reasoning. Artificial Analysis reports that Astra scores higher than Fable 5.1 on Terminal-Bench 4.0, which evaluates 66 difficult terminal tasks spanning software engineering, machine learning, science, operations, security, hardware, and media.
Astra is also particularly notable for its efficiency across different reasoning levels. Artificial Analysis reports that all five reasoning efforts of GPT-6 Astra sit on the Intelligence Index cost-efficiency frontier, giving users more flexibility to trade speed and cost against deeper reasoning. Its main limitation is therefore less about a lack of capability than about choosing the appropriate configuration for a given task. For difficult coding and agentic work, higher reasoning settings can provide substantially more capability, while simpler tasks may not justify the additional computation.
3. Claude Opus 5, for Deep Reasoning and High-Stakes Knowledge Work
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Claude Opus 5 | Strong performance across complex reasoning and knowledge-work tasks | More expensive and less efficient than lighter models | Research, analysis, software engineering, and complex professional work |
Claude Opus 5 is Anthropic’s second-highest-scoring model on the Artificial Analysis Intelligence Index, with a score of 51. That puts it just behind the 53-point scores of Claude Fable 5.1 and GPT-6 Astra and ahead of Claude Fable 5 at 50. Artificial Analysis’s v4.3 evaluation places particular emphasis on agentic, coding, general, and scientific reasoning, rather than relying primarily on traditional academic benchmarks.
Opus 5 is therefore positioned as a model for tasks where accuracy and depth are more important than raw throughput. It benefits from Anthropic’s strong performance on knowledge-work and scientific evaluations, although Fable 5.1 now sits above it on the overall Intelligence Index. For organizations that need a powerful model for research, complex analysis, coding, or extended reasoning but do not necessarily need the absolute highest-performing model, Opus 5 is a capable alternative within Anthropic’s lineup.
4. Claude Fable 5, for Advanced Reasoning at a Slightly Lower Tier
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Claude Fable 5 | Strong frontier-level reasoning with a lower capability tier than Fable 5.1 | Outperformed by newer flagship models | Research, writing, coding, and complex analysis |
Claude Fable 5 is among the highest-performing general-purpose models in our comparison. Its Intelligence Index score is lower than Claude Fable 5.1 and Claude Opus 5, but it remains firmly within the frontier tier. While the newer 5.1 version outperforms Fable 5, the latter can still provide very strong reasoning at a more enticing price point.
For demanding writing, research, coding, and analytical work, Fable 5 is considerably more capable than the average general-purpose model. However, with Fable 5.1 now in the mix, users choosing between the two should consider the additional cost and performance requirements of their particular workload.
5. Muse Spark 1.3, for High-Speed Multimodal Work
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Muse Spark 1.3 | Strong intelligence combined with impressively high output speed | Not as strong as the top frontier models on the most demanding tasks | High-volume analysis, multimodal workflows, and fast content generation |
Muse Spark 1.3 has an Artificial Analysis Intelligence Index with a score of 48, placing it below the 50-plus scores of the leading Claude and GPT models but ahead of many other current frontier and open-weight systems. The September v4.3 index evaluates models across agentic work, coding, general knowledge, and scientific reasoning, making the 48-point result a broad measure of this model’s overall capability.
Its biggest practical advantage is speed. Muse Spark 1.3 is designed for workloads where users may need to generate or process large amounts of material quickly, making it ideal for high-volume applications. For the hardest research or software-engineering tasks, however, the models above it on the Intelligence Index provide a higher ceiling.
6. GPT-5.6 Sol, for Efficient High-End General-Purpose Performance
| Model | Biggest Pro | Biggest Con | Best Use Case |
| GPT-5.6 Sol | Strong frontier performance at a substantially lower cost per task than the top models | Does not match the highest Intelligence Index scores | Coding, research, analysis, and general-purpose professional work |
GPT-5.6 Sol remains one of the strongest models in the comparison despite being on the older side of things. Its Intelligence Index score puts it comfortably within the upper tier, and it also scores 39.9% on Terminal-Bench 4.0, making it particularly competitive for terminal-based software engineering and other tasks that require models to interact with development environments.
Sol’s other advantage is cost efficiency. The maximum configuration is estimated at $1.99 per Intelligence Index task, substantially below the $5-plus costs associated with several of the highest-capability models. OpenAI also offers lower reasoning configurations, allowing users to trade some performance for lower cost and latency. That’s a big green flag for organizations that need strong reasoning and coding performance at scale rather than the absolute highest score available on the Intelligence Index.
7. Qwen3.8 Max, for Multimodal Reasoning and Open-Model Ecosystems
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Qwen3.8 Max | Strong reasoning with multimodal capabilities and a large context window | Relatively high cost for its Intelligence Index score | Engineering, research, multimodal analysis, and long-context work |
Qwen3.8 Max supports image input and offers a context window of approximately 984,000 tokens, putting it close to the 1-million-token class used by the leading models. Artificial Analysis also identifies Qwen3.8 Max as particularly capable in engineering and finance-related professional evaluations.
The main tradeoff is cost. Artificial Analysis estimates approximately $5.41 per Intelligence Index task for Qwen3.8 Max, which is substantially more than GPT-5.6 Sol, GLM-5.3, or the Flash models lower in this ranking. Its value is therefore strongest when its particular combination of multimodal reasoning, long context, and engineering-oriented performance are important. For routine high-volume inference, less expensive Qwen variants are likely more economical.
8. GLM-5.3, for Strong Open-Weight Performance
| Model | Biggest Pro | Biggest Con | Best Use Case |
| GLM-5.3 | Combines strong frontier-level performance with open weights | Not as capable as the newest proprietary flagships | Self-hosted AI, coding, engineering, and enterprise applications |
GLM-5.3 scores 45 on Artificial Analysis’s Intelligence Index, putting it level with Qwen3.8 Max and ahead of several newer models with lower scores. It has a 1-million-token context window and scores 42% on Terminal-Bench 4.0, making it particularly competitive for software engineering and other terminal-based workflows. Artificial Analysis also reports strong results for GLM-5.3 in its healthcare and engineering professional evaluations.
The model’s open-weight availability is an added bonus. Unlike proprietary systems such as GPT-5.6 Sol or Claude Fable 5.1, GLM-5.3 can be deployed in environments where organizations want greater control over infrastructure and model behavior. At $2.01 per Intelligence Index task, it is not exactly a low-cost alternative, but its appeal comes from combining relatively high capability with the flexibility of an open-weight model. For organizations with the infrastructure to operate it, it’s worth exploring for technical and enterprise workloads.
9. Grok 4.6, for Fast Reasoning and Multimodal Applications
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Grok 4.6 | Strong performance with fast output and multimodal capabilities | Smaller context window than most models near the top of the ranking | General reasoning, research, real-time applications, and multimodal tasks |
Grok 4.6 supports image input and has a 500,000-token context window, which is substantial but notably smaller than the 1-million-token windows offered by many of the models above it. Artificial Analysis also reports multiple reasoning configurations for Grok 4.6, with Intelligence Index scores ranging from 35 to 44.
Grok 4.6 is particularly notable for speed. Its high configuration is listed at about 54 tokens per second, at a cost of $1.86 per Intelligence Index task. That combination makes it potentially appealing for users who need strong reasoning without the latency or expense associated with the highest-end reasoning configurations. Its main limitation is that the 500K context window gives it less room for extremely large documents or long-running context-heavy workflows than the 1-million-token models dominating the upper portion of this comparison.
10. Kimi K3, for Long-Context Reasoning and Open-Weight Flexibility
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Kimi K3 | Strong reasoning with a 1-million-token-plus context window | High-end reasoning is relatively slow and expensive | Long-context research, coding, analysis, and self-hosted AI |
Kimi K3’s 1.05-million-token context window is among the largest in this comparison, giving it enough capacity for very large document collections, codebases, and extended research workflows. Artificial Analysis also reports strong results on scientific reasoning and long-context evaluations, including a 59% score on SciCode and 89% on AA-LCR v1.1.
Kimi K3 is also available with open weights, which distinguishes it from most of the proprietary models at this level. Of course, there’s a catch. At $2.00 per Intelligence Index task at maximum reasoning, with output speeds around 37 tokens per second, it’s the slowest model we evaluated. Kimi K3 is better suited to substantial reasoning tasks rather than to use cases where speed matters.
11. GLM-5.3 Flash, for Low-Cost High-Volume AI
| Model | Biggest Pro | Biggest Con | Best Use Case |
| GLM-5.3 Flash | Exceptional cost efficiency with strong reasoning performance | Lower peak capability than the larger GLM-5.3 | High-volume inference, coding, automation, and self-hosted applications |
GLM-5.3 Flash matches GPT-5.6 Terra on the Artificial Analysis Intelligence Index while using considerably fewer resources. It has a 1-million-token context window and performs well on professional knowledge-work evaluations, including a 1,655 AA-briefcase score.
At just $0.25 per Intelligence Index task, it’s the most budget-friendly model on our list. It is also an open-weight model, giving developers the option to deploy it within their own infrastructure. The combination of a 42 Intelligence Index score, 117-token-per-second output speed, and very low task cost makes Flash a great all-around option.
12. GPT-5.6 Terra, for Fast, Cost-Efficient Professional Work
| Model | Biggest Pro | Biggest Con | Best Use Case |
| GPT-5.6 Terra | Strong performance at a relatively low cost per task | Lower peak intelligence than GPT-5.6 Sol | Business analysis, coding, automation, and high-volume professional workloads |
GPT-5.6 Terra reaches an Intelligence Index score of 42 at maximum reasoning, putting it below GPT-5.6 Sol but alongside GLM-5.3 Flash. Its 1-million-token context window gives it ample room for long documents and complex workflows, while its Terminal-Bench 4.0 result of 21.5% shows strong performance on terminal-based coding and computing tasks.
Terra’s strongest practical advantage is its cost and speed profile. Artificial Analysis estimates about $1.40 per Intelligence Index task at maximum reasoning, while lower reasoning settings substantially reduce both cost and latency. The maximum configuration is relatively slow because of the amount of reasoning it performs, so organizations can choose lower settings for less demanding work. That makes Terra a useful middle ground for those who need more capability than a lightweight model but do not require the highest-end reasoning available.
13. Gemini 3.8 Flash, for Speed and Multimodal Processing
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Gemini 3.8 Flash | Extremely fast output with broad multimodal support | Lower overall reasoning score than the leading flagship models | Real-time applications, multimodal analysis, and high-throughput AI |
Gemini 3.8 Flash may not have the highest Intelligence Index score, but its capabilities extend beyond basic text input. The model accepts text, images, speech, and video, while maintaining a 1-million-token context window. Artificial Analysis reports strong results on scientific and long-context evaluations, including 57% on SciCode and 81% on AA-LCR v1.1.
Speed is where Gemini 3.8 Flash really excels. At 341 tokens per second at the high reasoning setting, it’s substantially faster than most models in this comparison. Its cost of $1.24 per Intelligence Index task is also relatively moderate, although higher than the most economical open-weight Flash models. Overall, Gemini 3.8 Flash is a good choice for applications where latency, multimodal input, and throughput matter as much as reasoning quality.
14. Qwen3.8 2.4T A95B, for Large-Scale Open-Weight Reasoning
| Model | Biggest Pro | Biggest Con | Best Use Case |
| Qwen3.8 2.4T A95B | Large-scale open-weight model with strong general reasoning | More expensive to run than smaller open-weight alternatives | Enterprise AI, coding, research, and self-hosted deployments |
Qwen3.8 2.4T A95B has approximately a 984,000-token context window and supports image input, giving it the capacity for large documents and multimodal analysis. Its overall Intelligence Index score reflects reliable performance across the broader set of 10 evaluations.
The model’s open-weight status is a major part of its appeal. Developers and organizations can deploy and adapt the model without relying exclusively on a proprietary hosted API. That does come at a cost of $2.16 per Intelligence Index task, which is considerably more than GLM-5.3 Flash or DeepSeek V4.1 Flash. Qwen3.8 2.4T A95B therefore makes the most sense when organizations value (and are willing to pay for) the ability to control deployment.
15. DeepSeek V4.1 Flash, for Low-Cost Open-Weight Reasoning
| Model | Biggest Pro | Biggest Con | Best Use Case |
| DeepSeek V4.1 Flash | Very low cost with strong reasoning and high output speed | Lower overall intelligence than the leading frontier models | High-volume inference, coding, research, and self-hosted AI |
DeepSeek V4.1 Flash scores 40 on the Artificial Analysis Intelligence Index at maximum reasoning, placing it alongside Qwen3.8 2.4T A95B. It has a 1-million-token context window, accepts both text and image input, and is available with open weights. This model offers a great combination of cost and speed. At $0.27 per Intelligence Index task and roughly 216 tokens per second of output, it is substantially cheaper and faster than many models with similar overall scores.
It does require more output and reasoning tokens than some alternatives, but its low per-task price makes that computational overhead relatively inexpensive. For developers looking for an open-weight model that can handle substantial reasoning workloads without the cost of a premium proprietary model, V4.1 Flash is designed for exactly that role.
Top AI Models by Use Case
Top AI Models for Coding and Software Engineering
| Rank | Model | Terminal-Bench 4.0 | Key Coding Strength |
| 1 | GPT-6 Astra | 59% | Complex terminal-based software engineering and agentic coding |
| 2 | Claude Fable 5.1 | 52% | Strong terminal coding; leading overall reasoning and agentic knowledge-work performance |
| 3 | GLM-5.3 | 42% | Strong open-weight coding model with a 1M-token context window |
| 4 | GPT-5.6 Sol | 40% | Strong general-purpose coding with competitive terminal and tool-use performance |
| 5 | GLM-5.3 Flash | 33% | Much lower-cost alternative for high-volume coding and automation workloads |
Top AI Models for Scientific Reasoning and Research
| Rank | Model | Intelligence Index | Key Reasoning Strength |
| 1 | Claude Fable 5.1 | 53 | Strongest overall II score; 63% SciCode and 59% Humanity’s Last Exam |
| 2 | GPT-6 Astra | 53 | Broad reasoning performance; 56% SciCode and 55% Humanity’s Last Exam |
| 3 | Claude Opus 5 | 51 | High-end reasoning across scientific, knowledge, and professional tasks |
| 4 | Claude Fable 5 | 50 | Strong scientific reasoning; 61% SciCode and 55% Humanity’s Last Exam |
| 5 | GPT-5.6 Sol | 47 | Strong scientific and general reasoning at a lower cost than the newest frontier models |
Top AI Models for Budget-Conscious Teams
| Rank | Model | Cost per Int. Index Task | Performance Justification |
| 1 | GLM-5.3 Flash | $0.25 | 42 Intelligence Index score, 60% AutomationBench-AA, and 33% Terminal-Bench 4.0; open-weight |
| 2 | DeepSeek V4.1 Flash | $0.27 | 40 Intelligence Index score with very high output speed and a 1M-token context window |
| 3 | GPT-6 Astra (low) | $0.82 | 46 Intelligence Index score at low reasoning; much lower task cost than its high-reasoning configurations |
| 4 | Grok 4.6 (low) | $0.48 | 35 Intelligence Index score at low reasoning, with very low cost per task and high output speed |
| 5 | GPT-5.6 Terra (low) | $0.14 | Lower-capability configuration, but exceptionally low estimated cost per Intelligence Index task |
- Artificial Analysis LLM Leaderboard, September 2026 — https://artificialanalysis.ai/leaderboards/models
- Artificial Analysis Intelligence Index v4.3, September 2026 — https://artificialanalysis.ai/articles/artificial-analysis-intelligence-index-v4-3
- Artificial Analysis — Humanity’s Last Exam Benchmark Leaderboard, September 2026 — https://artificialanalysis.ai/evaluations/humanitys-last-exam
- Artificial Analysis — Language Model Releases, September 2026 — https://artificialanalysis.ai/models/releases
- LM Council — AI Model Benchmarks, September 2026 — https://lmcouncil.ai/benchmarks
- Vals.ai’s — Standardized SWE-bench harness, September 2026 — https://www.vals.ai/benchmarks
- Anthropic System Card — Claude Fable 5.1 and Mythos 5.1, September 2026 — https://www.anthropic.com/claude-fable-and-mythos-5-1
- Anthropic System Card — Claude Model System Cards, September 2026 — https://www.anthropic.com/system-cards
- OpenAI — GPT-6 Astra Model Documentation, 2026 — https://openai.com/index/gpt-6-astra
- OpenAI — GPT-5.6 Sol Model Documentation, 2026 — https://developers.openai.com/api/docs/models/gpt-5.6-sol
- OpenAI — GPT-5.6 Terra Model Documentation, 2026 — https://developers.openai.com/api/docs/models/gpt-5.6-terra
- OpenAI — GPT-5.6 System Card, September 2026 — https://deploymentsafety.openai.com/gpt-5-6
- Google — Gemini 3.8 Flash API Documentation, September 2026 — https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash
- DeepSeek — “Introducing DeepSeek-V4.1-Flash,” September 2026 https://deepseek.com/en/news/deepseek-v4-1-flash/
- SpaceXAI — Grok 4.6 Developer Documentation, 2026 — https://docs.x.ai/developers/grok-4-6
- Localaimaster.com — “SWE-bench Explained: Complete Guide to AI Coding Benchmarks 2026,” updated August 2026 — https://localaimaster.com/models/swe-bench-explained-ai-benchmarks
- Albato — “Grok, ChatGPT, Gemini, Claude: Full Comparison of Top AI Chatbots (2026),” updated September 2026 — https://albato.com/blog/publications/grok-chatgpt-gemini-claude-overview



