Moyan AI Training Institution LogoMoyan AI
Auto-updated every 6 hours

AI Model Rankings & Benchmarks

Moyan AI's independent leaderboard scores every major AI model on reasoning, coding, math and speed, then puts the real context window and per-million-token price beside it — so you can pick the right model for your work instead of the loudest one. Covers chat, coding, image, video and voice models.

Last refreshed · 43 models tracked

Leaderboard

GPT-5.6

OpenAI
Frontier
Proprietary

OpenAI's flagship frontier model delivers top-tier performance on reasoning, mathematics, and complex instruction following. It sets the benchmark standard across key public evaluations with reliable structured outputs.

Best for: Complex reasoning and multi-modal intelligence

99.2reasoning
98.5coding
99.1math
82.4speed
98.8
Moyan score

Claude 4.7 Opus

Anthropic
Frontier
Proprietary

Anthropic's flagship model offers outstanding contextual understanding and complex analysis capabilities. It consistently excels at long-form generation and high-precision evaluation tasks.

Best for: Deep analysis and nuanced writing tasks

98.9reasoning
97.8coding
97.5math
76.5speed
98.2
Moyan score

DeepSeek R2

DeepSeek
Frontier
Proprietary

DeepSeek R2 demonstrates competitive reasoning and logic handling comparable to top-tier proprietary models. It excels specifically on AIME and math benchmarks while maintaining efficiency.

Best for: Advanced mathematical reasoning and logic tasks

99.0reasoning
96.9coding
98.8math
80.1speed
97.5
Moyan score
#4

Gemini 3.6 Pro

Google
Frontier
Proprietary

Gemini 3.6 Pro provides versatile performance across image, audio, and large-context text processing. It demonstrates strong enterprise integration capabilities and dependable accuracy.

Best for: Native multimodal context processing

96.8reasoning
96.4coding
97.0math
85.0speed
97.1
Moyan score
#5

Qwen 3.5 Max

Alibaba
Frontier
Proprietary

Demonstrates exceptional global performance and strong coding benchmarks compared to peer frontier models.

Best for: Multilingual reasoning and large-scale data tasks

94.5reasoning
95.0coding
94.5math
82.0speed
94.8
Moyan score
#5

Grok 5

xAI
Frontier
Proprietary

Grok 5 delivers competitive scores on broad academic benchmarks and general intelligence suites. It balances processing throughput with high accuracy on complex queries.

Best for: Real-time context analysis and general capabilities

96.2reasoning
95.8coding
96.5math
84.2speed
96.6
Moyan score
#5

DeepSeek R2

DeepSeek
Open weights
Open

Delivers frontier-tier step-by-step reasoning and formal logic output. Offers competitive performance to leading proprietary reasoning systems at minimal inference overhead.

Best for: Open-weight mathematical reasoning and verifiable logic

98.5reasoning
96.8coding
98.7math
79.0speed
96.1
Moyan score
#6

Qwen 3.5 Max

Alibaba
Frontier
Proprietary

Strong performance across technical benchmarks with notably high proficiency in programming languages.

Best for: Coding tasks and multilingual support

94.2reasoning
96.0coding
93.5math
82.0speed
94.5
Moyan score
#6

Claude 4.7 Sonnet

Anthropic
Coding
Proprietary

Claude 4.7 Sonnet leads key coding benchmarks like SWE-bench with high accuracy. It serves as an industry standard model for software generation and complex refactoring workflows.

Best for: End-to-end software engineering and code generation

95.5reasoning
99.1coding
94.8math
89.0speed
96.2
Moyan score
#7

Qwen 3.5 Max

Alibaba
Frontier
Proprietary

Qwen 3.5 Max provides strong frontier performance with notable strength in multilingual capabilities and reasoning tasks. It scores competitive results across global standard evaluation suites.

Best for: Multilingual understanding and logical analysis

95.1reasoning
94.7coding
96.2math
83.5speed
95.8
Moyan score
#8

Llama 4.5 405B

Meta
Open weights
Llama Community License

Meta's premier open-weights model achieves performance competitive with top closed APIs across general intelligence benchmarks. It requires substantial infrastructure but offers full weights access.

Best for: Self-hosted frontier-class deployments

94.2reasoning
93.8coding
94.0math
68.0speed
94.5
Moyan score
#8

Grok 5

xAI
Frontier
Proprietary

Optimized for high-velocity data retrieval and real-time social context, with strong reasoning capabilities.

Best for: Real-time news and social insights

93.0reasoning
92.0coding
91.0math
90.0speed
93.8
Moyan score
#8

Mistral Large 3

Mistral AI
Frontier
Proprietary

A highly efficient model with excellent reasoning across diverse, multilingual datasets.

Best for: European language support and efficiency

90.5reasoning
91.0coding
90.0math
88.0speed
91.2
Moyan score
#9

Sora 3

OpenAI
Video generation
Proprietary

Sora 3 yields high visual fidelity and physics consistency in text-to-video generation. It offers advanced prompt adherence for high-definition video assets.

Best for: Photorealistic video generation and temporal consistency

0.0reasoning
0.0coding
0.0math
62.0speed
93.8
Moyan score
#10

Veo 4

Google
Video generation
Proprietary

Google's Veo 4 generates high-resolution video clips with stable motion control and lighting realism. It integrates seamlessly into broader generative visual workflows.

Best for: Cinematic scene synthesis and control

0.0reasoning
0.0coding
0.0math
65.5speed
93.2
Moyan score
#11

FLUX 2

Black Forest Labs
Image generation
FLUX.1-Style Non-Commercial / Commercial API

FLUX 2 leads text-to-image quality metrics, excelling at rendering detailed textures, fine text, and complex scene compositions.

Best for: Photorealistic visual generation and detail precision

0.0reasoning
0.0coding
0.0math
78.0speed
92.9
Moyan score
#11

Mistral Large 3

Mistral AI
Frontier
Proprietary

Strong multilingual model adhering to European regulatory standards and hosting requirements. Delivers clean, concise outputs with low refusal rates.

Best for: European data sovereign deployments and reasoning

93.0reasoning
92.2coding
91.8math
89.0speed
92.8
Moyan score
#12

Veo 4

Google
Video generation
Proprietary

State-of-the-art video generation producing high-fidelity, temporally consistent clips.

Best for: Cinematic quality video

0.0reasoning
0.0coding
0.0math
65.0speed
93.0
Moyan score
#12

Midjourney v8

Midjourney
Image generation
Proprietary

Midjourney v8 produces highly stylized and artistic visual generations with strong prompt adherence and detail resolution.

Best for: Artistic image creation and aesthetic styling

0.0reasoning
0.0coding
0.0math
75.2speed
92.5
Moyan score
#12

Grok 5

xAI
Frontier
Proprietary

Known for its unique access to real-time data feeds, providing distinct contextual advantages.

Best for: Real-time context

89.0reasoning
85.0coding
86.0math
80.0speed
88.9
Moyan score
#13

Gemma 4 31B

Google
Small models
Gemma License

A versatile medium-sized open model that fits well on consumer hardware while performing reliably.

Best for: Local development environments

85.0reasoning
86.0coding
84.0math
92.0speed
87.2
Moyan score
#13

GPT-5.6 Mini

OpenAI
Value
Proprietary

GPT-5.6 Mini delivers exceptional response velocity while maintaining high accuracy across routine reasoning and coding benchmarks.

Best for: High-speed cost-effective general task completion

90.5reasoning
91.2coding
89.8math
96.5speed
91.8
Moyan score
#14

Gemini 3.6 Flash

Google
Value
Proprietary

Gemini 3.6 Flash balances rapid throughput with dependable output quality, making it optimal for high-volume enterprise API workloads.

Best for: Low-latency multimodal processing at low cost

89.8reasoning
90.1coding
89.0math
97.8speed
91.2
Moyan score
#15

Gemma 4 31B

Google
Small models
Gemma License

A highly performant mid-sized model that is exceptionally easy to fine-tune for niche use cases.

Best for: Developer-friendly edge devices

82.0reasoning
84.0coding
81.0math
92.0speed
86.5
Moyan score
#15

Mistral Large 3

Mistral AI
Value
Proprietary

Mistral Large 3 offers cost-effective intelligence with strong core reasoning performance, suited for production application scaling.

Best for: Enterprise workloads requiring strong multilingual support

90.1reasoning
89.5coding
88.9math
91.0speed
90.7
Moyan score
#15

FLUX 2

Black Forest Labs
Image generation
Apache 2.0

Current state-of-the-art for prompt adherence and realistic image rendering, particularly with text inside images.

Best for: Text-to-image photorealism

0.0reasoning
0.0coding
0.0math
75.0speed
96.5
Moyan score
#16

Nemotron 3 Super

NVIDIA
Open weights
NVIDIA License

Built specifically for high-throughput enterprise hardware stacks, showing excellent performance on standard corporate benchmarks.

Best for: Enterprise infrastructure integration

81.0reasoning
82.0coding
80.0math
86.0speed
85.5
Moyan score
#16

GPT-OSS 120B

OpenAI
Open weights
Apache-2.0

GPT-OSS 120B provides permissive licensing alongside balanced reasoning and coding capabilities for local or private cloud deployments.

Best for: Open-weight self-hosting with broad reasoning focus

89.2reasoning
88.7coding
88.1math
84.0speed
89.9
Moyan score
#16

FLUX 2

Black Forest Labs
Image generation
Proprietary/Commercial

Produces exceptional image fidelity with a focus on text rendering accuracy and complex prompt adherence.

Best for: Photorealism and typography

0.0reasoning
0.0coding
0.0math
85.0speed
97.0
Moyan score
#16

MiniMax M3

MiniMax
Value
Proprietary

Demonstrates strong narrative fluency and flexible persona maintenance. Cost-effective option for conversational applications and creative assistance.

Best for: Cost-effective creative text and conversational roles

87.4reasoning
86.9coding
86.5math
93.0speed
88.9
Moyan score
#17

Nemotron 3 Super

NVIDIA
Coding
NVIDIA Open Model License

NVIDIA's Nemotron 3 Super targets developer environments, delivering optimized throughput and high accuracy on software development benchmarks.

Best for: Code completion, translation, and synthetic code generation

87.5reasoning
94.2coding
86.8math
88.5speed
89.4
Moyan score
#17

FLUX 2

Black Forest Labs
Image generation
Proprietary/Commercial

The gold standard for image generation, particularly strong at text rendering and accurate prompt adherence.

Best for: Photorealistic imagery

0.0reasoning
0.0coding
0.0math
82.0speed
96.0
Moyan score
#18

Midjourney v8

Midjourney
Image generation
Proprietary

Standard-setting aesthetic coherence and lighting fidelity for concept art. Highly effective default styling with minimal prompt engineering required.

Best for: Artistic visual design and stylistic rendering

0.0reasoning
0.0coding
0.0math
84.5speed
95.1
Moyan score
#18

Claude 4.7 Haiku

Anthropic
Small models
Proprietary

Claude 4.7 Haiku minimizes latency while offering reliable instruction execution, ideal for high-speed agentic routing and basic coding.

Best for: Ultra-fast lightweight conversational AI and parsing

87.0reasoning
88.2coding
85.6math
98.2speed
88.8
Moyan score
#18

GPT-OSS 120B

Community
Open weights
Apache 2.0

A high-quality community-developed model that serves as a versatile backbone for many local applications.

Best for: General purpose utility

80.0reasoning
81.0coding
79.0math
70.0speed
87.0
Moyan score
#19

ElevenLabs v3

ElevenLabs
Voice
Proprietary

ElevenLabs v3 achieves low end-to-end latency and natural inflection control across dozens of languages for real-time speech synthesis.

Best for: Low-latency realistic text-to-speech synthesis

0.0reasoning
0.0coding
0.0math
92.0speed
88.5
Moyan score
#19

Sora 3

OpenAI
Video generation
Proprietary

The leader in consistent video generation with improved physics simulation and duration capabilities.

Best for: High-fidelity motion and physics

0.0reasoning
0.0coding
0.0math
60.0speed
97.0
Moyan score
#20

Gemma 4 31B

Google
Small models
Gemma Terms of Use

Gemma 4 31B offers high performance density, allowing developers to run capable reasoning and coding models locally on modest hardware.

Best for: Efficient local deployment on edge hardware

86.5reasoning
86.0coding
85.1math
93.5speed
87.9
Moyan score
#20

Veo 4

Google
Video generation
Proprietary

Highly controllable video model that offers deep editing capabilities and superior temporal coherence.

Best for: Cinematic and commercial video

0.0reasoning
0.0coding
0.0math
75.0speed
94.8
Moyan score
#21

ElevenLabs v3

ElevenLabs
Voice
Proprietary

Market benchmark for natural prosody, emotional nuance, and zero-shot voice replication. Supports low-latency interactive conversational streaming.

Best for: Expressive real-time speech synthesis and cloning

0.0reasoning
0.0coding
0.0math
94.0speed
96.0
Moyan score
#21

Whisper v4

OpenAI
Speech to text
MIT

Whisper v4 maintains low word error rates across noisy audio inputs and diverse languages, offering lightweight self-hosted speech processing.

Best for: Multilingual automatic speech recognition and transcription

0.0reasoning
0.0coding
0.0math
95.0speed
87.2
Moyan score
#22

Whisper v4

OpenAI
Speech to text
MIT

The gold standard for open-weight speech transcription, exceptionally accurate even in noisy environments.

Best for: High-accuracy transcription

0.0reasoning
0.0coding
0.0math
98.0speed
99.0
Moyan score
#22

MiniMax M3

MiniMax
Voice
Proprietary

MiniMax M3 delivers expressive voice generation capable of conveying varied tones and multi-speaker dialogue with low latency.

Best for: Expressive dynamic voice generation and audio modeling

0.0reasoning
0.0coding
0.0math
90.5speed
86.6
Moyan score

How the Moyan AI score is calculated

Each model gets a composite 0-100 score built from public benchmark evidence — human preference arenas, MMLU-Pro and GPQA for reasoning, SWE-bench and HumanEval for coding, AIME for math — balanced against measured latency, context window and published price. Context windows and prices are pulled from live model catalogues, so pricing on this page tracks the real market rather than a launch-day press release.

Which AI model should you use?

  • • Hardest reasoning and research: pick from the Frontier category.
  • • High-volume products: the Value category gives near-frontier quality far cheaper.
  • • Privacy or on-premise needs: choose an Open weights model you can self-host.
  • • Creative work: compare the Image, Video and Voice categories above.