Model Leaderboard

Compare AI models by capability and cost-effectiveness

Popular Comparisons

Logical Reasoning

71/251 models

HLE: Complex reasoning and problem-solving

Use Cases: Complex decision-making, multi-step analysis, logical reasoning

Knowledge Q&A

62/251 models

MMLU Pro: Broad knowledge assessment

Use Cases: Expert Q&A, fact-checking, educational tutoring

Scientific Research

74/251 models

GPQA: Graduate-level science questions

Use Cases: Academic research, scientific writing, experiment design

AI Agent

41/251 models

Tau2: Autonomous task completion

Use Cases: Automated workflows, multi-tool invocation, complex task decomposition

SciCode

55/251 models

SciCode: Scientific coding challenges

Use Cases: Scientific computing, research code, data analysis scripts

Terminal

66/251 models

Terminal-Bench 2.1: agentic coding & terminal operations

Use Cases: Agentic coding, shell scripting, DevOps automation

Instruction

53/251 models

IFEval: Instruction following accuracy

Use Cases: Precise task execution, format compliance, constraint adherence

Disclaimer: Rankings are for reference only and do not represent precise test results or constitute any purchase or usage advice. We do not guarantee the accuracy, completeness, or timeliness of the data.

Data Sources: Rankings are based on official technical reports and public evaluations from model providers.