Skip to content
UtilPick.
My toolbox
All tools
Favorites
Saved tools

불러오는 중…

Recently used

열어 본 도구가 여기에 표시돼요.

Get the desktop app
Settings
UtilPick.
All tools
All tools/AI Model Leaderboard

AI Model Leaderboard

Compare AI models by category scores, value and speed from LiveBench and OpenRouter data, refreshed daily with the date shown, plus image and video rankings.

Loading the tool…

User reviews

No reviews yet. If you’ve used this tool, be the first to leave one.

User reviews

No reviews yet.

Write a review

Checking your sign-in status…

    How to use

    1. 01

      Pick a benchmark tab and the companies you want to see

    2. 02

      Watch the top 12 models race to their places by score

    3. 03

      See the top 30 in the table and click a model name to gather its ranks across every benchmark

    How it works and good to know

    Once a day the server fetches LiveBench benchmarks and OpenRouter model data and re-ranks the models; the date is shown at the top. If the source data can’t be fetched, the built-in ranking is shown with a notice. For score benchmarks the track starts a little below the lowest value instead of zero so differences near the finish line are visible; price, cost, speed and context use a log scale. Tick marks on the track show real values so you can judge the exaggeration. Price is US dollars per million tokens, blended 3 input : 1 output. Logos and characters refer to each company’s trademarks; UtilPick is not affiliated with any of them.

    Related tools

    View all
    Chart MakerTime & Work
    Character CounterText
    Unit Price ComparisonMoney & Math
    © 2026 UtilPickSimple tools for a complicated day. A free collection of online utilities.Get the desktop appAbout UtilPick Request a toolReport a problem
    Privacy Policy (Korean)Terms of Use (Korean)Contact[email protected]

    Today’s leaders

    • Today’s changes…
    01

    Model ranking race

    Fetching the latest rankings…

    LiveBench overall
    Intelligence index
    Coding index
    Agentic index
    Average score across LiveBench’s 7 categories, refreshed monthly with new questions so they rarely leak into training dataHigher is betterSource LiveBench

    LiveBench overall: #1 Claude Fable 5.1 · #1 overall · 83.4%

    Axis doesn’t start at 0
    Start 77.0%80.2%Finish 83.4%
    3
    Claude Fable 5Anthropic · 2026-06-09
    83.0%avg. accuracy
    1
    Claude Fable 5.1NEWAnthropic · 2026-09-02
    83.4%avg. accuracy
    9
    Claude Opus 5Anthropic · 2026-07-25
    80.1%avg. accuracy
    2
    Claude Opus 5.5NEWAnthropic · 2026-09-23
    83.2%avg. accuracy
    6
    DeepSeek V4.1 FlashNEWDeepSeek · 2026-09-10
    81.1%avg. accuracy
    12
    Gemini 3.7 FlashGoogle DeepMind · 2026-08-14
    78.8%avg. accuracy
    8
    GPT-5.5OpenAI · 2026-04-25
    80.2%avg. accuracy
    6
    GPT-5.6 SolOpenAI · 2026-07-09
    81.1%avg. accuracy
    4
    GPT-6 AstraNEWOpenAI · 2026-09-05
    82.2%avg. accuracy
    10
    GPT-6 SolNEWOpenAI · 2026-09-23
    79.2%avg. accuracy
    10
    Kimi K3Moonshot · 2026-07-17
    79.2%avg. accuracy
    5
    Muse Spark 1.3NEWMeta AI · 2026-09-03
    81.6%avg. accuracy

    Today’s changes

    Daily rankings are collected and compared with yesterday

    Rankings started collecting today. From tomorrow, rises, drops and new entries compared with yesterday appear here.

    Value at a glance

    Value based on the real cost of solving LiveBench questions. Switch categories to compare.

    Models on the blue dashed line (value frontier) have no rival scoring higher for the same cost. Click a dot to see that model at a glance.
    Overall · cheapest cost per success first
    1. $0.016 · 65.5 pts
    2. $0.024 · 67.8 pts
    3. $0.026 · 72.0 pts
    4. $0.029 · 81.1 pts
    5. $0.030 · 71.6 pts
    6. $0.042 · 76.2 pts
    7. $0.044 · 77.4 pts
    8. $0.050 · 71.6 pts
    02

    Ranking table

    Average score across LiveBench’s 7 categories, refreshed monthly with new questions so they rarely leak into training data · Higher is better · click a model name to see every benchmark

    Swipe the table sideways to see more columns.

    AI model ranking by LiveBench overall, as of 2026-09-24
    RankModelCompanyavg. accuracyReleased
    1 NEWAnthropic83.4% · max2026-09-02
    2 NEWAnthropic83.2% · max2026-09-23
    3Anthropic83.0% · max2026-06-09
    4 NEWOpenAI82.2% · max2026-09-05
    5 NEWMeta AI81.6% · xhigh2026-09-03
    6 NEWDeepSeek81.1% · max2026-09-10
    6OpenAI81.1% · max2026-07-09
    8OpenAI80.2% · xhigh2026-04-25
    9Anthropic80.1% · max2026-07-25
    10 NEWOpenAI79.2% · max2026-09-23
    10Moonshot79.2%2026-07-17
    12Google DeepMind78.8% · high2026-08-14
    Show sources and method
    As of
    2026-09-24 · source data is fetched and re-ranked once a day
    Source
    • LiveBench.AILiveBench overall; reasoning, coding, agentic coding, math, data analysis, language and instruction following; cost per success and the value chart
    • via OpenRouterPrice and context length, last-day uptime, last-30-minute output speed and time to first token, intelligence/coding/agentic indexes (Artificial Analysis)
    • Artificial AnalysisArena Elo scores and release months for text to image, image editing, text to video and image to video, plus the image and video races in AI War Replay
    Reading the track
    The top 12 models run on the track; see the rest with Show all in the table. Score benchmarks start a little below the lowest value so differences are visible. Price and context use a log scale. Don’t read distance as a multiple — check the tick values.
    Measurement conditions
    Benchmarks use each model’s highest reasoning effort (max, xhigh, high). Price is USD per million tokens, blended 3 input : 1 output. Cost per success is LiveBench’s (cost per question ÷ score) × 100. Release dates are when a model appeared on OpenRouter; models not on OpenRouter are left blank. NEW marks models released within 30 days of the reference date. Image and video scores are Elo ratings from blind side-by-side votes, so differences of a few dozen points are effectively a tie. Releases are known only to the month and counted from the 1st, so NEW is given only to models certainly within 30 days.
    Trademarks
    Logos and characters refer to each company’s trademarks, which belong to their owners. UtilPick is not affiliated with, sponsored or endorsed by any of them. Logo icons: Lobe Icons (MIT).
    Other AI rankingsImage Generation AI RankingVideo Generation AI Ranking