AI Coding Agent Benchmark

Compare coding agents by their verified resolution rate on real GitHub issues in SWE-bench Verified.

Top 20#Project / model / agentContextHistoryPrimary metric
  1. 1
    live-SWE-agent + Claude 4.5 Opus medium (20251101)About this entrylive-SWE-agent + Claude 4.5 Opus medium (20251101) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    UIUC
    live-SWE-agent
    Recording79.2% resolvedMetric noteThe primary metric is 79.2% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  2. 2
    Sonar Foundation Agent + Claude 4.5 OpusAbout this entrySonar Foundation Agent + Claude 4.5 Opus is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    Sonar
    Sonar Foundation Agent
    Recording79.2% resolvedMetric noteThe primary metric is 79.2% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  3. 3
    TRAE + Doubao-Seed-CodeAbout this entryTRAE + Doubao-Seed-Code is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    ByteDance
    TRAE
    Recording78.8% resolvedMetric noteThe primary metric is 78.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  4. 4
    live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)About this entrylive-SWE-agent + Gemini 3 Pro Preview (2025-11-18) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    UIUC
    live-SWE-agent
    Recording77.4% resolvedMetric noteThe primary metric is 77.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  5. 5
    Atlassian Rovo Dev (2025-09-02)About this entryAtlassian Rovo Dev (2025-09-02) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    Atlassian
    Atlassian Rovo Dev
    Recording76.8% resolvedMetric noteThe primary metric is 76.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  6. 6
    EPAM AI/Run Developer Agent v20250719 + Claude 4 SonnetAbout this entryEPAM AI/Run Developer Agent v20250719 + Claude 4 Sonnet is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    EPAM Systems, Inc.
    EPAM AI/Run Developer Agent
    Recording76.8% resolvedMetric noteThe primary metric is 76.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  7. 7
    mini-SWE-agent + Claude 4.5 Opus (high)About this entrymini-SWE-agent + Claude 4.5 Opus (high) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    SWE-agent
    mini-SWE-agent
    Recording76.8% resolvedMetric noteThe primary metric is 76.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  8. 8
    ACoderAbout this entryACoder is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    ACoder
    ACoder
    Recording76.4% resolvedMetric noteThe primary metric is 76.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  9. 9
    mini-SWE-agent + Gemini 3 Flash (high)About this entrymini-SWE-agent + Gemini 3 Flash (high) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    SWE-agent
    mini-SWE-agent
    Recording75.8% resolvedMetric noteThe primary metric is 75.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  10. 10
    mini-SWE-agent + MiniMax M2.5 (high)About this entrymini-SWE-agent + MiniMax M2.5 (high) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    SWE-agent
    mini-SWE-agent
    Recording75.8% resolvedMetric noteThe primary metric is 75.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  11. 11
    WarpAbout this entryWarp is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    Warp
    Warp
    Recording75.6% resolvedMetric noteThe primary metric is 75.6% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  12. 12
    mini-SWE-agent + Claude 4.6 OpusAbout this entrymini-SWE-agent + Claude 4.6 Opus is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    SWE-agent
    mini-SWE-agent
    Recording75.6% resolvedMetric noteThe primary metric is 75.6% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  13. 13
    TRAE + Claude Sonnet 4 + Opus 4 + Sonnet 3.7 + Gemini 2.5 ProAbout this entryTRAE + Claude Sonnet 4 + Opus 4 + Sonnet 3.7 + Gemini 2.5 Pro is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    TRAE
    TRAE
    Recording75.2% resolvedMetric noteThe primary metric is 75.2% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  14. 14
    Harness AIAbout this entryHarness AI is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    Harness
    Harness AI
    Recording74.8% resolvedMetric noteThe primary metric is 74.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  15. 15
    Sonar Foundation Agent + Claude 4.5 SonnetAbout this entrySonar Foundation Agent + Claude 4.5 Sonnet is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    Sonar
    Sonar Foundation Agent
    Recording74.8% resolvedMetric noteThe primary metric is 74.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  16. 16
    Lingxi-v1.5_claude-4-sonnet-20250514About this entryLingxi-v1.5_claude-4-sonnet-20250514 is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    Lingxi
    Lingxi-v1.5
    Recording74.6% resolvedMetric noteThe primary metric is 74.6% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  17. 17
    JoyCode + Claude 4 Sonnet + GPT-4.1About this entryJoyCode + Claude 4 Sonnet + GPT-4.1 is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    JoyCode
    JoyCode
    Recording74.6% resolvedMetric noteThe primary metric is 74.6% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  18. 18
    Refact.ai Agent + Claude 4 Sonnet + o4-miniAbout this entryRefact.ai Agent + Claude 4 Sonnet + o4-mini is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    Refact.ai
    Refact.ai Agent
    Recording74.4% resolvedMetric noteThe primary metric is 74.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  19. 19
    Prometheus-v1.2.1 + GPT-5About this entryPrometheus-v1.2.1 + GPT-5 is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    EuniAI
    Prometheus-v1.2.1
    Recording74.4% resolvedMetric noteThe primary metric is 74.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
  20. 20
    mini-SWE-agent + Claude 4.5 Opus (20251101) (medium)About this entrymini-SWE-agent + Claude 4.5 Opus (20251101) (medium) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.
    SWE-agent
    mini-SWE-agent
    Recording74.4% resolvedMetric noteThe primary metric is 74.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.

How to read this ranking

This list compares entries only within the same source and metric. It should not be used to compare unlike signals such as stars, request volume, and benchmark resolution rates.

Review the original source