AI Coding Agent Benchmark
Compare coding agents by their verified resolution rate on real GitHub issues in SWE-bench Verified.
- 1live-SWE-agent + Claude 4.5 Opus medium (20251101)About this entrylive-SWE-agent + Claude 4.5 Opus medium (20251101) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.UIUClive-SWE-agentRecording79.2% resolvedMetric noteThe primary metric is 79.2% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 2Sonar Foundation Agent + Claude 4.5 OpusAbout this entrySonar Foundation Agent + Claude 4.5 Opus is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.SonarSonar Foundation AgentRecording79.2% resolvedMetric noteThe primary metric is 79.2% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 3TRAE + Doubao-Seed-CodeAbout this entryTRAE + Doubao-Seed-Code is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.ByteDanceTRAERecording78.8% resolvedMetric noteThe primary metric is 78.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 4live-SWE-agent + Gemini 3 Pro Preview (2025-11-18)About this entrylive-SWE-agent + Gemini 3 Pro Preview (2025-11-18) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.UIUClive-SWE-agentRecording77.4% resolvedMetric noteThe primary metric is 77.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 5Atlassian Rovo Dev (2025-09-02)About this entryAtlassian Rovo Dev (2025-09-02) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.AtlassianAtlassian Rovo DevRecording76.8% resolvedMetric noteThe primary metric is 76.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 6EPAM AI/Run Developer Agent v20250719 + Claude 4 SonnetAbout this entryEPAM AI/Run Developer Agent v20250719 + Claude 4 Sonnet is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.EPAM Systems, Inc.EPAM AI/Run Developer AgentRecording76.8% resolvedMetric noteThe primary metric is 76.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 7mini-SWE-agent + Claude 4.5 Opus (high)About this entrymini-SWE-agent + Claude 4.5 Opus (high) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.SWE-agentmini-SWE-agentRecording76.8% resolvedMetric noteThe primary metric is 76.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 8ACoderAbout this entryACoder is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.ACoderACoderRecording76.4% resolvedMetric noteThe primary metric is 76.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 9mini-SWE-agent + Gemini 3 Flash (high)About this entrymini-SWE-agent + Gemini 3 Flash (high) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.SWE-agentmini-SWE-agentRecording75.8% resolvedMetric noteThe primary metric is 75.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 10mini-SWE-agent + MiniMax M2.5 (high)About this entrymini-SWE-agent + MiniMax M2.5 (high) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.SWE-agentmini-SWE-agentRecording75.8% resolvedMetric noteThe primary metric is 75.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 11WarpAbout this entryWarp is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.WarpWarpRecording75.6% resolvedMetric noteThe primary metric is 75.6% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 12mini-SWE-agent + Claude 4.6 OpusAbout this entrymini-SWE-agent + Claude 4.6 Opus is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.SWE-agentmini-SWE-agentRecording75.6% resolvedMetric noteThe primary metric is 75.6% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 13TRAE + Claude Sonnet 4 + Opus 4 + Sonnet 3.7 + Gemini 2.5 ProAbout this entryTRAE + Claude Sonnet 4 + Opus 4 + Sonnet 3.7 + Gemini 2.5 Pro is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.TRAETRAERecording75.2% resolvedMetric noteThe primary metric is 75.2% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 14Harness AIAbout this entryHarness AI is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.HarnessHarness AIRecording74.8% resolvedMetric noteThe primary metric is 74.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 15Sonar Foundation Agent + Claude 4.5 SonnetAbout this entrySonar Foundation Agent + Claude 4.5 Sonnet is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.SonarSonar Foundation AgentRecording74.8% resolvedMetric noteThe primary metric is 74.8% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 16Lingxi-v1.5_claude-4-sonnet-20250514About this entryLingxi-v1.5_claude-4-sonnet-20250514 is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.LingxiLingxi-v1.5Recording74.6% resolvedMetric noteThe primary metric is 74.6% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 17JoyCode + Claude 4 Sonnet + GPT-4.1About this entryJoyCode + Claude 4 Sonnet + GPT-4.1 is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.JoyCodeJoyCodeRecording74.6% resolvedMetric noteThe primary metric is 74.6% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 18Refact.ai Agent + Claude 4 Sonnet + o4-miniAbout this entryRefact.ai Agent + Claude 4 Sonnet + o4-mini is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.Refact.aiRefact.ai AgentRecording74.4% resolvedMetric noteThe primary metric is 74.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 19Prometheus-v1.2.1 + GPT-5About this entryPrometheus-v1.2.1 + GPT-5 is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.EuniAIPrometheus-v1.2.1Recording74.4% resolvedMetric noteThe primary metric is 74.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
- 20mini-SWE-agent + Claude 4.5 Opus (20251101) (medium)About this entrymini-SWE-agent + Claude 4.5 Opus (20251101) (medium) is an Agent or coding system evaluated on SWE-bench Verified. This row represents benchmark performance, not general popularity or production usage.SWE-agentmini-SWE-agentRecording74.4% resolvedMetric noteThe primary metric is 74.4% resolved, sourced from the SWE-bench public benchmark. It is a benchmark resolution rate, not real-world Agent usage volume.
How to read this ranking
This list compares entries only within the same source and metric. It should not be used to compare unlike signals such as stars, request volume, and benchmark resolution rates.
Review the original source