Skip to content
aicoolies logo
SWE-bench logo

SWE-bench

Benchmark for evaluating AI coding agents on real GitHub issues

SWE-bench is a benchmark from Princeton NLP that evaluates AI coding agents by testing their ability to resolve real GitHub issues from popular open-source projects. Each task provides an issue description and repository state, and the agent must produce a working patch that passes the project's test suite. With 4,600+ GitHub stars, it has become the standard yardstick for comparing autonomous coding tools like Devin, Claude Code, and OpenHands.

About SWE-bench

SWE-bench is a benchmark dataset and evaluation framework created by researchers at Princeton's NLP group that measures how well AI systems can solve real-world software engineering tasks. Unlike synthetic coding benchmarks that test isolated function generation, SWE-bench uses actual GitHub issues from twelve popular Python repositories including Django, Flask, scikit-learn, and matplotlib. Each task consists of a natural language issue description, the repository at the commit where the issue was filed, and a set of tests that validate whether the proposed fix actually resolves the problem.

The benchmark has become the de facto standard for evaluating autonomous coding agents. When companies like Cognition demonstrate Devin, or when Anthropic reports Claude Code's capabilities, or when OpenHands publishes agent performance data, they reference SWE-bench scores. The evaluation uses Docker-based reproducible environments to ensure that test results are consistent across different hardware and software configurations. SWE-bench Verified is a curated subset of 500 tasks reviewed by software engineers to confirm that each task is unambiguous and solvable.

Beyond raw benchmarking, SWE-bench has influenced how the industry thinks about AI coding evaluation. It demonstrated that generating syntactically correct code is insufficient — agents must understand project architecture, navigate large codebases, identify relevant files, and produce patches that integrate correctly with existing test suites. The benchmark is MIT licensed with over 4,600 GitHub stars and continues to be maintained with new evaluation variants. For teams building or selecting AI coding tools, SWE-bench provides the only widely accepted, reproducible metric for comparing agent capabilities on realistic software engineering work.

Pricing & Platform Specs

Pricing Summary

Free and 100% open source under the MIT license by Princeton NLP. SWE-bench has no platform or benchmark fees; researchers and developers only provide their own Docker compute and LLM inference tokens.

full pricing breakdown →

Supported Platforms

Python, Docker (evaluation framework)

Explore categories, tags & use cases

Categories

Apache-2.0 Python framework for repeatable LLM, RAG, agent, MCP, and safety evaluation workflows.

DeepEval is an Apache-2.0 Python framework for evaluating LLM apps, RAG systems, agents, MCP workflows, and safety behavior with repeatable test cases. It works locally and in CI/CD, then connects to Confident AI for hosted reports, observability, red teaming, and governance when teams need shared evidence instead of ad-hoc prompt reviews and manual QA.

freemiumOpen Source

LLM testing and evaluation toolkit

Promptfoo is an OpenAI-owned open-source toolkit for evaluating, red-teaming and securing LLM applications. It supports config-driven prompt/model tests, CI regression gates, red-team scans, guardrails, model security workflows, MCP Proxy, code scanning and evaluations across prompts, agents and RAG pipelines.

freemiumOpen Source

LLM evaluation and tracking with RAG triad metrics

TruLens is an open-source framework for evaluating and tracking LLM experiments with feedback functions, RAG triad metrics (answer relevance, context relevance, groundedness), and Honest/Harmless/Helpful evaluations. Features a unified Metric API for systematic evaluation of RAG pipelines and AI agents. 3,200+ GitHub stars, MIT licensed. Snowflake partnership adds enterprise integration. Supports LangChain, LlamaIndex, and custom LLM applications.

Open Source

LLM evaluation and prompt engineering platform

Braintrust is an AI observability and evaluation platform for tracing LLM applications, building datasets, running prompt/model experiments, scoring outputs and turning production feedback into regression tests. It fits teams that need repeatable quality gates for AI releases rather than one-off prompt demos.

freemium

Community experience

Sources & verification

Sources checked
Content verified

Verification dates are editorial checks. Routine CMS saves and automatic updatedAt timestamps do not advance them.

FAQ

What is SWE-bench?

SWE-bench is a benchmark from Princeton NLP that evaluates AI coding agents by testing their ability to resolve real GitHub issues from popular open-source projects. Each task provides an issue description and repository state, and the agent must produce a working patch that passes the project's test suite. With 4,600+ GitHub stars, it has become the standard yardstick for comparing autonomous coding tools like Devin, Claude Code, and OpenHands.

Is SWE-bench free?

Yes — SWE-bench is open source and free to use. Free and 100% open source under the MIT license by Princeton NLP. SWE-bench has no platform or benchmark fees; researchers and developers only provide their own Docker compute and LLM inference tokens.

Is SWE-bench open source?

Yes — SWE-bench is open source.

Is SWE-bench still maintained?

Yes — SWE-bench is active. Its listing was last verified on September 6, 2026.

What are the best SWE-bench alternatives?

The first editor-selected SWE-bench alternatives are DeepEval, Promptfoo, TruLens, and more.