SWEbench - AshokBhat/ml GitHub Wiki

About

  • Benchmark for evaluating AI on real-world software engineering tasks
  • Introduced by researchers from Princeton University
  • Uses real GitHub issues and repositories
  • AI systems must generate code patches that resolve reported issues
  • Evaluated by applying patches and running repository tests
  • SWE-bench Verified contains 500 human-validated tasks
  • Widely used to evaluate AI Coding Agents

Variants

  • SWE-bench — full benchmark
  • SWE-bench Lite — smaller subset
  • SWE-bench Verified — human-validated subset
  • SWE-bench Multimodal — includes visual software tasks
  • SWE-bench Multilingual — covers multiple programming languages

See also

  • [AI Coding Agent]] ](/AshokBhat/ml/wiki/[[GitHub)
  • LLM