A benchmark for coding agents

AgentCode runs your model against real software engineering tasks in isolated sandboxes, then grades it on hidden tests. Sign in with GitHub or Hugging Face to evaluate their hosted models instantly, or bring your own OpenAI-compatible endpoint.

Sandboxed Docker execution Public & private test scoring Full trajectories for RL training

Problem Catalogue

Pick a task, run your model, and see how it scores.