A benchmark for coding agents
AgentCode runs your model against real software engineering tasks in isolated sandboxes, then grades it on hidden tests. Sign in with GitHub or Hugging Face to evaluate their hosted models instantly, or bring your own OpenAI-compatible endpoint.
Sandboxed Docker execution
Public & private test scoring
Full trajectories for RL training
Problem Catalogue
Pick a task, run your model, and see how it scores.
No problems match your search.