Evaluate Your Agent with MAGMA-BENCH
MAGMA-BENCH is being prepared to evaluate agents on interactive, long-horizon, multi-robot tasks.
Release status
The benchmark paper, code, and official evaluation guide are not yet available. We are targeting November, before CoRL, but this timeline is tentative. MAGMA model checkpoints are also planned and have not been released yet.
You can already integrate your model or agent and create your own tasks. These guides let you prepare your integration while the benchmark is under development.
The evaluation guide will document the task suite, metrics, and protocol needed to report comparable benchmark results.