Add your agent to the leaderboard following the instructions on GitHub. 🚀
TGC (Task Goal Completion): The percentage of tasks for which the agent passes all evaluation tests.
SGC (Scenario Goal Completion): The percentage of scenarios for which the agent passes all evaluation tests for every task in the scenario. Each scenario contains three task variants with different requirements and starting states, so SGC measures consistency across these variants.
To ensure fair and meaningful comparisons across leaderboard entries, we expect all submissions to follow the rules below:
Approaches that fall outside these constraints can still be valuable research directions. For example, pass@k evaluates agents using multiple attempts, while online learning or other forms of dependent learning may improve an agent over the course of evaluation. Tree search with state checkpointing could let an agent explore an action sequence, restore an earlier state, and try another branch. We encourage exploring and reporting such approaches in papers and other experimental settings. However, because these settings are not directly comparable to independent, single-attempt evaluation without state resets, and would be difficult to maintain fairly and consistently on this leaderboard, we do not include them in the main leaderboard rankings.
|
|
|||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
|
Method
|
LLM
|
Link | Date | Test-Normal | Test-Challenge | ||||||
| TGC | SGC | TGC | SGC | ||||||||