AppWorld Icon

Leaderboard

Agent Scores

Add your agent to the leaderboard following the instructions on GitHub. 🚀

Evaluation Metrics

TGC (Task Goal Completion): The percentage of tasks for which the agent passes all evaluation tests.

SGC (Scenario Goal Completion): The percentage of scenarios for which the agent passes all evaluation tests for every task in the scenario. Each scenario contains three task variants with different requirements and starting states, so SGC measures consistency across these variants.

Leaderboard Submission Policy (please read before submitting)

To ensure fair and meaningful comparisons across leaderboard entries, we expect all submissions to follow the rules below:

  • Report pass@1. Each agent should be allowed to attempt each task only once. Submissions should not retry tasks based on oracle evaluation results or other knowledge of whether a previous attempt succeeded.
  • Do not reset to previous states. Agents must not arbitrarily revert the environment to an earlier state. This is possible to do in AppWorld via database state checkpointing. But such resets can undo actions that are irreversible in real life (e.g., sending money), giving an unrealistic advantage.
  • Do not learn from or optimize on the test set. Each test task should be attempted independently by the agent, without using information obtained from other test tasks or their evaluation results. This also rules out manually inspecting test tasks or their evaluation reports to revise prompts or add hints or tips. Training, tuning, or learning from the train and development sets is allowed.

Approaches that fall outside these constraints can still be valuable research directions. For example, pass@k evaluates agents using multiple attempts, while online learning or other forms of dependent learning may improve an agent over the course of evaluation. Tree search with state checkpointing could let an agent explore an action sequence, restore an earlier state, and try another branch. We encourage exploring and reporting such approaches in papers and other experimental settings. However, because these settings are not directly comparable to independent, single-attempt evaluation without state resets, and would be difficult to maintain fairly and consistently on this leaderboard, we do not include them in the main leaderboard rankings.

Method
LLM
Link Date Test-Normal Test-Challenge
TGC SGC TGC SGC