Introduction
UseDesktop helps you evaluate agents in resettable professional software environments. Pick an environment, run your model through Desktop, then review traces, grader results, and eval history in Desktop.
The eval loop
Choose an environment
Start from a resettable professional software environment with tasks and graders.
Run a rollout
Use Desktop to connect a model and execute the task.
Review evidence
Use Desktop to compare runs, inspect failures, monitor weak evals, and prepare training data.