The same seven RoboLab pick-and-place tasks on a Franka DROID arm, three repetitions each, seed 0, DeepSeek policy. Single agent: one Native agent with direct set_gripper control (EXP10A baseline, commit 85a14ac1be16). Delegated: the parent has no set_gripper and hands each object's pick-and-place to a fresh child agent through grasp.engage (config <task>-pickplace-v1, commit 79e1421). Every verdict is the environment's own.
| Task | Single agent | Single | Delegated | Delegated | Change |
|---|
Scores are successes over valid runs. Undecided runs count as valid non-successes on both sides. The one infrastructure-invalid delegated run is excluded.
Choose a repetition to load both runs into the player above. Single-agent attempts are in time order. Repetition n on one side has no special link to repetition n on the other; they are simply the n-th attempt of each.
| Task | Repetition 1 | Repetition 2 | Repetition 3 |
|---|
finish_episode before the environment decided. The single-agent archive calls these "No verdict".