Is the simulator itself any good?
Baseline: a prompt-based user simulator (PBUS) that adds the behavior descriptions to the τ-bench prompt, with no extra modules. Tested with two behaviors at once. PBUS barely moves the agent; ours does, at equal or better initial goal alignment (IGA), so the difficulty comes from the behaviors, not from missing information.
| Benchmark | Simulator | COL | IMP+UNA | TAN+INC | TAN+UNA | IMP+TAN | INC+UNA | INC+IMP |
| SR | IGA | SR | IGA | SR | IGA | SR | IGA | SR | IGA | SR | IGA | SR | IGA |
| MultiWOZ | PBUS | 93.5 | 97.8 | 84.6 | 91.6 | 87.6 | 98.6 | 84.6 | 95.5 | 90.2 | 96.9 | 90.4 | 89.0 | 92.4 | 97.5 |
| Ours | 92.7 | 97.8 | 82.3 | 89.9 | 76.1 | 98.0 | 86.0 | 94.9 | 82.9 | 97.5 | 78.1 | 91.0 | 80.1 | 96.3 |
| τ-bench | PBUS | 38.9 | 87.8 | 40.9 | 92.4 | 44.4 | 95.8 | 39.2 | 97.9 | 43.3 | 96.0 | 39.3 | 92.3 | 45.1 | 88.3 |
| Ours | 45.5 | 97.5 | 40.9 | 98.1 | 34.6 | 96.4 | 36.8 | 99.5 | 33.8 | 97.7 | 40.0 | 97.7 | 38.1 | 93.8 |
Table 4 of the paper, agent GPT-4.1-mini. COL collaborative, IMP impatience, UNA unavailable service, TAN tangential, INC incomplete utterance.
Human evaluation
Nine annotators compared dialogues pairwise. Ours wins about 70% overall and every single-behavior mode:
Unavailable service30.0%70.0%
Incomplete utterance100.0%
PBUS preferredOurs preferred
Figure 5 of the paper, single-behavior mode.