Agent Evaluation and Human-Agent Collaboration
Series · 1 posts
-
XYEval: Why agents follow bad advice
Advanced Agent runtime, safety, and evaluationA critical reading of Wu et al.'s XYEval (arXiv 2609.23939 v1): controlled XY mutations across five models and six benchmark suites, with trace analyses of how misleading suggestions affect task completion, communication, and tool trajectories—and the limits of generators and judges.
Understand it in 90 seconds
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact