Agent Evaluation
Series · 4 posts
-
OSReward Deep Read: Why Agent Success Cannot Be Judged by Another Model Alone
Advanced Agent runtime, safety, and evaluationA complete reading of OSReward's data construction, 27 VLM judges, Hard and Multi subsets, error and cost analyses, OS-Shepherd-100K training, and a deployable hybrid verification architecture.
Understand it in 90 seconds
-
ContextWeave Deep Read: Does Memory Actually Make Agents Better at Work?
Advanced Agent runtime, safety, and evaluationA close reading of how ContextWeave reconstructs multi-month workflows into an executable benchmark and measures memory's effect on workspace outcomes, preference adherence, continuity, and misleading recall.
Understand it in 90 seconds
-
Real-Time Detection and Repair of LLM Agent Failures: A Deep Read of AgentTrajectorySentinel
Advanced Agent runtime, safety, and evaluationA critical reading of AgentTrajectorySentinel's low-cost healthy-only temporal monitor, deterministic verification, and rollback-and-retry loop, separating measured detection and repair gains from calibration dependence, content blind spots, and reproducibility limits.
Understand it in 90 seconds
-
PAST-Bench: What Did a Persistent Agent Actually Learn from the Past?
Advanced Agent runtime, safety, and evaluationA deep read of how PAST-Bench uses fresh-session task families, matched persistence controls, and trace-level mechanism evidence to separate genuine retained-experience gains from higher scores with unrelated causes.
Understand it in 90 seconds
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact