← Back to Paper Reading

  • BTS-AgentBench: Compiling Read-Only Telemetry into Replayable Agent Episodes

    Advanced
    Telemetry to Agent Evaluation: Part 1 , Note: Sep 17, 2026 , Paper: 2026 , AI Engineering

    A deep read of Jeong-Yoon Kim's BTS-AgentBench (arXiv:2608.27334 v1): a deterministic path from building telemetry to read-only tools, executable tasks, bounded interaction contracts, and evidence-grounded evaluation—strong on replay consistency, bounded beyond production safety or arbitrary-domain portability.

    Understand it in 90 seconds
    Problem
    Buildings accumulate years of sensor and equipment telemetry, but a raw history is not an executable multi-turn task for an Agent. Hand-authoring each task makes it difficult to preserve a site's vocabulary, source-derived answers, split identity, and evidence links at the same time.
    Core insight
    Treat benchmark construction as a replayable compiler. First place metadata and histories behind read-only tools; then build a static executable task with fixed golds; finally wrap that computation in a typed, bounded interaction contract. Clarification, goal revision, nearest-timestamp policy, quality decisions, and evidence follow-ups can change the surface, but the source computation and its gold must be re-executed together.
    Strongest evidence
    Two independent raw-to-episode builds match all 11 logical tool-store exports and regenerate the BTS 356/87/89 train/dev/test release row by row; all 532 released episodes pass coded contract preflight. This supports construction consistency, not operator realism or production deployment (paper Table 7 and Appendix A.3).
    Main boundary
    BTS-AgentBench is a read-only, offline, bounded building-telemetry benchmark. Its zero controller success is a construction-exclusion condition, not an independent hardness estimate; XAI4HEAT's 41/41 result shows execution on a second telemetry corpus, not portability to arbitrary event logs or physical control.
    Read the full deep dive

For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.

Speaking & contact