• Raven Paper Reading: How a Harness of Harnesses Plans Multi-Agent Work
    Paper Reading, AI Systems · Oct 1, 2026

    Raven composes model–harness pairs as specialists and asks a Host Agent to plan a dependency DAG. This reading unpacks harness composition, the MAOB benchmark, and what its +10.4/+10.5 percentage-point result actually measures: graph planning match, not successful worker execution.

  • Holo4: One Agent Across GUIs, Code, and Tools—with Different Licenses
    Blog, AI Engineering · Sep 30, 2026

    A closer look at H Company’s cross-interface agent and long-horizon harness, its benchmark claims, public traces, and the licensing split between checkpoints.

  • Kumo Tabular: Pretrain on Synthetic Tables, Learn New Tasks from Examples
    Blog, AI Engineering · Sep 30, 2026

    NVIDIA Kumo Tabular reframes prediction as in-context learning: a model pretrained on synthetic tables predicts new rows from labeled examples. We examine its method, vendor-reported leaderboards, and enterprise validation requirements.

  • EfficientAgent Reading: KV-Cache Offloading for Concurrent Agents
    Paper Reading, AI Systems · Sep 30, 2026

    EfficientAgent asks when moving a KV cache from GPU memory to host memory actually pays off. The answer depends not on one request alone, but on the reuse working set created by concurrent agents between uses. This reading examines capacity prediction, write admission, SWE-bench replay, and hardware boundaries.

  • NVIDIA Open Agent Safety Platform: Verify Runtime Controls Separately from External Monitoring
    Blog, AI Engineering · Sep 29, 2026

    A technical analysis of OpenShell's inspectable agent-runtime controls and NVIDIA's BlueField-4/Sentry reference design, separating their responsibilities from the efficacy evidence still needed.

  • Target Retail Product Search: Constrain Semantic Recall with Precision and Experiments
    Blog, AI Engineering · Sep 29, 2026

    An engineering analysis of Target's lexical and vector retrieval, attribute controls, weighted interleaving, and the evidence limits behind its reported gains.

  • Stale-Document Poisoning: How RAG Can Recognize Evidence That No Longer Applies
    Paper Reading, NLP · Sep 29, 2026

    Authentic, once-correct documents can make RAG overturn an answer the model already got right. This reading examines 317 cross-domain knowledge reversals, the difference between dates and validity intervals, causal interventions, and a reranker whose gains depend on trustworthy metadata.

  • Completed Pairs Hide Capped Failures: Stopping Rules and Unknown Outcomes in Paired Evaluation
    Paper Reading, AI Systems · Sep 29, 2026

    A single ReVerPi source-reading campaign shows how a runner can suppress a companion arm after the first arm reaches its request cap. Finite-frame bounds and stratified cost accounting reveal what completed-pair summaries omit, without making a population claim about context projection.

  • Turning Agent Risks into Runtime Policy: ASSERT Evaluation and ACS Enforcement
    Blog, AI Engineering · Sep 28, 2026

    How Microsoft’s run-assert-eval connects Clarity risk discovery, ASSERT behavior tests, and ACS runtime policy, while measuring unsafe behavior separately from over-refusal.

  • Docker Sandbox Kit: Versioning Agent Authority Requests with Software
    Blog, Cloud & Platform · Sep 28, 2026

    A close look at how Docker Sandbox Kit v3 places agent network, credential, and mixin declarations in an OCI artifact—and why a descriptor remains a request whose enforcement depends on the runtime.

  • Project Swap: Agents Can Trade Without Knowing What Their People Want
    Blog, Industry Pulse · Sep 28, 2026

    Anthropic's low-stakes employee book exchange found that preference estimates constrained outcomes more than bargaining rules, making preference understanding a separate test for delegated agents.

  • Who Signs Off? Let Specifications, Not Agents, Decide When a Task Is Done
    Paper Reading, AI Agent · Sep 28, 2026

    A deep reading of how SpecHarness compiles agent-visible instructions into source-linked obligations and commits state only from qualified evidence, with an examination of SkillsBench and GuideBench results, runtime costs, and limits on attributing gains to sign-off alone.