Tag: Evaluation
Posts with this tag
- Holo4: One Agent Across GUIs, Code, and Tools—with Different Licenses
A closer look at H Company’s cross-interface agent and long-horizon harness, its benchmark claims, public traces, and the licensing split between checkpoints.
- Target Retail Product Search: Constrain Semantic Recall with Precision and Experiments
An engineering analysis of Target's lexical and vector retrieval, attribute controls, weighted interleaving, and the evidence limits behind its reported gains.
- Turning Agent Risks into Runtime Policy: ASSERT Evaluation and ACS Enforcement
How Microsoft’s run-assert-eval connects Clarity risk discovery, ASSERT behavior tests, and ACS runtime policy, while measuring unsafe behavior separately from over-refusal.
- Project Swap: Agents Can Trade Without Knowing What Their People Want
Anthropic's low-stakes employee book exchange found that preference estimates constrained outcomes more than bargaining rules, making preference understanding a separate test for delegated agents.
- AI Agents Should Do More Than Agree: How the XY Problem Derails a Fix
XYEval shows how a plausible but misplaced user suggestion can lower agent success; the practical response is to verify the goal, then explain a better path with evidence.
- What Evidence Should an AI Agent's Vulnerability Report Include? MobileCybench Replays Executable Probes
MobileCybench replays Android agent exploits in an isolated environment, then checks trusted state with executable probes to see which security properties were violated.
- RSIAgent: Can an Agent Improve Without Updating Model Weights?
A focused engineering reading of RSIAgent's curriculum, actor, verifier, and broad-to-deep exploration loop, with a careful audit of frozen experience, benchmark reporting, and reproducibility limits.
- What Should We Measure When AI Starts Doing AI R&D? Anthropic's Three Dashboards
Anthropic proposes three measurements for AI-led R&D: automation, agent oversight, and safety compute; this article separates its internal self-report from methodology limits and cross-lab comparability.
- GitHub Agentic Workflows: Turning Runnable Agent CI into a Runtime Contract
A focused engineering reading of gh-aw v0.89.17: how log audits, MCP Gateway and firewall boundaries, grading, model-cost signals, and incident monitoring move Agentic CI beyond a demo.
- Pydantic AI v2.45: Durable Agent Reliability Is a Session and Trace Contract
A technical reading of how Pydantic AI v2.45.0 aligns DynamicToolset, MCP sessions, tool history, and usage spans with durable runs, plus the adoption boundaries for TypeSafeModel and Bedrock effort handling.
- Gemini Antigravity Agent 09-2026: An Agent Runtime Protocol Migration
A practical breakdown of Antigravity Agent 09-2026's remote/local compatibility boundary, built-in tool contract changes, adapter design, contract tests, and the migration risk before 05-2026 shuts down on October 5, 2026.
- Amazon SageMaker HyperPod Inference Gateway: Where GPU-Aware Routing Helps
An engineering analysis of how Amazon SageMaker HyperPod Inference Gateway uses KV cache, queue depth, LoRA, and prefix-cache signals, with the EKS add-on, CRD, failure, benchmark, and vendor-claim boundaries made explicit.
- How Claude Speeds Up Biomolecular Models: FlashPairformer and Reversible Inference Kits
An engineering reading of Anthropic's Claude-assisted optimization of more than 30 biomolecular and genomics models, from FlashPairformer and Big mode to stock/exact/fast contracts, cost curves, and evidence limits.
- TypeSafe AI and Jev: Turning AI into a Calibrated Decision Primitive
An engineering reading of TypeSafe AI's System One model and Jev: typed decisions, probability-aware workflows, evaluation claims, and the limits of replacing text generation with decision primitives.
- Jev in the Agent Runtime: Confidence-Gated Routing, Fan-Out, and Community Experiments
A practical architecture guide to placing Jev between agents, tools, and human review through confidence-gated routing, speculative fan-out, composite scoring, and careful evaluation.
- Gemini 3.8 Flash Coding-Agent Workflow: Routing Uncertainty from Planning to Execution
Using Astra planning and Flash execution as an example, this article adds a SPEC verification gate, escalation rules, and a careful reading of DeepSWE costs.
- Unified Knowledge Graph RAG: GraphRAG and LightRAG Are Query Policies, Not Global Switches
A systems reading of AWS's Unified Knowledge Graph RAG reference stack: how GraphRAG and LightRAG share ingestion, graph, hybrid retrieval, and lineage infrastructure while selecting a query strategy per question.
- Automated Alignment Researchers: Why Agentic Post-Training Needs Integrity Gates
Anthropic's automated alignment researcher experiment shows how agents can search and iterate on post-training methods while benchmarks, capability floors, data isolation, and integrity review remain outside the agent's authority.
- What LLM Inference Costs: DeepSeek-R1 Benchmarks and GPU Rental Prices
Read public DeepSeek-R1 MLPerf logs for 8 B200 and B300 GPUs, then apply sourced GPU rental quotes while separating measured latency, arithmetic cost scenarios, and API prices.
- How to Read AI Agent Papers: From CoT and WebGPT to ReAct
One diagram shows how CoT and WebGPT merge into ReAct, then connects Gorilla and IPI to the rest of the agent-systems reading path.
- How to Read RAG Papers: From Dense Retrieval (DPR) to Lewis RAG
One diagram shows how DPR and Lewis RAG connect to the retrieval papers already on this site.
- TREC RAG 2026: Why RAG Evaluation Is Adding Agents
Use TREC RAG 2026 to explain how RAG evaluation moved from document QA to agent-in-the-loop. This article covers direction and task design, not enterprise harness implementation.
- How to Build an Enterprise RAG Evaluation Harness (TREC RAG 2026)
Using TREC RAG 2026 and RAGDoll as references, design a replayable enterprise RAG evaluation harness: data model, citations, agent traces, judge calibration, and launch gates.
- What Is Anthropic Agent Memory: Cross-Session Memory vs Dreaming
Untangle Anthropic Agent Memory vs Dreaming: which handles cross-session recall, which runs overnight batches—and do not treat them as the same thing.
- Why Reinforcement Learning Breaks Small Language Models
Gradient freezing, numerical overflow, and policy collapse when aligning 70M–500M SLMs with PPO, and what the capacity-headroom hypothesis explains. This article covers failure modes—not a slogan about moving toward robustness.
- What Is AgentEscapeBench: Measuring Out-of-Domain Tool Reasoning
A deep read of AgentEscapeBench: why agents fail on out-of-domain, long tool chains, and what this benchmark can and cannot show.
- GPT-5.6 Prompting Checklist: From Shorter Prompts to Tool Orchestration
Turns GPT-5.6 official prompting into a checklist: lean prompts, tool orchestration, and when not to encode policy in prompts. This is the checklist—not the 15% article.
- Self-Scaffolding for Agentic Coding: Ornith 1.0 Training and Evaluation Limits
Read Ornith 1.0's self-scaffolding: what scaffolding solves in agentic coding, what benchmarks can prove, and where trust boundaries sit. Ornith is the case study, not the search entry point.
- AI Agent Guide: Architecture, Tools, Evaluation, and Production
A practical guide to agents versus workflows, single- and multi-agent architecture, tools and MCP, state and memory, evaluation, security, and the path from PoC to production.
- Enterprise RAG Guide: Retrieval Architecture, Evaluation, and Production Delivery
A practical framework for enterprise RAG data pipelines, hybrid search, reranking, GraphRAG, agentic RAG, evaluation, access governance, failure diagnosis, and operations.
- GPT-5.6 Sol Prompting: Why Shorter Prompts Score Higher
Reads OpenAI's official GPT-5.6 Sol prompting guidance: what the 15% shorter-prompt figure can and cannot support. This article explains why to write less—not a rules checklist.
- What Is GPT-5.6 Sol: Routing, Pricing, and Benchmarks
A product-oriented read of GPT-5.6 Sol: tiered pricing, model routing, and how to interpret benchmarks. This is not the architecture paper.
- OpenAI Deployment Simulation: Predicting LLM Safety Before Launch
Read OpenAI Deployment Simulation: why offline eval and real deployment diverge, and what this simulation can and cannot predict.
- Phil Schmid: Why Agent Harness Is More Important Than Model Leaderboards in 2026
Deep dive into the Jan 2026 article: durability, OS analogies, system evaluation gaps, lightweight Harness under the Bitter Lesson, and hill climbing alongside training-inference convergence.
- Harness Design for Long-Running AI Engineering: Generation, Evaluation, and Verification Chains
Based on Anthropic's 'Harness design for long-running application development': Improving the reliability and controllability of long-running tasks through generator-evaluator separation, external evaluation, and QA contracts.
- Why Reasoning Models 'Cannot Control Their Own Train of Thought' — And Why That's Good News for AI Safety
OpenAI's latest research reveals that current frontier reasoning models are almost completely unable to hide or alter their Chain of Thought (CoT) based on instructions, with maximum controllability at only 15.4%. This 'flaw' is not a problem, but rather the key reason why current CoT monitoring mechanisms can be trusted.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact