Search
Find blog posts, paper notes, and projects
- SWE-Bench ProMax: Can Large-Scale Multilingual Refactoring Measure Long-Horizon Coding Agents?
A deep reading of SWE-Bench ProMax, which uses 170 cross-file, multilingual, behavior-preserving refactoring tasks to test whether coding agents can complete large changes rather than merely fix a nearby test.
- AI Software Development Environments: Choosing Between Vibe Coding and Verified Agent Workflows
InfoWorld surveys GitHub Copilot, Google Antigravity, JetBrains Air, Kiro, Zed, and Zenflow; this article turns that tour into a selection framework based on autonomy, context, isolation, and verification.
- Agentic Configuration Management: Treating Agent Systems as Governed Configuration, Not Just One Execution
A deep reading of how ACM uses a framework-independent Configuration Graph, immutable revisions, dependency-aware impact propagation, and runtime provenance to govern heterogeneous agent configurations across LangGraph, CrewAI, and the OpenAI Agents SDK.
- ADIAS: Turning Agent Self-Improvement into Traceable Issue Repair
A deep reading of ADIAS: persistent issue state organizes failure evidence across optimization rounds so a full-code agent designer can remember what was tried, what regressed, and when a repair is actually confirmed.
- DocMemo: Letting Long-Document RAG Recover from a Bad First Retrieval
A deep reading of DocMemo: document schema, page belief, and question episodic memory preserve retrieval state across rounds, while Bayesian updates, Thompson sampling, and adaptive granularity recover missed evidence.
- GitHub Copilot MCP Governance: From Allowlist Semantics to Agent Adoption Telemetry
GitHub has added enterprise MCP allow and deny lists alongside third-party agent activity in usage metrics; this article explains the policy semantics, telemetry boundary, and a practical rollout loop.
- Claude Managed Agents Become a Governed Runtime: Budgets, Delegation, Locality, and Inference Hooks
Anthropic's Managed Agents now expose session budgets, advisors, inference geography, repository skills, and inference hooks as runtime controls; this article maps their value and remaining boundaries.
- FinRank: Hard-Negative Retrieval Evaluation for Financial-Document RAG
A deep reading of FinRank: how company, year, and disclosure boundaries create deceptively plausible evidence, and why pooled retrieval, hard negatives, and metadata filters must be evaluated together.
- A²E: A Traceable, Re-Evaluable Engine for Agent Auditing
A deep reading of A²E: ATP aligns benchmarks with agent harnesses, span-based traces preserve execution causality, and lifecycle-aligned metrics analyze correctness, tools, cost, and safety.
- TREC RAG 2026: When RAG Evaluation Becomes an Agent Workflow
A concise entry point to TREC RAG 2026, ClimbMix-400b, and RAGDoll, with a bridge to a technical guide for building a replayable enterprise RAG evaluation harness.
- TREC RAG 2026 Technical Deep Dive: From Evidence Lineage to a Replayable RAG Evaluation Harness
A practical design for an enterprise RAG evaluation harness based on TREC RAG 2026 and RAGDoll: data models, execution stages, citation support, agent traces, judge calibration, and production gates.
- Cloudflare's Open Agentic Internet Blueprint: Readable, Discoverable, Callable, and Payable Web
An architectural deep dive into Cloudflare's proposed Agentic Internet framework, spanning Web Bot Auth, Markdown for Agents, WebMCP browser tool exposure, and x402 micro-payments.
Huahua checked every note
Nothing matches this search yet.
Try a broader keyword, switch the content type, or continue from one of the curated paths below.