AI Frontier Research & Engineering
Tracking emerging technology, dissecting system architecture, and publishing what I actually build and evaluate.
Article archive — page 3
- What LLM Inference Costs: DeepSeek-R1 Benchmarks and GPU Rental Prices
Read public DeepSeek-R1 MLPerf logs for 8 B200 and B300 GPUs, then apply sourced GPU rental quotes while separating measured latency, arithmetic cost scenarios, and API prices.
- How to Read AI Agent Papers: From CoT and WebGPT to ReAct
One diagram shows how CoT and WebGPT merge into ReAct, then connects Gorilla and IPI to the rest of the agent-systems reading path.
- How to Read RAG Papers: From Dense Retrieval (DPR) to Lewis RAG
One diagram shows how DPR and Lewis RAG connect to the retrieval papers already on this site.
- Agentic AI Platform Contract: The Control Plane You Must Wire Before Production
A copyable Agentic platform contract: what the platform provides, what projects must wire (E·P·J·T), seven non-bypass rules, and where the IT 100-question evidence stops.
- AIPOCH Open Science: Turning Scientific Agents into a Governable Workbench
Break down AIPOCH Open Science v0.19.0: skills, notebook dependencies, OAuth, and artifact provenance. The version number lives in the body; the search entry point is the product name and scientific agents.
- AI Software Development Environments: Choosing Between Vibe Coding and Verified Agent Workflows
InfoWorld surveys GitHub Copilot, Google Antigravity, JetBrains Air, Kiro, Zed, and Zenflow; this article turns that tour into a selection framework based on autonomy, context, isolation, and verification.
- GitHub Copilot MCP Governance: From Allowlist Semantics to Agent Adoption Telemetry
GitHub has added enterprise MCP allow and deny lists alongside third-party agent activity in usage metrics; this article explains the policy semantics, telemetry boundary, and a practical rollout loop.
- Claude Managed Agents Become a Governed Runtime: Budgets, Delegation, Locality, and Inference Hooks
Anthropic's Managed Agents now expose session budgets, advisors, inference geography, repository skills, and inference hooks as runtime controls; this article maps their value and remaining boundaries.
- TREC RAG 2026: Why RAG Evaluation Is Adding Agents
Use TREC RAG 2026 to explain how RAG evaluation moved from document QA to agent-in-the-loop. This article covers direction and task design, not enterprise harness implementation.
- How to Build an Enterprise RAG Evaluation Harness (TREC RAG 2026)
Using TREC RAG 2026 and RAGDoll as references, design a replayable enterprise RAG evaluation harness: data model, citations, agent traces, judge calibration, and launch gates.
- Cloudflare's Open Agentic Internet Blueprint: Readable, Discoverable, Callable, and Payable Web
An architectural deep dive into Cloudflare's proposed Agentic Internet framework, spanning Web Bot Auth, Markdown for Agents, WebMCP browser tool exposure, and x402 micro-payments.
- Stanford CS Professor Chris Piech: AI Can Code, So Why Should You Still Learn Programming?
Stanford CS Professor Chris Piech breaks down software engineering education in the AI era: while LLMs handle syntax, problem-solving, architectural thinking, and human empathy remain irreplaceable multipliers.
- Claude Dreaming: How Agents Consolidate Long-Term Memory Offline
Clarifying the scope of Anthropic's Dreaming feature, what asynchronous memory consolidation can and cannot solve, and how to design a reviewable local Dream Gate.
- What Is Anthropic Agent Memory: Cross-Session Memory vs Dreaming
Untangle Anthropic Agent Memory vs Dreaming: which handles cross-session recall, which runs overnight batches—and do not treat them as the same thing.
- What Changed in OKF 0.2: Provenance and Attested Computation
Compares OKF v0.1 and v0.2: provenance, attested computation, source reputation, and the verified family. This article covers what changed—not an OKF primer.
- Kimi-K3 On-Premises Enterprise Cost: GPU, Power, and TCO
GPU memory topologies, server budgets, 3-phase power, liquid cooling, and a five-stage TCO framework for on-premises Kimi-K3 2.8T MoE deployment. Focuses on datacenter procurement—not the 80× RTX 5090 consumer GPU cluster path.
- 80 RTX 5090 GPUs for Kimi-K3: Consumer Hardware Ledger
Hardware topology, why 44 GPUs fall short, build costs, and power estimates for running Kimi-K3 2.8T MoE on an 80× RTX 5090 consumer GPU cluster. Focuses on the consumer-card proof-of-concept—not enterprise datacenter TCO.
- GPT-5.6 Architecture: Frontier Intelligence and Intelligence per Token
Reads the architecture layer of GPT-5.6 technical docs: Frontier Intelligence, intelligence per token, kernels, and harness routing. Architecture—not the Sol price sheet.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact