Tag: AI Agent
Posts with this tag
- AI Software Development Environments: Choosing Between Vibe Coding and Verified Agent Workflows
InfoWorld surveys GitHub Copilot, Google Antigravity, JetBrains Air, Kiro, Zed, and Zenflow; this article turns that tour into a selection framework based on autonomy, context, isolation, and verification.
- GitHub Copilot MCP Governance: From Allowlist Semantics to Agent Adoption Telemetry
GitHub has added enterprise MCP allow and deny lists alongside third-party agent activity in usage metrics; this article explains the policy semantics, telemetry boundary, and a practical rollout loop.
- Claude Managed Agents Become a Governed Runtime: Budgets, Delegation, Locality, and Inference Hooks
Anthropic's Managed Agents now expose session budgets, advisors, inference geography, repository skills, and inference hooks as runtime controls; this article maps their value and remaining boundaries.
- TREC RAG 2026: When RAG Evaluation Becomes an Agent Workflow
A concise entry point to TREC RAG 2026, ClimbMix-400b, and RAGDoll, with a bridge to a technical guide for building a replayable enterprise RAG evaluation harness.
- TREC RAG 2026 Technical Deep Dive: From Evidence Lineage to a Replayable RAG Evaluation Harness
A practical design for an enterprise RAG evaluation harness based on TREC RAG 2026 and RAGDoll: data models, execution stages, citation support, agent traces, judge calibration, and production gates.
- Cloudflare's Open Agentic Internet Blueprint: Readable, Discoverable, Callable, and Payable Web
An architectural deep dive into Cloudflare's proposed Agentic Internet framework, spanning Web Bot Auth, Markdown for Agents, WebMCP browser tool exposure, and x402 micro-payments.
- Stanford CS Professor Chris Piech: AI Can Code, So Why Should You Still Learn Programming?
Stanford CS Professor Chris Piech breaks down software engineering education in the AI era: while LLMs handle syntax, problem-solving, architectural thinking, and human empathy remain irreplaceable multipliers.
- Andrej Karpathy Just Fixed Claude Code's Biggest Weakness: Letting AI "Dream"
Anthropic introduces 'Dreaming' for Claude Code, solving the context amnesia and split-focus issues of in-band memory by allowing agents to consolidate and refine long-term memory overnight.
- Anthropic's Memory and Dreaming for Continuous Agent Learning
Exploring Anthropic's underlying memory systems and the asynchronous Dreaming process for Claude agents to solve context management and continuous learning in multi-agent environments.
- Google Cloud Releases OKF v0.2: A Complete Breakdown of the Open Knowledge Format Upgrade for AI Agents
An in-depth comparative analysis of Google Cloud's Open Knowledge Format (OKF) v0.2 vs v0.1. Explore Attested Computation, provenance credibility signals, structured verified mappings, and staleness lifecycle controls designed to eliminate hallucinations in enterprise AI agents.
- Kimi-K3 Enterprise On-Premises Deployment TCO: GPU Topologies, Power Demands, and Infrastructure Realities for a 2.8T MoE
A comprehensive breakdown of GPU memory topologies, server hardware budgets, 3-phase power, liquid cooling requirements, and a 5-stage TCO evaluation framework for deploying Moonshot's Kimi-K3 2.8T MoE on-premises.
- Running Kimi-K3 on 80 RTX 5090 GPUs: Hardware Ledger and Engineering Trade-Offs of Consumer GPU Clustering
A deep dive into deploying Moonshot Kimi-K3 2.8T MoE on 80 consumer RTX 5090 GPUs: why 44 GPUs fall short, node topologies, 25GbE networking, TCO ledger, and power constraints.
- Analyzing OpenAI GPT-5.6: Frontier Intelligence and the Architectural Shift to 'Intelligence per Token'
A comprehensive deep dive into OpenAI's official GPT-5.6 release: Sol/Terra/Luna tiering, pricing matrices, Terminal-Bench 2.1 data, Triton kernels, and Agentic Harness routing.
- Towards Robust Reinforcement Learning for Small-Scale Language Model Agents: Architecture and Practical Insights
An in-depth analysis of three common PPO failure modes (gradient freezing, numerical overflow, policy collapse) for 70M-500M SLMs, and the Capacity-Headroom Hypothesis (PPL < 20).
- AgentEscapeBench Deep Dive: Benchmarking Out-of-Domain Tool-Grounded Reasoning in LLM Agents
A technical analysis of arXiv:2605.07926 (AgentEscapeBench). Exploring how LLM agent performance degrades across long-range DAG tool dependency graphs, exposing bottlenecks in clue adherence, intermediate output propagation, and state tracking.
- The New Rules of Context Engineering for Claude 5 Models: Trimming 80% of System Prompts
Anthropic removed over 80% of Claude Code's system prompt for Claude Opus 5 and Claude Fable 5 without losing benchmark performance. Explore the paradigm shift from rigid instruction constraints to progressive disclosure, auto-memory, and rich reference harness engineering.
- OpenAI GPT-5.6 Prompting Guidance: From Lean Prompts to Programmatic Tool Orchestration
Unpacking OpenAI's official GPT-5.6 model guidance. Discover how leaner system prompts boost evaluation scores by 15% and cut costs by 67%, alongside autonomy boundaries, text.verbosity, Programmatic Tool Calling, and Pro mode.
- OpenAI Presence Announced: Redefining Enterprise AI Agent Governance and High-Stakes Workflows
Unpacking OpenAI Presence, a managed enterprise AI agent platform announced in July 2026. Explore its job-scoped design, governance guardrails, human escalation checkpoints, Codex-driven post-launch improvement loop, and dogfooding via OpenAI's AI Phone Support.
- Deep Dive into TAKT: Declarative AI Coding Workflows with Agent Coordination Topology
An architectural deep dive into TAKT, an open-source CLI using YAML workflows, isolated worktrees, and strict review loops to orchestrate AI coding agents.
- Ornith 1.0 and Self-Scaffolding: Training, Evaluation, and Trust Boundaries for Agentic Coding
A structured look at Ornith-1.0's Self-Scaffolding, anti-reward-hacking controls, and Pipeline-RL design, plus the evaluation and trust boundaries needed for agentic coding systems.
- In-Depth Analysis of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber: A New Model Architecture for Agentic Applications
Google introduces a brand new Gemini model lineup, featuring the comprehensively upgraded 3.6 Flash, the high-throughput 3.5 Flash-Lite, and the cybersecurity-focused 3.5 Flash Cyber, fully embracing the era of large-scale AI Agent applications.
- Gemini Enterprise Agent Platform: Google Cloud's Build, Scale, Govern, and Optimize Architecture
A structured look at Gemini Enterprise Agent Platform's development, runtime, governance, and evaluation capabilities, with the integration boundaries and adoption conditions that matter in enterprise deployments.
- AI Agent Guide: Architecture, Tools, Evaluation, and Enterprise Delivery
A practical guide to agents versus workflows, single- and multi-agent architecture, tools and MCP, state and memory, evaluation, security, and the path from PoC to production.
- Building Enterprise-Grade AI Agents on AWS: Bedrock AgentCore and HoyaBit's Journey from POC to Production
A summary of the AWS × HoyaBit session: the four major pain points of Enterprise Agentic AI, the Amazon Bedrock AgentCore (Runtime/Memory/Gateway/Governance) technology stack, and how Taiwan's FSC-compliant exchange pushed their voice-trading Agent and enterprise brain platform into production.
- New Enterprise Governance Challenges in the AI Agent Era: The Dual-Platform Path of FinOps × Agent Governance (OmiFin & MAIAH)
A summary of eCloudvalley Enterprise Solution Architect Elmer's talk: In the face of exponential growth in AI application costs, how to balance innovation with controllable operations using FinOps (OmiFin) and AI Agent governance (MAIAH Platform); covering governance/compliance concepts, the PPT (People/Process/Cloud & Platform) cloud governance model, FOCUS billing standards, and BYOA/Guardrails/Token controls.
- From Multi-Agent Architecture to Recruiting an AI Employee in Two Minutes: AWS × Super 8 (ORRA) Enterprise Implementation
A summary of the AWS and Super 8 (ORRA) session: the single agent decision loop, three major multi-agent orchestration patterns (Graph/Swarm/Workflow), A2A communication, Amazon Bedrock AgentCore core components, and how ORRA allows business users to build and deploy AI employees in two minutes using Job Descriptions.
- Securely Implementing Multi-tenant AI Agents on AWS EKS: From Sandbox Isolation to BitoClaw Real-world Practice
A summary of the salon sharing by AWS Solutions Architect HC and Bito Group Operations Manager Michael: OWASP LLM Top 10 security red lines, runC / gVisor / Kata sandbox comparison, Kubernetes multi-tenant isolation levels, and how BitoClaw builds a compliant, low-cost AI Agent platform using EKS, KEDA Scale-to-Zero, Pod Identity, and Network Policy.
- From Alexa for Shopping to Agentic Commerce: How Amazon × TapPay Make AI Actually Checkout for You
A summary of the talk by Amazon and TapPay VP Joseph: How Alexa for Shopping (formerly Rufus) achieved $12B revenue with a single-agent architecture, the difference between Agentic Commerce and traditional shopping guides, and AI autonomous shopping safety guardrails like one-time virtual cards, intent validation, and limit management.
- Siri AI Hands-on Experience: Its Impact and Practical Judgment on Our Daily Way of Using the iPhone
An in-depth analysis of The Verge's first-hand review of Siri AI in the first iOS 27 Public Beta. From new onscreen awareness capabilities and smart calendar parsing to the Entities and Intents architecture developers must implement, we comprehensively dissect the future of Apple's voice intelligence.
- 24% Conversion Boost! DoorDash Unveils the Underlying Architecture of its Ask Assistant Smart Shopping Agent
An in-depth look at how delivery giant DoorDash combined LLMs, domain-specific AI Agents, Model Context Protocol (MCP), and a three-tier memory system to build an enterprise-grade AI shopping assistant capable of running 2,000 automated evaluations daily.
- From Vibe Coding to Harness Engineering: A Comprehensive Guide to Google's New 50-Page SDLC Whitepaper
An in-depth review of Google's latest 50-page whitepaper, 'The New SDLC With Vibe Coding', released in 2026. This article breaks down the transformation of the Software Development Life Cycle in the AI era, the Model + Harness framework, automated feedback loops, and the essential skills for developers transitioning into 'quality arbitrators'.
- What Is Grok 4.5? Capabilities, Benchmarks, and Adoption Decisions
A source-backed review of Grok 4.5's API specifications, coding benchmarks, availability, and the limitations teams should validate before adoption.
- GPT-5.6 Sol Is Generally Available: Routing, Pricing, and Benchmark Caveats
An updated guide to Sol, Terra, and Luna after GPT-5.6 moved from limited preview to general availability, covering API specifications, pricing, benchmark limits, and adoption decisions.
- 2026 Google Cloud Day Taipei: Developer Tech Track Highlights, Moving Fully Towards the Agentic AI Era
Direct insights from the Google Cloud Day Taipei Tech Track! From the underlying TPU hardware and diverse Gemini model lineup to the custom-tailored Anti Gravity 2.0 platform and MCP protocol for developers, explore how Google is building a complete Agent development ecosystem.
- Google Cloud and CyberLink: The Impact and Practical Judgment of AI Engineering on the Multimedia Creation Market
Explore Google Cloud's latest multimedia AI technologies (Imagen 3, Veo, etc.) and how CyberLink uses Promeo to transform these powerful underlying technologies into user-friendly AI Agents, bringing unprecedented commercial competitive advantages to creators and small and medium-sized businesses.
- Google ADK 2.0: Workflow Graphs, Task Collaboration, and HITL Boundaries
A source-grounded analysis of how ADK 2.0 separates deterministic routing from LLM reasoning, with production boundaries for workflows, tasks, human approval, and durable state.
- Enterprise AI Agent Security: Threat Model, Control Plane, and Rollout Checklist
Build a testable defense-in-depth architecture around prompt injection, tool authorization, exfiltration, memory and identity, supply chain, and observability boundaries.
- Google Cloud Data Agent Kit: Skills, MCP, and Data Workflow Adoption
A practical guide to Data Agent Kit's preview status, open-source artifacts, Skills/MCP/plugin architecture, and enterprise responsibilities for access, cost, validation, and incident recovery.
- Ultimate Cloud Computing Liberation: Google Colab CLI Officially Launched, The Strongest Assistant for AI Agents and Developers
Google announces the release of the brand-new Colab CLI tool! Breaking the barrier between local and cloud GPUs, you can instantly invoke powerful computing power through simple terminal commands. It is also the perfect tool for the automated execution of next-generation AI Agents.
- GPT-Live Voice Architecture: Full-Duplex Interaction, Delegation, and API Boundaries
A grounded look at GPT-Live's full-duplex and background-delegation design in ChatGPT Voice, how it differs from the Realtime API, and which failure modes voice teams should test.
- LangChain OpenWiki: The Automated Code Documentation Manager Tailored for AI Agents
An in-depth exploration of LangChain's latest open-source tool, OpenWiki. From the underlying Git Diffs tracking mechanism to the brand-new 'OpenWiki Brains' proactive memory, comprehensively analyzing how to build an exclusive codebase documentation system that reduces Token consumption for AI Coding Agents.
- MCP 2026-07-28: Stateless Core, Tasks, Apps, and Migration Decisions
A practical reading of the Model Context Protocol 2026-07-28 breaking changes, official Tasks and Apps extensions, and the compatibility and security work required for enterprise migration.
- GitLab Orbit In-Depth: Building a Software Development Lifecycle Knowledge Graph for the AI Era
GitLab Orbit is a contextual graph for the software development lifecycle, providing unified, queryable development data for AI Agents and human developers. This article explores Orbit's underlying graph schema, ClickHouse/DuckDB deployment options, and how it combines with MCP to deliver a powerful AI development experience.
- Anthropic Introduces Claude Tag: Making Claude a Permanent AI Teammate for Your Team
Anthropic has released Claude Tag, designed specifically for team collaboration. By tagging @Claude in Slack, AI becomes a virtual teammate that proactively participates in discussions, executes asynchronous tasks, and continuously learns. This article details its core features, usage, target audience, and billing model.
- Step into the Agent Era: Deconstructing the Four Core Pillars of Cursor / Claude Code / Codex
An in-depth analysis of the four Harness mechanisms of modern AI editors—Skills, Subagents, Commands, and Hooks. Clarify the actual configuration formats, trigger timings, and collaborative relationships of each platform, evolving from 'prompt engineering' to 'AI workflow architect'.
- Google Releases Agentic Resource Discovery Specification: The 'Yellow Pages of Capabilities' for the AI Agent Era
An in-depth analysis of the open specification Agentic Resource Discovery (ARD) released by Google in June 2026. This specification aims to standardize how AI Agents discover, verify, and connect with tools, skills, and other Agents in distributed systems, solving the core pain point of multi-agent collaboration: 'How do I find a trusted partner?'
- Anthropic's Latest Research: The State of Agentic Coding and the Persistent Value of Domain Expertise
Anthropic releases a privacy-preserving analysis of 400,000 Claude Code interactions. The research reveals the true division of labor for AI coding agents: humans decide 'what to do', while AI decides 'how to do it'. More importantly, success depends not on 'coding ability', but on 'domain expertise'. This has profound implications for the future of knowledge work.
- PixelRAG: Web Screenshots Beat Text Retrieval! An In-Depth Analysis of a Million-Pixel Native RAG System
Unpacking the PixelRAG system proposed by UC Berkeley and other institutions. An analysis of its custom Chromium rendering, GPU-accelerated preprocessing, LoRA dual-tower visual embedding, and Text Warmup training recipe. Also, learn how to implement it as a web visual reading skill for Claude Code.
- Google Cloud Launches Open Knowledge Format (OKF): An Open Standard for AI Agents to Understand Enterprise Knowledge
An in-depth analysis of the Open Knowledge Format (OKF) specification introduced by Google Cloud. Exploring how it standardizes the traditional LLM-wiki model, breaks down enterprise knowledge silos through concise Markdown, YAML frontmatter, and interconnected structures, and provides a portable, highly interoperable knowledge foundation for AI Agents.
- Harness Engineering Guide
The starting point for the Harness Engineering section on this site: concepts, a full series article index, and reading paths based on roles and scenarios.
- Martin Fowler on Harness: Building Trust in Coding Agents with Control Loops
A deep dive into Thoughtworks' analysis: guides/sensors, computational/inferential, three types of regulation, and the behavior harness gap. By comparing with OpenAI's practical experience and common Agent failure modes, we present an actionable Harness checklist.
- LangChain Dissects Agent Harness: From Model Capabilities to Deliverable Work Engines
A deep dive into LangChain's long-form article: The formal definition of Harness, the component chain deduced from expected behaviors (Files, Bash, Sandbox, Memory, Context Rot, Ralph Loop), and the insights from model-harness co-training and Terminal Bench.
- Mitchell Hashimoto's Six Stages of AI Adoption: From Dropping Chatbots to Harness Engineering
A deep dive into Hashimoto's firsthand journey: three stages of tool adoption, redoing commits for practice, off-peak and slam dunk delegation, AGENTS.md and verifiable tools, plus the current state and limitations of 'always having an Agent running'.
- 16 Parallel Claudes Building a C Compiler: Anthropic's Agent Teams and Long-Running Harness Experiments
A deep dive into Nicholas Carlini's experiment: nearly 2,000 sessions, about $20,000 in API costs, and a 100,000-line Rust compiler capable of compiling Linux 6.9—exploring task locking, test harnesses, GCC oracle, multi-role specialization, and capability boundaries.
- Phil Schmid: Why Agent Harness Is More Important Than Model Leaderboards in 2026
Deep dive into the Jan 2026 article: durability, OS analogies, system evaluation gaps, lightweight Harness under the Bitter Lesson, and hill climbing alongside training-inference convergence.
- Parallel.ai Popular Science: Agent Harness is the Entire Lifecycle Beyond the Model
Deep dive into Parallel's long article: From intent capture, tool execution, context compilation to verification and persistence, clarifying the differences between Harness, orchestrator, and framework, and comparing with Anthropic / LangChain examples.
- Ignorance.ai Playbook: The Converging Harness Practices of OpenAI, Stripe, and OpenClaw
An in-depth review of the February 2026 horizontal roundup: an engineer's work splitting into 'building the environment' and 'managing Agents', architecture as guardrails, tools as feedback, AGENTS.md as system records, and the separation of planning and execution.
- HumanLayer: The Skill Issue of Coding Agents—Practical Implementation of Five Types of Harness Configurations
In-depth read of HumanLayer's long article: Failure is often a configuration issue, not a model one. Clarifying AGENTS.md, MCP, Skills, Sub-agents, Hooks, and back-pressure, and responding to ETH research and post-training overfitting debates.
- The New Rules of Startups in 2026: Why 'Building Capability' Is No Longer the Core Competency
Standing at the startup scene in 2026, we are witnessing an unprecedented paradigm shift. In the AI-native era, development costs and time are extremely compressed, and the bottleneck for startups is no longer 'building capability,' but 'selection capability.' This article reveals the most disruptive core insights in the AI-driven startup ecosystem.
- Agentic RAG: When Vector Search Meets Agentic Reasoning
A summary of my core insights from the report 'RAG 2026: When Vector Search Meets Agentic Reasoning': why pure vector RAG gets stuck in context blindness, why the direction for 2026 is coarse vector filtering plus deep agentic reading, and how enterprises can implement a verifiable, governable hybrid architecture.
- Harness Design for Long-Running AI Engineering: Generation, Evaluation, and Verification Chains
Based on Anthropic's 'Harness design for long-running application development': Improving the reliability and controllability of long-running tasks through generator-evaluator separation, external evaluation, and QA contracts.
- Long-Running Agent Harnesses: Handoffs, Verification, and Recovery
A source-grounded analysis of Anthropic's initializer, progress artifacts, feature inventory, Git checkpoints, and end-to-end verification pattern for work spanning context windows.
- Building Effective AI Agents: An Overview of Architecture Patterns and Implementation Strategies
Adapted from Anthropic's 'Building Effective AI Agents': From single agents to multi-agent collaboration, common architectural patterns, workflow design, and how to choose the right architecture based on control requirements, problem complexity, and resources.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact