Tag: Enterprise AI
Posts with this tag
- NVIDIA Open Agent Safety Platform: Verify Runtime Controls Separately from External Monitoring
A technical analysis of OpenShell's inspectable agent-runtime controls and NVIDIA's BlueField-4/Sentry reference design, separating their responsibilities from the efficacy evidence still needed.
- Target Retail Product Search: Constrain Semantic Recall with Precision and Experiments
An engineering analysis of Target's lexical and vector retrieval, attribute controls, weighted interleaving, and the evidence limits behind its reported gains.
- Turning Agent Risks into Runtime Policy: ASSERT Evaluation and ACS Enforcement
How Microsoft’s run-assert-eval connects Clarity risk discovery, ASSERT behavior tests, and ACS runtime policy, while measuring unsafe behavior separately from over-refusal.
- Docker Sandbox Kit: Versioning Agent Authority Requests with Software
A close look at how Docker Sandbox Kit v3 places agent network, credential, and mixin declarations in an OCI artifact—and why a descriptor remains a request whose enforcement depends on the runtime.
- Project Swap: Agents Can Trade Without Knowing What Their People Want
Anthropic's low-stakes employee book exchange found that preference estimates constrained outcomes more than bargaining rules, making preference understanding a separate test for delegated agents.
- How AlloyDB Isolates Agent Bursts from OLTP: MicroVMs, MCP, and Dedicated Storage Segments
A close look at AlloyDB for agents’ read-only nodes, dedicated Colossus segments, and elastic compute pool, including isolation, freshness, cold starts, cost, and transaction boundaries.
- Turnstile Spin: Let an Agent Install CAPTCHA, Keep Verification at the Security Boundary
A practical look at how Cloudflare Turnstile Spin installs, repairs, or migrates a widget with human approval, while Siteverify, credential isolation, and replay testing close the server-side security loop.
- How Benchling Isolates Multi-Tenant Code Execution with AgentCore, STS, and DNS Firewall
A layered design for running AI-generated code in a separate AWS account, with per-job credentials, S3 endpoint policies, DNS allow-listing, and continuous exfiltration tests.
- Cross-Account MCP with AWS AgentCore Gateway: Keep Data in Each LOB Account
Trace how AWS’s multi-account AgentCore Gateway pattern separates user JWT authorization through Cedar from downstream M2M OAuth, and identify the sample’s production hardening gaps.
- Self-Modifying Agents Need Inspectable Release Provenance: What Ouroboros v7.4.4 Shows
A focused engineering reading of Ouroboros v7.4.4: how durable identity, memory, self-modification, and managed subagents can be connected to SHA256, SBOMs, GitHub build provenance, and smoke receipts before deployment.
- Pydantic AI v2.45: Durable Agent Reliability Is a Session and Trace Contract
A technical reading of how Pydantic AI v2.45.0 aligns DynamicToolset, MCP sessions, tool history, and usage spans with durable runs, plus the adoption boundaries for TypeSafeModel and Bedrock effort handling.
- AWS Bedrock AgentCore and Documentation Drift: Code as Authority, MCP at the Write Boundary
A source-grounded analysis of AWS and Corley’s Eutelsat case: use code as the source of truth, RAG for domain context, and human-reviewed MCP writes for governed documentation publishing.
- TypeSafe AI and Jev: Turning AI into a Calibrated Decision Primitive
An engineering reading of TypeSafe AI's System One model and Jev: typed decisions, probability-aware workflows, evaluation claims, and the limits of replacing text generation with decision primitives.
- Jev in the Agent Runtime: Confidence-Gated Routing, Fan-Out, and Community Experiments
A practical architecture guide to placing Jev between agents, tools, and human review through confidence-gated routing, speculative fan-out, composite scoring, and careful evaluation.
- AWS Bedrock AgentCore Consent Portal: Agent OAuth Is More Than a Connect Button
A close reading of Amazon Bedrock AgentCore Consent Portal’s end-user OAuth flow: separating the corporate IdP, Gateway, GitHub or Slack outbound provider, callbacks, session binding, token vault, and CloudTrail from the security guarantees AWS does not claim.
- Redpanda Agentic Data Plane v0.2.61: Push Credentials, Context, and Write Authority to the Agent Boundary
A close reading of Redpanda Agentic Data Plane v0.2.61 and v0.2.60: credential passthrough, context estimates, activity filtering, Pylon capability gates, and the earlier v0.2.58 audit semantics.
- Unified Knowledge Graph RAG: GraphRAG and LightRAG Are Query Policies, Not Global Switches
A systems reading of AWS's Unified Knowledge Graph RAG reference stack: how GraphRAG and LightRAG share ingestion, graph, hybrid retrieval, and lineage infrastructure while selecting a query strategy per question.
- AWS Step Functions × Bedrock AgentCore: Validate Agent Proposals Before Critical State Changes
A practical control boundary for multi-agent workflows: separate AgentCore proposals from deterministic validation, human approval, idempotent execution, and durable audit history in Step Functions.
- GitSpawn: When Repository Git Config Runs Before a Coding Agent's Trust Boundary
A technical analysis of how repository-local core.fsmonitor can trigger background commands before coding-agent workspace trust and approvals, with the archive delivery condition, patch matrix, safe lab boundary, and enterprise controls.
- Forge v0.18.1: Open-Source Multi-User MCP Auth and Agent Governance
A runtime-contract analysis of how Forge v0.18.1 connects multi-user MCP identity, OAuth consent, tenancy, policy, egress, and per-invocation audit—and where independent evidence is still missing.
- AG2 v1.0.3: MCP 2.0 Migration with Deterministic Agent Governance
A technical review of AG2 v1.0.3's MCP 2.0 breaking migration, TealTiger's deterministic governance path, and a testable rollout boundary for agent runtimes.
- Model Hardware Standard: an MCP-shaped interface for physical devices
An evidence-led look at Anthropic's Model Hardware Standard research preview: standardized drivers, discovery, and device control for agents, while identity, authorization, approvals, and physical safety remain platform responsibilities.
- WeKnora Architecture: From Document Ingestion to a Governed Agent Knowledge Platform
A systems reading of Tencent WeKnora: its three-process core, ingestion and retrieval flows, and how Agents, MCP, sandboxed Skills, memory, and governance share one control plane.
- What LLM Inference Costs: DeepSeek-R1 Benchmarks and GPU Rental Prices
Read public DeepSeek-R1 MLPerf logs for 8 B200 and B300 GPUs, then apply sourced GPU rental quotes while separating measured latency, arithmetic cost scenarios, and API prices.
- Agentic AI Platform Contract: The Control Plane You Must Wire Before Production
A copyable Agentic platform contract: what the platform provides, what projects must wire (E·P·J·T), seven non-bypass rules, and where the IT 100-question evidence stops.
- GitHub Copilot MCP Governance: From Allowlist Semantics to Agent Adoption Telemetry
GitHub has added enterprise MCP allow and deny lists alongside third-party agent activity in usage metrics; this article explains the policy semantics, telemetry boundary, and a practical rollout loop.
- TREC RAG 2026: Why RAG Evaluation Is Adding Agents
Use TREC RAG 2026 to explain how RAG evaluation moved from document QA to agent-in-the-loop. This article covers direction and task design, not enterprise harness implementation.
- How to Build an Enterprise RAG Evaluation Harness (TREC RAG 2026)
Using TREC RAG 2026 and RAGDoll as references, design a replayable enterprise RAG evaluation harness: data model, citations, agent traces, judge calibration, and launch gates.
- Cloudflare's Open Agentic Internet Blueprint: Readable, Discoverable, Callable, and Payable Web
An architectural deep dive into Cloudflare's proposed Agentic Internet framework, spanning Web Bot Auth, Markdown for Agents, WebMCP browser tool exposure, and x402 micro-payments.
- What Is Anthropic Agent Memory: Cross-Session Memory vs Dreaming
Untangle Anthropic Agent Memory vs Dreaming: which handles cross-session recall, which runs overnight batches—and do not treat them as the same thing.
- Kimi-K3 On-Premises Enterprise Cost: GPU, Power, and TCO
GPU memory topologies, server budgets, 3-phase power, liquid cooling, and a five-stage TCO framework for on-premises Kimi-K3 2.8T MoE deployment. Focuses on datacenter procurement—not the 80× RTX 5090 consumer GPU cluster path.
- 80 RTX 5090 GPUs for Kimi-K3: Consumer Hardware Ledger
Hardware topology, why 44 GPUs fall short, build costs, and power estimates for running Kimi-K3 2.8T MoE on an 80× RTX 5090 consumer GPU cluster. Focuses on the consumer-card proof-of-concept—not enterprise datacenter TCO.
- GPT-5.6 Architecture: Frontier Intelligence and Intelligence per Token
Reads the architecture layer of GPT-5.6 technical docs: Frontier Intelligence, intelligence per token, kernels, and harness routing. Architecture—not the Sol price sheet.
- OpenAI Presence Announced: Redefining Enterprise AI Agent Governance and High-Stakes Workflows
Unpacking OpenAI Presence, a managed enterprise AI agent platform announced in July 2026. Explore its job-scoped design, governance guardrails, human escalation checkpoints, Codex-driven post-launch improvement loop, and dogfooding via OpenAI's AI Phone Support.
- Gemini Enterprise Agent Platform: Google Cloud's Build, Scale, Govern, and Optimize Architecture
A structured look at Gemini Enterprise Agent Platform's development, runtime, governance, and evaluation capabilities, with the integration boundaries and adoption conditions that matter in enterprise deployments.
- AI Agent Guide: Architecture, Tools, Evaluation, and Production
A practical guide to agents versus workflows, single- and multi-agent architecture, tools and MCP, state and memory, evaluation, security, and the path from PoC to production.
- Enterprise RAG Guide: Retrieval Architecture, Evaluation, and Production Delivery
A practical framework for enterprise RAG data pipelines, hybrid search, reranking, GraphRAG, agentic RAG, evaluation, access governance, failure diagnosis, and operations.
- Building Enterprise-Grade AI Agents on AWS: Bedrock AgentCore and HoyaBit's Journey from POC to Production
A summary of the AWS × HoyaBit session: the four major pain points of Enterprise Agentic AI, the Amazon Bedrock AgentCore (Runtime/Memory/Gateway/Governance) technology stack, and how Taiwan's FSC-compliant exchange pushed their voice-trading Agent and enterprise brain platform into production.
- New Enterprise Governance Challenges in the AI Agent Era: The Dual-Platform Path of FinOps × Agent Governance (OmiFin & MAIAH)
A source-checked reading of eCloudvalley's FinOps and agent-governance talk, grounded in public FOCUS, FinOps Framework, AWS, and MAIAH material.
- From Multi-Agent Architecture to Recruiting an AI Employee in Two Minutes: AWS × Super 8 (ORRA) Enterprise Implementation
A summary of the AWS and Super 8 (ORRA) session: the single agent decision loop, three major multi-agent orchestration patterns (Graph/Swarm/Workflow), A2A communication, Amazon Bedrock AgentCore core components, and how ORRA allows business users to build and deploy AI employees in two minutes using Job Descriptions.
- Decompose with Care: Architecture Patterns, Engineering Disciplines, and Hard Lessons in Banking System Modernization
A comprehensive summary of the AWS ProServ senior consultant's presentation 'Decompose with Care': How a leading Southeast Asian bank modernized its omni-channel monolithic platform serving 20 million active users to AWS cloud-native microservices with zero downtime. Covers four major challenges, Strangler Fig pattern, 3-Tier Facade, Contract-First / Mock-First approaches, and engineering guardrails in the AI era.
- Securely Implementing Multi-tenant AI Agents on AWS EKS: From Sandbox Isolation to BitoClaw Real-world Practice
A summary of the salon sharing by AWS Solutions Architect HC and Bito Group Operations Manager Michael: OWASP LLM Top 10 security red lines, runC / gVisor / Kata sandbox comparison, Kubernetes multi-tenant isolation levels, and how BitoClaw builds a compliant, low-cost AI Agent platform using EKS, KEDA Scale-to-Zero, Pod Identity, and Network Policy.
- From Alexa for Shopping to Agentic Commerce: How Amazon × TapPay Make AI Actually Checkout for You
A source-bounded review of low-latency shopping agents and payment guardrails from the Amazon × TapPay session, checked against first-party Alexa for Shopping and AgentCore material.
- DoorDash Ask Assistant Architecture: Where the 24% Conversion Lift Comes From
A breakdown of DoorDash Ask Assistant architecture and evaluation framing. The 24% is their published conversion figure—read it with the experiment scope, not as a headline guarantee.
- Enterprise AI Agent Security: Threat Model, Control Plane, and Rollout Checklist
Build a testable defense-in-depth architecture around prompt injection, tool authorization, exfiltration, memory and identity, supply chain, and observability boundaries.
- LangChain OpenWiki: The Automated Code Documentation Manager Tailored for AI Agents
An in-depth exploration of LangChain's latest open-source tool, OpenWiki. From the underlying Git Diffs tracking mechanism to the brand-new 'OpenWiki Brains' proactive memory, comprehensively analyzing how to build an exclusive codebase documentation system that reduces Token consumption for AI Coding Agents.
- GraphRAG In-Depth Analysis: How to Build Smarter AI Retrieval Workflows Using Knowledge Graphs?
Explore the highlights of Cassie Shum's talk at QCon AI. Learn from the ground up how GraphRAG solves enterprise RAG pain points through Global Context, Multi-hop Reasoning, and Cypher queries, with practical architectural implementations.
- GitLab Orbit: Querying Code and SDLC Relationships for AI Agents
GitLab Orbit builds a queryable graph from code and software-lifecycle data; this article clarifies Remote, Local, MCP, and the current Beta and Experiment boundaries.
- Financial-Grade Enterprise Agentic AI Architecture Design: From Demo to Agentic Operating System
AI Summit Recap: Enterprise AI Control Plane, 15+ Agents responsibility breakdown, 4-stage runtime workflow for wealth managers, 3-layer security boundaries, LLM-as-a-Judge quality governance, and E·P·J·T reusable capability foundation.
- Financial AI Engineering Platform Engineering: Building Operational Agentic AI with Cloud-Native Architecture
Summary of my Cloud Summit sharing: Three lifelines for financial AI deployment, why PoCs get stuck, the three-tier architecture of Cloud Native AI Runtime, MCP tool governance, Hybrid Search and Agentic RAG, and why accuracy is a workflow property rather than a model feature.
- Harness Engineering Explained: Martin Fowler's AI Coding Workflow
Harness Engineering, in Martin Fowler's terms, means outer control loops around coding agents—feedforward guides plus feedback sensors—so trust is designed control, not a feeling.
- LangChain Dissects Agent Harness: From Model Capabilities to Deliverable Work Engines
A deep dive into LangChain's long-form article: The formal definition of Harness, the component chain deduced from expected behaviors (Files, Bash, Sandbox, Memory, Context Rot, Ralph Loop), and the insights from model-harness co-training and Terminal Bench.
- Mitchell Hashimoto's Six Stages of AI Adoption: From Dropping Chatbots to Harness Engineering
A deep dive into Hashimoto's firsthand journey: three stages of tool adoption, redoing commits for practice, off-peak and slam dunk delegation, AGENTS.md and verifiable tools, plus the current state and limitations of 'always having an Agent running'.
- Phil Schmid: Why Agent Harness Is More Important Than Model Leaderboards in 2026
Deep dive into the Jan 2026 article: durability, OS analogies, system evaluation gaps, lightweight Harness under the Bitter Lesson, and hill climbing alongside training-inference convergence.
- Parallel.ai Popular Science: Agent Harness is the Entire Lifecycle Beyond the Model
Deep dive into Parallel's long article: From intent capture, tool execution, context compilation to verification and persistence, clarifying the differences between Harness, orchestrator, and framework, and comparing with Anthropic / LangChain examples.
- Ignorance.ai Playbook: The Converging Harness Practices of OpenAI, Stripe, and OpenClaw
An in-depth review of the February 2026 horizontal roundup: an engineer's work splitting into 'building the environment' and 'managing Agents', architecture as guardrails, tools as feedback, AGENTS.md as system records, and the separation of planning and execution.
- HumanLayer: The Skill Issue of Coding Agents—Practical Implementation of Five Types of Harness Configurations
In-depth read of HumanLayer's long article: Failure is often a configuration issue, not a model one. Clarifying AGENTS.md, MCP, Skills, Sub-agents, Hooks, and back-pressure, and responding to ETH research and post-training overfitting debates.
- Agentic RAG: When Vector Search Meets Agentic Reasoning
Core insights from the report 'RAG 2026: When Vector Search Meets Agentic Reasoning', plus the site's shipped IT knowledge Q&A case: hybrid retrieval, context validation, rule-first routing, frozen 100-question weighted 98% with 0 unsafe answers; 2026 direction is coarse vector filter plus deep agentic reading.
- How Anthropic Builds Effective Agents: Architecture Patterns and Tactics
An architectural read of Anthropic's Building Effective Agents: workflows, tools, evaluation, and when not to build an agent. This is not the site's complete AI Agent guide.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact