Tag: Architecture Patterns
Posts with this tag
- Target Retail Product Search: Constrain Semantic Recall with Precision and Experiments
An engineering analysis of Target's lexical and vector retrieval, attribute controls, weighted interleaving, and the evidence limits behind its reported gains.
- Turnstile Spin: Let an Agent Install CAPTCHA, Keep Verification at the Security Boundary
A practical look at how Cloudflare Turnstile Spin installs, repairs, or migrates a widget with human approval, while Siteverify, credential isolation, and replay testing close the server-side security loop.
- Self-Modifying Agents Need Inspectable Release Provenance: What Ouroboros v7.4.4 Shows
A focused engineering reading of Ouroboros v7.4.4: how durable identity, memory, self-modification, and managed subagents can be connected to SHA256, SBOMs, GitHub build provenance, and smoke receipts before deployment.
- RSIAgent: Can an Agent Improve Without Updating Model Weights?
A focused engineering reading of RSIAgent's curriculum, actor, verifier, and broad-to-deep exploration loop, with a careful audit of frozen experience, benchmark reporting, and reproducibility limits.
- Unified Knowledge Graph RAG: GraphRAG and LightRAG Are Query Policies, Not Global Switches
A systems reading of AWS's Unified Knowledge Graph RAG reference stack: how GraphRAG and LightRAG share ingestion, graph, hybrid retrieval, and lineage infrastructure while selecting a query strategy per question.
- AWS Step Functions × Bedrock AgentCore: Validate Agent Proposals Before Critical State Changes
A practical control boundary for multi-agent workflows: separate AgentCore proposals from deterministic validation, human approval, idempotent execution, and durable audit history in Step Functions.
- How to Read AI Agent Papers: From CoT and WebGPT to ReAct
One diagram shows how CoT and WebGPT merge into ReAct, then connects Gorilla and IPI to the rest of the agent-systems reading path.
- How to Read RAG Papers: From Dense Retrieval (DPR) to Lewis RAG
One diagram shows how DPR and Lewis RAG connect to the retrieval papers already on this site.
- Agentic AI Platform Contract: The Control Plane You Must Wire Before Production
A copyable Agentic platform contract: what the platform provides, what projects must wire (E·P·J·T), seven non-bypass rules, and where the IT 100-question evidence stops.
- Cloudflare's Open Agentic Internet Blueprint: Readable, Discoverable, Callable, and Payable Web
An architectural deep dive into Cloudflare's proposed Agentic Internet framework, spanning Web Bot Auth, Markdown for Agents, WebMCP browser tool exposure, and x402 micro-payments.
- Kimi-K3 On-Premises Enterprise Cost: GPU, Power, and TCO
GPU memory topologies, server budgets, 3-phase power, liquid cooling, and a five-stage TCO framework for on-premises Kimi-K3 2.8T MoE deployment. Focuses on datacenter procurement—not the 80× RTX 5090 consumer GPU cluster path.
- 80 RTX 5090 GPUs for Kimi-K3: Consumer Hardware Ledger
Hardware topology, why 44 GPUs fall short, build costs, and power estimates for running Kimi-K3 2.8T MoE on an 80× RTX 5090 consumer GPU cluster. Focuses on the consumer-card proof-of-concept—not enterprise datacenter TCO.
- GPT-5.6 Architecture: Frontier Intelligence and Intelligence per Token
Reads the architecture layer of GPT-5.6 technical docs: Frontier Intelligence, intelligence per token, kernels, and harness routing. Architecture—not the Sol price sheet.
- Modern LLM Architecture Comparison: Memory and Routing Trade-offs from DeepSeek V3 to Kimi K2
A comparison of how open-weight models such as DeepSeek V3, Gemma 3, and Kimi K2 use MLA, MoE, GQA, and sliding-window attention to trade off quality, KV cache, throughput, and deployment complexity.
- AI Agent Guide: Architecture, Tools, Evaluation, and Production
A practical guide to agents versus workflows, single- and multi-agent architecture, tools and MCP, state and memory, evaluation, security, and the path from PoC to production.
- Enterprise RAG Guide: Retrieval Architecture, Evaluation, and Production Delivery
A practical framework for enterprise RAG data pipelines, hybrid search, reranking, GraphRAG, agentic RAG, evaluation, access governance, failure diagnosis, and operations.
- Building Enterprise-Grade AI Agents on AWS: Bedrock AgentCore and HoyaBit's Journey from POC to Production
A summary of the AWS × HoyaBit session: the four major pain points of Enterprise Agentic AI, the Amazon Bedrock AgentCore (Runtime/Memory/Gateway/Governance) technology stack, and how Taiwan's FSC-compliant exchange pushed their voice-trading Agent and enterprise brain platform into production.
- Migrating from K8s to Amazon GameLift: Palworld Multiplayer Server Architecture in Practice
A source-bounded review of Pocketpair's Palworld persistent-world migration, with AWS documentation used to verify GameLift scale tests, Spot behavior, and UDP ping beacons.
- Hardware-Software Co-Optimization in Practice: How AWS Custom Silicon × Tomofun Furbo Slashed AI Inference Costs by 82%
A summary of the talk by AWS's Howard and Tomofun (Furbo)'s Ricky, supplemented with the official AWS tech blog: from the vertical integration of Trainium/Inferentia/Graviton/Nitro, to Furbo porting BLIP to Inferentia 2, two-tier Auto Scaling, and AMI cold start optimization, successfully saving 81.6%–83% in AI computing costs.
- From Multi-Agent Architecture to Recruiting an AI Employee in Two Minutes: AWS × Super 8 (ORRA) Enterprise Implementation
A summary of the AWS and Super 8 (ORRA) session: the single agent decision loop, three major multi-agent orchestration patterns (Graph/Swarm/Workflow), A2A communication, Amazon Bedrock AgentCore core components, and how ORRA allows business users to build and deploy AI employees in two minutes using Job Descriptions.
- Decompose with Care: Architecture Patterns, Engineering Disciplines, and Hard Lessons in Banking System Modernization
A comprehensive summary of the AWS ProServ senior consultant's presentation 'Decompose with Care': How a leading Southeast Asian bank modernized its omni-channel monolithic platform serving 20 million active users to AWS cloud-native microservices with zero downtime. Covers four major challenges, Strangler Fig pattern, 3-Tier Facade, Contract-First / Mock-First approaches, and engineering guardrails in the AI era.
- Securely Implementing Multi-tenant AI Agents on AWS EKS: From Sandbox Isolation to BitoClaw Real-world Practice
A summary of the salon sharing by AWS Solutions Architect HC and Bito Group Operations Manager Michael: OWASP LLM Top 10 security red lines, runC / gVisor / Kata sandbox comparison, Kubernetes multi-tenant isolation levels, and how BitoClaw builds a compliant, low-cost AI Agent platform using EKS, KEDA Scale-to-Zero, Pod Identity, and Network Policy.
- From Alexa for Shopping to Agentic Commerce: How Amazon × TapPay Make AI Actually Checkout for You
A source-bounded review of low-latency shopping agents and payment guardrails from the Amazon × TapPay session, checked against first-party Alexa for Shopping and AgentCore material.
- Google ADK 2.0: Workflow Graphs, Task Collaboration, and HITL Boundaries
A source-grounded analysis of how ADK 2.0 separates deterministic routing from LLM reasoning, with production boundaries for workflows, tasks, human approval, and durable state.
- Enterprise AI Agent Security: Threat Model, Control Plane, and Rollout Checklist
Build a testable defense-in-depth architecture around prompt injection, tool authorization, exfiltration, memory and identity, supply chain, and observability boundaries.
- MCP Spec Update: Stateless Core, Tasks, and Apps
What changed in the MCP 2026-07-28 spec versus earlier versions: stateless core, Tasks, Apps, and whether to migrate. The date belongs in the body, not the title.
- Financial-Grade Enterprise Agentic AI Architecture Design: From Demo to Agentic Operating System
AI Summit Recap: Enterprise AI Control Plane, 15+ Agents responsibility breakdown, 4-stage runtime workflow for wealth managers, 3-layer security boundaries, LLM-as-a-Judge quality governance, and E·P·J·T reusable capability foundation.
- Financial AI Engineering Platform Engineering: Building Operational Agentic AI with Cloud-Native Architecture
Summary of my Cloud Summit sharing: Three lifelines for financial AI deployment, why PoCs get stuck, the three-tier architecture of Cloud Native AI Runtime, MCP tool governance, Hybrid Search and Agentic RAG, and why accuracy is a workflow property rather than a model feature.
- Harness Engineering Explained: Martin Fowler's AI Coding Workflow
Harness Engineering, in Martin Fowler's terms, means outer control loops around coding agents—feedforward guides plus feedback sensors—so trust is designed control, not a feeling.
- Ignorance.ai Playbook: The Converging Harness Practices of OpenAI, Stripe, and OpenClaw
An in-depth review of the February 2026 horizontal roundup: an engineer's work splitting into 'building the environment' and 'managing Agents', architecture as guardrails, tools as feedback, AGENTS.md as system records, and the separation of planning and execution.
- Harness Design for Long-Running AI Engineering: Generation, Evaluation, and Verification Chains
Based on Anthropic's 'Harness design for long-running application development': Improving the reliability and controllability of long-running tasks through generator-evaluator separation, external evaluation, and QA contracts.
- Long-Running Agent Harnesses: Handoffs, Verification, and Recovery
A source-grounded analysis of Anthropic's initializer, progress artifacts, feature inventory, Git checkpoints, and end-to-end verification pattern for work spanning context windows.
- How Anthropic Builds Effective Agents: Architecture Patterns and Tactics
An architectural read of Anthropic's Building Effective Agents: workflows, tools, evaluation, and when not to build an agent. This is not the site's complete AI Agent guide.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact