AI Frontier Research & Engineering
Tracking emerging technology, dissecting system architecture, and publishing what I actually build and evaluate.
Article archive — page 4
- Why Reinforcement Learning Breaks Small Language Models
Gradient freezing, numerical overflow, and policy collapse when aligning 70M–500M SLMs with PPO, and what the capacity-headroom hypothesis explains. This article covers failure modes—not a slogan about moving toward robustness.
- Modern LLM Architecture Comparison: Memory and Routing Trade-offs from DeepSeek V3 to Kimi K2
A comparison of how open-weight models such as DeepSeek V3, Gemma 3, and Kimi K2 use MLA, MoE, GQA, and sliding-window attention to trade off quality, KV cache, throughput, and deployment complexity.
- What Is AgentEscapeBench: Measuring Out-of-Domain Tool Reasoning
A deep read of AgentEscapeBench: why agents fail on out-of-domain, long tool chains, and what this benchmark can and cannot show.
- The New Rules of Context Engineering for Claude 5 Models: Trimming 80% of System Prompts
Anthropic removed over 80% of Claude Code's system prompt for Claude Opus 5 and Claude Fable 5 without losing benchmark performance. Explore the paradigm shift from rigid instruction constraints to progressive disclosure, auto-memory, and rich reference harness engineering.
- GPT-5.6 Prompting Checklist: From Shorter Prompts to Tool Orchestration
Turns GPT-5.6 official prompting into a checklist: lean prompts, tool orchestration, and when not to encode policy in prompts. This is the checklist—not the 15% article.
- OpenAI Presence Announced: Redefining Enterprise AI Agent Governance and High-Stakes Workflows
Unpacking OpenAI Presence, a managed enterprise AI agent platform announced in July 2026. Explore its job-scoped design, governance guardrails, human escalation checkpoints, Codex-driven post-launch improvement loop, and dogfooding via OpenAI's AI Phone Support.
- How Multi-Agent Coding Teams Coordinate: TAKT's YAML Workflows
TAKT schedules multiple coding agents with YAML workflow topology. This article covers coordination control points—not another generic deep dive.
- Self-Scaffolding for Agentic Coding: Ornith 1.0 Training and Evaluation Limits
Read Ornith 1.0's self-scaffolding: what scaffolding solves in agentic coding, what benchmarks can prove, and where trust boundaries sit. Ornith is the case study, not the search entry point.
- Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber: Roles, Pricing, and Adoption Boundaries
A grounded comparison of the official positioning, current pricing, vendor benchmarks, and availability of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, with adoption checks for model routing and cyber permissions.
- Gemini Enterprise Agent Platform: Google Cloud's Build, Scale, Govern, and Optimize Architecture
A structured look at Gemini Enterprise Agent Platform's development, runtime, governance, and evaluation capabilities, with the integration boundaries and adoption conditions that matter in enterprise deployments.
- Bundesliga Partners with AWS: Generative AI Creates a Global Fan Experience
How the Bundesliga is leveraging AWS and Generative AI to bring an innovative, closer-to-the-action experience and automated content production to its 1 billion fans worldwide.
- AI Agent Guide: Architecture, Tools, Evaluation, and Production
A practical guide to agents versus workflows, single- and multi-agent architecture, tools and MCP, state and memory, evaluation, security, and the path from PoC to production.
- Enterprise RAG Guide: Retrieval Architecture, Evaluation, and Production Delivery
A practical framework for enterprise RAG data pipelines, hybrid search, reranking, GraphRAG, agentic RAG, evaluation, access governance, failure diagnosis, and operations.
- Building Enterprise-Grade AI Agents on AWS: Bedrock AgentCore and HoyaBit's Journey from POC to Production
A summary of the AWS × HoyaBit session: the four major pain points of Enterprise Agentic AI, the Amazon Bedrock AgentCore (Runtime/Memory/Gateway/Governance) technology stack, and how Taiwan's FSC-compliant exchange pushed their voice-trading Agent and enterprise brain platform into production.
- Migrating from K8s to Amazon GameLift: Palworld Multiplayer Server Architecture in Practice
A source-bounded review of Pocketpair's Palworld persistent-world migration, with AWS documentation used to verify GameLift scale tests, Spot behavior, and UDP ping beacons.
- New Enterprise Governance Challenges in the AI Agent Era: The Dual-Platform Path of FinOps × Agent Governance (OmiFin & MAIAH)
A source-checked reading of eCloudvalley's FinOps and agent-governance talk, grounded in public FOCUS, FinOps Framework, AWS, and MAIAH material.
- Hardware-Software Co-Optimization in Practice: How AWS Custom Silicon × Tomofun Furbo Slashed AI Inference Costs by 82%
A summary of the talk by AWS's Howard and Tomofun (Furbo)'s Ricky, supplemented with the official AWS tech blog: from the vertical integration of Trainium/Inferentia/Graviton/Nitro, to Furbo porting BLIP to Inferentia 2, two-tier Auto Scaling, and AMI cold start optimization, successfully saving 81.6%–83% in AI computing costs.
- From Multi-Agent Architecture to Recruiting an AI Employee in Two Minutes: AWS × Super 8 (ORRA) Enterprise Implementation
A summary of the AWS and Super 8 (ORRA) session: the single agent decision loop, three major multi-agent orchestration patterns (Graph/Swarm/Workflow), A2A communication, Amazon Bedrock AgentCore core components, and how ORRA allows business users to build and deploy AI employees in two minutes using Job Descriptions.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact