AI Frontier Research & Engineering
Tracking emerging technology, dissecting system architecture, and publishing what I actually build and evaluate.
Article archive — page 2
- Pydantic AI v2.45: Durable Agent Reliability Is a Session and Trace Contract
A technical reading of how Pydantic AI v2.45.0 aligns DynamicToolset, MCP sessions, tool history, and usage spans with durable runs, plus the adoption boundaries for TypeSafeModel and Bedrock effort handling.
- Gemini Antigravity Agent 09-2026: An Agent Runtime Protocol Migration
A practical breakdown of Antigravity Agent 09-2026's remote/local compatibility boundary, built-in tool contract changes, adapter design, contract tests, and the migration risk before 05-2026 shuts down on October 5, 2026.
- AWS Bedrock AgentCore and Documentation Drift: Code as Authority, MCP at the Write Boundary
A source-grounded analysis of AWS and Corley’s Eutelsat case: use code as the source of truth, RAG for domain context, and human-reviewed MCP writes for governed documentation publishing.
- Amazon SageMaker HyperPod Inference Gateway: Where GPU-Aware Routing Helps
An engineering analysis of how Amazon SageMaker HyperPod Inference Gateway uses KV cache, queue depth, LoRA, and prefix-cache signals, with the EKS add-on, CRD, failure, benchmark, and vendor-claim boundaries made explicit.
- How Claude Speeds Up Biomolecular Models: FlashPairformer and Reversible Inference Kits
An engineering reading of Anthropic's Claude-assisted optimization of more than 30 biomolecular and genomics models, from FlashPairformer and Big mode to stock/exact/fast contracts, cost curves, and evidence limits.
- TypeSafe AI and Jev: Turning AI into a Calibrated Decision Primitive
An engineering reading of TypeSafe AI's System One model and Jev: typed decisions, probability-aware workflows, evaluation claims, and the limits of replacing text generation with decision primitives.
- Jev in the Agent Runtime: Confidence-Gated Routing, Fan-Out, and Community Experiments
A practical architecture guide to placing Jev between agents, tools, and human review through confidence-gated routing, speculative fan-out, composite scoring, and careful evaluation.
- AWS Bedrock AgentCore Consent Portal: Agent OAuth Is More Than a Connect Button
A close reading of Amazon Bedrock AgentCore Consent Portal’s end-user OAuth flow: separating the corporate IdP, Gateway, GitHub or Slack outbound provider, callbacks, session binding, token vault, and CloudTrail from the security guarantees AWS does not claim.
- Redpanda Agentic Data Plane v0.2.61: Push Credentials, Context, and Write Authority to the Agent Boundary
A close reading of Redpanda Agentic Data Plane v0.2.61 and v0.2.60: credential passthrough, context estimates, activity filtering, Pylon capability gates, and the earlier v0.2.58 audit semantics.
- Gemini 3.8 Flash Coding-Agent Workflow: Routing Uncertainty from Planning to Execution
Using Astra planning and Flash execution as an example, this article adds a SPEC verification gate, escalation rules, and a careful reading of DeepSWE costs.
- Unified Knowledge Graph RAG: GraphRAG and LightRAG Are Query Policies, Not Global Switches
A systems reading of AWS's Unified Knowledge Graph RAG reference stack: how GraphRAG and LightRAG share ingestion, graph, hybrid retrieval, and lineage infrastructure while selecting a query strategy per question.
- AWS Step Functions × Bedrock AgentCore: Validate Agent Proposals Before Critical State Changes
A practical control boundary for multi-agent workflows: separate AgentCore proposals from deterministic validation, human approval, idempotent execution, and durable audit history in Step Functions.
- GitSpawn: When Repository Git Config Runs Before a Coding Agent's Trust Boundary
A technical analysis of how repository-local core.fsmonitor can trigger background commands before coding-agent workspace trust and approvals, with the archive delivery condition, patch matrix, safe lab boundary, and enterprise controls.
- Forge v0.18.1: Open-Source Multi-User MCP Auth and Agent Governance
A runtime-contract analysis of how Forge v0.18.1 connects multi-user MCP identity, OAuth consent, tenancy, policy, egress, and per-invocation audit—and where independent evidence is still missing.
- AG2 v1.0.3: MCP 2.0 Migration with Deterministic Agent Governance
A technical review of AG2 v1.0.3's MCP 2.0 breaking migration, TealTiger's deterministic governance path, and a testable rollout boundary for agent runtimes.
- Model Hardware Standard: an MCP-shaped interface for physical devices
An evidence-led look at Anthropic's Model Hardware Standard research preview: standardized drivers, discovery, and device control for agents, while identity, authorization, approvals, and physical safety remain platform responsibilities.
- Automated Alignment Researchers: Why Agentic Post-Training Needs Integrity Gates
Anthropic's automated alignment researcher experiment shows how agents can search and iterate on post-training methods while benchmarks, capability floors, data isolation, and integrity review remain outside the agent's authority.
- WeKnora Architecture: From Document Ingestion to a Governed Agent Knowledge Platform
A systems reading of Tencent WeKnora: its three-process core, ingestion and retrieval flows, and how Agents, MCP, sandboxed Skills, memory, and governance share one control plane.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact