Tag: Platform Engineering
Posts with this tag
- Docker Sandbox Kit: Versioning Agent Authority Requests with Software
A close look at how Docker Sandbox Kit v3 places agent network, credential, and mixin declarations in an OCI artifact—and why a descriptor remains a request whose enforcement depends on the runtime.
- How AlloyDB Isolates Agent Bursts from OLTP: MicroVMs, MCP, and Dedicated Storage Segments
A close look at AlloyDB for agents’ read-only nodes, dedicated Colossus segments, and elastic compute pool, including isolation, freshness, cold starts, cost, and transaction boundaries.
- How Benchling Isolates Multi-Tenant Code Execution with AgentCore, STS, and DNS Firewall
A layered design for running AI-generated code in a separate AWS account, with per-job credentials, S3 endpoint policies, DNS allow-listing, and continuous exfiltration tests.
- GitHub Agentic Workflows: Turning Runnable Agent CI into a Runtime Contract
A focused engineering reading of gh-aw v0.89.17: how log audits, MCP Gateway and firewall boundaries, grading, model-cost signals, and incident monitoring move Agentic CI beyond a demo.
- Pydantic AI v2.45: Durable Agent Reliability Is a Session and Trace Contract
A technical reading of how Pydantic AI v2.45.0 aligns DynamicToolset, MCP sessions, tool history, and usage spans with durable runs, plus the adoption boundaries for TypeSafeModel and Bedrock effort handling.
- Gemini Antigravity Agent 09-2026: An Agent Runtime Protocol Migration
A practical breakdown of Antigravity Agent 09-2026's remote/local compatibility boundary, built-in tool contract changes, adapter design, contract tests, and the migration risk before 05-2026 shuts down on October 5, 2026.
- Amazon SageMaker HyperPod Inference Gateway: Where GPU-Aware Routing Helps
An engineering analysis of how Amazon SageMaker HyperPod Inference Gateway uses KV cache, queue depth, LoRA, and prefix-cache signals, with the EKS add-on, CRD, failure, benchmark, and vendor-claim boundaries made explicit.
- How Claude Speeds Up Biomolecular Models: FlashPairformer and Reversible Inference Kits
An engineering reading of Anthropic's Claude-assisted optimization of more than 30 biomolecular and genomics models, from FlashPairformer and Big mode to stock/exact/fast contracts, cost curves, and evidence limits.
- TypeSafe AI and Jev: Turning AI into a Calibrated Decision Primitive
An engineering reading of TypeSafe AI's System One model and Jev: typed decisions, probability-aware workflows, evaluation claims, and the limits of replacing text generation with decision primitives.
- Jev in the Agent Runtime: Confidence-Gated Routing, Fan-Out, and Community Experiments
A practical architecture guide to placing Jev between agents, tools, and human review through confidence-gated routing, speculative fan-out, composite scoring, and careful evaluation.
- Redpanda Agentic Data Plane v0.2.61: Push Credentials, Context, and Write Authority to the Agent Boundary
A close reading of Redpanda Agentic Data Plane v0.2.61 and v0.2.60: credential passthrough, context estimates, activity filtering, Pylon capability gates, and the earlier v0.2.58 audit semantics.
- Forge v0.18.1: Open-Source Multi-User MCP Auth and Agent Governance
A runtime-contract analysis of how Forge v0.18.1 connects multi-user MCP identity, OAuth consent, tenancy, policy, egress, and per-invocation audit—and where independent evidence is still missing.
- AG2 v1.0.3: MCP 2.0 Migration with Deterministic Agent Governance
A technical review of AG2 v1.0.3's MCP 2.0 breaking migration, TealTiger's deterministic governance path, and a testable rollout boundary for agent runtimes.
- WeKnora Architecture: From Document Ingestion to a Governed Agent Knowledge Platform
A systems reading of Tencent WeKnora: its three-process core, ingestion and retrieval flows, and how Agents, MCP, sandboxed Skills, memory, and governance share one control plane.
- What LLM Inference Costs: DeepSeek-R1 Benchmarks and GPU Rental Prices
Read public DeepSeek-R1 MLPerf logs for 8 B200 and B300 GPUs, then apply sourced GPU rental quotes while separating measured latency, arithmetic cost scenarios, and API prices.
- AIPOCH Open Science: Turning Scientific Agents into a Governable Workbench
Break down AIPOCH Open Science v0.19.0: skills, notebook dependencies, OAuth, and artifact provenance. The version number lives in the body; the search entry point is the product name and scientific agents.
- GitHub Copilot MCP Governance: From Allowlist Semantics to Agent Adoption Telemetry
GitHub has added enterprise MCP allow and deny lists alongside third-party agent activity in usage metrics; this article explains the policy semantics, telemetry boundary, and a practical rollout loop.
- Claude Managed Agents Become a Governed Runtime: Budgets, Delegation, Locality, and Inference Hooks
Anthropic's Managed Agents now expose session budgets, advisors, inference geography, repository skills, and inference hooks as runtime controls; this article maps their value and remaining boundaries.
- Cloudflare's Open Agentic Internet Blueprint: Readable, Discoverable, Callable, and Payable Web
An architectural deep dive into Cloudflare's proposed Agentic Internet framework, spanning Web Bot Auth, Markdown for Agents, WebMCP browser tool exposure, and x402 micro-payments.
- Kimi-K3 On-Premises Enterprise Cost: GPU, Power, and TCO
GPU memory topologies, server budgets, 3-phase power, liquid cooling, and a five-stage TCO framework for on-premises Kimi-K3 2.8T MoE deployment. Focuses on datacenter procurement—not the 80× RTX 5090 consumer GPU cluster path.
- 80 RTX 5090 GPUs for Kimi-K3: Consumer Hardware Ledger
Hardware topology, why 44 GPUs fall short, build costs, and power estimates for running Kimi-K3 2.8T MoE on an 80× RTX 5090 consumer GPU cluster. Focuses on the consumer-card proof-of-concept—not enterprise datacenter TCO.
- GPT-5.6 Architecture: Frontier Intelligence and Intelligence per Token
Reads the architecture layer of GPT-5.6 technical docs: Frontier Intelligence, intelligence per token, kernels, and harness routing. Architecture—not the Sol price sheet.
- Gemini Enterprise Agent Platform: Google Cloud's Build, Scale, Govern, and Optimize Architecture
A structured look at Gemini Enterprise Agent Platform's development, runtime, governance, and evaluation capabilities, with the integration boundaries and adoption conditions that matter in enterprise deployments.
- Migrating from K8s to Amazon GameLift: Palworld Multiplayer Server Architecture in Practice
A source-bounded review of Pocketpair's Palworld persistent-world migration, with AWS documentation used to verify GameLift scale tests, Spot behavior, and UDP ping beacons.
- New Enterprise Governance Challenges in the AI Agent Era: The Dual-Platform Path of FinOps × Agent Governance (OmiFin & MAIAH)
A source-checked reading of eCloudvalley's FinOps and agent-governance talk, grounded in public FOCUS, FinOps Framework, AWS, and MAIAH material.
- Hardware-Software Co-Optimization in Practice: How AWS Custom Silicon × Tomofun Furbo Slashed AI Inference Costs by 82%
A summary of the talk by AWS's Howard and Tomofun (Furbo)'s Ricky, supplemented with the official AWS tech blog: from the vertical integration of Trainium/Inferentia/Graviton/Nitro, to Furbo porting BLIP to Inferentia 2, two-tier Auto Scaling, and AMI cold start optimization, successfully saving 81.6%–83% in AI computing costs.
- DoorDash Ask Assistant Architecture: Where the 24% Conversion Lift Comes From
A breakdown of DoorDash Ask Assistant architecture and evaluation framing. The 24% is their published conversion figure—read it with the experiment scope, not as a headline guarantee.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact