Tag: Harness Engineering
Posts with this tag
- AI Software Development Environments: Choosing Between Vibe Coding and Verified Agent Workflows
InfoWorld surveys GitHub Copilot, Google Antigravity, JetBrains Air, Kiro, Zed, and Zenflow; this article turns that tour into a selection framework based on autonomy, context, isolation, and verification.
- What Is AgentEscapeBench: Measuring Out-of-Domain Tool Reasoning
A deep read of AgentEscapeBench: why agents fail on out-of-domain, long tool chains, and what this benchmark can and cannot show.
- The New Rules of Context Engineering for Claude 5 Models: Trimming 80% of System Prompts
Anthropic removed over 80% of Claude Code's system prompt for Claude Opus 5 and Claude Fable 5 without losing benchmark performance. Explore the paradigm shift from rigid instruction constraints to progressive disclosure, auto-memory, and rich reference harness engineering.
- GPT-5.6 Prompting Checklist: From Shorter Prompts to Tool Orchestration
Turns GPT-5.6 official prompting into a checklist: lean prompts, tool orchestration, and when not to encode policy in prompts. This is the checklist—not the 15% article.
- How Multi-Agent Coding Teams Coordinate: TAKT's YAML Workflows
TAKT schedules multiple coding agents with YAML workflow topology. This article covers coordination control points—not another generic deep dive.
- Self-Scaffolding for Agentic Coding: Ornith 1.0 Training and Evaluation Limits
Read Ornith 1.0's self-scaffolding: what scaffolding solves in agentic coding, what benchmarks can prove, and where trust boundaries sit. Ornith is the case study, not the search entry point.
- GPT-5.6 Sol Prompting: Why Shorter Prompts Score Higher
Reads OpenAI's official GPT-5.6 Sol prompting guidance: what the 15% shorter-prompt figure can and cannot support. This article explains why to write less—not a rules checklist.
- From Vibe Coding to Harness Engineering: A Comprehensive Guide to Google's New 50-Page SDLC Whitepaper
An in-depth review of Google's latest 50-page whitepaper, 'The New SDLC With Vibe Coding', released in 2026. This article breaks down the transformation of the Software Development Life Cycle in the AI era, the Model + Harness framework, automated feedback loops, and the essential skills for developers transitioning into 'quality arbitrators'.
- Step into the Agent Era: Deconstructing the Four Core Pillars of Cursor / Claude Code / Codex
An in-depth analysis of the four Harness mechanisms of modern AI editors—Skills, Subagents, Commands, and Hooks. Clarify the actual configuration formats, trigger timings, and collaborative relationships of each platform, evolving from 'prompt engineering' to 'AI workflow architect'.
- How to Read Harness Engineering: Setup and Verification for Long-Running Agents
A reading map for long-running agent harnesses—setup, verification, handoff—and how to read related notes on this site.
- Harness Engineering Explained: Martin Fowler's AI Coding Workflow
Harness Engineering, in Martin Fowler's terms, means outer control loops around coding agents—feedforward guides plus feedback sensors—so trust is designed control, not a feeling.
- LangChain Dissects Agent Harness: From Model Capabilities to Deliverable Work Engines
A deep dive into LangChain's long-form article: The formal definition of Harness, the component chain deduced from expected behaviors (Files, Bash, Sandbox, Memory, Context Rot, Ralph Loop), and the insights from model-harness co-training and Terminal Bench.
- Mitchell Hashimoto's Six Stages of AI Adoption: From Dropping Chatbots to Harness Engineering
A deep dive into Hashimoto's firsthand journey: three stages of tool adoption, redoing commits for practice, off-peak and slam dunk delegation, AGENTS.md and verifiable tools, plus the current state and limitations of 'always having an Agent running'.
- 16 Parallel Claudes Building a C Compiler: Anthropic's Agent Teams and Long-Running Harness Experiments
A deep dive into Nicholas Carlini's experiment: nearly 2,000 sessions, about $20,000 in API costs, and a 100,000-line Rust compiler capable of compiling Linux 6.9—exploring task locking, test harnesses, GCC oracle, multi-role specialization, and capability boundaries.
- Phil Schmid: Why Agent Harness Is More Important Than Model Leaderboards in 2026
Deep dive into the Jan 2026 article: durability, OS analogies, system evaluation gaps, lightweight Harness under the Bitter Lesson, and hill climbing alongside training-inference convergence.
- Parallel.ai Popular Science: Agent Harness is the Entire Lifecycle Beyond the Model
Deep dive into Parallel's long article: From intent capture, tool execution, context compilation to verification and persistence, clarifying the differences between Harness, orchestrator, and framework, and comparing with Anthropic / LangChain examples.
- Ignorance.ai Playbook: The Converging Harness Practices of OpenAI, Stripe, and OpenClaw
An in-depth review of the February 2026 horizontal roundup: an engineer's work splitting into 'building the environment' and 'managing Agents', architecture as guardrails, tools as feedback, AGENTS.md as system records, and the separation of planning and execution.
- HumanLayer: The Skill Issue of Coding Agents—Practical Implementation of Five Types of Harness Configurations
In-depth read of HumanLayer's long article: Failure is often a configuration issue, not a model one. Clarifying AGENTS.md, MCP, Skills, Sub-agents, Hooks, and back-pressure, and responding to ETH research and post-training overfitting debates.
- Harness Design for Long-Running AI Engineering: Generation, Evaluation, and Verification Chains
Based on Anthropic's 'Harness design for long-running application development': Improving the reliability and controllability of long-running tasks through generator-evaluator separation, external evaluation, and QA contracts.
- Long-Running Agent Harnesses: Handoffs, Verification, and Recovery
A source-grounded analysis of Anthropic's initializer, progress artifacts, feature inventory, Git checkpoints, and end-to-end verification pattern for work spanning context windows.
- Harness Engineering: Making a Codex Repository Legible, Verifiable, and Governable
An analysis of OpenAI's agent-first engineering field report: repository knowledge, observable environments, enforced invariants, and continuous cleanup as a compounding control system.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact