AI Frontier Research & Engineering
Tracking emerging technology, dissecting system architecture, and publishing what I actually build and evaluate.
Article archive — page 7
- How to Use Zellij: Shortcuts, Clipboard, and Sessions
A current-default reference for Zellij Pane, Tab, Resize, Scroll, Session, clipboard, and mouse operations, including the impact of custom keybindings.
- How to Read Harness Engineering: Setup and Verification for Long-Running Agents
A reading map for long-running agent harnesses—setup, verification, handoff—and how to read related notes on this site.
- Harness Engineering Explained: Martin Fowler's AI Coding Workflow
Harness Engineering, in Martin Fowler's terms, means outer control loops around coding agents—feedforward guides plus feedback sensors—so trust is designed control, not a feeling.
- LangChain Dissects Agent Harness: From Model Capabilities to Deliverable Work Engines
A deep dive into LangChain's long-form article: The formal definition of Harness, the component chain deduced from expected behaviors (Files, Bash, Sandbox, Memory, Context Rot, Ralph Loop), and the insights from model-harness co-training and Terminal Bench.
- Mitchell Hashimoto's Six Stages of AI Adoption: From Dropping Chatbots to Harness Engineering
A deep dive into Hashimoto's firsthand journey: three stages of tool adoption, redoing commits for practice, off-peak and slam dunk delegation, AGENTS.md and verifiable tools, plus the current state and limitations of 'always having an Agent running'.
- 16 Parallel Claudes Building a C Compiler: Anthropic's Agent Teams and Long-Running Harness Experiments
A deep dive into Nicholas Carlini's experiment: nearly 2,000 sessions, about $20,000 in API costs, and a 100,000-line Rust compiler capable of compiling Linux 6.9—exploring task locking, test harnesses, GCC oracle, multi-role specialization, and capability boundaries.
- Phil Schmid: Why Agent Harness Is More Important Than Model Leaderboards in 2026
Deep dive into the Jan 2026 article: durability, OS analogies, system evaluation gaps, lightweight Harness under the Bitter Lesson, and hill climbing alongside training-inference convergence.
- Parallel.ai Popular Science: Agent Harness is the Entire Lifecycle Beyond the Model
Deep dive into Parallel's long article: From intent capture, tool execution, context compilation to verification and persistence, clarifying the differences between Harness, orchestrator, and framework, and comparing with Anthropic / LangChain examples.
- Ignorance.ai Playbook: The Converging Harness Practices of OpenAI, Stripe, and OpenClaw
An in-depth review of the February 2026 horizontal roundup: an engineer's work splitting into 'building the environment' and 'managing Agents', architecture as guardrails, tools as feedback, AGENTS.md as system records, and the separation of planning and execution.
- HumanLayer: The Skill Issue of Coding Agents—Practical Implementation of Five Types of Harness Configurations
In-depth read of HumanLayer's long article: Failure is often a configuration issue, not a model one. Clarifying AGENTS.md, MCP, Skills, Sub-agents, Hooks, and back-pressure, and responding to ETH research and post-training overfitting debates.
- The New Rules of Startups in 2026: Why 'Building Capability' Is No Longer the Core Competency
Standing at the startup scene in 2026, we are witnessing an unprecedented paradigm shift. In the AI-native era, development costs and time are extremely compressed, and the bottleneck for startups is no longer 'building capability,' but 'selection capability.' This article reveals the most disruptive core insights in the AI-driven startup ecosystem.
- Agentic RAG: When Vector Search Meets Agentic Reasoning
Core insights from the report 'RAG 2026: When Vector Search Meets Agentic Reasoning', plus the site's shipped IT knowledge Q&A case: hybrid retrieval, context validation, rule-first routing, frozen 100-question weighted 98% with 0 unsafe answers; 2026 direction is coarse vector filter plus deep agentic reading.
- Harness Design for Long-Running AI Engineering: Generation, Evaluation, and Verification Chains
Based on Anthropic's 'Harness design for long-running application development': Improving the reliability and controllability of long-running tasks through generator-evaluator separation, external evaluation, and QA contracts.
- Long-Running Agent Harnesses: Handoffs, Verification, and Recovery
A source-grounded analysis of Anthropic's initializer, progress artifacts, feature inventory, Git checkpoints, and end-to-end verification pattern for work spanning context windows.
- Harness Engineering: Making a Codex Repository Legible, Verifiable, and Governable
An analysis of OpenAI's agent-first engineering field report: repository knowledge, observable environments, enforced invariants, and continuous cleanup as a compounding control system.
- Efficient Academic Paper Reading: The Three-Pass Approach
Turning the 'three-pass reading method' into an actionable workflow: 5–10 minutes for screening, 1 hour to grasp methods and evidence, and 'virtually re-implementing' to master details. This article integrates Keshav's three-pass method with Mu Li's practical tips, complete with a checklist and literature review guide.
- The Impact of AI on the Labor Market: A New Measure from 'Theoretical Capability' to 'Observed Exposure'
A summary based on Anthropic's 'Labor market impacts of AI: A new measure and early evidence': introduces the 'observed exposure' metric, explains which occupations are most exposed to AI, its relationship with employment growth and unemployment rates, and implications for policy, businesses, and individual careers.
- Why Reasoning Models 'Cannot Control Their Own Train of Thought' — And Why That's Good News for AI Safety
OpenAI's latest research reveals that current frontier reasoning models are almost completely unable to hide or alter their Chain of Thought (CoT) based on instructions, with maximum controllability at only 15.4%. This 'flaw' is not a problem, but rather the key reason why current CoT monitoring mechanisms can be trusted.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact