Tag: OpenAI
Posts with this tag
- Gemini 3.8 Flash Coding-Agent Workflow: Routing Uncertainty from Planning to Execution
Using Astra planning and Flash execution as an example, this article adds a SPEC verification gate, escalation rules, and a careful reading of DeepSWE costs.
- GPT-5.6 Architecture: Frontier Intelligence and Intelligence per Token
Reads the architecture layer of GPT-5.6 technical docs: Frontier Intelligence, intelligence per token, kernels, and harness routing. Architecture—not the Sol price sheet.
- GPT-5.6 Prompting Checklist: From Shorter Prompts to Tool Orchestration
Turns GPT-5.6 official prompting into a checklist: lean prompts, tool orchestration, and when not to encode policy in prompts. This is the checklist—not the 15% article.
- OpenAI Presence Announced: Redefining Enterprise AI Agent Governance and High-Stakes Workflows
Unpacking OpenAI Presence, a managed enterprise AI agent platform announced in July 2026. Explore its job-scoped design, governance guardrails, human escalation checkpoints, Codex-driven post-launch improvement loop, and dogfooding via OpenAI's AI Phone Support.
- GPT-5.6 Sol Prompting: Why Shorter Prompts Score Higher
Reads OpenAI's official GPT-5.6 Sol prompting guidance: what the 15% shorter-prompt figure can and cannot support. This article explains why to write less—not a rules checklist.
- What Is GPT-5.6 Sol: Routing, Pricing, and Benchmarks
A product-oriented read of GPT-5.6 Sol: tiered pricing, model routing, and how to interpret benchmarks. This is not the architecture paper.
- GPT-Live Voice Architecture: Full-Duplex Interaction, Delegation, and API Boundaries
A grounded look at GPT-Live's full-duplex and background-delegation design in ChatGPT Voice, how it differs from the Realtime API, and which failure modes voice teams should test.
- Latest Research from OpenAI: How Reinforcement Learning (RL) Makes AI Systems More Aligned and Resilient
An in-depth analysis of OpenAI's latest research on reinforcement learning (RL) and AI alignment. Exploring how models demonstrate broad generalization across more than 40 unseen alignment benchmarks through training focused on 'beneficial traits', and exhibit strong persistence and resilience under malicious fine-tuning and adversarial prompts.
- OpenAI Deployment Simulation: Predicting LLM Safety Before Launch
Read OpenAI Deployment Simulation: why offline eval and real deployment diverge, and what this simulation can and cannot predict.
- Phil Schmid: Why Agent Harness Is More Important Than Model Leaderboards in 2026
Deep dive into the Jan 2026 article: durability, OS analogies, system evaluation gaps, lightweight Harness under the Bitter Lesson, and hill climbing alongside training-inference convergence.
- Why Reasoning Models 'Cannot Control Their Own Train of Thought' — And Why That's Good News for AI Safety
OpenAI's latest research reveals that current frontier reasoning models are almost completely unable to hide or alter their Chain of Thought (CoT) based on instructions, with maximum controllability at only 15.4%. This 'flaw' is not a problem, but rather the key reason why current CoT monitoring mechanisms can be trusted.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact