Tag: AI Safety
Posts with this tag
- NVIDIA Open Agent Safety Platform: Verify Runtime Controls Separately from External Monitoring
A technical analysis of OpenShell's inspectable agent-runtime controls and NVIDIA's BlueField-4/Sentry reference design, separating their responsibilities from the efficacy evidence still needed.
- Turning Agent Risks into Runtime Policy: ASSERT Evaluation and ACS Enforcement
How Microsoft’s run-assert-eval connects Clarity risk discovery, ASSERT behavior tests, and ACS runtime policy, while measuring unsafe behavior separately from over-refusal.
- Turnstile Spin: Let an Agent Install CAPTCHA, Keep Verification at the Security Boundary
A practical look at how Cloudflare Turnstile Spin installs, repairs, or migrates a widget with human approval, while Siteverify, credential isolation, and replay testing close the server-side security loop.
- AI Agents Should Do More Than Agree: How the XY Problem Derails a Fix
XYEval shows how a plausible but misplaced user suggestion can lower agent success; the practical response is to verify the goal, then explain a better path with evidence.
- What Evidence Should an AI Agent's Vulnerability Report Include? MobileCybench Replays Executable Probes
MobileCybench replays Android agent exploits in an isolated environment, then checks trusted state with executable probes to see which security properties were violated.
- Self-Modifying Agents Need Inspectable Release Provenance: What Ouroboros v7.4.4 Shows
A focused engineering reading of Ouroboros v7.4.4: how durable identity, memory, self-modification, and managed subagents can be connected to SHA256, SBOMs, GitHub build provenance, and smoke receipts before deployment.
- AWS Bedrock AgentCore Consent Portal: Agent OAuth Is More Than a Connect Button
A close reading of Amazon Bedrock AgentCore Consent Portal’s end-user OAuth flow: separating the corporate IdP, Gateway, GitHub or Slack outbound provider, callbacks, session binding, token vault, and CloudTrail from the security guarantees AWS does not claim.
- GitSpawn: When Repository Git Config Runs Before a Coding Agent's Trust Boundary
A technical analysis of how repository-local core.fsmonitor can trigger background commands before coding-agent workspace trust and approvals, with the archive delivery condition, patch matrix, safe lab boundary, and enterprise controls.
- Model Hardware Standard: an MCP-shaped interface for physical devices
An evidence-led look at Anthropic's Model Hardware Standard research preview: standardized drivers, discovery, and device control for agents, while identity, authorization, approvals, and physical safety remain platform responsibilities.
- Automated Alignment Researchers: Why Agentic Post-Training Needs Integrity Gates
Anthropic's automated alignment researcher experiment shows how agents can search and iterate on post-training methods while benchmarks, capability floors, data isolation, and integrity review remain outside the agent's authority.
- Gemini 3.6 Flash, 3.5 Flash-Lite, and Flash Cyber: Roles, Pricing, and Adoption Boundaries
A grounded comparison of the official positioning, current pricing, vendor benchmarks, and availability of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, with adoption checks for model routing and cyber permissions.
- Enterprise AI Agent Security: Threat Model, Control Plane, and Rollout Checklist
Build a testable defense-in-depth architecture around prompt injection, tool authorization, exfiltration, memory and identity, supply chain, and observability boundaries.
- Financial-Grade Enterprise Agentic AI Architecture Design: From Demo to Agentic Operating System
AI Summit Recap: Enterprise AI Control Plane, 15+ Agents responsibility breakdown, 4-stage runtime workflow for wealth managers, 3-layer security boundaries, LLM-as-a-Judge quality governance, and E·P·J·T reusable capability foundation.
- Latest Research from OpenAI: How Reinforcement Learning (RL) Makes AI Systems More Aligned and Resilient
An in-depth analysis of OpenAI's latest research on reinforcement learning (RL) and AI alignment. Exploring how models demonstrate broad generalization across more than 40 unseen alignment benchmarks through training focused on 'beneficial traits', and exhibit strong persistence and resilience under malicious fine-tuning and adversarial prompts.
- OpenAI Deployment Simulation: Predicting LLM Safety Before Launch
Read OpenAI Deployment Simulation: why offline eval and real deployment diverge, and what this simulation can and cannot predict.
- Why Reasoning Models 'Cannot Control Their Own Train of Thought' — And Why That's Good News for AI Safety
OpenAI's latest research reveals that current frontier reasoning models are almost completely unable to hide or alter their Chain of Thought (CoT) based on instructions, with maximum controllability at only 15.4%. This 'flaw' is not a problem, but rather the key reason why current CoT monitoring mechanisms can be trusted.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact