Tag: AI Alignment
Posts with this tag
- Latest Research from OpenAI: How Reinforcement Learning (RL) Makes AI Systems More Aligned and Resilient
An in-depth analysis of OpenAI's latest research on reinforcement learning (RL) and AI alignment. Exploring how models demonstrate broad generalization across more than 40 unseen alignment benchmarks through training focused on 'beneficial traits', and exhibit strong persistence and resilience under malicious fine-tuning and adversarial prompts.
- OpenAI Deployment Simulation: Predicting LLM Safety Before Launch
Read OpenAI Deployment Simulation: why offline eval and real deployment diverge, and what this simulation can and cannot predict.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact