Tag: Machine Learning
Posts with this tag
- Kumo Tabular: Pretrain on Synthetic Tables, Learn New Tasks from Examples
NVIDIA Kumo Tabular reframes prediction as in-context learning: a model pretrained on synthetic tables predicts new rows from labeled examples. We examine its method, vendor-reported leaderboards, and enterprise validation requirements.
- Target Retail Product Search: Constrain Semantic Recall with Precision and Experiments
An engineering analysis of Target's lexical and vector retrieval, attribute controls, weighted interleaving, and the evidence limits behind its reported gains.
- Why Reinforcement Learning Breaks Small Language Models
Gradient freezing, numerical overflow, and policy collapse when aligning 70M–500M SLMs with PPO, and what the capacity-headroom hypothesis explains. This article covers failure modes—not a slogan about moving toward robustness.
- Self-Scaffolding for Agentic Coding: Ornith 1.0 Training and Evaluation Limits
Read Ornith 1.0's self-scaffolding: what scaffolding solves in agentic coding, what benchmarks can prove, and where trust boundaries sit. Ornith is the case study, not the search entry point.
- DoorDash Ask Assistant Architecture: Where the 24% Conversion Lift Comes From
A breakdown of DoorDash Ask Assistant architecture and evaluation framing. The 24% is their published conversion figure—read it with the experiment scope, not as a headline guarantee.
- What Is Grok 4.5? Capabilities, Benchmarks, and Adoption Decisions
A source-backed review of Grok 4.5's API specifications, coding benchmarks, availability, and the limitations teams should validate before adoption.
- What Is GPT-5.6 Sol: Routing, Pricing, and Benchmarks
A product-oriented read of GPT-5.6 Sol: tiered pricing, model routing, and how to interpret benchmarks. This is not the architecture paper.
- Google Colab CLI: A Cloud Terminal for AI Agents
Google Colab CLI lets agents and developers run terminal commands in the cloud. This post covers what it can do, permissions, and limits—not a "compute liberation" launch piece.
- Kaggle Titanic: From 0.74 to 0.816, Feature Engineering Outperforms Parameter Tuning
A complete practical record of the Titanic survival prediction competition: progressive feature engineering, CatBoost and RF ensembling, decoupling CV from Public LB, strict notebook porting, and knowing when to stop. Final Public LB 0.81578.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact