Tag: Data Engineering
Posts with this tag
- Kumo Tabular: Pretrain on Synthetic Tables, Learn New Tasks from Examples
NVIDIA Kumo Tabular reframes prediction as in-context learning: a model pretrained on synthetic tables predicts new rows from labeled examples. We examine its method, vendor-reported leaderboards, and enterprise validation requirements.
- Google Cloud Data Agent Kit: Skills, MCP, and Data Workflow Adoption
A practical guide to Data Agent Kit's preview status, open-source artifacts, Skills/MCP/plugin architecture, and enterprise responsibilities for access, cost, validation, and incident recovery.
- Kaggle Titanic: From 0.74 to 0.816, Feature Engineering Outperforms Parameter Tuning
A complete practical record of the Titanic survival prediction competition: progressive feature engineering, CatBoost and RF ensembling, decoupling CV from Public LB, strict notebook porting, and knowing when to stop. Final Public LB 0.81578.
For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.
Speaking & contact