← Lab · Competitions

Lab note

Titanic — Machine Learning from Disaster

Kaggle · CatBoost · Feature Engineering · StratifiedKFold

View on GitHub →

Titanic — Machine Learning from Disaster Lab Bloss0m Note 000

Kaggle Titanic is a classic practice for introductory tabular classification: predicting survival based on passenger features. This project went through a complete ML mini-cycle of research → progressive implementation → submission → wrap-up, ultimately achieving a Public LB of 0.81578 (approximately 341 / 418 correct), landing at the practical upper bound of legitimate solutions.

For the complete methodology and step-by-step lessons, please see the blog post: Kaggle Titanic: From 0.74 to 0.816, Feature Engineering is More Important than Hyperparameter Tuning.


Context

Small-sample tabular binary classification (891 training, 418 test), with the evaluation metric being Accuracy. The goal is not to chase a 1.0 on the leaderboard (mostly cheating by looking up the actual answers), but rather to verify the gap between feature engineering, cross-validation, and the leaderboard under the premise of being reproducible and explainable.

Challenge

  • Early Tier 1 features + switching models (RF → CatBoost → voting) could only push the LB to ~0.78, while CV was often higher than the LB.
  • Bruteforce OneHot encoding of Deck or fitting imputation statistics outside of CV would lead to overly optimistic OOF scores and LB drop-offs.
  • Optuna boosted CV to 0.847, but the Public LB actually dropped to 0.794.
  • In Notebook porting, if scaler/encoder were not fitted separately, CV could differ by ~7%.

Solution

Single entry point train.py, modular features_*.py:

data/train.csv, test.csv
    → features_kaggle815.py   (Step 5 — 815 notebook features)
    → features_geeky837b.py   (Step 7b — strict ipynb porting)
    → CatBoost / RF + StratifiedKFold
    → submission_step_blend.csv (default: average of two survival probabilities)
Submission FileMethodCVPublic LB
submission_step5.csv815 notebook FE + CatBoost0.8240.81578
submission_step7b.csvStrict Geeky porting + RF0.8360.81578
submission_step_blend.csvAverage of two probabilities0.8330.81578

Public demonstration primarily focuses on Step 5; the blend is an ensemble experiment (it changed 6 hard labels compared to Step 5, but the score remained the same—correct and incorrect changes offset each other).

Quick Start

pip install -r requirements.txt
python train.py              # Default: blend
python train.py --step 5     # CatBoost single model
python train.py --step 7b    # RF strict porting

Impact

  • Public LB 0.81578, an improvement of about 5 percentage points over the gender baseline (~0.765).
  • The mature feature recipe of Step 5 alone raised the LB from ~0.782 to 0.816 (+3.3%).
  • Three independent pipelines achieved the same score on the public leaderboard, verifying that “when the recipe is right, switching the model has a low marginal benefit.”

Key Lessons

  1. Feature Engineering > Tuning — A mature FE recipe immediately widens the gap; Optuna is prone to overfitting the folds on small data.
  2. High CV ≠ High LB — The sample split of 891 vs 418 often decouples OOF from the leaderboard.
  3. Strict Reproduction — Whether scaler/encoder are fitted separately can result in a ~7% difference in CV.
  4. Knowing When to Stop — 0.81578 is practically the ceiling for honest solutions.

Further Reading

For speaking invitations, internal engineering sessions, or architecture exchange, see the topics and public work I can bring into the conversation.

Speaking & contact