Skip to main content

Command Palette

Search for a command to run...

Feature Engineering

Predicting Titanic Survival: A Machine Learning Journey from Raw Data to Insights

Updated
•3 min read•View as Markdown
Feature Engineering

Can we predict who survived the Titanic disaster? Let's find out!

Introduction

Why Predict Titanic Survival?

The Titanic disaster remains one of history's most tragic maritime accidents. But beyond the human stories, it's also a fascinating dataset for machine learning beginners. In this walkthrough, we'll:

  • Load and explore the Titanic dataset

  • Clean and preprocess the data

  • Engineer meaningful features

  • Train and compare models

  • Optimize our best model

No prior ML experience needed, I'll explain everything step by step!

Step 1: Loading the Data – Meet Our Passengers

We start by importing our tools and loading the data directly from GitHub:

import pandas as pd

# Load the dataset
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
titanic = pd.read_csv(url)

# Peek at the first 5 passengers
print(titanic.head())

What we're looking at:

  • Each row represents a passenger

  • Columns include age, ticket class, fare, and most importantly, whether they survived.

Step 2: Selecting the Right Features

Not all data is useful! We'll focus on:

features = ['Pclass', 'Sex', 'Age', 'Fare', 'Embarked', 'Survived']
titanic = titanic[features]

Why these?

  • Pclass: 1st class passengers had higher survival rates

  • Sex: "Women and children first" policy mattered

  • Age: Children were prioritized

  • Fare: Correlates with class and survival

  • Embarked: Port of boarding might reveal patterns

Step 3: Handling Missing Data - Filling the Gaps

Real-world data is messy. Let's clean it up:

# Age: Fill missing values with median age
titanic['Age'].fillna(titanic['Age'].median(), inplace=True)

# Embarked: Fill missing ports with the most common one 
titanic['Embarked'].fillna(titanic['Embarked'].mode()[0], inplace=True)

The strategy:

  • For numbers (Age): Use median (resists outliers)

  • For categories (Embarked): Use mode (most frequent value)

Step 4: Preprocessing – Getting the Data Ready for ML

Machines need numbers, not words. We'll:

  • Scale numerical features (Age, Fare)

  • Encode categories (Sex, Embarked)

from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer

preprocessor = ColumnTransformer(
    transformers=[
        ('num', StandardScaler(), ['Age', 'Fare']),  # Scales to mean=0, std=1
        ('cat', OneHotEncoder(), ['Pclass', 'Sex', 'Embarked'])  # Converts categories to numbers
    ]
)

# Separate features from target
X = titanic.drop('Survived', axis=1)
y = titanic['Survived']

# Apply preprocessing
X_processed = preprocessor.fit_transform(X)

What this does:

  • StandardScaler: Normalizes ages and fares so they're comparable

  • OneHotEncoder: Turns "Sex" (male/female) into 0s and 1s

Step 5: Training Models – Logistic Regression vs. Random Forest

Let's test two popular algorithms:

Option 1: Logistic Regression

from sklearn.linear_model import LogisticRegression

model_lr = LogisticRegression(max_iter=1000)
scores = cross_val_score(model_lr, X_processed, y, cv=5)
print(f"Logistic Regression Accuracy: {scores.mean():.2f}") # Result: ~79% accuracy

Option 2: Random Forest

from sklearn.ensemble import RandomForestClassifier

model_rf = RandomForestClassifier(random_state=42)
scores = cross_val_score(model_rf, X_processed, y, cv=5)
print(f"Random Forest Accuracy: {scores.mean():.2f}") # Result: ~81% accuracy

The Results:

Key insight:

  • The Random Forest performs slightly better out of the box!

Step 6: Hyperparameter Tuning – Boosting Our Model

Let's squeeze out extra performance by tuning the Random Forest:

param_grid = {
    'n_estimators': [50, 100, 200],  # Number of trees
    'max_depth': [None, 10, 20],     # How deep trees can grow
    'min_samples_split': [2, 5, 10]  # Minimum samples to split a node
}

grid_search = GridSearchCV(
    RandomForestClassifier(random_state=42),
    param_grid,
    cv=5,
    n_jobs=-1
)
grid_search.fit(X_processed, y)

print(f"Best parameters: {grid_search.best_params_}")
print(f"Best accuracy: {grid_search.best_score_:.2f}")

The Result:

  • Best accuracy improved to ~83%
  • Optimal parameters:
    • max_depth: 10
    • min_samples_split: 5
    • n_estimators: 100

Conclusion.

We've built a model that can predict Titanic survival with 83% accuracy!

Key takeaways:

  • Feature engineering is crucial – The right preprocessing boosted our accuracy
  • Random Forests are powerful – They outperformed Logistic Regression
  • Hyperparameter tuning helps – Gained an extra 2% through optimization

What would you like to see next? Let me know in the comments!

Bonus: Here's the full code to practice and master Feature Engineering on GitHub: https://github.com/karmat-1/Feature-Engineering/