Feature Engineering
Predicting Titanic Survival: A Machine Learning Journey from Raw Data to Insights

Can we predict who survived the Titanic disaster? Let's find out!
Introduction
Why Predict Titanic Survival?
The Titanic disaster remains one of history's most tragic maritime accidents. But beyond the human stories, it's also a fascinating dataset for machine learning beginners. In this walkthrough, we'll:
Load and explore the Titanic dataset
Clean and preprocess the data
Engineer meaningful features
Train and compare models
Optimize our best model
No prior ML experience needed, I'll explain everything step by step!
Step 1: Loading the Data – Meet Our Passengers
We start by importing our tools and loading the data directly from GitHub:
import pandas as pd
# Load the dataset
url = "https://raw.githubusercontent.com/datasciencedojo/datasets/master/titanic.csv"
titanic = pd.read_csv(url)
# Peek at the first 5 passengers
print(titanic.head())
What we're looking at:
Each row represents a passenger
Columns include age, ticket class, fare, and most importantly, whether they survived.

Step 2: Selecting the Right Features
Not all data is useful! We'll focus on:
features = ['Pclass', 'Sex', 'Age', 'Fare', 'Embarked', 'Survived']
titanic = titanic[features]
Why these?
Pclass: 1st class passengers had higher survival rates
Sex: "Women and children first" policy mattered
Age: Children were prioritized
Fare: Correlates with class and survival
Embarked: Port of boarding might reveal patterns
Step 3: Handling Missing Data - Filling the Gaps
Real-world data is messy. Let's clean it up:
# Age: Fill missing values with median age
titanic['Age'].fillna(titanic['Age'].median(), inplace=True)
# Embarked: Fill missing ports with the most common one
titanic['Embarked'].fillna(titanic['Embarked'].mode()[0], inplace=True)
The strategy:
For numbers (Age): Use median (resists outliers)
For categories (Embarked): Use mode (most frequent value)
Step 4: Preprocessing – Getting the Data Ready for ML
Machines need numbers, not words. We'll:
Scale numerical features (Age, Fare)
Encode categories (Sex, Embarked)
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.compose import ColumnTransformer
preprocessor = ColumnTransformer(
transformers=[
('num', StandardScaler(), ['Age', 'Fare']), # Scales to mean=0, std=1
('cat', OneHotEncoder(), ['Pclass', 'Sex', 'Embarked']) # Converts categories to numbers
]
)
# Separate features from target
X = titanic.drop('Survived', axis=1)
y = titanic['Survived']
# Apply preprocessing
X_processed = preprocessor.fit_transform(X)
What this does:
StandardScaler: Normalizes ages and fares so they're comparable
OneHotEncoder: Turns "Sex" (male/female) into 0s and 1s
Step 5: Training Models – Logistic Regression vs. Random Forest
Let's test two popular algorithms:
Option 1: Logistic Regression
from sklearn.linear_model import LogisticRegression
model_lr = LogisticRegression(max_iter=1000)
scores = cross_val_score(model_lr, X_processed, y, cv=5)
print(f"Logistic Regression Accuracy: {scores.mean():.2f}") # Result: ~79% accuracy
Option 2: Random Forest
from sklearn.ensemble import RandomForestClassifier
model_rf = RandomForestClassifier(random_state=42)
scores = cross_val_score(model_rf, X_processed, y, cv=5)
print(f"Random Forest Accuracy: {scores.mean():.2f}") # Result: ~81% accuracy
The Results:

Key insight:
- The Random Forest performs slightly better out of the box!
Step 6: Hyperparameter Tuning – Boosting Our Model
Let's squeeze out extra performance by tuning the Random Forest:
param_grid = {
'n_estimators': [50, 100, 200], # Number of trees
'max_depth': [None, 10, 20], # How deep trees can grow
'min_samples_split': [2, 5, 10] # Minimum samples to split a node
}
grid_search = GridSearchCV(
RandomForestClassifier(random_state=42),
param_grid,
cv=5,
n_jobs=-1
)
grid_search.fit(X_processed, y)
print(f"Best parameters: {grid_search.best_params_}")
print(f"Best accuracy: {grid_search.best_score_:.2f}")
The Result:
- Best accuracy improved to ~83%
- Optimal parameters:
- max_depth: 10
- min_samples_split: 5
- n_estimators: 100

Conclusion.
We've built a model that can predict Titanic survival with 83% accuracy!
Key takeaways:
- Feature engineering is crucial – The right preprocessing boosted our accuracy
- Random Forests are powerful – They outperformed Logistic Regression
- Hyperparameter tuning helps – Gained an extra 2% through optimization
What would you like to see next? Let me know in the comments!
Bonus: Here's the full code to practice and master Feature Engineering on GitHub: https://github.com/karmat-1/Feature-Engineering/



