Titanic Dataset Exploratory Data Analysis (EDA) Report

1. Basic Information Overview

The dataset contains 891 rows (passengers) and 12 columns (features).

1.1 Data Types and Non-Null Counts:


RangeIndex: 891 entries, 0 to 890
Data columns (total 12 columns):
 #   Column       Non-Null Count  Dtype  
---  ------       --------------  -----  
 0   PassengerId  891 non-null    int64  
 1   Survived     891 non-null    int64  
 2   Pclass       891 non-null    int64  
 3   Name         891 non-null    object 
 4   Sex          891 non-null    object 
 5   Age          714 non-null    float64
 6   SibSp        891 non-null    int64  
 7   Parch        891 non-null    int64  
 8   Ticket       891 non-null    object 
 9   Fare         891 non-null    float64
 10  Cabin        204 non-null    object 
 11  Embarked     889 non-null    object 
dtypes: float64(2), int64(5), object(5)
memory usage: 83.7+ KB

1.2 Missing Values Analysis:

The following columns have missing values:

Count Percentage
Cabin 687 77.104377
Age 177 19.865320
Embarked 2 0.224467

Insights:

'Cabin' is missing in 77.1% of the data. This column might need significant feature engineering (e.g., extracting the deck) or might be dropped.

'Age' is missing in a significant portion of the data. Since age is likely important for survival prediction, imputation strategies will be necessary.

'Embarked' has a small number of missing values, which can be easily filled using the mode (most frequent port).

1.3 Summary Statistics:

Numerical Features:

PassengerId Survived Pclass Age SibSp Parch Fare
count 891.000000 891.000000 891.000000 714.000000 891.000000 891.000000 891.000000
mean 446.000000 0.383838 2.308642 29.699118 0.523008 0.381594 32.204208
std 257.353842 0.486592 0.836071 14.526497 1.102743 0.806057 49.693429
min 1.000000 0.000000 1.000000 0.420000 0.000000 0.000000 0.000000
25% 223.500000 0.000000 2.000000 20.125000 0.000000 0.000000 7.910400
50% 446.000000 0.000000 3.000000 28.000000 0.000000 0.000000 14.454200
75% 668.500000 1.000000 3.000000 38.000000 1.000000 0.000000 31.000000
max 891.000000 1.000000 3.000000 80.000000 8.000000 6.000000 512.329200

Categorical Features:

Name Sex Ticket Cabin Embarked
count 891 891 891 204 889
unique 891 2 681 147 3
top Braund, Mr. Owen Harris male 347082 B96 B98 S
freq 1 577 7 4 644

2. Univariate Analysis (Distribution of Features)

2.1 Target Variable: Survived

The overall survival rate is approximately 38.38%. The dataset is imbalanced, with significantly more passengers dying than surviving.

2.2 Passenger Class (Pclass)

The majority of passengers were in 3rd class, followed by 1st and then 2nd class.

2.3 Sex

There were significantly more male passengers than female passengers.

2.4 Age

The age distribution is slightly right-skewed. Most passengers were young adults (20s and 30s), with a notable number of infants and children.

2.5 Fare

The fare distribution is heavily right-skewed. Most tickets were inexpensive, but a few were very costly.

3. Bivariate Analysis (Features vs. Survival)

3.1 Survival Rate by Pclass

A strong correlation between Pclass and Survival is evident. 1st class passengers had a much higher survival rate (>60%), while 3rd class passengers had a very low survival rate (<25%). Socioeconomic status played a significant role.

3.2 Survival Rate by Sex

Sex is a critical predictor. Females had a very high survival rate (around 75%), while males had a very low survival rate (below 20%). This reflects the 'women and children first' policy.

3.3 Age Distribution by Survival

The violin plots show differences in age distribution. Notably, a higher proportion of infants and young children survived compared to adults.

3.4 Fare Distribution by Survival

Passengers who paid higher fares were more likely to survive. This is correlated with Pclass, as 1st class tickets were more expensive.

3.5 Survival Rate by Embarked Port

Passengers who embarked from Cherbourg (C) had the highest survival rate. This might be because Cherbourg passengers were more likely to be in 1st class.

4. Correlation Analysis

Insights from Correlation: