Introduction - Parkinson's Disease Background & Analysis¶
Parkinson’s disease is a neurological condition that gradually affects movement, speech, and daily activities. Because its early symptoms, like slight tremors or changes in voice, can be easy to overlook, diagnosing the disease early is often difficult. Detecting it sooner could make a big difference in how patients manage their symptoms and maintain their quality of life.
In this project, I’m using a clinical dataset of Parkinson’s patients to explore patterns in their symptoms and clinical measurements. My interest in this topic was sparked by a family member being diagnosed with Parkinson’s. The goal is to identify the key factors that indicate the presence or progression of the disease and to explore ways that this information could support early detection.
Key Investigative Questions¶
Which aspects of a person’s voice - pitch, tremors, or variations in loudness - seem to be most connected to Parkinson’s disease?
Using the voice and signal measurements in this dataset, is it possible to predict whether someone has Parkinson’s?
Do patterns in the vocal and signal features emerge that clearly separate healthy individuals from those with Parkinson’s?
Are there certain complexity measures of the voice, like how regular or irregular the pitch is, that can help identify the disease?
Can we figure out which features are the most important for detecting Parkinson’s early, so doctors could potentially use them for quicker diagnosis?
Clinical Parkinson's Dataset Overview¶
For this project, I’m working with the Clinical Parkinson’s Dataset from Kaggle: Clinical Parkinson Dataset.
The dataset contains information collected from patients’ voices. Patients with Parkinson’s disease often are affected by speech and vocal patterns. By analyzing these vocal characteristics, researchers can look for patterns that may help distinguish people with Parkinson’s from those without it.
There are many more columns that include different aspects of the voice. Some track pitch (average, highest, and lowest frequencies), while others measure stability, like tiny changes in pitch (jitter) or loudness (shimmer). Features such as NHR and HNR tell us how clear or noisy the voice sounds. More advanced measures like RPDE, DFA, spread values, D2, and PPE summarize how complex, irregular, or unpredictable the voice signal is.
To better understand the visualizations and analyses, it’s helpful to know what each feature represents. The dataset includes measurements of voice, signal, and clinical information that are linked to Parkinson’s disease. Below are simple explanations for the features used, so anyone — regardless of technical background — can understand!
Jitter Measures (the pitch variation)¶
jitter_percent– How much the pitch varies from cycle to cycle (higher = more unstable).jitter_abs– Absolute amount of pitch variation.jitter_rap– Short term pitch variation.jitter_ppq– Pitch variation over five cycles.jitter_ddp– Difference of differences in pitch variation.
Shimmer Measures (loudness variation)¶
shimmer– Cycle variation in loudness.shimmer_db– Loudness variation in decibels.shimmer_apq3– The average amplitude variation over 3 cycles.shimmer_apq5– The average amplitude variation over 5 cycles.shimmer_apq– Average amplitude variation overall.shimmer_dda– Change of differences in amplitude variation.
Noise & Signal Quality Measures¶
nhr– Noise to harmonics ratio (higher = more noise in voice).hnr– Harmonics to noise ratio (higher = cleaner voice signal).
Complexity & Irregularity Measures¶
rpde– Measures how unpredictable or irregular the voice sounds.dfa– Shows patterns in pitch over time, like whether fluctuations are consistent or random.ppe– Measures how unpredictable the pitch is from one cycle to the next.
Other Vocal Features¶
spread_1– Frequency spread of the voice (1st dimension).spread_2– Frequency spread of the voice (2nd dimension).detrended_fluctuation– Overall fluctuation of pitch after removing trends.
Target Variable¶
parkinson_status– 0 = Healthy, 1 = Parkinson’s disease
Data Preprocessing¶
The dataset we used was already fairly clean. Before diving into analysis and visualizations, I performed some basic checks to make sure everything was ready to use:
- Checked for missing values: Luckily, all columns had complete data, so no filling or imputation was needed.
- Checked data types: Ensured numeric columns were actually numeric and categorical variables were correctly formatted.
- Handled duplicates: The dataset contained repeated recordings for some patients. These are expected and valid - so I kept them.
- Verified target variable: Made sure the
parkinson_statuscolumn was converted to integers (0 = Healthy, 1 = Parkinson’s) to make analysis easier.
Overall there were only small adjustments needed because the data was already well prepared. The dataset was ready for visualizations and modeling.
import pandas as pd
df = pd.read_csv("/Users/chasepatterson/Library/Mobile Documents/com~apple~CloudDocs/parkinson_cleaned.csv")
df.head() # Checks for a proper read out of the data from csv file
| recording_id | fundamental_freq_hz | max_freq_hz | min_freq_hz | jitter_percent | jitter_abs | jitter_rap | jitter_ppq | jitter_ddp | shimmer | ... | spread_1 | spread_2 | detrended_fluctuation | ppe | subject_id | age | gender | test_time | motor_updrs_score | total_updrs_score | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | phon_R01_S01_1 | 119.992 | 157.302 | 74.997 | 0.00784 | 0.00007 | 0.0037 | 0.00554 | 0.01109 | 0.04374 | ... | -4.813031 | 0.266482 | 2.301442 | 0.284654 | 1 | 72 | Female | 5.6431 | 28.199 | 34.398 |
| 1 | phon_R01_S01_1 | 119.992 | 157.302 | 74.997 | 0.00784 | 0.00007 | 0.0037 | 0.00554 | 0.01109 | 0.04374 | ... | -4.813031 | 0.266482 | 2.301442 | 0.284654 | 1 | 72 | Female | 12.6660 | 28.447 | 34.894 |
| 2 | phon_R01_S01_1 | 119.992 | 157.302 | 74.997 | 0.00784 | 0.00007 | 0.0037 | 0.00554 | 0.01109 | 0.04374 | ... | -4.813031 | 0.266482 | 2.301442 | 0.284654 | 1 | 72 | Female | 19.6810 | 28.695 | 35.389 |
| 3 | phon_R01_S01_1 | 119.992 | 157.302 | 74.997 | 0.00784 | 0.00007 | 0.0037 | 0.00554 | 0.01109 | 0.04374 | ... | -4.813031 | 0.266482 | 2.301442 | 0.284654 | 1 | 72 | Female | 25.6470 | 28.905 | 35.810 |
| 4 | phon_R01_S01_1 | 119.992 | 157.302 | 74.997 | 0.00784 | 0.00007 | 0.0037 | 0.00554 | 0.01109 | 0.04374 | ... | -4.813031 | 0.266482 | 2.301442 | 0.284654 | 1 | 72 | Female | 33.6420 | 29.187 | 36.375 |
5 rows × 30 columns
# The Imports and the Global settings for all visualizations and normalizations - for frequencies
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (confusion_matrix, classification_report, roc_curve, auc,
RocCurveDisplay, PrecisionRecallDisplay, accuracy_score)
from sklearn.inspection import permutation_importance
from sklearn.pipeline import Pipeline
import os
print(df.columns.tolist()) #Confirms the column names for visualizations
['recording_id', 'fundamental_freq_hz', 'max_freq_hz', 'min_freq_hz', 'jitter_percent', 'jitter_abs', 'jitter_rap', 'jitter_ppq', 'jitter_ddp', 'shimmer', 'shimmer_db', 'shimmer_apq3', 'shimmer_apq5', 'shimmer_apq', 'shimmer_dda', 'nhr', 'hnr', 'parkinson_status', 'rpde', 'dfa', 'spread_1', 'spread_2', 'detrended_fluctuation', 'ppe', 'subject_id', 'age', 'gender', 'test_time', 'motor_updrs_score', 'total_updrs_score']
df = pd.read_csv("/Users/chasepatterson/Library/Mobile Documents/com~apple~CloudDocs/parkinson_cleaned.csv")
# Make sure the target variable is integer type for the analysis (0 for healthy and 1 for Parkinson's)
df['parkinson_status'] = df['parkinson_status'].astype(int)
# Counts how many healthy vs Parkinson's samples are in the dataset
counts = df['parkinson_status'].value_counts().sort_index()
print("Counts of each class (0=Healthy, 1=Parkinson's):")
print(counts)
# Shows the total number of samples in the dataset
print("\nTotal samples:", len(df))
Counts of each class (0=Healthy, 1=Parkinson's): parkinson_status 0 4290 1 19551 Name: count, dtype: int64 Total samples: 23841
# Checks for any duplicate rows
num_duplicates = df.duplicated().sum()
print(f"Number of duplicate rows: {num_duplicates}")
# I'm checking any for missing values
null_counts = df.isnull().sum()
print("\nNumber of missing values-")
print(null_counts[null_counts > 0] if null_counts.any() else "No missing values")
# Quick overview of data types and completeness
print("\nData types and the non null counts:")
print(df.info())
# Basic descriptive statistics for numeric columns
print("\nBasic statistics for the numeric columns:")
print(df.describe())
# Confirms unique values in the target variable
print("\nUnique values in parkinson_status:", df['parkinson_status'].unique())
print("Counts per class-")
print(df['parkinson_status'].value_counts())
Number of duplicate rows: 12822
Number of missing values-
No missing values
Data types and the non null counts:
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 23841 entries, 0 to 23840
Data columns (total 30 columns):
# Column Non-Null Count Dtype
--- ------ -------------- -----
0 recording_id 23841 non-null object
1 fundamental_freq_hz 23841 non-null float64
2 max_freq_hz 23841 non-null float64
3 min_freq_hz 23841 non-null float64
4 jitter_percent 23841 non-null float64
5 jitter_abs 23841 non-null float64
6 jitter_rap 23841 non-null float64
7 jitter_ppq 23841 non-null float64
8 jitter_ddp 23841 non-null float64
9 shimmer 23841 non-null float64
10 shimmer_db 23841 non-null float64
11 shimmer_apq3 23841 non-null float64
12 shimmer_apq5 23841 non-null float64
13 shimmer_apq 23841 non-null float64
14 shimmer_dda 23841 non-null float64
15 nhr 23841 non-null float64
16 hnr 23841 non-null float64
17 parkinson_status 23841 non-null int64
18 rpde 23841 non-null float64
19 dfa 23841 non-null float64
20 spread_1 23841 non-null float64
21 spread_2 23841 non-null float64
22 detrended_fluctuation 23841 non-null float64
23 ppe 23841 non-null float64
24 subject_id 23841 non-null int64
25 age 23841 non-null int64
26 gender 23841 non-null object
27 test_time 23841 non-null float64
28 motor_updrs_score 23841 non-null float64
29 total_updrs_score 23841 non-null float64
dtypes: float64(25), int64(3), object(2)
memory usage: 5.5+ MB
None
Basic statistics for the numeric columns:
fundamental_freq_hz max_freq_hz min_freq_hz jitter_percent \
count 23841.000000 23841.000000 23841.000000 23841.000000
mean 157.257589 196.077131 119.373145 0.006609
std 42.544988 84.224334 46.348230 0.005329
min 88.333000 102.145000 65.476000 0.001680
25% 119.992000 137.871000 83.961000 0.003460
50% 151.989000 189.398000 104.680000 0.005050
75% 188.620000 223.982000 147.226000 0.007610
max 260.105000 588.518000 239.170000 0.033160
jitter_abs jitter_rap jitter_ppq jitter_ddp shimmer \
count 23841.000000 23841.000000 23841.000000 23841.000000 23841.000000
mean 0.000046 0.003552 0.003680 0.010656 0.031626
std 0.000038 0.003271 0.003059 0.009812 0.020023
min 0.000007 0.000680 0.000920 0.002040 0.009540
25% 0.000020 0.001660 0.001820 0.004980 0.017060
50% 0.000040 0.002600 0.002830 0.007800 0.024980
75% 0.000060 0.003980 0.004220 0.011930 0.040240
max 0.000260 0.021440 0.019580 0.064330 0.119080
shimmer_db ... dfa spread_1 spread_2 \
count 23841.000000 ... 23841.000000 23841.000000 23841.000000
mean 0.301494 ... 0.718128 -5.598732 0.234541
std 0.209046 ... 0.056515 1.167403 0.085984
min 0.085000 ... 0.574282 -7.964984 0.006274
25% 0.154000 ... 0.676023 -6.471427 0.177551
50% 0.228000 ... 0.722085 -5.557447 0.233070
75% 0.370000 ... 0.762726 -4.813031 0.299111
max 1.302000 ... 0.825288 -2.434031 0.450493
detrended_fluctuation ppe subject_id age \
count 23841.000000 23841.000000 23841.000000 23841.000000
mean 2.405992 0.214835 20.453379 64.639612
std 0.393156 0.096557 12.137185 8.572230
min 1.423287 0.044539 1.000000 36.000000
25% 2.108873 0.136390 8.000000 58.000000
50% 2.398422 0.214075 21.000000 66.000000
75% 2.642276 0.268144 32.000000 72.000000
max 3.671155 0.527367 42.000000 76.000000
test_time motor_updrs_score total_updrs_score
count 23841.000000 23841.000000 23841.000000
mean 91.780388 21.364941 29.374836
std 52.972240 8.810362 12.182424
min -4.262500 5.037700 7.000000
25% 45.800000 13.256000 19.000000
50% 89.637000 21.931000 28.634000
75% 137.780000 28.415000 39.088000
max 202.430000 39.511000 54.992000
[8 rows x 28 columns]
Unique values in parkinson_status: [1 0]
Counts per class-
parkinson_status
1 19551
0 4290
Name: count, dtype: int64
# Counts the total duplicate rows across all the columns
num_duplicates = df.duplicated().sum()
print(f"Total duplicate rows- {num_duplicates}")
# Counts the duplicates ignoring the identifier columns such as the recording_id or the subject_id.
feature_cols = df.drop(columns=['recording_id', 'subject_id', 'test_time']).columns
true_duplicates = df.duplicated(subset=feature_cols).sum()
print(f"Number of rows with identical features (ignoring IDs): {true_duplicates}")
# My explanation of what was read out from the code above
print("\nNote: These duplicates are expected because the dataset includes multiple recordings per subject.")
print("They represent legitimate repeated measurements - so we'll keep all rows for analysis.")
Total duplicate rows- 12822 Number of rows with identical features (ignoring IDs): 18495 Note: These duplicates are expected because the dataset includes multiple recordings per subject. They represent legitimate repeated measurements - so we'll keep all rows for analysis.
Data Visualizations & Understanding¶
# Visualizing the number of healthy compared to the Parkinson's samples in the dataset
plt.figure(figsize=(6,4))
sns.countplot(x='parkinson_status', data=df)
plt.xticks([0,1], ['Healthy (0)', "Parkinson's (1)"])
plt.title('Class Distribution')
plt.show()
# Showing how key vocal features like the pitch and voice quality differ between healthy individuals and those with Parkinson's
key_feats = [c for c in df.columns if any(k in c.lower() for k in ['jitter','shimmer','nhr','hnr'])][:8]
# Plots the violon plots for a more appealing visualizaion compared to the standard points.
num_feats = len(key_feats)
cols = 2
rows = (num_feats + 1) // cols
plt.figure(figsize=(12, 4*rows))
for i, feat in enumerate(key_feats, 1):
plt.subplot(rows, cols, i)
sns.violinplot(x='parkinson_status', y=feat, data=df, hue='parkinson_status', palette=['green','red'], legend=False)
plt.xticks([0,1], ['Healthy', "Parkinson's"])
plt.title(feat)
plt.xlabel('')
plt.ylabel('')
plt.tight_layout()
plt.suptitle('Vocal Feature Distributions by Parkinson Status', fontsize=16, y=1.02)
plt.show()
# I wanted to ignore Seaborn Future Warnings for a cleaner output so just the plots would show :)
warnings.simplefilter(action='ignore', category=FutureWarning)
# Showing which vocal and signal features are most strongly related to Parkinson's disease
# by computing correlations with the Parkinson's status and visualizing the top correlations
df['parkinson_status'] = df['parkinson_status'].astype(int)
key_features = [
'jitter_percent', 'jitter_abs', 'jitter_rap', 'jitter_ppq', 'jitter_ddp',
'shimmer', 'shimmer_db', 'shimmer_apq3', 'shimmer_apq5', 'shimmer_apq',
'shimmer_dda', 'nhr', 'hnr', 'rpde', 'dfa', 'spread_1', 'spread_2',
'detrended_fluctuation', 'ppe'
]
corr = df[key_features + ['parkinson_status']].corr()
target_corr = corr['parkinson_status'].drop('parkinson_status').sort_values(ascending=False)
plt.figure(figsize=(10,6))
sns.barplot(x=target_corr.values, y=target_corr.index, palette='Reds_r')
plt.xlabel("Correlation with Parkinson's Status")
plt.title("Top Vocal & Signal Features Correlated with Parkinson's Disease")
plt.xlim(0, 1)
plt.tight_layout()
plt.show()
# Remove columns that aren’t useful for analysis which are the IDs and target
# because PCA only works on numeric features
features = df.drop(columns=['recording_id','subject_id','test_time','parkinson_status'])
numeric_features = features.select_dtypes(include='number')
# Standardize all the numbers so features on different scales don't overpower each other
X_scaled = StandardScaler().fit_transform(numeric_features)
y = df['parkinson_status'].astype(int)
# Reduce all the many voice measurements down to 2 main dimensions
# so we can plotting them can be easier and more clean
pca = PCA(n_components=2, random_state=42)
X_pca = pca.fit_transform(X_scaled)
plot_df = pd.DataFrame(X_pca, columns=['PC1','PC2'])
plot_df['Status'] = y
# Create a scatter plot to visualize the main patterns
# Green = Healthy, Red = Parkinson's
plt.figure(figsize=(8,6))
sns.scatterplot(
x='PC1', y='PC2',
hue='Status',
data=plot_df,
palette={0:'green', 1:'red'},
alpha=0.6,
s=40
)
# Label axes with how much of the original information they capture
plt.xlabel(f"PC1 ({pca.explained_variance_ratio_[0]*100:.1f}% variance)")
plt.ylabel(f"PC2 ({pca.explained_variance_ratio_[1]*100:.1f}% variance)")
# Title and legend to make the chart easy to interpret for all readers
plt.title("Principal Component Analysis - 2 Components - Healthy vs Parkinson's")
plt.legend(title='Status', labels=['Healthy (0)', "Parkinson's (1)"])
plt.grid(True)
plt.tight_layout()
plt.show()
# Make sure the parkinson status works with the colors
df['parkinson_status'] = df['parkinson_status'].astype(str)
# These are the complexity measures that capture how irregular or unpredictable the voice is
complexity_feats = ['rpde', 'dfa', 'ppe']
plt.figure(figsize=(12,5))
for i, feat in enumerate(complexity_feats):
plt.subplot(1, len(complexity_feats), i+1)
# Ensure each feature is numeric to prepare for any formatting issues
df[feat] = pd.to_numeric(df[feat], errors='coerce')
# Create a boxplot to compare healthy vs Parkinson's for this feature
sns.boxplot(
x='parkinson_status',
y=feat,
data=df,
palette={'0':'green', '1':'red'} # green = healthy, red = Parkinson's
)
#reduce clutter for graphs
plt.xlabel('')
plt.ylabel(feat)
plt.title(feat)
plt.suptitle("Complexity Measures by Parkinson's vs Healthier Status")
plt.tight_layout()
plt.show()
Story Telling & Insights¶
As I explored the Clinical Parkinson’s Dataset, a few interesting things stood out:
1. Dataset distribution
The first graph shows how many samples are from healthy people versus people with Parkinson’s. There are more samples from people with Parkinson’s, which helps when looking for patterns.
2. Voice patterns
Features like jitter and shimmer show that people with Parkinson’s tend to have higher and more spread-out values. Healthy people’s voices are more consistent, which makes sense because Parkinson’s affects voice stability.
3. Important features
Features such as spread_1, spread_2, RPDE, PPE, and jitter_abs show the strongest connection to Parkinson’s. Other features have weaker relationships but still provide insight.
4. Principal Component Analysis (PCA)
The PCA plot summarizes all features into two main components. PC1 explains about 53.8% of the variation and PC2 explains about 11%. The plot shows that healthy and Parkinson’s samples begin to form separate clusters, highlighting meaningful differences in voice features.
5. Complexity measures
RPDE, DFA, and PPE measure how irregular or unpredictable the voice is. People with Parkinson’s generally have higher values here which can mean their voices are less steady than healthy people.
What I learned:
These visualizations confirm that certain voice features, especially jitter, shimmer, spread_1, spread_2, RPDE, and PPE, are strongly linked to Parkinson’s. They help identify which aspects of speech are affected and could support earlier detection of the disease!!
Impact Section¶
This analysis provides useful insights but does not capture the full picture! Below are some key considerations to keep in mind when interpreting the results and drawing conclusions:
Potential benefits:
- The visualizations and feature analysis could help researchers or doctors better understand which voice characteristics are affected by Parkinson’s.
- Identifying key features like jitter, shimmer, RPDE, and PPE could support earlier detection or monitoring of the disease.
Potential risks or harm:
- This analysis is based only on the dataset provided, which may not represent the full diversity of patients. For example, age, gender, accents, or other health conditions might influence voice features but aren’t fully captured from the data I personally visualized and studied.
- Misinterpretation of the results could lead to overconfidence or other cognitive biases in diagnosing Parkinson’s from voice alone. THIS SHOULD NOT BE USED FOR OFFICIAL MEDICAL EVALUATION
- Visualizations might unintentionally exaggerate differences if the audience doesn’t understand the context and is new to this feild of study which can possibly lead to more bias or misunderstanding.
Missing data or perspectives:
- There could be other important clinical or lifestyle factors that influence voice patterns but are not included in the dataset.
- Longitudinal tracking of voice changes over time could/would provide better understanding and prediction of voice fequencies.
Conclusion:
My aim for this project is to demonstrate the potential of using voice features for Parkinson’s research. Careful interpretation and awareness of dataset limitations are of most importance to avoid any misuse, incorrect conclusions, or harm. I hope you found this report interesting and insighful to hopefully spark or deepen your understanding of others around you who may have been diagnosed with Parkinson's disease. Thank you for reading! :)
References¶
DevPress. (2022, December 10). Python Matplotlib: How to create histogram plot in Python. Hive. Retrieved from https://hive.blog/hive-122108/@devpress/python-matplotlib-how-to-create-histogram-plot-in-python
Mayo Clinic Staff. (2024, September 27). Parkinson's disease - Symptoms and causes. Mayo Clinic. Retrieved from https://www.mayoclinic.org/diseases-conditions/parkinsons-disease/symptoms-causes/syc-20376055
Purohit, R. (2023). Clinical Parkinson’s Dataset. Kaggle. Retrieved from https://www.kaggle.com/datasets/rohanpurohit0705/clinical-parkinson-dataset
University of California, Berkeley. (n.d.). Data visualization with Python. Berkeley Data Science Education Program. Retrieved September 12, 2025, from [https://guides.lib.berkeley.edu/data-visualization]
Tech With Tim. (2020, September 15). Python Data Visualization Full Course – Data Visualization with Python [Video]. YouTube. Retrieved from https://www.youtube.com/watch?v=q68Qundmans
Data Science Tutorials. (2022, March 8). Master Python Plotly in 1.5 Hours: From Basics to Advanced [Video]. YouTube. Retrieved from https://www.youtube.com/watch?v=W_qQTKupZpY