Introduction - Parkinson's Disease Background & Analysis¶

Parkinson’s disease is a neurological condition that gradually affects movement, speech, and daily activities. Because its early symptoms, like slight tremors or changes in voice, can be easy to overlook, diagnosing the disease early is often difficult. Detecting it sooner could make a big difference in how patients manage their symptoms and maintain their quality of life.

In this project, I’m using a clinical dataset of Parkinson’s patients to explore patterns in their symptoms and clinical measurements. My interest in this topic was sparked by a family member being diagnosed with Parkinson’s. The goal is to identify the key factors that indicate the presence or progression of the disease and to explore ways that this information could support early detection.

Key Investigative Questions¶

  1. Which aspects of a person’s voice - pitch, tremors, or variations in loudness - seem to be most connected to Parkinson’s disease?

  2. Using the voice and signal measurements in this dataset, is it possible to predict whether someone has Parkinson’s?

  3. Do patterns in the vocal and signal features emerge that clearly separate healthy individuals from those with Parkinson’s?

  4. Are there certain complexity measures of the voice, like how regular or irregular the pitch is, that can help identify the disease?

  5. Can we figure out which features are the most important for detecting Parkinson’s early, so doctors could potentially use them for quicker diagnosis?

Clinical Parkinson's Dataset Overview¶

For this project, I’m working with the Clinical Parkinson’s Dataset from Kaggle: Clinical Parkinson Dataset.

The dataset contains information collected from patients’ voices. Patients with Parkinson’s disease often are affected by speech and vocal patterns. By analyzing these vocal characteristics, researchers can look for patterns that may help distinguish people with Parkinson’s from those without it.

There are many more columns that include different aspects of the voice. Some track pitch (average, highest, and lowest frequencies), while others measure stability, like tiny changes in pitch (jitter) or loudness (shimmer). Features such as NHR and HNR tell us how clear or noisy the voice sounds. More advanced measures like RPDE, DFA, spread values, D2, and PPE summarize how complex, irregular, or unpredictable the voice signal is.

To better understand the visualizations and analyses, it’s helpful to know what each feature represents. The dataset includes measurements of voice, signal, and clinical information that are linked to Parkinson’s disease. Below are simple explanations for the features used, so anyone — regardless of technical background — can understand!


Jitter Measures (the pitch variation)¶

  • jitter_percent – How much the pitch varies from cycle to cycle (higher = more unstable).
  • jitter_abs – Absolute amount of pitch variation.
  • jitter_rap – Short term pitch variation.
  • jitter_ppq – Pitch variation over five cycles.
  • jitter_ddp – Difference of differences in pitch variation.

Shimmer Measures (loudness variation)¶

  • shimmer – Cycle variation in loudness.
  • shimmer_db – Loudness variation in decibels.
  • shimmer_apq3 – The average amplitude variation over 3 cycles.
  • shimmer_apq5 – The average amplitude variation over 5 cycles.
  • shimmer_apq – Average amplitude variation overall.
  • shimmer_dda – Change of differences in amplitude variation.

Noise & Signal Quality Measures¶

  • nhr – Noise to harmonics ratio (higher = more noise in voice).
  • hnr – Harmonics to noise ratio (higher = cleaner voice signal).

Complexity & Irregularity Measures¶

  • rpde – Measures how unpredictable or irregular the voice sounds.
  • dfa – Shows patterns in pitch over time, like whether fluctuations are consistent or random.
  • ppe – Measures how unpredictable the pitch is from one cycle to the next.

Other Vocal Features¶

  • spread_1 – Frequency spread of the voice (1st dimension).
  • spread_2 – Frequency spread of the voice (2nd dimension).
  • detrended_fluctuation – Overall fluctuation of pitch after removing trends.

Target Variable¶

  • parkinson_status – 0 = Healthy, 1 = Parkinson’s disease

Data Preprocessing¶

The dataset we used was already fairly clean. Before diving into analysis and visualizations, I performed some basic checks to make sure everything was ready to use:

  • Checked for missing values: Luckily, all columns had complete data, so no filling or imputation was needed.
  • Checked data types: Ensured numeric columns were actually numeric and categorical variables were correctly formatted.
  • Handled duplicates: The dataset contained repeated recordings for some patients. These are expected and valid - so I kept them.
  • Verified target variable: Made sure the parkinson_status column was converted to integers (0 = Healthy, 1 = Parkinson’s) to make analysis easier.

Overall there were only small adjustments needed because the data was already well prepared. The dataset was ready for visualizations and modeling.

In [6]:
import pandas as pd
df = pd.read_csv("/Users/chasepatterson/Library/Mobile Documents/com~apple~CloudDocs/parkinson_cleaned.csv")
df.head() # Checks for a proper read out of the data from csv file
Out[6]:
recording_id fundamental_freq_hz max_freq_hz min_freq_hz jitter_percent jitter_abs jitter_rap jitter_ppq jitter_ddp shimmer ... spread_1 spread_2 detrended_fluctuation ppe subject_id age gender test_time motor_updrs_score total_updrs_score
0 phon_R01_S01_1 119.992 157.302 74.997 0.00784 0.00007 0.0037 0.00554 0.01109 0.04374 ... -4.813031 0.266482 2.301442 0.284654 1 72 Female 5.6431 28.199 34.398
1 phon_R01_S01_1 119.992 157.302 74.997 0.00784 0.00007 0.0037 0.00554 0.01109 0.04374 ... -4.813031 0.266482 2.301442 0.284654 1 72 Female 12.6660 28.447 34.894
2 phon_R01_S01_1 119.992 157.302 74.997 0.00784 0.00007 0.0037 0.00554 0.01109 0.04374 ... -4.813031 0.266482 2.301442 0.284654 1 72 Female 19.6810 28.695 35.389
3 phon_R01_S01_1 119.992 157.302 74.997 0.00784 0.00007 0.0037 0.00554 0.01109 0.04374 ... -4.813031 0.266482 2.301442 0.284654 1 72 Female 25.6470 28.905 35.810
4 phon_R01_S01_1 119.992 157.302 74.997 0.00784 0.00007 0.0037 0.00554 0.01109 0.04374 ... -4.813031 0.266482 2.301442 0.284654 1 72 Female 33.6420 29.187 36.375

5 rows × 30 columns

In [33]:
# The Imports and the Global settings for all visualizations and normalizations - for frequencies
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns 
from sklearn.model_selection import train_test_split, cross_val_score, StratifiedKFold
from sklearn.preprocessing import StandardScaler
from sklearn.decomposition import PCA
from sklearn.ensemble import RandomForestClassifier
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import (confusion_matrix, classification_report, roc_curve, auc,
                             RocCurveDisplay, PrecisionRecallDisplay, accuracy_score)
from sklearn.inspection import permutation_importance
from sklearn.pipeline import Pipeline
import os
In [10]:
print(df.columns.tolist()) #Confirms the column names for visualizations 
['recording_id', 'fundamental_freq_hz', 'max_freq_hz', 'min_freq_hz', 'jitter_percent', 'jitter_abs', 'jitter_rap', 'jitter_ppq', 'jitter_ddp', 'shimmer', 'shimmer_db', 'shimmer_apq3', 'shimmer_apq5', 'shimmer_apq', 'shimmer_dda', 'nhr', 'hnr', 'parkinson_status', 'rpde', 'dfa', 'spread_1', 'spread_2', 'detrended_fluctuation', 'ppe', 'subject_id', 'age', 'gender', 'test_time', 'motor_updrs_score', 'total_updrs_score']
In [74]:
df = pd.read_csv("/Users/chasepatterson/Library/Mobile Documents/com~apple~CloudDocs/parkinson_cleaned.csv")

# Make sure the target variable is integer type for the analysis (0 for healthy and 1 for Parkinson's)
df['parkinson_status'] = df['parkinson_status'].astype(int)

# Counts how many healthy vs Parkinson's samples are in the dataset
counts = df['parkinson_status'].value_counts().sort_index()
print("Counts of each class (0=Healthy, 1=Parkinson's):")
print(counts)

# Shows the total number of samples in the dataset
print("\nTotal samples:", len(df))
Counts of each class (0=Healthy, 1=Parkinson's):
parkinson_status
0     4290
1    19551
Name: count, dtype: int64

Total samples: 23841
In [76]:
# Checks for any duplicate rows
num_duplicates = df.duplicated().sum()
print(f"Number of duplicate rows: {num_duplicates}")

# I'm checking any for missing values
null_counts = df.isnull().sum()
print("\nNumber of missing values-")
print(null_counts[null_counts > 0] if null_counts.any() else "No missing values")

# Quick overview of data types and completeness
print("\nData types and the non null counts:")
print(df.info())

# Basic descriptive statistics for numeric columns
print("\nBasic statistics for the numeric columns:")
print(df.describe())

# Confirms unique values in the target variable
print("\nUnique values in parkinson_status:", df['parkinson_status'].unique())
print("Counts per class-")
print(df['parkinson_status'].value_counts())
Number of duplicate rows: 12822

Number of missing values-
No missing values

Data types and the non null counts:
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 23841 entries, 0 to 23840
Data columns (total 30 columns):
 #   Column                 Non-Null Count  Dtype  
---  ------                 --------------  -----  
 0   recording_id           23841 non-null  object 
 1   fundamental_freq_hz    23841 non-null  float64
 2   max_freq_hz            23841 non-null  float64
 3   min_freq_hz            23841 non-null  float64
 4   jitter_percent         23841 non-null  float64
 5   jitter_abs             23841 non-null  float64
 6   jitter_rap             23841 non-null  float64
 7   jitter_ppq             23841 non-null  float64
 8   jitter_ddp             23841 non-null  float64
 9   shimmer                23841 non-null  float64
 10  shimmer_db             23841 non-null  float64
 11  shimmer_apq3           23841 non-null  float64
 12  shimmer_apq5           23841 non-null  float64
 13  shimmer_apq            23841 non-null  float64
 14  shimmer_dda            23841 non-null  float64
 15  nhr                    23841 non-null  float64
 16  hnr                    23841 non-null  float64
 17  parkinson_status       23841 non-null  int64  
 18  rpde                   23841 non-null  float64
 19  dfa                    23841 non-null  float64
 20  spread_1               23841 non-null  float64
 21  spread_2               23841 non-null  float64
 22  detrended_fluctuation  23841 non-null  float64
 23  ppe                    23841 non-null  float64
 24  subject_id             23841 non-null  int64  
 25  age                    23841 non-null  int64  
 26  gender                 23841 non-null  object 
 27  test_time              23841 non-null  float64
 28  motor_updrs_score      23841 non-null  float64
 29  total_updrs_score      23841 non-null  float64
dtypes: float64(25), int64(3), object(2)
memory usage: 5.5+ MB
None

Basic statistics for the numeric columns:
       fundamental_freq_hz   max_freq_hz   min_freq_hz  jitter_percent  \
count         23841.000000  23841.000000  23841.000000    23841.000000   
mean            157.257589    196.077131    119.373145        0.006609   
std              42.544988     84.224334     46.348230        0.005329   
min              88.333000    102.145000     65.476000        0.001680   
25%             119.992000    137.871000     83.961000        0.003460   
50%             151.989000    189.398000    104.680000        0.005050   
75%             188.620000    223.982000    147.226000        0.007610   
max             260.105000    588.518000    239.170000        0.033160   

         jitter_abs    jitter_rap    jitter_ppq    jitter_ddp       shimmer  \
count  23841.000000  23841.000000  23841.000000  23841.000000  23841.000000   
mean       0.000046      0.003552      0.003680      0.010656      0.031626   
std        0.000038      0.003271      0.003059      0.009812      0.020023   
min        0.000007      0.000680      0.000920      0.002040      0.009540   
25%        0.000020      0.001660      0.001820      0.004980      0.017060   
50%        0.000040      0.002600      0.002830      0.007800      0.024980   
75%        0.000060      0.003980      0.004220      0.011930      0.040240   
max        0.000260      0.021440      0.019580      0.064330      0.119080   

         shimmer_db  ...           dfa      spread_1      spread_2  \
count  23841.000000  ...  23841.000000  23841.000000  23841.000000   
mean       0.301494  ...      0.718128     -5.598732      0.234541   
std        0.209046  ...      0.056515      1.167403      0.085984   
min        0.085000  ...      0.574282     -7.964984      0.006274   
25%        0.154000  ...      0.676023     -6.471427      0.177551   
50%        0.228000  ...      0.722085     -5.557447      0.233070   
75%        0.370000  ...      0.762726     -4.813031      0.299111   
max        1.302000  ...      0.825288     -2.434031      0.450493   

       detrended_fluctuation           ppe    subject_id           age  \
count           23841.000000  23841.000000  23841.000000  23841.000000   
mean                2.405992      0.214835     20.453379     64.639612   
std                 0.393156      0.096557     12.137185      8.572230   
min                 1.423287      0.044539      1.000000     36.000000   
25%                 2.108873      0.136390      8.000000     58.000000   
50%                 2.398422      0.214075     21.000000     66.000000   
75%                 2.642276      0.268144     32.000000     72.000000   
max                 3.671155      0.527367     42.000000     76.000000   

          test_time  motor_updrs_score  total_updrs_score  
count  23841.000000       23841.000000       23841.000000  
mean      91.780388          21.364941          29.374836  
std       52.972240           8.810362          12.182424  
min       -4.262500           5.037700           7.000000  
25%       45.800000          13.256000          19.000000  
50%       89.637000          21.931000          28.634000  
75%      137.780000          28.415000          39.088000  
max      202.430000          39.511000          54.992000  

[8 rows x 28 columns]

Unique values in parkinson_status: [1 0]
Counts per class-
parkinson_status
1    19551
0     4290
Name: count, dtype: int64
In [77]:
# Counts the total duplicate rows across all the columns
num_duplicates = df.duplicated().sum()
print(f"Total duplicate rows- {num_duplicates}")

# Counts the duplicates ignoring the identifier columns  such as the recording_id or the subject_id.
feature_cols = df.drop(columns=['recording_id', 'subject_id', 'test_time']).columns
true_duplicates = df.duplicated(subset=feature_cols).sum()
print(f"Number of rows with identical features (ignoring IDs): {true_duplicates}")

# My explanation of what was read out from the code above
print("\nNote: These duplicates are expected because the dataset includes multiple recordings per subject.")
print("They represent legitimate repeated measurements - so we'll keep all rows for analysis.")
Total duplicate rows- 12822
Number of rows with identical features (ignoring IDs): 18495

Note: These duplicates are expected because the dataset includes multiple recordings per subject.
They represent legitimate repeated measurements - so we'll keep all rows for analysis.

Data Visualizations & Understanding¶

In [78]:
# Visualizing the number of healthy compared to the Parkinson's samples in the dataset
plt.figure(figsize=(6,4))
sns.countplot(x='parkinson_status', data=df)
plt.xticks([0,1], ['Healthy (0)', "Parkinson's (1)"])
plt.title('Class Distribution')
plt.show()
No description has been provided for this image
In [80]:
# Showing how key vocal features like the pitch and voice quality differ between healthy individuals and those with Parkinson's
key_feats = [c for c in df.columns if any(k in c.lower() for k in ['jitter','shimmer','nhr','hnr'])][:8]

# Plots the violon plots for a more appealing visualizaion compared to the standard points.
num_feats = len(key_feats)
cols = 2
rows = (num_feats + 1) // cols
plt.figure(figsize=(12, 4*rows))

for i, feat in enumerate(key_feats, 1):
    plt.subplot(rows, cols, i)
    sns.violinplot(x='parkinson_status', y=feat, data=df, hue='parkinson_status', palette=['green','red'], legend=False)
    plt.xticks([0,1], ['Healthy', "Parkinson's"])
    plt.title(feat)
    plt.xlabel('')
    plt.ylabel('')

plt.tight_layout()
plt.suptitle('Vocal Feature Distributions by Parkinson Status', fontsize=16, y=1.02)
plt.show()

# I wanted to ignore Seaborn Future Warnings for a cleaner output so just the plots would show :)
warnings.simplefilter(action='ignore', category=FutureWarning)
No description has been provided for this image
In [82]:
# Showing which vocal and signal features are most strongly related to Parkinson's disease
# by computing correlations with the Parkinson's status and visualizing the top correlations

df['parkinson_status'] = df['parkinson_status'].astype(int)

key_features = [
    'jitter_percent', 'jitter_abs', 'jitter_rap', 'jitter_ppq', 'jitter_ddp',
    'shimmer', 'shimmer_db', 'shimmer_apq3', 'shimmer_apq5', 'shimmer_apq',
    'shimmer_dda', 'nhr', 'hnr', 'rpde', 'dfa', 'spread_1', 'spread_2', 
    'detrended_fluctuation', 'ppe'
]

corr = df[key_features + ['parkinson_status']].corr()
target_corr = corr['parkinson_status'].drop('parkinson_status').sort_values(ascending=False)

plt.figure(figsize=(10,6))
sns.barplot(x=target_corr.values, y=target_corr.index, palette='Reds_r')
plt.xlabel("Correlation with Parkinson's Status")
plt.title("Top Vocal & Signal Features Correlated with Parkinson's Disease")
plt.xlim(0, 1)
plt.tight_layout()
plt.show()
No description has been provided for this image
In [93]:
# Remove columns that aren’t useful for analysis which are the IDs and target
# because PCA only works on numeric features
features = df.drop(columns=['recording_id','subject_id','test_time','parkinson_status'])
numeric_features = features.select_dtypes(include='number')

# Standardize all the numbers so features on different scales don't overpower each other
X_scaled = StandardScaler().fit_transform(numeric_features)
y = df['parkinson_status'].astype(int)

# Reduce all the many voice measurements down to 2 main dimensions 
# so we can plotting them can be easier and more clean 
pca = PCA(n_components=2, random_state=42)
X_pca = pca.fit_transform(X_scaled)
plot_df = pd.DataFrame(X_pca, columns=['PC1','PC2'])
plot_df['Status'] = y

# Create a scatter plot to visualize the main patterns
# Green = Healthy, Red = Parkinson's
plt.figure(figsize=(8,6))
sns.scatterplot(
    x='PC1', y='PC2', 
    hue='Status', 
    data=plot_df, 
    palette={0:'green', 1:'red'}, 
    alpha=0.6, 
    s=40
)

# Label axes with how much of the original information they capture
plt.xlabel(f"PC1 ({pca.explained_variance_ratio_[0]*100:.1f}% variance)")
plt.ylabel(f"PC2 ({pca.explained_variance_ratio_[1]*100:.1f}% variance)")

# Title and legend to make the chart easy to interpret for all readers
plt.title("Principal Component Analysis - 2 Components - Healthy vs Parkinson's")
plt.legend(title='Status', labels=['Healthy (0)', "Parkinson's (1)"])
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image
In [94]:
# Make sure the parkinson status works with the colors
df['parkinson_status'] = df['parkinson_status'].astype(str)

# These are the complexity measures that capture how irregular or unpredictable the voice is
complexity_feats = ['rpde', 'dfa', 'ppe']

plt.figure(figsize=(12,5))

for i, feat in enumerate(complexity_feats):
    plt.subplot(1, len(complexity_feats), i+1)
    
    # Ensure each feature is numeric to prepare for any formatting issues
    df[feat] = pd.to_numeric(df[feat], errors='coerce')
    
    # Create a boxplot to compare healthy vs Parkinson's for this feature
    sns.boxplot(
        x='parkinson_status',
        y=feat,
        data=df,
        palette={'0':'green', '1':'red'}  # green = healthy, red = Parkinson's
    )
    #reduce clutter for graphs
    plt.xlabel('')  
    plt.ylabel(feat) 
    plt.title(feat)  

plt.suptitle("Complexity Measures by Parkinson's vs Healthier Status")

plt.tight_layout()
plt.show()
No description has been provided for this image

Story Telling & Insights¶

As I explored the Clinical Parkinson’s Dataset, a few interesting things stood out:

1. Dataset distribution
The first graph shows how many samples are from healthy people versus people with Parkinson’s. There are more samples from people with Parkinson’s, which helps when looking for patterns.

2. Voice patterns
Features like jitter and shimmer show that people with Parkinson’s tend to have higher and more spread-out values. Healthy people’s voices are more consistent, which makes sense because Parkinson’s affects voice stability.

3. Important features
Features such as spread_1, spread_2, RPDE, PPE, and jitter_abs show the strongest connection to Parkinson’s. Other features have weaker relationships but still provide insight.

4. Principal Component Analysis (PCA)
The PCA plot summarizes all features into two main components. PC1 explains about 53.8% of the variation and PC2 explains about 11%. The plot shows that healthy and Parkinson’s samples begin to form separate clusters, highlighting meaningful differences in voice features.

5. Complexity measures
RPDE, DFA, and PPE measure how irregular or unpredictable the voice is. People with Parkinson’s generally have higher values here which can mean their voices are less steady than healthy people.

What I learned:
These visualizations confirm that certain voice features, especially jitter, shimmer, spread_1, spread_2, RPDE, and PPE, are strongly linked to Parkinson’s. They help identify which aspects of speech are affected and could support earlier detection of the disease!!

Impact Section¶

This analysis provides useful insights but does not capture the full picture! Below are some key considerations to keep in mind when interpreting the results and drawing conclusions:

Potential benefits:

  • The visualizations and feature analysis could help researchers or doctors better understand which voice characteristics are affected by Parkinson’s.
  • Identifying key features like jitter, shimmer, RPDE, and PPE could support earlier detection or monitoring of the disease.

Potential risks or harm:

  • This analysis is based only on the dataset provided, which may not represent the full diversity of patients. For example, age, gender, accents, or other health conditions might influence voice features but aren’t fully captured from the data I personally visualized and studied.
  • Misinterpretation of the results could lead to overconfidence or other cognitive biases in diagnosing Parkinson’s from voice alone. THIS SHOULD NOT BE USED FOR OFFICIAL MEDICAL EVALUATION
  • Visualizations might unintentionally exaggerate differences if the audience doesn’t understand the context and is new to this feild of study which can possibly lead to more bias or misunderstanding.

Missing data or perspectives:

  • There could be other important clinical or lifestyle factors that influence voice patterns but are not included in the dataset.
  • Longitudinal tracking of voice changes over time could/would provide better understanding and prediction of voice fequencies.

Conclusion:
My aim for this project is to demonstrate the potential of using voice features for Parkinson’s research. Careful interpretation and awareness of dataset limitations are of most importance to avoid any misuse, incorrect conclusions, or harm. I hope you found this report interesting and insighful to hopefully spark or deepen your understanding of others around you who may have been diagnosed with Parkinson's disease. Thank you for reading! :)

References¶

  • DevPress. (2022, December 10). Python Matplotlib: How to create histogram plot in Python. Hive. Retrieved from https://hive.blog/hive-122108/@devpress/python-matplotlib-how-to-create-histogram-plot-in-python

  • Mayo Clinic Staff. (2024, September 27). Parkinson's disease - Symptoms and causes. Mayo Clinic. Retrieved from https://www.mayoclinic.org/diseases-conditions/parkinsons-disease/symptoms-causes/syc-20376055

  • Purohit, R. (2023). Clinical Parkinson’s Dataset. Kaggle. Retrieved from https://www.kaggle.com/datasets/rohanpurohit0705/clinical-parkinson-dataset

  • University of California, Berkeley. (n.d.). Data visualization with Python. Berkeley Data Science Education Program. Retrieved September 12, 2025, from [https://guides.lib.berkeley.edu/data-visualization]

  • Tech With Tim. (2020, September 15). Python Data Visualization Full Course – Data Visualization with Python [Video]. YouTube. Retrieved from https://www.youtube.com/watch?v=q68Qundmans

  • Data Science Tutorials. (2022, March 8). Master Python Plotly in 1.5 Hours: From Basics to Advanced [Video]. YouTube. Retrieved from https://www.youtube.com/watch?v=W_qQTKupZpY