Demystifying Mushroom Classification with Machine Learning & Unsupervised Taxa Discovery
Can data science determine whether a mushroom is edible or deadly? Using the iconic UCI Mushroom Dataset containing 8,124 samples across 23 features, this exploratory Google Colab notebook explores data visualization, morphological clustering, and random forest classification to accurately identify fungal traits.
1. Environment Setup & Data Loading
First, we fetch the dataset directly from Kaggle using kagglehub and load it into a Pandas DataFrame.
2. Visualizing Key Morphological Features
Understanding which visual traits signal toxicity is crucial. Below, we compare key traits such as Odor, Gill Color, Ring Type, and Spore Print Color between Edible (e) and Poisonous (p) mushrooms.
Key Insight: Odor is a distinct indicator. For instance, mushrooms with an almond or anise odor (a, l) are predominantly edible, while foul odors (f) indicate poisonous specimens.
3. Unsupervised Taxa Discovery & Supervised Classification
We extract critical taxonomic features, apply One-Hot Encoding, and group the mushrooms into 5 morphive taxa clusters using K-Means Clustering. Additionally, we train a RandomForestClassifier to evaluate taxonomic family predictability based on spore print colors.
Classification Performance
The Random Forest model achieves 100% Precision, Recall, and F1-Score across all mapped spore families on the test set:
| Estimated Taxa Family | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| Agaricaceae (Black Spored) | 1.00 | 1.00 | 1.00 | 373 |
| Agaricaceae (Brown Spored) | 1.00 | 1.00 | 1.00 | 404 |
| Amanitaceae / Lepiotaceae (White Spored) | 1.00 | 1.00 | 1.00 | 452 |
| Bolbitiaceae (Chocolate Spored) | 1.00 | 1.00 | 1.00 | 338 |
| Coprinaceae (Buff Spored) | 1.00 | 1.00 | 1.00 | 8 |
| Cortinariaceae (Orange Spored) | 1.00 | 1.00 | 1.00 | 9 |
| Entolomataceae (Purple Spored) | 1.00 | 1.00 | 1.00 | 14 |
| Russulaceae (Yellow Spored) | 1.00 | 1.00 | 1.00 | 13 |
| Strophariaceae (Green Spored) | 1.00 | 1.00 | 1.00 | 14 |
4. Visualizing Clusters: PCA vs t-SNE Projections
To inspect cluster boundaries in 2D space, linear dimensional reduction (PCA) and non-linear manifold learning (t-SNE) are applied.
Takeaway: PCA retains 33.7% of total variance in 2D space and separates general linear groupings, whereas t-SNE provides clear, distinct clusters that perfectly isolate poisonous species from edible ones within sub-clusters.





No comments:
Post a Comment