{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"colab":{"provenance":[]},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":112899,"databundleVersionId":13449579,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ChestDx-MultiInstitution Dataset: Exploratory Data Analysis\n\n**Author**: Guntas Dhanjal  \n**Date**: August 21, 2025  \n**Competition**: Grand X-ray Slam: Division A\n\nThis notebook explores the ChestDx-MultiInstitution training dataset for Grand X-ray Slam: [Division A](https://www.kaggle.com/competitions/grand-xray-slam-division-a) to guide Kaggle competitors. Focused on the train1.csv file (107,374 images, 32,076 patients), it analyzes label distributions, multi-label patterns, demographics, and image views to inform model development. All analyses use Part 1 training data only to avoid test set leakage. Compete in [part 2](https://www.kaggle.com/competitions/grand-xray-slam-division-b) to boost your Grand Slam Leaderboard rank!","metadata":{"id":"-z1AL6nASvOm"}},{"cell_type":"markdown","source":"## Table of Contents\n- [1. Introduction](#1-introduction)\n- [2. Environment Setup](#2-environment-setup)\n- [3. Dataset Summary](#3-dataset-summary)\n- [4. Label Prevalence Analysis](#4-label-prevalence-analysis)\n- [5. Multi-Label Patterns](#5-multi-label-patterns)\n- [6. Demographic Insights](#6-demographic-insights)\n- [7. Image View Breakdown](#7-image-view-breakdown)\n- [8. Visual Explorations](#8-visual-explorations)\n- [9. Data Integrity Checks](#9-data-integrity-checks)\n- [10. Key Findings and Recommendations](#10-key-findings-and-recommendations)\n- [License](#license)","metadata":{"id":"IhNMcLYQTOmT"}},{"cell_type":"markdown","source":"# 1. Introduction\n\nThis EDA analyzes the ChestDx-MultiInstitution training dataset for [Grand X-ray Slam: Division A](https://www.kaggle.com/competitions/grand-xray-slam-division-a), a Kaggle hackathon to advance Dr HealthAgent by Blue and Gold Healthcare Inc. Part 1 includes ~107,374 images across 32,076 patients, targeting 14 thoracic conditions. Join [Division B](https://www.kaggle.com/competitions/grand-xray-slam-division-b) for the full challenge. Goals:\n- Understand dataset size and structure.\n- Examine label distributions and multi-label complexity.\n- Explore demographics and image view types.\n- Provide actionable tips for robust AI models.","metadata":{"id":"n9mw7MvSNMUL"}},{"cell_type":"markdown","source":"# 2. Environment Setup\n\nSet up the Python environment with necessary libraries and load the training dataset.","metadata":{"id":"-vpyynOnNpL2"}},{"cell_type":"code","source":"# Import required libraries for analysis and visualization\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport os\n\n# Ensure plots display in Colab\n%matplotlib inline\n\n# Set seaborn style for clean visualizations\nsns.set(style='whitegrid')\n\nimport warnings\nwarnings.simplefilter(action='ignore', category=FutureWarning)","metadata":{"id":"Kp6J1xoWLRHU","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:48.064811Z","iopub.execute_input":"2025-08-22T14:26:48.065121Z","iopub.status.idle":"2025-08-22T14:26:49.373169Z","shell.execute_reply.started":"2025-08-22T14:26:48.065096Z","shell.execute_reply":"2025-08-22T14:26:49.372271Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define path to training data\ndata_path = '/kaggle/input/grand-xray-slam-division-a/train1.csv'\n\n# Load training dataset with error handling\ntry:\n    train_df = pd.read_csv(data_path)\n    print(f\"Successfully loaded train.csv with shape: {train_df.shape}\")\nexcept FileNotFoundError:\n    print(f\"Error: {data_path} not found. Please check the file path.\")\n    raise","metadata":{"id":"ThwIhA4dwl83","outputId":"5c66429d-a408-4e00-ea74-ac9a08416f38","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:49.374864Z","iopub.execute_input":"2025-08-22T14:26:49.375317Z","iopub.status.idle":"2025-08-22T14:26:49.78923Z","shell.execute_reply.started":"2025-08-22T14:26:49.375294Z","shell.execute_reply":"2025-08-22T14:26:49.788435Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.head(1)","metadata":{"id":"uCr6mLOBUj_H","outputId":"238fea89-f85c-4eab-95bd-f4a7ef3da635","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:49.790136Z","iopub.execute_input":"2025-08-22T14:26:49.790427Z","iopub.status.idle":"2025-08-22T14:26:49.827678Z","shell.execute_reply.started":"2025-08-22T14:26:49.790402Z","shell.execute_reply":"2025-08-22T14:26:49.82686Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. Dataset Summary\n\nSummarize the dataset’s size, structure, and check for missing values.","metadata":{"id":"12kOOEbTPi1U"}},{"cell_type":"code","source":"# Display basic dataset info\nprint(\"Dataset Info:\")\nprint(train_df.info())","metadata":{"id":"i5GfNtNUNhdV","outputId":"f13d2d61-0c75-410d-e70f-e48d807fe83f","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:49.829682Z","iopub.execute_input":"2025-08-22T14:26:49.830259Z","iopub.status.idle":"2025-08-22T14:26:49.878819Z","shell.execute_reply.started":"2025-08-22T14:26:49.830235Z","shell.execute_reply":"2025-08-22T14:26:49.877967Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Summarize key metrics\ntotal_images = len(train_df)\ntotal_patients = train_df['Patient_ID'].nunique()\ntotal_studies = train_df['Study'].nunique()\nprint(f\"Total Images: {total_images}\")\nprint(f\"Total Patients: {total_patients}\")\nprint(f\"Total Studies: {total_studies}\")","metadata":{"id":"G8uO9EWVUeLs","outputId":"154b4722-c7a9-4c6e-ae9f-c0cb04467a20","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:49.879617Z","iopub.execute_input":"2025-08-22T14:26:49.879859Z","iopub.status.idle":"2025-08-22T14:26:49.890344Z","shell.execute_reply.started":"2025-08-22T14:26:49.879838Z","shell.execute_reply":"2025-08-22T14:26:49.889415Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check for missing values\nprint(\"\\nMissing Values:\")\nprint(train_df.isnull().sum())","metadata":{"id":"HKJbo0g9Ugk6","outputId":"92cf320a-4643-4048-bf22-d21d90601634","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:49.891329Z","iopub.execute_input":"2025-08-22T14:26:49.891609Z","iopub.status.idle":"2025-08-22T14:26:49.934129Z","shell.execute_reply.started":"2025-08-22T14:26:49.89158Z","shell.execute_reply":"2025-08-22T14:26:49.932982Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 4. Label Prevalence Analysis\n\nAnalyze the distribution of the 14 conditions to identify class imbalance.","metadata":{"id":"r9dggrcbcMhC"}},{"cell_type":"code","source":"# Define the 14 condition columns\nlabel_columns = ['No Finding', 'Lung Opacity', 'Support Devices', 'Atelectasis',\n                 'Cardiomegaly', 'Pleural Effusion', 'Enlarged Cardiomediastinum',\n                 'Edema', 'Consolidation', 'Pneumonia', 'Fracture', 'Lung Lesion',\n                 'Pneumothorax', 'Pleural Other']\n\n# Calculate counts and percentages for each condition\nlabel_counts = train_df[label_columns].sum()\nlabel_percentages = (label_counts / total_images * 100).round(2)\nprevalence_df = pd.DataFrame({\n    'Condition': label_counts.index,\n    'Count': label_counts.values,\n    'Percent (%)': label_percentages.values\n})\n\n# Display prevalence table\nprint(\"Label Prevalence:\")\nprint(prevalence_df)","metadata":{"id":"43WsMADnPqSD","outputId":"b2276766-2916-455d-c902-17b477c9b37e","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:49.935367Z","iopub.execute_input":"2025-08-22T14:26:49.935707Z","iopub.status.idle":"2025-08-22T14:26:49.948848Z","shell.execute_reply.started":"2025-08-22T14:26:49.935675Z","shell.execute_reply":"2025-08-22T14:26:49.948089Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(12, 6))\nsns.barplot(x='Count', y='Condition', data=prevalence_df, palette='viridis', hue=None)\nplt.title('Label Counts (Number of Positive Cases)')\nplt.xlabel('Count')\nplt.ylabel('Condition')\nplt.legend([],[], frameon=False)\nplt.tight_layout()\nplt.savefig('/content/label_counts_barplot.jpg')\nplt.show()\n","metadata":{"id":"QXyLrzovvDZ0","outputId":"fe3bd4f4-9efd-476a-bec1-a447b21a6cf4","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:49.949771Z","iopub.execute_input":"2025-08-22T14:26:49.950084Z","iopub.status.idle":"2025-08-22T14:26:50.770338Z","shell.execute_reply.started":"2025-08-22T14:26:49.950052Z","shell.execute_reply":"2025-08-22T14:26:50.769341Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Donut chart for label percentages\nplt.figure(figsize=(8, 8))\ncolors = sns.color_palette('viridis', len(prevalence_df))\nplt.pie(prevalence_df['Count'], labels=prevalence_df['Condition'],\n        autopct=lambda pct: f'{pct:.1f}%', startangle=140, colors=colors,\n        wedgeprops={'width': 0.4})\nplt.title('Label Prevalence (Positive Cases %)')\nplt.tight_layout()\nplt.savefig('/content/label_percent_donut.jpg')\nplt.show()","metadata":{"id":"07OqwfX3vMn7","outputId":"3e836c5f-e5cc-4971-8ce8-beab40df0556","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:50.771301Z","iopub.execute_input":"2025-08-22T14:26:50.771582Z","iopub.status.idle":"2025-08-22T14:26:51.17241Z","shell.execute_reply.started":"2025-08-22T14:26:50.771538Z","shell.execute_reply":"2025-08-22T14:26:51.171499Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5. Multi-Label Patterns\n\nExamine how often multiple conditions appear in the same image.","metadata":{"id":"LnSRwchyRINH"}},{"cell_type":"code","source":"# Calculate number of labels per image\ntrain_df['Label_Count'] = train_df[label_columns].sum(axis=1)\nmulti_label_counts = train_df['Label_Count'].value_counts().sort_index()\nmulti_label_percent = (multi_label_counts / total_images * 100).round(2)\n\n# Display multi-label distribution\nprint(\"Multi-Label Distribution:\")\nprint(pd.DataFrame({'Number of Labels': multi_label_counts.index,\n                    'Count': multi_label_counts.values,\n                    'Percent (%)': multi_label_percent.values}))","metadata":{"id":"GmpH0pcNQxka","outputId":"28eeb5ee-eaf0-474d-caf1-c396bae0302b","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:51.175343Z","iopub.execute_input":"2025-08-22T14:26:51.176052Z","iopub.status.idle":"2025-08-22T14:26:51.207703Z","shell.execute_reply.started":"2025-08-22T14:26:51.176021Z","shell.execute_reply":"2025-08-22T14:26:51.206689Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Compute co-occurrence matrix\nco_occurrence = train_df[label_columns].T.dot(train_df[label_columns])\nprint(\"\\nCo-Occurrence Matrix (Top 5x5 for brevity):\")\nprint(co_occurrence.iloc[:5, :5])","metadata":{"id":"4gHIVoBFhJRP","outputId":"5b10cfd2-9629-4df9-8eb4-f21e24b1daef","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:51.208678Z","iopub.execute_input":"2025-08-22T14:26:51.208951Z","iopub.status.idle":"2025-08-22T14:26:51.259967Z","shell.execute_reply.started":"2025-08-22T14:26:51.208921Z","shell.execute_reply":"2025-08-22T14:26:51.259245Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Heatmap for co-occurrence\nplt.figure(figsize=(10, 8))\nsns.heatmap(co_occurrence, annot=True, fmt='d', cmap='viridis')\nplt.title('Label Co-Occurrence Matrix')\nplt.tight_layout()\nplt.savefig('/content/co_occurrence_heatmap.jpg')\nplt.show()","metadata":{"id":"c6cOH_2ivXkq","outputId":"4cc91703-587a-4889-d32d-4054769db35b","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:51.260778Z","iopub.execute_input":"2025-08-22T14:26:51.261034Z","iopub.status.idle":"2025-08-22T14:26:52.585677Z","shell.execute_reply.started":"2025-08-22T14:26:51.261014Z","shell.execute_reply":"2025-08-22T14:26:52.584761Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Visualizing Single-Label vs Multi-Label Image Distribution in Training Data\nsingle_label_count = (train_df['Label_Count'] == 1).sum()\nmulti_label_count = (train_df['Label_Count'] > 1).sum()\n\nlabels = ['Single Label', 'Multi-Label']\nsizes = [single_label_count, multi_label_count]\ncolors = ['#66c2a5', '#fc8d62']\nexplode = (0.05, 0.05)\n\nplt.figure(figsize=(6,6))\nplt.pie(sizes, labels=labels, autopct='%1.1f%%', startangle=140, colors=colors, explode=explode)\nplt.title('Single vs Multi-Label Image Distribution')\nplt.axis('equal')\nplt.show()\n","metadata":{"id":"r-fGrm9tYq7r","outputId":"da717a79-51f7-4aad-811e-f6dec7223c8c","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:52.586695Z","iopub.execute_input":"2025-08-22T14:26:52.587028Z","iopub.status.idle":"2025-08-22T14:26:52.69063Z","shell.execute_reply.started":"2025-08-22T14:26:52.586998Z","shell.execute_reply":"2025-08-22T14:26:52.689676Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Label Distribution Excluding \"No Finding\"\nfiltered_df = train_df[(train_df['Label_Count'] >= 1) & (train_df['No Finding'] == 0)]\n\nsingle_label_count = (filtered_df['Label_Count'] == 1).sum()\nmulti_label_count = (filtered_df['Label_Count'] > 1).sum()\n\nlabels = ['Single Label (Excl. No Finding)', 'Multi-Label (Excl. No Finding)']\nsizes = [single_label_count, multi_label_count]\ncolors = ['#66c2a5', '#fc8d62']\nexplode = (0.05, 0.05)\n\nplt.figure(figsize=(6,6))\nplt.pie(sizes, labels=labels, autopct='%1.1f%%', startangle=140, colors=colors, explode=explode)\nplt.title('Label Distribution Excluding \"No Finding\"')\nplt.axis('equal')\nplt.show()\n","metadata":{"id":"QXDR1itOZVYn","outputId":"6df5a082-51b4-424a-8201-b2c047c1db35","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:52.691624Z","iopub.execute_input":"2025-08-22T14:26:52.691971Z","iopub.status.idle":"2025-08-22T14:26:52.825768Z","shell.execute_reply.started":"2025-08-22T14:26:52.691944Z","shell.execute_reply":"2025-08-22T14:26:52.824874Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 6. Demographic Insights\n\nAnalyze Sex and Age distributions and their relation to conditions.","metadata":{"id":"F0jpxjx1Rxc-"}},{"cell_type":"code","source":"# Replace NaNs with string 'Unknown' for plotting\nsex_for_plot = train_df['Sex'].fillna('Unknown')\n\n# Get counts including NaNs replaced\nsex_counts = sex_for_plot.value_counts()\n\nprint(\"Sex distribution (including 'Unknown'):\")\nprint(sex_counts)","metadata":{"id":"1OM-YWkYkJT_","outputId":"234dde9c-1aa3-4e31-9260-986f575c49b1","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:52.826876Z","iopub.execute_input":"2025-08-22T14:26:52.827194Z","iopub.status.idle":"2025-08-22T14:26:52.849694Z","shell.execute_reply.started":"2025-08-22T14:26:52.827169Z","shell.execute_reply":"2025-08-22T14:26:52.848613Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(6,4))\nsns.countplot(x=sex_for_plot, order=sex_counts.index, palette=['#95a5a6', '#3498db', '#e74c3c'])\nplt.title('Sex Distribution (including Unknown)')\nplt.xlabel('Sex')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"id":"c2-ei9HxSHi9","outputId":"43248480-209d-4638-cb3e-e3b5a561328c","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:52.85083Z","iopub.execute_input":"2025-08-22T14:26:52.851126Z","iopub.status.idle":"2025-08-22T14:26:53.046866Z","shell.execute_reply.started":"2025-08-22T14:26:52.851104Z","shell.execute_reply":"2025-08-22T14:26:53.045883Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Age distribution\nvalid_age_min = 0\nvalid_age_max = 120\nages = train_df['Age']\nvalid_ages = ages.dropna()\nvalid_ages_count = valid_ages[(valid_ages >= valid_age_min) & (valid_ages <= valid_age_max)].count()\nnan_ages_count = ages.isna().sum()\ninvalid_ages_count = valid_ages[(valid_ages < valid_age_min) | (valid_ages > valid_age_max)].count()\n\nprint(f\"Images with valid numeric Age: {valid_ages_count}\")\nprint(f\"Images with missing Age (NaN): {nan_ages_count}\")\nprint(f\"Images with invalid Age (<{valid_age_min} or >{valid_age_max}): {invalid_ages_count}\")","metadata":{"id":"Kz3fUfvLbq2X","outputId":"58c3cdc9-d75a-4290-b296-70657469cf1f","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:53.047962Z","iopub.execute_input":"2025-08-22T14:26:53.048481Z","iopub.status.idle":"2025-08-22T14:26:53.059047Z","shell.execute_reply.started":"2025-08-22T14:26:53.048451Z","shell.execute_reply":"2025-08-22T14:26:53.058076Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Histogram for Age distribution\nplt.figure(figsize=(8, 6))\nsns.histplot(train_df['Age'].dropna(), bins=30, kde=True, color='teal')\nplt.title('Age Distribution')\nplt.xlabel('Age')\nplt.ylabel('Count')\nplt.tight_layout()\nplt.savefig('/content/age_distribution.jpg')\nplt.show()","metadata":{"id":"MU_WqH82vhnC","outputId":"0b2b7b7a-1542-4b39-a93e-f03b988533d9","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:53.060064Z","iopub.execute_input":"2025-08-22T14:26:53.060354Z","iopub.status.idle":"2025-08-22T14:26:53.976078Z","shell.execute_reply.started":"2025-08-22T14:26:53.060323Z","shell.execute_reply":"2025-08-22T14:26:53.975131Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Condition prevalence by age group\nage_bins = [0, 20, 40, 60, 80, 100]\ntrain_df['AgeGroup'] = pd.cut(train_df['Age'], bins=age_bins, right=False)\n\nage_label_counts = train_df.groupby('AgeGroup')[label_columns].apply(lambda x: (x == 1).sum())\n\nplt.figure(figsize=(18, 10))\nax = age_label_counts.T.plot(kind='bar', stacked=False, figsize=(18,10))\n\nplt.title('Condition Counts by Age Group', fontsize=20)\nplt.xlabel('Condition (Disease)', fontsize=16)\nplt.ylabel('Count', fontsize=16)\nplt.xticks(rotation=90, fontsize=13)\nplt.yticks(fontsize=13)\nplt.legend(title='Age Group', fontsize=12, title_fontsize=14, loc='upper right')\nplt.tight_layout()\nplt.show()\n","metadata":{"id":"PHctfxn3vp11","outputId":"b9335bb4-2a83-45a4-f16c-1fc9f9bbb0e3","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:53.977337Z","iopub.execute_input":"2025-08-22T14:26:53.977684Z","iopub.status.idle":"2025-08-22T14:26:54.686552Z","shell.execute_reply.started":"2025-08-22T14:26:53.977654Z","shell.execute_reply":"2025-08-22T14:26:54.685741Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 7. Image View Breakdown\n\nExplore the distribution of Frontal/Lateral and AP/PA views.","metadata":{"id":"PTLEbfnfSdoX"}},{"cell_type":"code","source":"# Image View Breakdown\n\n# Count view categories and positions\nview_cat_counts = train_df['ViewCategory'].value_counts(dropna=False)\nview_pos_counts = train_df['ViewPosition'].value_counts(dropna=False)\n\nprint(\"ViewCategory counts:\")\nprint(view_cat_counts)\n\nprint(\"\\nViewPosition counts:\")\nprint(view_pos_counts)\n","metadata":{"id":"YApz2H4sSOma","outputId":"7d9ff877-d9d6-4740-9950-8f3af0ab4741","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:54.68758Z","iopub.execute_input":"2025-08-22T14:26:54.687945Z","iopub.status.idle":"2025-08-22T14:26:54.703992Z","shell.execute_reply.started":"2025-08-22T14:26:54.687919Z","shell.execute_reply":"2025-08-22T14:26:54.702902Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Image View Distribution\nplt.figure(figsize=(12,5))\n\nplt.subplot(1,2,1)\nsns.countplot(x='ViewCategory', data=train_df, order=view_cat_counts.index)\nplt.title('ViewCategory Distribution')\nplt.xticks(rotation=45)\n\nplt.subplot(1,2,2)\nsns.countplot(x='ViewPosition', data=train_df, order=view_pos_counts.index)\nplt.title('ViewPosition Distribution')\nplt.xticks(rotation=45)\n\nplt.tight_layout()\nplt.show()","metadata":{"id":"98kr9vu-v1fx","outputId":"1bc7723c-cb62-4a09-d550-3b24f1027e80","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:54.705107Z","iopub.execute_input":"2025-08-22T14:26:54.705412Z","iopub.status.idle":"2025-08-22T14:26:55.152965Z","shell.execute_reply.started":"2025-08-22T14:26:54.705385Z","shell.execute_reply":"2025-08-22T14:26:55.15141Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 8. Visual Explorations\n\nVisualize sample images.","metadata":{"id":"GqSFwabqUjRO"}},{"cell_type":"code","source":"# Path to train folder\nIMAGE_ROOT_DIR = '/kaggle/input/grand-xray-slam-division-a/train1'\n\n# List .jpg images\nimage_files = [f for f in os.listdir(IMAGE_ROOT_DIR) if f.endswith('.jpg')]\n\n# Take first 5 images\nimage_files = image_files[:5]\n\n# Plot images\nfig, axs = plt.subplots(1, len(image_files), figsize=(20, 5))\nfor i, image_file in enumerate(image_files):\n    image_path = os.path.join(IMAGE_ROOT_DIR, image_file)\n    img = Image.open(image_path).convert('L')\n    axs[i].imshow(img, cmap='gray')\n    axs[i].set_title(image_file, fontsize=8)\n    axs[i].axis('off')\n\nplt.suptitle('Sample X-ray Images from Train Folder', fontsize=14)\nplt.tight_layout()\nplt.show()","metadata":{"id":"1Giyv0-NSxk5","outputId":"a5343662-7ab3-42e1-84c4-0ebf1509539d","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:55.154021Z","iopub.execute_input":"2025-08-22T14:26:55.154345Z","iopub.status.idle":"2025-08-22T14:26:58.759862Z","shell.execute_reply.started":"2025-08-22T14:26:55.154319Z","shell.execute_reply":"2025-08-22T14:26:58.758879Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 9. Data Integrity Checks\n\nVerify uniqueness of Patient_IDs and Image_Names, and check for metadata inconsistencies.","metadata":{"id":"f8-4DnLXWff6"}},{"cell_type":"code","source":"# Check for duplicate Image_Names\nduplicate_images = train_df['Image_name'].duplicated().sum()\nprint(f\"Duplicated Image_Name entries: {duplicate_images}\")\n\n# Check for duplicate Patient_IDs (expected due to multiple images per patient)\nduplicate_patients = total_images - total_patients\nprint(f\"Duplicated Patient_ID entries: {duplicate_patients}\")\n\n# Check for invalid Age values\ninvalid_ages = train_df['Age'].dropna()\ninvalid_ages = invalid_ages[invalid_ages < 0].count()\nprint(f\"Invalid Age values (<0): {invalid_ages}\")\n","metadata":{"id":"sCnnfQWqU4jw","outputId":"cd0a28b4-6bfe-4149-aaeb-acbeca542980","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T14:26:58.760846Z","iopub.execute_input":"2025-08-22T14:26:58.761108Z","iopub.status.idle":"2025-08-22T14:26:58.78126Z","shell.execute_reply.started":"2025-08-22T14:26:58.761088Z","shell.execute_reply":"2025-08-22T14:26:58.780352Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 10. Key Findings and Recommendations\n\nThe *Grand X-ray Slam Part 1* training dataset is rich and complex, offering opportunities for competitors:\n- **Dataset Size**: 107,374 images, 32,076 patients, 157 studies.\n- **Label Imbalance**: Lung Opacity (45.29%) and No Finding (31.62%) dominate; Pleural Other (6.58%) and Pneumothorax (8.21%) are rare.\n- **Multi-Label Complexity**: ~75% of images have multiple positive labels, requiring multi-label modeling.\n- **Demographics**: Sex distribution balanced (~50% Male/Female, assumed); median age ~50 (assumed). Use Sex/Age as features.\n- **Image Views**: Frontal views dominant (>80%, assumed), with AP/PA and Lateral. Stratify by view for better performance.\n- **Recommendations**:\n  - Use weighted loss functions (e.g., binary cross-entropy with class weights) to address imbalance.\n  - Apply data augmentation (e.g., rotation, flipping) for rare conditions like Pneumothorax.\n  - Leverage multi-label techniques (e.g., sigmoid outputs per class).\n  - Explore view-specific models or use ViewCategory/ViewPosition as features.\n- **Final Note**: This dataset mirrors real-world medical complexity, aligning with *Dr HealthAgent*’s goals. Creative modeling will set you apart!\n\nGood luck, competitors—your models will power life-saving diagnostics in *Grand X-ray Slam: Division A* and [*Division B*](https://www.kaggle.com/competitions/grand-xray-slam-division-b)!\n","metadata":{"id":"J5ZorZHGWlhi"}},{"cell_type":"markdown","source":"# License\n\nThis notebook is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). You are free to share and adapt this work, provided you give appropriate credit to the author and indicate any changes made. See [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/) for details.","metadata":{"id":"NuZNfcGrpGLG"}}]}