{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"colab":{"provenance":[]},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":113002,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# ChestDx-MultiInstitution Dataset: Exploratory Data Analysis\n\n**Author**: Guntas Dhanjal  \n**Date**: August 22, 2025  \n**Competition**: Grand X-ray Slam: Division B  \n\nThis notebook explores the *ChestDx-MultiInstitution* training dataset for [*Grand X-ray Slam: Division B*](https://www.kaggle.com/competitions/grand-xray-slam-division-b) to guide Kaggle competitors. Focused on the `train2.csv` file (108,494 images, 32,077 patients), it analyzes label distributions, multi-label patterns, demographics, and image views to inform model development. All analyses use Part 2 training data only to avoid test set leakage. Compete in [*Division A*](https://www.kaggle.com/competitions/grand-xray-slam-division-a) to boost your Grand Slam Leaderboard rank!\n","metadata":{"id":"-z1AL6nASvOm"}},{"cell_type":"markdown","source":"## Table of Contents\n- [1. Introduction](#1-introduction)\n- [2. Environment Setup](#2-environment-setup)\n- [3. Dataset Summary](#3-dataset-summary)\n- [4. Label Prevalence Analysis](#4-label-prevalence-analysis)\n- [5. Multi-Label Patterns](#5-multi-label-patterns)\n- [6. Demographic Insights](#6-demographic-insights)\n- [7. Image View Breakdown](#7-image-view-breakdown)\n- [8. Visual Explorations](#8-visual-explorations)\n- [9. Data Integrity Checks](#9-data-integrity-checks)\n- [10. Key Findings and Recommendations](#10-key-findings-and-recommendations)\n- [License](#license)","metadata":{"id":"IhNMcLYQTOmT"}},{"cell_type":"markdown","source":"# 1. Introduction\n\nThis EDA analyzes the *ChestDx-MultiInstitution* training dataset for *Grand X-ray Slam: Division B*[](https://www.kaggle.com/competitions/grand-xray-slam-division-b), a Kaggle hackathon to advance *Dr HealthAgent* by *Blue and Gold Healthcare Inc.*. Division B includes 108,494 images across 32,077 patients, targeting 14 thoracic conditions. Join [*Division A*](https://www.kaggle.com/competitions/grand-xray-slam-division-a) for the full challenge. Goals:\n- Understand dataset size and structure.\n- Examine label distributions and multi-label complexity.\n- Explore demographics and image view types.\n- Provide actionable tips for robust AI models.","metadata":{"id":"n9mw7MvSNMUL"}},{"cell_type":"markdown","source":"# 2. Environment Setup\n\nSet up the Python environment with necessary libraries and load the training dataset.","metadata":{"id":"-vpyynOnNpL2"}},{"cell_type":"code","source":"# Import required libraries for analysis and visualization\nimport pandas as pd\nimport numpy as np\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom PIL import Image\nimport os\n\n# Ensure plots display in Colab\n%matplotlib inline\n\n# Set seaborn style for clean visualizations\nsns.set(style='whitegrid')\n\n# Set seaborn style for clean visualizations\nsns.set(style='whitegrid')\n","metadata":{"id":"Kp6J1xoWLRHU","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:41.323692Z","iopub.execute_input":"2025-08-22T15:41:41.324298Z","iopub.status.idle":"2025-08-22T15:41:42.567537Z","shell.execute_reply.started":"2025-08-22T15:41:41.324269Z","shell.execute_reply":"2025-08-22T15:41:42.566824Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Define path to training data\ndata_path = '/kaggle/input/grand-xray-slam-division-b/train2.csv'\n\n# Load training dataset with error handling\ntry:\n    train_df = pd.read_csv(data_path)\n    print(f\"Successfully loaded train.csv with shape: {train_df.shape}\")\nexcept FileNotFoundError:\n    print(f\"Error: {data_path} not found. Please check the file path.\")\n    raise","metadata":{"id":"ThwIhA4dwl83","outputId":"f2970b4e-11c7-401c-cba9-ef598604e3cb","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:42.568732Z","iopub.execute_input":"2025-08-22T15:41:42.569346Z","iopub.status.idle":"2025-08-22T15:41:42.963706Z","shell.execute_reply.started":"2025-08-22T15:41:42.569322Z","shell.execute_reply":"2025-08-22T15:41:42.962902Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df.head(1)","metadata":{"id":"uCr6mLOBUj_H","outputId":"e83a46c1-19c1-4f64-dfac-222255f3758e","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:42.964395Z","iopub.execute_input":"2025-08-22T15:41:42.964666Z","iopub.status.idle":"2025-08-22T15:41:43.000241Z","shell.execute_reply.started":"2025-08-22T15:41:42.964644Z","shell.execute_reply":"2025-08-22T15:41:42.999304Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 3. Dataset Summary\n\nSummarize the dataset’s size, structure, and check for missing values.","metadata":{"id":"12kOOEbTPi1U"}},{"cell_type":"code","source":"# Display basic dataset info\nprint(\"Dataset Info:\")\nprint(train_df.info())","metadata":{"id":"i5GfNtNUNhdV","outputId":"daa3b58a-54cd-4966-a5fd-17d44e79edb9","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:43.002181Z","iopub.execute_input":"2025-08-22T15:41:43.00244Z","iopub.status.idle":"2025-08-22T15:41:43.04922Z","shell.execute_reply.started":"2025-08-22T15:41:43.002419Z","shell.execute_reply":"2025-08-22T15:41:43.048338Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Summarize key metrics\ntotal_images = len(train_df)\ntotal_patients = train_df['Patient_ID'].nunique()\ntotal_studies = train_df['Study'].nunique()\nprint(f\"Total Images: {total_images}\")\nprint(f\"Total Patients: {total_patients}\")\nprint(f\"Total Studies: {total_studies}\")","metadata":{"id":"G8uO9EWVUeLs","outputId":"7921df79-f258-472f-c964-a8d4079404c5","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:43.050024Z","iopub.execute_input":"2025-08-22T15:41:43.050318Z","iopub.status.idle":"2025-08-22T15:41:43.060513Z","shell.execute_reply.started":"2025-08-22T15:41:43.05029Z","shell.execute_reply":"2025-08-22T15:41:43.059656Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check for missing values\nprint(\"\\nMissing Values:\")\nprint(train_df.isnull().sum())","metadata":{"id":"HKJbo0g9Ugk6","outputId":"0e75e6a1-e7a5-4568-fd5e-96fdb777cc28","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:43.061478Z","iopub.execute_input":"2025-08-22T15:41:43.061783Z","iopub.status.idle":"2025-08-22T15:41:43.102083Z","shell.execute_reply.started":"2025-08-22T15:41:43.061754Z","shell.execute_reply":"2025-08-22T15:41:43.101023Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 4. Label Prevalence Analysis\n\nAnalyze the distribution of the 14 conditions to identify class imbalance.","metadata":{"id":"r9dggrcbcMhC"}},{"cell_type":"code","source":"# Define the 14 condition columns\nlabel_columns = ['No Finding', 'Lung Opacity', 'Support Devices', 'Atelectasis',\n                 'Cardiomegaly', 'Pleural Effusion', 'Enlarged Cardiomediastinum',\n                 'Edema', 'Consolidation', 'Pneumonia', 'Fracture', 'Lung Lesion',\n                 'Pneumothorax', 'Pleural Other']\n\n# Calculate counts and percentages for each condition\nlabel_counts = train_df[label_columns].sum()\nlabel_percentages = (label_counts / total_images * 100).round(2)\nprevalence_df = pd.DataFrame({\n    'Condition': label_counts.index,\n    'Count': label_counts.values,\n    'Percent (%)': label_percentages.values\n})\n\n# Display prevalence table\nprint(\"Label Prevalence:\")\nprint(prevalence_df)","metadata":{"id":"43WsMADnPqSD","outputId":"1e391fc9-85c2-43c7-d9da-1fc1506f858d","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:43.102845Z","iopub.execute_input":"2025-08-22T15:41:43.103073Z","iopub.status.idle":"2025-08-22T15:41:43.115886Z","shell.execute_reply.started":"2025-08-22T15:41:43.103051Z","shell.execute_reply":"2025-08-22T15:41:43.115055Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Barplot for label prevalence\nplt.figure(figsize=(12, 6))\nsns.barplot(x='Count', y='Condition', data=prevalence_df, palette='viridis', hue=None)\nplt.title('Label Counts (Number of Positive Cases)')\nplt.xlabel('Count')\nplt.ylabel('Condition')\nplt.legend([],[], frameon=False)\nplt.tight_layout()\nplt.savefig('/content/label_counts_barplot.jpg')\nplt.show()","metadata":{"id":"QXyLrzovvDZ0","outputId":"20da1c0e-9fa7-472f-dcb0-43c99dffe198","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:43.116705Z","iopub.execute_input":"2025-08-22T15:41:43.116982Z","iopub.status.idle":"2025-08-22T15:41:43.851614Z","shell.execute_reply.started":"2025-08-22T15:41:43.11696Z","shell.execute_reply":"2025-08-22T15:41:43.850701Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Donut chart for label percentages\nplt.figure(figsize=(8, 8))\ncolors = sns.color_palette('viridis', len(prevalence_df))\nplt.pie(prevalence_df['Count'], labels=prevalence_df['Condition'],\n        autopct=lambda pct: f'{pct:.1f}%', startangle=140, colors=colors,\n        wedgeprops={'width': 0.4})\nplt.title('Label Prevalence (Positive Cases %)')\nplt.tight_layout()\nplt.savefig('/content/label_percent_donut.jpg')\nplt.show()","metadata":{"id":"07OqwfX3vMn7","outputId":"e8bbdac3-9ea0-47ea-f3af-4c76f1d13710","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:43.85246Z","iopub.execute_input":"2025-08-22T15:41:43.852733Z","iopub.status.idle":"2025-08-22T15:41:44.239306Z","shell.execute_reply.started":"2025-08-22T15:41:43.852707Z","shell.execute_reply":"2025-08-22T15:41:44.238477Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 5. Multi-Label Patterns\n\nExamine how often multiple conditions appear in the same image.","metadata":{"id":"LnSRwchyRINH"}},{"cell_type":"code","source":"# Calculate number of labels per image\ntrain_df['Label_Count'] = train_df[label_columns].sum(axis=1)\nmulti_label_counts = train_df['Label_Count'].value_counts().sort_index()\nmulti_label_percent = (multi_label_counts / total_images * 100).round(2)\n\n# Display multi-label distribution\nprint(\"Multi-Label Distribution:\")\nprint(pd.DataFrame({'Number of Labels': multi_label_counts.index,\n                    'Count': multi_label_counts.values,\n                    'Percent (%)': multi_label_percent.values}))","metadata":{"id":"GmpH0pcNQxka","outputId":"09f04cf1-751a-4e02-a94f-a259618cb58d","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:44.241624Z","iopub.execute_input":"2025-08-22T15:41:44.2419Z","iopub.status.idle":"2025-08-22T15:41:44.273765Z","shell.execute_reply.started":"2025-08-22T15:41:44.24188Z","shell.execute_reply":"2025-08-22T15:41:44.273038Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Compute co-occurrence matrix\nco_occurrence = train_df[label_columns].T.dot(train_df[label_columns])\nprint(\"\\nCo-Occurrence Matrix (Top 5x5 for brevity):\")\nprint(co_occurrence.iloc[:5, :5])","metadata":{"id":"4gHIVoBFhJRP","outputId":"7ec0f93e-5a80-4649-fc07-9d3452d821b2","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:44.274614Z","iopub.execute_input":"2025-08-22T15:41:44.274881Z","iopub.status.idle":"2025-08-22T15:41:44.331015Z","shell.execute_reply.started":"2025-08-22T15:41:44.27486Z","shell.execute_reply":"2025-08-22T15:41:44.330164Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Heatmap for co-occurrence\nplt.figure(figsize=(10, 8))\nsns.heatmap(co_occurrence, annot=True, fmt='d', cmap='viridis')\nplt.title('Label Co-Occurrence Matrix')\nplt.tight_layout()\nplt.savefig('/content/co_occurrence_heatmap.jpg')\nplt.show()","metadata":{"id":"c6cOH_2ivXkq","outputId":"bc4b117e-1631-4906-9982-470fcbe6ea69","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:44.332127Z","iopub.execute_input":"2025-08-22T15:41:44.332398Z","iopub.status.idle":"2025-08-22T15:41:45.589491Z","shell.execute_reply.started":"2025-08-22T15:41:44.332373Z","shell.execute_reply":"2025-08-22T15:41:45.588572Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Visualizing Single-Label vs Multi-Label Image Distribution in Training Data\nsingle_label_count = (train_df['Label_Count'] == 1).sum()\nmulti_label_count = (train_df['Label_Count'] > 1).sum()\n\nlabels = ['Single Label', 'Multi-Label']\nsizes = [single_label_count, multi_label_count]\ncolors = ['#66c2a5', '#fc8d62']\nexplode = (0.05, 0.05)\n\nplt.figure(figsize=(6,6))\nplt.pie(sizes, labels=labels, autopct='%1.1f%%', startangle=140, colors=colors, explode=explode)\nplt.title('Single vs Multi-Label Image Distribution')\nplt.axis('equal')\nplt.show()\n","metadata":{"id":"r-fGrm9tYq7r","outputId":"51643141-f847-4fbb-c4d4-ca7efbec616a","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:45.590527Z","iopub.execute_input":"2025-08-22T15:41:45.590858Z","iopub.status.idle":"2025-08-22T15:41:45.689643Z","shell.execute_reply.started":"2025-08-22T15:41:45.590835Z","shell.execute_reply":"2025-08-22T15:41:45.688727Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Label Distribution Excluding \"No Finding\"\nfiltered_df = train_df[(train_df['Label_Count'] >= 1) & (train_df['No Finding'] == 0)]\n\nsingle_label_count = (filtered_df['Label_Count'] == 1).sum()\nmulti_label_count = (filtered_df['Label_Count'] > 1).sum()\n\nlabels = ['Single Label (Excl. No Finding)', 'Multi-Label (Excl. No Finding)']\nsizes = [single_label_count, multi_label_count]\ncolors = ['#66c2a5', '#fc8d62']\nexplode = (0.05, 0.05)\n\nplt.figure(figsize=(6,6))\nplt.pie(sizes, labels=labels, autopct='%1.1f%%', startangle=140, colors=colors, explode=explode)\nplt.title('Label Distribution Excluding \"No Finding\"')\nplt.axis('equal')\nplt.show()\n","metadata":{"id":"QXDR1itOZVYn","outputId":"f7e127eb-675f-46e7-e090-068c8fd3f989","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:45.690458Z","iopub.execute_input":"2025-08-22T15:41:45.69068Z","iopub.status.idle":"2025-08-22T15:41:45.822309Z","shell.execute_reply.started":"2025-08-22T15:41:45.690662Z","shell.execute_reply":"2025-08-22T15:41:45.82146Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 6. Demographic Insights\n\nAnalyze Sex and Age distributions and their relation to conditions.","metadata":{"id":"F0jpxjx1Rxc-"}},{"cell_type":"code","source":"# Replace NaNs with string 'Unknown' for plotting\nsex_for_plot = train_df['Sex'].fillna('Unknown')\n\n# Get counts including NaNs replaced\nsex_counts = sex_for_plot.value_counts()\n\nprint(\"Sex distribution (including 'Unknown'):\")\nprint(sex_counts)","metadata":{"id":"1OM-YWkYkJT_","outputId":"c84ba4fd-5c80-44cb-deb3-2753b7882339","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:45.823274Z","iopub.execute_input":"2025-08-22T15:41:45.823523Z","iopub.status.idle":"2025-08-22T15:41:45.84504Z","shell.execute_reply.started":"2025-08-22T15:41:45.823503Z","shell.execute_reply":"2025-08-22T15:41:45.844279Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.figure(figsize=(6,4))\nsns.countplot(x=sex_for_plot, order=sex_counts.index, palette=['#95a5a6', '#3498db', '#e74c3c'])\nplt.title('Sex Distribution (including Unknown)')\nplt.xlabel('Sex')\nplt.ylabel('Count')\nplt.show()\n","metadata":{"id":"c2-ei9HxSHi9","outputId":"bbcb004f-b2ae-4ffa-ed87-98a7b2a9c229","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:45.846008Z","iopub.execute_input":"2025-08-22T15:41:45.846309Z","iopub.status.idle":"2025-08-22T15:41:46.032849Z","shell.execute_reply.started":"2025-08-22T15:41:45.846288Z","shell.execute_reply":"2025-08-22T15:41:46.031866Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Age distribution\nvalid_age_min = 0\nvalid_age_max = 120\nages = train_df['Age']\nvalid_ages = ages.dropna()\nvalid_ages_count = valid_ages[(valid_ages >= valid_age_min) & (valid_ages <= valid_age_max)].count()\nnan_ages_count = ages.isna().sum()\ninvalid_ages_count = valid_ages[(valid_ages < valid_age_min) | (valid_ages > valid_age_max)].count()\n\nprint(f\"Images with valid numeric Age: {valid_ages_count}\")\nprint(f\"Images with missing Age (NaN): {nan_ages_count}\")\nprint(f\"Images with invalid Age (<{valid_age_min} or >{valid_age_max}): {invalid_ages_count}\")","metadata":{"id":"Kz3fUfvLbq2X","outputId":"4072a45d-4604-48f1-ceb8-89871b7c5c26","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:46.033753Z","iopub.execute_input":"2025-08-22T15:41:46.034077Z","iopub.status.idle":"2025-08-22T15:41:46.044809Z","shell.execute_reply.started":"2025-08-22T15:41:46.034051Z","shell.execute_reply":"2025-08-22T15:41:46.043924Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Histogram for Age distribution\nplt.figure(figsize=(8, 6))\nsns.histplot(train_df['Age'].dropna(), bins=30, kde=True, color='teal')\nplt.title('Age Distribution')\nplt.xlabel('Age')\nplt.ylabel('Count')\nplt.tight_layout()\nplt.savefig('/content/age_distribution.jpg')\nplt.show()","metadata":{"id":"MU_WqH82vhnC","outputId":"e28317bc-f289-4bae-998b-a6832c95afcc","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:46.045763Z","iopub.execute_input":"2025-08-22T15:41:46.046067Z","iopub.status.idle":"2025-08-22T15:41:46.918057Z","shell.execute_reply.started":"2025-08-22T15:41:46.046041Z","shell.execute_reply":"2025-08-22T15:41:46.917173Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Condition prevalence by age group\nage_bins = [0, 20, 40, 60, 80, 100]\ntrain_df['AgeGroup'] = pd.cut(train_df['Age'], bins=age_bins, right=False)\n\nage_label_counts = train_df.groupby('AgeGroup')[label_columns].apply(lambda x: (x == 1).sum())\n\nplt.figure(figsize=(18, 10))\nax = age_label_counts.T.plot(kind='bar', stacked=False, figsize=(18,10))\n\nplt.title('Condition Counts by Age Group', fontsize=20)\nplt.xlabel('Condition (Disease)', fontsize=16)\nplt.ylabel('Count', fontsize=16)\nplt.xticks(rotation=90, fontsize=13)\nplt.yticks(fontsize=13)\nplt.legend(title='Age Group', fontsize=12, title_fontsize=14, loc='upper right')\nplt.tight_layout()\nplt.show()\n","metadata":{"id":"PHctfxn3vp11","outputId":"bbe30944-5ab0-4e37-a88c-89f8979b73e0","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:46.918895Z","iopub.execute_input":"2025-08-22T15:41:46.919124Z","iopub.status.idle":"2025-08-22T15:41:47.604704Z","shell.execute_reply.started":"2025-08-22T15:41:46.919106Z","shell.execute_reply":"2025-08-22T15:41:47.603748Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 7. Image View Breakdown\n\nExplore the distribution of Frontal/Lateral and AP/PA views.","metadata":{"id":"PTLEbfnfSdoX"}},{"cell_type":"code","source":"# Image View Breakdown\n\n# Count view categories and positions\nview_cat_counts = train_df['ViewCategory'].value_counts(dropna=False)\nview_pos_counts = train_df['ViewPosition'].value_counts(dropna=False)\n\nprint(\"ViewCategory counts:\")\nprint(view_cat_counts)\n\nprint(\"\\nViewPosition counts:\")\nprint(view_pos_counts)\n","metadata":{"id":"YApz2H4sSOma","outputId":"5decbee0-fe01-4245-9778-b3ca154d09fe","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:47.6059Z","iopub.execute_input":"2025-08-22T15:41:47.606156Z","iopub.status.idle":"2025-08-22T15:41:47.620405Z","shell.execute_reply.started":"2025-08-22T15:41:47.606136Z","shell.execute_reply":"2025-08-22T15:41:47.619505Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Image View Distribution\nplt.figure(figsize=(12,5))\n\nplt.subplot(1,2,1)\nsns.countplot(x='ViewCategory', data=train_df, order=view_cat_counts.index)\nplt.title('ViewCategory Distribution')\nplt.xticks(rotation=45)\n\nplt.subplot(1,2,2)\nsns.countplot(x='ViewPosition', data=train_df, order=view_pos_counts.index)\nplt.title('ViewPosition Distribution')\nplt.xticks(rotation=45)\n\nplt.tight_layout()\nplt.show()","metadata":{"id":"98kr9vu-v1fx","outputId":"fa976851-3e80-4710-ed86-28f7a364b281","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:47.621694Z","iopub.execute_input":"2025-08-22T15:41:47.622513Z","iopub.status.idle":"2025-08-22T15:41:48.040864Z","shell.execute_reply.started":"2025-08-22T15:41:47.622463Z","shell.execute_reply":"2025-08-22T15:41:48.039977Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 8. Visual Explorations\n\nVisualize sample images.","metadata":{"id":"GqSFwabqUjRO"}},{"cell_type":"code","source":"# Path to train folder\nIMAGE_ROOT_DIR = '/kaggle/input/grand-xray-slam-division-b/train2'\n\n# List .jpg images\nimage_files = [f for f in os.listdir(IMAGE_ROOT_DIR) if f.endswith('.jpg')]\n\n# Take first 5 images\nimage_files = image_files[:5]\n\n# Plot images\nfig, axs = plt.subplots(1, len(image_files), figsize=(20, 5))\nfor i, image_file in enumerate(image_files):\n    image_path = os.path.join(IMAGE_ROOT_DIR, image_file)\n    img = Image.open(image_path).convert('L')\n    axs[i].imshow(img, cmap='gray')\n    axs[i].set_title(image_file, fontsize=8)\n    axs[i].axis('off')\n\nplt.suptitle('Sample X-ray Images from Train Folder', fontsize=14)\nplt.tight_layout()\nplt.show()","metadata":{"id":"1Giyv0-NSxk5","outputId":"18a92691-a795-4f35-804c-ab137f4a1142","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:48.041702Z","iopub.execute_input":"2025-08-22T15:41:48.042004Z","iopub.status.idle":"2025-08-22T15:41:51.602484Z","shell.execute_reply.started":"2025-08-22T15:41:48.041981Z","shell.execute_reply":"2025-08-22T15:41:51.601614Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 9. Data Integrity Checks\n\nVerify uniqueness of Patient_IDs and Image_Names, and check for metadata inconsistencies.","metadata":{"id":"f8-4DnLXWff6"}},{"cell_type":"code","source":"# Check for duplicate Image_Names\nduplicate_images = train_df['Image_name'].duplicated().sum()\nprint(f\"Duplicated Image_Name entries: {duplicate_images}\")\n\n# Check for duplicate Patient_IDs (expected due to multiple images per patient)\nduplicate_patients = total_images - total_patients\nprint(f\"Duplicated Patient_ID entries: {duplicate_patients}\")\n\n# Check for invalid Age values\ninvalid_ages = train_df['Age'].dropna()\ninvalid_ages = invalid_ages[invalid_ages < 0].count()\nprint(f\"Invalid Age values (<0): {invalid_ages}\")","metadata":{"id":"sCnnfQWqU4jw","outputId":"4eb10b6f-0abb-4812-c3ae-18693efcaf04","trusted":true,"execution":{"iopub.status.busy":"2025-08-22T15:41:51.603475Z","iopub.execute_input":"2025-08-22T15:41:51.604033Z","iopub.status.idle":"2025-08-22T15:41:51.62781Z","shell.execute_reply.started":"2025-08-22T15:41:51.604007Z","shell.execute_reply":"2025-08-22T15:41:51.626982Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# 10. Key Findings and Recommendations\n\nThe *Grand X-ray Slam: division B* training dataset is rich and complex, offering opportunities for competitors:\n- **Dataset Size**: 108,494 images, 32,077 patients, 184 studies.\n- **Label Imbalance**: Lung Opacity (45.18%) and No Finding (31.67%) dominate; Pleural Other (6.39%) and Pneumothorax (8.05%) are rare.\n- **Multi-Label Complexity**: ~75% of images have multiple positive labels, requiring multi-label modeling.\n- **Demographics**: Sex distribution balanced (~50% Male/Female, assumed); median age ~50 (assumed). Use Sex/Age as features.\n- **Image Views**: Frontal views dominant (>80%, assumed), with AP/PA and Lateral. Stratify by view for better performance.\n- **Recommendations**:\n  - Use weighted loss functions (e.g., binary cross-entropy with class weights) to address imbalance.\n  - Apply data augmentation (e.g., rotation, flipping) for rare conditions like Pneumothorax.\n  - Leverage multi-label techniques (e.g., sigmoid outputs per class).\n  - Explore view-specific models or use ViewCategory/ViewPosition as features.\n- **Final Note**: This dataset mirrors real-world medical complexity, aligning with *Dr HealthAgent*’s goals. Creative modeling will set you apart!\n\nGood luck, competitors, your models will power life-saving diagnostics in *Grand X-ray Slam: Division B* and [*Division A*](https://www.kaggle.com/competitions/grand-xray-slam-division-a)!","metadata":{"id":"J5ZorZHGWlhi"}},{"cell_type":"markdown","source":"# License\n\nThis notebook is licensed under the Creative Commons Attribution 4.0 International License (CC-BY 4.0). You are free to share and adapt this work, provided you give appropriate credit to the author and indicate any changes made. See [https://creativecommons.org/licenses/by/4.0/](https://creativecommons.org/licenses/by/4.0/) for details.","metadata":{"id":"NuZNfcGrpGLG"}}]}