{"metadata":{"colab":{"provenance":[]},"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":37333,"databundleVersionId":3949526,"sourceType":"competition"},{"sourceId":145998,"sourceType":"modelInstanceVersion","modelInstanceId":123817,"modelId":146879},{"sourceId":146617,"sourceType":"modelInstanceVersion","modelInstanceId":124360,"modelId":147407}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Classification of Blood Clot Origins in Ischemic Strokes 🩸**","metadata":{"id":"ziQaJdvqh88F"}},{"cell_type":"markdown","source":"This notebook explores the task of classifying the etiology of blood clots in whole-slide digital pathology images, specifically identifying whether they are of Cardioembolic (CE) or Large Artery Atherosclerosis (LAA) origin. Through an extensive exploratory data analysis (EDA), we describe the dataset, analyze missing and duplicate values, examine the distribution of image sizes, classify variables, and review the label distribution for the training set, along with plenty of other analysis. This EDA provides a comprehensive understanding of the data and helps identify potential preprocessing steps for optimal model performance.\n\nFollowing the EDA, we preprocess the images to standardize them for model input. The preprocessing involves resizing, converting images to grayscale, normalizing pixel values, and applying Gaussian blur to reduce noise. These steps ensure that the images are suitable for a Convolutional Neural Network (CNN) by preparing them with a consistent size, format, and reduced noise, enabling more efficient training and improved classification accuracy.","metadata":{"id":"2kK9OsnqiC9h"}},{"cell_type":"markdown","source":"**Authors:**\n- [Daniel Valdez](https://github.com/Danval-003)\n- [Emilio Solano](https://github.com/emiliosolanoo21)\n- [Adrian Flores](https://github.com/adrianRFlores)\n- [Andrea Ramírez](https://github.com/Andrea-gt)","metadata":{"id":"CXBxh8WPiG7b"}},{"cell_type":"markdown","source":"***","metadata":{"id":"sQGNhxwEiJKw"}},{"cell_type":"markdown","source":"## **(1) Import Libraries** ⬇️","metadata":{"id":"63DGMm3MiMEh"}},{"cell_type":"code","source":"# Data manipulation and visualization\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nfrom sklearn.model_selection import train_test_split\nimport subprocess\nfrom PIL import Image\nimport tifffile as tifi\nimport cv2\n\n# Set the maximum allowable pixels to a higher number.\nImage.MAX_IMAGE_PIXELS = None\n\n# Standard libraries\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# ===== ===== Reproducibility Seed ===== =====\n# Set a fixed seed for the random number generator for reproducibility\nrandom_state = 42\n\n# Set matplotlib inline\n%matplotlib inline\n\n# Set default figure size\nplt.rcParams['figure.figsize'] = (6, 4)\n\n# Define custom color palette\npalette = sns.color_palette(\"viridis\", 12)\n\n# Set the style of seaborn\nsns.set(style=\"whitegrid\")","metadata":{"id":"AP02dNX3cF31","execution":{"iopub.status.busy":"2024-10-25T05:37:46.476965Z","iopub.execute_input":"2024-10-25T05:37:46.477607Z","iopub.status.idle":"2024-10-25T05:37:48.3828Z","shell.execute_reply.started":"2024-10-25T05:37:46.477559Z","shell.execute_reply":"2024-10-25T05:37:48.381395Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **(2) Data Upload** 📄","metadata":{"id":"fsI-gEUliSPC"}},{"cell_type":"code","source":"# Note:\n# -----\n# For this specific task, we are only using the training data (`train.csv`) throughout most of our analysis.\n# The test dataset (`test.csv`) will not be used until the very end when making final predictions.\n# We will train, validate, and tune your model using the training data only.\n#\n# The reason the test data is only used at the very end is to prevent any bias during model training and evaluation.\n# We want the test data to remain completely unseen until after we've finalized our model so that it can serve\n# as a true evaluation of our model's performance.\n# ---------------------------------------------\n\ndf = pd.read_csv('../input/mayo-clinic-strip-ai/train.csv')  # Load the training data\ndf.head()  # Display the first 5 rows of the DataFrame for a quick inspection of the data","metadata":{"id":"WIj3lNgYmBA7","outputId":"689b0ef1-4fc0-4ed5-f5a1-84629e718226","execution":{"iopub.status.busy":"2024-10-25T05:37:51.350206Z","iopub.execute_input":"2024-10-25T05:37:51.350759Z","iopub.status.idle":"2024-10-25T05:37:51.392066Z","shell.execute_reply.started":"2024-10-25T05:37:51.350719Z","shell.execute_reply":"2024-10-25T05:37:51.390497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### **Descripción de los Datos**","metadata":{}},{"cell_type":"code","source":"# Assuming your DataFrame is named df\nbase_path = \"../input/mayo-clinic-strip-ai/train/\"\n\n# Add the full path to the df\ndf['image_path'] = base_path + df['image_id'] + '.tif'\n\n# Preview the DataFrame to ensure the new column is added correctly\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-25T05:38:10.510792Z","iopub.execute_input":"2024-10-25T05:38:10.511211Z","iopub.status.idle":"2024-10-25T05:38:10.526216Z","shell.execute_reply.started":"2024-10-25T05:38:10.511172Z","shell.execute_reply":"2024-10-25T05:38:10.524967Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **(3) Image Preprocessing 📷**","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\nimport cv2\nfrom sklearn.model_selection import StratifiedKFold\nfrom multiprocessing import Pool\nimport openslide\nimport albumentations as A\nfrom albumentations.pytorch import ToTensorV2\n\n# Definir aumentaciones avanzadas para mayor variabilidad en los datos\nminority_augmentations = A.Compose([\n    A.HorizontalFlip(p=0.8),\n    A.VerticalFlip(p=0.8),\n    A.RandomRotate90(p=0.8),\n    A.ColorJitter(brightness=0.4, contrast=0.4, saturation=0.4, hue=0.2, p=0.8),\n    A.Perspective(p=0.5),\n    A.RandomBrightnessContrast(p=0.6),\n    A.CoarseDropout(max_holes=10, max_height=20, max_width=20, min_holes=1, p=0.5),\n    A.Resize(256, 256),\n    ToTensorV2()\n])\n\nmajority_augmentations = A.Compose([\n    A.HorizontalFlip(p=0.5),\n    A.VerticalFlip(p=0.5),\n    A.RandomRotate90(p=0.5),\n    A.ColorJitter(brightness=0.3, contrast=0.3, saturation=0.3, hue=0.2, p=0.5),\n    A.Resize(256, 256),\n    ToTensorV2()\n])\n\n# Verificación de fondo para asegurar parches informativos\ndef is_background_patch(img, std_threshold=15, mean_diff_threshold=10):\n    gray = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n    if np.std(gray) < std_threshold:\n        return True\n    mean_r, mean_g, mean_b = img[..., 0].mean(), img[..., 1].mean(), img[..., 2].mean()\n    return max(abs(mean_r - mean_g), abs(mean_r - mean_b), abs(mean_g - mean_b)) < mean_diff_threshold\n\n# Verificar si un parche es válido\ndef is_valid_patch(img, threshold=15):\n    grayscale = cv2.cvtColor(img, cv2.COLOR_RGB2GRAY)\n    return np.std(grayscale) >= threshold\n\n# Procesar parches y aplicar aumentaciones balanceadas\ndef preprocess_image(args):\n    image_path, label, idx, num_patches = args\n    print(f\"Procesando imagen en el índice {idx}\")\n    \n    slide = openslide.OpenSlide(image_path)\n    width, height = slide.dimensions\n    patches = []\n    \n    patch_size = 512\n    step_size = patch_size // 2\n    max_attempts = num_patches * 30\n    \n    attempts = 0\n    for y in range(0, height - patch_size + 1, step_size):\n        for x in range(0, width - patch_size + 1, step_size):\n            if len(patches) >= num_patches:\n                break\n            \n            img = slide.read_region((x, y), 0, (patch_size, patch_size)).convert(\"RGB\")\n            img = np.array(img)\n            \n            if is_background_patch(img):\n                continue\n            \n            if is_valid_patch(img):\n                if label == 'LAA':  # Clase minoritaria\n                    augmented = minority_augmentations(image=img)\n                    patches.extend([augmented['image']] * 2)  # Añadir duplicados\n                else:  # Clase mayoritaria\n                    augmented = majority_augmentations(image=img)\n                    patches.append(augmented['image'])\n                \n                attempts = 0\n            else:\n                attempts += 1\n            \n            if attempts >= max_attempts:\n                print(f\"Excedido el máximo de intentos en imagen {image_path}\")\n                break\n    \n    if len(patches) < num_patches:\n        print(f\"No se encontraron suficientes parches válidos para la imagen en {image_path}\")\n        return None, None\n    \n    patches_tensor = np.stack(patches[:num_patches], axis=0)  # Equilibrar cantidad final de parches\n    return patches_tensor, label\n\n# Procesar y guardar conjunto de datos en fragmentos\ndef create_and_save_dataset_in_chunks(df, dataset_type, num_patches_per_image=15, chunk_size=5000):\n    images_dir = '/kaggle/input/mayo-clinic-strip-ai/test' if dataset_type == 'test' else '/kaggle/input/mayo-clinic-strip-ai/train'\n    \n    all_images = []\n    all_labels = []\n    total_samples = 0\n    chunk_index = 0\n    \n    args_list = []\n    for idx, row in df.iterrows():\n        image_path = os.path.join(images_dir, f\"{row['image_id']}.tif\") if dataset_type == 'test' else row['image_path']\n        label = None if dataset_type == 'test' else row['label']\n        args_list.append((image_path, label, idx, num_patches_per_image))\n    \n    with Pool(processes=os.cpu_count()) as pool:\n        for patches, label in pool.imap(preprocess_image, args_list):\n            if patches is not None:\n                all_images.append(patches)\n                if dataset_type != 'test':\n                    all_labels.append(label)\n                \n                total_samples += 1\n                \n                if total_samples >= chunk_size:\n                    if dataset_type != 'test':\n                        save_chunk(np.array(all_images), np.array(all_labels), dataset_type, chunk_index)\n                    else:\n                        save_test_chunk(np.array(all_images), chunk_index)\n                    chunk_index += 1\n                    all_images = []\n                    all_labels = []\n                    total_samples = 0\n    \n    if total_samples > 0:\n        if dataset_type != 'test':\n            save_chunk(np.array(all_images), np.array(all_labels), dataset_type, chunk_index)\n        else:\n            save_test_chunk(np.array(all_images), chunk_index)\n\n# Guardar fragmentos de entrenamiento y validación\ndef save_chunk(images, labels, dataset_type, chunk_index):\n    np.save(f'X_{dataset_type}_chunk_{chunk_index}.npy', images)\n    np.save(f'y_{dataset_type}_chunk_{chunk_index}.npy', labels)\n    print(f'Guardado {len(images)} muestras en X_{dataset_type}_chunk_{chunk_index}.npy y y_{dataset_type}_chunk_{chunk_index}.npy')\n\n# Guardar fragmentos del conjunto de prueba\ndef save_test_chunk(images, chunk_index):\n    np.save(f'X_test_chunk_{chunk_index}.npy', images)\n    print(f'Guardado {len(images)} muestras en X_test_chunk_{chunk_index}.npy')\n\n# Cargar el CSV de entrenamiento\ntrain_csv_path = '/kaggle/input/mayo-clinic-strip-ai/train.csv'\ndf_train = pd.read_csv(train_csv_path)\ntrain_images_dir = '/kaggle/input/mayo-clinic-strip-ai/train'\ndf_train['image_path'] = df_train['image_id'].apply(lambda x: os.path.join(train_images_dir, f\"{x}.tif\"))\ndf_train.rename(columns={'target': 'label'}, inplace=True)\n\n# Crear K-Folds estratificados en lugar de dividir una vez\nkf = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\nfor fold, (train_idx, val_idx) in enumerate(kf.split(df_train, df_train['label'])):\n    print(f'Procesando fold {fold + 1}')\n    df_train_fold = df_train.iloc[train_idx]\n    df_val_fold = df_train.iloc[val_idx]\n    \n    create_and_save_dataset_in_chunks(df_train_fold, dataset_type=f'train_fold{fold}', num_patches_per_image=15, chunk_size=5000)\n    create_and_save_dataset_in_chunks(df_val_fold, dataset_type=f'val_fold{fold}', num_patches_per_image=15, chunk_size=5000)\n\n# Procesar el conjunto de prueba\ntest_csv_path = '/kaggle/input/mayo-clinic-strip-ai/test.csv'\ndf_test = pd.read_csv(test_csv_path)\ncreate_and_save_dataset_in_chunks(df_test, dataset_type='test', num_patches_per_image=15, chunk_size=5000)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-25T07:00:13.957035Z","iopub.execute_input":"2024-10-25T07:00:13.958475Z","iopub.status.idle":"2024-10-25T07:00:32.344378Z","shell.execute_reply.started":"2024-10-25T07:00:13.958412Z","shell.execute_reply":"2024-10-25T07:00:32.343032Z"},"trusted":true},"execution_count":null,"outputs":[]}]}