{"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.11.11","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":39272,"databundleVersionId":4629629,"sourceType":"competition"},{"sourceId":4866520,"sourceType":"datasetVersion","datasetId":2820722}],"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true},"colab":{"provenance":[]}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **A Beacon of Hope: Harnessing AI for Breast Cancer Detection**","metadata":{"id":"clVQpa5MNt4X"}},{"cell_type":"markdown","source":"### **Chapter 1: The Call to Action**\n**Every year, millions of women undergo mammography screening. For many, the results bring comfort and reassurance; for others, the diagnosis of breast cancer changes their lives. Radiologists like Dr. Amal Ali have long dedicated themselves to early detection, knowing that accurate diagnosis can save lives. Yet, despite her expertise, Dr. Amal knew that even the most skilled eyes can sometimes miss subtle signs—or mistakenly raise false alarms that cause undue stress.**\n\n`What if there were an assistant, a tireless ally, that could help reduce both missed cases and false positives?`\n\n    — Dr. Amal Ali  \n\n**Driven by this vision, Dr. Amal partnered with a team of data scientists to create an AI-powered tool. Their mission was clear: develop a model that could sift through thousands of radiographic breast images and accurately flag potential cases of cancer, while minimizing unnecessary worry.**","metadata":{"id":"-qIgHdLFNt4h"}},{"cell_type":"markdown","source":"### **Chapter 2: Unveiling the Dataset**\n**The team embarked on their journey by assembling a rich [dataset](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/data) of radiographic breast images. The data was organized as follows:**\n- `Data Source:` train_images/[patient_id]/[image_id].dcm\n- `Subjects:` Female patients undergoing screening exams\n- `Patients:` Approximately 8,000 unique patients\n- `Images per Patient:` Usually, but not always, 4 images\n- `Image Format:` **DICOM** (with many images encoded in JPEG 2000)\n- `train.csv:` a csv file that have corresponding labels (0 for benign, 1 for malignant).\n\n**This diverse and challenging dataset held the promise of training a robust model. However, the team knew that working with DICOM files—and handling JPEG 2000 images in particular—would require special care.**","metadata":{"id":"Y9edfrLKNt4j"}},{"cell_type":"markdown","source":"### **Chapter 3: Preparing the Data 👨‍💻**\n**Before the AI could learn to detect cancer, the images needed to be carefully processed. Using the pydicom library, the team set out to load and preprocess the data. They also accounted for the fact that some images were stored in JPEG 2000 format, ensuring the proper libraries were in place.**","metadata":{"id":"q_IzuPIBNt4l"}},{"cell_type":"markdown","source":"**At first he imported needed modules**","metadata":{"id":"a6STEzyoNt4l"}},{"cell_type":"code","source":"import os\nimport pandas as pd\nimport numpy as np\nimport pydicom\nfrom PIL import Image\nimport torch\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nfrom PIL import Image\nimport matplotlib.pyplot as plt\nimport torch.nn as nn\nimport torchvision.models as models\nimport torch.optim as optim\nfrom tqdm import tqdm\nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nimport matplotlib.pyplot as plt\nfrom sklearn.model_selection import train_test_split\nimport timm  # For pre-trained models like EfficientNet\n\nimport tensorflow as tf\nfrom tensorflow.keras.applications import EfficientNetB5, EfficientNetB0\nfrom tensorflow.keras import layers, models\nfrom tensorflow.keras.callbacks import EarlyStopping, ModelCheckpoint\n\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","trusted":true,"id":"82W3kT0oNt4m","execution":{"iopub.status.busy":"2025-04-28T13:22:37.716494Z","iopub.execute_input":"2025-04-28T13:22:37.7176Z","iopub.status.idle":"2025-04-28T13:22:37.725384Z","shell.execute_reply.started":"2025-04-28T13:22:37.717565Z","shell.execute_reply":"2025-04-28T13:22:37.724562Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**And here, you can convert `.dcm` files into `.jpg` files, you can skip this step if you need and use this preprocessed [datatset](https://www.kaggle.com/datasets/paulbacher/rsna-bcd-1024x512-preprocessed/data)**","metadata":{"id":"Cy03kL-sNt4o"}},{"cell_type":"code","source":"# def convert_dicom_to_jpg(dicom_path, jpg_path):\n#     \"\"\"Converts a single DICOM file to a JPEG file with error handling.\"\"\"\n#     try:\n#         # Read the DICOM file\n#         ds = pydicom.dcmread(dicom_path)\n\n#         # Extract pixel data\n#         image = ds.pixel_array.astype(float)\n\n#         # Apply rescale slope and intercept if available (commonly for CT images)\n#         if 'RescaleSlope' in ds and 'RescaleIntercept' in ds:\n#             image = image * float(ds.RescaleSlope) + float(ds.RescaleIntercept)\n\n#         # Normalize the pixel values to the range 0-255\n#         min_val = np.min(image)\n#         max_val = np.max(image)\n#         if max_val - min_val != 0:\n#             image_normalized = (image - min_val) / (max_val - min_val) * 255.0\n#         else:\n#             image_normalized = np.zeros_like(image)\n\n#         image_normalized = image_normalized.astype(np.uint8)\n\n#         # Create a PIL image and save as JPEG\n#         im = Image.fromarray(image_normalized)\n#         im.save(jpg_path)\n#         print(f\"Converted {dicom_path} to {jpg_path}\")\n\n#     except RuntimeError as e:\n#         # This error likely indicates that decompression failed because of missing plugins.\n#         print(f\"RuntimeError for file {dicom_path}: {e}. Skipping this file.\")\n#     except Exception as e:\n#         # Catch any other unexpected errors.\n#         print(f\"An error occurred while converting {dicom_path}: {e}. Skipping this file.\")","metadata":{"trusted":true,"id":"4cQaJ18vNt4p"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# def process_folder(input_folder, output_folder):\n#     \"\"\"\n#     Recursively processes the input_folder, converts all .dcm files,\n#     and saves them into the output_folder, preserving the subfolder structure.\n#     \"\"\"\n#     for root, _, files in os.walk(input_folder):\n#         for file in files:\n#             if file.lower().endswith('.dcm'):\n#                 # Construct the full input file path\n#                 dicom_path = os.path.join(root, file)\n\n#                 # Determine the relative path to recreate folder structure in output_folder\n#                 relative_path = os.path.relpath(root, input_folder)\n#                 output_dir = os.path.join(output_folder, relative_path)\n#                 os.makedirs(output_dir, exist_ok=True)\n\n#                 # Create output file path by replacing .dcm with .jpg\n#                 jpg_filename = os.path.splitext(file)[0] + '.jpg'\n#                 jpg_path = os.path.join(output_dir, jpg_filename)\n\n#                 # Convert and save the image, with error handling inside the conversion function\n#                 convert_dicom_to_jpg(dicom_path, jpg_path)","metadata":{"trusted":true,"id":"9mMoGNk5Nt4q"},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# # Define the input and output directories\n# input_folder = ''         # folder containing subfolders with DICOM files\n# output_folder = ''        # folder to store converted JPEG images\n\n# # Process the folder\n# process_folder(input_folder, output_folder)","metadata":{"trusted":true,"id":"sR4Zxn8qNt4r"},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Then the team read data and store it in dataframe. The data was in a folder 📂 `train_images` and there is a `train.csv` file to get the labels, so the team stored images into `df`.**","metadata":{"id":"b4V9nf-6Nt4s"}},{"cell_type":"code","source":"import pandas as pd\nimport os\n\ndf = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\n\nbase_image_path = '/kaggle/input/rsna-bcd-1024x512-preprocessed/train_images'\ndf['image_path'] = base_image_path + '/' + df['patient_id'].astype(str) + '/' + df['image_id'].astype(str) + '.png'\n\nprint(df['image_path'].head())\n\nmissing_files = []\nfor path in df['image_path']:\n    if not os.path.exists(path):\n        missing_files.append(path)\n\nprint(f\"Num_missing_images: {len(missing_files)}\")\nif missing_files:\n    print(\"missing_images:\", missing_files[:5])\n","metadata":{"trusted":true,"id":"iUEO_AC9Nt4s","execution":{"iopub.status.busy":"2025-04-28T13:22:55.777606Z","iopub.execute_input":"2025-04-28T13:22:55.778241Z","iopub.status.idle":"2025-04-28T13:22:55.913563Z","shell.execute_reply.started":"2025-04-28T13:22:55.778213Z","shell.execute_reply":"2025-04-28T13:22:55.912756Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**After that they split `df` into `train_df`, `valid_df`, and `test_df` with 80% for training and 10% for validation and 10% for testing**","metadata":{"id":"mdEl_QRmNt4s"}},{"cell_type":"code","source":"train_df, temp_df = train_test_split(df, test_size=0.2, random_state=42, stratify=df['invasive'])\nvalid_df, test_df = train_test_split(temp_df, test_size=0.5, random_state=42, stratify=temp_df['invasive'])\nprint(\"Train_length:\", len(train_df), \"\\nValidation_length:\", len(valid_df), \"\\nTest_length:\", len(test_df))","metadata":{"trusted":true,"id":"P5exTqH2Nt4t","execution":{"iopub.status.busy":"2025-04-28T13:22:59.618391Z","iopub.execute_input":"2025-04-28T13:22:59.618936Z","iopub.status.idle":"2025-04-28T13:22:59.66302Z","shell.execute_reply.started":"2025-04-28T13:22:59.618909Z","shell.execute_reply":"2025-04-28T13:22:59.662401Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**At this time he was able to convert train_df, valid_df, and test_df into tensors with `batch_size = 32` for train and `batch_size = 16` for valid and test, and `img_size = (512, 512)`**","metadata":{"id":"LSFKXNOGNt4t"}},{"cell_type":"code","source":"import tensorflow as tf\n\nIMG_SIZE =  (512, 512)\nBATCH_SIZE_TRAIN = 32\nBATCH_SIZE_VAL_TEST = 16\n\ndef load_image(image_path, label):\n    image = tf.io.read_file(image_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, IMG_SIZE)\n    image = image / 255.0  # Normalization to [0,1]\n    return image, label\n\ndef df_to_dataset(df, batch_size, shuffle=True):\n    paths = df['image_path'].values\n    labels = df['invasive'].values  \n    dataset = tf.data.Dataset.from_tensor_slices((paths, labels))\n    dataset = dataset.map(load_image, num_parallel_calls=tf.data.AUTOTUNE)\n    if shuffle:\n        dataset = dataset.shuffle(buffer_size=len(df))\n    dataset = dataset.batch(batch_size).prefetch(buffer_size=tf.data.AUTOTUNE)\n    return dataset\n\ntrain_dataset = df_to_dataset(train_df, batch_size=BATCH_SIZE_TRAIN, shuffle=True)\nvalid_dataset = df_to_dataset(valid_df, batch_size=BATCH_SIZE_VAL_TEST, shuffle=True)\ntest_dataset  = df_to_dataset(test_df,  batch_size=BATCH_SIZE_VAL_TEST, shuffle=True)","metadata":{"trusted":true,"id":"LerLlRxvNt4t","execution":{"iopub.status.busy":"2025-04-28T13:23:06.584049Z","iopub.execute_input":"2025-04-28T13:23:06.584566Z","iopub.status.idle":"2025-04-28T13:23:06.667655Z","shell.execute_reply.started":"2025-04-28T13:23:06.584543Z","shell.execute_reply":"2025-04-28T13:23:06.666889Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**To show a sample of data, they tried to use a python script to just showing a sample of the data**","metadata":{"id":"yBnn7LOWNt4u"}},{"cell_type":"code","source":"def load_image(image_path):\n    image = tf.io.read_file(image_path)\n    image = tf.image.decode_jpeg(image, channels=3)\n    image = tf.image.resize(image, (512, 512))\n    image = image / 255.0  # Normalization to [0,1]\n    return image\n\n# Show random sample of 5 images\ndef show_sample_images(df, num_samples=5):\n    # Randomly select image paths\n    sample_paths = df.sample(num_samples)['image_path'].values\n\n    plt.figure(figsize=(15, 15))\n    for i, path in enumerate(sample_paths):\n        image = load_image(path)\n        plt.subplot(1, num_samples, i+1)\n        plt.imshow(image.numpy())  \n        plt.axis('off')  \n        plt.title(f\"Image {i+1}\")\n    plt.show()\n\n# Show sample images\nshow_sample_images(df, num_samples=5)","metadata":{"trusted":true,"id":"SQBTnvUENt4u","execution":{"iopub.status.busy":"2025-04-28T13:23:14.810179Z","iopub.execute_input":"2025-04-28T13:23:14.811038Z","iopub.status.idle":"2025-04-28T13:23:15.243146Z","shell.execute_reply.started":"2025-04-28T13:23:14.811007Z","shell.execute_reply":"2025-04-28T13:23:15.242318Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Chapter 4: Forging the AI Model 🤖**\n**With the data prepped, the team moved on to building the AI model. They chose a Convolutional Neural Network (CNN) architecture specially `EfficientNetB5` as it is a well-suited for image classification tasks like this. Given the critical nature of the diagnosis, the model needed to be both sensitive to true cases of cancer and cautious enough to avoid excessive false positives.**","metadata":{"id":"bQh5YRy3Nt4u"}},{"cell_type":"code","source":"base_model = EfficientNetB5(weights='imagenet', include_top=False, input_shape=(512, 512, 3))\nbase_model.trainable = False\nmodel = models.Sequential([\n    base_model,\n    layers.GlobalAveragePooling2D(),\n    layers.Dense(1024, activation='relu'),\n    layers.Dropout(0.5),\n    layers.Dense(1, activation='sigmoid')  \n])\nmodel.compile(optimizer=tf.keras.optimizers.Adam(),\n              loss='binary_crossentropy',  \n              metrics=['accuracy'])\nmodel.summary()","metadata":{"trusted":true,"id":"_65v1lzWNt4v","execution":{"iopub.status.busy":"2025-04-28T13:23:19.086542Z","iopub.execute_input":"2025-04-28T13:23:19.087072Z","iopub.status.idle":"2025-04-28T13:23:21.529533Z","shell.execute_reply.started":"2025-04-28T13:23:19.087049Z","shell.execute_reply":"2025-04-28T13:23:21.528937Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**This neural network was designed to extract hierarchical features from the mammograms, allowing the model to learn the subtle differences that indicate cancerous changes.**","metadata":{"id":"dW9qFZhjNt4v"}},{"cell_type":"markdown","source":"### **Chapter 5: Training the Champion**\n**The next step was the battle of training. The model was fed batches of preprocessed images along with their labels. With each epoch, the model learned to differentiate between healthy tissue and suspicious findings. Special care was taken to monitor performance, ensuring that false positives were minimized.**","metadata":{"id":"gyTqd-8qNt4v"}},{"cell_type":"code","source":"early_stopping = EarlyStopping(monitor='val_loss', patience=5, restore_best_weights=True)\nmodel_checkpoint = ModelCheckpoint('/kaggle/working/best_model.keras', save_best_only=True, monitor='val_loss', mode='min')\n\nhistory = model.fit(\n    train_dataset,\n    epochs=25, \n    validation_data=valid_dataset,\n    callbacks=[early_stopping, model_checkpoint]\n)","metadata":{"trusted":true,"id":"_65v1lzWNt4v","execution":{"iopub.status.busy":"2025-04-28T13:23:47.134249Z","iopub.execute_input":"2025-04-28T13:23:47.135068Z","execution_failed":"2025-04-28T13:36:30.139Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"model.save('/kaggle/working/final_model.keras')\ntest_loss, test_accuracy = model.evaluate(test_dataset)\nprint(f\"Test Accuracy: {test_accuracy * 100:.2f}%\")","metadata":{"trusted":true,"id":"_65v1lzWNt4v","execution":{"iopub.status.busy":"2025-04-28T13:37:11.060653Z","iopub.execute_input":"2025-04-28T13:37:11.060979Z","iopub.status.idle":"2025-04-28T13:37:11.137947Z","shell.execute_reply.started":"2025-04-28T13:37:11.060953Z","shell.execute_reply":"2025-04-28T13:37:11.137024Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plt.plot(history.history['loss'], label='Train Loss')\nplt.plot(history.history['val_loss'], label='Validation Loss')\nplt.legend()\nplt.title('Loss during training')\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.show()\n\nplt.plot(history.history['accuracy'], label='Train Accuracy')\nplt.plot(history.history['val_accuracy'], label='Validation Accuracy')\nplt.legend()\nplt.title('Accuracy during training')\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.show()","metadata":{"trusted":true,"id":"_65v1lzWNt4v","execution":{"execution_failed":"2025-04-28T13:22:21.062Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Chapter 6: The Triumph and Beyond**\n**After rigorous training and validation, the model was put to the test on unseen data. Its performance was promising—a testament to the tireless work of Dr. Amal and her team. The AI assistant demonstrated its potential to serve as a second pair of eyes, supporting radiologists in making more accurate diagnoses.😀**","metadata":{"id":"FelJreyYNt4w"}},{"cell_type":"code","source":"# Evaluate the model on the unseen test set\ntest_loss, test_accuracy, test_auc = model.evaluate(test_dataset)\n\nprint(f\"Test Loss: {test_loss:.4f}\")\nprint(f\"Test Accuracy: {test_accuracy:.4f}\")\nprint(f\"Test AUC: {test_auc:.4f}\")","metadata":{"id":"D6N2wNGjNt4x","trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### **Epilogue: A New Dawn of Hope**\n**In the quiet hum of the radiology lab, as mammograms cycled through the imaging machines, a new kind of vigilance took shape. The AI assistant, born from collaboration and fueled by data, was more than just an algorithm—it was a beacon of hope for countless women.**\n\n**Dr. Amal often reflected on the journey:**\n    \n`\"Every image processed, every diagnosis aided, brings us one step closer to a future where early detection is the norm rather than the exception. This is not just technology; it's a promise of better care and a brighter tomorrow.\"`","metadata":{"id":"Xzb4D10YNt4x"}},{"cell_type":"markdown","source":"#### **And so, the story continues—each breakthrough and every line of code reinforcing the commitment to save lives, one image at a time.❤️**","metadata":{"id":"INxnRGo6Nt4x"}}]}