{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.7.12","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":39272,"databundleVersionId":4629629,"sourceType":"competition"},{"sourceId":4696088,"sourceType":"datasetVersion","datasetId":2687741}],"dockerImageVersionId":30302,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In order to get up and running quickly in this competition, let's look at:\n\n* [An overview of data](#section-one)\n* [Training a fast.ai model](#section-two)\n* [Making a first submission](#section-three)\n\n## Other resources you might find useful:\n\n* [💡 how to process DICOM images to PNGs](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs)\n* [🤖 [fast.ai starter pack] train + inference 🚀](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference)\n\nLet's get started! 🚀","metadata":{}},{"cell_type":"markdown","source":"<a id=\"section-one\"></a>\n# Overview of data","metadata":{}},{"cell_type":"markdown","source":"In this competition our goal is to predict the presence or absence of cancer in mammography images.\n\n## How the data is organized\n\nThe images are organized in directories by patient id.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"code","source":"ls ../input/rsna-breast-cancer-detection/train_images | head -1 -n 4","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:22:38.519851Z","iopub.execute_input":"2024-04-29T11:22:38.521274Z","iopub.status.idle":"2024-04-29T11:22:43.549641Z","shell.execute_reply.started":"2024-04-29T11:22:38.521156Z","shell.execute_reply":"2024-04-29T11:22:43.548457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each directory contains several images.","metadata":{}},{"cell_type":"code","source":"ls ../input/rsna-breast-cancer-detection/train_images/57175","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:22:52.289345Z","iopub.execute_input":"2024-04-29T11:22:52.289793Z","iopub.status.idle":"2024-04-29T11:22:53.252902Z","shell.execute_reply.started":"2024-04-29T11:22:52.289747Z","shell.execute_reply":"2024-04-29T11:22:53.251515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nfrom matplotlib import pyplot as plt\n\nsample_sub = pd.read_csv('../input/rsna-breast-cancer-detection/sample_submission.csv')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-29T11:22:57.104138Z","iopub.execute_input":"2024-04-29T11:22:57.105062Z","iopub.status.idle":"2024-04-29T11:22:57.124594Z","shell.execute_reply.started":"2024-04-29T11:22:57.10501Z","shell.execute_reply":"2024-04-29T11:22:57.123624Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For each patient, we are to predict the presence or absence of cancer in the left and right breast.\n\nHere is the submission format.","metadata":{}},{"cell_type":"code","source":"sample_sub","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:23:03.524041Z","iopub.execute_input":"2024-04-29T11:23:03.524429Z","iopub.status.idle":"2024-04-29T11:23:03.551539Z","shell.execute_reply.started":"2024-04-29T11:23:03.524394Z","shell.execute_reply":"2024-04-29T11:23:03.550531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are 11913 patients in the train set.","metadata":{}},{"cell_type":"code","source":"!ls -l ../input/rsna-breast-cancer-detection/train_images | wc -l # wc -l outputs one a count of one row too many when used like this! hence we need to subtract one to get the true count","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:23:17.078668Z","iopub.execute_input":"2024-04-29T11:23:17.079112Z","iopub.status.idle":"2024-04-29T11:23:19.500949Z","shell.execute_reply.started":"2024-04-29T11:23:17.079075Z","shell.execute_reply":"2024-04-29T11:23:19.499736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In test, we have access to just a single example!","metadata":{}},{"cell_type":"code","source":"ls ../input/rsna-breast-cancer-detection/test_images/10008","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:23:26.864978Z","iopub.execute_input":"2024-04-29T11:23:26.865381Z","iopub.status.idle":"2024-04-29T11:23:27.831103Z","shell.execute_reply.started":"2024-04-29T11:23:26.865341Z","shell.execute_reply":"2024-04-29T11:23:27.829809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"\nHowever, once we submit from our notebook, the full test set will get mounted.\n\nThe images we will work with are in the DICOM format. Let us take a closer look at it.","metadata":{"execution":{"iopub.status.busy":"2022-11-29T00:15:53.09451Z","iopub.execute_input":"2022-11-29T00:15:53.095003Z","iopub.status.idle":"2022-11-29T00:15:53.103826Z","shell.execute_reply.started":"2022-11-29T00:15:53.094963Z","shell.execute_reply":"2022-11-29T00:15:53.101863Z"}}},{"cell_type":"markdown","source":"## DICOM image format overview\n\nA DICOM file consists of a header and image data sets packed into a single file. The information within the header is organized as a constant and standardized series of tags. By extracting data from these tags one can access important information regarding the patient demographics, study parameters, etc.\n\nLet's look at a couple of images to gain a better intution on what types of images we are working here with.","metadata":{}},{"cell_type":"code","source":"import pydicom\nimport numpy as np\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport os\nfrom pathlib import Path\nimport glob","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-29T11:23:54.883677Z","iopub.execute_input":"2024-04-29T11:23:54.88422Z","iopub.status.idle":"2024-04-29T11:23:55.620095Z","shell.execute_reply.started":"2024-04-29T11:23:54.884174Z","shell.execute_reply":"2024-04-29T11:23:55.619032Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"First, let's start with the metadata. Each image contains a rich metadata that might guide is in how we split the data for training (we might also want to provide some of this information directly to our model!)","metadata":{}},{"cell_type":"code","source":"example = '../input/rsna-breast-cancer-detection/train_images/10006/1459541791.dcm'\npydicom.dcmread(example)","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:24:00.243442Z","iopub.execute_input":"2024-04-29T11:24:00.243957Z","iopub.status.idle":"2024-04-29T11:24:00.405057Z","shell.execute_reply.started":"2024-04-29T11:24:00.24392Z","shell.execute_reply":"2024-04-29T11:24:00.404092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And now let us look at some images.","metadata":{"execution":{"iopub.status.busy":"2022-11-29T00:23:32.976227Z","iopub.execute_input":"2022-11-29T00:23:32.977415Z","iopub.status.idle":"2022-11-29T00:23:32.984473Z","shell.execute_reply.started":"2022-11-29T00:23:32.977365Z","shell.execute_reply":"2022-11-29T00:23:32.982948Z"}}},{"cell_type":"code","source":"# source: https://www.kaggle.com/code/allunia/rsna-csf-cervical-spine-fracture-eda/notebook\ndef rescale_img_to_hu(dcm_ds):\n    \"\"\"Rescales the image to Hounsfield unit.\"\"\"\n    data = dcm_ds.pixel_array\n    if dcm_ds.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    return data * dcm_ds.RescaleSlope + dcm_ds.RescaleIntercept","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-29T11:24:09.149024Z","iopub.execute_input":"2024-04-29T11:24:09.149905Z","iopub.status.idle":"2024-04-29T11:24:09.155653Z","shell.execute_reply.started":"2024-04-29T11:24:09.149864Z","shell.execute_reply":"2024-04-29T11:24:09.15462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def show_images_for_patient(patient_id):\n    patient_dir = os.path.join('../input/rsna-breast-cancer-detection/train_images', str(patient_id))\n    num_images = len(glob.glob(f\"{patient_dir}/*\"))\n    print(f\"Number of images for patient: {num_images}\")\n    fig, axs = plt.subplots(2, 2, figsize=(24,15))\n    axs = axs.flatten()\n    for i, img_path in enumerate(list(Path(patient_dir).iterdir())):\n        ds = pydicom.dcmread(img_path)\n        axs[i].imshow(rescale_img_to_hu(ds), cmap=\"bone\")","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2024-04-29T11:24:24.627659Z","iopub.execute_input":"2024-04-29T11:24:24.62868Z","iopub.status.idle":"2024-04-29T11:24:24.638033Z","shell.execute_reply.started":"2024-04-29T11:24:24.628629Z","shell.execute_reply":"2024-04-29T11:24:24.636783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"show_images_for_patient(10006)","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:24:32.004784Z","iopub.execute_input":"2024-04-29T11:24:32.005263Z","iopub.status.idle":"2024-04-29T11:24:50.101972Z","shell.execute_reply.started":"2024-04-29T11:24:32.005222Z","shell.execute_reply":"2024-04-29T11:24:50.100946Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we see from the plots, the images are of considerable size.\n\nWe see that the breast occupies only a small portion of the image.\n\nAs such (especially given the limited compute budget!) it will most likely be a very good idea to find ways of cropping out the portions of the image htat do not contain any information! The divider line can serve an important role here.","metadata":{}},{"cell_type":"markdown","source":"## The competition metric\n\nIt is always a good idea to understand exactly the type of predictions our model will be required to deliver.\n\nThe metric that the organizers opted for here is the probabilistic F1 score:\n\n$pF_1 = 2\\frac{pPrecision \\cdot pRecall}{pPrecision+pRecall}$\n\nwith:\n\n$pPrecision = \\frac{pTP}{pTP+pFP}$\n\n$pRecall = \\frac{pTP}{pTP+pFN}$\n\nYou can find a Python implementation of the metric [here](https://www.kaggle.com/code/sohier/probabilistic-f-score)\n\nOur model should output the likelihood of cancer in the corresponding image.\n\nSo what are the labels we will train on?","metadata":{}},{"cell_type":"code","source":"train_csv = pd.read_csv('../input/rsna-breast-cancer-detection/train.csv')\ntrain_csv.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:24:50.103572Z","iopub.execute_input":"2024-04-29T11:24:50.103976Z","iopub.status.idle":"2024-04-29T11:24:50.203846Z","shell.execute_reply.started":"2024-04-29T11:24:50.103946Z","shell.execute_reply":"2024-04-29T11:24:50.202781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The targets are stored in the `train.csv` file along with some additional metadata.\n\nThey are stored in the `cancer` column. Additionally, the file provides mapping of images to `laterality` (whether an image is of the left or right breast -- this can be very useful in training).\n\nAnd where can we find information on `laterality` (and other metadata) in test?\n\nIt is contained in the `test.csv` file.","metadata":{}},{"cell_type":"code","source":"test_csv = pd.read_csv('../input/rsna-breast-cancer-detection/test.csv')\ntest_csv","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:26:24.35329Z","iopub.execute_input":"2024-04-29T11:26:24.353691Z","iopub.status.idle":"2024-04-29T11:26:24.378007Z","shell.execute_reply.started":"2024-04-29T11:26:24.353655Z","shell.execute_reply":"2024-04-29T11:26:24.376978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, the test file doesn't contain the `cancer` column.\n\nSo how many images do we have to train on?","metadata":{"execution":{"iopub.status.busy":"2022-11-29T00:49:29.188171Z","iopub.execute_input":"2022-11-29T00:49:29.1886Z","iopub.status.idle":"2022-11-29T00:49:29.218817Z","shell.execute_reply.started":"2022-11-29T00:49:29.188566Z","shell.execute_reply":"2022-11-29T00:49:29.217571Z"}}},{"cell_type":"code","source":"train_csv.shape[0], train_csv.patient_id.nunique()","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:26:49.615988Z","iopub.execute_input":"2024-04-29T11:26:49.616389Z","iopub.status.idle":"2024-04-29T11:26:49.629775Z","shell.execute_reply.started":"2024-04-29T11:26:49.616348Z","shell.execute_reply":"2024-04-29T11:26:49.62877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a total of 54706 images in train across 11913 patients.","metadata":{}},{"cell_type":"markdown","source":"## Distribution of labels","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(20,6))\nplt.subplot(1,2,1)\nax1 = sns.countplot(data=train_csv, x='cancer')\nfor container in ax1.containers:\n    ax1.bar_label(container)\nplt.title('Distribution of targets');","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:26:53.38852Z","iopub.execute_input":"2024-04-29T11:26:53.389389Z","iopub.status.idle":"2024-04-29T11:26:53.606528Z","shell.execute_reply.started":"2024-04-29T11:26:53.38934Z","shell.execute_reply":"2024-04-29T11:26:53.605495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The classes are highly unbalanced! Most likely, we will need to address this in training (possibly via upsampling of the `cancer` class, downsampling of the `cancer absent`, class weights, etc).","metadata":{}},{"cell_type":"markdown","source":"<a id=\"section-two\"></a>\n# Training a fast.ai model","metadata":{}},{"cell_type":"markdown","source":"Let's now put all the pieces in place that we will need to trian a fast.ai model\n\nTo facilitate quick experimentation, I have processed [train images to PNGs](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369282) and shared them here: [RSNA Mammography -- images as PNGs (256px & 512px)](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs).\n\nHere is everything that we need in order to get up and runnig quickly with the fast.ai library.","metadata":{}},{"cell_type":"markdown","source":"## Creating dataloaders","metadata":{}},{"cell_type":"code","source":"from fastai.data.all import *\nfrom fastai.vision.all import *\n\npath = '../input/rsna-mammography-images-as-pngs/images_as_pngs/train_images_processed'\n\ntrain_csv = pd.read_csv('../input/rsna-breast-cancer-detection/train.csv')\nfn2label = {fn: cancer_or_not for fn, cancer_or_not in zip(train_csv['image_id'].astype('str'), train_csv['cancer'])}\n\ndef label_func(path):\n    return fn2label[path.stem]\n\ndblock = DataBlock(\n    blocks    = (ImageBlock, CategoryBlock),\n    get_items = get_image_files,\n    get_y = label_func,\n    splitter  = RandomSplitter()\n)\ndsets = dblock.datasets(path)\ndls = dblock.dataloaders(path)","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:27:10.398882Z","iopub.execute_input":"2024-04-29T11:27:10.399276Z","iopub.status.idle":"2024-04-29T11:29:53.991871Z","shell.execute_reply.started":"2024-04-29T11:27:10.399242Z","shell.execute_reply":"2024-04-29T11:29:53.990676Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And this is what a batch of our data will look like!","metadata":{}},{"cell_type":"code","source":"dls.show_batch()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating the learner and training for a single epoch","metadata":{}},{"cell_type":"code","source":"learn = vision_learner(dls, resnet18, metrics=error_rate, pretrained=False)\nlearn.fit_one_cycle(1, 1e-2)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There is a lot of functionality in the fast.ai toolkit that can be particularly useful in this competition:\n\n* training with one cycle (to make the best use of resources on Kaggle and to train better models)\n* the [medical imaging module](https://docs.fast.ai/medical.imaging.html) that we can use for reading DICOM files directly\n* image augmentation on the GPU (augmenting images is a compute-intensive process, on Kaggle we only have 2 CPU cores on GPU VMs so offloading augmenting our data to the super computer that a GPU is might lead to faster training)","metadata":{}},{"cell_type":"markdown","source":"<a id=\"section-three\"></a>\n# First submission\n\nLet's take a look at how to create a first submission to understand what format of predictions is expected (and how to output them).\n\nIn the notebook, we don't have access to the actual test set, that is until we submit! Once we finish editing our notebook and submit, the full test set will get mounted to our notebook.\n\nIt will follow the structure of the stub test set we have mounted currently, but will contain more than just the single directory.\n\nHere is what is currently mounted as the test set:","metadata":{}},{"cell_type":"code","source":"ls ../input/rsna-breast-cancer-detection/test_images","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:31:36.802213Z","iopub.execute_input":"2024-04-29T11:31:36.802639Z","iopub.status.idle":"2024-04-29T11:31:37.812423Z","shell.execute_reply.started":"2024-04-29T11:31:36.802603Z","shell.execute_reply":"2024-04-29T11:31:37.81095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's put all of the machinery in place we will need for outputting predictions (and test that we got it right by outputting a prediction consisting of random floats between 0 and 1).\n\nIf our submission gets scored, we will know we got this right.","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame(data={'prediction_id': test_csv['prediction_id'], 'cancer': np.random.rand(test_csv.shape[0])}).drop_duplicates(subset='prediction_id')\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:34:07.484554Z","iopub.execute_input":"2024-04-29T11:34:07.485155Z","iopub.status.idle":"2024-04-29T11:34:07.50807Z","shell.execute_reply.started":"2024-04-29T11:34:07.485104Z","shell.execute_reply":"2024-04-29T11:34:07.506793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-29T11:34:20.920528Z","iopub.execute_input":"2024-04-29T11:34:20.920956Z","iopub.status.idle":"2024-04-29T11:34:20.92789Z","shell.execute_reply.started":"2024-04-29T11:34:20.920918Z","shell.execute_reply":"2024-04-29T11:34:20.926792Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And that's it! Thank you very much for reading! 🙂\n\n**If you enjoyed the notebook, please upvote! 🙏 Thank you, appreciate your support!**\n\nHappy Kaggling 🥳\n","metadata":{}}]}