{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.11.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":39272,"databundleVersionId":4629629,"sourceType":"competition"}],"dockerImageVersionId":31089,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## TLDR\nThis program identifies cases of breast cancer in mammograms. This notebook will walk through the techniques I used and my thought process behind them. \n\n## Neural networks\nI will be using neural networks to analyze the data. Neural networks are the gold standard for image recognition. Skip to the data section of the notebook if you already know what a neural network is. \n\nNeural networks are similar to linear regression. In linear regression you are given a set of data points and attempt to find a line (of the form y = Mx + b) that models the data. Neural networks are a very general algorithm that finds an arbitrary equation to model the data (eg it does not have to be y = Mx + b: it could be y = Mx_1 + Zx_2 + b or really anything). The idea is that, given a large enough equation, you can classify almost everything including images.  \n","metadata":{}},{"cell_type":"markdown","source":"## DATA","metadata":{}},{"cell_type":"code","source":"#basic imports \nimport torch \nfrom torch.utils.data import Dataset, DataLoader\nfrom torchvision import transforms\nimport matplotlib.pyplot as plt\nimport pandas as pd\nimport pydicom\nimport warnings\nimport os\nimport numpy as np\nwarnings.filterwarnings(\"ignore\", category=RuntimeWarning)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-06T21:13:40.086183Z","iopub.execute_input":"2025-08-06T21:13:40.086609Z","iopub.status.idle":"2025-08-06T21:13:40.094603Z","shell.execute_reply.started":"2025-08-06T21:13:40.086573Z","shell.execute_reply":"2025-08-06T21:13:40.093009Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"The dataset is a collection of mammograms. Each mammogram is accompanied by metadata listing stuff like the patient's age, the machine ID, and whether or not the patient had cancer. \n\nBellow shows an example datapoint:","metadata":{}},{"cell_type":"code","source":"#metadata\ntrain_df = pd.read_csv('/kaggle/input/rsna-breast-cancer-detection/train.csv')\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-06T20:45:14.874546Z","iopub.execute_input":"2025-08-06T20:45:14.875344Z","iopub.status.idle":"2025-08-06T20:45:15.056493Z","shell.execute_reply.started":"2025-08-06T20:45:14.875235Z","shell.execute_reply":"2025-08-06T20:45:15.055512Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#image\npath = '/kaggle/input/rsna-breast-cancer-detection/train_images/10006/1459541791.dcm'\ndicom = pydicom.dcmread(path)\n\nplt.imshow(dicom.pixel_array, cmap='gray')\nplt.title(\"example image\")\nplt.axis('off')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-06T20:45:19.0941Z","iopub.execute_input":"2025-08-06T20:45:19.09445Z","iopub.status.idle":"2025-08-06T20:45:24.065496Z","shell.execute_reply.started":"2025-08-06T20:45:19.094425Z","shell.execute_reply":"2025-08-06T20:45:24.064209Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"A good first step is to determine if the dataset is balanced (i.e., are there the same number of positive cases as negative cases). Neural networks are lazy - if 99% of the mammograms are negative, then it won't bother to learn anything about the data and just guess negative every time. ","metadata":{}},{"cell_type":"code","source":"#checks % of positive vs negative in train\nprint(\"counts: \")\ncnts = train_df['cancer'].value_counts()\nprint(f\"positive: {cnts[1]}, negative: {cnts[0]}\")\n\nprint(\"\") #newline\n\nprint(\"%'s': \")\npres = train_df['cancer'].value_counts(normalize=True)\nprint(f\"positive: {pres[1]*100:.2f}%, negative: {pres[0]*100:.2f}%\")\n\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-06T20:51:09.488636Z","iopub.execute_input":"2025-08-06T20:51:09.488937Z","iopub.status.idle":"2025-08-06T20:51:09.497628Z","shell.execute_reply.started":"2025-08-06T20:51:09.488919Z","shell.execute_reply":"2025-08-06T20:51:09.496416Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"About 98% of cases are negative.\n\nI will be using random downsampling to balance the dataset. Essentially, I will discard a handful of random negative cases to make the quantities the same.","metadata":{}},{"cell_type":"code","source":"#random downsample\npositive_count = cnts[1]\ntrain_df = pd.concat([train_df[train_df['cancer'] == 1], train_df[train_df['cancer'] == 0].sample(n=positive_count)]).sample(frac=1)  \n\nprint(\"updated counts: \")\ncnts = train_df['cancer'].value_counts()\nprint(f\"positive: {cnts[1]}, negative: {cnts[0]}\")\n\nprint(\"\") #newline\n\nprint(\"updated %'s': \")\npres = train_df['cancer'].value_counts(normalize=True)\nprint(f\"positive: {pres[1]*100:.2f}%, negative: {pres[0]*100:.2f}%\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-06T20:54:22.373499Z","iopub.execute_input":"2025-08-06T20:54:22.373927Z","iopub.status.idle":"2025-08-06T20:54:22.388162Z","shell.execute_reply.started":"2025-08-06T20:54:22.373891Z","shell.execute_reply":"2025-08-06T20:54:22.386949Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"1158 is a very small sample size, but oh well. \n\nThe next step is to set up a way to load the data. (esentially the database is too big to store in ram so we need to load it bit by bit)","metadata":{}},{"cell_type":"code","source":"class RSNADataset(Dataset):\n    def __init__(self, df, image_dir, transform=None):\n        self.df = df\n        self.image_dir = image_dir\n        self.transform = transform\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx):\n        row = self.df.iloc[idx]\n        patient_id = row['patient_id']\n        image_id = row['image_id']\n        label = row['cancer']\n\n        image_path = os.path.join(self.image_dir, str(patient_id), f\"{image_id}.dcm\")\n        dicom = pydicom.dcmread(image_path)\n        image = dicom.pixel_array.astype(np.float32)\n\n        image -= image.min()\n        image /= image.max()\n\n        image = torch.tensor(image).unsqueeze(0)\n\n        if self.transform:\n            image = self.transform(image)\n\n        return image, torch.tensor(label, dtype=torch.float32)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-06T21:04:15.346674Z","iopub.execute_input":"2025-08-06T21:04:15.347002Z","iopub.status.idle":"2025-08-06T21:04:15.354802Z","shell.execute_reply.started":"2025-08-06T21:04:15.346976Z","shell.execute_reply":"2025-08-06T21:04:15.35371Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"image_dir = \"/kaggle/input/rsna-breast-cancer-detection/train_images\"\n\ndataset = RSNADataset(train_df, image_dir)\ndataloader = DataLoader(dataset, batch_size=16, shuffle=True, num_workers=2)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-06T21:06:55.697167Z","iopub.execute_input":"2025-08-06T21:06:55.69767Z","iopub.status.idle":"2025-08-06T21:06:55.702799Z","shell.execute_reply.started":"2025-08-06T21:06:55.697643Z","shell.execute_reply":"2025-08-06T21:06:55.701855Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"for images, labels in dataloader:\n    print(images.shape)  # [batch_size, 1, H, W]\n    print(labels)\n    break","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2025-08-06T21:09:49.391741Z","iopub.execute_input":"2025-08-06T21:09:49.392039Z","iopub.status.idle":"2025-08-06T21:10:07.048635Z","shell.execute_reply.started":"2025-08-06T21:09:49.392018Z","shell.execute_reply":"2025-08-06T21:10:07.046507Z"}},"outputs":[],"execution_count":null}]}