{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        #print(os.path.join(dirname, filename))\n        pass\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-03T21:05:30.116076Z","iopub.execute_input":"2023-01-03T21:05:30.116808Z","iopub.status.idle":"2023-01-03T21:06:02.523109Z","shell.execute_reply.started":"2023-01-03T21:05:30.116698Z","shell.execute_reply":"2023-01-03T21:06:02.522109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%load_ext autoreload\n%autoreload 2","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:02.525125Z","iopub.execute_input":"2023-01-03T21:06:02.525873Z","iopub.status.idle":"2023-01-03T21:06:02.560479Z","shell.execute_reply.started":"2023-01-03T21:06:02.525834Z","shell.execute_reply":"2023-01-03T21:06:02.559617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom pathlib import Path","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:02.561741Z","iopub.execute_input":"2023-01-03T21:06:02.562465Z","iopub.status.idle":"2023-01-03T21:06:02.590064Z","shell.execute_reply.started":"2023-01-03T21:06:02.562425Z","shell.execute_reply":"2023-01-03T21:06:02.589088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- First I have tried to work on preprocessed dicom images, which is published [here](https://www.kaggle.com/datasets/remekkinas/rsna-breast-cancer-detection-poi-images)\n- In my local laptop which has 4 GB GPU it took around 40 minutes to run one epoch. I thought to create a smaller dataset with smaller image size and smaller number of images.\n_ Those smaller data can be found [here](https://www.kaggle.com/datasets/hasangoni/rsna-small-for-faster-experimentation).\n- With those data it took me around 1minutes for one epoch. Now I can create a quick pipeline, to see every thing works or not. \n","metadata":{}},{"cell_type":"code","source":"!pip install -Uq fastai\n!pip install timm","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:02.592866Z","iopub.execute_input":"2023-01-03T21:06:02.593247Z","iopub.status.idle":"2023-01-03T21:06:25.151246Z","shell.execute_reply.started":"2023-01-03T21:06:02.593212Z","shell.execute_reply":"2023-01-03T21:06:25.150145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from pathlib import Path\nfrom fastai.vision.all import *","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:25.153259Z","iopub.execute_input":"2023-01-03T21:06:25.153663Z","iopub.status.idle":"2023-01-03T21:06:28.539531Z","shell.execute_reply.started":"2023-01-03T21:06:25.153622Z","shell.execute_reply":"2023-01-03T21:06:28.538401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Defining paths","metadata":{}},{"cell_type":"code","source":"root_path = Path(r'/kaggle/input/')\n\nactual_path = Path(fr'{root_path}/rsna-breast-cancer-detection')\nsmall_image_path = Path(fr'{root_path}/rsna-small-for-faster-experimentation/data_small')\n                 \n# checking how much files are avilable    \nsmall_image_path.ls()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:28.541634Z","iopub.execute_input":"2023-01-03T21:06:28.542418Z","iopub.status.idle":"2023-01-03T21:06:28.609179Z","shell.execute_reply.started":"2023-01-03T21:06:28.542377Z","shell.execute_reply":"2023-01-03T21:06:28.608338Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fn = get_image_files(small_image_path)\nlen(fn)","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:28.611246Z","iopub.execute_input":"2023-01-03T21:06:28.611932Z","iopub.status.idle":"2023-01-03T21:06:29.658273Z","shell.execute_reply.started":"2023-01-03T21:06:28.611893Z","shell.execute_reply":"2023-01-03T21:06:29.657363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"actual_path.ls()","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:29.665754Z","iopub.execute_input":"2023-01-03T21:06:29.668161Z","iopub.status.idle":"2023-01-03T21:06:29.739354Z","shell.execute_reply.started":"2023-01-03T21:06:29.668116Z","shell.execute_reply":"2023-01-03T21:06:29.738355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Reading Training csv files and process little bit ","metadata":{}},{"cell_type":"code","source":"df_trn = pd.read_csv(f'{actual_path}/train.csv')\ndf_trn=(\n    df_trn\n    .assign(image_id = lambda df: df['image_id'].astype(str))\n)\ndf_trn.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:29.743751Z","iopub.execute_input":"2023-01-03T21:06:29.745903Z","iopub.status.idle":"2023-01-03T21:06:30.018874Z","shell.execute_reply.started":"2023-01-03T21:06:29.745865Z","shell.execute_reply":"2023-01-03T21:06:30.017995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating a function to get label name from the file path and training dataframe","metadata":{}},{"cell_type":"code","source":"fn_id = [ i.stem.split('_')[1] for i in fn]\nlabel_df = [df_trn.loc[df_trn['image_id'] == i, 'cancer'].values[0] for i in fn_id]\ndef get_y_label(x):\n    return df_trn.loc[df_trn['image_id'] == x.stem.split('_')[1], 'cancer'].values[0]","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:30.02522Z","iopub.execute_input":"2023-01-03T21:06:30.027502Z","iopub.status.idle":"2023-01-03T21:06:50.26083Z","shell.execute_reply.started":"2023-01-03T21:06:30.027463Z","shell.execute_reply":"2023-01-03T21:06:50.259814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating a dataloader to train ","metadata":{}},{"cell_type":"code","source":"dls = ImageDataLoaders.from_path_func(\n                                small_image_path,# path of the image files\n                                fnames=fn,# file names\n                                label_func=get_y_label,# function to get the label\n                                valid_pct=0.2,# percentage of validation data\n                                seed=42,# now seed to get repetitive result\n                                item_tfms=Resize(166, method='squish'),# resize image using squash method\n                                batch_tfms=aug_transforms(size=128, min_scale=0.75) # in Gpu batch augmenation\n                                )\ndls.device = default_device() # I guess it is not required, it automatically select the default_device, but just making sure.","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:06:50.315225Z","iopub.execute_input":"2023-01-03T21:06:50.316186Z","iopub.status.idle":"2023-01-03T21:07:14.488063Z","shell.execute_reply.started":"2023-01-03T21:06:50.316144Z","shell.execute_reply":"2023-01-03T21:07:14.487052Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"default_device()","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:07:14.489432Z","iopub.execute_input":"2023-01-03T21:07:14.489805Z","iopub.status.idle":"2023-01-03T21:07:14.535618Z","shell.execute_reply.started":"2023-01-03T21:07:14.489768Z","shell.execute_reply":"2023-01-03T21:07:14.534649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Seeing whether our batch works","metadata":{}},{"cell_type":"code","source":"dls.show_batch(max_n=6)","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:07:14.538507Z","iopub.execute_input":"2023-01-03T21:07:14.538798Z","iopub.status.idle":"2023-01-03T21:07:16.09156Z","shell.execute_reply.started":"2023-01-03T21:07:14.538772Z","shell.execute_reply":"2023-01-03T21:07:16.090547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Creating a dataloader telling structure, metrics, and use floating point 16","metadata":{}},{"cell_type":"code","source":"learn = vision_learner(dls, 'resnet26d', metrics=error_rate, path='.').to_fp16()","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:07:16.092672Z","iopub.execute_input":"2023-01-03T21:07:16.092981Z","iopub.status.idle":"2023-01-03T21:07:20.263032Z","shell.execute_reply.started":"2023-01-03T21:07:16.092952Z","shell.execute_reply":"2023-01-03T21:07:20.261991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Searching for learing rate. Nice and easy function developed by fastai","metadata":{}},{"cell_type":"code","source":"learn.lr_find(suggest_funcs=(valley, slide))","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:07:20.264471Z","iopub.execute_input":"2023-01-03T21:07:20.264856Z","iopub.status.idle":"2023-01-03T21:08:17.796887Z","shell.execute_reply.started":"2023-01-03T21:07:20.264812Z","shell.execute_reply":"2023-01-03T21:08:17.795752Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Training our learner","metadata":{}},{"cell_type":"code","source":"learn.fine_tune(3, 0.012)","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:08:36.46241Z","iopub.execute_input":"2023-01-03T21:08:36.463572Z","iopub.status.idle":"2023-01-03T21:11:25.080578Z","shell.execute_reply.started":"2023-01-03T21:08:36.463514Z","shell.execute_reply":"2023-01-03T21:11:25.079348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Saving our learner to make offline submission","metadata":{}},{"cell_type":"code","source":"learn.export()","metadata":{"execution":{"iopub.status.busy":"2023-01-03T21:13:12.372828Z","iopub.execute_input":"2023-01-03T21:13:12.373699Z","iopub.status.idle":"2023-01-03T21:13:12.621474Z","shell.execute_reply.started":"2023-01-03T21:13:12.373651Z","shell.execute_reply":"2023-01-03T21:13:12.620436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- with learn.export() a pickle file is generated and we need to load it during inference","metadata":{}},{"cell_type":"code","source":"ss = pd.read_csv(f'{actual_path}/sample_submission.csv')\ndf_tst = pd.read_csv(f'{actual_path}/test.csv')\ndf_tst.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Now we can do first experimentation, because each epoch requires approximate 1 minute","metadata":{}},{"cell_type":"markdown","source":"# Just seeing test images and reading submission files","metadata":{}},{"cell_type":"code","source":"ss.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ls ../input/rsna-breast-cancer-detection/test_images/10008","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame(\n                         data={\n                             'prediction_id': df_tst['prediction_id'], \n                             'cancer': np.random.rand(df_tst.shape[0])}\n                        ).drop_duplicates(subset='prediction_id')\nsubmission.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"learn.save(r'/kaggle/working/learner_first')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- One problem will be during submission, is that we need offline inference.\n- So we need to save timm library.\n- I have used fastkaggle library to save the timm library. So we use first with internet connection pip install fastkaggle library, then save the timm libray\n- [Here](https://www.kaggle.com/code/hasangoni/timm-library-for-offline-submission) one can find the timm library.\n- Then one just need to add those data in the offline notebook and use pip install.\n- For offline submission, I have published [here](https://www.kaggle.com/code/hasangoni/offline-inference/edit/run/115595817), I am getting ```Submission Scoring Error``. Don't know what is happening here. -> solution is actually to process the test images from dicom files to files for the learner. Right now the offline submission works. Thanks to @radek1 for his useful tips.\n","metadata":{"execution":{"iopub.status.busy":"2023-01-05T19:55:12.63809Z","iopub.execute_input":"2023-01-05T19:55:12.638596Z","iopub.status.idle":"2023-01-05T19:55:12.659331Z","shell.execute_reply.started":"2023-01-05T19:55:12.638553Z","shell.execute_reply":"2023-01-05T19:55:12.656188Z"}}},{"cell_type":"markdown","source":"- IF this notebook is helpful, please try to upvote it. This will help me to keep me motivated and publish more notebook like this.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}