{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom pathlib import Path\nfrom typing import List, Union, Any, Tuple, Dict\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport sys\nfrom glob import glob\nfrom fastai.vision.all import *\nimport matplotlib.pyplot as plt\n\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-01-08T07:29:06.342016Z","iopub.execute_input":"2023-01-08T07:29:06.342387Z","iopub.status.idle":"2023-01-08T07:29:06.349358Z","shell.execute_reply.started":"2023-01-08T07:29:06.342356Z","shell.execute_reply":"2023-01-08T07:29:06.348336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"working = Path(r'/kaggle/input/timm-library-for-offline-submission')\nsys.path.insert(1, working)","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:29:33.396204Z","iopub.execute_input":"2023-01-08T07:29:33.396603Z","iopub.status.idle":"2023-01-08T07:29:33.401297Z","shell.execute_reply.started":"2023-01-08T07:29:33.396575Z","shell.execute_reply":"2023-01-08T07:29:33.400547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- timm library is required for offline inference","metadata":{}},{"cell_type":"code","source":"!pip install {working}/timm-0.6.12-py3-none-any.whl\n!pip install {working}/pydicom-2.3.1-py3-none-any.whl\n!pip install {working}/dicomsdl-0.109.1-cp37-cp37m-manylinux_2_12_x86_64.manylinux2010_x86_64.whl","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:29:33.685771Z","iopub.execute_input":"2023-01-08T07:29:33.686398Z","iopub.status.idle":"2023-01-08T07:30:59.975859Z","shell.execute_reply.started":"2023-01-08T07:29:33.686363Z","shell.execute_reply":"2023-01-08T07:30:59.974639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- I have created first trained model [here](https://www.kaggle.com/code/hasangoni/first-submission-with-smaller-dataset/notebook).\n- It is a first iteration, just creating a pipeline to see everything is working or not.\n- Here we will used that learner(trained model) to make inference and submit our model inference.\n- I have not used actual competition data, but some proecessed data. One can have a look on the dataset added to this notebook. The dataset can be found by name rsna-breast-cancer-detection-poi-images. From this ROI images, I have created another dataset which is smaller and one iteration took almost 1 minute. I have also published those smaller dataset [here](https://www.kaggle.com/datasets/hasangoni/rsna-small-for-faster-experimentation).\n- Smaller dataset is just train-test split with same distribution of actual dataset. just 10 times smaller.","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"#import pydicom\nimport dicomsdl\nimport numpy as np\nimport cv2\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.notebook import tqdm\nfrom pathlib import Path\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\nfrom glob import glob\nfrom functools import partial\nimport multiprocessing as mp\nfrom fastcore import parallel","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:30:59.977356Z","iopub.execute_input":"2023-01-08T07:30:59.977651Z","iopub.status.idle":"2023-01-08T07:30:59.983733Z","shell.execute_reply.started":"2023-01-08T07:30:59.977624Z","shell.execute_reply":"2023-01-08T07:30:59.982494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import timm\nfrom pathlib import Path\nfrom fastai.vision.all import *","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:12:23.152803Z","iopub.execute_input":"2023-01-08T07:12:23.153148Z","iopub.status.idle":"2023-01-08T07:12:23.1672Z","shell.execute_reply.started":"2023-01-08T07:12:23.153122Z","shell.execute_reply":"2023-01-08T07:12:23.166415Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Defining Paths","metadata":{}},{"cell_type":"code","source":"\npath = Path(r'/kaggle/input/first-submission-with-smaller-dataset')\ncompetition_path = Path(r'/kaggle/input/rsna-breast-cancer-detection')\ntest_image_path = Path(competition_path/'test_images')\nsmall_test = Path(r'/kaggle/input/rsna-breast-cancer-detection-poi-images/bc_768_roi/bc_768_roi/test')\ntmp_dir = Path(r'/kaggle/temp/')\ntmp_dir.mkdir(exist_ok=True, parents=True)\ntmp_dir","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:34.508538Z","iopub.execute_input":"2023-01-08T07:31:34.508951Z","iopub.status.idle":"2023-01-08T07:31:34.51756Z","shell.execute_reply.started":"2023-01-08T07:31:34.508919Z","shell.execute_reply":"2023-01-08T07:31:34.516091Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading learner\n- When we load the learner, we need all the functions we used during traiing. Therefore adding those functions and variables","metadata":{}},{"cell_type":"code","source":"df_trn = pd.read_csv(f'{competition_path}/train.csv')\ndf_trn=(\n    df_trn\n    .assign(image_id = lambda df: df['image_id'].astype(str))\n)\ndf_trn.head()\ndef get_y_label(x):\n    return df_trn.loc[df_trn['image_id'] == x.stem.split('_')[1], 'cancer'].values[0]\n\nlearn_inf = load_learner(path/'export.pkl')","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:36.122214Z","iopub.execute_input":"2023-01-08T07:31:36.122572Z","iopub.status.idle":"2023-01-08T07:31:36.256711Z","shell.execute_reply.started":"2023-01-08T07:31:36.122542Z","shell.execute_reply":"2023-01-08T07:31:36.25555Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_tst = pd.read_csv(competition_path/'test.csv')\ndf_tst=(\n    df_tst\n    .assign(patient_id = lambda df_: df_['patient_id'].astype(str),\n           image_id = lambda df_: df_['image_id'].astype(str),\n           prediction_id = lambda df_: df_['patient_id'].astype(str)+\"_\"+df_['laterality'])\n)\ndf_tst","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:37.617004Z","iopub.execute_input":"2023-01-08T07:31:37.617441Z","iopub.status.idle":"2023-01-08T07:31:37.642876Z","shell.execute_reply.started":"2023-01-08T07:31:37.617408Z","shell.execute_reply":"2023-01-08T07:31:37.641736Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_pred_id_from_file_name(\n                               test_path:Path=competition_path/'test_images',\n                               df:pd.DataFrame=df_tst)->str:\n    patient_id = Path(test_path).parent.name\n    image_id = Path(test_path).stem\n\n    \n    lat = df.query('patient_id == @patient_id & image_id == @image_id')['laterality'].values[0]\n    return f'{patient_id}_{lat}'","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:38.51304Z","iopub.execute_input":"2023-01-08T07:31:38.513479Z","iopub.status.idle":"2023-01-08T07:31:38.519658Z","shell.execute_reply.started":"2023-01-08T07:31:38.513445Z","shell.execute_reply":"2023-01-08T07:31:38.518809Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#dir(learn_inf)","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:39.674126Z","iopub.execute_input":"2023-01-08T07:31:39.674671Z","iopub.status.idle":"2023-01-08T07:31:39.67877Z","shell.execute_reply.started":"2023-01-08T07:31:39.674637Z","shell.execute_reply":"2023-01-08T07:31:39.677517Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Processing test images","metadata":{}},{"cell_type":"code","source":"# stolen from here https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs\n\n\ndef image_resize(\n                   image,\n                   width=None,\n                   height=None,\n                   inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n\n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n\n    return resized\n\n\ndef dcm_array(path:Path,\n              im_size:Tuple[int,int]\n              )->np.ndarray:\n    dcm_file = dicomsdl.open(str(path))\n    data = dcm_file.pixelData()\n\n    data = (data - data.min()) / (data.max() - data.min())\n\n    if dcm_file.getPixelDataInfo()['PhotometricInterpretation'] == \"MONOCHROME1\":\n        data = 1 - data\n\n    data = cv2.resize(data, im_size)\n    data = (data * 255).astype(np.uint8)\n    return data\n\n\ndef dcm_to_array(path:Union[Path,str], # dicom files name w\n                 im_size:int # image size e.g. (1024, 1024)\n                ):\n    dicom = pydicom.dcmread(path)\n    data = dicom.pixel_array\n    data = apply_voi_lut(dicom.pixel_array, dicom)\n       \n    data = (data - data.min()) / (data.max() - data.min())\n    \n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = 1 - data\n    \n    h, w = data.shape\n    if w > h:\n         data = image_resize(data, width = im_size)\n    else:\n         data = image_resize(data, height = im_size)\n    \n    data = (data * 255).astype(np.uint8)\n    return data\n\n","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:41.082954Z","iopub.execute_input":"2023-01-08T07:31:41.083582Z","iopub.status.idle":"2023-01-08T07:31:41.095938Z","shell.execute_reply.started":"2023-01-08T07:31:41.083542Z","shell.execute_reply":"2023-01-08T07:31:41.094783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from glob import glob","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:42.110579Z","iopub.execute_input":"2023-01-08T07:31:42.111257Z","iopub.status.idle":"2023-01-08T07:31:42.115588Z","shell.execute_reply.started":"2023-01-08T07:31:42.11122Z","shell.execute_reply":"2023-01-08T07:31:42.114794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def get_dcm_files(\n                    path:Union[Path, str]=test_image_path)->List[Path]:\n    return glob(f'{path}/**/*.dcm', recursive=True)\n    ","metadata":{"execution":{"iopub.status.busy":"2023-01-08T08:01:57.767189Z","iopub.execute_input":"2023-01-08T08:01:57.767511Z","iopub.status.idle":"2023-01-08T08:01:57.772924Z","shell.execute_reply.started":"2023-01-08T08:01:57.767485Z","shell.execute_reply":"2023-01-08T08:01:57.77198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_images = get_dcm_files()","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:46.298987Z","iopub.execute_input":"2023-01-08T07:31:46.299386Z","iopub.status.idle":"2023-01-08T07:31:46.306887Z","shell.execute_reply.started":"2023-01-08T07:31:46.299353Z","shell.execute_reply":"2023-01-08T07:31:46.305776Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Some try to work on parallel processing\n- With Ipython notebook parallel working is little bit difficult\n- Sometimes it hangs for a longer period.\n- See for details [here](https://stackoverflow.com/questions/34086112/python-multiprocessing-pool-stuck)","metadata":{}},{"cell_type":"code","source":"def pr_image(file_name:Path)->float:\n    im = dcm_array(file_name, (1024, 1024))\n    idx, _, prob = learn_inf.predict(im)\n    pr_id = get_pred_id_from_file_name(file_name,df_tst)\n    return prob.numpy()[1], pr_id","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:31:50.843351Z","iopub.execute_input":"2023-01-08T07:31:50.843746Z","iopub.status.idle":"2023-01-08T07:31:50.848326Z","shell.execute_reply.started":"2023-01-08T07:31:50.843712Z","shell.execute_reply":"2023-01-08T07:31:50.847589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from concurrent.futures import ProcessPoolExecutor, ThreadPoolExecutor","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:45:01.62531Z","iopub.execute_input":"2023-01-08T07:45:01.62573Z","iopub.status.idle":"2023-01-08T07:45:01.631737Z","shell.execute_reply.started":"2023-01-08T07:45:01.625698Z","shell.execute_reply":"2023-01-08T07:45:01.630228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n#with ThreadPoolExecutor(max_workers=60) as exec_:\n    #res = exec_.map(pr_image, test_images)\n#for i in res:\n    #print(i)","metadata":{"execution":{"iopub.status.busy":"2023-01-08T08:01:02.586359Z","iopub.execute_input":"2023-01-08T08:01:02.586756Z","iopub.status.idle":"2023-01-08T08:01:02.594347Z","shell.execute_reply.started":"2023-01-08T08:01:02.58672Z","shell.execute_reply":"2023-01-08T08:01:02.593097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#sys.modules['__main__'].__file__ = 'ipython'","metadata":{"execution":{"iopub.status.busy":"2023-01-08T07:58:18.032108Z","iopub.execute_input":"2023-01-08T07:58:18.032536Z","iopub.status.idle":"2023-01-08T07:58:18.037645Z","shell.execute_reply.started":"2023-01-08T07:58:18.032498Z","shell.execute_reply":"2023-01-08T07:58:18.036773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n#for i in test_images:\n    #print(pr_image(i))","metadata":{"execution":{"iopub.status.busy":"2023-01-08T08:01:11.966968Z","iopub.execute_input":"2023-01-08T08:01:11.967404Z","iopub.status.idle":"2023-01-08T08:01:11.973898Z","shell.execute_reply.started":"2023-01-08T08:01:11.967366Z","shell.execute_reply":"2023-01-08T08:01:11.972513Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Without Parallel processing creating final submission df","metadata":{"execution":{"iopub.status.busy":"2023-01-08T08:08:04.870793Z","iopub.execute_input":"2023-01-08T08:08:04.871198Z","iopub.status.idle":"2023-01-08T08:08:04.875432Z","shell.execute_reply.started":"2023-01-08T08:08:04.871168Z","shell.execute_reply":"2023-01-08T08:08:04.874554Z"}}},{"cell_type":"code","source":"def process_test_image(\n                       tst_files:List[Path], # test files \n                       df_tst:pd.DataFrame, # dataframe containing information \n                        )->pd.DataFrame: # dataframe containing prediciton id and prediciton \n                        \n    \n    if 'prediction_id' not in df_tst.columns:\n         df_tst = df_tst.assign(\n                              prediction_id = lambda df_: df_['patient_id'].astype(str)+\"_\"+df_['laterality']\n                               )\n    prediction_ids, probs = [],[]\n    for i in tst_files:\n        \n        probs.append(pr_image(i))\n        pred_id = get_pred_id_from_file_name(i, df_tst)\n        prediction_ids.append(pred_id )\n    return (\n        pd.DataFrame(prediction_ids)\n        .assign(\n            cancer = lambda df:probs)\n        .rename(columns={0:'prediction_id'})\n        .groupby('prediction_id', as_index=False)['cancer'].mean()\n       )","metadata":{"execution":{"iopub.status.busy":"2023-01-08T08:07:09.706291Z","iopub.execute_input":"2023-01-08T08:07:09.706634Z","iopub.status.idle":"2023-01-08T08:07:09.714578Z","shell.execute_reply.started":"2023-01-08T08:07:09.706607Z","shell.execute_reply":"2023-01-08T08:07:09.713553Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ss = process_test_image(test_images, df_tst)","metadata":{"execution":{"iopub.status.busy":"2023-01-08T08:07:16.196084Z","iopub.execute_input":"2023-01-08T08:07:16.196407Z","iopub.status.idle":"2023-01-08T08:07:18.54472Z","shell.execute_reply.started":"2023-01-08T08:07:16.196382Z","shell.execute_reply":"2023-01-08T08:07:18.543589Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ss.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-08T08:07:21.038228Z","iopub.execute_input":"2023-01-08T08:07:21.03877Z","iopub.status.idle":"2023-01-08T08:07:21.048751Z","shell.execute_reply.started":"2023-01-08T08:07:21.038742Z","shell.execute_reply":"2023-01-08T08:07:21.047652Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ss.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-07T13:29:59.729468Z","iopub.execute_input":"2023-01-07T13:29:59.729935Z","iopub.status.idle":"2023-01-07T13:29:59.739627Z","shell.execute_reply.started":"2023-01-07T13:29:59.729884Z","shell.execute_reply":"2023-01-07T13:29:59.738725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- At last I got it working. Thanks to @radek1 for his comments\n- For a test purpose I first read sample submission file and then make is output to submission.csv, then it was working.\n- I then tried to use poi dataset test images to infer, but this is wrong[which should be very normal]. I just didnot think enough on it. As sample submission was working I thought something wrong with my submission file.\n- Right now this notebook using actual competition test data (not poi images anymore) as a dcm format, then process it for my model inference.\n- Upto now I just wanted to make sure, this pipeline is working for me or not from starting to submission. Now that I know it's working, I can perform all my experiments.\n- Huge thanks to @radek1, without him may be I just got frustated and left the competition (just like I have done last 2 years, sevaral times)\n","metadata":{}},{"cell_type":"markdown","source":"# todo\n\n- Competition metric addition in train pipeline\n- Use all tricks for CV\n  - test time augmentaion\n  - Gradient accumulation\n  - progressive resizing\n  - Augmentation like (cut mix)\n- how to handle imbalance set\n  - weighting is loss\n  - just upscaling the images\n- kfold validation\n- ensemble of different models\n\n- may be try also some image patching\n- see CAM map of success models","metadata":{"execution":{"iopub.status.busy":"2023-01-08T08:13:18.154901Z","iopub.execute_input":"2023-01-08T08:13:18.155249Z","iopub.status.idle":"2023-01-08T08:13:18.160668Z","shell.execute_reply.started":"2023-01-08T08:13:18.15522Z","shell.execute_reply":"2023-01-08T08:13:18.158733Z"}}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"markdown","source":"- Feel free to upvote if it helps you. It will keep me motivating.","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}