{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"This notebook will walk you through how to create very large datasets (>20GB) directly from inside a notebook. This is a little bit more complex than generating the dataset from the output, but it's more flexible and not capped at 20GB; in fact, you can create multiple datasets as well.","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true}},{"cell_type":"markdown","source":"shutout to these two notebooks for helping me out:\n\nhttps://www.kaggle.com/code/xhlulu/how-to-create-very-large-datasets-from-a-notebook on how to upload big datasets on kaggle\n\nhttps://www.kaggle.com/code/awsaf49/rsna-bcd-efficientnet-tf-tpu-1vm-infer fast preprocessing using dicomsdl","metadata":{}},{"cell_type":"markdown","source":"## Preliminary steps\n\nBefore starting, you will first need to get an API token. To do that, navigate to \"Account\" -> \"API\":\n\n![](https://i.imgur.com/ZKxDCvB.png)\n\nThen, click on \"Create New API Token\":\n\n![](https://i.imgur.com/46OOLkv.png)\n\nIt will download a JSON file. Open the JSON file; you will have a *username* and a *key*. Now, click on \"Add-ons\" -> \"Secrets\":\n\n![](https://i.imgur.com/sSIRm8X.png)\n\nClick on \"Add a new secret\" and set the value of `KAGGLE_KEY` to the key you just got, and `KAGGLE_USERNAME` to the username you got.\n\n![](https://i.imgur.com/juigYPK.png)\n![](https://i.imgur.com/g0UIB2l.png)","metadata":{}},{"cell_type":"code","source":"import os\nimport json\n\nfrom kaggle_secrets import UserSecretsClient","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:27:46.401052Z","iopub.execute_input":"2023-01-20T15:27:46.401697Z","iopub.status.idle":"2023-01-20T15:27:46.418027Z","shell.execute_reply.started":"2023-01-20T15:27:46.401661Z","shell.execute_reply":"2023-01-20T15:27:46.417267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Setting up environment variables\n\nAt this point, you will need to set up the environment variables in order to use the `kaggle` CLI. We will retrieve them from the user secrets that we just created.","metadata":{}},{"cell_type":"code","source":"secrets = UserSecretsClient()\n\nos.environ['KAGGLE_USERNAME'] = secrets.get_secret(\"KAGGLE_USERNAME\")\nos.environ['KAGGLE_KEY'] = secrets.get_secret(\"KAGGLE_KEY\")","metadata":{"execution":{"iopub.status.busy":"2023-01-20T16:01:28.504744Z","iopub.execute_input":"2023-01-20T16:01:28.505122Z","iopub.status.idle":"2023-01-20T16:01:57.189109Z","shell.execute_reply.started":"2023-01-20T16:01:28.505093Z","shell.execute_reply":"2023-01-20T16:01:57.188142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# make sure dataset dir does not exist\n!rm -rf /kaggle/dataset/","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:34:05.51644Z","iopub.execute_input":"2023-01-20T15:34:05.517277Z","iopub.status.idle":"2023-01-20T15:34:06.228069Z","shell.execute_reply.started":"2023-01-20T15:34:05.51724Z","shell.execute_reply":"2023-01-20T15:34:06.226903Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Creating the metadata file\n\nNow, let's create the folder that will hold our dataset and create the metadata file:","metadata":{}},{"cell_type":"code","source":"os.makedirs('/kaggle/dataset/', exist_ok=True)\n\n\n# Change below\nmeta = dict(\n    id=f\"{os.getenv('KAGGLE_USERNAME')}/png-1536x960-dataset\",\n    title=\"RSNA Mammography PNG Dataset\",\n    isPrivate=True,\n    licenses=[dict(name=\"other\")]\n)\n\nwith open('/kaggle/dataset/dataset-metadata.json', 'w') as f:\n    json.dump(meta, f)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T16:02:27.555202Z","iopub.execute_input":"2023-01-20T16:02:27.55564Z","iopub.status.idle":"2023-01-20T16:08:31.362585Z","shell.execute_reply.started":"2023-01-20T16:02:27.555606Z","shell.execute_reply":"2023-01-20T16:08:31.361597Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above, we are manually creating the metadata file. Alternatively, you can call `!kaggle init -p /kaggle/dataset` but you will still need to edit the file using the `json` python library.\n\n## Dataset generation","metadata":{}},{"cell_type":"code","source":"# installs\n\n!pip install -q python_gdcm\n!pip install -q pylibjpeg\n!pip install -q dicomsdl\n!apt-get -qq update && apt-get -qq install -y python3-opencv\n!pip install opencv-python","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:36:11.030923Z","iopub.execute_input":"2023-01-20T15:36:11.03184Z","iopub.status.idle":"2023-01-20T15:36:28.005315Z","shell.execute_reply.started":"2023-01-20T15:36:11.031779Z","shell.execute_reply":"2023-01-20T15:36:28.004205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import os\nimport pandas as pd, numpy as np, random, shutil\nimport matplotlib.pyplot as plt\nimport yaml\nfrom IPython import display as ipd\nfrom glob import glob\nfrom tqdm import tqdm","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:36:32.16056Z","iopub.execute_input":"2023-01-20T15:36:32.16102Z","iopub.status.idle":"2023-01-20T15:36:32.166727Z","shell.execute_reply.started":"2023-01-20T15:36:32.160986Z","shell.execute_reply":"2023-01-20T15:36:32.165956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('np:', np.__version__)\nprint('pd:', pd.__version__)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:36:34.02783Z","iopub.execute_input":"2023-01-20T15:36:34.0287Z","iopub.status.idle":"2023-01-20T15:36:34.033547Z","shell.execute_reply.started":"2023-01-20T15:36:34.028659Z","shell.execute_reply":"2023-01-20T15:36:34.032766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class CFG:\n    # seed for data-split, layer init, augs\n    seed = 42\n\n    # dicom to png size\n    resize_dim = 1536\n    aspect_ratio = True\n    \n    # size of training image\n    img_size = [1536, 960]","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:36:37.100565Z","iopub.execute_input":"2023-01-20T15:36:37.101007Z","iopub.status.idle":"2023-01-20T15:36:37.105884Z","shell.execute_reply.started":"2023-01-20T15:36:37.100972Z","shell.execute_reply":"2023-01-20T15:36:37.1051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def seeding(SEED):\n    np.random.seed(SEED)\n    random.seed(SEED)\n    os.environ['PYTHONHASHSEED'] = str(SEED)\n    print('seeding done!!!')\nseeding(CFG.seed)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:37:16.201845Z","iopub.execute_input":"2023-01-20T15:37:16.202873Z","iopub.status.idle":"2023-01-20T15:37:16.209929Z","shell.execute_reply.started":"2023-01-20T15:37:16.20283Z","shell.execute_reply":"2023-01-20T15:37:16.208814Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"BASE_PATH = '/kaggle/input/rsna-breast-cancer-detection'\nIMG_DIR = '/kaggle/dataset'","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:37:18.380554Z","iopub.execute_input":"2023-01-20T15:37:18.381489Z","iopub.status.idle":"2023-01-20T15:37:18.386002Z","shell.execute_reply.started":"2023-01-20T15:37:18.381444Z","shell.execute_reply":"2023-01-20T15:37:18.385018Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# train\ndf = pd.read_csv(f'{BASE_PATH}/train.csv')\ndf['dicom_path'] = f'{BASE_PATH}/train_images'\\\n                    + '/' + df.patient_id.astype(str)\\\n                    + '/' + df.image_id.astype(str)\\\n                    + '.dcm'\ndf['image_path'] = df.dicom_path.str.replace('.dcm','.png').str.replace(BASE_PATH, IMG_DIR)\nprint('Train:')\ndisplay(df.head(2))\n\n# test\ntest_df = pd.read_csv(f'{BASE_PATH}/test.csv')\ntest_df['dicom_path'] = f'{BASE_PATH}/test_images'\\\n                    + '/' + test_df.patient_id.astype(str)\\\n                    + '/' + test_df.image_id.astype(str)\\\n                    + '.dcm'\ntest_df['image_path'] = test_df.dicom_path.str.replace('.dcm','.png').str.replace(BASE_PATH, IMG_DIR)\nprint('\\nTest:')\ndisplay(test_df.head(2))","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:37:21.060049Z","iopub.execute_input":"2023-01-20T15:37:21.061016Z","iopub.status.idle":"2023-01-20T15:37:21.510756Z","shell.execute_reply.started":"2023-01-20T15:37:21.060973Z","shell.execute_reply":"2023-01-20T15:37:21.509662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('train_files:',df.shape[0])\nprint('test_files:',test_df.shape[0])","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:37:24.481279Z","iopub.execute_input":"2023-01-20T15:37:24.482678Z","iopub.status.idle":"2023-01-20T15:37:24.488309Z","shell.execute_reply.started":"2023-01-20T15:37:24.482636Z","shell.execute_reply":"2023-01-20T15:37:24.487192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"os.makedirs(os.path.join(IMG_DIR, 'train_images'), exist_ok = True)\nos.makedirs(os.path.join(IMG_DIR, 'test_images'), exist_ok = True)","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:37:26.29913Z","iopub.execute_input":"2023-01-20T15:37:26.29949Z","iopub.status.idle":"2023-01-20T15:37:26.305569Z","shell.execute_reply.started":"2023-01-20T15:37:26.299462Z","shell.execute_reply":"2023-01-20T15:37:26.304547Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"***Preprocessing***\n\nofc this can be changed to anything you want, this is just a place holder","metadata":{}},{"cell_type":"code","source":"import cv2\n\ndef img2roi(img):\n    # Binarize the image\n    bin_img = cv2.threshold(img, 20, 255, cv2.THRESH_BINARY)[1]\n\n    # Make contours around the binarized image, keep only the largest contour\n    contours, _ = cv2.findContours(bin_img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n    contour = max(contours, key=cv2.contourArea)\n\n    # Find ROI from largest contour\n    ys = contour.squeeze()[:, 0]\n    xs = contour.squeeze()[:, 1]\n    roi =  img[np.min(xs):np.max(xs), np.min(ys):np.max(ys)]\n    \n    return roi","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:38:00.960729Z","iopub.execute_input":"2023-01-20T15:38:00.961155Z","iopub.status.idle":"2023-01-20T15:38:01.8741Z","shell.execute_reply.started":"2023-01-20T15:38:00.961124Z","shell.execute_reply":"2023-01-20T15:38:01.873113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# import pydicom\n# from pydicom.pixel_data_handlers.util import apply_voi_lut\n\n\nimport dicomsdl\n\n# logic to read data\ndef read_xray(path, fix_monochrome = True):\n    dicom = dicomsdl.open(path)\n    data = dicom.pixelData(storedvalue=False)  # storedvalue = True for int16 return otherwise float32\n    data = data - np.min(data)\n    data = data / np.max(data)\n    if fix_monochrome and dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = 1.0 - data\n    return data\n\n# logic to preprocess and save\ndef resize_and_save(file_path):\n    img = read_xray(file_path)\n    h, w = img.shape[:2]  # orig hw\n    if CFG.aspect_ratio:\n        r = CFG.resize_dim / max(h, w)  # resize image to img_size\n        interp = cv2.INTER_LINEAR\n        if r != 1:  # always resize down, only resize up if training with augmentation\n            img = cv2.resize(img, (int(w * r), int(h * r)), interpolation=interp)\n    else:\n        img = cv2.resize(img, (CFG.resize_dim, CFG.resize_dim), cv2.INTER_LINEAR)\n    \n    img = (img * 255).astype(np.uint8)\n    img = img2roi(img)\n    img = cv2.resize(img, CFG.img_size[::-1], cv2.INTER_LINEAR)\n    \n    sub_path = file_path.split(\"/\",4)[-1].split('.dcm')[0] + '.png'\n    infos = sub_path.split('/')\n    pid = infos[-2]\n    iid = infos[-1]; iid = iid.replace('.png','')\n    new_path = os.path.join(IMG_DIR, sub_path)\n    os.makedirs(new_path.rsplit('/',1)[0], exist_ok=True)\n    cv2.imwrite(new_path, img)\n    return pid,iid,w,h","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:38:04.369085Z","iopub.execute_input":"2023-01-20T15:38:04.369489Z","iopub.status.idle":"2023-01-20T15:38:04.410646Z","shell.execute_reply.started":"2023-01-20T15:38:04.369457Z","shell.execute_reply":"2023-01-20T15:38:04.409874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!pip install joblib\n!pip install kaggle","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:38:08.240472Z","iopub.execute_input":"2023-01-20T15:38:08.240895Z","iopub.status.idle":"2023-01-20T15:38:12.365387Z","shell.execute_reply.started":"2023-01-20T15:38:08.240863Z","shell.execute_reply":"2023-01-20T15:38:12.364407Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\nfrom joblib import Parallel, delayed\n# go for main df instead of test\nfile_paths = df.dicom_path.tolist()\n# backend loky and -1 is important on TPU\nimgsize = Parallel(n_jobs=-1,backend='loky')(delayed(resize_and_save)(file_path)\\\n                                             for file_path in tqdm(file_paths, leave=True, position=0))","metadata":{"execution":{"iopub.status.busy":"2023-01-20T15:38:15.000394Z","iopub.execute_input":"2023-01-20T15:38:15.000824Z","iopub.status.idle":"2023-01-20T16:00:00.485994Z","shell.execute_reply.started":"2023-01-20T15:38:15.000788Z","shell.execute_reply":"2023-01-20T16:00:00.484786Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Uploading the dataset to Kaggle\n\nNow, it's time to upload the dataset we created inside `/kaggle/dataset`. If it's the first time you will need to run this:","metadata":{}},{"cell_type":"code","source":"!kaggle datasets create -p \"/kaggle/dataset\" --dir-mode zip","metadata":{"execution":{"iopub.status.busy":"2023-01-20T16:13:40.705368Z","iopub.execute_input":"2023-01-20T16:13:40.705818Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If you have already created the dataset before and just want to push a new version, run this:","metadata":{}},{"cell_type":"code","source":"# !kaggle datasets version -p \"/kaggle/dataset\" -m \"Updated via notebook\" --dir-mode zip","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are done! For more information, read the [docs on the Kaggle API](https://github.com/Kaggle/kaggle-api).","metadata":{}}]}