{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"hey!\n\nIn a post where I shared my dataset: [📸 DICOM files converted to PNGs [314.72 GB -> 921 MB]](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369282#2049882) there has been a question on how I created it 🙂\n\nLet me give an answer in this notebook.\n\nThe dataset in question is this: [RSNA Mammography -- images as PNGs (256px & 512px)](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs) (most likely will add additional resolutions soon).\n\nLet's first take a look at a couple of the images.\n\n**Update:** In my experiments, processing with `cv2` can be 10% - 20% faster than processing with `PIL`. Both datasets should be equivalent. But to take no chances that there could be differences, I am uploading a version of the dataset processed by `cv2` (for now, for the 256px and 512px dimensions). I am including the code I used to convert images using `cv2` at the bottom of this notebook.\n\n**Dataset legend:**\n\n* filenames with no attributes: processed using PIL\n* filenames with `cv2`: processed with opencv\n* filenames with `vl`: with VOI LUT transformation applied (thank you [raddar](https://www.kaggle.com/code/radek1/how-to-process-dicom-images-to-pngs/comments#2056978)! 🙏)\n* filenames with `dicomsdl`: processed using `dicomsdl`, 2x faster vs `pydicom` (thank you [Remek Kinas](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033)! 🙏)\n* filenames with `asp`: longer side resized to 1024, aspect ratio kept, another suggestion from [Remek Kinas](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/371033)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Example images","metadata":{}},{"cell_type":"code","source":"# 256x256\n\nfrom pathlib import Path\nfrom PIL import Image\nfrom matplotlib import pyplot as plt\n\nfig, axs = plt.subplots(2, 2, figsize=(10,10))\nfor ax, path in zip(axs.flat, Path('../input/rsna-mammography-images-as-pngs/images_as_pngs/train_images_processed/15843').iterdir()):\n    img = Image.open(path)\n    ax.imshow(img, cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2022-12-11T13:49:10.817527Z","iopub.execute_input":"2022-12-11T13:49:10.817901Z","iopub.status.idle":"2022-12-11T13:49:11.424827Z","shell.execute_reply.started":"2022-12-11T13:49:10.81782Z","shell.execute_reply":"2022-12-11T13:49:11.423328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# 512x512\n\nfrom pathlib import Path\nfrom PIL import Image\nfrom matplotlib import pyplot as plt\n\nfig, axs = plt.subplots(2, 2, figsize=(10,10))\nfor ax, path in zip(axs.flat, Path('../input/rsna-mammography-images-as-pngs/images_as_pngs_512/train_images_processed_512/15843').iterdir()):\n    img = Image.open(path)\n    ax.imshow(img, cmap=\"bone\")","metadata":{"execution":{"iopub.status.busy":"2022-12-11T13:49:11.42666Z","iopub.execute_input":"2022-12-11T13:49:11.426956Z","iopub.status.idle":"2022-12-11T13:49:12.176395Z","shell.execute_reply.started":"2022-12-11T13:49:11.426932Z","shell.execute_reply":"2022-12-11T13:49:12.175015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# How I carried out the processing","metadata":{}},{"cell_type":"markdown","source":"I used a GCP N2 VM with 64 cores with 64 cores. Processing of all images to 256x256 took around 17 minutes. Here is the code I ran:","metadata":{}},{"cell_type":"code","source":"import pydicom\nfrom pathlib import Path\nimport numpy as np\nfrom PIL import Image\n\nRESIZE_TO = (512, 512)\n!rm -rf train_images_processed_{RESIZE_TO[0]}\n!mkdir train_images_processed_{RESIZE_TO[0]}\n\n# https://www.kaggle.com/code/tanlikesmath/brain-tumor-radiogenomic-classification-eda/notebook\ndef dicom_file_to_ary(path):\n    dicom = pydicom.read_file(path)\n    data = dicom.pixel_array\n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = np.amax(data) - data\n    data = data - np.min(data)\n    data = data / np.max(data)\n    data = (data * 255).astype(np.uint8)\n    return data\n\ndirectories = list(Path('train_images').iterdir())\n\ndef process_directory(directory_path):\n    parent_directory = str(directory_path).split('/')[-1]\n    !mkdir -p train_images_processed_{RESIZE_TO[0]}/{parent_directory}\n    for image_path in directory_path.iterdir():\n        processed_ary = dicom_file_to_ary(image_path)\n        im = Image.fromarray(processed_ary).resize(RESIZE_TO)\n        im.save(f'train_images_processed_{RESIZE_TO[0]}/{parent_directory}/{image_path.stem}.png')\n        \nimport multiprocessing as mp\n\nwith mp.Pool(64) as p:\n    p.map(process_directory, directories)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thank you for reading! If you found this useful, please upvote the dataset: [RSNA Mammography -- images as PNGs (256px & 512px)](https://www.kaggle.com/datasets/radek1/rsna-mammography-images-as-pngs)\n\nBest of luck in the competition! 🙂","metadata":{}},{"cell_type":"markdown","source":"**Code for processing with `opencv`**","metadata":{}},{"cell_type":"code","source":"import pydicom\nimport numpy as np\nimport cv2\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.notebook import tqdm\nfrom pathlib import Path\n\nRESIZE_TO = (512, 512)\n\n!rm -rf train_images_processed_cv2_{RESIZE_TO[0]}\n!mkdir train_images_processed_cv2_{RESIZE_TO[0]}\n\n# https://www.kaggle.com/code/tanlikesmath/brain-tumor-radiogenomic-classification-eda/notebook\ndef dicom_file_to_ary(path):\n    dicom = pydicom.dcmread(path)\n    data = dicom.pixel_array\n       \n    data = (data - data.min()) / (data.max() - data.min())\n    \n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = 1 - data\n        \n    data = cv2.resize(data, RESIZE_TO)\n    data = (data * 255).astype(np.uint8)\n    return data\n\ndirectories = list(Path('train_images').iterdir())\n\ndef process_directory(directory_path):\n    parent_directory = str(directory_path).split('/')[-1]\n    !mkdir -p train_images_processed_cv2_{RESIZE_TO[0]}/{parent_directory}\n    for image_path in directory_path.iterdir():\n        processed_ary = dicom_file_to_ary(image_path)\n        \n        cv2.imwrite(\n            f'train_images_processed_cv2_{RESIZE_TO[0]}/{parent_directory}/{image_path.stem}.png',\n            processed_ary\n        )\n        \nimport multiprocessing as mp\n\nwith mp.Pool(64) as p:\n    p.map(process_directory, directories)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Code for processing using opencv with VOI LUT transformation**","metadata":{}},{"cell_type":"code","source":"%%time\n\nimport pydicom\nimport numpy as np\nimport cv2\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.notebook import tqdm\nfrom pathlib import Path\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\nRESIZE_TO = (512, 512)\n\n!rm -rf train_images_processed_cv2_vl_{RESIZE_TO[0]}\n!mkdir train_images_processed_cv2_vl_{RESIZE_TO[0]}\n\n# https://www.kaggle.com/code/tanlikesmath/brain-tumor-radiogenomic-classification-eda/notebook\ndef dicom_file_to_ary(path):\n    dicom = pydicom.dcmread(path)\n    data = dicom.pixel_array\n    data = apply_voi_lut(dicom.pixel_array, dicom)\n       \n    data = (data - data.min()) / (data.max() - data.min())\n    \n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = 1 - data\n        \n    data = cv2.resize(data, RESIZE_TO)\n    data = (data * 255).astype(np.uint8)\n    return data\n\ndirectories = list(Path('train_images').iterdir())\n\ndef process_directory(directory_path):\n    parent_directory = str(directory_path).split('/')[-1]\n    !mkdir -p train_images_processed_cv2_vl_{RESIZE_TO[0]}/{parent_directory}\n    for image_path in directory_path.iterdir():\n        processed_ary = dicom_file_to_ary(image_path)\n        \n        cv2.imwrite(\n            f'train_images_processed_cv2_vl_{RESIZE_TO[0]}/{parent_directory}/{image_path.stem}.png',\n            processed_ary\n        )\n        \nimport multiprocessing as mp\n\nwith mp.Pool(64) as p:\n    p.map(process_directory, directories)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Code for processing using opencv dicomsdl (very fast 🔥🔥🔥 see [FAST dicom processing (1.6-2x faster) 💪💪](https://www.kaggle.com/code/remekkinas/fast-dicom-processing-1-6-2x-faster)**","metadata":{}},{"cell_type":"code","source":"%%time\n\nimport pydicom\nimport numpy as np\nimport cv2\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.notebook import tqdm\nfrom pathlib import Path\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\nimport dicomsdl\n\nRESIZE_TO = (512, 512)\n\n!rm -rf train_images_processed_cv2_dicomsdl_{RESIZE_TO[0]}\n!mkdir train_images_processed_cv2_dicomsdl_{RESIZE_TO[0]}\n\n# https://www.kaggle.com/code/tanlikesmath/brain-tumor-radiogenomic-classification-eda/notebook\ndef dicom_file_to_ary(path):\n    dcm_file = dicomsdl.open(str(path))\n    data = dcm_file.pixelData()\n\n    data = (data - data.min()) / (data.max() - data.min())\n\n    if dcm_file.getPixelDataInfo()['PhotometricInterpretation'] == \"MONOCHROME1\":\n        data = 1 - data\n\n    data = cv2.resize(data, RESIZE_TO)\n    data = (data * 255).astype(np.uint8)\n    return data\n\ndirectories = list(Path('train_images').iterdir())\n\ndef process_directory(directory_path):\n    parent_directory = str(directory_path).split('/')[-1]\n    !mkdir -p train_images_processed_cv2_dicomsdl_{RESIZE_TO[0]}/{parent_directory}\n    for image_path in directory_path.iterdir():\n        processed_ary = dicom_file_to_ary(image_path)\n        \n        cv2.imwrite(\n            f'train_images_processed_cv2_dicomsdl_{RESIZE_TO[0]}/{parent_directory}/{image_path.stem}.png',\n            processed_ary\n        )\n        \nimport multiprocessing as mp\n\nwith mp.Pool(64) as p:\n    p.map(process_directory, directories)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Code for processing with maintaining the aspect ratio, longer side resized to 1024**","metadata":{}},{"cell_type":"code","source":"%%time\n\nimport pydicom\nimport numpy as np\nimport cv2\nimport os\nfrom joblib import Parallel, delayed\nfrom tqdm.notebook import tqdm\nfrom pathlib import Path\nfrom pydicom.pixel_data_handlers.util import apply_voi_lut\n\nRESIZE_TO = 1024\n\n!rm -rf train_images_processed_cv2_vl_asp_{RESIZE_TO}\n!mkdir train_images_processed_cv2_vl_asp_{RESIZE_TO}\n\n\ndef image_resize(image, width = None, height = None, inter = cv2.INTER_LINEAR):\n\n    dim = None\n    (h, w) = image.shape[:2]\n\n    if width is None and height is None:\n        return image\n\n    if width is None:\n        r = height / float(h)\n        dim = (int(w * r), height)\n    else:\n        r = width / float(w)\n        dim = (width, int(h * r))\n    resized = cv2.resize(image, dim, interpolation = inter)\n\n    return resized\n\n\n# https://www.kaggle.com/code/tanlikesmath/brain-tumor-radiogenomic-classification-eda/notebook\ndef dicom_file_to_ary(path):\n    dicom = pydicom.dcmread(path)\n    data = dicom.pixel_array\n    data = apply_voi_lut(dicom.pixel_array, dicom)\n       \n    data = (data - data.min()) / (data.max() - data.min())\n    \n    if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n        data = 1 - data\n    \n    h, w = data.shape\n    if w > h:\n         data = image_resize(data, width = RESIZE_TO)\n    else:\n         data = image_resize(data, height = RESIZE_TO)\n    \n    data = (data * 255).astype(np.uint8)\n    return data\n\ndirectories = list(Path('train_images').iterdir())\n\ndef process_directory(directory_path):\n    parent_directory = str(directory_path).split('/')[-1]\n    !mkdir -p train_images_processed_cv2_vl_asp_{RESIZE_TO}/{parent_directory}\n    for image_path in directory_path.iterdir():\n        processed_ary = dicom_file_to_ary(image_path)\n        \n        cv2.imwrite(\n            f'train_images_processed_cv2_vl_asp_{RESIZE_TO}/{parent_directory}/{image_path.stem}.png',\n            processed_ary\n        )\n        \nimport multiprocessing as mp\n\nwith mp.Pool(64) as p:\n    p.map(process_directory, directories)","metadata":{},"execution_count":null,"outputs":[]}]}