{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## <center><strong>Introduction</strong></center>\n<p>\n    <strong>Image Preprocessing: </strong>The aim of pre-processing is to improve the quality of the image so that we can analyse it in a better way. By preprocessing we can suppress undesired distortions and enhance some features which are necessary for the particular application we are working for. Those features might vary for different applications.\n</p>\n<p>\n    From EDA, some basic properties are found for all the images. These are described as follows: <br>\n    For more insigts, you can visit my EDA notebook: <a href=\"https://www.kaggle.com/code/onkur7/eda-of-image-dataset-mammography-images\">EDA Notebook</a>\n    <ul>\n        <li>Format of the images are dcm which should be converted to either png or jpg as these formats are generally used for computer vision models.</li>\n        <li>It is also found that images has a lot of empty space which should be trimmed for bettwer classification result.</li>\n    </ul>\n</p>\n<p>In this notebook, I implement a class which can be used for preproceesing. The class converts the dicom images to png and then cut empty spaces from the image. Later, it is converted to a fixed size image and saved to the desired path and also can download it to your local system.</p>","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-12-17T08:32:25.423927Z","iopub.execute_input":"2022-12-17T08:32:25.424459Z","iopub.status.idle":"2022-12-17T08:32:26.571774Z","shell.execute_reply.started":"2022-12-17T08:32:25.424347Z","shell.execute_reply":"2022-12-17T08:32:26.570467Z"}}},{"cell_type":"code","source":"!pip install -qU python-gdcm pydicom pylibjpeg","metadata":{"execution":{"iopub.status.busy":"2022-12-21T08:16:17.659225Z","iopub.execute_input":"2022-12-21T08:16:17.659653Z","iopub.status.idle":"2022-12-21T08:16:33.939662Z","shell.execute_reply.started":"2022-12-21T08:16:17.65962Z","shell.execute_reply":"2022-12-21T08:16:33.938426Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import time\nimport random\nimport numpy as np\nimport matplotlib.pyplot as plt \nimport os\nfrom glob import glob\nfrom tqdm.notebook import tqdm\nimport pydicom\nimport cv2\nfrom joblib import Parallel, delayed","metadata":{"execution":{"iopub.status.busy":"2022-12-21T08:16:43.919194Z","iopub.execute_input":"2022-12-21T08:16:43.919645Z","iopub.status.idle":"2022-12-21T08:16:43.92542Z","shell.execute_reply.started":"2022-12-21T08:16:43.919593Z","shell.execute_reply":"2022-12-21T08:16:43.924113Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## <center><strong>Preprocessing: Conversion of Image Extension and Cut Empty Spaces (Modular Approach)</strong></center>","metadata":{}},{"cell_type":"markdown","source":"### <strong>Reading</strong>\n<p>\n    To understand which operation are performed for which reason you can go through the following articles:\n    <ul>\n        <li><a href=\"https://medium.com/analytics-vidhya/binarization-of-image-using-numpy-65df2b82e189\">Binarization of Image</a></li>\n        <li><a href=\"https://www.geeksforgeeks.org/find-and-draw-contours-using-opencv-python/\">Contour using OpenCV</a></li>\n        <li><a href=\"https://www.kaggle.com/code/fabiendaniel/dicom-cropped-resized-png-jpg\">Reference: Visit this notebook</a></li>\n    </ul>\n<p>\n<p>\n    <strong>Parameters of constructor: </strong>\n    <ul>\n        <li><strong>source_path: </strong>Path of the parent folder of all the images.</li>\n        <li><strong>dest_path: </strong>Path of the parent folder where processed image will be stored.</li>\n        <li><strong>size: </strong>Desired size of the output image.</li>\n    </ul>\n    <strong>Parameters of call: </strong>\n    <ul>\n        <li><strong>png: </strong>If true, image is converted to png format and if false then image is converted to jpg format.</li>\n        <li><strong>do_zip: </strong>Convert the output folder to a zip file.</li>\n    </ul>\n    <strong>Parameters of visualize_samples: </strong>\n    <ul>\n        <li><strong>k: </strong>Number of images to be visualized. k should be even.</li>\n    </ul>\n</p>\n<strong>Note: </strong>Don't use '_' in the dest_path. For some reason it isn't working properly.","metadata":{}},{"cell_type":"code","source":"class RSNAMammographyPreprocessor:\n    # Constructor\n    def __init__(self, source_path: str, dest_path: str, size: int):\n        self.source_path = source_path\n        self.dest_path = dest_path\n        # If destination folder isn't present then create it\n        os.makedirs(dest_path, exist_ok=True)\n        self.size = size\n    def get_image_path(self):\n        # Get all the dcm image paths\n        image_paths = glob(self.source_path + '/*/*.dcm')\n        return image_paths\n    def get_save_path(self, image_path: str, png: bool):\n        # Extract filename\n        path = image_path.split('/')\n        filename = path[-1]\n        # Add proper extension\n        if png: filename = filename.replace('dcm', 'png')\n        else: filename = filename.replace('decm', 'jpeg')\n        # Return the path of the output file name and path\n        return (os.path.join(self.dest_path, path[-2]), os.path.join(self.dest_path, path[-2], filename))\n    def image_format_converter(self, image_path: str):\n        # Read dicom image\n        dicom = pydicom.dcmread(image_path)\n        # Extract the pixel values of the image\n        img = dicom.pixel_array\n        # Process image to convert to png/jpg\n        img = (img - img.min()) / (img.max() - img.min())\n        img *= 255\n        img = np.uint8(img)\n        # If image is inverted then re-invert\n        if dicom.PhotometricInterpretation == \"MONOCHROME1\":\n            img = 1 - img\n        # Reaturn the image array to convert into jpg/png format\n        return img\n    def binarize_image(self, img: np.ndarray, thresh_val: int = 50):\n        # Depending on threshold convert the image to 1-0 image. Pixel either has value 255 or 0\n        _, bin_image= cv2.threshold(src=img, thresh=thresh_val, maxval=255, type=cv2.THRESH_BINARY)\n        return bin_image\n    def extract_contour(self, bin_img: np.ndarray):\n        # Extract all conctours\n        contours, hierarchy = cv2.findContours(bin_img, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_NONE)\n        # Find the contour with maximum area\n        return max(contours, key=cv2.contourArea)\n    def get_final_image(self, img: np.ndarray, contour: np.ndarray):\n        # Create a mask\n        mask = np.zeros(img.shape, np.uint8)\n        # Extract the image which intersects with largest contour\n        cv2.drawContours(mask, [contour], -1, 255, cv2.FILLED)\n        output = cv2.bitwise_and(img, mask)\n        return output\n    def crop_image(self, img: np.ndarray):\n        # Binarize image\n        bin_img = self.binarize_image(img)\n        # Extract desired contour\n        contour = self.extract_contour(bin_img)\n        # Getv the desired image\n        output = self.get_final_image(img, contour)\n        # Extract the desired part from the image\n        y1, y2 = np.min(contour[:, :, 1]), np.max(contour[:, :, 1])\n        x1, x2 = np.min(contour[:, :, 0]), np.max(contour[:, :, 0])\n        y1, y2 = np.min(contour[:, :, 1]), np.max(contour[:, :, 1])\n        x1, x2 = np.min(contour[:, :, 0]), np.max(contour[:, :, 0])\n        x1 = int(0.99 * x1)\n        x2 = int(1.01 * x2)\n        y1 = int(0.99 * y1)\n        y2 = int(1.01 * y2)\n        return output[y1:y2, x1:x2]\n    def process_single_image(self, path: str, png: bool):\n        dest_path, save_path = self.get_save_path(path, png)\n        img = self.image_format_converter(path)\n        os.makedirs(dest_path, exist_ok=True)\n        img = self.crop_image(img)\n        img = cv2.resize(img, (self.size, self.size))\n        cv2.imwrite(save_path, img)\n    def convert_all_images(self, png: bool):\n        # Convert the image from dicom to png\n        image_paths = self.get_image_path()\n        Parallel(n_jobs=4)(delayed(self.process_single_image)(path, png) for path in tqdm(image_paths, total=len(image_paths)))\n    def visualize_samples(self, k = 4):\n        # Visualize some samples\n        image_paths = glob(self.dest_path + '/*/*.png')\n        image_paths.extend(glob(self.dest_path + '/*/*.jpg'))\n        if len(image_paths) == 0: return\n        if k > len(image_paths): k = len(image_paths)\n        if k%2 == 1: k -= 1\n        samples = random.choices(image_paths, k=k)\n        columns = 2\n        rows = k // 2\n        idx = 1\n        fig, ax = plt.subplots(figsize=(rows * 12, columns * 8))\n        for img_path in tqdm(samples, total=len(samples)):\n            img = cv2.imread(img_path)\n            fig.add_subplot(rows, columns, idx)\n            idx += 1\n            plt.imshow(img, cmap='gray')\n            plt.axis('off')\n            plt.title(img_path.split('/')[-1][:-4])\n        plt.show()\n    def zipping(self):\n        # Zip preprocessed image\n        zip_filename = self.dest_path.split('/')[-1] +'.zip'\n        command = 'zip -q -r ' + zip_filename +' ' + self.dest_path\n        os.system(command)\n    def __call__(self, png: bool = True, do_zip: bool = True):\n        self.convert_all_images(png)\n        if do_zip: self.zipping()","metadata":{"execution":{"iopub.status.busy":"2022-12-21T08:30:36.795201Z","iopub.execute_input":"2022-12-21T08:30:36.79598Z","iopub.status.idle":"2022-12-21T08:30:36.826021Z","shell.execute_reply.started":"2022-12-21T08:30:36.795932Z","shell.execute_reply":"2022-12-21T08:30:36.824659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <strong>Test Parallel Processing</strong>","metadata":{}},{"cell_type":"code","source":"start = time.time()\nprocessor = RSNAMammographyPreprocessor('/kaggle/input/rsna-breast-cancer-detection/test_images',\n                                      os.path.join(os.getcwd(), 'prcocessedTestImageRSNA'), 512)\nprocessor()\nend = time.time()\nprint(\"Exceution Time:\", end-start)","metadata":{"execution":{"iopub.status.busy":"2022-12-21T08:30:58.211982Z","iopub.execute_input":"2022-12-21T08:30:58.212398Z","iopub.status.idle":"2022-12-21T08:31:03.617888Z","shell.execute_reply.started":"2022-12-21T08:30:58.212363Z","shell.execute_reply":"2022-12-21T08:31:03.616511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <strong>Process Train Images</strong>","metadata":{}},{"cell_type":"code","source":"start = time.time()\nprocessor = RSNAMammographyPreprocessor('/kaggle/input/rsna-breast-cancer-detection/train_images',\n                                      os.path.join(os.getcwd(), 'prcocessedTrainImageRSNA'), 512)\nprocessor()\nend = time.time()\nprint(\"Exceution Time:\", end-start)","metadata":{"execution":{"iopub.status.busy":"2022-12-21T08:32:52.609776Z","iopub.execute_input":"2022-12-21T08:32:52.611171Z","iopub.status.idle":"2022-12-21T08:33:18.767062Z","shell.execute_reply.started":"2022-12-21T08:32:52.611105Z","shell.execute_reply":"2022-12-21T08:33:18.765334Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### <center><strong>If you like my notebook: Please Upvote</strong></center>","metadata":{}}]}