{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# OVERVIEW\n\nThis notebook is to explore prepocessing options. \n\nMany of the the techniques changed through-out the course of the notebook. I have chosen to leave them in to display my progress and hopefully give inspiration to others.\n\nThe final goal is to create a dataset of cropped images that can be interchanged with the provided images. \n\nPNG|Dataset - [RSNA-ATD | preprocessed](https://www.kaggle.com/datasets/davidjohnmillard/rsna-atd-preprocessed)\n\nTFWriter - [RSNA-ATD | TFRecord 📓](https://www.kaggle.com/code/davidjohnmillard/rsna-atd-tfrecord/notebook) \n\nTFRecord|Dataset - (WORK IN PROGRESS)\n\n# UPDATES\n\nWorking with the full dataset (WORK IN PROGRESS)","metadata":{}},{"cell_type":"markdown","source":"# Imports/Setup\n\nImporting necessary libraries and setting up base dataframes.","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd\nimport pydicom\nimport matplotlib.pyplot as plt\nimport cv2\nimport seaborn as sns\nimport tensorflow as tf\nimport os","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:27:48.489428Z","iopub.execute_input":"2023-08-24T17:27:48.490409Z","iopub.status.idle":"2023-08-24T17:28:00.578825Z","shell.execute_reply.started":"2023-08-24T17:27:48.490353Z","shell.execute_reply":"2023-08-24T17:28:00.577791Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_series_meta.csv')\nlabels = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train.csv', index_col='patient_id')\ntrain_meta = pd.read_csv('/kaggle/input/rsna-2023-abdominal-trauma-detection/train_series_meta.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:00.58129Z","iopub.execute_input":"2023-08-24T17:28:00.582706Z","iopub.status.idle":"2023-08-24T17:28:00.645725Z","shell.execute_reply.started":"2023-08-24T17:28:00.582663Z","shell.execute_reply":"2023-08-24T17:28:00.644306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Full Training Dataset\n\nMaking a full training dataset with patient_id, series_id, instance_number, img_path, labels, aortic_hu, and incomplete_organ.","metadata":{}},{"cell_type":"code","source":"create_fulld = False\nfull_train_paths = []\n\ndef comp_train_imgs(row):\n    patient_id = int(row['patient_id'])\n    series_id = int(row['series_id'])\n    \n    for i in os.listdir(f'/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/{patient_id}/{series_id}'):\n        instance_number = i.split('.')[0]\n        img_path = f'/kaggle/input/rsna-2023-abdominal-trauma-detection/train_images/{patient_id}/{series_id}/{i}'\n        \n        x_data = [patient_id, series_id, instance_number, img_path]\n        y_data = labels.loc[patient_id, :].values.tolist()\n        \n        tmp = train_meta[(train_meta['patient_id'] == patient_id)]\n        tmps = tmp[tmp['series_id'] == series_id]\n        meta_data = [tmps['aortic_hu'].values[0], tmps['incomplete_organ'].values[0]]\n        \n        data = x_data + y_data + meta_data\n        \n        full_train_paths.append(data)\n        \nif create_fulld:\n    _ = train.apply(comp_train_imgs, axis=1)\n    columns = ['patient_id', 'series_id', 'instance_number', 'img_path'] + labels.columns.tolist() + ['aortic_hu', 'incomplete_organ']\n    train = pd.DataFrame(full_train_paths, columns=columns)\n    train.to_parquet('full_train.parquet')\nelse: \n    train = pd.read_parquet('/kaggle/input/rsna-atd-full-train-meta/full_train.parquet')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:00.647413Z","iopub.execute_input":"2023-08-24T17:28:00.647812Z","iopub.status.idle":"2023-08-24T17:28:02.633418Z","shell.execute_reply.started":"2023-08-24T17:28:00.647768Z","shell.execute_reply":"2023-08-24T17:28:02.632133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:02.637044Z","iopub.execute_input":"2023-08-24T17:28:02.637666Z","iopub.status.idle":"2023-08-24T17:28:02.674718Z","shell.execute_reply.started":"2023-08-24T17:28:02.637628Z","shell.execute_reply":"2023-08-24T17:28:02.673482Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA\n\nTaking a basic look at the data and labels to see what work needs to be done.\n\n### Labels\n\nThe labels are multi-class and are not independent (each instance can have multiple 1 labels).","metadata":{}},{"cell_type":"code","source":"labels","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:02.676427Z","iopub.execute_input":"2023-08-24T17:28:02.676892Z","iopub.status.idle":"2023-08-24T17:28:02.697646Z","shell.execute_reply.started":"2023-08-24T17:28:02.676849Z","shell.execute_reply":"2023-08-24T17:28:02.696169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Just from a basic view of each mean (healthy vs injury) there is clear imbalance.*","metadata":{}},{"cell_type":"code","source":"labels.loc[: , labels.columns!='patient_id'].describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:02.699482Z","iopub.execute_input":"2023-08-24T17:28:02.700295Z","iopub.status.idle":"2023-08-24T17:28:02.763476Z","shell.execute_reply.started":"2023-08-24T17:28:02.700239Z","shell.execute_reply":"2023-08-24T17:28:02.762147Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Example Instance\n\nTaking a look at the the imaging of one instance (of one series, of one patient).\n\n*Working with small sample of each series of every patient for time reasons.*","metadata":{}},{"cell_type":"code","source":"sub_train = train.groupby(by=['patient_id', 'series_id']).sample(n=1, random_state=42)\nfirst_ind = sub_train.index[0]","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:02.765326Z","iopub.execute_input":"2023-08-24T17:28:02.765825Z","iopub.status.idle":"2023-08-24T17:28:03.716446Z","shell.execute_reply.started":"2023-08-24T17:28:02.765792Z","shell.execute_reply":"2023-08-24T17:28:03.715189Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_train","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:03.717736Z","iopub.execute_input":"2023-08-24T17:28:03.718074Z","iopub.status.idle":"2023-08-24T17:28:03.746964Z","shell.execute_reply.started":"2023-08-24T17:28:03.718046Z","shell.execute_reply":"2023-08-24T17:28:03.745638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ex_inst = pydicom.dcmread(sub_train.loc[first_ind, 'img_path'])","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:03.748479Z","iopub.execute_input":"2023-08-24T17:28:03.748927Z","iopub.status.idle":"2023-08-24T17:28:03.772534Z","shell.execute_reply.started":"2023-08-24T17:28:03.748884Z","shell.execute_reply":"2023-08-24T17:28:03.771384Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ex_inst","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:03.778054Z","iopub.execute_input":"2023-08-24T17:28:03.778466Z","iopub.status.idle":"2023-08-24T17:28:03.78766Z","shell.execute_reply.started":"2023-08-24T17:28:03.778432Z","shell.execute_reply":"2023-08-24T17:28:03.78654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Looks like there is a decent amount of dead space.*","metadata":{}},{"cell_type":"code","source":"plt.imshow(ex_inst.pixel_array, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:03.788788Z","iopub.execute_input":"2023-08-24T17:28:03.789168Z","iopub.status.idle":"2023-08-24T17:28:04.208723Z","shell.execute_reply.started":"2023-08-24T17:28:03.789083Z","shell.execute_reply":"2023-08-24T17:28:04.207319Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looking at the dimensions of each image should tell us whether a lot of resizing is needed.","metadata":{}},{"cell_type":"code","source":"def row_size(row):\n    return pydicom.dcmread(row['img_path']).Rows\n\ndef col_size(row):\n    return pydicom.dcmread(row['img_path']).Columns\n\ndef interp(row):\n    return pydicom.dcmread(row['img_path']).PhotometricInterpretation\n\nsub_train['row'] = sub_train.apply(row_size, axis=1)\nsub_train['col'] = sub_train.apply(col_size, axis=1)\nsub_train['interp'] = sub_train.apply(interp, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:28:04.210472Z","iopub.execute_input":"2023-08-24T17:28:04.211502Z","iopub.status.idle":"2023-08-24T17:29:33.331563Z","shell.execute_reply.started":"2023-08-24T17:28:04.211458Z","shell.execute_reply":"2023-08-24T17:29:33.330501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The dimensions don't vary too much but there may be a lot of dead space. So the results from cropped dimensions will be the deciding factor.","metadata":{}},{"cell_type":"code","source":"sub_train['row'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.332916Z","iopub.execute_input":"2023-08-24T17:29:33.333309Z","iopub.status.idle":"2023-08-24T17:29:33.343026Z","shell.execute_reply.started":"2023-08-24T17:29:33.333278Z","shell.execute_reply":"2023-08-24T17:29:33.341852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_train['col'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.344573Z","iopub.execute_input":"2023-08-24T17:29:33.345271Z","iopub.status.idle":"2023-08-24T17:29:33.362306Z","shell.execute_reply.started":"2023-08-24T17:29:33.345235Z","shell.execute_reply":"2023-08-24T17:29:33.361248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Everything is MONOCHROME2 so no need to worry about inverting.*","metadata":{}},{"cell_type":"code","source":"sub_train['interp'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.36364Z","iopub.execute_input":"2023-08-24T17:29:33.363983Z","iopub.status.idle":"2023-08-24T17:29:33.380688Z","shell.execute_reply.started":"2023-08-24T17:29:33.363954Z","shell.execute_reply":"2023-08-24T17:29:33.379465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Basic Pre-Processing\n\nThe basic will include normalization, cropping using thresholding, and standardization given pixel representation == 1.\n\n**Normalization** : Truncated normalization from someones notebook in the previous competition. (DM me if it was you for credit)\n\n**Cropping with Thresholding** : Cropping the images using thresholding from someones notebook in the previous competition. (DM me if it was you for credit)\n\n**Standarization** : Standardization for Pixel Representation == 1 from the provided competition code.","metadata":{}},{"cell_type":"code","source":"n_blur = 21\n\ndef crop_coords(img):\n    blur = cv2.GaussianBlur(img, (n_blur, n_blur), 0)\n    _, mask = cv2.threshold(blur,0,255,cv2.THRESH_BINARY+cv2.THRESH_OTSU)\n    \n    cnts, _ = cv2.findContours(mask.astype(np.uint8), cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)\n    cnt = max(cnts, key = cv2.contourArea)\n    \n    x, y, w, h = cv2.boundingRect(cnt)\n    \n    return (x, y, w, h)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.382294Z","iopub.execute_input":"2023-08-24T17:29:33.382635Z","iopub.status.idle":"2023-08-24T17:29:33.393629Z","shell.execute_reply.started":"2023-08-24T17:29:33.382604Z","shell.execute_reply":"2023-08-24T17:29:33.392143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def truncation_normalization(img):\n    pmin = np.percentile(img[img!=0], 5)\n    pmax = np.percentile(img[img!=0], 99)\n    \n    truncated = np.clip(img, pmin, pmax)  \n    \n    normalized = (truncated - pmin)/(pmax - pmin)\n    normalized[img==0] = 0\n    \n    return normalized","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.395028Z","iopub.execute_input":"2023-08-24T17:29:33.395527Z","iopub.status.idle":"2023-08-24T17:29:33.405519Z","shell.execute_reply.started":"2023-08-24T17:29:33.395483Z","shell.execute_reply":"2023-08-24T17:29:33.404487Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def standardize_pixel_array(dcm):\n    pixel_array = dcm.pixel_array\n    \n    if dcm.PixelRepresentation == 1:\n        bit_shift = dcm.BitsAllocated - dcm.BitsStored\n        dtype = pixel_array.dtype \n        \n        new_array = (pixel_array << bit_shift).astype(dtype) >>  bit_shift\n        pixel_array = pydicom.pixel_data_handlers.util.apply_modality_lut(new_array, dcm)\n        \n    return pixel_array","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.407385Z","iopub.execute_input":"2023-08-24T17:29:33.407733Z","iopub.status.idle":"2023-08-24T17:29:33.417572Z","shell.execute_reply.started":"2023-08-24T17:29:33.407696Z","shell.execute_reply":"2023-08-24T17:29:33.416348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Puting it all together.","metadata":{}},{"cell_type":"code","source":"def preprocess(inst):\n    dcm = pydicom.dcmread(inst['img_path'])\n    \n    img = standardize_pixel_array(dcm)\n    img = truncation_normalization(img)\n    img = (img * 255).astype('uint8')\n    (x, y, w, h) = crop_coords(img)\n    img = img[y:y+h, x:x+w]\n    \n    return img","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.418886Z","iopub.execute_input":"2023-08-24T17:29:33.419284Z","iopub.status.idle":"2023-08-24T17:29:33.432039Z","shell.execute_reply.started":"2023-08-24T17:29:33.419253Z","shell.execute_reply":"2023-08-24T17:29:33.430575Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(preprocess(sub_train.loc[first_ind , :]), cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.434828Z","iopub.execute_input":"2023-08-24T17:29:33.435352Z","iopub.status.idle":"2023-08-24T17:29:33.868459Z","shell.execute_reply.started":"2023-08-24T17:29:33.435305Z","shell.execute_reply":"2023-08-24T17:29:33.8672Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.imshow(pydicom.dcmread(sub_train.loc[first_ind , :]['img_path']).pixel_array, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:33.870189Z","iopub.execute_input":"2023-08-24T17:29:33.870616Z","iopub.status.idle":"2023-08-24T17:29:34.236526Z","shell.execute_reply.started":"2023-08-24T17:29:33.870575Z","shell.execute_reply":"2023-08-24T17:29:34.235341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Crop Dimensions\n\nPre-Processing each image then saving their bounds. ","metadata":{}},{"cell_type":"code","source":"def cropped_sizes(row):\n    dcm = pydicom.dcmread(row['img_path'])\n    \n    img = standardize_pixel_array(dcm)\n    img = truncation_normalization(img)\n    img = (img * 255).astype('uint8')\n    (x, y, w, h) = crop_coords(img)\n        \n    return (x, y, w, h)\n\nsub_train['crop_raw_size'] = sub_train.apply(cropped_sizes, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:29:34.23802Z","iopub.execute_input":"2023-08-24T17:29:34.238499Z","iopub.status.idle":"2023-08-24T17:31:49.587143Z","shell.execute_reply.started":"2023-08-24T17:29:34.238454Z","shell.execute_reply":"2023-08-24T17:31:49.585877Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inst = sub_train.loc[first_ind, 'crop_raw_size']\nimg_cropped = ex_inst.pixel_array[inst[1]:inst[1]+inst[3], inst[0]:inst[0]+inst[2]]","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:49.588662Z","iopub.execute_input":"2023-08-24T17:31:49.589122Z","iopub.status.idle":"2023-08-24T17:31:49.597761Z","shell.execute_reply.started":"2023-08-24T17:31:49.589059Z","shell.execute_reply":"2023-08-24T17:31:49.596411Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Cropped example instance.*","metadata":{}},{"cell_type":"code","source":"plt.imshow(img_cropped, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:49.599492Z","iopub.execute_input":"2023-08-24T17:31:49.600076Z","iopub.status.idle":"2023-08-24T17:31:49.979344Z","shell.execute_reply.started":"2023-08-24T17:31:49.600031Z","shell.execute_reply":"2023-08-24T17:31:49.978149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Getting the width, height, and their ratio.","metadata":{}},{"cell_type":"code","source":"def get_w_data(row):\n    return row['crop_raw_size'][2]\n\ndef get_h_data(row):\n    return row['crop_raw_size'][3]\n\ndef get_ratio_data(row):\n    return row['crop_raw_size'][2] / row['crop_raw_size'][3]\n\nsub_train['crop_ratio'] = sub_train.apply(get_ratio_data, axis=1)\nsub_train['crop_width'] = sub_train.apply(get_w_data, axis=1)\nsub_train['crop_height'] = sub_train.apply(get_h_data, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:49.981071Z","iopub.execute_input":"2023-08-24T17:31:49.981803Z","iopub.status.idle":"2023-08-24T17:31:50.20794Z","shell.execute_reply.started":"2023-08-24T17:31:49.981759Z","shell.execute_reply":"2023-08-24T17:31:50.206627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_train['crop_ratio'].describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:50.209247Z","iopub.execute_input":"2023-08-24T17:31:50.209592Z","iopub.status.idle":"2023-08-24T17:31:50.224887Z","shell.execute_reply.started":"2023-08-24T17:31:50.209564Z","shell.execute_reply":"2023-08-24T17:31:50.223403Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*There is a peak at ~1.0 and a distribution to the right.*","metadata":{}},{"cell_type":"code","source":"sns.histplot(sub_train['crop_ratio'])","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:50.226609Z","iopub.execute_input":"2023-08-24T17:31:50.22724Z","iopub.status.idle":"2023-08-24T17:31:50.60828Z","shell.execute_reply.started":"2023-08-24T17:31:50.227202Z","shell.execute_reply":"2023-08-24T17:31:50.607062Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_train['crop_width'].describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:50.61532Z","iopub.execute_input":"2023-08-24T17:31:50.615743Z","iopub.status.idle":"2023-08-24T17:31:50.631997Z","shell.execute_reply.started":"2023-08-24T17:31:50.615709Z","shell.execute_reply":"2023-08-24T17:31:50.630645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Signicantly condensed at ~512 mark and a distribution to the left.*","metadata":{}},{"cell_type":"code","source":"sns.histplot(sub_train['crop_width'])","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:50.633866Z","iopub.execute_input":"2023-08-24T17:31:50.634981Z","iopub.status.idle":"2023-08-24T17:31:51.1039Z","shell.execute_reply.started":"2023-08-24T17:31:50.634941Z","shell.execute_reply":"2023-08-24T17:31:51.102497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_train['crop_height'].describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:51.105703Z","iopub.execute_input":"2023-08-24T17:31:51.10615Z","iopub.status.idle":"2023-08-24T17:31:51.11957Z","shell.execute_reply.started":"2023-08-24T17:31:51.106112Z","shell.execute_reply":"2023-08-24T17:31:51.118309Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Peak at ~ 502 and a distribution to the left.","metadata":{}},{"cell_type":"code","source":"sns.histplot(sub_train['crop_height'])","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:51.121069Z","iopub.execute_input":"2023-08-24T17:31:51.121518Z","iopub.status.idle":"2023-08-24T17:31:51.436377Z","shell.execute_reply.started":"2023-08-24T17:31:51.121484Z","shell.execute_reply":"2023-08-24T17:31:51.435192Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Taking a look at one of the outliers is the max cropped width.","metadata":{}},{"cell_type":"code","source":"sub_train[sub_train['crop_width'] == 1024.0].index[0]","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:51.439587Z","iopub.execute_input":"2023-08-24T17:31:51.440475Z","iopub.status.idle":"2023-08-24T17:31:51.450341Z","shell.execute_reply.started":"2023-08-24T17:31:51.440423Z","shell.execute_reply":"2023-08-24T17:31:51.449118Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"max_crop_width = 1024.0\nout_ind = sub_train[sub_train['crop_width'] == max_crop_width]\nimg_data = out_ind.loc[out_ind.index[0], :]\nimg = preprocess(img_data)\n\nplt.imshow(img, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:51.451747Z","iopub.execute_input":"2023-08-24T17:31:51.45216Z","iopub.status.idle":"2023-08-24T17:31:51.900379Z","shell.execute_reply.started":"2023-08-24T17:31:51.452119Z","shell.execute_reply":"2023-08-24T17:31:51.899126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Cropping the outlier with the width and height means.*","metadata":{}},{"cell_type":"code","source":"crop_means = (476, 377)\nimg_resize = cv2.resize(img, crop_means)\n\nplt.imshow(img_resize, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:51.90212Z","iopub.execute_input":"2023-08-24T17:31:51.902951Z","iopub.status.idle":"2023-08-24T17:31:52.211137Z","shell.execute_reply.started":"2023-08-24T17:31:51.902902Z","shell.execute_reply":"2023-08-24T17:31:52.209979Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compress (PNG)\n\nEncoding an image to png for TFRecords.","metadata":{}},{"cell_type":"code","source":"img = img_resize[..., tf.newaxis]\npng = tf.io.encode_png(img)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:52.212792Z","iopub.execute_input":"2023-08-24T17:31:52.213742Z","iopub.status.idle":"2023-08-24T17:31:52.350494Z","shell.execute_reply.started":"2023-08-24T17:31:52.213699Z","shell.execute_reply":"2023-08-24T17:31:52.349196Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Making sure its correct.*","metadata":{}},{"cell_type":"code","source":"plt.imshow(tf.io.decode_png(png), cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:52.351955Z","iopub.execute_input":"2023-08-24T17:31:52.352348Z","iopub.status.idle":"2023-08-24T17:31:52.666236Z","shell.execute_reply.started":"2023-08-24T17:31:52.352314Z","shell.execute_reply":"2023-08-24T17:31:52.665008Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# All-Together\n\nThe final ***(or so I thought it was)*** preprocessing function that can be fed into a TFRecord writer. ","metadata":{}},{"cell_type":"code","source":"def final_preprocess(inst):\n    dcm = pydicom.dcmread(inst['img_path'])\n    \n    img = standardize_pixel_array(dcm)\n    img = truncation_normalization(img)\n    img = (img * 255).astype('uint8')\n    (x, y, w, h) = crop_coords(img)\n    img = img[y:y+h, x:x+w]\n    \n    img_resize = cv2.resize(img, crop_means)\n    \n    img = img_resize[..., tf.newaxis]\n    png = tf.io.encode_png(img)\n    \n    return png","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:52.667799Z","iopub.execute_input":"2023-08-24T17:31:52.668178Z","iopub.status.idle":"2023-08-24T17:31:52.675827Z","shell.execute_reply.started":"2023-08-24T17:31:52.668142Z","shell.execute_reply":"2023-08-24T17:31:52.674803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Getting random int with tensorflow.*","metadata":{}},{"cell_type":"code","source":"def random_int(shape=[], minval=0, maxval=1):\n    return tf.random.uniform(shape=shape, minval=minval, maxval=maxval, dtype=tf.int32)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:52.677215Z","iopub.execute_input":"2023-08-24T17:31:52.677563Z","iopub.status.idle":"2023-08-24T17:31:52.691852Z","shell.execute_reply.started":"2023-08-24T17:31:52.677534Z","shell.execute_reply":"2023-08-24T17:31:52.690936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Display\n\nDisplay and interpret our preprocessed results by comparing random selections.","metadata":{}},{"cell_type":"code","source":"sub_train_ind = sub_train.index\n\ndef display_random_inst(inst_num, axarr):\n    rows = []\n    \n    for i in range(inst_num):\n        ind = int(random_int(maxval=sub_train.shape[0]-1))\n        dex = sub_train_ind[ind]\n        inst = sub_train.loc[dex, :]\n        \n        rows.append(dex)\n\n        axarr[i][1].imshow(tf.io.decode_png(final_preprocess(inst)), cmap='gray')\n        axarr[i][0].imshow(pydicom.dcmread(inst['img_path']).pixel_array, cmap='gray')\n        \n    return rows","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:52.693418Z","iopub.execute_input":"2023-08-24T17:31:52.694135Z","iopub.status.idle":"2023-08-24T17:31:52.706446Z","shell.execute_reply.started":"2023-08-24T17:31:52.694087Z","shell.execute_reply":"2023-08-24T17:31:52.705404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"inst_num = 20\n\nf, axarr = plt.subplots(inst_num, 2) \nf.set_figheight(inst_num*8)\nf.set_figwidth(15)\n\nrows = display_random_inst(inst_num, axarr)\n\nfor ax, col in zip(axarr[0], ['BASE', 'PREPROCESSED']):\n    ax.set_title(col, color='red')\n    \nfor ax, row in zip(axarr[:,0], rows):\n    ax.set_ylabel(row, rotation=0, size='large', color='red')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:31:52.707993Z","iopub.execute_input":"2023-08-24T17:31:52.708676Z","iopub.status.idle":"2023-08-24T17:32:08.328602Z","shell.execute_reply.started":"2023-08-24T17:31:52.708643Z","shell.execute_reply":"2023-08-24T17:32:08.327172Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks like there is some series artifacting.\n\nNext two things on agenda:\n1. Investiage the (dots) artifacts.\n> Fixed with different standarized technique.\n2. Work to reduce the circular lighting.\n> Work to tune the bottom percentile clip based on pixel intensities.\n3. Tweak Threshold so things don't get cut off.\n> First figure out if something not the main subject is even useful.","metadata":{}},{"cell_type":"markdown","source":"# Agenda-One\n\nAddressing everything from the first agenda.\n\n### Standardization (REVISION)\n\nShift the values to positive scale, standardize, then clip based on intesity. (called withing the next function) ","metadata":{}},{"cell_type":"code","source":"def norm(img, intesity):    \n    pix_min = img.min()\n    pix_max = img.max() + (-1 * pix_min)\n\n    img = img + (-1 * pix_min)\n    img =  img / pix_max\n    \n    pmin = np.percentile(img, intesity)\n    pmax = np.percentile(img, 100)\n    \n    img = (img * 255).astype('uint8')\n    \n    return img, intesity","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:32:08.330392Z","iopub.execute_input":"2023-08-24T17:32:08.330787Z","iopub.status.idle":"2023-08-24T17:32:08.338119Z","shell.execute_reply.started":"2023-08-24T17:32:08.330744Z","shell.execute_reply":"2023-08-24T17:32:08.336889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ind = 867323\ninst = sub_train.loc[ind, :]\nimg = pydicom.dcmread(inst['img_path']).pixel_array\n\nf, axes = plt.subplots(2, 2)\nf.set_figheight(16)\nf.set_figwidth(15)\n\naxes[0][0].imshow(img, cmap='gray')\naxes[1][0].imshow(truncation_normalization(img), cmap='gray')\naxes[0][1].imshow(img, cmap='gray')\nnorm_img, intensity = norm(img, 30)\naxes[1][1].imshow(img, cmap='gray')\n\nfor ax, col in zip(axes[0], ['PREVIOUS-METHOD', 'NEW-METHOD']):\n    ax.set_title(col, color='red')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:32:08.339627Z","iopub.execute_input":"2023-08-24T17:32:08.340003Z","iopub.status.idle":"2023-08-24T17:32:09.805819Z","shell.execute_reply.started":"2023-08-24T17:32:08.33997Z","shell.execute_reply":"2023-08-24T17:32:09.804554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Reactive Clipping\n\nGet the relative strength needed for the bottom percentile clip based on the difference between min and max pixel intensity.\n\nThe quality of the crop is greatly impacted by the clipping done.\n\n*Using a repeating norm to recompute min-max distance until it is at least greater than 50.*","metadata":{}},{"cell_type":"code","source":"def testing_inst(inst, intensity, minmax, only_norm=False):\n    dcm = pydicom.dcmread(inst['img_path'])\n    base_img = standardize_pixel_array(dcm)\n    img = np.copy(base_img)\n    \n    pix_min = img.min()\n    pix_max = img.max() + (-1 * pix_min)\n\n    img = img + (-1 * pix_min)\n    img =  img / pix_max\n            \n    if only_norm:\n        return img\n        \n    pmin = np.percentile(img, intensity)\n    pmax = np.percentile(img, 100)\n    \n    img = np.clip(img, pmin, pmax)  \n    img = (img * 255).astype('uint8')\n            \n    if minmax:\n        minmax_diff = (img.max() - img.min())\n\n        while not (minmax_diff > 205):\n            img, intensity = norm(np.copy(base_img), intensity-2)\n            minmax_diff = (img.max() - img.min())\n                    \n    (x, y, w, h) = crop_coords(img)\n    img = img[y:y+h, x:x+w]\n    \n    return img","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:32:09.807492Z","iopub.execute_input":"2023-08-24T17:32:09.807918Z","iopub.status.idle":"2023-08-24T17:32:09.81936Z","shell.execute_reply.started":"2023-08-24T17:32:09.80788Z","shell.execute_reply":"2023-08-24T17:32:09.818464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**NOTE:** This is an unused method in the final preprocessing. It is left here to show progression and give inspiration. The previous version had instances where it worked but I couldn't find more so don't mind the useless image.","metadata":{}},{"cell_type":"code","source":"ind = 23297 # 6794 # 10020 # 8378 # 7122 # 4875 # 11807 # 10886 # 8243\ninst = train.loc[ind, :]\n\nf, axes = plt.subplots(2, 2)\nf.set_figheight(16)\nf.set_figwidth(15)\n\naxes[0][0].imshow(testing_inst(inst, None, False, only_norm=True), cmap='gray')\naxes[1][0].imshow(testing_inst(inst, 50, False), cmap='gray')\naxes[0][1].imshow(testing_inst(inst, None, False, only_norm=True), cmap='gray')\naxes[1][1].imshow(testing_inst(inst, 50, True), cmap='gray')\n\nfor ax, col in zip(axes[0], ['NO-MINMAX', 'WITH-MINMAX']):\n    ax.set_title(col, color='red')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:32:09.820782Z","iopub.execute_input":"2023-08-24T17:32:09.821154Z","iopub.status.idle":"2023-08-24T17:32:11.299132Z","shell.execute_reply.started":"2023-08-24T17:32:09.821112Z","shell.execute_reply":"2023-08-24T17:32:11.297893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the NO-MINMAX there is is no drop in clipping value. \n\nIn the WITH-MINMAX there is a recurring 5 point drop while the min-max difference in pixel intensities is not greater than 50.\n\n**The quality of the crop has greatly improved.**","metadata":{}},{"cell_type":"markdown","source":"Next on agenda:\n1. Compute the best starter intensity that leads to least amount of secondary calls. (for complexity reasons)\n> To ensure the quality of the data there should be at least one call per instance.\n2. Put it all together. (again)\n> The results are ok but more work is never bad.","metadata":{}},{"cell_type":"markdown","source":"# All-Together (Part 2)\n\nThe ***final*** preprocessing function that can be fed into a TFRecord writer. \n\n*The new standardization function to drop clipping intensity.*","metadata":{}},{"cell_type":"code","source":"def standardize(base_img):\n    img = np.copy(base_img)\n    \n    pix_min = img.min()\n    pix_max = img.max() + (-1 * pix_min)\n\n    img = img + (-1 * pix_min)\n    img =  img / pix_max\n    \n    return img","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:32:11.300798Z","iopub.execute_input":"2023-08-24T17:32:11.301251Z","iopub.status.idle":"2023-08-24T17:32:11.30881Z","shell.execute_reply.started":"2023-08-24T17:32:11.301209Z","shell.execute_reply":"2023-08-24T17:32:11.307541Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Clipping/Cropping (THIRD REVISION)\n\nAfter more work the above intensity approach has been dropped in favour for a scaling pmin percentile clip and a constant pmax percentile clip of 85.","metadata":{}},{"cell_type":"code","source":"def clipping(img):\n    pnc = 10\n    pmin = 0.0\n    \n    #while pmin < 0.15:\n    #    pmin = np.percentile(img, pnc)\n    #    pnc+=1\n        \n    pmin = np.percentile(img, 50)\n    pmax = np.percentile(img, 75)\n            \n    img = np.clip(img, pmin, pmax)  \n    img = (img * 255).astype('uint8')\n    \n    return img","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:32:11.310064Z","iopub.execute_input":"2023-08-24T17:32:11.310455Z","iopub.status.idle":"2023-08-24T17:32:11.32738Z","shell.execute_reply.started":"2023-08-24T17:32:11.31042Z","shell.execute_reply":"2023-08-24T17:32:11.325686Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Clip/Crop Visualizations\n\nExamining some past problem instances clips and following crops.","metadata":{}},{"cell_type":"code","source":"ind = 1037489 # 1467344 # 324809 # 1210155 # 1109208 # 79408 # 23297 # 518781 # 403688\ninst = train.loc[ind, :]\n\ndcm = pydicom.dcmread(inst['img_path'])\nimg = standardize_pixel_array(dcm)\nimg = standardize(img)\n\nimg_o = img\n\nloss_less_crop = np.copy(img)\nimg = clipping(img)\n\nimg_t = img\n\n(x, y, w, h) = crop_coords(img)\nimg = loss_less_crop[y:y+h, x:x+w]\nimg = img[..., tf.newaxis]\nimg = (img * 65535).astype('uint16')\n\nimg_th = img\n\nf, axes = plt.subplots(1, 3)\nf.set_figheight(30)\nf.set_figwidth(30)\n\naxes[0].imshow(img_o, cmap='gray')\naxes[1].imshow(img_t, cmap='gray')\naxes[2].imshow(img_th, cmap='gray')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:32:11.328763Z","iopub.execute_input":"2023-08-24T17:32:11.329138Z","iopub.status.idle":"2023-08-24T17:32:12.564928Z","shell.execute_reply.started":"2023-08-24T17:32:11.329081Z","shell.execute_reply":"2023-08-24T17:32:12.563617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Loss-Less Cropping\n\nThe new clipping method deteriorates the image so a lossless clip will remedy it.","metadata":{}},{"cell_type":"code","source":"def final_preprocess(inst, png=True):\n    dcm = pydicom.dcmread(inst['img_path'])\n    \n    img = standardize_pixel_array(dcm)\n    img = standardize(img)\n    \n    loss_less_crop = np.copy(img)\n    \n    img = clipping(img)\n    \n    try:\n        (x, y, w, h) = crop_coords(img)\n        img = loss_less_crop[y:y+h, x:x+w]\n    except ValueError:\n        img = loss_less_crop\n    \n    # img = 255 - img  # uncomment to invert\n    \n    #img_resize = cv2.resize(img, crop_means)\n    \n    img = img[..., tf.newaxis]\n    img = (img * 65535).astype('uint16')\n    \n    if png:\n        png = tf.io.encode_png(img)\n    else:\n        img = np.squeeze(img)\n        return img\n    \n    return png","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:32:12.566444Z","iopub.execute_input":"2023-08-24T17:32:12.566807Z","iopub.status.idle":"2023-08-24T17:32:12.575327Z","shell.execute_reply.started":"2023-08-24T17:32:12.566773Z","shell.execute_reply":"2023-08-24T17:32:12.574132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Displaying random instances using the same function from before.","metadata":{}},{"cell_type":"code","source":"inst_num = 50 \n\nf, axarr = plt.subplots(inst_num, 2) \nf.set_figheight(inst_num*8)\nf.set_figwidth(15)\n\nrows = display_random_inst(inst_num, axarr)\n\nfor ax, col in zip(axarr[0], ['BASE', 'PREPROCESSED']):\n    ax.set_title(col, color='red')\n    \nfor ax, row in zip(axarr[:,0], rows):\n    ax.set_ylabel(row, rotation=0, size='large', color='red')","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:33:46.238768Z","iopub.execute_input":"2023-08-24T17:33:46.239231Z","iopub.status.idle":"2023-08-24T17:34:25.30373Z","shell.execute_reply.started":"2023-08-24T17:33:46.239196Z","shell.execute_reply":"2023-08-24T17:34:25.302141Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Crop Means (REVISION)\n\nThe new crop sizes based on the improvements for cropping accuracy.","metadata":{}},{"cell_type":"code","source":"def cropped_sizes(row):\n    dcm = pydicom.dcmread(row['img_path'])\n    \n    img = standardize_pixel_array(dcm)\n    img = standardize(img)    \n    img = clipping(img)\n    \n    if img.sum() == 0:\n        return False\n    \n    (x, y, w, h) = crop_coords(img)\n    return (x, y, w, h)\n    \n\nsub_train['crop_raw_size'] = sub_train.apply(cropped_sizes, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:34:25.306443Z","iopub.execute_input":"2023-08-24T17:34:25.307362Z","iopub.status.idle":"2023-08-24T17:36:46.011111Z","shell.execute_reply.started":"2023-08-24T17:34:25.307309Z","shell.execute_reply":"2023-08-24T17:36:46.009758Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Some imcomplete images can be completely removed by the percentile clipping. Its best to just ignore these in the finaly dataset.","metadata":{}},{"cell_type":"code","source":"sub_train[sub_train['crop_raw_size'] == False]","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:46.013567Z","iopub.execute_input":"2023-08-24T17:36:46.01395Z","iopub.status.idle":"2023-08-24T17:36:46.040304Z","shell.execute_reply.started":"2023-08-24T17:36:46.013914Z","shell.execute_reply":"2023-08-24T17:36:46.039081Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Getting the statistics for the crop data.*","metadata":{}},{"cell_type":"code","source":"def get_w_data(row):\n    return (row['crop_raw_size'][2] if row['crop_raw_size'] else 256)\n\ndef get_h_data(row):\n    return (row['crop_raw_size'][3] if row['crop_raw_size'] else 256)\n\ndef get_ratio_data(row):\n    if not row['crop_raw_size']:\n        return 1.0\n    return row['crop_raw_size'][2] / row['crop_raw_size'][3]\n\nsub_train['crop_ratio'] = sub_train.apply(get_ratio_data, axis=1)\nsub_train['crop_width'] = sub_train.apply(get_w_data, axis=1)\nsub_train['crop_height'] = sub_train.apply(get_h_data, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:46.041662Z","iopub.execute_input":"2023-08-24T17:36:46.041995Z","iopub.status.idle":"2023-08-24T17:36:46.344674Z","shell.execute_reply.started":"2023-08-24T17:36:46.041965Z","shell.execute_reply":"2023-08-24T17:36:46.343627Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Taking a look at the crop data again.","metadata":{}},{"cell_type":"code","source":"sub_train['crop_ratio'].describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:46.347178Z","iopub.execute_input":"2023-08-24T17:36:46.347511Z","iopub.status.idle":"2023-08-24T17:36:46.361354Z","shell.execute_reply.started":"2023-08-24T17:36:46.347485Z","shell.execute_reply":"2023-08-24T17:36:46.360299Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.histplot(sub_train['crop_ratio'])","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:46.362793Z","iopub.execute_input":"2023-08-24T17:36:46.363165Z","iopub.status.idle":"2023-08-24T17:36:46.790508Z","shell.execute_reply.started":"2023-08-24T17:36:46.363118Z","shell.execute_reply":"2023-08-24T17:36:46.789462Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_train['crop_width'].describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:46.792015Z","iopub.execute_input":"2023-08-24T17:36:46.792371Z","iopub.status.idle":"2023-08-24T17:36:46.805588Z","shell.execute_reply.started":"2023-08-24T17:36:46.792342Z","shell.execute_reply":"2023-08-24T17:36:46.804416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.histplot(sub_train['crop_width'])","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:46.807181Z","iopub.execute_input":"2023-08-24T17:36:46.807535Z","iopub.status.idle":"2023-08-24T17:36:47.282376Z","shell.execute_reply.started":"2023-08-24T17:36:46.807505Z","shell.execute_reply":"2023-08-24T17:36:47.28119Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_train['crop_height'].describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:47.283797Z","iopub.execute_input":"2023-08-24T17:36:47.284275Z","iopub.status.idle":"2023-08-24T17:36:47.298992Z","shell.execute_reply.started":"2023-08-24T17:36:47.284239Z","shell.execute_reply":"2023-08-24T17:36:47.297667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.histplot(sub_train['crop_height'])","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:47.300722Z","iopub.execute_input":"2023-08-24T17:36:47.301058Z","iopub.status.idle":"2023-08-24T17:36:47.700426Z","shell.execute_reply.started":"2023-08-24T17:36:47.301028Z","shell.execute_reply":"2023-08-24T17:36:47.69909Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It looks like a safe resize would be (451, 303) though there is a significant amount of variation.\n\nNext things on agenda:\n1. Create a dataset from cropped images.\n\n2. Any more possible pre-processing???\n\n# JPEG Dataset\n\nCreating a dataset of cropped JPEG images using the preprocessing function.\n\n*This was intially a png dataset but I decided 1.5 million images of pngs would be too much space.*","metadata":{}},{"cell_type":"code","source":"for i in os.listdir('/kaggle/working/'):\n    if i == '.virtual_documents':\n        continue\n    os.remove(i)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:36:47.704328Z","iopub.execute_input":"2023-08-24T17:36:47.704835Z","iopub.status.idle":"2023-08-24T17:36:47.710889Z","shell.execute_reply.started":"2023-08-24T17:36:47.704786Z","shell.execute_reply":"2023-08-24T17:36:47.709349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def save_jpgs(row):\n    img = final_preprocess(row, png=False)\n    plt.imsave(\n        f\"/kaggle/working/{row['patient_id']}-{row['series_id']}-{row['instance_number']}.jpg\", \n        img, format='jpeg', cmap='gray'\n    )\n    \ntrain.apply(save_jpgs, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-24T17:37:36.142641Z","iopub.execute_input":"2023-08-24T17:37:36.143149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.to_csv('updated_train.csv')","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}