{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## A First Simple Baseline - Submit Weighted Mean Probability for Each Label\n\nAs a first starting point in the competition, we'll compute the mean probabilities for each tag/label that we have to predict. We'll use the information in **train.csv** to compute the mean probabilites for each injury type.\n\nFor each patient in the test set, we simply return the computed mean values.\n","metadata":{}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport os\nfrom sklearn.model_selection import train_test_split\nfrom rsna_2023_atd_metric import score","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-03T10:30:29.515495Z","iopub.execute_input":"2023-10-03T10:30:29.515891Z","iopub.status.idle":"2023-10-03T10:30:30.193288Z","shell.execute_reply.started":"2023-10-03T10:30:29.515863Z","shell.execute_reply":"2023-10-03T10:30:30.192076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We read in the **train.csv** file.","metadata":{}},{"cell_type":"code","source":"BASE_DIR = '/kaggle/input/rsna-2023-abdominal-trauma-detection/'\npatient_labels = pd.read_csv(BASE_DIR + 'train.csv', index_col = False)\npatient_labels = patient_labels.reset_index(drop = True)\npatient_labels.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:30:38.726492Z","iopub.execute_input":"2023-10-03T10:30:38.727816Z","iopub.status.idle":"2023-10-03T10:30:38.771963Z","shell.execute_reply.started":"2023-10-03T10:30:38.72774Z","shell.execute_reply":"2023-10-03T10:30:38.770705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_patients, validation_patients = train_test_split(patient_labels, test_size = 0.2, random_state = 42) ","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:47:53.871548Z","iopub.execute_input":"2023-10-03T10:47:53.872694Z","iopub.status.idle":"2023-10-03T10:47:53.87995Z","shell.execute_reply.started":"2023-10-03T10:47:53.872657Z","shell.execute_reply":"2023-10-03T10:47:53.87886Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We compute the mean values of the labels we are interested in.","metadata":{}},{"cell_type":"code","source":"target_labels = ['bowel_healthy', 'bowel_injury', 'extravasation_healthy',\n       'extravasation_injury', 'kidney_healthy', 'kidney_low', 'kidney_high',\n       'liver_healthy', 'liver_low', 'liver_high', 'spleen_healthy',\n       'spleen_low', 'spleen_high'] # get from patient_labels.columns, instead of typing\nmean_prob = train_patients[target_labels].mean()\nmean_prob","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:30:45.711725Z","iopub.execute_input":"2023-10-03T10:30:45.712798Z","iopub.status.idle":"2023-10-03T10:30:45.72717Z","shell.execute_reply.started":"2023-10-03T10:30:45.712735Z","shell.execute_reply":"2023-10-03T10:30:45.725475Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The sample weights imply that we care more about getting some types of injury predictions correct than others. So, for instance, we are penalized 3x higher if we miss detecting extravasation injury than if we miss bowel injury.\n\nTo do well on the competition metric, it'll be beneficial to rescale the mean probabilites using the injury sample weights. First, let us make a pandas Series from these sample weights.","metadata":{}},{"cell_type":"code","source":"#     1 for all healthy labels.\n#     2 for low grade solid organ injuries (liver, spleen, kidney).\n#     4 for high grade solid organ injuries.\n#     2 for bowel injuries.\n#     6 for extravasation.\n#     6 for the auto-generated any_injury label.\n\nw_bowel = [1, 2]\nw_extravasation = [1, 6]\nw_kidney = w_liver = w_spleen = [1, 2, 4]","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:33:55.649767Z","iopub.execute_input":"2023-10-03T10:33:55.650193Z","iopub.status.idle":"2023-10-03T10:33:55.655673Z","shell.execute_reply.started":"2023-10-03T10:33:55.650163Z","shell.execute_reply":"2023-10-03T10:33:55.65419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"all_w = [w_bowel] + [w_extravasation] + [w_kidney] + [w_liver] + [w_spleen]\nall_w = pd.Series([item for sublist in all_w for item in sublist], dtype = float, index = mean_prob.index)\nassert mean_prob.shape == all_w.shape\nall_w","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:34:29.249622Z","iopub.execute_input":"2023-10-03T10:34:29.250407Z","iopub.status.idle":"2023-10-03T10:34:29.259534Z","shell.execute_reply.started":"2023-10-03T10:34:29.250359Z","shell.execute_reply":"2023-10-03T10:34:29.25823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we rescale the mean proababilities by the class weights.","metadata":{}},{"cell_type":"code","source":"w_mean_prob = pd.Series([x*y for x, y in zip(mean_prob, all_w)], dtype = float, index = mean_prob.index)\nw_mean_prob","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:35:57.966653Z","iopub.execute_input":"2023-10-03T10:35:57.967134Z","iopub.status.idle":"2023-10-03T10:35:57.978105Z","shell.execute_reply.started":"2023-10-03T10:35:57.967098Z","shell.execute_reply":"2023-10-03T10:35:57.975148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compute Metric on the Validation Set","metadata":{}},{"cell_type":"markdown","source":"Let us now compute the score on our validation set to get an idea of how we'll perform on LB upon submission.\n\n**validation_preds** dataframe will contain predictions on the validation set. **validation_patients** contains the ground truth labels for the validation patients.","metadata":{}},{"cell_type":"code","source":"validation_preds = pd.DataFrame(np.empty((validation_patients.shape[0], validation_patients.shape[1]-1)))\nvalidation_preds.columns = validation_patients.columns[0:-1]\nvalidation_patients = validation_patients.reset_index(drop = True)\nvalidation_preds = validation_preds.reset_index(drop = True)","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:47:58.090225Z","iopub.execute_input":"2023-10-03T10:47:58.090578Z","iopub.status.idle":"2023-10-03T10:47:58.096848Z","shell.execute_reply.started":"2023-10-03T10:47:58.090552Z","shell.execute_reply":"2023-10-03T10:47:58.095552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We make the predictions using the weighted mean proabilities (same for all patients.)","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:38:45.085616Z","iopub.execute_input":"2023-10-03T10:38:45.086259Z","iopub.status.idle":"2023-10-03T10:38:45.093499Z","shell.execute_reply.started":"2023-10-03T10:38:45.086213Z","shell.execute_reply":"2023-10-03T10:38:45.091881Z"}}},{"cell_type":"code","source":"validation_preds['patient_id'] = validation_patients['patient_id'].copy()\nvalidation_preds.iloc[:, 1:] = pd.DataFrame(np.broadcast_to(w_mean_prob, (validation_patients.shape[0], 13)))\nvalidation_preds","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:48:00.282874Z","iopub.execute_input":"2023-10-03T10:48:00.283267Z","iopub.status.idle":"2023-10-03T10:48:00.310208Z","shell.execute_reply.started":"2023-10-03T10:48:00.283236Z","shell.execute_reply":"2023-10-03T10:48:00.308978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For scoring, we need to fill in the appropriate values for the different weight columns. So, *bowel_weight* column contains 2 if patient had bowel injury, else 1. The following simple logic does the job:","metadata":{}},{"cell_type":"code","source":"validation_patients['bowel_weight'] = pd.Series(np.maximum(1*validation_patients.bowel_healthy.to_numpy(), 2*validation_patients.bowel_injury.to_numpy()))\nvalidation_patients['extravasation_weight'] = pd.Series(np.maximum(1*validation_patients.extravasation_healthy.to_numpy(), 6*validation_patients.extravasation_injury.to_numpy()))\nvalidation_patients['kidney_weight'] = pd.Series(np.maximum.reduce([1*validation_patients.kidney_healthy.to_numpy(), 2*validation_patients.kidney_low.to_numpy(), 4*validation_patients.kidney_high.to_numpy()]))\nvalidation_patients['spleen_weight'] = pd.Series(np.maximum.reduce([1*validation_patients.spleen_healthy.to_numpy(), 2*validation_patients.spleen_low.to_numpy(), 4*validation_patients.spleen_high.to_numpy()]))\nvalidation_patients['liver_weight'] = pd.Series(np.maximum.reduce([1*validation_patients.liver_healthy.to_numpy(), 2*validation_patients.liver_low.to_numpy(), 4*validation_patients.liver_high.to_numpy()]))\nvalidation_patients['any_injury_weight'] = pd.Series([6]*validation_patients.shape[0])","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:48:02.681056Z","iopub.execute_input":"2023-10-03T10:48:02.681418Z","iopub.status.idle":"2023-10-03T10:48:02.694777Z","shell.execute_reply.started":"2023-10-03T10:48:02.681392Z","shell.execute_reply":"2023-10-03T10:48:02.693593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can now call the scoring script (here imported as a utility file) to get the score on the validation set.","metadata":{}},{"cell_type":"code","source":"score(validation_patients, validation_preds, 'patient_id')","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:49:51.187812Z","iopub.execute_input":"2023-10-03T10:49:51.188183Z","iopub.status.idle":"2023-10-03T10:49:51.981658Z","shell.execute_reply.started":"2023-10-03T10:49:51.188156Z","shell.execute_reply":"2023-10-03T10:49:51.979908Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Prediction on Test Set","metadata":{}},{"cell_type":"markdown","source":"We get a list of all the patients in the test set for whom we need to make predictions.","metadata":{}},{"cell_type":"code","source":"test_patient_ids = os.listdir(BASE_DIR + 'test_images')","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:48:11.337168Z","iopub.execute_input":"2023-10-03T10:48:11.337566Z","iopub.status.idle":"2023-10-03T10:48:11.344732Z","shell.execute_reply.started":"2023-10-03T10:48:11.337535Z","shell.execute_reply":"2023-10-03T10:48:11.343888Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We use the full data available (training + validation) to recompute the mean for better estimates.","metadata":{}},{"cell_type":"code","source":"mean_prob = patient_labels[target_labels].mean()\nw_mean_prob = pd.Series([x*y for x, y in zip(mean_prob, all_w)], dtype = float, index = mean_prob.index)","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:51:34.310067Z","iopub.execute_input":"2023-10-03T10:51:34.310461Z","iopub.status.idle":"2023-10-03T10:51:34.318386Z","shell.execute_reply.started":"2023-10-03T10:51:34.310431Z","shell.execute_reply":"2023-10-03T10:51:34.317275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We create a new dataframe with appropriate column names (based on what is asked in Evaluation section of competition).","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame(test_patient_ids, columns = ['patient_id'])\nfor label in target_labels:\n    submission[label] = w_mean_prob[label]\nsubmission","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:51:40.729838Z","iopub.execute_input":"2023-10-03T10:51:40.730229Z","iopub.status.idle":"2023-10-03T10:51:40.751558Z","shell.execute_reply.started":"2023-10-03T10:51:40.730198Z","shell.execute_reply":"2023-10-03T10:51:40.750772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We write our predictions to the submission csv file.","metadata":{}},{"cell_type":"code","source":"submission.to_csv('submission.csv', header = True, index = False)","metadata":{"execution":{"iopub.status.busy":"2023-10-03T10:51:45.845693Z","iopub.execute_input":"2023-10-03T10:51:45.846084Z","iopub.status.idle":"2023-10-03T10:51:45.853133Z","shell.execute_reply.started":"2023-10-03T10:51:45.846056Z","shell.execute_reply":"2023-10-03T10:51:45.852045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As expected, this solution is not very good; we get LB score of 0.75. We're now going to iterate upon our submission.","metadata":{}}]}