{"cells":[{"metadata":{},"cell_type":"markdown","source":"This is just a demo submission to understand the format of the submission. This has nothing to do with the actual submission. If you think that this clarifies your doubt regarding sample submission, have a look at this notebook,I hope you will understand this. "},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport glob, os\n","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","trusted":true},"cell_type":"code","source":"PATH = \"../input/rsna-str-pulmonary-embolism-detection/\"\n\ntrain_df = pd.read_csv(PATH + \"train.csv\")\ntest_df = pd.read_csv(PATH + \"test.csv\")\n\nTRAIN_PATH = PATH + \"train/\"\nTEST_PATH = PATH + \"test/\"\nsub = pd.read_csv(PATH + \"sample_submission.csv\")\ntrain_image_file_paths = glob.glob(TRAIN_PATH + '/*/*/*.dcm')\ntest_image_file_paths = glob.glob(TEST_PATH + '/*/*/*.dcm')\n\nprint(f'Train dataframe shape  :{train_df.shape}')\nprint(f'Test dataframe shape   :{test_df.shape}')\n\nprint(f'Number of train images : {len(train_image_file_paths)}')\nprint(f'Number of test images  : {len(test_image_file_paths)}')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"exam_level_features = ['negative_exam_for_pe', 'rv_lv_ratio_gte_1', 'rv_lv_ratio_lt_1',\n                       'leftsided_pe',         'chronic_pe',        'rightsided_pe', \n                       'acute_and_chronic_pe', 'central_pe',        'indeterminate']","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"Now let us have a look at the sample submission file so that we can understand what to predict from all the images. **"},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"from tqdm.notebook import tqdm\nprediction_counts = {}\nfor idx in tqdm(range(sub.shape[0])):\n    if len(sub['id'][idx][13:]) > 1:\n        key = sub['id'][idx][13:]\n    else:\n        key = 'pe_present_on_image'\n    prediction_counts[key] = prediction_counts.get(key, 0) + 1\nprint(f'Total row count in submission: {sub.shape[0]}')\nprediction_counts","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Submission size calculation for correctness: \n\nHere among the labels given in the training data, `pe_present_on_image` is the image level feature that needs to be predicted for all the images. \n\nAnd the rest of the features will be predicted for only the exam (the observation). In that case each exam will have multiple images. But for that whole group we will submit only one set of prediction for those following labels:\n \n * negative_exam_for_pe\n * rv_lv_ratio_gte_1\n * rv_lv_ratio_lt_1\n * leftsided_pe\n * chronic_pe\n * rightsided_pe\n * acute_and_chronic_pe\n * central_pe\n * indeterminate\n \nTherefore our prediction file should have a number of rows equal to: $$(N_{img}) + (N_{exams} * N_{examlevelfeature}).$$\n\nFor example, In the above calculation we have 650 examination and 146853 images in total. So the total number of rows in the submission file will "},{"metadata":{"trusted":true},"cell_type":"code","source":"N_img = len(test_image_file_paths)\nN_exams = len(os.listdir(TEST_PATH))\nN_exam_level_features = len(exam_level_features)\n\ntotal_rows_submission = N_img + (N_exams * N_exam_level_features)\nprint(f'Total row count in submission: {total_rows_submission}')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"StudyInstanceUIDs = os.listdir(TEST_PATH)\nSOPInstanceUIDs = [filename[-16:-4] for filename in test_image_file_paths]\n\nsubmission_rows = []\n\nfor exam_level_feature in exam_level_features:\n    for StudyInstanceUID in StudyInstanceUIDs:\n        submission_rows.append(StudyInstanceUID+'_'+exam_level_feature)\n\nsubmission_rows = submission_rows + SOPInstanceUIDs \nprint(f'Total row count in submission: {len(submission_rows)}')","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"len(submission_rows)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission_file = pd.DataFrame({'id': submission_rows, 'label': (np.zeros(len(submission_rows))+0.35)})","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"submission_file.head()\nsubmission_file.to_csv('submission.csv', index = False)","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}