{"cells":[{"metadata":{},"cell_type":"markdown","source":"## Only purpose of this notebook is to share a post-processing technique of making 14th class probability to 0 when we have found >0 bounding boxes.\n\n## To show that, I have used [@nxhong93](https://www.kaggle.com/nxhong93)'s public notebook output which you can find [here](https://www.kaggle.com/nxhong93/yolov5-chest-512). Thanks [@nxhong93](https://www.kaggle.com/nxhong93) for sharing the notebook. All credits goes to him.\n\n## So, for samples we don't have bounding boxes, we write 14 1 0 0 1 1 in the submission which means, we found 14th class i.e. \"No finding\" with probability of 1.\n\n## Now, if we found some bounding box, then intuitively we can make that probability to 0, if model predicts it non-zero.\n\n### And obviosly, I am not saying that, this will also work in private test-data, but chances are there that, it may work.\n\n### As there are 21 days to go for this competition to end, everyone should get a fair chance to use this if they want."},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"import pandas as pd\npd.set_option('max_colwidth', None)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# need to create a dataset of the CSV file, as that kernel was not appearing in search results while importing data.\nsub = pd.read_csv('../input/vinbigdata-0235-lb/submission_0.235_LB.csv')\nsub.shape","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"def divide(l, n):\n    '''\n    divide submission string into group of 6\n    '''\n    for i in range(0, len(l), n):  \n        yield l[i:i + n]\n\npreds = sub['PredictionString'].tolist()\ngrouped_preds = [list(divide(pred.split(), 6)) for pred in preds]\ngrouped_preds[:5]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"new_preds = []\n\nfor pred in grouped_preds:\n    temp = ''\n    # each box is a tuple of 6 i.e. (class, confidence, xmin, ymin, xmax, ymax)\n    for box in pred:\n        # if we found some bounding-box i.e. `len(pred) > 1` & class is \"No finding\".\n        if len(pred) > 1 and box[0] == '14':\n            # Make the probability 0.\n            box[1] = '0'\n        temp += ' '.join(box) + ' '\n    new_preds.append(temp.strip())\n    \nnew_preds[:5]","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub['PredictionString'] = new_preds\nsub.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"sub.to_csv('submission_postprocessed.csv', index=False)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## On public LB, we're getting improvement of 0.04 i.e. 0.235 to 0.239"}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":4}