{
  "id": 220295,
  "title": "What anchor size & aspect ratio should be used?",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/220295",
  "author_name": "corochann",
  "post_date": "2021-02-17T22:54:01.047000",
  "votes": 30,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>EDA</h1>\n<h2>Small bounding box</h2>\n<p>Now I understand that the challenge of this competition is to predict wide range size of the bounding boxes.<br>\nEspecially for <strong>small sized bounding box</strong>.</p>\n<p>When I check the AP by class, \"Calcification\", \"Nodule/Mass\" and \"Pneumothorax\" were the difficult categories.</p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/54411746/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..aLpE9X3UbjaJB2RiJORRpA.2jADfQTTkBT-Ad4s8Sus-JUCGl1jQT2clRnOTHU3FXUHvPwj7zvPvJG7SojA8Wc-ZFVhzwTmwLctRZf5XYcsnqy3kPrDO0w4SQdg8MvxcTKjrO2H0KLhQyZNTq1iJlFQDmFfYcaVN2pdFztCw4hLqPjhJcShBn-uayrUhEX25SyiFhDMK10TfDxsjM9HCXg1cnFdy2suxdzoxnGFp_3fJZkSi5XYXgonDqjfeteiDpQjc-mcT1x5h7Y6LoiyCi27RK5SPImPLiYpc1P4NmzX0VOA-jJrLTmTEbJAfidCDe505Qufx0BDyIBq75tvCLFXcC9ng_jVdn79W1mAW3X6-GITNM47UR4sykWik4TeWnBTlDrr5dgjK-GLlocupJIjoyfl-UDxhBriy276K5PCrqEK93DIv7ZpRcnmclBZ53mJtAIkUXfs4SbbCgusoqZeNCfTcDugkAnQjsGcdoJEjjYY3f94MkThxJx_F2Ro1kTWFdhqyswHJkG8OCrRNYOPeZu7BZg1RNj8_rBTCDpotXJezhxs5NhwVfu_kLpWcJ9Esc565K9ND9KrpAIamIqnzOhGgWj7yYW1IDCLAfo0cgTl6q58EmA6xRcNWJI3D9Jk1RuCe9zyI2RvRXpG0QqIVto14ycw1kK-qqPNN-EGG92PwGTWXNTfyb7HjrUIn4E.2Wp5qo9vOwE1SvZyzHsN-A/__results___files/__results___45_0.png\" alt=\"\"><br>\nFrom <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train#VinBigData-detectron2-train\" target=\"_blank\">📸VinBigData detectron2 train</a>.</p>\n<p>And I see that bounding box is really small in these categories.<br>\nFor example, this is the \"Nodule/Mass\" bounding boxes:<br>\n<img src=\"https://www.kaggleusercontent.com/kf/53546324/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..3JgopvaHUEk2iyxwddnyhA.IP2k4_iRRQEpZjcnkwq9z4WnWtdYRgMbeVw3pVMQzBcsfZrYo_mB48GxE6qxZQzi0iRnx6mLdMOax0z8YEkG41G_UXIPcIB_LTOFt4TeDVltsDAcE0WfSn2PlqsUWpqG7xhGHkU72mJGaJh3VBjQSIRwHQhbasiygWfF7VJ58ecDQ1h5Wtdsi-G_PtOPRSxp2o25R8xJvmeUdbGG8njC3i2micR7cDAw7Pdok6lA0U-Q7l3j2BI0ZgYBcI99jiHs3wbd13P-72WI8rCkUOkakb0yTxHljFASyT0ZruiipkAT6sova0YmMMDDR6l1Ov0zsGeZogz8qZNXEoKAVF9rhnlAmCTqlKlwjQgJW1Oe84tIXKJom9cFsxO0mK6W4dvM3MH3aKdEadc2IneccRx4qRgGKiAoTbdTR2SVRswNhGGuOImWlpl8Yofx1LWEM5P4_xcpeMIoplfHVPhaq_JxhVGTT468hjdKauGmG96ZNBPRLvFUqUX4h_Rali4qwA08XIxiv9sJKiMbdkENtK_kWvupCBVWHqeUewpS0eiWeUDbomWq3GtAVj05xpBLqqlbiW_iwkNQMt3aO7hMWJDO5N3JSjjvtktEQeI-5g_pKev1wpQX9W1DCJAu9sezHQblfEw6NM9Gu6MoWMxCoe-KW3-8tzTX6-Ygu-qaso29fn66R2xXTPv8QqEvOpD86d9-A_MM813Dv3CEcfnJiOoqBw.xl84Q5mm7DfCZ5pqFzC4UQ/__results___files/__results___61_0.png\" alt=\"\"><br>\nThe image from <a href=\"https://www.kaggle.com/dschettler8845/visual-in-depth-eda-vinbigdata-competition-data\" target=\"_blank\">Visual In-Depth EDA – VinBigData Competition Data</a> by <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a></p>\n<p>So it is important to consider how to predict such a small bounding boxes.<br>\nWe can specify \"anchor\" in the detection model (FasterRCNN is 2-step, while YOLO is 1-step model).</p>\n<p>If you used smaller-resized images, bounding box size also resized to much smaller value. So I think it is important to use larger size of the image in this competition.</p>\n<h2>High aspect ratio bounding box</h2>\n<p>Also high aspect ratio bounding box is included in the data, as you can also check in <a href=\"https://www.kaggle.com/dschettler8845/visual-in-depth-eda-vinbigdata-competition-data\" target=\"_blank\">Visual In-Depth EDA – VinBigData Competition Data</a> by <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>.<br>\nIt makes the problem difficult.</p>\n<p>Usual detection model sets the anchor aspect ratio range from 0.5 to 2.0.<br>\nHowever aspect ratio of 0.33 or 3.0 is included in the data (vertically/horizontally very long image) more than 10%.</p>\n<p>If you used square-resized images, original horizontally long bounding box becomes much more horizontally long when the original image height is longer than width.<br>\nSo it is important to predict horizontally long bounding box or you should use the preprocessed image with keeping aspect ratio.</p>\n<h1>Experiment</h1>\n<p>I have changed to use more smaller anchor size 16pixel, and used anchor aspect ratio of 0.33 &amp; 3.0.<br>\nThese are the results:</p>\n<table>\n<thead>\n<tr>\n<th>Exp setting</th>\n<th>Local validation score ap40</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Original anchor config</td>\n<td>28.18</td>\n</tr>\n<tr>\n<td>Additionally use small anchor size</td>\n<td>29.35</td>\n</tr>\n<tr>\n<td>Additionally use high aspect ratio</td>\n<td>28.92</td>\n</tr>\n<tr>\n<td>Additionally use small anchor size + high aspect ratio</td>\n<td>34.11</td>\n</tr>\n</tbody>\n</table>\n<p><a href=\"https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024\" target=\"_blank\">1024x1024 pixel size image</a> is used in the experiment.</p>\n<p>We can see that using <strong>both small anchor size + high aspect ratio</strong> is important!</p>\n<h1>Code</h1>\n<p>Above experiment is done based on the kernel <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train#VinBigData-detectron2-train\" target=\"_blank\">📸VinBigData detectron2 train</a>.<br>\nYou can change anchor configuration in the following way in <code>detectron2</code>.</p>\n<h2>anchor size</h2>\n<pre><code>cfg.MODEL.ANCHOR_GENERATOR.SIZES = [[16], [32], [64], [128], [256], [512], [1024]]\ncfg.MODEL.RPN.IN_FEATURES = ['p2', 'p2', 'p3', 'p4', 'p5', 'p6', 'p6']\n</code></pre>\n<p>Original configuration is <code>[[32], [64], [128], [256], [512]]</code> and <code>['p2', 'p3', 'p4', 'p5', 'p6']</code>.<br>\nYou need to change region proposal network features at the same time when you change anchor generator sizes.</p>\n<h2>anchor aspect ratio</h2>\n<pre><code>cfg.MODEL.ANCHOR_GENERATOR.ASPECT_RATIOS = [[0.33, 0.5, 1.0, 2.0, 3.0]]\n</code></pre>\n<p>Original anchor aspect ratios were <code>[[0.5, 1.0, 2.0]]</code>.</p>\n<h1>Other approaches</h1>\n<p>Using bigger image resolution helps predicting the small bounding boxes.<br>\nIt's worth to try training with original image size with keeping aspect ratio.</p>",
  "messages": [
    {
      "id": 1207537,
      "postDate": "2021-02-17T22:54:01.047Z",
      "content": "<h1>EDA</h1>\n<h2>Small bounding box</h2>\n<p>Now I understand that the challenge of this competition is to predict wide range size of the bounding boxes.<br>\nEspecially for <strong>small sized bounding box</strong>.</p>\n<p>When I check the AP by class, \"Calcification\", \"Nodule/Mass\" and \"Pneumothorax\" were the difficult categories.</p>\n<p><img src=\"https://www.kaggleusercontent.com/kf/54411746/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..aLpE9X3UbjaJB2RiJORRpA.2jADfQTTkBT-Ad4s8Sus-JUCGl1jQT2clRnOTHU3FXUHvPwj7zvPvJG7SojA8Wc-ZFVhzwTmwLctRZf5XYcsnqy3kPrDO0w4SQdg8MvxcTKjrO2H0KLhQyZNTq1iJlFQDmFfYcaVN2pdFztCw4hLqPjhJcShBn-uayrUhEX25SyiFhDMK10TfDxsjM9HCXg1cnFdy2suxdzoxnGFp_3fJZkSi5XYXgonDqjfeteiDpQjc-mcT1x5h7Y6LoiyCi27RK5SPImPLiYpc1P4NmzX0VOA-jJrLTmTEbJAfidCDe505Qufx0BDyIBq75tvCLFXcC9ng_jVdn79W1mAW3X6-GITNM47UR4sykWik4TeWnBTlDrr5dgjK-GLlocupJIjoyfl-UDxhBriy276K5PCrqEK93DIv7ZpRcnmclBZ53mJtAIkUXfs4SbbCgusoqZeNCfTcDugkAnQjsGcdoJEjjYY3f94MkThxJx_F2Ro1kTWFdhqyswHJkG8OCrRNYOPeZu7BZg1RNj8_rBTCDpotXJezhxs5NhwVfu_kLpWcJ9Esc565K9ND9KrpAIamIqnzOhGgWj7yYW1IDCLAfo0cgTl6q58EmA6xRcNWJI3D9Jk1RuCe9zyI2RvRXpG0QqIVto14ycw1kK-qqPNN-EGG92PwGTWXNTfyb7HjrUIn4E.2Wp5qo9vOwE1SvZyzHsN-A/__results___files/__results___45_0.png\" alt=\"\"><br>\nFrom <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train#VinBigData-detectron2-train\" target=\"_blank\">📸VinBigData detectron2 train</a>.</p>\n<p>And I see that bounding box is really small in these categories.<br>\nFor example, this is the \"Nodule/Mass\" bounding boxes:<br>\n<img src=\"https://www.kaggleusercontent.com/kf/53546324/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..3JgopvaHUEk2iyxwddnyhA.IP2k4_iRRQEpZjcnkwq9z4WnWtdYRgMbeVw3pVMQzBcsfZrYo_mB48GxE6qxZQzi0iRnx6mLdMOax0z8YEkG41G_UXIPcIB_LTOFt4TeDVltsDAcE0WfSn2PlqsUWpqG7xhGHkU72mJGaJh3VBjQSIRwHQhbasiygWfF7VJ58ecDQ1h5Wtdsi-G_PtOPRSxp2o25R8xJvmeUdbGG8njC3i2micR7cDAw7Pdok6lA0U-Q7l3j2BI0ZgYBcI99jiHs3wbd13P-72WI8rCkUOkakb0yTxHljFASyT0ZruiipkAT6sova0YmMMDDR6l1Ov0zsGeZogz8qZNXEoKAVF9rhnlAmCTqlKlwjQgJW1Oe84tIXKJom9cFsxO0mK6W4dvM3MH3aKdEadc2IneccRx4qRgGKiAoTbdTR2SVRswNhGGuOImWlpl8Yofx1LWEM5P4_xcpeMIoplfHVPhaq_JxhVGTT468hjdKauGmG96ZNBPRLvFUqUX4h_Rali4qwA08XIxiv9sJKiMbdkENtK_kWvupCBVWHqeUewpS0eiWeUDbomWq3GtAVj05xpBLqqlbiW_iwkNQMt3aO7hMWJDO5N3JSjjvtktEQeI-5g_pKev1wpQX9W1DCJAu9sezHQblfEw6NM9Gu6MoWMxCoe-KW3-8tzTX6-Ygu-qaso29fn66R2xXTPv8QqEvOpD86d9-A_MM813Dv3CEcfnJiOoqBw.xl84Q5mm7DfCZ5pqFzC4UQ/__results___files/__results___61_0.png\" alt=\"\"><br>\nThe image from <a href=\"https://www.kaggle.com/dschettler8845/visual-in-depth-eda-vinbigdata-competition-data\" target=\"_blank\">Visual In-Depth EDA – VinBigData Competition Data</a> by <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a></p>\n<p>So it is important to consider how to predict such a small bounding boxes.<br>\nWe can specify \"anchor\" in the detection model (FasterRCNN is 2-step, while YOLO is 1-step model).</p>\n<p>If you used smaller-resized images, bounding box size also resized to much smaller value. So I think it is important to use larger size of the image in this competition.</p>\n<h2>High aspect ratio bounding box</h2>\n<p>Also high aspect ratio bounding box is included in the data, as you can also check in <a href=\"https://www.kaggle.com/dschettler8845/visual-in-depth-eda-vinbigdata-competition-data\" target=\"_blank\">Visual In-Depth EDA – VinBigData Competition Data</a> by <a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a>.<br>\nIt makes the problem difficult.</p>\n<p>Usual detection model sets the anchor aspect ratio range from 0.5 to 2.0.<br>\nHowever aspect ratio of 0.33 or 3.0 is included in the data (vertically/horizontally very long image) more than 10%.</p>\n<p>If you used square-resized images, original horizontally long bounding box becomes much more horizontally long when the original image height is longer than width.<br>\nSo it is important to predict horizontally long bounding box or you should use the preprocessed image with keeping aspect ratio.</p>\n<h1>Experiment</h1>\n<p>I have changed to use more smaller anchor size 16pixel, and used anchor aspect ratio of 0.33 &amp; 3.0.<br>\nThese are the results:</p>\n<table>\n<thead>\n<tr>\n<th>Exp setting</th>\n<th>Local validation score ap40</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Original anchor config</td>\n<td>28.18</td>\n</tr>\n<tr>\n<td>Additionally use small anchor size</td>\n<td>29.35</td>\n</tr>\n<tr>\n<td>Additionally use high aspect ratio</td>\n<td>28.92</td>\n</tr>\n<tr>\n<td>Additionally use small anchor size + high aspect ratio</td>\n<td>34.11</td>\n</tr>\n</tbody>\n</table>\n<p><a href=\"https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024\" target=\"_blank\">1024x1024 pixel size image</a> is used in the experiment.</p>\n<p>We can see that using <strong>both small anchor size + high aspect ratio</strong> is important!</p>\n<h1>Code</h1>\n<p>Above experiment is done based on the kernel <a href=\"https://www.kaggle.com/corochann/vinbigdata-detectron2-train#VinBigData-detectron2-train\" target=\"_blank\">📸VinBigData detectron2 train</a>.<br>\nYou can change anchor configuration in the following way in <code>detectron2</code>.</p>\n<h2>anchor size</h2>\n<pre><code>cfg.MODEL.ANCHOR_GENERATOR.SIZES = [[16], [32], [64], [128], [256], [512], [1024]]\ncfg.MODEL.RPN.IN_FEATURES = ['p2', 'p2', 'p3', 'p4', 'p5', 'p6', 'p6']\n</code></pre>\n<p>Original configuration is <code>[[32], [64], [128], [256], [512]]</code> and <code>['p2', 'p3', 'p4', 'p5', 'p6']</code>.<br>\nYou need to change region proposal network features at the same time when you change anchor generator sizes.</p>\n<h2>anchor aspect ratio</h2>\n<pre><code>cfg.MODEL.ANCHOR_GENERATOR.ASPECT_RATIOS = [[0.33, 0.5, 1.0, 2.0, 3.0]]\n</code></pre>\n<p>Original anchor aspect ratios were <code>[[0.5, 1.0, 2.0]]</code>.</p>\n<h1>Other approaches</h1>\n<p>Using bigger image resolution helps predicting the small bounding boxes.<br>\nIt's worth to try training with original image size with keeping aspect ratio.</p>",
      "rawMarkdown": "# EDA\n\n## Small bounding box\nNow I understand that the challenge of this competition is to predict wide range size of the bounding boxes.\nEspecially for **small sized bounding box**.\n\nWhen I check the AP by class, \"Calcification\", \"Nodule/Mass\" and \"Pneumothorax\" were the difficult categories.\n\n![](https://www.kaggleusercontent.com/kf/54411746/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..aLpE9X3UbjaJB2RiJORRpA.2jADfQTTkBT-Ad4s8Sus-JUCGl1jQT2clRnOTHU3FXUHvPwj7zvPvJG7SojA8Wc-ZFVhzwTmwLctRZf5XYcsnqy3kPrDO0w4SQdg8MvxcTKjrO2H0KLhQyZNTq1iJlFQDmFfYcaVN2pdFztCw4hLqPjhJcShBn-uayrUhEX25SyiFhDMK10TfDxsjM9HCXg1cnFdy2suxdzoxnGFp_3fJZkSi5XYXgonDqjfeteiDpQjc-mcT1x5h7Y6LoiyCi27RK5SPImPLiYpc1P4NmzX0VOA-jJrLTmTEbJAfidCDe505Qufx0BDyIBq75tvCLFXcC9ng_jVdn79W1mAW3X6-GITNM47UR4sykWik4TeWnBTlDrr5dgjK-GLlocupJIjoyfl-UDxhBriy276K5PCrqEK93DIv7ZpRcnmclBZ53mJtAIkUXfs4SbbCgusoqZeNCfTcDugkAnQjsGcdoJEjjYY3f94MkThxJx_F2Ro1kTWFdhqyswHJkG8OCrRNYOPeZu7BZg1RNj8_rBTCDpotXJezhxs5NhwVfu_kLpWcJ9Esc565K9ND9KrpAIamIqnzOhGgWj7yYW1IDCLAfo0cgTl6q58EmA6xRcNWJI3D9Jk1RuCe9zyI2RvRXpG0QqIVto14ycw1kK-qqPNN-EGG92PwGTWXNTfyb7HjrUIn4E.2Wp5qo9vOwE1SvZyzHsN-A/__results___files/__results___45_0.png)\nFrom [📸VinBigData detectron2 train](https://www.kaggle.com/corochann/vinbigdata-detectron2-train#VinBigData-detectron2-train).\n\nAnd I see that bounding box is really small in these categories.\nFor example, this is the \"Nodule/Mass\" bounding boxes:\n![](https://www.kaggleusercontent.com/kf/53546324/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..3JgopvaHUEk2iyxwddnyhA.IP2k4_iRRQEpZjcnkwq9z4WnWtdYRgMbeVw3pVMQzBcsfZrYo_mB48GxE6qxZQzi0iRnx6mLdMOax0z8YEkG41G_UXIPcIB_LTOFt4TeDVltsDAcE0WfSn2PlqsUWpqG7xhGHkU72mJGaJh3VBjQSIRwHQhbasiygWfF7VJ58ecDQ1h5Wtdsi-G_PtOPRSxp2o25R8xJvmeUdbGG8njC3i2micR7cDAw7Pdok6lA0U-Q7l3j2BI0ZgYBcI99jiHs3wbd13P-72WI8rCkUOkakb0yTxHljFASyT0ZruiipkAT6sova0YmMMDDR6l1Ov0zsGeZogz8qZNXEoKAVF9rhnlAmCTqlKlwjQgJW1Oe84tIXKJom9cFsxO0mK6W4dvM3MH3aKdEadc2IneccRx4qRgGKiAoTbdTR2SVRswNhGGuOImWlpl8Yofx1LWEM5P4_xcpeMIoplfHVPhaq_JxhVGTT468hjdKauGmG96ZNBPRLvFUqUX4h_Rali4qwA08XIxiv9sJKiMbdkENtK_kWvupCBVWHqeUewpS0eiWeUDbomWq3GtAVj05xpBLqqlbiW_iwkNQMt3aO7hMWJDO5N3JSjjvtktEQeI-5g_pKev1wpQX9W1DCJAu9sezHQblfEw6NM9Gu6MoWMxCoe-KW3-8tzTX6-Ygu-qaso29fn66R2xXTPv8QqEvOpD86d9-A_MM813Dv3CEcfnJiOoqBw.xl84Q5mm7DfCZ5pqFzC4UQ/__results___files/__results___61_0.png)\nThe image from [Visual In-Depth EDA – VinBigData Competition Data](https://www.kaggle.com/dschettler8845/visual-in-depth-eda-vinbigdata-competition-data) by @dschettler8845\n\nSo it is important to consider how to predict such a small bounding boxes.\nWe can specify \"anchor\" in the detection model (FasterRCNN is 2-step, while YOLO is 1-step model).\n\nIf you used smaller-resized images, bounding box size also resized to much smaller value. So I think it is important to use larger size of the image in this competition.\n\n## High aspect ratio bounding box\n\nAlso high aspect ratio bounding box is included in the data, as you can also check in [Visual In-Depth EDA – VinBigData Competition Data](https://www.kaggle.com/dschettler8845/visual-in-depth-eda-vinbigdata-competition-data) by @dschettler8845.\nIt makes the problem difficult.\n\nUsual detection model sets the anchor aspect ratio range from 0.5 to 2.0.\nHowever aspect ratio of 0.33 or 3.0 is included in the data (vertically/horizontally very long image) more than 10%.\n\nIf you used square-resized images, original horizontally long bounding box becomes much more horizontally long when the original image height is longer than width.\nSo it is important to predict horizontally long bounding box or you should use the preprocessed image with keeping aspect ratio.\n\n# Experiment\nI have changed to use more smaller anchor size 16pixel, and used anchor aspect ratio of 0.33 & 3.0.\nThese are the results:\n\n| Exp setting | Local validation score ap40 |\n| --- | --- |\n| Original anchor config | 28.18 |\n| Additionally use small anchor size | 29.35 |\n| Additionally use high aspect ratio | 28.92 |\n| Additionally use small anchor size + high aspect ratio | 34.11 |\n\n[1024x1024 pixel size image](https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024) is used in the experiment.\n\nWe can see that using **both small anchor size + high aspect ratio** is important!\n\n# Code\n\nAbove experiment is done based on the kernel [📸VinBigData detectron2 train](https://www.kaggle.com/corochann/vinbigdata-detectron2-train#VinBigData-detectron2-train).\nYou can change anchor configuration in the following way in `detectron2`.\n\n## anchor size\n```\ncfg.MODEL.ANCHOR_GENERATOR.SIZES = [[16], [32], [64], [128], [256], [512], [1024]]\ncfg.MODEL.RPN.IN_FEATURES = ['p2', 'p2', 'p3', 'p4', 'p5', 'p6', 'p6']\n```\n\nOriginal configuration is `[[32], [64], [128], [256], [512]]` and `['p2', 'p3', 'p4', 'p5', 'p6']`.\nYou need to change region proposal network features at the same time when you change anchor generator sizes.\n\n## anchor aspect ratio\n```\ncfg.MODEL.ANCHOR_GENERATOR.ASPECT_RATIOS = [[0.33, 0.5, 1.0, 2.0, 3.0]]\n```\n\nOriginal anchor aspect ratios were `[[0.5, 1.0, 2.0]]`.\n\n# Other approaches\n\nUsing bigger image resolution helps predicting the small bounding boxes.\nIt's worth to try training with original image size with keeping aspect ratio.",
      "votes": 30
    },
    {
      "id": 1207634,
      "postDate": "2021-02-18T00:29:35.940Z",
      "content": "<p>The anchors used in yolov5:</p>\n<ul>\n<li>[10,13, 16,30, 33,23]  # P3/8</li>\n<li>[30,61, 62,45, 59,119]  # P4/16</li>\n<li>[116,90, 156,198, 373,326]  # P5/32</li>\n</ul>\n<p>If they don't fit well to the data, the model automatically calculates new ones.</p>",
      "rawMarkdown": "The anchors used in yolov5:\n  - [10,13, 16,30, 33,23]  # P3/8\n  - [30,61, 62,45, 59,119]  # P4/16\n  - [116,90, 156,198, 373,326]  # P5/32\n\nIf they don't fit well to the data, the model automatically calculates new ones.",
      "votes": 3,
      "replies": [
        {
          "id": 1207690,
          "postDate": "2021-02-18T01:15:26.197Z",
          "content": "<p>Thank you for useful comment! It seems yolov5 is more strong to detect small bbox in the default setting.</p>\n<blockquote>\n  <p>If they don't fit well to the data, the model automatically calculates new ones.</p>\n</blockquote>\n<p>Wow that's great! Maybe that's why many people uses this model and performs well!?</p>",
          "rawMarkdown": "Thank you for useful comment! It seems yolov5 is more strong to detect small bbox in the default setting.\n\n> If they don't fit well to the data, the model automatically calculates new ones.\n\nWow that's great! Maybe that's why many people uses this model and performs well!?",
          "votes": 2
        },
        {
          "id": 1208551,
          "postDate": "2021-02-18T10:36:03.837Z",
          "content": "<p>Yes, I think that can be a reason, although I don't know how often this is necessary, for this competition the predefined anchors work very well. Besides that integrated augmentation at train and test time and the option to ensemble different models are maybe other reasons. </p>",
          "rawMarkdown": "Yes, I think that can be a reason, although I don't know how often this is necessary, for this competition the predefined anchors work very well. Besides that integrated augmentation at train and test time and the option to ensemble different models are maybe other reasons. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1227068,
      "postDate": "2021-03-05T06:57:50.197Z",
      "content": "<p>For your experiment, did you just add the three lines above?  I've tried these configurations for multiple models and actually somehow in every case it greatly decreased CV and LB…wondering if I might have done something wrong or forgotten to do something here.</p>",
      "rawMarkdown": "For your experiment, did you just add the three lines above?  I've tried these configurations for multiple models and actually somehow in every case it greatly decreased CV and LB...wondering if I might have done something wrong or forgotten to do something here.",
      "replies": [
        {
          "id": 1227267,
          "postDate": "2021-03-05T11:18:19.500Z",
          "content": "<p>Yes, difference between \"Original anchor config\" and \"Additionally use small anchor size + high aspect ratio\" is only three lines.</p>\n<p>The baseline \"Original anchor config\" maybe different from uploaded latest kernel, I used 1024x1024 images and I might use 1-step prediction.</p>",
          "rawMarkdown": "Yes, difference between \"Original anchor config\" and \"Additionally use small anchor size + high aspect ratio\" is only three lines.\n\nThe baseline \"Original anchor config\" maybe different from uploaded latest kernel, I used 1024x1024 images and I might use 1-step prediction.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1207561,
      "postDate": "2021-02-17T23:19:53.973Z",
      "content": "<p>I uploaded original sized image too!</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/corochann/vinbigdata-chest-xray-original-png\" target=\"_blank\">vinbigdata-chest-xray-original-png</a> (<a href=\"https://www.kaggle.com/corochann/preprocessing-image-original-size-lossless-png\" target=\"_blank\">notebook</a> on kaggle fails due to disk limit)</li>\n</ul>",
      "rawMarkdown": "I uploaded original sized image too!\n - [vinbigdata-chest-xray-original-png](https://www.kaggle.com/corochann/vinbigdata-chest-xray-original-png) ([notebook](https://www.kaggle.com/corochann/preprocessing-image-original-size-lossless-png) on kaggle fails due to disk limit)"
    }
  ],
  "comments": [
    {
      "id": 1207634,
      "author_name": "Hannes Öhler",
      "author_url": "",
      "post_date": "2021-02-18T00:29:35.940000",
      "content": "<p>The anchors used in yolov5:</p>\n<ul>\n<li>[10,13, 16,30, 33,23]  # P3/8</li>\n<li>[30,61, 62,45, 59,119]  # P4/16</li>\n<li>[116,90, 156,198, 373,326]  # P5/32</li>\n</ul>\n<p>If they don't fit well to the data, the model automatically calculates new ones.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1207690,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2021-02-18T01:15:26.197000",
          "content": "<p>Thank you for useful comment! It seems yolov5 is more strong to detect small bbox in the default setting.</p>\n<blockquote>\n  <p>If they don't fit well to the data, the model automatically calculates new ones.</p>\n</blockquote>\n<p>Wow that's great! Maybe that's why many people uses this model and performs well!?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1208551,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-02-18T10:36:03.837000",
          "content": "<p>Yes, I think that can be a reason, although I don't know how often this is necessary, for this competition the predefined anchors work very well. Besides that integrated augmentation at train and test time and the option to ensemble different models are maybe other reasons. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1227068,
      "author_name": "Daniel Shan",
      "author_url": "",
      "post_date": "2021-03-05T06:57:50.197000",
      "content": "<p>For your experiment, did you just add the three lines above?  I've tried these configurations for multiple models and actually somehow in every case it greatly decreased CV and LB…wondering if I might have done something wrong or forgotten to do something here.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1227267,
          "author_name": "corochann",
          "author_url": "",
          "post_date": "2021-03-05T11:18:19.500000",
          "content": "<p>Yes, difference between \"Original anchor config\" and \"Additionally use small anchor size + high aspect ratio\" is only three lines.</p>\n<p>The baseline \"Original anchor config\" maybe different from uploaded latest kernel, I used 1024x1024 images and I might use 1-step prediction.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1207561,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2021-02-17T23:19:53.973000",
      "content": "<p>I uploaded original sized image too!</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/corochann/vinbigdata-chest-xray-original-png\" target=\"_blank\">vinbigdata-chest-xray-original-png</a> (<a href=\"https://www.kaggle.com/corochann/preprocessing-image-original-size-lossless-png\" target=\"_blank\">notebook</a> on kaggle fails due to disk limit)</li>\n</ul>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207537": "# EDA\n\n## Small bounding box\nNow I understand that the challenge of this competition is to predict wide range size of the bounding boxes.\nEspecially for **small sized bounding box**.\n\nWhen I check the AP by class, \"Calcification\", \"Nodule/Mass\" and \"Pneumothorax\" were the difficult categories.\n\n![](https://www.kaggleusercontent.com/kf/54411746/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..aLpE9X3UbjaJB2RiJORRpA.2jADfQTTkBT-Ad4s8Sus-JUCGl1jQT2clRnOTHU3FXUHvPwj7zvPvJG7SojA8Wc-ZFVhzwTmwLctRZf5XYcsnqy3kPrDO0w4SQdg8MvxcTKjrO2H0KLhQyZNTq1iJlFQDmFfYcaVN2pdFztCw4hLqPjhJcShBn-uayrUhEX25SyiFhDMK10TfDxsjM9HCXg1cnFdy2suxdzoxnGFp_3fJZkSi5XYXgonDqjfeteiDpQjc-mcT1x5h7Y6LoiyCi27RK5SPImPLiYpc1P4NmzX0VOA-jJrLTmTEbJAfidCDe505Qufx0BDyIBq75tvCLFXcC9ng_jVdn79W1mAW3X6-GITNM47UR4sykWik4TeWnBTlDrr5dgjK-GLlocupJIjoyfl-UDxhBriy276K5PCrqEK93DIv7ZpRcnmclBZ53mJtAIkUXfs4SbbCgusoqZeNCfTcDugkAnQjsGcdoJEjjYY3f94MkThxJx_F2Ro1kTWFdhqyswHJkG8OCrRNYOPeZu7BZg1RNj8_rBTCDpotXJezhxs5NhwVfu_kLpWcJ9Esc565K9ND9KrpAIamIqnzOhGgWj7yYW1IDCLAfo0cgTl6q58EmA6xRcNWJI3D9Jk1RuCe9zyI2RvRXpG0QqIVto14ycw1kK-qqPNN-EGG92PwGTWXNTfyb7HjrUIn4E.2Wp5qo9vOwE1SvZyzHsN-A/__results___files/__results___45_0.png)\nFrom [📸VinBigData detectron2 train](https://www.kaggle.com/corochann/vinbigdata-detectron2-train#VinBigData-detectron2-train).\n\nAnd I see that bounding box is really small in these categories.\nFor example, this is the \"Nodule/Mass\" bounding boxes:\n![](https://www.kaggleusercontent.com/kf/53546324/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..3JgopvaHUEk2iyxwddnyhA.IP2k4_iRRQEpZjcnkwq9z4WnWtdYRgMbeVw3pVMQzBcsfZrYo_mB48GxE6qxZQzi0iRnx6mLdMOax0z8YEkG41G_UXIPcIB_LTOFt4TeDVltsDAcE0WfSn2PlqsUWpqG7xhGHkU72mJGaJh3VBjQSIRwHQhbasiygWfF7VJ58ecDQ1h5Wtdsi-G_PtOPRSxp2o25R8xJvmeUdbGG8njC3i2micR7cDAw7Pdok6lA0U-Q7l3j2BI0ZgYBcI99jiHs3wbd13P-72WI8rCkUOkakb0yTxHljFASyT0ZruiipkAT6sova0YmMMDDR6l1Ov0zsGeZogz8qZNXEoKAVF9rhnlAmCTqlKlwjQgJW1Oe84tIXKJom9cFsxO0mK6W4dvM3MH3aKdEadc2IneccRx4qRgGKiAoTbdTR2SVRswNhGGuOImWlpl8Yofx1LWEM5P4_xcpeMIoplfHVPhaq_JxhVGTT468hjdKauGmG96ZNBPRLvFUqUX4h_Rali4qwA08XIxiv9sJKiMbdkENtK_kWvupCBVWHqeUewpS0eiWeUDbomWq3GtAVj05xpBLqqlbiW_iwkNQMt3aO7hMWJDO5N3JSjjvtktEQeI-5g_pKev1wpQX9W1DCJAu9sezHQblfEw6NM9Gu6MoWMxCoe-KW3-8tzTX6-Ygu-qaso29fn66R2xXTPv8QqEvOpD86d9-A_MM813Dv3CEcfnJiOoqBw.xl84Q5mm7DfCZ5pqFzC4UQ/__results___files/__results___61_0.png)\nThe image from [Visual In-Depth EDA – VinBigData Competition Data](https://www.kaggle.com/dschettler8845/visual-in-depth-eda-vinbigdata-competition-data) by @dschettler8845\n\nSo it is important to consider how to predict such a small bounding boxes.\nWe can specify \"anchor\" in the detection model (FasterRCNN is 2-step, while YOLO is 1-step model).\n\nIf you used smaller-resized images, bounding box size also resized to much smaller value. So I think it is important to use larger size of the image in this competition.\n\n## High aspect ratio bounding box\n\nAlso high aspect ratio bounding box is included in the data, as you can also check in [Visual In-Depth EDA – VinBigData Competition Data](https://www.kaggle.com/dschettler8845/visual-in-depth-eda-vinbigdata-competition-data) by @dschettler8845.\nIt makes the problem difficult.\n\nUsual detection model sets the anchor aspect ratio range from 0.5 to 2.0.\nHowever aspect ratio of 0.33 or 3.0 is included in the data (vertically/horizontally very long image) more than 10%.\n\nIf you used square-resized images, original horizontally long bounding box becomes much more horizontally long when the original image height is longer than width.\nSo it is important to predict horizontally long bounding box or you should use the preprocessed image with keeping aspect ratio.\n\n# Experiment\nI have changed to use more smaller anchor size 16pixel, and used anchor aspect ratio of 0.33 & 3.0.\nThese are the results:\n\n| Exp setting | Local validation score ap40 |\n| --- | --- |\n| Original anchor config | 28.18 |\n| Additionally use small anchor size | 29.35 |\n| Additionally use high aspect ratio | 28.92 |\n| Additionally use small anchor size + high aspect ratio | 34.11 |\n\n[1024x1024 pixel size image](https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024) is used in the experiment.\n\nWe can see that using **both small anchor size + high aspect ratio** is important!\n\n# Code\n\nAbove experiment is done based on the kernel [📸VinBigData detectron2 train](https://www.kaggle.com/corochann/vinbigdata-detectron2-train#VinBigData-detectron2-train).\nYou can change anchor configuration in the following way in `detectron2`.\n\n## anchor size\n```\ncfg.MODEL.ANCHOR_GENERATOR.SIZES = [[16], [32], [64], [128], [256], [512], [1024]]\ncfg.MODEL.RPN.IN_FEATURES = ['p2', 'p2', 'p3', 'p4', 'p5', 'p6', 'p6']\n```\n\nOriginal configuration is `[[32], [64], [128], [256], [512]]` and `['p2', 'p3', 'p4', 'p5', 'p6']`.\nYou need to change region proposal network features at the same time when you change anchor generator sizes.\n\n## anchor aspect ratio\n```\ncfg.MODEL.ANCHOR_GENERATOR.ASPECT_RATIOS = [[0.33, 0.5, 1.0, 2.0, 3.0]]\n```\n\nOriginal anchor aspect ratios were `[[0.5, 1.0, 2.0]]`.\n\n# Other approaches\n\nUsing bigger image resolution helps predicting the small bounding boxes.\nIt's worth to try training with original image size with keeping aspect ratio.",
    "1207634": "The anchors used in yolov5:\n  - [10,13, 16,30, 33,23]  # P3/8\n  - [30,61, 62,45, 59,119]  # P4/16\n  - [116,90, 156,198, 373,326]  # P5/32\n\nIf they don't fit well to the data, the model automatically calculates new ones.",
    "1227068": "For your experiment, did you just add the three lines above?  I've tried these configurations for multiple models and actually somehow in every case it greatly decreased CV and LB...wondering if I might have done something wrong or forgotten to do something here.",
    "1207561": "I uploaded original sized image too!\n - [vinbigdata-chest-xray-original-png](https://www.kaggle.com/corochann/vinbigdata-chest-xray-original-png) ([notebook](https://www.kaggle.com/corochann/preprocessing-image-original-size-lossless-png) on kaggle fails due to disk limit)"
  }
}