{
  "id": 447450,
  "title": "10th Place Solution",
  "url": "/competitions/rsna-2023-abdominal-trauma-detection/discussion/447450",
  "author_name": "yu4u",
  "post_date": "2023-10-16T00:05:59.627000",
  "votes": 38,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congrats to all prize and medal winners! I enjoyed this competition because there are a great variety of options for solving this task as with last year's RSNA competition. We share our team's solution ( <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> + <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> ).</p>\n<h1>Summary</h1>\n<p>Our solution is to first segment the organs, cut out each organ region, and build a dedicated model for each organ.<br>\nFor the bowel and extravasation classes, for which large regions must be explored, we do not perform segmentation to cut out the regions, but instead perform simple black region removal and input the large regions into the models.<br>\nThe results of each model are refined by the stacking model and submitted as the final result.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Ff5b5b27bf27dbd11ae78efca228ff017%2F1.png?generation=1697414699260439&amp;alt=media\" alt=\"\"></p>\n<h1>Segmentation Model</h1>\n<p>We used the 3D SwinUNETR model provided by MONAI. It works surprisingly well with even a small amount of training data. To reduce the computational cost, the entire voxel was resized into 128x128x128 before segmentation.</p>\n<h1>Organ Models</h1>\n<p>For liver, spleen, kidney, and bowel cropped regions, 2.5D CNN + LSTM models are used.</p>\n<h2>Liver and Spleen Models</h2>\n<p>Cropped region is resized into 16x386x386, and fed into dedicated models.</p>\n<h2>Kidney Model</h2>\n<p>Left and right kidney are independently cropped, resized, and concatinated along with horizontal axis. This enables us horizontal flip augmentation and TTA. Concatenated region becomes 16x224x448.</p>\n<h2>Bowel Model</h2>\n<p>Cropped region is resized into 64x224x224. In training bowel model, both patient level label and image level labels were used.</p>\n<h1>Stacking Model</h1>\n<p>The purpose of the stacking model is to directly optimize the average of the weighted logloss, which is the metric for this competition, including the logloss for any_injury. Each model is optimized for the weighted logloss for each of the injury type, but not for any_injury because it is automatically calculated from the probabilities of the other injuries. For any_injury, it is essential to optimize any_injury because the weight for any_injury is relatively large.<br>\nAs a stacking model, we use a simple 4-layer MLP, and trained with competition metric as loss function.</p>\n<p>The table below shows the CV evaluation results before and after stacking (The values below are out of date because we wrote this solution before the competition was extended. The final submission's CV is 0.3686).</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>bowel</th>\n<th>extravasation</th>\n<th>kidney</th>\n<th>liver</th>\n<th>spleen</th>\n<th>any</th>\n<th>overall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>w/o stacking</td>\n<td>0.1186</td>\n<td>0.4716</td>\n<td>0.2948</td>\n<td>0.4282</td>\n<td>0.4301</td>\n<td>0.5528</td>\n<td>0.3827</td>\n</tr>\n<tr>\n<td>with stacking</td>\n<td>0.1034</td>\n<td>0.4885</td>\n<td>0.2861</td>\n<td>0.4409</td>\n<td>0.4452</td>\n<td>0.4777</td>\n<td>0.3736</td>\n</tr>\n</tbody>\n</table>\n<h1>tattaka's Part</h1>\n<p>I was in charge of bowel and extravasation classification. Our solution followed <a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">the 2-stage approach of the RSNA competition 3 years ago</a>.  <br>\nThe basic setup is as follows</p>\n<h2>Image-Level Modeling (1st stage)</h2>\n<p>The input for the 1st stage is a 3-channel image including adjacent frames.  <br>\n1epoch training was performed on all labeled images.  </p>\n<ul>\n<li>backbone: resnetrs50  </li>\n<li>head: <ul>\n<li>Separate the head by bowel and extravasation</li></ul></li>\n</ul>\n<pre><code>  nn.Sequential(\n      nn.Conv2d(num_features[-], , kernel_size=),\n      nn.AdaptiveAvgPool2d((, )),\n      Flatten(),\n      nn.Dropout(, inplace=),\n      nn.Linear(, ),\n  )\n</code></pre>\n<p>In the 1st stage, students learn bowel and extravasation at the same time.  </p>\n<h2>Series-Level Modeling (2nd stage)</h2>\n<p>The input for the 2nd stage also followed the previous solution.  <br>\nUse the 512 dimensions after Flatten in the head of the 1st stage as image features.<br>\nImage features are created with stride=3 instead of all images, and the input sequence length is set to a maximum of 256 in the same way as the previous solution.   <br>\nThe differences between adjacent features are combined, and the input to the model is in the form (bs, 256, 1536). </p>\n<ul>\n<li>model:<ul>\n<li>Combines the attention pooling and max pooling of the BiGRU outputs in one layer to create an series-level prediction.</li>\n<li>The BiGRU output should also predict the image label. </li></ul></li>\n</ul>\n<p>Unlike the 1st stage, bowel and extravasation were optimized separately.</p>\n<h2>Tricks for Successful Training</h2>\n<p>There are a few tricks to successful learning in this competition.</p>\n<ul>\n<li>Because of large data imbalance, upsampling of the positive sample by a factor of 10 (high impact)<ul>\n<li>In addition, focal loss is used</li></ul></li>\n<li>After training stage1 at image level, max of the logit and gt are trained again as labels for the new image level (high impact)<ul>\n<li>repeated it twice</li>\n<li>Perhaps the image label is noisy</li></ul></li>\n<li>Rule-based removal of the outer black areas before the image is entered into the model<ul>\n<li>Because of the longer computation time when using a larger resolution, a size of 384x384 is used after removing the outer area.</li></ul></li>\n</ul>\n<pre><code> () -&gt; np.ndarray:\n    image_1ch = (img.mean() * ).astype(np.uint8)\n    kernel = np.ones((, ), np.uint8)\n    image_1ch = cv2.erode(image_1ch, kernel, iterations=)\n    mask = image_1ch &gt; \n     mask.() == :\n         img\n    rows = np.(mask, axis=)\n    cols = np.(mask, axis=)\n    y_min, y_max = np.where(rows)[][[, -]]\n    x_min, x_max = np.where(cols)[][[, -]]\n     (y_max - y_min) &gt;   (x_max - x_min) &gt; :\n        img = img[y_min:y_max, x_min:x_max]\n     img\n</code></pre>\n<p>The bowel and extravasation scores before stacking are</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>bowel logloss</th>\n<th>bowel auc</th>\n<th>ev logloss</th>\n<th>ev auc</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>stage1</td>\n<td>0.2719</td>\n<td>0.9314</td>\n<td>0.4602</td>\n<td>0.8287</td>\n</tr>\n<tr>\n<td>stage2</td>\n<td>0.1167</td>\n<td>0.9167</td>\n<td>0.4579</td>\n<td>0.8264</td>\n</tr>\n</tbody>\n</table>\n<h2>Not works</h2>\n<ul>\n<li>label smoothing</li>\n<li>resnet3dcsn<ul>\n<li>Not bad, but I didn't have time to tune in.</li></ul></li>\n<li>scaling logit</li>\n<li>GeM pooling</li>\n<li>Other backbone<ul>\n<li>ConvNeXt is slightly worse than resnetrs50</li>\n<li>I could not get the backbone of the transformer to work </li></ul></li>\n</ul>",
  "messages": [
    {
      "id": 2483722,
      "postDate": "2023-10-16T00:05:59.627Z",
      "content": "<p>Congrats to all prize and medal winners! I enjoyed this competition because there are a great variety of options for solving this task as with last year's RSNA competition. We share our team's solution ( <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a> + <a href=\"https://www.kaggle.com/tattaka\" target=\"_blank\">@tattaka</a> ).</p>\n<h1>Summary</h1>\n<p>Our solution is to first segment the organs, cut out each organ region, and build a dedicated model for each organ.<br>\nFor the bowel and extravasation classes, for which large regions must be explored, we do not perform segmentation to cut out the regions, but instead perform simple black region removal and input the large regions into the models.<br>\nThe results of each model are refined by the stacking model and submitted as the final result.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Ff5b5b27bf27dbd11ae78efca228ff017%2F1.png?generation=1697414699260439&amp;alt=media\" alt=\"\"></p>\n<h1>Segmentation Model</h1>\n<p>We used the 3D SwinUNETR model provided by MONAI. It works surprisingly well with even a small amount of training data. To reduce the computational cost, the entire voxel was resized into 128x128x128 before segmentation.</p>\n<h1>Organ Models</h1>\n<p>For liver, spleen, kidney, and bowel cropped regions, 2.5D CNN + LSTM models are used.</p>\n<h2>Liver and Spleen Models</h2>\n<p>Cropped region is resized into 16x386x386, and fed into dedicated models.</p>\n<h2>Kidney Model</h2>\n<p>Left and right kidney are independently cropped, resized, and concatinated along with horizontal axis. This enables us horizontal flip augmentation and TTA. Concatenated region becomes 16x224x448.</p>\n<h2>Bowel Model</h2>\n<p>Cropped region is resized into 64x224x224. In training bowel model, both patient level label and image level labels were used.</p>\n<h1>Stacking Model</h1>\n<p>The purpose of the stacking model is to directly optimize the average of the weighted logloss, which is the metric for this competition, including the logloss for any_injury. Each model is optimized for the weighted logloss for each of the injury type, but not for any_injury because it is automatically calculated from the probabilities of the other injuries. For any_injury, it is essential to optimize any_injury because the weight for any_injury is relatively large.<br>\nAs a stacking model, we use a simple 4-layer MLP, and trained with competition metric as loss function.</p>\n<p>The table below shows the CV evaluation results before and after stacking (The values below are out of date because we wrote this solution before the competition was extended. The final submission's CV is 0.3686).</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>bowel</th>\n<th>extravasation</th>\n<th>kidney</th>\n<th>liver</th>\n<th>spleen</th>\n<th>any</th>\n<th>overall</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>w/o stacking</td>\n<td>0.1186</td>\n<td>0.4716</td>\n<td>0.2948</td>\n<td>0.4282</td>\n<td>0.4301</td>\n<td>0.5528</td>\n<td>0.3827</td>\n</tr>\n<tr>\n<td>with stacking</td>\n<td>0.1034</td>\n<td>0.4885</td>\n<td>0.2861</td>\n<td>0.4409</td>\n<td>0.4452</td>\n<td>0.4777</td>\n<td>0.3736</td>\n</tr>\n</tbody>\n</table>\n<h1>tattaka's Part</h1>\n<p>I was in charge of bowel and extravasation classification. Our solution followed <a href=\"https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145\" target=\"_blank\">the 2-stage approach of the RSNA competition 3 years ago</a>.  <br>\nThe basic setup is as follows</p>\n<h2>Image-Level Modeling (1st stage)</h2>\n<p>The input for the 1st stage is a 3-channel image including adjacent frames.  <br>\n1epoch training was performed on all labeled images.  </p>\n<ul>\n<li>backbone: resnetrs50  </li>\n<li>head: <ul>\n<li>Separate the head by bowel and extravasation</li></ul></li>\n</ul>\n<pre><code>  nn.Sequential(\n      nn.Conv2d(num_features[-], , kernel_size=),\n      nn.AdaptiveAvgPool2d((, )),\n      Flatten(),\n      nn.Dropout(, inplace=),\n      nn.Linear(, ),\n  )\n</code></pre>\n<p>In the 1st stage, students learn bowel and extravasation at the same time.  </p>\n<h2>Series-Level Modeling (2nd stage)</h2>\n<p>The input for the 2nd stage also followed the previous solution.  <br>\nUse the 512 dimensions after Flatten in the head of the 1st stage as image features.<br>\nImage features are created with stride=3 instead of all images, and the input sequence length is set to a maximum of 256 in the same way as the previous solution.   <br>\nThe differences between adjacent features are combined, and the input to the model is in the form (bs, 256, 1536). </p>\n<ul>\n<li>model:<ul>\n<li>Combines the attention pooling and max pooling of the BiGRU outputs in one layer to create an series-level prediction.</li>\n<li>The BiGRU output should also predict the image label. </li></ul></li>\n</ul>\n<p>Unlike the 1st stage, bowel and extravasation were optimized separately.</p>\n<h2>Tricks for Successful Training</h2>\n<p>There are a few tricks to successful learning in this competition.</p>\n<ul>\n<li>Because of large data imbalance, upsampling of the positive sample by a factor of 10 (high impact)<ul>\n<li>In addition, focal loss is used</li></ul></li>\n<li>After training stage1 at image level, max of the logit and gt are trained again as labels for the new image level (high impact)<ul>\n<li>repeated it twice</li>\n<li>Perhaps the image label is noisy</li></ul></li>\n<li>Rule-based removal of the outer black areas before the image is entered into the model<ul>\n<li>Because of the longer computation time when using a larger resolution, a size of 384x384 is used after removing the outer area.</li></ul></li>\n</ul>\n<pre><code> () -&gt; np.ndarray:\n    image_1ch = (img.mean() * ).astype(np.uint8)\n    kernel = np.ones((, ), np.uint8)\n    image_1ch = cv2.erode(image_1ch, kernel, iterations=)\n    mask = image_1ch &gt; \n     mask.() == :\n         img\n    rows = np.(mask, axis=)\n    cols = np.(mask, axis=)\n    y_min, y_max = np.where(rows)[][[, -]]\n    x_min, x_max = np.where(cols)[][[, -]]\n     (y_max - y_min) &gt;   (x_max - x_min) &gt; :\n        img = img[y_min:y_max, x_min:x_max]\n     img\n</code></pre>\n<p>The bowel and extravasation scores before stacking are</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>bowel logloss</th>\n<th>bowel auc</th>\n<th>ev logloss</th>\n<th>ev auc</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>stage1</td>\n<td>0.2719</td>\n<td>0.9314</td>\n<td>0.4602</td>\n<td>0.8287</td>\n</tr>\n<tr>\n<td>stage2</td>\n<td>0.1167</td>\n<td>0.9167</td>\n<td>0.4579</td>\n<td>0.8264</td>\n</tr>\n</tbody>\n</table>\n<h2>Not works</h2>\n<ul>\n<li>label smoothing</li>\n<li>resnet3dcsn<ul>\n<li>Not bad, but I didn't have time to tune in.</li></ul></li>\n<li>scaling logit</li>\n<li>GeM pooling</li>\n<li>Other backbone<ul>\n<li>ConvNeXt is slightly worse than resnetrs50</li>\n<li>I could not get the backbone of the transformer to work </li></ul></li>\n</ul>",
      "rawMarkdown": "Congrats to all prize and medal winners! I enjoyed this competition because there are a great variety of options for solving this task as with last year's RSNA competition. We share our team's solution ( @ren4yu + @tattaka ).\n\n# Summary\nOur solution is to first segment the organs, cut out each organ region, and build a dedicated model for each organ.\nFor the bowel and extravasation classes, for which large regions must be explored, we do not perform segmentation to cut out the regions, but instead perform simple black region removal and input the large regions into the models.\nThe results of each model are refined by the stacking model and submitted as the final result.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Ff5b5b27bf27dbd11ae78efca228ff017%2F1.png?generation=1697414699260439&alt=media)\n\n# Segmentation Model\nWe used the 3D SwinUNETR model provided by MONAI. It works surprisingly well with even a small amount of training data. To reduce the computational cost, the entire voxel was resized into 128x128x128 before segmentation.\n\n\n# Organ Models\nFor liver, spleen, kidney, and bowel cropped regions, 2.5D CNN + LSTM models are used.\n\n## Liver and Spleen Models\nCropped region is resized into 16x386x386, and fed into dedicated models.\n\n## Kidney Model\nLeft and right kidney are independently cropped, resized, and concatinated along with horizontal axis. This enables us horizontal flip augmentation and TTA. Concatenated region becomes 16x224x448.\n\n## Bowel Model\nCropped region is resized into 64x224x224. In training bowel model, both patient level label and image level labels were used.\n\n# Stacking Model\nThe purpose of the stacking model is to directly optimize the average of the weighted logloss, which is the metric for this competition, including the logloss for any_injury. Each model is optimized for the weighted logloss for each of the injury type, but not for any_injury because it is automatically calculated from the probabilities of the other injuries. For any_injury, it is essential to optimize any_injury because the weight for any_injury is relatively large.\nAs a stacking model, we use a simple 4-layer MLP, and trained with competition metric as loss function.\n\nThe table below shows the CV evaluation results before and after stacking (The values below are out of date because we wrote this solution before the competition was extended. The final submission's CV is 0.3686).\n\n|                | bowel    | extravasation | kidney  | liver   | spleen  | any     | overall |\n|----------------|----------|---------------|---------|---------|---------|---------|---------|\n| w/o stacking   | 0.1186   | 0.4716        | 0.2948  | 0.4282  | 0.4301  | 0.5528  | 0.3827  |\n| with stacking  | 0.1034   | 0.4885        | 0.2861  | 0.4409  | 0.4452  | 0.4777  | 0.3736  |\n\n\n# tattaka's Part\nI was in charge of bowel and extravasation classification. Our solution followed [the 2-stage approach of the RSNA competition 3 years ago](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145).  \nThe basic setup is as follows\n\n## Image-Level Modeling (1st stage)\nThe input for the 1st stage is a 3-channel image including adjacent frames.  \n1epoch training was performed on all labeled images.  \n* backbone: resnetrs50  \n* head: \n  * Separate the head by bowel and extravasation\n  ```python\n  nn.Sequential(\n      nn.Conv2d(num_features[-1], 512, kernel_size=1),\n      nn.AdaptiveAvgPool2d((1, 1)),\n      Flatten(),\n      nn.Dropout(0.5, inplace=False),\n      nn.Linear(512, 1),\n  )\n  ```\nIn the 1st stage, students learn bowel and extravasation at the same time.  \n\n## Series-Level Modeling (2nd stage)\nThe input for the 2nd stage also followed the previous solution.  \nUse the 512 dimensions after Flatten in the head of the 1st stage as image features.\nImage features are created with stride=3 instead of all images, and the input sequence length is set to a maximum of 256 in the same way as the previous solution.   \nThe differences between adjacent features are combined, and the input to the model is in the form (bs, 256, 1536). \n* model:\n  * Combines the attention pooling and max pooling of the BiGRU outputs in one layer to create an series-level prediction.\n  * The BiGRU output should also predict the image label. \n  \nUnlike the 1st stage, bowel and extravasation were optimized separately.\n\n## Tricks for Successful Training\nThere are a few tricks to successful learning in this competition.\n\n* Because of large data imbalance, upsampling of the positive sample by a factor of 10 (high impact)\n  * In addition, focal loss is used\n* After training stage1 at image level, max of the logit and gt are trained again as labels for the new image level (high impact)\n  * repeated it twice\n  * Perhaps the image label is noisy\n* Rule-based removal of the outer black areas before the image is entered into the model\n  * Because of the longer computation time when using a larger resolution, a size of 384x384 is used after removing the outer area.\n\n```python\ndef region_crop(img: np.ndarray) -> np.ndarray:\n    image_1ch = (img.mean(2) * 255).astype(np.uint8)\n    kernel = np.ones((3, 3), np.uint8)\n    image_1ch = cv2.erode(image_1ch, kernel, iterations=2)\n    mask = image_1ch > 50\n    if mask.sum() == 0:\n        return img\n    rows = np.any(mask, axis=1)\n    cols = np.any(mask, axis=0)\n    y_min, y_max = np.where(rows)[0][[0, -1]]\n    x_min, x_max = np.where(cols)[0][[0, -1]]\n    if (y_max - y_min) > 5 and (x_max - x_min) > 5:\n        img = img[y_min:y_max, x_min:x_max]\n    return img\n```\nThe bowel and extravasation scores before stacking are\n|        | bowel logloss | bowel auc | ev logloss | ev auc |\n| ------ | ------------- | --------- | ---------- | ------ |\n| stage1 | 0.2719        | 0.9314    | 0.4602     | 0.8287 |\n| stage2 | 0.1167        | 0.9167    | 0.4579     | 0.8264 |\n\n## Not works\n* label smoothing\n* resnet3dcsn\n  * Not bad, but I didn't have time to tune in.\n* scaling logit\n* GeM pooling\n* Other backbone\n  * ConvNeXt is slightly worse than resnetrs50\n  * I could not get the backbone of the transformer to work \n",
      "votes": 38
    },
    {
      "id": 2484271,
      "postDate": "2023-10-16T10:59:08.120Z",
      "content": "<p>Congratulations on winning the competition at 10th position. Thanks for sharing the solution details.</p>",
      "rawMarkdown": "Congratulations on winning the competition at 10th position. Thanks for sharing the solution details."
    },
    {
      "id": 2483756,
      "postDate": "2023-10-16T00:53:56.400Z",
      "content": "<p>it's an impressive work.</p>",
      "rawMarkdown": "it's an impressive work."
    },
    {
      "id": 2483730,
      "postDate": "2023-10-16T00:17:30.900Z",
      "content": "<p>thx for sharing your great work.</p>",
      "rawMarkdown": "thx for sharing your great work."
    },
    {
      "id": 2483785,
      "postDate": "2023-10-16T01:50:50.053Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 2484271,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2023-10-16T10:59:08.120000",
      "content": "<p>Congratulations on winning the competition at 10th position. Thanks for sharing the solution details.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2483756,
      "author_name": "FLy2sky",
      "author_url": "",
      "post_date": "2023-10-16T00:53:56.400000",
      "content": "<p>it's an impressive work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2483730,
      "author_name": "dragon zhang",
      "author_url": "",
      "post_date": "2023-10-16T00:17:30.900000",
      "content": "<p>thx for sharing your great work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2483785,
      "author_name": "Sohail Hosseini",
      "author_url": "",
      "post_date": "2023-10-16T01:50:50.053000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2483722": "Congrats to all prize and medal winners! I enjoyed this competition because there are a great variety of options for solving this task as with last year's RSNA competition. We share our team's solution ( @ren4yu + @tattaka ).\n\n# Summary\nOur solution is to first segment the organs, cut out each organ region, and build a dedicated model for each organ.\nFor the bowel and extravasation classes, for which large regions must be explored, we do not perform segmentation to cut out the regions, but instead perform simple black region removal and input the large regions into the models.\nThe results of each model are refined by the stacking model and submitted as the final result.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F745525%2Ff5b5b27bf27dbd11ae78efca228ff017%2F1.png?generation=1697414699260439&alt=media)\n\n# Segmentation Model\nWe used the 3D SwinUNETR model provided by MONAI. It works surprisingly well with even a small amount of training data. To reduce the computational cost, the entire voxel was resized into 128x128x128 before segmentation.\n\n\n# Organ Models\nFor liver, spleen, kidney, and bowel cropped regions, 2.5D CNN + LSTM models are used.\n\n## Liver and Spleen Models\nCropped region is resized into 16x386x386, and fed into dedicated models.\n\n## Kidney Model\nLeft and right kidney are independently cropped, resized, and concatinated along with horizontal axis. This enables us horizontal flip augmentation and TTA. Concatenated region becomes 16x224x448.\n\n## Bowel Model\nCropped region is resized into 64x224x224. In training bowel model, both patient level label and image level labels were used.\n\n# Stacking Model\nThe purpose of the stacking model is to directly optimize the average of the weighted logloss, which is the metric for this competition, including the logloss for any_injury. Each model is optimized for the weighted logloss for each of the injury type, but not for any_injury because it is automatically calculated from the probabilities of the other injuries. For any_injury, it is essential to optimize any_injury because the weight for any_injury is relatively large.\nAs a stacking model, we use a simple 4-layer MLP, and trained with competition metric as loss function.\n\nThe table below shows the CV evaluation results before and after stacking (The values below are out of date because we wrote this solution before the competition was extended. The final submission's CV is 0.3686).\n\n|                | bowel    | extravasation | kidney  | liver   | spleen  | any     | overall |\n|----------------|----------|---------------|---------|---------|---------|---------|---------|\n| w/o stacking   | 0.1186   | 0.4716        | 0.2948  | 0.4282  | 0.4301  | 0.5528  | 0.3827  |\n| with stacking  | 0.1034   | 0.4885        | 0.2861  | 0.4409  | 0.4452  | 0.4777  | 0.3736  |\n\n\n# tattaka's Part\nI was in charge of bowel and extravasation classification. Our solution followed [the 2-stage approach of the RSNA competition 3 years ago](https://www.kaggle.com/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145).  \nThe basic setup is as follows\n\n## Image-Level Modeling (1st stage)\nThe input for the 1st stage is a 3-channel image including adjacent frames.  \n1epoch training was performed on all labeled images.  \n* backbone: resnetrs50  \n* head: \n  * Separate the head by bowel and extravasation\n  ```python\n  nn.Sequential(\n      nn.Conv2d(num_features[-1], 512, kernel_size=1),\n      nn.AdaptiveAvgPool2d((1, 1)),\n      Flatten(),\n      nn.Dropout(0.5, inplace=False),\n      nn.Linear(512, 1),\n  )\n  ```\nIn the 1st stage, students learn bowel and extravasation at the same time.  \n\n## Series-Level Modeling (2nd stage)\nThe input for the 2nd stage also followed the previous solution.  \nUse the 512 dimensions after Flatten in the head of the 1st stage as image features.\nImage features are created with stride=3 instead of all images, and the input sequence length is set to a maximum of 256 in the same way as the previous solution.   \nThe differences between adjacent features are combined, and the input to the model is in the form (bs, 256, 1536). \n* model:\n  * Combines the attention pooling and max pooling of the BiGRU outputs in one layer to create an series-level prediction.\n  * The BiGRU output should also predict the image label. \n  \nUnlike the 1st stage, bowel and extravasation were optimized separately.\n\n## Tricks for Successful Training\nThere are a few tricks to successful learning in this competition.\n\n* Because of large data imbalance, upsampling of the positive sample by a factor of 10 (high impact)\n  * In addition, focal loss is used\n* After training stage1 at image level, max of the logit and gt are trained again as labels for the new image level (high impact)\n  * repeated it twice\n  * Perhaps the image label is noisy\n* Rule-based removal of the outer black areas before the image is entered into the model\n  * Because of the longer computation time when using a larger resolution, a size of 384x384 is used after removing the outer area.\n\n```python\ndef region_crop(img: np.ndarray) -> np.ndarray:\n    image_1ch = (img.mean(2) * 255).astype(np.uint8)\n    kernel = np.ones((3, 3), np.uint8)\n    image_1ch = cv2.erode(image_1ch, kernel, iterations=2)\n    mask = image_1ch > 50\n    if mask.sum() == 0:\n        return img\n    rows = np.any(mask, axis=1)\n    cols = np.any(mask, axis=0)\n    y_min, y_max = np.where(rows)[0][[0, -1]]\n    x_min, x_max = np.where(cols)[0][[0, -1]]\n    if (y_max - y_min) > 5 and (x_max - x_min) > 5:\n        img = img[y_min:y_max, x_min:x_max]\n    return img\n```\nThe bowel and extravasation scores before stacking are\n|        | bowel logloss | bowel auc | ev logloss | ev auc |\n| ------ | ------------- | --------- | ---------- | ------ |\n| stage1 | 0.2719        | 0.9314    | 0.4602     | 0.8287 |\n| stage2 | 0.1167        | 0.9167    | 0.4579     | 0.8264 |\n\n## Not works\n* label smoothing\n* resnet3dcsn\n  * Not bad, but I didn't have time to tune in.\n* scaling logit\n* GeM pooling\n* Other backbone\n  * ConvNeXt is slightly worse than resnetrs50\n  * I could not get the backbone of the transformer to work \n",
    "2484271": "Congratulations on winning the competition at 10th position. Thanks for sharing the solution details.",
    "2483756": "it's an impressive work.",
    "2483730": "thx for sharing your great work.",
    "2483785": "Thanks for sharing!"
  }
}