{
  "id": 194145,
  "title": "1st place solution with code",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/194145",
  "author_name": "Guanshuo Xu",
  "post_date": "2020-10-30T21:07:01.465000",
  "votes": 150,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Congratulations to all the winners! Thanks to Kaggle and RSNA for hosting this competition and presenting us this interesting problem. The data size is big and of high quality and there is no shakeup. I’m glad I can win this one and I have learnt a lot during this journey.<br>\nSpecial thanks to <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for providing the topic introduction and useful input processing code. Also, credits should go to last year’s RSNA winners, lots of their ideas are incorporated in my solution.<br>\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117242\" target=\"_blank\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117242</a></p>\n<p>My solution is described below. <br>\nCode: <a href=\"https://github.com/GuanshuoXu/RSNA-STR-Pulmonary-Embolism-Detection\" target=\"_blank\">https://github.com/GuanshuoXu/RSNA-STR-Pulmonary-Embolism-Detection</a><br>\nInference kernel: <a href=\"https://www.kaggle.com/wowfattie/notebook6fff7ff27a?scriptVersionId=45476524\" target=\"_blank\">https://www.kaggle.com/wowfattie/notebook6fff7ff27a?scriptVersionId=45476524</a></p>\n<h1>Preprocessing</h1>\n<p>Early after I joined this competition, I noticed that increasing input image size from 512x512 to 640x640 improves the modeling performance. By browsing the training images, I further noticed that the lungs did not occupy large and consistent portions of the images. This is inefficient because we know input size matters and it’s not worthy to waste computing time on irrelevant things in the images, and this could also give the modeling unnecessary difficulty to learn large scale and shift invariance. So, it’s necessary to have a high-quality lung localizer. There are some existing pretrained lung localizer online, I did not try them because according to my observation it’s easy for a CNN to accurately localize the lung area from images as long as we have the bbox labels of the lungs. So, I annotated the train data and built a lung localizer with the bboxes and Efficientnet-b0 as the backbone. For simplicity I only annotated four images per study. The training and prediction process were also on only four images per study to save time. Some examples of this preprocessing are given below. The localizer is very robust even in some relatively difficult conditions. The idea of preprocessing the input is partly inspired from last year’s 2nd place solution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F478989%2F51ad8deb361d5f8c50f2265970a30d3c%2FPicture1.png?generation=1604080349102371&amp;alt=media\" alt=\"\"></p>\n<h1>Training/validation split</h1>\n<p>Since the provided data are big and of high quality, we don't have to do cross validation, a single training/validation split is reliable enough. In this competition, I randomly set aside 1000 studies for validation and used the rest 6200+ studies for training and hyperparameter tuning. For final LB submission I re-trained my models with the full training set and the optimized hyperparameters.</p>\n<h1>Image-level modeling</h1>\n<p>I used the same 2-stage training strategy as in last year’s RSNA competitions. For image-level modeling, the 3-channel input was the PE windows of the current image and its two direct neighbors. Using neighboring images has proved to be effective in last years 1st and 3rd place solutions. My experiments also confirmed that this input setting outperformed single images with 3 types of windows.</p>\n<p>Apart from predicting image-level labels, this year we are given various study-level labels. At first glance it appeared to me that, because the input of the study-level models are image embeddings, we need to use these study-level labels during image-level modeling so that the following study-level model could have sufficient knowledge to model and predict them. But after I tried lots of combinations of them and various loss masking tricks, the best performing  model in both the image-level and study-level stages was still the one trained with the image-level labels only. I’m a little puzzled how the image embeddings are encoded with the study-level labels, for example, the exact position labels (center, left, right) and the more refined acute and chronic, when the image-level models were not trained using any of those labels.</p>\n<p>The training loss was the vanilla BCE loss with linear lr scheduler. No special data sampling was applied. I found that a single epoch through the train data was the optimal for my settings. The best augmentations were </p>\n<pre><code>albumentations.RandomContrast(limit=0.2, p=1.0),\nalbumentations.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=20, border_mode=cv2.BORDER_CONSTANT, p=1.0),\nalbumentations.Cutout(num_holes=2, max_h_size=int(0.4*image_size), max_w_size=int(0.4*image_size), fill_value=0, always_apply=True, p=1.0),\n</code></pre>\n<p>My final ensemble were with one serexnext50 and one seresnext101. Their respective validation performance for image-level PE prediction was</p>\n<pre><code>                     Loss      AUC\nseresnext101        0.079    0.964\nseresnext50         0.080    0.962\n</code></pre>\n<p>Other good backbones are inception_resnet_v2 and efficientnets. Densenets and resnexts performed a lot worse. Input were resized to 576x576 after the lung localization, this was the largest size the models could finish running in the 9 hours.</p>\n<h1>Study-level modeling</h1>\n<p>Image embeddings of dimension 2048 served as the input to a RNN for both image-level and study-level modeling.  </p>\n<p>One thing we need to handle was that the number of images each study has could vary from 100+ to 1000+. As we don’t know the information of the private test data, it was hard to predefine an input sequence length for our RNN model if we want to predict all the images. Stacking all the images into a 3-D array and resizing it along the z-axis before generating image embeddings is an option, but it was not compatible to my inference pipeline. For convenience, I swapped the order of embedding generation and resizing, in other words, I chose to resize the features instead of images. For example, given a study which has N images, the input feature shape is Nx2048. If the max sequence length limit in the RNN is M, the cv2.resize function is applied to resize features to Mx2048 if N&gt;M, otherwise if N&lt;M zero-padding is used. The image-level labels and the predictions are zoomed in and out in the same way during training and inference. To find the best M, I ran a search in the step size of 32, and M=128 gave the best performance. In the train set, the majority of Ns is in the range of 200-250. This means downsizing across the z-pos first before sequence modeling improves the performance. In my final models, I actually set m=192 because I believed there might be more big Ns in the private test data.</p>\n<p>Inspired from last year's 2nd place solution, I also computed the difference of embeddings between current and the two direct neighbors and concatenate with the current features. So the input size was expanded to 2048x3. </p>\n<p>The exact RNN architecture is not very important, I settled down to only a single bidirectional GRU layer, with the study-level labels predicted by a concatenated attention weighted average pooling and max pooling over the sequence. My local validation loss was around 0.18, I have no idea why it is much higher than the LB scores.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F478989%2F5aabec8cb08e9bc19b65b9efd4c0dcdd%2FPicture5.png?generation=1604080838336664&amp;alt=media\" alt=\"\"></p>\n<h1>Postprocessing</h1>\n<p>The main purpose of the postprocessing step is to satisfy the consistency requirement of the labels. Since this consistency requirement agrees with how the data was labeled, a careful postprocessing could improve the performance. In my case, the local validation has a tiny improvement after postprocessing. The brief workflow of the postprocessing is </p>\n<pre><code>for each study:\n    if the original predictions satisfy the consistency requirement\n        do nothing\n    else\n        change the original predictions into consistent positive predictions, and compute loss between them\n        change the original predictions into consistent negative predictions, and compute loss between them\n        choose from the positive and negative predictions based on which causes the smaller loss\n</code></pre>\n<p>The weights of the loss function is almost same as the competition metric, except that the  q_i of image loss weight is replaced by a fixed 0.005 because we don’t have the ground truth of the test data. Code of this postprocessing can be found in my inference kernel.</p>",
  "messages": [
    {
      "id": 1065098,
      "postDate": "2020-10-30T21:07:01.467Z",
      "content": "<p>Congratulations to all the winners! Thanks to Kaggle and RSNA for hosting this competition and presenting us this interesting problem. The data size is big and of high quality and there is no shakeup. I’m glad I can win this one and I have learnt a lot during this journey.<br>\nSpecial thanks to <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for providing the topic introduction and useful input processing code. Also, credits should go to last year’s RSNA winners, lots of their ideas are incorporated in my solution.<br>\n<a href=\"https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117242\" target=\"_blank\">https://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117242</a></p>\n<p>My solution is described below. <br>\nCode: <a href=\"https://github.com/GuanshuoXu/RSNA-STR-Pulmonary-Embolism-Detection\" target=\"_blank\">https://github.com/GuanshuoXu/RSNA-STR-Pulmonary-Embolism-Detection</a><br>\nInference kernel: <a href=\"https://www.kaggle.com/wowfattie/notebook6fff7ff27a?scriptVersionId=45476524\" target=\"_blank\">https://www.kaggle.com/wowfattie/notebook6fff7ff27a?scriptVersionId=45476524</a></p>\n<h1>Preprocessing</h1>\n<p>Early after I joined this competition, I noticed that increasing input image size from 512x512 to 640x640 improves the modeling performance. By browsing the training images, I further noticed that the lungs did not occupy large and consistent portions of the images. This is inefficient because we know input size matters and it’s not worthy to waste computing time on irrelevant things in the images, and this could also give the modeling unnecessary difficulty to learn large scale and shift invariance. So, it’s necessary to have a high-quality lung localizer. There are some existing pretrained lung localizer online, I did not try them because according to my observation it’s easy for a CNN to accurately localize the lung area from images as long as we have the bbox labels of the lungs. So, I annotated the train data and built a lung localizer with the bboxes and Efficientnet-b0 as the backbone. For simplicity I only annotated four images per study. The training and prediction process were also on only four images per study to save time. Some examples of this preprocessing are given below. The localizer is very robust even in some relatively difficult conditions. The idea of preprocessing the input is partly inspired from last year’s 2nd place solution.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F478989%2F51ad8deb361d5f8c50f2265970a30d3c%2FPicture1.png?generation=1604080349102371&amp;alt=media\" alt=\"\"></p>\n<h1>Training/validation split</h1>\n<p>Since the provided data are big and of high quality, we don't have to do cross validation, a single training/validation split is reliable enough. In this competition, I randomly set aside 1000 studies for validation and used the rest 6200+ studies for training and hyperparameter tuning. For final LB submission I re-trained my models with the full training set and the optimized hyperparameters.</p>\n<h1>Image-level modeling</h1>\n<p>I used the same 2-stage training strategy as in last year’s RSNA competitions. For image-level modeling, the 3-channel input was the PE windows of the current image and its two direct neighbors. Using neighboring images has proved to be effective in last years 1st and 3rd place solutions. My experiments also confirmed that this input setting outperformed single images with 3 types of windows.</p>\n<p>Apart from predicting image-level labels, this year we are given various study-level labels. At first glance it appeared to me that, because the input of the study-level models are image embeddings, we need to use these study-level labels during image-level modeling so that the following study-level model could have sufficient knowledge to model and predict them. But after I tried lots of combinations of them and various loss masking tricks, the best performing  model in both the image-level and study-level stages was still the one trained with the image-level labels only. I’m a little puzzled how the image embeddings are encoded with the study-level labels, for example, the exact position labels (center, left, right) and the more refined acute and chronic, when the image-level models were not trained using any of those labels.</p>\n<p>The training loss was the vanilla BCE loss with linear lr scheduler. No special data sampling was applied. I found that a single epoch through the train data was the optimal for my settings. The best augmentations were </p>\n<pre><code>albumentations.RandomContrast(limit=0.2, p=1.0),\nalbumentations.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=20, border_mode=cv2.BORDER_CONSTANT, p=1.0),\nalbumentations.Cutout(num_holes=2, max_h_size=int(0.4*image_size), max_w_size=int(0.4*image_size), fill_value=0, always_apply=True, p=1.0),\n</code></pre>\n<p>My final ensemble were with one serexnext50 and one seresnext101. Their respective validation performance for image-level PE prediction was</p>\n<pre><code>                     Loss      AUC\nseresnext101        0.079    0.964\nseresnext50         0.080    0.962\n</code></pre>\n<p>Other good backbones are inception_resnet_v2 and efficientnets. Densenets and resnexts performed a lot worse. Input were resized to 576x576 after the lung localization, this was the largest size the models could finish running in the 9 hours.</p>\n<h1>Study-level modeling</h1>\n<p>Image embeddings of dimension 2048 served as the input to a RNN for both image-level and study-level modeling.  </p>\n<p>One thing we need to handle was that the number of images each study has could vary from 100+ to 1000+. As we don’t know the information of the private test data, it was hard to predefine an input sequence length for our RNN model if we want to predict all the images. Stacking all the images into a 3-D array and resizing it along the z-axis before generating image embeddings is an option, but it was not compatible to my inference pipeline. For convenience, I swapped the order of embedding generation and resizing, in other words, I chose to resize the features instead of images. For example, given a study which has N images, the input feature shape is Nx2048. If the max sequence length limit in the RNN is M, the cv2.resize function is applied to resize features to Mx2048 if N&gt;M, otherwise if N&lt;M zero-padding is used. The image-level labels and the predictions are zoomed in and out in the same way during training and inference. To find the best M, I ran a search in the step size of 32, and M=128 gave the best performance. In the train set, the majority of Ns is in the range of 200-250. This means downsizing across the z-pos first before sequence modeling improves the performance. In my final models, I actually set m=192 because I believed there might be more big Ns in the private test data.</p>\n<p>Inspired from last year's 2nd place solution, I also computed the difference of embeddings between current and the two direct neighbors and concatenate with the current features. So the input size was expanded to 2048x3. </p>\n<p>The exact RNN architecture is not very important, I settled down to only a single bidirectional GRU layer, with the study-level labels predicted by a concatenated attention weighted average pooling and max pooling over the sequence. My local validation loss was around 0.18, I have no idea why it is much higher than the LB scores.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F478989%2F5aabec8cb08e9bc19b65b9efd4c0dcdd%2FPicture5.png?generation=1604080838336664&amp;alt=media\" alt=\"\"></p>\n<h1>Postprocessing</h1>\n<p>The main purpose of the postprocessing step is to satisfy the consistency requirement of the labels. Since this consistency requirement agrees with how the data was labeled, a careful postprocessing could improve the performance. In my case, the local validation has a tiny improvement after postprocessing. The brief workflow of the postprocessing is </p>\n<pre><code>for each study:\n    if the original predictions satisfy the consistency requirement\n        do nothing\n    else\n        change the original predictions into consistent positive predictions, and compute loss between them\n        change the original predictions into consistent negative predictions, and compute loss between them\n        choose from the positive and negative predictions based on which causes the smaller loss\n</code></pre>\n<p>The weights of the loss function is almost same as the competition metric, except that the  q_i of image loss weight is replaced by a fixed 0.005 because we don’t have the ground truth of the test data. Code of this postprocessing can be found in my inference kernel.</p>",
      "rawMarkdown": "Congratulations to all the winners! Thanks to Kaggle and RSNA for hosting this competition and presenting us this interesting problem. The data size is big and of high quality and there is no shakeup. I’m glad I can win this one and I have learnt a lot during this journey.\nSpecial thanks to @vaillant for providing the topic introduction and useful input processing code. Also, credits should go to last year’s RSNA winners, lots of their ideas are incorporated in my solution.\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117242\n\nMy solution is described below. \nCode: https://github.com/GuanshuoXu/RSNA-STR-Pulmonary-Embolism-Detection\nInference kernel: https://www.kaggle.com/wowfattie/notebook6fff7ff27a?scriptVersionId=45476524\n\n# Preprocessing\n\nEarly after I joined this competition, I noticed that increasing input image size from 512x512 to 640x640 improves the modeling performance. By browsing the training images, I further noticed that the lungs did not occupy large and consistent portions of the images. This is inefficient because we know input size matters and it’s not worthy to waste computing time on irrelevant things in the images, and this could also give the modeling unnecessary difficulty to learn large scale and shift invariance. So, it’s necessary to have a high-quality lung localizer. There are some existing pretrained lung localizer online, I did not try them because according to my observation it’s easy for a CNN to accurately localize the lung area from images as long as we have the bbox labels of the lungs. So, I annotated the train data and built a lung localizer with the bboxes and Efficientnet-b0 as the backbone. For simplicity I only annotated four images per study. The training and prediction process were also on only four images per study to save time. Some examples of this preprocessing are given below. The localizer is very robust even in some relatively difficult conditions. The idea of preprocessing the input is partly inspired from last year’s 2nd place solution.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F478989%2F51ad8deb361d5f8c50f2265970a30d3c%2FPicture1.png?generation=1604080349102371&alt=media)\n\n# Training/validation split\n\nSince the provided data are big and of high quality, we don't have to do cross validation, a single training/validation split is reliable enough. In this competition, I randomly set aside 1000 studies for validation and used the rest 6200+ studies for training and hyperparameter tuning. For final LB submission I re-trained my models with the full training set and the optimized hyperparameters.\n\n# Image-level modeling\n\nI used the same 2-stage training strategy as in last year’s RSNA competitions. For image-level modeling, the 3-channel input was the PE windows of the current image and its two direct neighbors. Using neighboring images has proved to be effective in last years 1st and 3rd place solutions. My experiments also confirmed that this input setting outperformed single images with 3 types of windows.\n\nApart from predicting image-level labels, this year we are given various study-level labels. At first glance it appeared to me that, because the input of the study-level models are image embeddings, we need to use these study-level labels during image-level modeling so that the following study-level model could have sufficient knowledge to model and predict them. But after I tried lots of combinations of them and various loss masking tricks, the best performing  model in both the image-level and study-level stages was still the one trained with the image-level labels only. I’m a little puzzled how the image embeddings are encoded with the study-level labels, for example, the exact position labels (center, left, right) and the more refined acute and chronic, when the image-level models were not trained using any of those labels.\n\nThe training loss was the vanilla BCE loss with linear lr scheduler. No special data sampling was applied. I found that a single epoch through the train data was the optimal for my settings. The best augmentations were \n\n```\nalbumentations.RandomContrast(limit=0.2, p=1.0),\nalbumentations.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=20, border_mode=cv2.BORDER_CONSTANT, p=1.0),\nalbumentations.Cutout(num_holes=2, max_h_size=int(0.4*image_size), max_w_size=int(0.4*image_size), fill_value=0, always_apply=True, p=1.0),\n```\n\nMy final ensemble were with one serexnext50 and one seresnext101. Their respective validation performance for image-level PE prediction was\n\n```\n                     Loss      AUC\nseresnext101        0.079    0.964\nseresnext50         0.080    0.962\n```\n\nOther good backbones are inception_resnet_v2 and efficientnets. Densenets and resnexts performed a lot worse. Input were resized to 576x576 after the lung localization, this was the largest size the models could finish running in the 9 hours.\n\n# Study-level modeling\n\nImage embeddings of dimension 2048 served as the input to a RNN for both image-level and study-level modeling.  \n\nOne thing we need to handle was that the number of images each study has could vary from 100+ to 1000+. As we don’t know the information of the private test data, it was hard to predefine an input sequence length for our RNN model if we want to predict all the images. Stacking all the images into a 3-D array and resizing it along the z-axis before generating image embeddings is an option, but it was not compatible to my inference pipeline. For convenience, I swapped the order of embedding generation and resizing, in other words, I chose to resize the features instead of images. For example, given a study which has N images, the input feature shape is Nx2048. If the max sequence length limit in the RNN is M, the cv2.resize function is applied to resize features to Mx2048 if N>M, otherwise if N<M zero-padding is used. The image-level labels and the predictions are zoomed in and out in the same way during training and inference. To find the best M, I ran a search in the step size of 32, and M=128 gave the best performance. In the train set, the majority of Ns is in the range of 200-250. This means downsizing across the z-pos first before sequence modeling improves the performance. In my final models, I actually set m=192 because I believed there might be more big Ns in the private test data.\n\nInspired from last year's 2nd place solution, I also computed the difference of embeddings between current and the two direct neighbors and concatenate with the current features. So the input size was expanded to 2048x3. \n\nThe exact RNN architecture is not very important, I settled down to only a single bidirectional GRU layer, with the study-level labels predicted by a concatenated attention weighted average pooling and max pooling over the sequence. My local validation loss was around 0.18, I have no idea why it is much higher than the LB scores.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F478989%2F5aabec8cb08e9bc19b65b9efd4c0dcdd%2FPicture5.png?generation=1604080838336664&alt=media)\n\n# Postprocessing\n\nThe main purpose of the postprocessing step is to satisfy the consistency requirement of the labels. Since this consistency requirement agrees with how the data was labeled, a careful postprocessing could improve the performance. In my case, the local validation has a tiny improvement after postprocessing. The brief workflow of the postprocessing is \n\n```\nfor each study:\n    if the original predictions satisfy the consistency requirement\n        do nothing\n    else\n        change the original predictions into consistent positive predictions, and compute loss between them\n        change the original predictions into consistent negative predictions, and compute loss between them\n        choose from the positive and negative predictions based on which causes the smaller loss\n```\n\n\nThe weights of the loss function is almost same as the competition metric, except that the  q_i of image loss weight is replaced by a fixed 0.005 because we don’t have the ground truth of the test data. Code of this postprocessing can be found in my inference kernel.\n",
      "votes": 149
    },
    {
      "id": 1065212,
      "postDate": "2020-10-31T03:11:41.330Z",
      "content": "<p>Really waited for your solution. Thanks for sharing your details solution and code. Congrats on 1st place and 1 st global competitions ranking <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> !</p>",
      "rawMarkdown": "Really waited for your solution. Thanks for sharing your details solution and code. Congrats on 1st place and 1 st global competitions ranking @wowfattie !",
      "votes": 7
    },
    {
      "id": 1065229,
      "postDate": "2020-10-31T04:00:23.793Z",
      "content": "<p>Congratulations Guanshuo. Brilliant and simple solution, well done! I like how you located the lungs, that was smart. My best models used 320x320 center crops from 512x512 which looks like it corresponds to your lung bbox.</p>\n<blockquote>\n  <p>I actually set m=192 because I believed there might be more big Ns in the private test data.</p>\n</blockquote>\n<p>This is good intuition. I probed the private LB. The public LB has 0% rows associtated with exams of more than 350 images. The private LB has 10% rows associated with exams of more than 350 images. (The train data had around 8%)</p>",
      "rawMarkdown": "Congratulations Guanshuo. Brilliant and simple solution, well done! I like how you located the lungs, that was smart. My best models used 320x320 center crops from 512x512 which looks like it corresponds to your lung bbox.\n\n>  I actually set m=192 because I believed there might be more big Ns in the private test data.\n\nThis is good intuition. I probed the private LB. The public LB has 0% rows associtated with exams of more than 350 images. The private LB has 10% rows associated with exams of more than 350 images. (The train data had around 8%)",
      "votes": 8
    },
    {
      "id": 1065117,
      "postDate": "2020-10-30T21:38:08.553Z",
      "content": "<p>Congratulations on 1st place and competitions rank 1! Thanks for sharing your solution. </p>\n<blockquote>\n  <p>I’m a little puzzled how the image embeddings are encoded with the study-level labels, for example, the exact position labels (center, left, right) and the more refined acute and chronic, when the image-level models were not trained using any of those labels.</p>\n</blockquote>\n<p>This is interesting- I'm also surprised. I used the study-level labels while training the image-level model because I thought they were necessary to encode those features as well, but apparently not. </p>\n<p>How did you apply the bounding box predictions to the image? Did you use the largest box across all slices for each slice? Or individual boxes per slice? Some slices don't have lungs at all. </p>",
      "rawMarkdown": "Congratulations on 1st place and competitions rank 1! Thanks for sharing your solution. \n\n> I’m a little puzzled how the image embeddings are encoded with the study-level labels, for example, the exact position labels (center, left, right) and the more refined acute and chronic, when the image-level models were not trained using any of those labels.\n\nThis is interesting- I'm also surprised. I used the study-level labels while training the image-level model because I thought they were necessary to encode those features as well, but apparently not. \n\nHow did you apply the bounding box predictions to the image? Did you use the largest box across all slices for each slice? Or individual boxes per slice? Some slices don't have lungs at all. ",
      "votes": 4,
      "replies": [
        {
          "id": 1065159,
          "postDate": "2020-10-30T23:31:21.060Z",
          "content": "<p>Thanks. Congratulations to you too.<br>\nI computed the largest possible bbox from the four bboxes for each study/scan. The code from my inference kernel.</p>\n<pre><code>        xmin = np.round(min([bbox[0,0], bbox[1,0], bbox[2,0], bbox[3,0]])*512)\n        ymin = np.round(min([bbox[0,1], bbox[1,1], bbox[2,1], bbox[3,1]])*512)\n        xmax = np.round(max([bbox[0,2], bbox[1,2], bbox[2,2], bbox[3,2]])*512)\n        ymax = np.round(max([bbox[0,3], bbox[1,3], bbox[2,3], bbox[3,3]])*512)\n        bbox_dict[series_id] = [int(max(0, xmin)), int(max(0, ymin)), int(min(512, xmax)), int(min(512, ymax))]\n</code></pre>",
          "rawMarkdown": "Thanks. Congratulations to you too.\nI computed the largest possible bbox from the four bboxes for each study/scan. The code from my inference kernel.\n```\n        xmin = np.round(min([bbox[0,0], bbox[1,0], bbox[2,0], bbox[3,0]])*512)\n        ymin = np.round(min([bbox[0,1], bbox[1,1], bbox[2,1], bbox[3,1]])*512)\n        xmax = np.round(max([bbox[0,2], bbox[1,2], bbox[2,2], bbox[3,2]])*512)\n        ymax = np.round(max([bbox[0,3], bbox[1,3], bbox[2,3], bbox[3,3]])*512)\n        bbox_dict[series_id] = [int(max(0, xmin)), int(max(0, ymin)), int(min(512, xmax)), int(min(512, ymax))]\n```",
          "votes": 3
        }
      ]
    },
    {
      "id": 1065136,
      "postDate": "2020-10-30T22:12:22.847Z",
      "content": "<p>Congratulations and thanks for sharing your solution! I also started off using lungmasks and then finding the largest bounding box for a given study but it wasn't feasible given the code requirements and I couldn't think of training a new model like you did unfortunately. Do you know how much improvement lung cropping gave in your final solution or was it mainly helpful for using computational resources efficiently? </p>\n<p>We also noticed a huge gap (0.03) between our local CV (1456 studies for validation) vs the public/private LB, don't have an idea why it happens.</p>\n<p>Resizing sequence input and then resizing the output back to map the original image number was something I was wondering if possible, thanks for sharing it is!</p>\n<p>Pretty novel and clever postprocessing!</p>\n<p>What kind of attention mechanism is applied to GRU hidden outputs?</p>",
      "rawMarkdown": "Congratulations and thanks for sharing your solution! I also started off using lungmasks and then finding the largest bounding box for a given study but it wasn't feasible given the code requirements and I couldn't think of training a new model like you did unfortunately. Do you know how much improvement lung cropping gave in your final solution or was it mainly helpful for using computational resources efficiently? \n\nWe also noticed a huge gap (0.03) between our local CV (1456 studies for validation) vs the public/private LB, don't have an idea why it happens.\n\nResizing sequence input and then resizing the output back to map the original image number was something I was wondering if possible, thanks for sharing it is!\n\nPretty novel and clever postprocessing!\n\nWhat kind of attention mechanism is applied to GRU hidden outputs?",
      "votes": 2,
      "replies": [
        {
          "id": 1065161,
          "postDate": "2020-10-30T23:39:12.047Z",
          "content": "<p>I only have a result of a naive model, with lung localization the AUC improved from 0.945 to 0.953.</p>",
          "rawMarkdown": "I only have a result of a naive model, with lung localization the AUC improved from 0.945 to 0.953.",
          "votes": 2
        },
        {
          "id": 1065194,
          "postDate": "2020-10-31T02:20:18.283Z",
          "content": "<p>Thanks that's helpful!</p>",
          "rawMarkdown": "Thanks that's helpful!"
        }
      ]
    },
    {
      "id": 3434703,
      "postDate": "2026-04-03T06:22:18.973Z",
      "content": "<p>Congratulations Guanshuo. Brilliant and simple solution</p>",
      "rawMarkdown": "Congratulations Guanshuo. Brilliant and simple solution"
    },
    {
      "id": 2456806,
      "postDate": "2023-09-26T12:25:38.493Z",
      "content": "<p>this is a very detailed solution, thanks  </p>",
      "rawMarkdown": "this is a very detailed solution, thanks  "
    },
    {
      "id": 1851358,
      "postDate": "2022-07-11T07:27:25.697Z",
      "content": "<p>Would you please provide the trained weights?</p>",
      "rawMarkdown": "Would you please provide the trained weights?"
    },
    {
      "id": 1738699,
      "postDate": "2022-03-29T13:09:43.133Z",
      "content": "<p>Hello, author, congratulations on winning the first prize. I tried to reproduce your code content, but the final effect failed to reach your AUC of 0.96. Do you know the reason?</p>",
      "rawMarkdown": "Hello, author, congratulations on winning the first prize. I tried to reproduce your code content, but the final effect failed to reach your AUC of 0.96. Do you know the reason?"
    },
    {
      "id": 1124234,
      "postDate": "2020-12-23T18:57:17.100Z",
      "content": "<p>Thank you  for sharing, and congratulation for 1st place.</p>",
      "rawMarkdown": "Thank you  for sharing, and congratulation for 1st place."
    },
    {
      "id": 1090057,
      "postDate": "2020-11-25T03:14:37.813Z",
      "content": "<p>Hi thank you for posting this awesome code! I'm trying to run the lung localizer and getting \"KeyErrors\" like there are missing file names in the lung localization lung_bbox.csv file</p>\n<p>Any idea what might be going wrong?</p>",
      "rawMarkdown": "Hi thank you for posting this awesome code! I'm trying to run the lung localizer and getting \"KeyErrors\" like there are missing file names in the lung localization lung_bbox.csv file\n\nAny idea what might be going wrong?",
      "replies": [
        {
          "id": 1090158,
          "postDate": "2020-11-25T06:01:49.080Z",
          "content": "<p>actually getting this error:<br>\nKeyError: Caught KeyError in DataLoader worker process 0</p>\n<p>Maybe its a memory issue trying to run this large dataset on Kaggle?</p>",
          "rawMarkdown": "actually getting this error:\nKeyError: Caught KeyError in DataLoader worker process 0\n\nMaybe its a memory issue trying to run this large dataset on Kaggle?"
        }
      ]
    },
    {
      "id": 1076765,
      "postDate": "2020-11-12T21:08:09.330Z",
      "content": "<p>Well done on the achievement and thanks for this amazing breakdown and the accompanying github repo. I have one question (actually I would ask a hundred if I could 😜). Why do you use a max pool instead of average pool over the sequence dimension?</p>\n<pre><code>max_pool, _ = torch.max(h_lstm1, 1)\n</code></pre>\n<p>I was actually working on another project and realised max pool worked better than average in a similar scenario. But would love to get your thoughts.</p>",
      "rawMarkdown": "Well done on the achievement and thanks for this amazing breakdown and the accompanying github repo. I have one question (actually I would ask a hundred if I could 😜). Why do you use a max pool instead of average pool over the sequence dimension?\n\n```\nmax_pool, _ = torch.max(h_lstm1, 1)\n```\n\nI was actually working on another project and realised max pool worked better than average in a similar scenario. But would love to get your thoughts.",
      "replies": [
        {
          "id": 1076888,
          "postDate": "2020-11-13T02:39:26.723Z",
          "content": "<p>I used both max and attention pooling. The attention pooling is similar to average pooling.</p>",
          "rawMarkdown": "I used both max and attention pooling. The attention pooling is similar to average pooling.",
          "votes": 2
        },
        {
          "id": 1077099,
          "postDate": "2020-11-13T08:37:02.870Z",
          "content": "<p>Thank you! Yes, your work has finally convinced me to take some time to learn about attention. Nevertheless, any logic or anecdotal experience behind the choice? I ask because <a href=\"https://stats.stackexchange.com/questions/495954/what-are-ways-to-learn-a-classifier-for-labelling-a-series-of-images-rather-than\" target=\"_blank\">I've been on this topic for some days now</a></p>",
          "rawMarkdown": "Thank you! Yes, your work has finally convinced me to take some time to learn about attention. Nevertheless, any logic or anecdotal experience behind the choice? I ask because [I've been on this topic for some days now](https://stats.stackexchange.com/questions/495954/what-are-ways-to-learn-a-classifier-for-labelling-a-series-of-images-rather-than)"
        },
        {
          "id": 1077534,
          "postDate": "2020-11-13T17:30:11.877Z",
          "content": "<p>Sequence model may not be suitable for your problem. You can take a lot at some previous kaggle competitions, for example<br>\n<a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification\" target=\"_blank\">https://www.kaggle.com/c/yelp-restaurant-photo-classification</a><br>\n<a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge\" target=\"_blank\">https://www.kaggle.com/c/cdiscount-image-classification-challenge</a></p>",
          "rawMarkdown": "Sequence model may not be suitable for your problem. You can take a lot at some previous kaggle competitions, for example\nhttps://www.kaggle.com/c/yelp-restaurant-photo-classification\nhttps://www.kaggle.com/c/cdiscount-image-classification-challenge",
          "votes": 2
        }
      ]
    },
    {
      "id": 1070691,
      "postDate": "2020-11-06T04:13:10.640Z",
      "content": "<p>Nice work. Congrats on 1st place</p>",
      "rawMarkdown": "Nice work. Congrats on 1st place"
    },
    {
      "id": 1070619,
      "postDate": "2020-11-06T00:57:06.480Z",
      "content": "<p>Nice solution</p>",
      "rawMarkdown": "Nice solution"
    },
    {
      "id": 1069890,
      "postDate": "2020-11-05T04:03:02.087Z",
      "content": "<p>Congrats! Very good work.</p>",
      "rawMarkdown": "Congrats! Very good work."
    },
    {
      "id": 1068578,
      "postDate": "2020-11-03T14:29:59.087Z",
      "content": "<p>Amazing.Stuff</p>",
      "rawMarkdown": "Amazing.Stuff\n"
    },
    {
      "id": 1068313,
      "postDate": "2020-11-03T09:11:52.827Z",
      "content": "<p>congratulations and thanks for sharing the solution. Thanks for the guidance.</p>",
      "rawMarkdown": "congratulations and thanks for sharing the solution. Thanks for the guidance."
    },
    {
      "id": 1068237,
      "postDate": "2020-11-03T07:46:54.387Z",
      "content": "<p>Nice! Congratulations on the first place :)</p>",
      "rawMarkdown": "Nice! Congratulations on the first place :)"
    },
    {
      "id": 1067998,
      "postDate": "2020-11-03T00:51:23.720Z",
      "content": "<p>Congratulations! Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congratulations! Thanks for sharing your solution."
    },
    {
      "id": 1067049,
      "postDate": "2020-11-02T10:10:34.107Z",
      "content": "<blockquote>\n  <p>So, I annotated the train data and built a lung localizer with the bboxes and Efficientnet-b0 as the backbone.</p>\n</blockquote>\n<p>Awesome !! </p>",
      "rawMarkdown": ">  So, I annotated the train data and built a lung localizer with the bboxes and Efficientnet-b0 as the backbone.\n\nAwesome !! "
    },
    {
      "id": 1065177,
      "postDate": "2020-10-31T00:49:47.710Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1738701,
      "postDate": "2022-03-29T13:11:48.393Z",
      "content": "<p>Thank you for your reply!</p>",
      "rawMarkdown": "Thank you for your reply!"
    }
  ],
  "comments": [
    {
      "id": 1065212,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-10-31T03:11:41.330000",
      "content": "<p>Really waited for your solution. Thanks for sharing your details solution and code. Congrats on 1st place and 1 st global competitions ranking <a href=\"https://www.kaggle.com/wowfattie\" target=\"_blank\">@wowfattie</a> !</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1065229,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-10-31T04:00:23.793000",
      "content": "<p>Congratulations Guanshuo. Brilliant and simple solution, well done! I like how you located the lungs, that was smart. My best models used 320x320 center crops from 512x512 which looks like it corresponds to your lung bbox.</p>\n<blockquote>\n  <p>I actually set m=192 because I believed there might be more big Ns in the private test data.</p>\n</blockquote>\n<p>This is good intuition. I probed the private LB. The public LB has 0% rows associtated with exams of more than 350 images. The private LB has 10% rows associated with exams of more than 350 images. (The train data had around 8%)</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 1065117,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-10-30T21:38:08.553000",
      "content": "<p>Congratulations on 1st place and competitions rank 1! Thanks for sharing your solution. </p>\n<blockquote>\n  <p>I’m a little puzzled how the image embeddings are encoded with the study-level labels, for example, the exact position labels (center, left, right) and the more refined acute and chronic, when the image-level models were not trained using any of those labels.</p>\n</blockquote>\n<p>This is interesting- I'm also surprised. I used the study-level labels while training the image-level model because I thought they were necessary to encode those features as well, but apparently not. </p>\n<p>How did you apply the bounding box predictions to the image? Did you use the largest box across all slices for each slice? Or individual boxes per slice? Some slices don't have lungs at all. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1065159,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-10-30T23:31:21.060000",
          "content": "<p>Thanks. Congratulations to you too.<br>\nI computed the largest possible bbox from the four bboxes for each study/scan. The code from my inference kernel.</p>\n<pre><code>        xmin = np.round(min([bbox[0,0], bbox[1,0], bbox[2,0], bbox[3,0]])*512)\n        ymin = np.round(min([bbox[0,1], bbox[1,1], bbox[2,1], bbox[3,1]])*512)\n        xmax = np.round(max([bbox[0,2], bbox[1,2], bbox[2,2], bbox[3,2]])*512)\n        ymax = np.round(max([bbox[0,3], bbox[1,3], bbox[2,3], bbox[3,3]])*512)\n        bbox_dict[series_id] = [int(max(0, xmin)), int(max(0, ymin)), int(min(512, xmax)), int(min(512, ymax))]\n</code></pre>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1065136,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-10-30T22:12:22.847000",
      "content": "<p>Congratulations and thanks for sharing your solution! I also started off using lungmasks and then finding the largest bounding box for a given study but it wasn't feasible given the code requirements and I couldn't think of training a new model like you did unfortunately. Do you know how much improvement lung cropping gave in your final solution or was it mainly helpful for using computational resources efficiently? </p>\n<p>We also noticed a huge gap (0.03) between our local CV (1456 studies for validation) vs the public/private LB, don't have an idea why it happens.</p>\n<p>Resizing sequence input and then resizing the output back to map the original image number was something I was wondering if possible, thanks for sharing it is!</p>\n<p>Pretty novel and clever postprocessing!</p>\n<p>What kind of attention mechanism is applied to GRU hidden outputs?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1065161,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-10-30T23:39:12.047000",
          "content": "<p>I only have a result of a naive model, with lung localization the AUC improved from 0.945 to 0.953.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1065194,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-10-31T02:20:18.283000",
          "content": "<p>Thanks that's helpful!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3434703,
      "author_name": "Mahmoud Alyosify",
      "author_url": "",
      "post_date": "2026-04-03T06:22:18.973000",
      "content": "<p>Congratulations Guanshuo. Brilliant and simple solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2456806,
      "author_name": "Abdulquadri Adegbiji",
      "author_url": "",
      "post_date": "2023-09-26T12:25:38.493000",
      "content": "<p>this is a very detailed solution, thanks  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1851358,
      "author_name": "Hi-AI",
      "author_url": "",
      "post_date": "2022-07-11T07:27:25.697000",
      "content": "<p>Would you please provide the trained weights?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1738699,
      "author_name": "li huatao",
      "author_url": "",
      "post_date": "2022-03-29T13:09:43.133000",
      "content": "<p>Hello, author, congratulations on winning the first prize. I tried to reproduce your code content, but the final effect failed to reach your AUC of 0.96. Do you know the reason?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1124234,
      "author_name": "Ammar Nassr Mohammed",
      "author_url": "",
      "post_date": "2020-12-23T18:57:17.100000",
      "content": "<p>Thank you  for sharing, and congratulation for 1st place.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1090057,
      "author_name": "rivesis",
      "author_url": "",
      "post_date": "2020-11-25T03:14:37.813000",
      "content": "<p>Hi thank you for posting this awesome code! I'm trying to run the lung localizer and getting \"KeyErrors\" like there are missing file names in the lung localization lung_bbox.csv file</p>\n<p>Any idea what might be going wrong?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1090158,
          "author_name": "rivesis",
          "author_url": "",
          "post_date": "2020-11-25T06:01:49.080000",
          "content": "<p>actually getting this error:<br>\nKeyError: Caught KeyError in DataLoader worker process 0</p>\n<p>Maybe its a memory issue trying to run this large dataset on Kaggle?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1076765,
      "author_name": "Alexander Soare",
      "author_url": "",
      "post_date": "2020-11-12T21:08:09.330000",
      "content": "<p>Well done on the achievement and thanks for this amazing breakdown and the accompanying github repo. I have one question (actually I would ask a hundred if I could 😜). Why do you use a max pool instead of average pool over the sequence dimension?</p>\n<pre><code>max_pool, _ = torch.max(h_lstm1, 1)\n</code></pre>\n<p>I was actually working on another project and realised max pool worked better than average in a similar scenario. But would love to get your thoughts.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1076888,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-11-13T02:39:26.723000",
          "content": "<p>I used both max and attention pooling. The attention pooling is similar to average pooling.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1077099,
          "author_name": "Alexander Soare",
          "author_url": "",
          "post_date": "2020-11-13T08:37:02.870000",
          "content": "<p>Thank you! Yes, your work has finally convinced me to take some time to learn about attention. Nevertheless, any logic or anecdotal experience behind the choice? I ask because <a href=\"https://stats.stackexchange.com/questions/495954/what-are-ways-to-learn-a-classifier-for-labelling-a-series-of-images-rather-than\" target=\"_blank\">I've been on this topic for some days now</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1077534,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2020-11-13T17:30:11.877000",
          "content": "<p>Sequence model may not be suitable for your problem. You can take a lot at some previous kaggle competitions, for example<br>\n<a href=\"https://www.kaggle.com/c/yelp-restaurant-photo-classification\" target=\"_blank\">https://www.kaggle.com/c/yelp-restaurant-photo-classification</a><br>\n<a href=\"https://www.kaggle.com/c/cdiscount-image-classification-challenge\" target=\"_blank\">https://www.kaggle.com/c/cdiscount-image-classification-challenge</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1070691,
      "author_name": "Naga Munindra Pasupuleti",
      "author_url": "",
      "post_date": "2020-11-06T04:13:10.640000",
      "content": "<p>Nice work. Congrats on 1st place</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1070619,
      "author_name": "Nguyen Nhat Minh",
      "author_url": "",
      "post_date": "2020-11-06T00:57:06.480000",
      "content": "<p>Nice solution</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1069890,
      "author_name": "David Llanio",
      "author_url": "",
      "post_date": "2020-11-05T04:03:02.087000",
      "content": "<p>Congrats! Very good work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1068578,
      "author_name": "BipanGharti",
      "author_url": "",
      "post_date": "2020-11-03T14:29:59.087000",
      "content": "<p>Amazing.Stuff</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1068313,
      "author_name": "keerthi pakka",
      "author_url": "",
      "post_date": "2020-11-03T09:11:52.827000",
      "content": "<p>congratulations and thanks for sharing the solution. Thanks for the guidance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1068237,
      "author_name": "kagglebaggledaggle",
      "author_url": "",
      "post_date": "2020-11-03T07:46:54.387000",
      "content": "<p>Nice! Congratulations on the first place :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1067998,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2020-11-03T00:51:23.720000",
      "content": "<p>Congratulations! Thanks for sharing your solution.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1067049,
      "author_name": "Vee",
      "author_url": "",
      "post_date": "2020-11-02T10:10:34.107000",
      "content": "<blockquote>\n  <p>So, I annotated the train data and built a lung localizer with the bboxes and Efficientnet-b0 as the backbone.</p>\n</blockquote>\n<p>Awesome !! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1065177,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-31T00:49:47.710000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1738701,
      "author_name": "li huatao",
      "author_url": "",
      "post_date": "2022-03-29T13:11:48.393000",
      "content": "<p>Thank you for your reply!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1065098": "Congratulations to all the winners! Thanks to Kaggle and RSNA for hosting this competition and presenting us this interesting problem. The data size is big and of high quality and there is no shakeup. I’m glad I can win this one and I have learnt a lot during this journey.\nSpecial thanks to @vaillant for providing the topic introduction and useful input processing code. Also, credits should go to last year’s RSNA winners, lots of their ideas are incorporated in my solution.\nhttps://www.kaggle.com/c/rsna-intracranial-hemorrhage-detection/discussion/117242\n\nMy solution is described below. \nCode: https://github.com/GuanshuoXu/RSNA-STR-Pulmonary-Embolism-Detection\nInference kernel: https://www.kaggle.com/wowfattie/notebook6fff7ff27a?scriptVersionId=45476524\n\n# Preprocessing\n\nEarly after I joined this competition, I noticed that increasing input image size from 512x512 to 640x640 improves the modeling performance. By browsing the training images, I further noticed that the lungs did not occupy large and consistent portions of the images. This is inefficient because we know input size matters and it’s not worthy to waste computing time on irrelevant things in the images, and this could also give the modeling unnecessary difficulty to learn large scale and shift invariance. So, it’s necessary to have a high-quality lung localizer. There are some existing pretrained lung localizer online, I did not try them because according to my observation it’s easy for a CNN to accurately localize the lung area from images as long as we have the bbox labels of the lungs. So, I annotated the train data and built a lung localizer with the bboxes and Efficientnet-b0 as the backbone. For simplicity I only annotated four images per study. The training and prediction process were also on only four images per study to save time. Some examples of this preprocessing are given below. The localizer is very robust even in some relatively difficult conditions. The idea of preprocessing the input is partly inspired from last year’s 2nd place solution.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F478989%2F51ad8deb361d5f8c50f2265970a30d3c%2FPicture1.png?generation=1604080349102371&alt=media)\n\n# Training/validation split\n\nSince the provided data are big and of high quality, we don't have to do cross validation, a single training/validation split is reliable enough. In this competition, I randomly set aside 1000 studies for validation and used the rest 6200+ studies for training and hyperparameter tuning. For final LB submission I re-trained my models with the full training set and the optimized hyperparameters.\n\n# Image-level modeling\n\nI used the same 2-stage training strategy as in last year’s RSNA competitions. For image-level modeling, the 3-channel input was the PE windows of the current image and its two direct neighbors. Using neighboring images has proved to be effective in last years 1st and 3rd place solutions. My experiments also confirmed that this input setting outperformed single images with 3 types of windows.\n\nApart from predicting image-level labels, this year we are given various study-level labels. At first glance it appeared to me that, because the input of the study-level models are image embeddings, we need to use these study-level labels during image-level modeling so that the following study-level model could have sufficient knowledge to model and predict them. But after I tried lots of combinations of them and various loss masking tricks, the best performing  model in both the image-level and study-level stages was still the one trained with the image-level labels only. I’m a little puzzled how the image embeddings are encoded with the study-level labels, for example, the exact position labels (center, left, right) and the more refined acute and chronic, when the image-level models were not trained using any of those labels.\n\nThe training loss was the vanilla BCE loss with linear lr scheduler. No special data sampling was applied. I found that a single epoch through the train data was the optimal for my settings. The best augmentations were \n\n```\nalbumentations.RandomContrast(limit=0.2, p=1.0),\nalbumentations.ShiftScaleRotate(shift_limit=0.2, scale_limit=0.2, rotate_limit=20, border_mode=cv2.BORDER_CONSTANT, p=1.0),\nalbumentations.Cutout(num_holes=2, max_h_size=int(0.4*image_size), max_w_size=int(0.4*image_size), fill_value=0, always_apply=True, p=1.0),\n```\n\nMy final ensemble were with one serexnext50 and one seresnext101. Their respective validation performance for image-level PE prediction was\n\n```\n                     Loss      AUC\nseresnext101        0.079    0.964\nseresnext50         0.080    0.962\n```\n\nOther good backbones are inception_resnet_v2 and efficientnets. Densenets and resnexts performed a lot worse. Input were resized to 576x576 after the lung localization, this was the largest size the models could finish running in the 9 hours.\n\n# Study-level modeling\n\nImage embeddings of dimension 2048 served as the input to a RNN for both image-level and study-level modeling.  \n\nOne thing we need to handle was that the number of images each study has could vary from 100+ to 1000+. As we don’t know the information of the private test data, it was hard to predefine an input sequence length for our RNN model if we want to predict all the images. Stacking all the images into a 3-D array and resizing it along the z-axis before generating image embeddings is an option, but it was not compatible to my inference pipeline. For convenience, I swapped the order of embedding generation and resizing, in other words, I chose to resize the features instead of images. For example, given a study which has N images, the input feature shape is Nx2048. If the max sequence length limit in the RNN is M, the cv2.resize function is applied to resize features to Mx2048 if N>M, otherwise if N<M zero-padding is used. The image-level labels and the predictions are zoomed in and out in the same way during training and inference. To find the best M, I ran a search in the step size of 32, and M=128 gave the best performance. In the train set, the majority of Ns is in the range of 200-250. This means downsizing across the z-pos first before sequence modeling improves the performance. In my final models, I actually set m=192 because I believed there might be more big Ns in the private test data.\n\nInspired from last year's 2nd place solution, I also computed the difference of embeddings between current and the two direct neighbors and concatenate with the current features. So the input size was expanded to 2048x3. \n\nThe exact RNN architecture is not very important, I settled down to only a single bidirectional GRU layer, with the study-level labels predicted by a concatenated attention weighted average pooling and max pooling over the sequence. My local validation loss was around 0.18, I have no idea why it is much higher than the LB scores.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F478989%2F5aabec8cb08e9bc19b65b9efd4c0dcdd%2FPicture5.png?generation=1604080838336664&alt=media)\n\n# Postprocessing\n\nThe main purpose of the postprocessing step is to satisfy the consistency requirement of the labels. Since this consistency requirement agrees with how the data was labeled, a careful postprocessing could improve the performance. In my case, the local validation has a tiny improvement after postprocessing. The brief workflow of the postprocessing is \n\n```\nfor each study:\n    if the original predictions satisfy the consistency requirement\n        do nothing\n    else\n        change the original predictions into consistent positive predictions, and compute loss between them\n        change the original predictions into consistent negative predictions, and compute loss between them\n        choose from the positive and negative predictions based on which causes the smaller loss\n```\n\n\nThe weights of the loss function is almost same as the competition metric, except that the  q_i of image loss weight is replaced by a fixed 0.005 because we don’t have the ground truth of the test data. Code of this postprocessing can be found in my inference kernel.\n",
    "1065212": "Really waited for your solution. Thanks for sharing your details solution and code. Congrats on 1st place and 1 st global competitions ranking @wowfattie !",
    "1065229": "Congratulations Guanshuo. Brilliant and simple solution, well done! I like how you located the lungs, that was smart. My best models used 320x320 center crops from 512x512 which looks like it corresponds to your lung bbox.\n\n>  I actually set m=192 because I believed there might be more big Ns in the private test data.\n\nThis is good intuition. I probed the private LB. The public LB has 0% rows associtated with exams of more than 350 images. The private LB has 10% rows associated with exams of more than 350 images. (The train data had around 8%)",
    "1065117": "Congratulations on 1st place and competitions rank 1! Thanks for sharing your solution. \n\n> I’m a little puzzled how the image embeddings are encoded with the study-level labels, for example, the exact position labels (center, left, right) and the more refined acute and chronic, when the image-level models were not trained using any of those labels.\n\nThis is interesting- I'm also surprised. I used the study-level labels while training the image-level model because I thought they were necessary to encode those features as well, but apparently not. \n\nHow did you apply the bounding box predictions to the image? Did you use the largest box across all slices for each slice? Or individual boxes per slice? Some slices don't have lungs at all. ",
    "1065136": "Congratulations and thanks for sharing your solution! I also started off using lungmasks and then finding the largest bounding box for a given study but it wasn't feasible given the code requirements and I couldn't think of training a new model like you did unfortunately. Do you know how much improvement lung cropping gave in your final solution or was it mainly helpful for using computational resources efficiently? \n\nWe also noticed a huge gap (0.03) between our local CV (1456 studies for validation) vs the public/private LB, don't have an idea why it happens.\n\nResizing sequence input and then resizing the output back to map the original image number was something I was wondering if possible, thanks for sharing it is!\n\nPretty novel and clever postprocessing!\n\nWhat kind of attention mechanism is applied to GRU hidden outputs?",
    "3434703": "Congratulations Guanshuo. Brilliant and simple solution",
    "2456806": "this is a very detailed solution, thanks  ",
    "1851358": "Would you please provide the trained weights?",
    "1738699": "Hello, author, congratulations on winning the first prize. I tried to reproduce your code content, but the final effect failed to reach your AUC of 0.96. Do you know the reason?",
    "1124234": "Thank you  for sharing, and congratulation for 1st place.",
    "1090057": "Hi thank you for posting this awesome code! I'm trying to run the lung localizer and getting \"KeyErrors\" like there are missing file names in the lung localization lung_bbox.csv file\n\nAny idea what might be going wrong?",
    "1076765": "Well done on the achievement and thanks for this amazing breakdown and the accompanying github repo. I have one question (actually I would ask a hundred if I could 😜). Why do you use a max pool instead of average pool over the sequence dimension?\n\n```\nmax_pool, _ = torch.max(h_lstm1, 1)\n```\n\nI was actually working on another project and realised max pool worked better than average in a similar scenario. But would love to get your thoughts.",
    "1070691": "Nice work. Congrats on 1st place",
    "1070619": "Nice solution",
    "1069890": "Congrats! Very good work.",
    "1068578": "Amazing.Stuff\n",
    "1068313": "congratulations and thanks for sharing the solution. Thanks for the guidance.",
    "1068237": "Nice! Congratulations on the first place :)",
    "1067998": "Congratulations! Thanks for sharing your solution.",
    "1067049": ">  So, I annotated the train data and built a lung localizer with the bboxes and Efficientnet-b0 as the backbone.\n\nAwesome !! ",
    "1065177": "",
    "1738701": "Thank you for your reply!"
  }
}