{
  "id": 193970,
  "title": "4th place solution",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193970",
  "author_name": "kazumax",
  "post_date": "2020-10-29T23:49:08.355000",
  "votes": 45,
  "comment_count": 5,
  "views": 0,
  "content": "<p>First of all, Thank you to the organizers and Congrats to all the participants and winners !</p>\n<p>My solution is very simple.\n・train CNN backbone (Stage-1)\n・extract embedding\n・train LSTMs (Stage-2)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2F37a328ab2727b4b5c02b459c978f80fb%2FRSNA2020%20solution.png?generation=1604015018732684&amp;alt=media\" alt=\"\"></p>\n<p><strong>Stage-1:</strong>\n　Split 3fold and train CNN which predict pre_present_on_image label for each images.\nAs a result of exploring several backbones, I decided to use rexnet200. Efficientnet-b4 and b5 got nan loss with mixed precision so I gave up using.\nFull-size(512x512) jpeg images preprocessed by Ian Pan’s windowing are fed into CNN.\nFor augmentation, I used the following:\n ・Horizontal Flip\n ・ShiftScaleRotate\n ・One of(Cutout, GridDopout)\nSince the ratio of pe present on image was quite small, I also use focal loss.\nIn Stage-1, I created 6 models. (3-fold BCE and 3-fold Focal) </p>\n<p><strong>Stage-2:</strong>\n　The embedding was dumped and sorted in z-axis order and stored in a disk, and then LSTMs were trained using it. I implemented a mini-batch training of variable length series since the number of images for each study is different. Each series is filled with invalid values until the maximum series length (1083) of the training data is reached, and the invalid part is ignored when calculating the loss.\n　For each study, embeddings were put into BiLSTMx2 and each LSTM’s outputs went into two branches: one to predict pe_present_on_image and the other to predict the exam level label. In pe_present_on_image branch, the two outputs were simply added and transformed by Linear Layer. In exam_level branch,  features were aggregated using attention layer and  transformed by Linear Layer.\nFor Loss, I used simple BCE.</p>\n<p><strong>Inference:</strong>\nI implemented a model that connects backbone and LSTMs for inference without dumping embedding to disk. Inference is done by batch_size=1. There were not enough time to apply TTA.</p>\n<p><strong>Submit model:</strong>\n・6-model average ensemble. \n　3-fold with backbone trained by BCE Loss + 3-fold with backbone trained by Focal Loss.\n　public LB: 0.159 private LB: 0.152\n・3-model average ensemble.\n　3-fold with backbone trained by BCE Loss\n　public LB: 0.158 private LB: 0.153</p>\n<p><strong>Thank you !</strong>\nThis is my first gold medal. I’m super glad to finally be a kaggle master!\nAnd this also is  my first time of posting discussion. \nI'm sorry if my English is not good enough to convey my solution.</p>\n<p>The code link will appeare here once I clean it up.</p>\n<p>Update:\n<a href=\"https://github.com/piwafp0720/RSNA-STR-Pulmonary-Embolism-Detection\" target=\"_blank\">repository</a>\n<a href=\"https://www.kaggle.com/kazumax0720/rsna-str-pulmonary-embolism-detection-inference\" target=\"_blank\">inference kernel</a></p>",
  "messages": [
    {
      "id": 1064289,
      "postDate": "2020-10-29T23:49:08.357Z",
      "content": "<p>First of all, Thank you to the organizers and Congrats to all the participants and winners !</p>\n<p>My solution is very simple.\n・train CNN backbone (Stage-1)\n・extract embedding\n・train LSTMs (Stage-2)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2F37a328ab2727b4b5c02b459c978f80fb%2FRSNA2020%20solution.png?generation=1604015018732684&amp;alt=media\" alt=\"\"></p>\n<p><strong>Stage-1:</strong>\n　Split 3fold and train CNN which predict pre_present_on_image label for each images.\nAs a result of exploring several backbones, I decided to use rexnet200. Efficientnet-b4 and b5 got nan loss with mixed precision so I gave up using.\nFull-size(512x512) jpeg images preprocessed by Ian Pan’s windowing are fed into CNN.\nFor augmentation, I used the following:\n ・Horizontal Flip\n ・ShiftScaleRotate\n ・One of(Cutout, GridDopout)\nSince the ratio of pe present on image was quite small, I also use focal loss.\nIn Stage-1, I created 6 models. (3-fold BCE and 3-fold Focal) </p>\n<p><strong>Stage-2:</strong>\n　The embedding was dumped and sorted in z-axis order and stored in a disk, and then LSTMs were trained using it. I implemented a mini-batch training of variable length series since the number of images for each study is different. Each series is filled with invalid values until the maximum series length (1083) of the training data is reached, and the invalid part is ignored when calculating the loss.\n　For each study, embeddings were put into BiLSTMx2 and each LSTM’s outputs went into two branches: one to predict pe_present_on_image and the other to predict the exam level label. In pe_present_on_image branch, the two outputs were simply added and transformed by Linear Layer. In exam_level branch,  features were aggregated using attention layer and  transformed by Linear Layer.\nFor Loss, I used simple BCE.</p>\n<p><strong>Inference:</strong>\nI implemented a model that connects backbone and LSTMs for inference without dumping embedding to disk. Inference is done by batch_size=1. There were not enough time to apply TTA.</p>\n<p><strong>Submit model:</strong>\n・6-model average ensemble. \n　3-fold with backbone trained by BCE Loss + 3-fold with backbone trained by Focal Loss.\n　public LB: 0.159 private LB: 0.152\n・3-model average ensemble.\n　3-fold with backbone trained by BCE Loss\n　public LB: 0.158 private LB: 0.153</p>\n<p><strong>Thank you !</strong>\nThis is my first gold medal. I’m super glad to finally be a kaggle master!\nAnd this also is  my first time of posting discussion. \nI'm sorry if my English is not good enough to convey my solution.</p>\n<p>The code link will appeare here once I clean it up.</p>\n<p>Update:\n<a href=\"https://github.com/piwafp0720/RSNA-STR-Pulmonary-Embolism-Detection\" target=\"_blank\">repository</a>\n<a href=\"https://www.kaggle.com/kazumax0720/rsna-str-pulmonary-embolism-detection-inference\" target=\"_blank\">inference kernel</a></p>",
      "rawMarkdown": "First of all, Thank you to the organizers and Congrats to all the participants and winners !\n\nMy solution is very simple.\n・train CNN backbone (Stage-1)\n・extract embedding\n・train LSTMs (Stage-2)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2F37a328ab2727b4b5c02b459c978f80fb%2FRSNA2020%20solution.png?generation=1604015018732684&alt=media)\n\n**Stage-1:**\n　Split 3fold and train CNN which predict pre_present_on_image label for each images.\nAs a result of exploring several backbones, I decided to use rexnet200. Efficientnet-b4 and b5 got nan loss with mixed precision so I gave up using.\nFull-size(512x512) jpeg images preprocessed by Ian Pan’s windowing are fed into CNN.\nFor augmentation, I used the following:\n ・Horizontal Flip\n ・ShiftScaleRotate\n ・One of(Cutout, GridDopout)\nSince the ratio of pe present on image was quite small, I also use focal loss.\nIn Stage-1, I created 6 models. (3-fold BCE and 3-fold Focal) \n\n**Stage-2:**\n　The embedding was dumped and sorted in z-axis order and stored in a disk, and then LSTMs were trained using it. I implemented a mini-batch training of variable length series since the number of images for each study is different. Each series is filled with invalid values until the maximum series length (1083) of the training data is reached, and the invalid part is ignored when calculating the loss.\n　For each study, embeddings were put into BiLSTMx2 and each LSTM’s outputs went into two branches: one to predict pe_present_on_image and the other to predict the exam level label. In pe_present_on_image branch, the two outputs were simply added and transformed by Linear Layer. In exam_level branch,  features were aggregated using attention layer and  transformed by Linear Layer.\nFor Loss, I used simple BCE.\n\n**Inference:**\nI implemented a model that connects backbone and LSTMs for inference without dumping embedding to disk. Inference is done by batch_size=1. There were not enough time to apply TTA.\n\n**Submit model:**\n・6-model average ensemble. \n　3-fold with backbone trained by BCE Loss + 3-fold with backbone trained by Focal Loss.\n　public LB: 0.159 private LB: 0.152\n・3-model average ensemble.\n　3-fold with backbone trained by BCE Loss\n　public LB: 0.158 private LB: 0.153\n\n**Thank you !**\nThis is my first gold medal. I’m super glad to finally be a kaggle master!\nAnd this also is  my first time of posting discussion. \nI'm sorry if my English is not good enough to convey my solution.\n\nThe code link will appeare here once I clean it up.\n\nUpdate:\n[repository](https://github.com/piwafp0720/RSNA-STR-Pulmonary-Embolism-Detection)\n[inference kernel](https://www.kaggle.com/kazumax0720/rsna-str-pulmonary-embolism-detection-inference)",
      "votes": 45
    },
    {
      "id": 1065037,
      "postDate": "2020-10-30T19:24:03.197Z",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kazumax0720\" target=\"_blank\">@kazumax0720</a> thanks for the nice write up. I gave up on resnet type models as I thought they would be too heavy on memory at inference. Will know for next time 😁 looking forward to you solution to see how you implement focal loss. </p>",
      "rawMarkdown": "Great work @kazumax0720 thanks for the nice write up. I gave up on resnet type models as I thought they would be too heavy on memory at inference. Will know for next time 😁 looking forward to you solution to see how you implement focal loss. ",
      "votes": 1,
      "replies": [
        {
          "id": 1065919,
          "postDate": "2020-11-01T03:09:26.150Z",
          "content": "<p>for focal loss, I used this <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/65938\" target=\"_blank\">discussion</a>. Kaggle community helps me learn a lot. thanks !</p>",
          "rawMarkdown": "for focal loss, I used this [discussion](https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/65938). Kaggle community helps me learn a lot. thanks !",
          "votes": 1
        }
      ]
    },
    {
      "id": 1066161,
      "postDate": "2020-11-01T12:11:47.510Z",
      "content": "<p>Awesome! I was actually looking for an explained tutorial on how to combine CNNs and LSTMs efficiently, but none of which I previously found actually had such detailed code snippets. <br>\nThank you for your work!</p>",
      "rawMarkdown": "Awesome! I was actually looking for an explained tutorial on how to combine CNNs and LSTMs efficiently, but none of which I previously found actually had such detailed code snippets. \nThank you for your work!",
      "votes": 2
    },
    {
      "id": 1065202,
      "postDate": "2020-10-31T02:51:45.553Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 1065921,
          "postDate": "2020-11-01T03:10:28.023Z",
          "content": "<p>I used NVIDIA DGX station.<br>\nFor submission, please check my inference kernel.</p>",
          "rawMarkdown": "I used NVIDIA DGX station.\nFor submission, please check my inference kernel."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1065037,
      "author_name": "Darragh",
      "author_url": "",
      "post_date": "2020-10-30T19:24:03.197000",
      "content": "<p>Great work <a href=\"https://www.kaggle.com/kazumax0720\" target=\"_blank\">@kazumax0720</a> thanks for the nice write up. I gave up on resnet type models as I thought they would be too heavy on memory at inference. Will know for next time 😁 looking forward to you solution to see how you implement focal loss. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1065919,
          "author_name": "kazumax",
          "author_url": "",
          "post_date": "2020-11-01T03:09:26.150000",
          "content": "<p>for focal loss, I used this <a href=\"https://www.kaggle.com/c/tgs-salt-identification-challenge/discussion/65938\" target=\"_blank\">discussion</a>. Kaggle community helps me learn a lot. thanks !</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1066161,
      "author_name": "Grigore Lucian",
      "author_url": "",
      "post_date": "2020-11-01T12:11:47.510000",
      "content": "<p>Awesome! I was actually looking for an explained tutorial on how to combine CNNs and LSTMs efficiently, but none of which I previously found actually had such detailed code snippets. <br>\nThank you for your work!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1065202,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-31T02:51:45.553000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1065921,
          "author_name": "kazumax",
          "author_url": "",
          "post_date": "2020-11-01T03:10:28.023000",
          "content": "<p>I used NVIDIA DGX station.<br>\nFor submission, please check my inference kernel.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1064289": "First of all, Thank you to the organizers and Congrats to all the participants and winners !\n\nMy solution is very simple.\n・train CNN backbone (Stage-1)\n・extract embedding\n・train LSTMs (Stage-2)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2422331%2F37a328ab2727b4b5c02b459c978f80fb%2FRSNA2020%20solution.png?generation=1604015018732684&alt=media)\n\n**Stage-1:**\n　Split 3fold and train CNN which predict pre_present_on_image label for each images.\nAs a result of exploring several backbones, I decided to use rexnet200. Efficientnet-b4 and b5 got nan loss with mixed precision so I gave up using.\nFull-size(512x512) jpeg images preprocessed by Ian Pan’s windowing are fed into CNN.\nFor augmentation, I used the following:\n ・Horizontal Flip\n ・ShiftScaleRotate\n ・One of(Cutout, GridDopout)\nSince the ratio of pe present on image was quite small, I also use focal loss.\nIn Stage-1, I created 6 models. (3-fold BCE and 3-fold Focal) \n\n**Stage-2:**\n　The embedding was dumped and sorted in z-axis order and stored in a disk, and then LSTMs were trained using it. I implemented a mini-batch training of variable length series since the number of images for each study is different. Each series is filled with invalid values until the maximum series length (1083) of the training data is reached, and the invalid part is ignored when calculating the loss.\n　For each study, embeddings were put into BiLSTMx2 and each LSTM’s outputs went into two branches: one to predict pe_present_on_image and the other to predict the exam level label. In pe_present_on_image branch, the two outputs were simply added and transformed by Linear Layer. In exam_level branch,  features were aggregated using attention layer and  transformed by Linear Layer.\nFor Loss, I used simple BCE.\n\n**Inference:**\nI implemented a model that connects backbone and LSTMs for inference without dumping embedding to disk. Inference is done by batch_size=1. There were not enough time to apply TTA.\n\n**Submit model:**\n・6-model average ensemble. \n　3-fold with backbone trained by BCE Loss + 3-fold with backbone trained by Focal Loss.\n　public LB: 0.159 private LB: 0.152\n・3-model average ensemble.\n　3-fold with backbone trained by BCE Loss\n　public LB: 0.158 private LB: 0.153\n\n**Thank you !**\nThis is my first gold medal. I’m super glad to finally be a kaggle master!\nAnd this also is  my first time of posting discussion. \nI'm sorry if my English is not good enough to convey my solution.\n\nThe code link will appeare here once I clean it up.\n\nUpdate:\n[repository](https://github.com/piwafp0720/RSNA-STR-Pulmonary-Embolism-Detection)\n[inference kernel](https://www.kaggle.com/kazumax0720/rsna-str-pulmonary-embolism-detection-inference)",
    "1065037": "Great work @kazumax0720 thanks for the nice write up. I gave up on resnet type models as I thought they would be too heavy on memory at inference. Will know for next time 😁 looking forward to you solution to see how you implement focal loss. ",
    "1066161": "Awesome! I was actually looking for an explained tutorial on how to combine CNNs and LSTMs efficiently, but none of which I previously found actually had such detailed code snippets. \nThank you for your work!",
    "1065202": ""
  }
}