{
  "id": 193506,
  "title": "Provisional 8th Place Solution - Monai x EfficientNets x LGBM",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193506",
  "author_name": "Jan Bre",
  "post_date": "2020-10-27T12:07:25.678000",
  "votes": 27,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Github: <a href=\"https://github.com/lisosia/kaggle-rsna-str\" target=\"_blank\">https://github.com/lisosia/kaggle-rsna-str</a></p>\n<h3>What a competition!</h3>\n<p>First of all Congratulations to everyone and especially to <strong>Yuji-san</strong> ( <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> ) and <strong>Yama-san</strong> <br>\n( <a href=\"https://www.kaggle.com/lisosia\" target=\"_blank\">@lisosia</a> ) for hopefully -fingers crossed- archiving <strong>Master status</strong>. <br>\nAnd a heartfelt Thank you to the organizers and the Kaggle Team for an awesome competition.</p>\n<h1>Solution Overview</h1>\n<p>In the big scope our solution is split into (1) image and (2) exam level predictions, which then are (3) ensembled in a Decision Tree. </p>\n<p>On an image level we predict Pe present on image, left,central and right. <br>\nOn the exam level we predict Left/Central/right, RV/LV Ratio and Acute/Chronic. <br>\nWe feed the predictions into multiple Decision Trees each fintuned with specific features, which then outputs the final predictions. <br>\nWe skipped predicting Indeterminate due to inference time purposes.</p>\n<p><img src=\"https://storage.cloud.google.com/kaggleimages/kagglersna.JPG?generation=1603760751579416&amp;alt=media\" alt=\"\"></p>\n<h2>PreProcess</h2>\n<p>Ian Pans ( <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> )dataset was a great start. But we very early on created a 512*512 image jpg dataset without jpg compression. With jpg compression our score was always worse.  For this we used Ian Pans Script for preprocessing with the same window sizes. </p>\n<h2>Image Level</h2>\n<p>In our final solution we used 2 Efficient Net models to predict  on image level.<br>\nThey were trained on all images to predict PE present on Image and the efficient Net b0 also to predict left/central/right . </p>\n<p>Important Edit: <br>\nOur Image level models predictions were transformed by calibrating the predicted probability of pe_present_on_image using:</p>\n<pre><code>def calib_p(arr, factor):  # set factor&gt;1 to enhance positive prob\n    return arr * factor / (arr * factor + (1-arr))\n</code></pre>\n<p>It is conducted to equalize each folds pe_present_on_image predictions before stacking with LGBM . The factor for each fold is determined so that the per-fold validation weighted-logloss is minimized. Yama-san came up with this idea and it boosted our image level predictionstogether with LGBM by a lot(see below).</p>\n<h2>Exam Level</h2>\n<p>Our pipeline here is very much based on the awesome kernel of <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> .<br>\nImages are center cropped and first and last 20% of images in the z axis removed and resized to 100<em>100</em>100 spatial size.  We now see some clever approaches for determining the Heart level (for example by Ian Pan).<br>\nThe 3D model is very bad at predicting positive_exams. Therefore rv_lv_ratio gte was combined to one target. <br>\nWe have trained the exam level model only on exams with PE. The main target for us initially was to predict RV/LV ratio but the model is also surprisingly good at predicting left/central/right and acute/chronic, which also gave us a boost for these features. </p>\n<h2>LGBM</h2>\n<h5>Image Level LGBM</h5>\n<p>For Image level Predictions we used a LGBM which for each image got the last 10 and the next 10 images as input to predict each image. This essentially simulates the CNN+LSTM method many competitors used. But in our case CNN + LGBM always outperformed CNN + LSTM.</p>\n<p>CV pe present on image: 0.12 (raw prediction ) -&gt; 0.105 (with LGBM + calibration)</p>\n<h5>Exam Level LGBM</h5>\n<p>We used the raw predictions of the Monai model and derived features from the image level models. These features mostly consist of percentiles [30, 50, 70, 80, 90, 95, 99] and number of images over a certain threshold  [0.1, 0.2 ,, ..0.9] , also  mean and max of image level models predictions. </p>\n<h2>Post Process.</h2>\n<ol>\n<li>We clipped right/left/center predictions to average(right) &lt; ave(pe_present)</li>\n<li>We simply set Inderminate to pos_exam * MEAN_WHEN_POS * (1-pos) * MEAN_WHEN_NOT_POS  as we had no time left in inference and it had a low weight assigned to it.</li>\n<li>Consistency Requirement: we have adjusted the exam level predictions to fit the consistency requirement with the lowest weighted difference (using the official metric weights).</li>\n</ol>\n<h2>Final Score</h2>\n<p>5 Fold CV:<br>\n(exam level unweighted)</p>\n<table>\n<thead>\n<tr>\n<th>Target</th>\n<th>LogLoss</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pe_present_on_image</td>\n<td>0.105</td>\n</tr>\n<tr>\n<td>rv_lv_ratio_gte_1</td>\n<td>0.231</td>\n</tr>\n<tr>\n<td>rv_lv_ratio_lt_1</td>\n<td>0.334</td>\n</tr>\n<tr>\n<td>leftsided_pe</td>\n<td>0.282</td>\n</tr>\n<tr>\n<td>central_pe</td>\n<td>0.114</td>\n</tr>\n<tr>\n<td>rightsided_pe</td>\n<td>0.285</td>\n</tr>\n<tr>\n<td>acute_and_chronic_pe</td>\n<td>0.087</td>\n</tr>\n<tr>\n<td>chronic_pe</td>\n<td>0.160</td>\n</tr>\n</tbody>\n</table>\n<h2>Things that did not work</h2>\n<p>We have tried many different architectures. Sequence models always yielded worse results for us compared to image level prediction ensembles with LGBM. I suppose this may be due to the very inconsistent number of images per exam compared to last years challenge. </p>\n<h2>Thank you!</h2>\n<p>2 days before submission deadline we did not have a final submission, which made things realy close to the end with Kaggle commits being slower to start and commit. I have never witnessed this issue before but it is good to keep this in mind for future competitions. Our final best submission finished just in time. </p>",
  "messages": [
    {
      "id": 1061885,
      "postDate": "2020-10-27T12:07:25.677Z",
      "content": "<p>Github: <a href=\"https://github.com/lisosia/kaggle-rsna-str\" target=\"_blank\">https://github.com/lisosia/kaggle-rsna-str</a></p>\n<h3>What a competition!</h3>\n<p>First of all Congratulations to everyone and especially to <strong>Yuji-san</strong> ( <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> ) and <strong>Yama-san</strong> <br>\n( <a href=\"https://www.kaggle.com/lisosia\" target=\"_blank\">@lisosia</a> ) for hopefully -fingers crossed- archiving <strong>Master status</strong>. <br>\nAnd a heartfelt Thank you to the organizers and the Kaggle Team for an awesome competition.</p>\n<h1>Solution Overview</h1>\n<p>In the big scope our solution is split into (1) image and (2) exam level predictions, which then are (3) ensembled in a Decision Tree. </p>\n<p>On an image level we predict Pe present on image, left,central and right. <br>\nOn the exam level we predict Left/Central/right, RV/LV Ratio and Acute/Chronic. <br>\nWe feed the predictions into multiple Decision Trees each fintuned with specific features, which then outputs the final predictions. <br>\nWe skipped predicting Indeterminate due to inference time purposes.</p>\n<p><img src=\"https://storage.cloud.google.com/kaggleimages/kagglersna.JPG?generation=1603760751579416&amp;alt=media\" alt=\"\"></p>\n<h2>PreProcess</h2>\n<p>Ian Pans ( <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> )dataset was a great start. But we very early on created a 512*512 image jpg dataset without jpg compression. With jpg compression our score was always worse.  For this we used Ian Pans Script for preprocessing with the same window sizes. </p>\n<h2>Image Level</h2>\n<p>In our final solution we used 2 Efficient Net models to predict  on image level.<br>\nThey were trained on all images to predict PE present on Image and the efficient Net b0 also to predict left/central/right . </p>\n<p>Important Edit: <br>\nOur Image level models predictions were transformed by calibrating the predicted probability of pe_present_on_image using:</p>\n<pre><code>def calib_p(arr, factor):  # set factor&gt;1 to enhance positive prob\n    return arr * factor / (arr * factor + (1-arr))\n</code></pre>\n<p>It is conducted to equalize each folds pe_present_on_image predictions before stacking with LGBM . The factor for each fold is determined so that the per-fold validation weighted-logloss is minimized. Yama-san came up with this idea and it boosted our image level predictionstogether with LGBM by a lot(see below).</p>\n<h2>Exam Level</h2>\n<p>Our pipeline here is very much based on the awesome kernel of <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> .<br>\nImages are center cropped and first and last 20% of images in the z axis removed and resized to 100<em>100</em>100 spatial size.  We now see some clever approaches for determining the Heart level (for example by Ian Pan).<br>\nThe 3D model is very bad at predicting positive_exams. Therefore rv_lv_ratio gte was combined to one target. <br>\nWe have trained the exam level model only on exams with PE. The main target for us initially was to predict RV/LV ratio but the model is also surprisingly good at predicting left/central/right and acute/chronic, which also gave us a boost for these features. </p>\n<h2>LGBM</h2>\n<h5>Image Level LGBM</h5>\n<p>For Image level Predictions we used a LGBM which for each image got the last 10 and the next 10 images as input to predict each image. This essentially simulates the CNN+LSTM method many competitors used. But in our case CNN + LGBM always outperformed CNN + LSTM.</p>\n<p>CV pe present on image: 0.12 (raw prediction ) -&gt; 0.105 (with LGBM + calibration)</p>\n<h5>Exam Level LGBM</h5>\n<p>We used the raw predictions of the Monai model and derived features from the image level models. These features mostly consist of percentiles [30, 50, 70, 80, 90, 95, 99] and number of images over a certain threshold  [0.1, 0.2 ,, ..0.9] , also  mean and max of image level models predictions. </p>\n<h2>Post Process.</h2>\n<ol>\n<li>We clipped right/left/center predictions to average(right) &lt; ave(pe_present)</li>\n<li>We simply set Inderminate to pos_exam * MEAN_WHEN_POS * (1-pos) * MEAN_WHEN_NOT_POS  as we had no time left in inference and it had a low weight assigned to it.</li>\n<li>Consistency Requirement: we have adjusted the exam level predictions to fit the consistency requirement with the lowest weighted difference (using the official metric weights).</li>\n</ol>\n<h2>Final Score</h2>\n<p>5 Fold CV:<br>\n(exam level unweighted)</p>\n<table>\n<thead>\n<tr>\n<th>Target</th>\n<th>LogLoss</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pe_present_on_image</td>\n<td>0.105</td>\n</tr>\n<tr>\n<td>rv_lv_ratio_gte_1</td>\n<td>0.231</td>\n</tr>\n<tr>\n<td>rv_lv_ratio_lt_1</td>\n<td>0.334</td>\n</tr>\n<tr>\n<td>leftsided_pe</td>\n<td>0.282</td>\n</tr>\n<tr>\n<td>central_pe</td>\n<td>0.114</td>\n</tr>\n<tr>\n<td>rightsided_pe</td>\n<td>0.285</td>\n</tr>\n<tr>\n<td>acute_and_chronic_pe</td>\n<td>0.087</td>\n</tr>\n<tr>\n<td>chronic_pe</td>\n<td>0.160</td>\n</tr>\n</tbody>\n</table>\n<h2>Things that did not work</h2>\n<p>We have tried many different architectures. Sequence models always yielded worse results for us compared to image level prediction ensembles with LGBM. I suppose this may be due to the very inconsistent number of images per exam compared to last years challenge. </p>\n<h2>Thank you!</h2>\n<p>2 days before submission deadline we did not have a final submission, which made things realy close to the end with Kaggle commits being slower to start and commit. I have never witnessed this issue before but it is good to keep this in mind for future competitions. Our final best submission finished just in time. </p>",
      "rawMarkdown": "Github: [https://github.com/lisosia/kaggle-rsna-str](https://github.com/lisosia/kaggle-rsna-str)\n\n### What a competition!\nFirst of all Congratulations to everyone and especially to **Yuji-san** ( @yujiariyasu ) and **Yama-san** \n( @lisosia ) for hopefully -fingers crossed- archiving **Master status**. \nAnd a heartfelt Thank you to the organizers and the Kaggle Team for an awesome competition.\n\n#Solution Overview\n\nIn the big scope our solution is split into (1) image and (2) exam level predictions, which then are (3) ensembled in a Decision Tree. \n\nOn an image level we predict Pe present on image, left,central and right. \nOn the exam level we predict Left/Central/right, RV/LV Ratio and Acute/Chronic. \nWe feed the predictions into multiple Decision Trees each fintuned with specific features, which then outputs the final predictions. \nWe skipped predicting Indeterminate due to inference time purposes.\n\n\n ![](https://storage.cloud.google.com/kaggleimages/kagglersna.JPG?generation=1603760751579416&alt=media)\n\n## PreProcess\nIan Pans ( @vaillant )dataset was a great start. But we very early on created a 512*512 image jpg dataset without jpg compression. With jpg compression our score was always worse.  For this we used Ian Pans Script for preprocessing with the same window sizes. \n\n## Image Level\n\nIn our final solution we used 2 Efficient Net models to predict  on image level.\nThey were trained on all images to predict PE present on Image and the efficient Net b0 also to predict left/central/right . \n\nImportant Edit: \nOur Image level models predictions were transformed by calibrating the predicted probability of pe_present_on_image using:\n\n```\ndef calib_p(arr, factor):  # set factor>1 to enhance positive prob\n    return arr * factor / (arr * factor + (1-arr))\n```\n\nIt is conducted to equalize each folds pe_present_on_image predictions before stacking with LGBM . The factor for each fold is determined so that the per-fold validation weighted-logloss is minimized. Yama-san came up with this idea and it boosted our image level predictionstogether with LGBM by a lot(see below).\n\n## Exam Level\n\nOur pipeline here is very much based on the awesome kernel of @boliu0 .\nImages are center cropped and first and last 20% of images in the z axis removed and resized to 100*100*100 spatial size.  We now see some clever approaches for determining the Heart level (for example by Ian Pan).\nThe 3D model is very bad at predicting positive_exams. Therefore rv_lv_ratio gte was combined to one target. \nWe have trained the exam level model only on exams with PE. The main target for us initially was to predict RV/LV ratio but the model is also surprisingly good at predicting left/central/right and acute/chronic, which also gave us a boost for these features. \n\n## LGBM \n#####Image Level LGBM\nFor Image level Predictions we used a LGBM which for each image got the last 10 and the next 10 images as input to predict each image. This essentially simulates the CNN+LSTM method many competitors used. But in our case CNN + LGBM always outperformed CNN + LSTM.\n\nCV pe present on image: 0.12 (raw prediction ) -> 0.105 (with LGBM + calibration)\n\n#####Exam Level LGBM\nWe used the raw predictions of the Monai model and derived features from the image level models. These features mostly consist of percentiles [30, 50, 70, 80, 90, 95, 99] and number of images over a certain threshold  [0.1, 0.2 ,, ..0.9] , also  mean and max of image level models predictions. \n\n## Post Process. \n1. We clipped right/left/center predictions to average(right) < ave(pe_present)\n2. We simply set Inderminate to pos_exam * MEAN_WHEN_POS * (1-pos) * MEAN_WHEN_NOT_POS  as we had no time left in inference and it had a low weight assigned to it.\n3. Consistency Requirement: we have adjusted the exam level predictions to fit the consistency requirement with the lowest weighted difference (using the official metric weights).\n\n##Final Score\n\n5 Fold CV:\n(exam level unweighted)\n\n| Target| LogLoss|\n| --- | --- |\n|  pe_present_on_image|   0.105|\n|  rv_lv_ratio_gte_1|  0.231| \n| rv_lv_ratio_lt_1 | 0.334 |\n| leftsided_pe| 0.282 | \n| central_pe| 0.114 | \n| rightsided_pe| 0.285 | \n| acute_and_chronic_pe | 0.087 | \n| chronic_pe | 0.160| \n\n\n\n##Things that did not work\n\nWe have tried many different architectures. Sequence models always yielded worse results for us compared to image level prediction ensembles with LGBM. I suppose this may be due to the very inconsistent number of images per exam compared to last years challenge. \n\n## Thank you!\n\n2 days before submission deadline we did not have a final submission, which made things realy close to the end with Kaggle commits being slower to start and commit. I have never witnessed this issue before but it is good to keep this in mind for future competitions. Our final best submission finished just in time. \n",
      "votes": 27
    },
    {
      "id": 1063382,
      "postDate": "2020-10-28T20:06:52.177Z",
      "content": "<p>Congratulations Jan <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> , Yuji-san <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> , Yama-san <a href=\"https://www.kaggle.com/lisosia\" target=\"_blank\">@lisosia</a> on building a great model and winning 8th place Gold. Congrats Yuji-san and   Yama-san for obtaining competition master status. </p>\n<p>This is a creative solution. I like your use of stacking with LGBM and your use of 3D model. Using <code>def calib_p(arr, factor)</code> is a great trick to adjust to this competition's sample weighted log loss for image level predictions.</p>\n<p>Great job!</p>",
      "rawMarkdown": "Congratulations Jan @jpbremer , Yuji-san @yujiariyasu , Yama-san @lisosia on building a great model and winning 8th place Gold. Congrats Yuji-san and   Yama-san for obtaining competition master status. \n\nThis is a creative solution. I like your use of stacking with LGBM and your use of 3D model. Using `def calib_p(arr, factor)` is a great trick to adjust to this competition's sample weighted log loss for image level predictions.\n\nGreat job!",
      "votes": 2,
      "replies": [
        {
          "id": 1063418,
          "postDate": "2020-10-28T22:13:33.643Z",
          "content": "<p>Thanks Chris<br>\nI learned so much from you since starting Kaggle. Thank you for your contributions to the platform!</p>",
          "rawMarkdown": "Thanks Chris\nI learned so much from you since starting Kaggle. Thank you for your contributions to the platform!",
          "votes": 1
        },
        {
          "id": 1063739,
          "postDate": "2020-10-29T09:01:26.113Z",
          "content": "<p>Thanks!<br>\nI'm so happy to have a great kaggler like you appreciate our solution. I look forward to seeing you at another competition.</p>",
          "rawMarkdown": "Thanks!\nI'm so happy to have a great kaggler like you appreciate our solution. I look forward to seeing you at another competition."
        }
      ]
    },
    {
      "id": 1062310,
      "postDate": "2020-10-27T18:15:22.547Z",
      "content": "<blockquote>\n  <p>2 days before submission deadline we did not have a final submission, which made things realy close to the end with Kaggle commits being slower to start and commit.</p>\n</blockquote>\n<p>Kaggle is way more exciting like this, isn't it ?</p>\n<p>Jokes aside, congratulations on the strong result !</p>",
      "rawMarkdown": "> 2 days before submission deadline we did not have a final submission, which made things realy close to the end with Kaggle commits being slower to start and commit.\n\nKaggle is way more exciting like this, isn't it ?\n\nJokes aside, congratulations on the strong result !",
      "replies": [
        {
          "id": 1062407,
          "postDate": "2020-10-27T19:45:43.313Z",
          "content": "<p>Thanks Theo. </p>\n<p>Yeah, Kaggle would be boring without the emotional roller coasters. :D</p>",
          "rawMarkdown": "Thanks Theo. \n\nYeah, Kaggle would be boring without the emotional roller coasters. :D",
          "votes": 1
        }
      ]
    },
    {
      "id": 1061917,
      "postDate": "2020-10-27T12:39:05.953Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> and team on 9th place. Thanks for sharing solution! </p>",
      "rawMarkdown": "Congrats @jpbremer and team on 9th place. Thanks for sharing solution! ",
      "replies": [
        {
          "id": 1062409,
          "postDate": "2020-10-27T19:46:05.700Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1063382,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-10-28T20:06:52.177000",
      "content": "<p>Congratulations Jan <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> , Yuji-san <a href=\"https://www.kaggle.com/yujiariyasu\" target=\"_blank\">@yujiariyasu</a> , Yama-san <a href=\"https://www.kaggle.com/lisosia\" target=\"_blank\">@lisosia</a> on building a great model and winning 8th place Gold. Congrats Yuji-san and   Yama-san for obtaining competition master status. </p>\n<p>This is a creative solution. I like your use of stacking with LGBM and your use of 3D model. Using <code>def calib_p(arr, factor)</code> is a great trick to adjust to this competition's sample weighted log loss for image level predictions.</p>\n<p>Great job!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1063418,
          "author_name": "Jan Bre",
          "author_url": "",
          "post_date": "2020-10-28T22:13:33.643000",
          "content": "<p>Thanks Chris<br>\nI learned so much from you since starting Kaggle. Thank you for your contributions to the platform!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1063739,
          "author_name": "YujiAriyasu",
          "author_url": "",
          "post_date": "2020-10-29T09:01:26.113000",
          "content": "<p>Thanks!<br>\nI'm so happy to have a great kaggler like you appreciate our solution. I look forward to seeing you at another competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1062310,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-10-27T18:15:22.547000",
      "content": "<blockquote>\n  <p>2 days before submission deadline we did not have a final submission, which made things realy close to the end with Kaggle commits being slower to start and commit.</p>\n</blockquote>\n<p>Kaggle is way more exciting like this, isn't it ?</p>\n<p>Jokes aside, congratulations on the strong result !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1062407,
          "author_name": "Jan Bre",
          "author_url": "",
          "post_date": "2020-10-27T19:45:43.313000",
          "content": "<p>Thanks Theo. </p>\n<p>Yeah, Kaggle would be boring without the emotional roller coasters. :D</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1061917,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-10-27T12:39:05.953000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> and team on 9th place. Thanks for sharing solution! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 1062409,
          "author_name": "Jan Bre",
          "author_url": "",
          "post_date": "2020-10-27T19:46:05.700000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1061885": "Github: [https://github.com/lisosia/kaggle-rsna-str](https://github.com/lisosia/kaggle-rsna-str)\n\n### What a competition!\nFirst of all Congratulations to everyone and especially to **Yuji-san** ( @yujiariyasu ) and **Yama-san** \n( @lisosia ) for hopefully -fingers crossed- archiving **Master status**. \nAnd a heartfelt Thank you to the organizers and the Kaggle Team for an awesome competition.\n\n#Solution Overview\n\nIn the big scope our solution is split into (1) image and (2) exam level predictions, which then are (3) ensembled in a Decision Tree. \n\nOn an image level we predict Pe present on image, left,central and right. \nOn the exam level we predict Left/Central/right, RV/LV Ratio and Acute/Chronic. \nWe feed the predictions into multiple Decision Trees each fintuned with specific features, which then outputs the final predictions. \nWe skipped predicting Indeterminate due to inference time purposes.\n\n\n ![](https://storage.cloud.google.com/kaggleimages/kagglersna.JPG?generation=1603760751579416&alt=media)\n\n## PreProcess\nIan Pans ( @vaillant )dataset was a great start. But we very early on created a 512*512 image jpg dataset without jpg compression. With jpg compression our score was always worse.  For this we used Ian Pans Script for preprocessing with the same window sizes. \n\n## Image Level\n\nIn our final solution we used 2 Efficient Net models to predict  on image level.\nThey were trained on all images to predict PE present on Image and the efficient Net b0 also to predict left/central/right . \n\nImportant Edit: \nOur Image level models predictions were transformed by calibrating the predicted probability of pe_present_on_image using:\n\n```\ndef calib_p(arr, factor):  # set factor>1 to enhance positive prob\n    return arr * factor / (arr * factor + (1-arr))\n```\n\nIt is conducted to equalize each folds pe_present_on_image predictions before stacking with LGBM . The factor for each fold is determined so that the per-fold validation weighted-logloss is minimized. Yama-san came up with this idea and it boosted our image level predictionstogether with LGBM by a lot(see below).\n\n## Exam Level\n\nOur pipeline here is very much based on the awesome kernel of @boliu0 .\nImages are center cropped and first and last 20% of images in the z axis removed and resized to 100*100*100 spatial size.  We now see some clever approaches for determining the Heart level (for example by Ian Pan).\nThe 3D model is very bad at predicting positive_exams. Therefore rv_lv_ratio gte was combined to one target. \nWe have trained the exam level model only on exams with PE. The main target for us initially was to predict RV/LV ratio but the model is also surprisingly good at predicting left/central/right and acute/chronic, which also gave us a boost for these features. \n\n## LGBM \n#####Image Level LGBM\nFor Image level Predictions we used a LGBM which for each image got the last 10 and the next 10 images as input to predict each image. This essentially simulates the CNN+LSTM method many competitors used. But in our case CNN + LGBM always outperformed CNN + LSTM.\n\nCV pe present on image: 0.12 (raw prediction ) -> 0.105 (with LGBM + calibration)\n\n#####Exam Level LGBM\nWe used the raw predictions of the Monai model and derived features from the image level models. These features mostly consist of percentiles [30, 50, 70, 80, 90, 95, 99] and number of images over a certain threshold  [0.1, 0.2 ,, ..0.9] , also  mean and max of image level models predictions. \n\n## Post Process. \n1. We clipped right/left/center predictions to average(right) < ave(pe_present)\n2. We simply set Inderminate to pos_exam * MEAN_WHEN_POS * (1-pos) * MEAN_WHEN_NOT_POS  as we had no time left in inference and it had a low weight assigned to it.\n3. Consistency Requirement: we have adjusted the exam level predictions to fit the consistency requirement with the lowest weighted difference (using the official metric weights).\n\n##Final Score\n\n5 Fold CV:\n(exam level unweighted)\n\n| Target| LogLoss|\n| --- | --- |\n|  pe_present_on_image|   0.105|\n|  rv_lv_ratio_gte_1|  0.231| \n| rv_lv_ratio_lt_1 | 0.334 |\n| leftsided_pe| 0.282 | \n| central_pe| 0.114 | \n| rightsided_pe| 0.285 | \n| acute_and_chronic_pe | 0.087 | \n| chronic_pe | 0.160| \n\n\n\n##Things that did not work\n\nWe have tried many different architectures. Sequence models always yielded worse results for us compared to image level prediction ensembles with LGBM. I suppose this may be due to the very inconsistent number of images per exam compared to last years challenge. \n\n## Thank you!\n\n2 days before submission deadline we did not have a final submission, which made things realy close to the end with Kaggle commits being slower to start and commit. I have never witnessed this issue before but it is good to keep this in mind for future competitions. Our final best submission finished just in time. \n",
    "1063382": "Congratulations Jan @jpbremer , Yuji-san @yujiariyasu , Yama-san @lisosia on building a great model and winning 8th place Gold. Congrats Yuji-san and   Yama-san for obtaining competition master status. \n\nThis is a creative solution. I like your use of stacking with LGBM and your use of 3D model. Using `def calib_p(arr, factor)` is a great trick to adjust to this competition's sample weighted log loss for image level predictions.\n\nGreat job!",
    "1062310": "> 2 days before submission deadline we did not have a final submission, which made things realy close to the end with Kaggle commits being slower to start and commit.\n\nKaggle is way more exciting like this, isn't it ?\n\nJokes aside, congratulations on the strong result !",
    "1061917": "Congrats @jpbremer and team on 9th place. Thanks for sharing solution! "
  }
}