{
  "id": 193475,
  "title": "5th Place Solution Overview",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193475",
  "author_name": "Darragh",
  "post_date": "2020-10-27T08:17:13.473000",
  "votes": 52,
  "comment_count": 11,
  "views": 0,
  "content": "<p><img src=\"https://raw.githubusercontent.com/darraghdog/rsnastr/main/docs/architecture.jpg\" alt=\"\"><br>\n<strong>Code base</strong> : <a href=\"https://github.com/darraghdog/rsnastr\" target=\"_blank\">https://github.com/darraghdog/rsnastr</a> <br>\n<strong>Kaggle submission</strong> : <a href=\"https://www.kaggle.com/darraghdog/rsnastr2020-prediction\" target=\"_blank\">https://www.kaggle.com/darraghdog/rsnastr2020-prediction</a></p>\n<p>Thanks to all the organisers for hosting this competition. And congratulations to all the participants and winners - these competitions and solutions are getting more and more advanced 😁  </p>\n<p>My compute was a v100 on AWS initially with 320X320 images, with prototyping on Macbook (a lot can be checked on CPU). Then in the last 3 weeks I moved to an A100 to run 512X512 (thanks DoubleYard) - 40GB GPU memory per card - one card was enough. <br>\nOverall the solution pipeline ended up being pretty similar to last year, but the journey to get there was a lot different. The metric used by the competition was difficult to simulate, and all parts of the pipeline needed to be finalised/integrated to get feedback on the competition benchmark.  </p>\n<p><strong>Preprocessing</strong><br>\nUse the <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/preprocessing/dicom_to_jpeg.py#L59-L75\" target=\"_blank\">windowing</a> from Ian Pan's <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182930\" target=\"_blank\">post</a> pretty much as is, just <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/preprocessing/dicom_to_jpeg.py#L77-L118\" target=\"_blank\">parallelised</a> it to speed it up. Some dicom's fell out which I saw later was due to the way pydicom was used, but I think the majority were good. <br>\nLight <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L72-L74\" target=\"_blank\">augmentations</a>, maybe more would have helped but did not want to lose information which side the PE lay on. </p>\n<p><strong>Image Level</strong><br>\nEfficientb5 seemed to be better than anything else I tried both in terms of speed and loss. Given the time and compute limits on submission, speed was important. <br>\nThe datasampler took at least <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L112-L117\" target=\"_blank\">two images from each study, in each epoch, and positive images were oversampled</a> to give a rate of 4:1 -ve:+ve, with image batchsize of 48 (using half point precision - amp). <br>\n<a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L132-L138\" target=\"_blank\">Weighted study loss</a> was calculated at each step on the positive samples only (weights as per competition metric weights excl. <code>negative_exam_for_pe</code>), image loss was calculated on all samples. Final loss per step was <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L200-L209\" target=\"_blank\">sum of both averages</a>; image loss was tuned so that image loss and study loss would be roughly equal. <br>\nEach epoch took about 5 mins and ran ~15 epochs, final solution used three of five folds.</p>\n<p><strong>Study Level</strong><br>\nExtracted gap layer of each image to disk. These were fed into <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_sequence_classifier.py#L113-L124\" target=\"_blank\">three independent sequence models</a>. <br>\n1) 2X Bi-LSTM, more or less <a href=\"https://github.com/darraghdog/rsna\" target=\"_blank\">same architecture</a> as last year.<br>\n2) Bert transformer model with 1 layer. <br>\n2) Bert transformer model with 2 layers. <br>\nEach of these used the same loss <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_sequence_classifier.py#L150-L170\" target=\"_blank\">(or at least as close as I could get)</a> as the competition metric on each step, study batchsize 64, no accumulation (again with amp).<br>\nFinally both models and all folds were averaged. <br>\nI submitted a few more transformer models on last day which I had high hopes for but most broke the consistency check - I only allowed submission, if it passed - and then time ran out. (Thanks <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> for the tip to only drop to disk if it passed). Actually I think 7 of my last 10 submissions failed. </p>\n<p><strong>Inference</strong><br>\nA few tricks to help speed up and bring RAM and GPU memory down - <a href=\"https://www.kaggle.com/darraghdog/rsna512-effnetb5-fold-all-exam-xfrmr-validated?scriptVersionId=45523025\" target=\"_blank\">final sub</a>.<br>\nBatchsize of 1. Feed sequences of images into image level model in chunks of 128 images, but only normalize (uint8-&gt;float32) when that chunk is to be passed into image model. Then concat all output (GAP layer) chunks to make the sequence.  <br>\nConvert all models and input data to half point precision <code>model=model.half()</code> - this halved the memory, but results were just as good. <br>\nWith the above, I managed to squeeze three folds out of the above pipeline - four folds also worked, but it failed on consistency check on last day. Interested to see optimisations from other competitiors. </p>\n<p>A final note, about two weeks ago, I was stuck on CV could not get the metric working, and decided to kick the competition. Luckily I stuck with it to try a first submission which got in at LB 0.181 single fold :) For such a challenging metric and large dataset the timeline was tight, I think with another couple of weeks the community could have got to 0.12x or so on the leaderboard - but maybe that would overfit the dataset.   </p>",
  "messages": [
    {
      "id": 1061699,
      "postDate": "2020-10-27T08:17:13.473Z",
      "content": "<p><img src=\"https://raw.githubusercontent.com/darraghdog/rsnastr/main/docs/architecture.jpg\" alt=\"\"><br>\n<strong>Code base</strong> : <a href=\"https://github.com/darraghdog/rsnastr\" target=\"_blank\">https://github.com/darraghdog/rsnastr</a> <br>\n<strong>Kaggle submission</strong> : <a href=\"https://www.kaggle.com/darraghdog/rsnastr2020-prediction\" target=\"_blank\">https://www.kaggle.com/darraghdog/rsnastr2020-prediction</a></p>\n<p>Thanks to all the organisers for hosting this competition. And congratulations to all the participants and winners - these competitions and solutions are getting more and more advanced 😁  </p>\n<p>My compute was a v100 on AWS initially with 320X320 images, with prototyping on Macbook (a lot can be checked on CPU). Then in the last 3 weeks I moved to an A100 to run 512X512 (thanks DoubleYard) - 40GB GPU memory per card - one card was enough. <br>\nOverall the solution pipeline ended up being pretty similar to last year, but the journey to get there was a lot different. The metric used by the competition was difficult to simulate, and all parts of the pipeline needed to be finalised/integrated to get feedback on the competition benchmark.  </p>\n<p><strong>Preprocessing</strong><br>\nUse the <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/preprocessing/dicom_to_jpeg.py#L59-L75\" target=\"_blank\">windowing</a> from Ian Pan's <a href=\"https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182930\" target=\"_blank\">post</a> pretty much as is, just <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/preprocessing/dicom_to_jpeg.py#L77-L118\" target=\"_blank\">parallelised</a> it to speed it up. Some dicom's fell out which I saw later was due to the way pydicom was used, but I think the majority were good. <br>\nLight <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L72-L74\" target=\"_blank\">augmentations</a>, maybe more would have helped but did not want to lose information which side the PE lay on. </p>\n<p><strong>Image Level</strong><br>\nEfficientb5 seemed to be better than anything else I tried both in terms of speed and loss. Given the time and compute limits on submission, speed was important. <br>\nThe datasampler took at least <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L112-L117\" target=\"_blank\">two images from each study, in each epoch, and positive images were oversampled</a> to give a rate of 4:1 -ve:+ve, with image batchsize of 48 (using half point precision - amp). <br>\n<a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L132-L138\" target=\"_blank\">Weighted study loss</a> was calculated at each step on the positive samples only (weights as per competition metric weights excl. <code>negative_exam_for_pe</code>), image loss was calculated on all samples. Final loss per step was <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L200-L209\" target=\"_blank\">sum of both averages</a>; image loss was tuned so that image loss and study loss would be roughly equal. <br>\nEach epoch took about 5 mins and ran ~15 epochs, final solution used three of five folds.</p>\n<p><strong>Study Level</strong><br>\nExtracted gap layer of each image to disk. These were fed into <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_sequence_classifier.py#L113-L124\" target=\"_blank\">three independent sequence models</a>. <br>\n1) 2X Bi-LSTM, more or less <a href=\"https://github.com/darraghdog/rsna\" target=\"_blank\">same architecture</a> as last year.<br>\n2) Bert transformer model with 1 layer. <br>\n2) Bert transformer model with 2 layers. <br>\nEach of these used the same loss <a href=\"https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_sequence_classifier.py#L150-L170\" target=\"_blank\">(or at least as close as I could get)</a> as the competition metric on each step, study batchsize 64, no accumulation (again with amp).<br>\nFinally both models and all folds were averaged. <br>\nI submitted a few more transformer models on last day which I had high hopes for but most broke the consistency check - I only allowed submission, if it passed - and then time ran out. (Thanks <a href=\"https://www.kaggle.com/yuval6967\" target=\"_blank\">@yuval6967</a> for the tip to only drop to disk if it passed). Actually I think 7 of my last 10 submissions failed. </p>\n<p><strong>Inference</strong><br>\nA few tricks to help speed up and bring RAM and GPU memory down - <a href=\"https://www.kaggle.com/darraghdog/rsna512-effnetb5-fold-all-exam-xfrmr-validated?scriptVersionId=45523025\" target=\"_blank\">final sub</a>.<br>\nBatchsize of 1. Feed sequences of images into image level model in chunks of 128 images, but only normalize (uint8-&gt;float32) when that chunk is to be passed into image model. Then concat all output (GAP layer) chunks to make the sequence.  <br>\nConvert all models and input data to half point precision <code>model=model.half()</code> - this halved the memory, but results were just as good. <br>\nWith the above, I managed to squeeze three folds out of the above pipeline - four folds also worked, but it failed on consistency check on last day. Interested to see optimisations from other competitiors. </p>\n<p>A final note, about two weeks ago, I was stuck on CV could not get the metric working, and decided to kick the competition. Luckily I stuck with it to try a first submission which got in at LB 0.181 single fold :) For such a challenging metric and large dataset the timeline was tight, I think with another couple of weeks the community could have got to 0.12x or so on the leaderboard - but maybe that would overfit the dataset.   </p>",
      "rawMarkdown": "![](https://raw.githubusercontent.com/darraghdog/rsnastr/main/docs/architecture.jpg)\n**Code base** : https://github.com/darraghdog/rsnastr \n**Kaggle submission** : https://www.kaggle.com/darraghdog/rsnastr2020-prediction\n\nThanks to all the organisers for hosting this competition. And congratulations to all the participants and winners - these competitions and solutions are getting more and more advanced 😁  \n\nMy compute was a v100 on AWS initially with 320X320 images, with prototyping on Macbook (a lot can be checked on CPU). Then in the last 3 weeks I moved to an A100 to run 512X512 (thanks DoubleYard) - 40GB GPU memory per card - one card was enough. \nOverall the solution pipeline ended up being pretty similar to last year, but the journey to get there was a lot different. The metric used by the competition was difficult to simulate, and all parts of the pipeline needed to be finalised/integrated to get feedback on the competition benchmark.  \n\n**Preprocessing**\nUse the [windowing](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/preprocessing/dicom_to_jpeg.py#L59-L75) from Ian Pan's [post](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182930) pretty much as is, just [parallelised](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/preprocessing/dicom_to_jpeg.py#L77-L118) it to speed it up. Some dicom's fell out which I saw later was due to the way pydicom was used, but I think the majority were good. \nLight [augmentations](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L72-L74), maybe more would have helped but did not want to lose information which side the PE lay on. \n\n**Image Level**\nEfficientb5 seemed to be better than anything else I tried both in terms of speed and loss. Given the time and compute limits on submission, speed was important. \nThe datasampler took at least [two images from each study, in each epoch, and positive images were oversampled](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L112-L117) to give a rate of 4:1 -ve:+ve, with image batchsize of 48 (using half point precision - amp). \n[Weighted study loss](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L132-L138) was calculated at each step on the positive samples only (weights as per competition metric weights excl. `negative_exam_for_pe`), image loss was calculated on all samples. Final loss per step was [sum of both averages](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L200-L209); image loss was tuned so that image loss and study loss would be roughly equal. \nEach epoch took about 5 mins and ran ~15 epochs, final solution used three of five folds.\n\n**Study Level**\nExtracted gap layer of each image to disk. These were fed into [three independent sequence models](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_sequence_classifier.py#L113-L124). \n1) 2X Bi-LSTM, more or less [same architecture](https://github.com/darraghdog/rsna) as last year.\n2) Bert transformer model with 1 layer. \n2) Bert transformer model with 2 layers. \nEach of these used the same loss [(or at least as close as I could get)](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_sequence_classifier.py#L150-L170) as the competition metric on each step, study batchsize 64, no accumulation (again with amp).\nFinally both models and all folds were averaged. \nI submitted a few more transformer models on last day which I had high hopes for but most broke the consistency check - I only allowed submission, if it passed - and then time ran out. (Thanks @yuval6967 for the tip to only drop to disk if it passed). Actually I think 7 of my last 10 submissions failed. \n\n**Inference**\nA few tricks to help speed up and bring RAM and GPU memory down - [final sub](https://www.kaggle.com/darraghdog/rsna512-effnetb5-fold-all-exam-xfrmr-validated?scriptVersionId=45523025).\nBatchsize of 1. Feed sequences of images into image level model in chunks of 128 images, but only normalize (uint8->float32) when that chunk is to be passed into image model. Then concat all output (GAP layer) chunks to make the sequence.  \nConvert all models and input data to half point precision `model=model.half()` - this halved the memory, but results were just as good. \nWith the above, I managed to squeeze three folds out of the above pipeline - four folds also worked, but it failed on consistency check on last day. Interested to see optimisations from other competitiors. \n\nA final note, about two weeks ago, I was stuck on CV could not get the metric working, and decided to kick the competition. Luckily I stuck with it to try a first submission which got in at LB 0.181 single fold :) For such a challenging metric and large dataset the timeline was tight, I think with another couple of weeks the community could have got to 0.12x or so on the leaderboard - but maybe that would overfit the dataset.   ",
      "votes": 52
    },
    {
      "id": 1062255,
      "postDate": "2020-10-27T17:21:46.827Z",
      "content": "<p>Congratz !</p>\n<p>I kind of expected that your solution would be similar to the one you came up with last year since it was already really good. It's even more impressive that you've managed to perform really well here as well.</p>",
      "rawMarkdown": "Congratz !\n\nI kind of expected that your solution would be similar to the one you came up with last year since it was already really good. It's even more impressive that you've managed to perform really well here as well.",
      "votes": 1,
      "replies": [
        {
          "id": 1062275,
          "postDate": "2020-10-27T17:39:11.750Z",
          "content": "<p>Thanks Theo, congrats on your solution also! I found it was handling the metric which was the brain cruncher here --but its good that we are being forced to build a solution which makes sense clinically. </p>",
          "rawMarkdown": "Thanks Theo, congrats on your solution also! I found it was handling the metric which was the brain cruncher here --but its good that we are being forced to build a solution which makes sense clinically. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1061968,
      "postDate": "2020-10-27T13:20:44.483Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>, your speedup tricks are great and using an A100 definitely helps :) <br>\nWhat's your wmetric for the image level model?</p>",
      "rawMarkdown": "Congrats @darraghdog, your speedup tricks are great and using an A100 definitely helps :) \nWhat's your wmetric for the image level model?",
      "votes": 1,
      "replies": [
        {
          "id": 1061992,
          "postDate": "2020-10-27T13:41:12.100Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/Giba\" target=\"_blank\">@Giba</a>, and yes it was first time I tried A100 - compute constraints disappear more or less<br>\nLinky <a href=\"https://github.com/darraghdog/rsnastr/blob/70fd97dbba23278232aa02b7b73826b5d04d1cdb/training/pipeline/train_image_mask_classifier.py#L229-L242\" target=\"_blank\">here</a> how image loss was set up. For each step, I got</p>\n<p>1) Image loss : was just <code>pe_present_on_image</code> loss using BCE for each image. <br>\n2) Study Loss : for each image in the batch get the weighted BCE study loss as hosts weighted them, but excluded <code>negative_exam_for_pe</code> as it is constant for +ve images. Before getting the mean study loss, I masked it to only include +ve  images in the batch.</p>\n<p>Weights for the above are <a href=\"https://github.com/darraghdog/rsnastr/blob/70fd97dbba23278232aa02b7b73826b5d04d1cdb/configs/512/effnetb5_lr5e4_multi.json#L28-L31\" target=\"_blank\">here</a>.  I just tuned the weights for each loss so image and study loss would be close roughly equal when converging. Final loss was sum of 1) and 2).</p>\n<p>This was not ideal, as I got different number of positives each batch, and some batches had no +ve images (I skipped those steps - no backprop). Ideally the sampler would have sampled an equal number of positives each batch - but time was running so I just made batchsize large enough = 48 images with sample of 4:1 for -ve to +ve.   </p>",
          "rawMarkdown": "Thanks @Giba, and yes it was first time I tried A100 - compute constraints disappear more or less\nLinky [here](https://github.com/darraghdog/rsnastr/blob/70fd97dbba23278232aa02b7b73826b5d04d1cdb/training/pipeline/train_image_mask_classifier.py#L229-L242) how image loss was set up. For each step, I got\n\n1) Image loss : was just `pe_present_on_image` loss using BCE for each image. \n2) Study Loss : for each image in the batch get the weighted BCE study loss as hosts weighted them, but excluded `negative_exam_for_pe` as it is constant for +ve images. Before getting the mean study loss, I masked it to only include +ve  images in the batch.\n\nWeights for the above are [here](https://github.com/darraghdog/rsnastr/blob/70fd97dbba23278232aa02b7b73826b5d04d1cdb/configs/512/effnetb5_lr5e4_multi.json#L28-L31).  I just tuned the weights for each loss so image and study loss would be close roughly equal when converging. Final loss was sum of 1) and 2).\n\nThis was not ideal, as I got different number of positives each batch, and some batches had no +ve images (I skipped those steps - no backprop). Ideally the sampler would have sampled an equal number of positives each batch - but time was running so I just made batchsize large enough = 48 images with sample of 4:1 for -ve to +ve.   ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1061918,
      "postDate": "2020-10-27T12:39:12.173Z",
      "content": "<p>Why only 1 transformer layer?</p>",
      "rawMarkdown": "Why only 1 transformer layer?",
      "votes": 1,
      "replies": [
        {
          "id": 1061950,
          "postDate": "2020-10-27T13:05:40.277Z",
          "content": "<p>Good question… the next submissions I was adding included two layers, but it failed on submission with consistency check. I'm running now to see how it would have done. <br>\nI tried 4 layers and it was heavy on memory.  </p>",
          "rawMarkdown": "Good question... the next submissions I was adding included two layers, but it failed on submission with consistency check. I'm running now to see how it would have done. \nI tried 4 layers and it was heavy on memory.  ",
          "votes": 1
        },
        {
          "id": 1062936,
          "postDate": "2020-10-28T10:34:36.907Z",
          "content": "<p>FYI - slight bump by adding a 2-layer transformer in addition to the above - it would get to 0.154 private and public which would not change LB positions. But I'll probably put it in the final solution. </p>",
          "rawMarkdown": "FYI - slight bump by adding a 2-layer transformer in addition to the above - it would get to 0.154 private and public which would not change LB positions. But I'll probably put it in the final solution. "
        }
      ]
    },
    {
      "id": 1061832,
      "postDate": "2020-10-27T11:14:17.517Z",
      "content": "<p>Congratulations. Very interesting model!</p>",
      "rawMarkdown": "Congratulations. Very interesting model!",
      "votes": 1
    },
    {
      "id": 1061776,
      "postDate": "2020-10-27T10:14:37.793Z",
      "content": "<p>Wow. Reading these solutions makes me realise the brains and hard work that goes into these competitions. Really great work!</p>",
      "rawMarkdown": "Wow. Reading these solutions makes me realise the brains and hard work that goes into these competitions. Really great work!",
      "votes": 1
    },
    {
      "id": 1061734,
      "postDate": "2020-10-27T09:18:09.927Z",
      "content": "<p>Congrats on results <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> and thanks for the writeup solution. </p>",
      "rawMarkdown": "Congrats on results @darraghdog and thanks for the writeup solution. ",
      "votes": 1
    },
    {
      "id": 1062107,
      "postDate": "2020-10-27T15:25:13.570Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1062255,
      "author_name": "Theo Viel",
      "author_url": "",
      "post_date": "2020-10-27T17:21:46.827000",
      "content": "<p>Congratz !</p>\n<p>I kind of expected that your solution would be similar to the one you came up with last year since it was already really good. It's even more impressive that you've managed to perform really well here as well.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1062275,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2020-10-27T17:39:11.750000",
          "content": "<p>Thanks Theo, congrats on your solution also! I found it was handling the metric which was the brain cruncher here --but its good that we are being forced to build a solution which makes sense clinically. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1061968,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-10-27T13:20:44.483000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a>, your speedup tricks are great and using an A100 definitely helps :) <br>\nWhat's your wmetric for the image level model?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1061992,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2020-10-27T13:41:12.100000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/Giba\" target=\"_blank\">@Giba</a>, and yes it was first time I tried A100 - compute constraints disappear more or less<br>\nLinky <a href=\"https://github.com/darraghdog/rsnastr/blob/70fd97dbba23278232aa02b7b73826b5d04d1cdb/training/pipeline/train_image_mask_classifier.py#L229-L242\" target=\"_blank\">here</a> how image loss was set up. For each step, I got</p>\n<p>1) Image loss : was just <code>pe_present_on_image</code> loss using BCE for each image. <br>\n2) Study Loss : for each image in the batch get the weighted BCE study loss as hosts weighted them, but excluded <code>negative_exam_for_pe</code> as it is constant for +ve images. Before getting the mean study loss, I masked it to only include +ve  images in the batch.</p>\n<p>Weights for the above are <a href=\"https://github.com/darraghdog/rsnastr/blob/70fd97dbba23278232aa02b7b73826b5d04d1cdb/configs/512/effnetb5_lr5e4_multi.json#L28-L31\" target=\"_blank\">here</a>.  I just tuned the weights for each loss so image and study loss would be close roughly equal when converging. Final loss was sum of 1) and 2).</p>\n<p>This was not ideal, as I got different number of positives each batch, and some batches had no +ve images (I skipped those steps - no backprop). Ideally the sampler would have sampled an equal number of positives each batch - but time was running so I just made batchsize large enough = 48 images with sample of 4:1 for -ve to +ve.   </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1061918,
      "author_name": "Andrés Miguel Torrubia Sáez",
      "author_url": "",
      "post_date": "2020-10-27T12:39:12.173000",
      "content": "<p>Why only 1 transformer layer?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1061950,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2020-10-27T13:05:40.277000",
          "content": "<p>Good question… the next submissions I was adding included two layers, but it failed on submission with consistency check. I'm running now to see how it would have done. <br>\nI tried 4 layers and it was heavy on memory.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1062936,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2020-10-28T10:34:36.907000",
          "content": "<p>FYI - slight bump by adding a 2-layer transformer in addition to the above - it would get to 0.154 private and public which would not change LB positions. But I'll probably put it in the final solution. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1061832,
      "author_name": "Eisa",
      "author_url": "",
      "post_date": "2020-10-27T11:14:17.517000",
      "content": "<p>Congratulations. Very interesting model!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061776,
      "author_name": "Mark P",
      "author_url": "",
      "post_date": "2020-10-27T10:14:37.793000",
      "content": "<p>Wow. Reading these solutions makes me realise the brains and hard work that goes into these competitions. Really great work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1061734,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-10-27T09:18:09.927000",
      "content": "<p>Congrats on results <a href=\"https://www.kaggle.com/darraghdog\" target=\"_blank\">@darraghdog</a> and thanks for the writeup solution. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062107,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-27T15:25:13.570000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1061699": "![](https://raw.githubusercontent.com/darraghdog/rsnastr/main/docs/architecture.jpg)\n**Code base** : https://github.com/darraghdog/rsnastr \n**Kaggle submission** : https://www.kaggle.com/darraghdog/rsnastr2020-prediction\n\nThanks to all the organisers for hosting this competition. And congratulations to all the participants and winners - these competitions and solutions are getting more and more advanced 😁  \n\nMy compute was a v100 on AWS initially with 320X320 images, with prototyping on Macbook (a lot can be checked on CPU). Then in the last 3 weeks I moved to an A100 to run 512X512 (thanks DoubleYard) - 40GB GPU memory per card - one card was enough. \nOverall the solution pipeline ended up being pretty similar to last year, but the journey to get there was a lot different. The metric used by the competition was difficult to simulate, and all parts of the pipeline needed to be finalised/integrated to get feedback on the competition benchmark.  \n\n**Preprocessing**\nUse the [windowing](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/preprocessing/dicom_to_jpeg.py#L59-L75) from Ian Pan's [post](https://www.kaggle.com/c/rsna-str-pulmonary-embolism-detection/discussion/182930) pretty much as is, just [parallelised](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/preprocessing/dicom_to_jpeg.py#L77-L118) it to speed it up. Some dicom's fell out which I saw later was due to the way pydicom was used, but I think the majority were good. \nLight [augmentations](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L72-L74), maybe more would have helped but did not want to lose information which side the PE lay on. \n\n**Image Level**\nEfficientb5 seemed to be better than anything else I tried both in terms of speed and loss. Given the time and compute limits on submission, speed was important. \nThe datasampler took at least [two images from each study, in each epoch, and positive images were oversampled](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L112-L117) to give a rate of 4:1 -ve:+ve, with image batchsize of 48 (using half point precision - amp). \n[Weighted study loss](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L132-L138) was calculated at each step on the positive samples only (weights as per competition metric weights excl. `negative_exam_for_pe`), image loss was calculated on all samples. Final loss per step was [sum of both averages](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_image_classifier.py#L200-L209); image loss was tuned so that image loss and study loss would be roughly equal. \nEach epoch took about 5 mins and ran ~15 epochs, final solution used three of five folds.\n\n**Study Level**\nExtracted gap layer of each image to disk. These were fed into [three independent sequence models](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_sequence_classifier.py#L113-L124). \n1) 2X Bi-LSTM, more or less [same architecture](https://github.com/darraghdog/rsna) as last year.\n2) Bert transformer model with 1 layer. \n2) Bert transformer model with 2 layers. \nEach of these used the same loss [(or at least as close as I could get)](https://github.com/darraghdog/rsnastr/blob/14c8516d3a81bc26c5101afff004ce1e99c3f5a9/training/pipeline/train_sequence_classifier.py#L150-L170) as the competition metric on each step, study batchsize 64, no accumulation (again with amp).\nFinally both models and all folds were averaged. \nI submitted a few more transformer models on last day which I had high hopes for but most broke the consistency check - I only allowed submission, if it passed - and then time ran out. (Thanks @yuval6967 for the tip to only drop to disk if it passed). Actually I think 7 of my last 10 submissions failed. \n\n**Inference**\nA few tricks to help speed up and bring RAM and GPU memory down - [final sub](https://www.kaggle.com/darraghdog/rsna512-effnetb5-fold-all-exam-xfrmr-validated?scriptVersionId=45523025).\nBatchsize of 1. Feed sequences of images into image level model in chunks of 128 images, but only normalize (uint8->float32) when that chunk is to be passed into image model. Then concat all output (GAP layer) chunks to make the sequence.  \nConvert all models and input data to half point precision `model=model.half()` - this halved the memory, but results were just as good. \nWith the above, I managed to squeeze three folds out of the above pipeline - four folds also worked, but it failed on consistency check on last day. Interested to see optimisations from other competitiors. \n\nA final note, about two weeks ago, I was stuck on CV could not get the metric working, and decided to kick the competition. Luckily I stuck with it to try a first submission which got in at LB 0.181 single fold :) For such a challenging metric and large dataset the timeline was tight, I think with another couple of weeks the community could have got to 0.12x or so on the leaderboard - but maybe that would overfit the dataset.   ",
    "1062255": "Congratz !\n\nI kind of expected that your solution would be similar to the one you came up with last year since it was already really good. It's even more impressive that you've managed to perform really well here as well.",
    "1061968": "Congrats @darraghdog, your speedup tricks are great and using an A100 definitely helps :) \nWhat's your wmetric for the image level model?",
    "1061918": "Why only 1 transformer layer?",
    "1061832": "Congratulations. Very interesting model!",
    "1061776": "Wow. Reading these solutions makes me realise the brains and hard work that goes into these competitions. Really great work!",
    "1061734": "Congrats on results @darraghdog and thanks for the writeup solution. ",
    "1062107": ""
  }
}