{
  "id": 391208,
  "title": "4th place solution",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/391208",
  "author_name": "Dieter",
  "post_date": "2023-02-28T17:36:02.926000",
  "votes": 84,
  "comment_count": 19,
  "views": 0,
  "content": "<h2>Introduction</h2>\n<p>Congrats all and thank you RSNA for another great challenge. Our solution is below. Code and slides will be posted over the coming weeks.  </p>\n<h2>Preprocessing</h2>\n<p>We used the same preprocessing and resizing for all models, which enabled us to use more models.</p>\n<p>Pixel values for each image were windowed using the width and center from the DICOM metadata. A linear window was used regardless of the specified function in the metadata to decrease processing time. We applied a coarse CNN with minimum filter to crop the image to the breast area, and hopefully any noise/text, from the image. </p>\n<p>The cancer was often small and the images large, so when resizing it was important to lose as little detail as possible. <a href=\"https://arxiv.org/pdf/2104.11222.pdf\" target=\"_blank\">This paper</a> , concludes that the best resize method is PIL with lanczos, or tensorflow with antialias. There is a good example (Fig 1. in the paper) of how resizing can lose information, comparing cv2, PIL, pytorch and others. We used PIL, which was slow but results outperformed any cv2 method. Using the <code>.thumbnail()</code> function seemed to help speed this up. The trade off was we did not generate any different sized images for modeling, but only did one down size of the image. <br>\nAll images were reduced to 1152 dim. In addition, during training we applied augmentations prior to downsizing, as they would lose more information if applied to a small image. The augmentations are mainly cv2 based. Also, we kept the original filtered image aspect ratio and padded to a square image. </p>\n<h2>Augmentations</h2>\n<p>Train augmentations before downsizing were vflip, hflip,transpose, shift, scale, rotate, grid distortion &amp; affine. After downsizing, we used one of random grid shuffle &amp; coarse dropout. In train, we random cropped to 1024, and in val we center cropped. <br>\nExample batch below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fc28321c365bb1d679c2c35eb1c03e8f7%2FScreenshot%202023-02-28%20at%2018.09.06.png?generation=1677604161886568&amp;alt=media\" alt=\"\"></p>\n<h2>Models</h2>\n<h3>Model type 1: Breast level feature combination 1D-CNN with CNN backbone</h3>\n<p>Auxiliary loss improved time to convergence a lot. For each backbone, stage 1 model trained for 9 epochs, stage 2 model trained for 2 epochs with frozen backbone. Heavy dropout (0.5) was applied on the linear output. <br>\nThe 1D-CNN was set up with a filter size of 2 and no padding, so we get a (2, n_features) shape input and a (1, n_features) output. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F59667ff80e38d33611d00b97752a6aaa%2FScreenshot%202023-03-01%20at%2010.33.04.png?generation=1677663219589043&amp;alt=media\" alt=\"\"></p>\n<p>In our final ensemble we ran the above architecture with efficientnet b3 b4 b5 v2s and v2m.</p>\n<p>Some of the solution was inspired by Bo’s great write up on the <a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">SIIM melanoma competition 1st place solution</a>.</p>\n<h3>Model type 2: Patient-level multi-view-multi-lateral transformer with CNN backbone</h3>\n<p>Inspired by the paper <a href=\"https://arxiv.org/pdf/2206.10096.pdf\" target=\"_blank\">Transformers Improve Breast Cancer Diagnosis from Unregistered Multi-View Mammograms</a>, we designed another 2nd stage approach, but on a patient level. </p>\n<h4>Stage 1:</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F809528c5c906c90e6752f2f6729f1f5c%2FScreenshot%202023-02-28%20at%2013.38.19.png?generation=1677604933375806&amp;alt=media\" alt=\"\"></p>\n<p>In the first stage we train an image level CNN which is trained with an auxiliary segmentation loss by predicting masks we got from training a Yolo_v7 on CBIS dataset.</p>\n<h4>Stage 2:</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fa908b5aad5f9bf7b0b1dc8ce77461eab%2FScreenshot%202023-02-28%20at%2013.38.10.png?generation=1677605002279914&amp;alt=media\" alt=\"\"></p>\n<p>We take the CNN trained in stage 1 and aggregate to patient level by considering four images per patient. Two views for each breast, i.e L-CC, L-MLO, R-CC, R-MLO. Using a 2-view/ 2-lateral input, the transformer can hopefully learn: </p>\n<ul>\n<li>consistency between views </li>\n<li>differences in laterals</li>\n</ul>\n<p>We extract the output feature maps of the image level CNN for each view ending up with a tensors of size (4,out_channels, 32,32), which is projected to a predefined grid-size (4,hidden_dim, 16,16) using a conv layer. Then the 4x16x16 features, which can be seen as patch-tokens are spatially flattened and concatenated. The result is a sequence (hidden_dim, 1024) which is put into a transformer. Output of that transformer are two cancer predictions, one for each breast.</p>\n<p>We first froze the CNN backbone for a few epochs to let the pretrained vision transformer \"adjust\" before fine-tuning the whole 2-stage model end2end. <br>\nIf a patient has more than one image per view we randomly sample one image per view for training and use multiple 4-view combinations for inference which are then averaged. <br>\nBackbone-wise we used pretrained seresnext50 and convnext tiny as CNN and pretrained deit-tiny-patch16-224 as transformer.</p>\n<h2>Ensembling</h2>\n<p>So in total we used single seed fullfits of the following models</p>\n<ul>\n<li>Effnet v2_s + 1D-CNN</li>\n<li>Effnet v2_m + 1D-CNN</li>\n<li>Effnet b3 + 1D-CNN</li>\n<li>Effnet b4 + 1D-CNN</li>\n<li>Effnet b5 + 1D-CNN</li>\n<li>SE-ResNext50 + deit-tiny-patch16-224</li>\n<li>ConvNext_tiny + deit-tiny-patch16-224</li>\n</ul>\n<p>A threshold to convert to binary output was selected based on the CV of the blend. <br>\nOur two highest submissions according to Public LB were also our two highest blends on CV. They were not the highest private, but close to it. For the shake up, it probably helped that we did not start blending the models until the last few days of the competition, so did not get distracted by high Public LB scores. </p>\n<h2>Things that did not work/were ineffective:</h2>\n<ul>\n<li>Pretraining on DDSM/VinDr-Mammo</li>\n<li>Training a ROI extractor and training subsequent models on focused ROIs rather than whole images</li>\n<li>Higher resolution images (similar to worse performance, longer inference times)</li>\n</ul>",
  "messages": [
    {
      "id": 2163266,
      "postDate": "2023-02-28T17:36:02.927Z",
      "content": "<h2>Introduction</h2>\n<p>Congrats all and thank you RSNA for another great challenge. Our solution is below. Code and slides will be posted over the coming weeks.  </p>\n<h2>Preprocessing</h2>\n<p>We used the same preprocessing and resizing for all models, which enabled us to use more models.</p>\n<p>Pixel values for each image were windowed using the width and center from the DICOM metadata. A linear window was used regardless of the specified function in the metadata to decrease processing time. We applied a coarse CNN with minimum filter to crop the image to the breast area, and hopefully any noise/text, from the image. </p>\n<p>The cancer was often small and the images large, so when resizing it was important to lose as little detail as possible. <a href=\"https://arxiv.org/pdf/2104.11222.pdf\" target=\"_blank\">This paper</a> , concludes that the best resize method is PIL with lanczos, or tensorflow with antialias. There is a good example (Fig 1. in the paper) of how resizing can lose information, comparing cv2, PIL, pytorch and others. We used PIL, which was slow but results outperformed any cv2 method. Using the <code>.thumbnail()</code> function seemed to help speed this up. The trade off was we did not generate any different sized images for modeling, but only did one down size of the image. <br>\nAll images were reduced to 1152 dim. In addition, during training we applied augmentations prior to downsizing, as they would lose more information if applied to a small image. The augmentations are mainly cv2 based. Also, we kept the original filtered image aspect ratio and padded to a square image. </p>\n<h2>Augmentations</h2>\n<p>Train augmentations before downsizing were vflip, hflip,transpose, shift, scale, rotate, grid distortion &amp; affine. After downsizing, we used one of random grid shuffle &amp; coarse dropout. In train, we random cropped to 1024, and in val we center cropped. <br>\nExample batch below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fc28321c365bb1d679c2c35eb1c03e8f7%2FScreenshot%202023-02-28%20at%2018.09.06.png?generation=1677604161886568&amp;alt=media\" alt=\"\"></p>\n<h2>Models</h2>\n<h3>Model type 1: Breast level feature combination 1D-CNN with CNN backbone</h3>\n<p>Auxiliary loss improved time to convergence a lot. For each backbone, stage 1 model trained for 9 epochs, stage 2 model trained for 2 epochs with frozen backbone. Heavy dropout (0.5) was applied on the linear output. <br>\nThe 1D-CNN was set up with a filter size of 2 and no padding, so we get a (2, n_features) shape input and a (1, n_features) output. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F59667ff80e38d33611d00b97752a6aaa%2FScreenshot%202023-03-01%20at%2010.33.04.png?generation=1677663219589043&amp;alt=media\" alt=\"\"></p>\n<p>In our final ensemble we ran the above architecture with efficientnet b3 b4 b5 v2s and v2m.</p>\n<p>Some of the solution was inspired by Bo’s great write up on the <a href=\"https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412\" target=\"_blank\">SIIM melanoma competition 1st place solution</a>.</p>\n<h3>Model type 2: Patient-level multi-view-multi-lateral transformer with CNN backbone</h3>\n<p>Inspired by the paper <a href=\"https://arxiv.org/pdf/2206.10096.pdf\" target=\"_blank\">Transformers Improve Breast Cancer Diagnosis from Unregistered Multi-View Mammograms</a>, we designed another 2nd stage approach, but on a patient level. </p>\n<h4>Stage 1:</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F809528c5c906c90e6752f2f6729f1f5c%2FScreenshot%202023-02-28%20at%2013.38.19.png?generation=1677604933375806&amp;alt=media\" alt=\"\"></p>\n<p>In the first stage we train an image level CNN which is trained with an auxiliary segmentation loss by predicting masks we got from training a Yolo_v7 on CBIS dataset.</p>\n<h4>Stage 2:</h4>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fa908b5aad5f9bf7b0b1dc8ce77461eab%2FScreenshot%202023-02-28%20at%2013.38.10.png?generation=1677605002279914&amp;alt=media\" alt=\"\"></p>\n<p>We take the CNN trained in stage 1 and aggregate to patient level by considering four images per patient. Two views for each breast, i.e L-CC, L-MLO, R-CC, R-MLO. Using a 2-view/ 2-lateral input, the transformer can hopefully learn: </p>\n<ul>\n<li>consistency between views </li>\n<li>differences in laterals</li>\n</ul>\n<p>We extract the output feature maps of the image level CNN for each view ending up with a tensors of size (4,out_channels, 32,32), which is projected to a predefined grid-size (4,hidden_dim, 16,16) using a conv layer. Then the 4x16x16 features, which can be seen as patch-tokens are spatially flattened and concatenated. The result is a sequence (hidden_dim, 1024) which is put into a transformer. Output of that transformer are two cancer predictions, one for each breast.</p>\n<p>We first froze the CNN backbone for a few epochs to let the pretrained vision transformer \"adjust\" before fine-tuning the whole 2-stage model end2end. <br>\nIf a patient has more than one image per view we randomly sample one image per view for training and use multiple 4-view combinations for inference which are then averaged. <br>\nBackbone-wise we used pretrained seresnext50 and convnext tiny as CNN and pretrained deit-tiny-patch16-224 as transformer.</p>\n<h2>Ensembling</h2>\n<p>So in total we used single seed fullfits of the following models</p>\n<ul>\n<li>Effnet v2_s + 1D-CNN</li>\n<li>Effnet v2_m + 1D-CNN</li>\n<li>Effnet b3 + 1D-CNN</li>\n<li>Effnet b4 + 1D-CNN</li>\n<li>Effnet b5 + 1D-CNN</li>\n<li>SE-ResNext50 + deit-tiny-patch16-224</li>\n<li>ConvNext_tiny + deit-tiny-patch16-224</li>\n</ul>\n<p>A threshold to convert to binary output was selected based on the CV of the blend. <br>\nOur two highest submissions according to Public LB were also our two highest blends on CV. They were not the highest private, but close to it. For the shake up, it probably helped that we did not start blending the models until the last few days of the competition, so did not get distracted by high Public LB scores. </p>\n<h2>Things that did not work/were ineffective:</h2>\n<ul>\n<li>Pretraining on DDSM/VinDr-Mammo</li>\n<li>Training a ROI extractor and training subsequent models on focused ROIs rather than whole images</li>\n<li>Higher resolution images (similar to worse performance, longer inference times)</li>\n</ul>",
      "rawMarkdown": "## Introduction\n\nCongrats all and thank you RSNA for another great challenge. Our solution is below. Code and slides will be posted over the coming weeks.  \n\n## Preprocessing \n\nWe used the same preprocessing and resizing for all models, which enabled us to use more models.\n\nPixel values for each image were windowed using the width and center from the DICOM metadata. A linear window was used regardless of the specified function in the metadata to decrease processing time. We applied a coarse CNN with minimum filter to crop the image to the breast area, and hopefully any noise/text, from the image. \n\nThe cancer was often small and the images large, so when resizing it was important to lose as little detail as possible. [This paper](https://arxiv.org/pdf/2104.11222.pdf) , concludes that the best resize method is PIL with lanczos, or tensorflow with antialias. There is a good example (Fig 1. in the paper) of how resizing can lose information, comparing cv2, PIL, pytorch and others. We used PIL, which was slow but results outperformed any cv2 method. Using the `.thumbnail()` function seemed to help speed this up. The trade off was we did not generate any different sized images for modeling, but only did one down size of the image. \nAll images were reduced to 1152 dim. In addition, during training we applied augmentations prior to downsizing, as they would lose more information if applied to a small image. The augmentations are mainly cv2 based. Also, we kept the original filtered image aspect ratio and padded to a square image. \n\n## Augmentations\n\nTrain augmentations before downsizing were vflip, hflip,transpose, shift, scale, rotate, grid distortion & affine. After downsizing, we used one of random grid shuffle & coarse dropout. In train, we random cropped to 1024, and in val we center cropped. \nExample batch below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fc28321c365bb1d679c2c35eb1c03e8f7%2FScreenshot%202023-02-28%20at%2018.09.06.png?generation=1677604161886568&alt=media)\n\n## Models\n\n### Model type 1: Breast level feature combination 1D-CNN with CNN backbone\n\nAuxiliary loss improved time to convergence a lot. For each backbone, stage 1 model trained for 9 epochs, stage 2 model trained for 2 epochs with frozen backbone. Heavy dropout (0.5) was applied on the linear output. \nThe 1D-CNN was set up with a filter size of 2 and no padding, so we get a (2, n_features) shape input and a (1, n_features) output. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F59667ff80e38d33611d00b97752a6aaa%2FScreenshot%202023-03-01%20at%2010.33.04.png?generation=1677663219589043&alt=media)\n\nIn our final ensemble we ran the above architecture with efficientnet b3 b4 b5 v2s and v2m.\n\nSome of the solution was inspired by Bo’s great write up on the [SIIM melanoma competition 1st place solution](https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412).\n\n### Model type 2: Patient-level multi-view-multi-lateral transformer with CNN backbone\n\nInspired by the paper [Transformers Improve Breast Cancer Diagnosis from Unregistered Multi-View Mammograms](https://arxiv.org/pdf/2206.10096.pdf), we designed another 2nd stage approach, but on a patient level. \n\n#### Stage 1:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F809528c5c906c90e6752f2f6729f1f5c%2FScreenshot%202023-02-28%20at%2013.38.19.png?generation=1677604933375806&alt=media)\n\nIn the first stage we train an image level CNN which is trained with an auxiliary segmentation loss by predicting masks we got from training a Yolo_v7 on CBIS dataset.\n\n#### Stage 2:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fa908b5aad5f9bf7b0b1dc8ce77461eab%2FScreenshot%202023-02-28%20at%2013.38.10.png?generation=1677605002279914&alt=media)\n\nWe take the CNN trained in stage 1 and aggregate to patient level by considering four images per patient. Two views for each breast, i.e L-CC, L-MLO, R-CC, R-MLO. Using a 2-view/ 2-lateral input, the transformer can hopefully learn: \n\n- consistency between views \n- differences in laterals\n\nWe extract the output feature maps of the image level CNN for each view ending up with a tensors of size (4,out_channels, 32,32), which is projected to a predefined grid-size (4,hidden_dim, 16,16) using a conv layer. Then the 4x16x16 features, which can be seen as patch-tokens are spatially flattened and concatenated. The result is a sequence (hidden_dim, 1024) which is put into a transformer. Output of that transformer are two cancer predictions, one for each breast.\n\nWe first froze the CNN backbone for a few epochs to let the pretrained vision transformer \"adjust\" before fine-tuning the whole 2-stage model end2end. \nIf a patient has more than one image per view we randomly sample one image per view for training and use multiple 4-view combinations for inference which are then averaged. \nBackbone-wise we used pretrained seresnext50 and convnext tiny as CNN and pretrained deit-tiny-patch16-224 as transformer.\n\n##Ensembling\n\nSo in total we used single seed fullfits of the following models\n\n- Effnet v2_s + 1D-CNN\n- Effnet v2_m + 1D-CNN\n- Effnet b3 + 1D-CNN\n- Effnet b4 + 1D-CNN\n- Effnet b5 + 1D-CNN\n- SE-ResNext50 + deit-tiny-patch16-224\n- ConvNext_tiny + deit-tiny-patch16-224\n\nA threshold to convert to binary output was selected based on the CV of the blend. \nOur two highest submissions according to Public LB were also our two highest blends on CV. They were not the highest private, but close to it. For the shake up, it probably helped that we did not start blending the models until the last few days of the competition, so did not get distracted by high Public LB scores. \n\n##Things that did not work/were ineffective:\n\n- Pretraining on DDSM/VinDr-Mammo\n- Training a ROI extractor and training subsequent models on focused ROIs rather than whole images\n- Higher resolution images (similar to worse performance, longer inference times)\n\n",
      "votes": 84
    },
    {
      "id": 2163674,
      "postDate": "2023-03-01T01:50:25.167Z",
      "content": "<p>Very cool and beautiful solution, and congrats for your team!</p>\n<p>Could you offer more details about <strong>Model type 1 Stage 2</strong>, I've the following question:</p>\n<ol>\n<li><p>the input \"images from the breast\". In the stage 1, the model input is single image, but here, for the same model, the input becomes multi images, I assume that you are using the <code>batch size dimension</code> to compose the multi images?</p></li>\n<li><p>I do not quite understand the \"2D CNN for each 2D combination with the 1D output of same dim\", could please tell me more about this part?</p></li>\n<li><p>for \"Model type 2 stage 1\". This stage seems like has not have a strong correlation with \" \"Model type 2 stage 2\". I want to know if you have some experiments on how much this stage helps with the CV score? </p></li>\n</ol>\n<p>Thanks in advance.</p>",
      "rawMarkdown": "Very cool and beautiful solution, and congrats for your team!\n\nCould you offer more details about **Model type 1 Stage 2**, I've the following question:\n\n1. the input \"images from the breast\". In the stage 1, the model input is single image, but here, for the same model, the input becomes multi images, I assume that you are using the `batch size dimension` to compose the multi images?\n\n2. I do not quite understand the \"2D CNN for each 2D combination with the 1D output of same dim\", could please tell me more about this part?\n\n3. for \"Model type 2 stage 1\". This stage seems like has not have a strong correlation with \" \"Model type 2 stage 2\". I want to know if you have some experiments on how much this stage helps with the CV score? \n\nThanks in advance.",
      "votes": 3,
      "replies": [
        {
          "id": 2164030,
          "postDate": "2023-03-01T08:55:17.370Z",
          "content": "<p>For Q2. Lets say we have 3 images, <code>img1, img2, img3</code>, we get the gap layer for each and get the two way pairs, (eg. <code>img1&amp;img2</code>, <code>img1&amp;img3</code>, <code>img2&amp;img3</code>) and pass them to a 2D CNN. Then average the results of each pair. The intuition here is that the feature maps in the same position on each gap layer should correspond; and a 2DCNN can directly compare on the feature map level. <br>\nQ3. The CNN backbone is the most parameters in any of the above models… so getting the backbone trained to learn what cancer looks like, helps a lot before adding more advanced layer on top of it. Also, stage 1 would have more data (~56K images), as only one image is passed in at a time. Whereas stage2 only has the ~11K patients as samples. </p>",
          "rawMarkdown": "For Q2. Lets say we have 3 images, `img1, img2, img3`, we get the gap layer for each and get the two way pairs, (eg. `img1&img2`, `img1&img3`, `img2&img3`) and pass them to a 2D CNN. Then average the results of each pair. The intuition here is that the feature maps in the same position on each gap layer should correspond; and a 2DCNN can directly compare on the feature map level. \nQ3. The CNN backbone is the most parameters in any of the above models... so getting the backbone trained to learn what cancer looks like, helps a lot before adding more advanced layer on top of it. Also, stage 1 would have more data (~56K images), as only one image is passed in at a time. Whereas stage2 only has the ~11K patients as samples. ",
          "votes": 2,
          "replies": [
            {
              "id": 2164043,
              "postDate": "2023-03-01T09:07:41.237Z",
              "content": "<p>Thanks very much for answering. I am still confused about \"2D CNN\", what is that specifically? Is that a <code>nn.Conv2d</code>?</p>\n<p>After the pooling layer, for the 3 images, we can get 3 feature vectors each. I want to know how do you \"pair\" those vectors?</p>",
              "rawMarkdown": "Thanks very much for answering. I am still confused about \"2D CNN\", what is that specifically? Is that a `nn.Conv2d`?\n\nAfter the pooling layer, for the 3 images, we can get 3 feature vectors each. I want to know how do you \"pair\" those vectors?",
              "votes": 1
            },
            {
              "id": 2164055,
              "postDate": "2023-03-01T09:21:00.787Z",
              "content": "<p>Oops, sorry, your right, its a 1DCNN, need to update the docu 🤦‍♂️ </p>\n<pre><code>backbone_out = 1408\nfea_conv = nn.Sequential( nn.Conv1d(backbone_out, backbone_out, kernel_size=2, padding=0), nn.ReLU())\n</code></pre>\n<p>So lets say we have 8 pairs in a batch…</p>\n<pre><code>x = torch.zeros(8,backbone_out,2)\nxout = fea_conv(x)\nxout.shape # [8, 1408, 1]\n</code></pre>",
              "rawMarkdown": "Oops, sorry, your right, its a 1DCNN, need to update the docu 🤦‍♂️ \n```\nbackbone_out = 1408\nfea_conv = nn.Sequential( nn.Conv1d(backbone_out, backbone_out, kernel_size=2, padding=0), nn.ReLU())\n```\nSo lets say we have 8 pairs in a batch...\n```\n\nx = torch.zeros(8,backbone_out,2)\nxout = fea_conv(x)\nxout.shape # [8, 1408, 1]\n```",
              "votes": 1
            },
            {
              "id": 2164200,
              "postDate": "2023-03-01T11:53:42.780Z",
              "content": "<p>That's it, thanks so much for the explaination.</p>",
              "rawMarkdown": "That's it, thanks so much for the explaination."
            }
          ]
        }
      ]
    },
    {
      "id": 2170877,
      "postDate": "2023-03-06T11:10:20.120Z",
      "content": "<p>The transformer part is fancy. Congratulations and thank you! 🎉</p>",
      "rawMarkdown": "The transformer part is fancy. Congratulations and thank you! 🎉",
      "votes": 1
    },
    {
      "id": 2163303,
      "postDate": "2023-02-28T17:55:28.007Z",
      "content": "<p>Congratulations!!!, I'd like to know if it's allow to use Yolo during the competition.</p>\n<p>Thank in advance</p>",
      "rawMarkdown": "Congratulations!!!, I'd like to know if it's allow to use Yolo during the competition.\n\nThank in advance",
      "votes": 1
    },
    {
      "id": 2192544,
      "postDate": "2023-03-22T18:04:33.130Z",
      "content": "<p>Hey, any idea of sharing your Source code? That would be really helpful. <br>\nThank You.</p>",
      "rawMarkdown": "Hey, any idea of sharing your Source code? That would be really helpful. \nThank You."
    },
    {
      "id": 2170876,
      "postDate": "2023-03-06T11:08:51.193Z",
      "content": "<p>Congratulations!<br>\nI have two questions:</p>\n<ol>\n<li>What device did you use? Is it kaggle device or you own machine?</li>\n<li>What size of batch_size did you use?</li>\n</ol>",
      "rawMarkdown": "Congratulations!\nI have two questions:\n1. What device did you use? Is it kaggle device or you own machine?\n2. What size of batch_size did you use?"
    },
    {
      "id": 2170063,
      "postDate": "2023-03-05T16:47:15.017Z",
      "content": "<p>I'm sorry if this is a dumb question, but how are the axillary losses supposed to help the model? it would be, for example, \"Can cancer show up differently for mammography Birads/Density type \"A\" and type \"B\" ? . What I'm trying to understand is what we're saying to the model when we use auxiliary losses along with the main loss (Cancer). Sorry if my question is not clear. Thanks !!! </p>",
      "rawMarkdown": "\nI'm sorry if this is a dumb question, but how are the axillary losses supposed to help the model? it would be, for example, \"Can cancer show up differently for mammography Birads/Density type \"A\" and type \"B\" ? . What I'm trying to understand is what we're saying to the model when we use auxiliary losses along with the main loss (Cancer). Sorry if my question is not clear. Thanks !!! ",
      "replies": [
        {
          "id": 2170172,
          "postDate": "2023-03-05T18:27:44.217Z",
          "content": "<p>I am not sure which aux labels helped more or less; and I see other teams also used <code>age</code> as an aux label. Some of them may not have helped - we did not test this too much. <code>birads</code> was an indication of cancer risk (I think the C <code>birads</code> class was obvioulsy cancer free), and especially <code>difficult_negative_case</code> also indicated dicoms which deserved more attention. <br>\nAs the <code>cancer</code> label is highly imbalanced (many batches had no positive label), my assumption is that the model was better able to learn the type of dicoms it should pay more attention to by providing it more labels like the above, especially in early epochs. This is why it learned a lot faster; and I guess this lead to more stable learning. In later epochs, we downweighted the aux to roughly 50/50 cancer/auxilliiary. <br>\nI believe 2nd placed team, used much larger batch size so every batch had enough cancer to learn in a stable manner, this may also have worked well. <br>\nHope this makes sense. </p>",
          "rawMarkdown": "I am not sure which aux labels helped more or less; and I see other teams also used `age` as an aux label. Some of them may not have helped - we did not test this too much. `birads` was an indication of cancer risk (I think the C `birads` class was obvioulsy cancer free), and especially `difficult_negative_case` also indicated dicoms which deserved more attention. \nAs the `cancer` label is highly imbalanced (many batches had no positive label), my assumption is that the model was better able to learn the type of dicoms it should pay more attention to by providing it more labels like the above, especially in early epochs. This is why it learned a lot faster; and I guess this lead to more stable learning. In later epochs, we downweighted the aux to roughly 50/50 cancer/auxilliiary. \nI believe 2nd placed team, used much larger batch size so every batch had enough cancer to learn in a stable manner, this may also have worked well. \nHope this makes sense. ",
          "votes": 1,
          "replies": [
            {
              "id": 2170312,
              "postDate": "2023-03-05T20:51:46.583Z",
              "content": "<p>I really liked your explanation. Now I understand the use of axillary losses, thank you very much !!! 😄</p>",
              "rawMarkdown": "\nI really liked your explanation. Now I understand the use of axillary losses, thank you very much !!! 😄",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2169074,
      "postDate": "2023-03-04T18:57:58.463Z",
      "content": "<p>Congratulations team! Great job.</p>",
      "rawMarkdown": "Congratulations team! Great job."
    },
    {
      "id": 2168540,
      "postDate": "2023-03-04T10:04:10.920Z",
      "content": "<p>good project</p>",
      "rawMarkdown": "good project"
    },
    {
      "id": 2165116,
      "postDate": "2023-03-02T02:12:43.307Z",
      "content": "<p>Thanks for sharing and congrats to your team!</p>",
      "rawMarkdown": "Thanks for sharing and congrats to your team!"
    },
    {
      "id": 2164450,
      "postDate": "2023-03-01T15:28:04.680Z",
      "content": "<p>Congrats to your team and thank you for sharing the solution (looking forward to reading the code!). Loved the recommendation about PIL and TF antialias for resizing. Will read that paper!</p>\n<p>Thanks so much for sharing!</p>",
      "rawMarkdown": "Congrats to your team and thank you for sharing the solution (looking forward to reading the code!). Loved the recommendation about PIL and TF antialias for resizing. Will read that paper!\n\nThanks so much for sharing!\n\n"
    },
    {
      "id": 2164165,
      "postDate": "2023-03-01T11:07:32.053Z",
      "content": "<p>Congratulations, a very strong solution!<br>\nCould you share how the CV for the stage2 model was arranged?</p>",
      "rawMarkdown": "Congratulations, a very strong solution!\nCould you share how the CV for the stage2 model was arranged?"
    },
    {
      "id": 2163589,
      "postDate": "2023-02-28T23:45:01.660Z",
      "content": "<p>Thanks for sharing and many congratulations to the entire team!</p>\n<p>Any chance you could share your CV score for the different models/ensembling?</p>",
      "rawMarkdown": "Thanks for sharing and many congratulations to the entire team!\n\nAny chance you could share your CV score for the different models/ensembling?"
    },
    {
      "id": 2163399,
      "postDate": "2023-02-28T19:23:57.737Z",
      "content": "<blockquote>\n  <p>The cancer was often small and the images large, so when resizing it was important to lose as little detail as possible. This paper , concludes that the best resize method is PIL with lanczos, or tensorflow with antialias. There is a good example (Fig 1. in the paper) of how resizing can lose information, comparing cv2, PIL, pytorch and others</p>\n</blockquote>\n<p>Thank you for this research highlight!</p>\n<p>And congratulations for the 4th place solution to you and your team!</p>",
      "rawMarkdown": ">The cancer was often small and the images large, so when resizing it was important to lose as little detail as possible. This paper , concludes that the best resize method is PIL with lanczos, or tensorflow with antialias. There is a good example (Fig 1. in the paper) of how resizing can lose information, comparing cv2, PIL, pytorch and others\n\nThank you for this research highlight!\n\nAnd congratulations for the 4th place solution to you and your team!"
    }
  ],
  "comments": [
    {
      "id": 2163674,
      "author_name": "Chenglu",
      "author_url": "",
      "post_date": "2023-03-01T01:50:25.167000",
      "content": "<p>Very cool and beautiful solution, and congrats for your team!</p>\n<p>Could you offer more details about <strong>Model type 1 Stage 2</strong>, I've the following question:</p>\n<ol>\n<li><p>the input \"images from the breast\". In the stage 1, the model input is single image, but here, for the same model, the input becomes multi images, I assume that you are using the <code>batch size dimension</code> to compose the multi images?</p></li>\n<li><p>I do not quite understand the \"2D CNN for each 2D combination with the 1D output of same dim\", could please tell me more about this part?</p></li>\n<li><p>for \"Model type 2 stage 1\". This stage seems like has not have a strong correlation with \" \"Model type 2 stage 2\". I want to know if you have some experiments on how much this stage helps with the CV score? </p></li>\n</ol>\n<p>Thanks in advance.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2164030,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2023-03-01T08:55:17.370000",
          "content": "<p>For Q2. Lets say we have 3 images, <code>img1, img2, img3</code>, we get the gap layer for each and get the two way pairs, (eg. <code>img1&amp;img2</code>, <code>img1&amp;img3</code>, <code>img2&amp;img3</code>) and pass them to a 2D CNN. Then average the results of each pair. The intuition here is that the feature maps in the same position on each gap layer should correspond; and a 2DCNN can directly compare on the feature map level. <br>\nQ3. The CNN backbone is the most parameters in any of the above models… so getting the backbone trained to learn what cancer looks like, helps a lot before adding more advanced layer on top of it. Also, stage 1 would have more data (~56K images), as only one image is passed in at a time. Whereas stage2 only has the ~11K patients as samples. </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2164043,
              "author_name": "Chenglu",
              "author_url": "",
              "post_date": "2023-03-01T09:07:41.237000",
              "content": "<p>Thanks very much for answering. I am still confused about \"2D CNN\", what is that specifically? Is that a <code>nn.Conv2d</code>?</p>\n<p>After the pooling layer, for the 3 images, we can get 3 feature vectors each. I want to know how do you \"pair\" those vectors?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2164055,
              "author_name": "Darragh",
              "author_url": "",
              "post_date": "2023-03-01T09:21:00.787000",
              "content": "<p>Oops, sorry, your right, its a 1DCNN, need to update the docu 🤦‍♂️ </p>\n<pre><code>backbone_out = 1408\nfea_conv = nn.Sequential( nn.Conv1d(backbone_out, backbone_out, kernel_size=2, padding=0), nn.ReLU())\n</code></pre>\n<p>So lets say we have 8 pairs in a batch…</p>\n<pre><code>x = torch.zeros(8,backbone_out,2)\nxout = fea_conv(x)\nxout.shape # [8, 1408, 1]\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2164200,
              "author_name": "Chenglu",
              "author_url": "",
              "post_date": "2023-03-01T11:53:42.780000",
              "content": "<p>That's it, thanks so much for the explaination.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2170877,
      "author_name": "ForcewithMe",
      "author_url": "",
      "post_date": "2023-03-06T11:10:20.120000",
      "content": "<p>The transformer part is fancy. Congratulations and thank you! 🎉</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2163303,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2023-02-28T17:55:28.007000",
      "content": "<p>Congratulations!!!, I'd like to know if it's allow to use Yolo during the competition.</p>\n<p>Thank in advance</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2192544,
      "author_name": "tharun_01",
      "author_url": "",
      "post_date": "2023-03-22T18:04:33.130000",
      "content": "<p>Hey, any idea of sharing your Source code? That would be really helpful. <br>\nThank You.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2170876,
      "author_name": "Veronika Chernova",
      "author_url": "",
      "post_date": "2023-03-06T11:08:51.193000",
      "content": "<p>Congratulations!<br>\nI have two questions:</p>\n<ol>\n<li>What device did you use? Is it kaggle device or you own machine?</li>\n<li>What size of batch_size did you use?</li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2170063,
      "author_name": "adriel cabral",
      "author_url": "",
      "post_date": "2023-03-05T16:47:15.017000",
      "content": "<p>I'm sorry if this is a dumb question, but how are the axillary losses supposed to help the model? it would be, for example, \"Can cancer show up differently for mammography Birads/Density type \"A\" and type \"B\" ? . What I'm trying to understand is what we're saying to the model when we use auxiliary losses along with the main loss (Cancer). Sorry if my question is not clear. Thanks !!! </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2170172,
          "author_name": "Darragh",
          "author_url": "",
          "post_date": "2023-03-05T18:27:44.217000",
          "content": "<p>I am not sure which aux labels helped more or less; and I see other teams also used <code>age</code> as an aux label. Some of them may not have helped - we did not test this too much. <code>birads</code> was an indication of cancer risk (I think the C <code>birads</code> class was obvioulsy cancer free), and especially <code>difficult_negative_case</code> also indicated dicoms which deserved more attention. <br>\nAs the <code>cancer</code> label is highly imbalanced (many batches had no positive label), my assumption is that the model was better able to learn the type of dicoms it should pay more attention to by providing it more labels like the above, especially in early epochs. This is why it learned a lot faster; and I guess this lead to more stable learning. In later epochs, we downweighted the aux to roughly 50/50 cancer/auxilliiary. <br>\nI believe 2nd placed team, used much larger batch size so every batch had enough cancer to learn in a stable manner, this may also have worked well. <br>\nHope this makes sense. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2170312,
              "author_name": "adriel cabral",
              "author_url": "",
              "post_date": "2023-03-05T20:51:46.583000",
              "content": "<p>I really liked your explanation. Now I understand the use of axillary losses, thank you very much !!! 😄</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2169074,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-03-04T18:57:58.463000",
      "content": "<p>Congratulations team! Great job.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2168540,
      "author_name": "SHOHRUH BEGMATOV",
      "author_url": "",
      "post_date": "2023-03-04T10:04:10.920000",
      "content": "<p>good project</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2165116,
      "author_name": "ZEROO",
      "author_url": "",
      "post_date": "2023-03-02T02:12:43.307000",
      "content": "<p>Thanks for sharing and congrats to your team!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2164450,
      "author_name": "Ebc",
      "author_url": "",
      "post_date": "2023-03-01T15:28:04.680000",
      "content": "<p>Congrats to your team and thank you for sharing the solution (looking forward to reading the code!). Loved the recommendation about PIL and TF antialias for resizing. Will read that paper!</p>\n<p>Thanks so much for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2164165,
      "author_name": "Arthur Mulikhov",
      "author_url": "",
      "post_date": "2023-03-01T11:07:32.053000",
      "content": "<p>Congratulations, a very strong solution!<br>\nCould you share how the CV for the stage2 model was arranged?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2163589,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-02-28T23:45:01.660000",
      "content": "<p>Thanks for sharing and many congratulations to the entire team!</p>\n<p>Any chance you could share your CV score for the different models/ensembling?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2163399,
      "author_name": "Antti Isosalo",
      "author_url": "",
      "post_date": "2023-02-28T19:23:57.737000",
      "content": "<blockquote>\n  <p>The cancer was often small and the images large, so when resizing it was important to lose as little detail as possible. This paper , concludes that the best resize method is PIL with lanczos, or tensorflow with antialias. There is a good example (Fig 1. in the paper) of how resizing can lose information, comparing cv2, PIL, pytorch and others</p>\n</blockquote>\n<p>Thank you for this research highlight!</p>\n<p>And congratulations for the 4th place solution to you and your team!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2163266": "## Introduction\n\nCongrats all and thank you RSNA for another great challenge. Our solution is below. Code and slides will be posted over the coming weeks.  \n\n## Preprocessing \n\nWe used the same preprocessing and resizing for all models, which enabled us to use more models.\n\nPixel values for each image were windowed using the width and center from the DICOM metadata. A linear window was used regardless of the specified function in the metadata to decrease processing time. We applied a coarse CNN with minimum filter to crop the image to the breast area, and hopefully any noise/text, from the image. \n\nThe cancer was often small and the images large, so when resizing it was important to lose as little detail as possible. [This paper](https://arxiv.org/pdf/2104.11222.pdf) , concludes that the best resize method is PIL with lanczos, or tensorflow with antialias. There is a good example (Fig 1. in the paper) of how resizing can lose information, comparing cv2, PIL, pytorch and others. We used PIL, which was slow but results outperformed any cv2 method. Using the `.thumbnail()` function seemed to help speed this up. The trade off was we did not generate any different sized images for modeling, but only did one down size of the image. \nAll images were reduced to 1152 dim. In addition, during training we applied augmentations prior to downsizing, as they would lose more information if applied to a small image. The augmentations are mainly cv2 based. Also, we kept the original filtered image aspect ratio and padded to a square image. \n\n## Augmentations\n\nTrain augmentations before downsizing were vflip, hflip,transpose, shift, scale, rotate, grid distortion & affine. After downsizing, we used one of random grid shuffle & coarse dropout. In train, we random cropped to 1024, and in val we center cropped. \nExample batch below.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fc28321c365bb1d679c2c35eb1c03e8f7%2FScreenshot%202023-02-28%20at%2018.09.06.png?generation=1677604161886568&alt=media)\n\n## Models\n\n### Model type 1: Breast level feature combination 1D-CNN with CNN backbone\n\nAuxiliary loss improved time to convergence a lot. For each backbone, stage 1 model trained for 9 epochs, stage 2 model trained for 2 epochs with frozen backbone. Heavy dropout (0.5) was applied on the linear output. \nThe 1D-CNN was set up with a filter size of 2 and no padding, so we get a (2, n_features) shape input and a (1, n_features) output. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F59667ff80e38d33611d00b97752a6aaa%2FScreenshot%202023-03-01%20at%2010.33.04.png?generation=1677663219589043&alt=media)\n\nIn our final ensemble we ran the above architecture with efficientnet b3 b4 b5 v2s and v2m.\n\nSome of the solution was inspired by Bo’s great write up on the [SIIM melanoma competition 1st place solution](https://www.kaggle.com/competitions/siim-isic-melanoma-classification/discussion/175412).\n\n### Model type 2: Patient-level multi-view-multi-lateral transformer with CNN backbone\n\nInspired by the paper [Transformers Improve Breast Cancer Diagnosis from Unregistered Multi-View Mammograms](https://arxiv.org/pdf/2206.10096.pdf), we designed another 2nd stage approach, but on a patient level. \n\n#### Stage 1:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2F809528c5c906c90e6752f2f6729f1f5c%2FScreenshot%202023-02-28%20at%2013.38.19.png?generation=1677604933375806&alt=media)\n\nIn the first stage we train an image level CNN which is trained with an auxiliary segmentation loss by predicting masks we got from training a Yolo_v7 on CBIS dataset.\n\n#### Stage 2:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1424766%2Fa908b5aad5f9bf7b0b1dc8ce77461eab%2FScreenshot%202023-02-28%20at%2013.38.10.png?generation=1677605002279914&alt=media)\n\nWe take the CNN trained in stage 1 and aggregate to patient level by considering four images per patient. Two views for each breast, i.e L-CC, L-MLO, R-CC, R-MLO. Using a 2-view/ 2-lateral input, the transformer can hopefully learn: \n\n- consistency between views \n- differences in laterals\n\nWe extract the output feature maps of the image level CNN for each view ending up with a tensors of size (4,out_channels, 32,32), which is projected to a predefined grid-size (4,hidden_dim, 16,16) using a conv layer. Then the 4x16x16 features, which can be seen as patch-tokens are spatially flattened and concatenated. The result is a sequence (hidden_dim, 1024) which is put into a transformer. Output of that transformer are two cancer predictions, one for each breast.\n\nWe first froze the CNN backbone for a few epochs to let the pretrained vision transformer \"adjust\" before fine-tuning the whole 2-stage model end2end. \nIf a patient has more than one image per view we randomly sample one image per view for training and use multiple 4-view combinations for inference which are then averaged. \nBackbone-wise we used pretrained seresnext50 and convnext tiny as CNN and pretrained deit-tiny-patch16-224 as transformer.\n\n##Ensembling\n\nSo in total we used single seed fullfits of the following models\n\n- Effnet v2_s + 1D-CNN\n- Effnet v2_m + 1D-CNN\n- Effnet b3 + 1D-CNN\n- Effnet b4 + 1D-CNN\n- Effnet b5 + 1D-CNN\n- SE-ResNext50 + deit-tiny-patch16-224\n- ConvNext_tiny + deit-tiny-patch16-224\n\nA threshold to convert to binary output was selected based on the CV of the blend. \nOur two highest submissions according to Public LB were also our two highest blends on CV. They were not the highest private, but close to it. For the shake up, it probably helped that we did not start blending the models until the last few days of the competition, so did not get distracted by high Public LB scores. \n\n##Things that did not work/were ineffective:\n\n- Pretraining on DDSM/VinDr-Mammo\n- Training a ROI extractor and training subsequent models on focused ROIs rather than whole images\n- Higher resolution images (similar to worse performance, longer inference times)\n\n",
    "2163674": "Very cool and beautiful solution, and congrats for your team!\n\nCould you offer more details about **Model type 1 Stage 2**, I've the following question:\n\n1. the input \"images from the breast\". In the stage 1, the model input is single image, but here, for the same model, the input becomes multi images, I assume that you are using the `batch size dimension` to compose the multi images?\n\n2. I do not quite understand the \"2D CNN for each 2D combination with the 1D output of same dim\", could please tell me more about this part?\n\n3. for \"Model type 2 stage 1\". This stage seems like has not have a strong correlation with \" \"Model type 2 stage 2\". I want to know if you have some experiments on how much this stage helps with the CV score? \n\nThanks in advance.",
    "2170877": "The transformer part is fancy. Congratulations and thank you! 🎉",
    "2163303": "Congratulations!!!, I'd like to know if it's allow to use Yolo during the competition.\n\nThank in advance",
    "2192544": "Hey, any idea of sharing your Source code? That would be really helpful. \nThank You.",
    "2170876": "Congratulations!\nI have two questions:\n1. What device did you use? Is it kaggle device or you own machine?\n2. What size of batch_size did you use?",
    "2170063": "\nI'm sorry if this is a dumb question, but how are the axillary losses supposed to help the model? it would be, for example, \"Can cancer show up differently for mammography Birads/Density type \"A\" and type \"B\" ? . What I'm trying to understand is what we're saying to the model when we use auxiliary losses along with the main loss (Cancer). Sorry if my question is not clear. Thanks !!! ",
    "2169074": "Congratulations team! Great job.",
    "2168540": "good project",
    "2165116": "Thanks for sharing and congrats to your team!",
    "2164450": "Congrats to your team and thank you for sharing the solution (looking forward to reading the code!). Loved the recommendation about PIL and TF antialias for resizing. Will read that paper!\n\nThanks so much for sharing!\n\n",
    "2164165": "Congratulations, a very strong solution!\nCould you share how the CV for the stage2 model was arranged?",
    "2163589": "Thanks for sharing and many congratulations to the entire team!\n\nAny chance you could share your CV score for the different models/ensembling?",
    "2163399": ">The cancer was often small and the images large, so when resizing it was important to lose as little detail as possible. This paper , concludes that the best resize method is PIL with lanczos, or tensorflow with antialias. There is a good example (Fig 1. in the paper) of how resizing can lose information, comparing cv2, PIL, pytorch and others\n\nThank you for this research highlight!\n\nAnd congratulations for the 4th place solution to you and your team!"
  }
}