{
  "id": 376879,
  "title": "[LB 0.47] MONAI training and inference pipeline",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/376879",
  "author_name": "Yiheng Wang",
  "post_date": "2023-01-09T03:26:25.631000",
  "votes": 71,
  "comment_count": 61,
  "views": 0,
  "content": "<p>In this topic, I will share my experiences in this challenge, and open source all of my works on how to get a 0.47 (the number may change) model.</p>\n<p>Inference code link:<br>\n<a href=\"https://www.kaggle.com/code/yiheng/monai-baseline\" target=\"_blank\">https://www.kaggle.com/code/yiheng/monai-baseline</a><br>\nTraining code link:<br>\n<a href=\"https://www.kaggle.com/code/yiheng/monai-pipeline-training/notebook\" target=\"_blank\">https://www.kaggle.com/code/yiheng/monai-pipeline-training/notebook</a></p>\n<p>you can also clone the training pipeline from github:</p>\n<pre><code>git clone git@github.com:yiheng-wang-nv/rsna_breast_monai_solution.git\n</code></pre>\n<p>The following is my local and lb results (will update if having any progresses):</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>pf1 local</th>\n<th>bin pf1 local</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold 0, agg by max</td>\n<td>0.2753</td>\n<td>0.3352</td>\n<td>0.38</td>\n</tr>\n<tr>\n<td>fold 0, agg by mean</td>\n<td>0.2042</td>\n<td>0.3392</td>\n<td>0.42</td>\n</tr>\n<tr>\n<td>fold 0, agg by mean, 2 tta</td>\n<td>0.2045</td>\n<td>0.3879</td>\n<td>0.44</td>\n</tr>\n<tr>\n<td>4 folds, agg by mean, 2 tta</td>\n<td>0.2234</td>\n<td>0.3129</td>\n<td>0.47</td>\n</tr>\n</tbody>\n</table>",
  "messages": [
    {
      "id": 2092068,
      "postDate": "2023-01-09T03:26:25.630Z",
      "content": "<p>In this topic, I will share my experiences in this challenge, and open source all of my works on how to get a 0.47 (the number may change) model.</p>\n<p>Inference code link:<br>\n<a href=\"https://www.kaggle.com/code/yiheng/monai-baseline\" target=\"_blank\">https://www.kaggle.com/code/yiheng/monai-baseline</a><br>\nTraining code link:<br>\n<a href=\"https://www.kaggle.com/code/yiheng/monai-pipeline-training/notebook\" target=\"_blank\">https://www.kaggle.com/code/yiheng/monai-pipeline-training/notebook</a></p>\n<p>you can also clone the training pipeline from github:</p>\n<pre><code>git clone git@github.com:yiheng-wang-nv/rsna_breast_monai_solution.git\n</code></pre>\n<p>The following is my local and lb results (will update if having any progresses):</p>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>pf1 local</th>\n<th>bin pf1 local</th>\n<th>LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>fold 0, agg by max</td>\n<td>0.2753</td>\n<td>0.3352</td>\n<td>0.38</td>\n</tr>\n<tr>\n<td>fold 0, agg by mean</td>\n<td>0.2042</td>\n<td>0.3392</td>\n<td>0.42</td>\n</tr>\n<tr>\n<td>fold 0, agg by mean, 2 tta</td>\n<td>0.2045</td>\n<td>0.3879</td>\n<td>0.44</td>\n</tr>\n<tr>\n<td>4 folds, agg by mean, 2 tta</td>\n<td>0.2234</td>\n<td>0.3129</td>\n<td>0.47</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "In this topic, I will share my experiences in this challenge, and open source all of my works on how to get a 0.47 (the number may change) model.\n\nInference code link:\nhttps://www.kaggle.com/code/yiheng/monai-baseline\nTraining code link:\nhttps://www.kaggle.com/code/yiheng/monai-pipeline-training/notebook\n\nyou can also clone the training pipeline from github:\n```\ngit clone git@github.com:yiheng-wang-nv/rsna_breast_monai_solution.git\n```\n\nThe following is my local and lb results (will update if having any progresses):\n\n| | pf1 local | bin pf1 local | LB |\n| --- | --- | --- | --- |\n| fold 0, agg by max | 0.2753 | 0.3352 | 0.38 |\n| fold 0, agg by mean | 0.2042 | 0.3392 | 0.42 |\n| fold 0, agg by mean, 2 tta | 0.2045 | 0.3879 | 0.44 |\n| 4 folds, agg by mean, 2 tta | 0.2234 | 0.3129 | 0.47 |",
      "votes": 71
    },
    {
      "id": 2092074,
      "postDate": "2023-01-09T03:46:48.203Z",
      "content": "<p>[Update 1] My purpose is to make use of MONAI, as well as other libs (such as timm for network, dali for acceleration) to build a decent model.<br>\nFirst of all, I wish to thank <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> , I followed <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369769\" target=\"_blank\">your topic</a> and <a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-inference\" target=\"_blank\">your code</a> to design my training pipeline, which can be summarized as the following settings:</p>\n<ul>\n<li>group stratified 4 fold data split</li>\n<li>512 * 512 input size</li>\n<li>h flip augmentation</li>\n<li>4 epoch, lr decay from 3e-4</li>\n<li>timm, efnv2s network</li>\n</ul>\n<p>Based on these settings, I started to do experiments, and my initial purpose is to reach pf1 ~0.083 and auc 0.755 (these two values are from theoviel's open sourced notebook I mentioned above). What I tuned include:</p>\n<ol>\n<li>number of epochs</li>\n<li>pos sample weights on bce loss</li>\n<li>drop rate of efficientnet (this idea is achieved from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> in <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333\" target=\"_blank\">his topic</a><br>\nafter ~ 8 experiments, I got pf1 0.10475 on one fold locally. Then I switched to use larger input size: 1024 * 1024 and do all 4 folds experiments. In this step, I did not tune any parameters, but just try to finish training, and then submit to LB.</li>\n</ol>",
      "rawMarkdown": "[Update 1] My purpose is to make use of MONAI, as well as other libs (such as timm for network, dali for acceleration) to build a decent model.\nFirst of all, I wish to thank @theoviel , I followed [your topic](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369769) and [your code](https://www.kaggle.com/code/theoviel/rsna-breast-baseline-inference) to design my training pipeline, which can be summarized as the following settings:\n\n- group stratified 4 fold data split\n- 512 * 512 input size\n- h flip augmentation\n- 4 epoch, lr decay from 3e-4\n- timm, efnv2s network\n\nBased on these settings, I started to do experiments, and my initial purpose is to reach pf1 ~0.083 and auc 0.755 (these two values are from theoviel's open sourced notebook I mentioned above). What I tuned include:\n\n1. number of epochs\n2. pos sample weights on bce loss\n3. drop rate of efficientnet (this idea is achieved from @hengck23 in [his topic](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333)\nafter ~ 8 experiments, I got pf1 0.10475 on one fold locally. Then I switched to use larger input size: 1024 * 1024 and do all 4 folds experiments. In this step, I did not tune any parameters, but just try to finish training, and then submit to LB.",
      "votes": 13,
      "replies": [
        {
          "id": 2092075,
          "postDate": "2023-01-09T03:49:50.580Z",
          "content": "<p>My next plan:</p>\n<ol>\n<li>try to train a 2048 * 2048 model, which I expect a ~0.51 LB result as <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared : )</li>\n<li>adjust the input, since there are many black area in each figure, which may be useless.</li>\n</ol>",
          "rawMarkdown": "My next plan:\n1. try to train a 2048 * 2048 model, which I expect a ~0.51 LB result as @hengck23 shared : )\n2. adjust the input, since there are many black area in each figure, which may be useless.",
          "votes": 6,
          "replies": [
            {
              "id": 2092470,
              "postDate": "2023-01-09T11:03:51.207Z",
              "content": "<p>\"adjust the input, since there are many black area in each figure\"<br>\n<a href=\"https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region\" target=\"_blank\">https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region</a></p>\n<p>i think this will take about 10 min to process the all kaggle test images</p>",
              "rawMarkdown": "\"adjust the input, since there are many black area in each figure\"\nhttps://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region\n\ni think this will take about 10 min to process the all kaggle test images",
              "votes": 3
            },
            {
              "id": 2095384,
              "postDate": "2023-01-11T10:14:51.273Z",
              "content": "<p>Now I give up to train 2048 size model because of limited resources, and I will focus on roi based method as others suggested.<br>\nTo get a quick experiment, I will try to use the dataset provided by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> in his topic: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369754\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369754</a> (Thanks Remek!) locally, and if getting any progresses I will enhance my inference notebook and share my results.</p>",
              "rawMarkdown": "Now I give up to train 2048 size model because of limited resources, and I will focus on roi based method as others suggested.\nTo get a quick experiment, I will try to use the dataset provided by @remekkinas in his topic: https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369754 (Thanks Remek!) locally, and if getting any progresses I will enhance my inference notebook and share my results.",
              "votes": 1
            },
            {
              "id": 2096719,
              "postDate": "2023-01-12T08:19:59.647Z",
              "content": "<p>if you are using crop strategy, do remember to take care of the case if no box can be detected. i spend a few days to debug this submission error.</p>\n<pre><code>box = run_box_detection_model(image)\nif box is None:\n      use whole image\nelse:\n       crop image\n</code></pre>",
              "rawMarkdown": "if you are using crop strategy, do remember to take care of the case if no box can be detected. i spend a few days to debug this submission error.\n\n```\nbox = run_box_detection_model(image)\nif box is None:\n      use whole image\nelse:\n       crop image\n\n\n```",
              "votes": 4
            },
            {
              "id": 2097348,
              "postDate": "2023-01-12T15:57:13.713Z",
              "content": "<p>and … use </p>\n<pre><code>try:\n....\nexcept Exception as e \n... rollback operations in try\n</code></pre>\n<p>I have one place in code which is unstable - probably in test DS is dicom file which break rules from train. I have spent a lot of time debugging - have not found any problem. Then decided to wrap code in try except and for this image rollback all operations - use oryginal image instead of cropped and postprocessed.</p>",
              "rawMarkdown": "and ... use \n\n```\ntry:\n....\nexcept Exception as e \n... rollback operations in try\n```\nI have one place in code which is unstable - probably in test DS is dicom file which break rules from train. I have spent a lot of time debugging - have not found any problem. Then decided to wrap code in try except and for this image rollback all operations - use oryginal image instead of cropped and postprocessed."
            },
            {
              "id": 2101521,
              "postDate": "2023-01-16T01:11:18.770Z",
              "content": "<p>Which part of Heng's code did you wrap? I tried it but sub still fails. </p>",
              "rawMarkdown": "Which part of Heng's code did you wrap? I tried it but sub still fails. "
            }
          ]
        },
        {
          "id": 2092868,
          "postDate": "2023-01-09T17:01:04.390Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yiheng\" target=\"_blank\">@yiheng</a> </p>\n<p>Thanks for sharing your insights.. Can you mention the basis for your CV split.. </p>\n<p>I am currently using StratifiedGroupKFold (5 folds) with patient_id for grouping</p>",
          "rawMarkdown": "Hi @yiheng \n\nThanks for sharing your insights.. Can you mention the basis for your CV split.. \n\nI am currently using StratifiedGroupKFold (5 folds) with patient_id for grouping",
          "replies": [
            {
              "id": 2093695,
              "postDate": "2023-01-10T08:19:05.527Z",
              "content": "<p>Similarly, I also used StratifiedGroupKFold with k=4 (the same as Theo does in his notebook)</p>",
              "rawMarkdown": "Similarly, I also used StratifiedGroupKFold with k=4 (the same as Theo does in his notebook)",
              "votes": 2
            }
          ]
        },
        {
          "id": 2100708,
          "postDate": "2023-01-15T11:27:02.840Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yiheng\" target=\"_blank\">@yiheng</a> I want to know when ensembling 4 folds, which validation set do you use to choose threshold to binarize ?</p>",
          "rawMarkdown": "Hi @yiheng I want to know when ensembling 4 folds, which validation set do you use to choose threshold to binarize ?",
          "replies": [
            {
              "id": 2101783,
              "postDate": "2023-01-16T06:54:52.707Z",
              "content": "<p>Hi, the whole dataset. I used each fold's model to do predictions on each validation set, and then combine all predictions together to search the threshold.</p>",
              "rawMarkdown": "Hi, the whole dataset. I used each fold's model to do predictions on each validation set, and then combine all predictions together to search the threshold.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2100083,
      "postDate": "2023-01-14T23:32:17.640Z",
      "content": "<p>Great Job!</p>",
      "rawMarkdown": "Great Job!",
      "votes": 1
    },
    {
      "id": 2092163,
      "postDate": "2023-01-09T06:42:59.137Z",
      "content": "<p>Are you using Kaggle for training, some other cloud provider, or local GPU setup?</p>",
      "rawMarkdown": "Are you using Kaggle for training, some other cloud provider, or local GPU setup?",
      "votes": 1,
      "replies": [
        {
          "id": 2092179,
          "postDate": "2023-01-09T07:17:24.047Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/aadimator\" target=\"_blank\">@aadimator</a> , I use local GPUs</p>",
          "rawMarkdown": "Hi @aadimator , I use local GPUs",
          "votes": 1
        }
      ]
    },
    {
      "id": 2104971,
      "postDate": "2023-01-18T07:17:55.217Z",
      "content": "<p>[Update 2] I used <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's trained model (in: <a href=\"https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region\" target=\"_blank\">https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region</a>) to infer the bounding boxes for all training images and then do the following preprocessing:<br>\ndicom ---&gt; crop by bounding box ---&gt; resize to 1024 * 1024<br>\nthe resized dataset is uploaded to: <a href=\"https://www.kaggle.com/datasets/yiheng/rsna-breast-1024\" target=\"_blank\">https://www.kaggle.com/datasets/yiheng/rsna-breast-1024</a></p>\n<p>Now I'm training with this dataset, if having any progresses, I will update the topic and the notebook.<br>\nThe next plans are:</p>\n<ol>\n<li>train with different backbone</li>\n<li>train with different input size</li>\n<li>optimize the inference time</li>\n</ol>",
      "rawMarkdown": "[Update 2] I used @hengck23 's trained model (in: https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region) to infer the bounding boxes for all training images and then do the following preprocessing:\ndicom ---> crop by bounding box ---> resize to 1024 * 1024\nthe resized dataset is uploaded to: https://www.kaggle.com/datasets/yiheng/rsna-breast-1024\n\nNow I'm training with this dataset, if having any progresses, I will update the topic and the notebook.\nThe next plans are:\n1. train with different backbone\n2. train with different input size\n3. optimize the inference time",
      "replies": [
        {
          "id": 2110936,
          "postDate": "2023-01-22T14:39:57.640Z",
          "content": "<p>remember to use try except to capture dengerate case at submissiom<br>\n<a href=\"https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example/comments#2104692\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example/comments#2104692</a></p>",
          "rawMarkdown": "remember to use try except to capture dengerate case at submissiom\nhttps://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example/comments#2104692",
          "votes": 1
        }
      ]
    },
    {
      "id": 2102690,
      "postDate": "2023-01-16T19:37:33.720Z",
      "content": "<p>Hi! Thanks for your work!<br>\nI tried to reproduce and set batch_size to 16 due to my single GPU… and I find that the eval loss is increasing instead of decreasing…<br>\nShould I decreas the lr for my smaller batch_size?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi! Thanks for your work!\nI tried to reproduce and set batch_size to 16 due to my single GPU... and I find that the eval loss is increasing instead of decreasing...\nShould I decreas the lr for my smaller batch_size?\n\nThanks!",
      "replies": [
        {
          "id": 2102835,
          "postDate": "2023-01-16T21:29:17.630Z",
          "content": "<p>I tried this as well and failed, the model didn't converge, to make effnetv2 train I had to enable gradient checkpointing and keep batch size at 64, and the easier thing that I didn't test is to lower the learning rate, maybe it'll work comparatively to keeping batch size large</p>",
          "rawMarkdown": "I tried this as well and failed, the model didn't converge, to make effnetv2 train I had to enable gradient checkpointing and keep batch size at 64, and the easier thing that I didn't test is to lower the learning rate, maybe it'll work comparatively to keeping batch size large",
          "replies": [
            {
              "id": 2102910,
              "postDate": "2023-01-16T22:06:47.737Z",
              "content": "<p>I tried many models and I found effnetv2 is really hard to train… The eval loss often didn't converge…<br>\nI will try to decrease the lr from 3e-4 to 1e-4, then maybe 5e-5 to have a look.<br>\nBut as you said, i don't know if this only works for large batch size</p>",
              "rawMarkdown": "I tried many models and I found effnetv2 is really hard to train... The eval loss often didn't converge...\nI will try to decrease the lr from 3e-4 to 1e-4, then maybe 5e-5 to have a look.\nBut as you said, i don't know if this only works for large batch size"
            }
          ]
        },
        {
          "id": 2103174,
          "postDate": "2023-01-17T04:00:31.857Z",
          "content": "<p>could you try large batch size with smaller image size first? Or maybe you can upsample the positive samples, I guess a small batch size may make the task hard since there may have many batches that have no pos samples.</p>",
          "rawMarkdown": "could you try large batch size with smaller image size first? Or maybe you can upsample the positive samples, I guess a small batch size may make the task hard since there may have many batches that have no pos samples.",
          "replies": [
            {
              "id": 2103803,
              "postDate": "2023-01-17T11:58:52.540Z",
              "content": "<p>Thanks! now I'm trying to use upsampling and also gradient accumulation to have a try!</p>",
              "rawMarkdown": "Thanks! now I'm trying to use upsampling and also gradient accumulation to have a try!"
            },
            {
              "id": 2104315,
              "postDate": "2023-01-17T18:40:59.550Z",
              "content": "<p>I tried with upsampling the pos to 1/8 of each batch and gradient accumulation to get 64 a batch.<br>\nBut I still get the overfitting problems and cannot even close to your score…</p>",
              "rawMarkdown": "I tried with upsampling the pos to 1/8 of each batch and gradient accumulation to get 64 a batch.\nBut I still get the overfitting problems and cannot even close to your score..."
            },
            {
              "id": 2104979,
              "postDate": "2023-01-18T07:22:45.210Z",
              "content": "<p>what is your val score? Have you tried to train on 512 * 512 size? It should get a ~0.1 val score.</p>",
              "rawMarkdown": "what is your val score? Have you tried to train on 512 * 512 size? It should get a ~0.1 val score."
            }
          ]
        }
      ]
    },
    {
      "id": 2099092,
      "postDate": "2023-01-14T07:01:16.593Z",
      "content": "<p>Great Job!<br>\nSimple &amp; Impressive for your sharing.</p>\n<p>Could you get my 3 simple questions?</p>\n<ol>\n<li><p>As following your training config files, the pos_weight in loss function set to 1.<br>\nDid your 0.47 LB score come from it? or other value?<br>\nI'm just wondering because the pos_weight in most of other codes set to range of 10~50.</p></li>\n<li><p>In validation for each folds, how did you choose the best epoch weight?<br>\n2 scores(bin pfbeta1(thres 0.1) &amp; auc) can be printed.<br>\nwhich scores did you follow?</p></li>\n<li><p>In your experiment, Was just 10 epoch training enough to get the result?</p></li>\n</ol>\n<p>Thank you!</p>",
      "rawMarkdown": "Great Job!\nSimple & Impressive for your sharing.\n\nCould you get my 3 simple questions?\n1. As following your training config files, the pos_weight in loss function set to 1.\nDid your 0.47 LB score come from it? or other value?\nI'm just wondering because the pos_weight in most of other codes set to range of 10~50.\n\n2. In validation for each folds, how did you choose the best epoch weight?\n2 scores(bin pfbeta1(thres 0.1) & auc) can be printed.\nwhich scores did you follow?\n\n3. In your experiment, Was just 10 epoch training enough to get the result?\n\nThank you!",
      "replies": [
        {
          "id": 2101788,
          "postDate": "2023-01-16T06:56:48.293Z",
          "content": "<ol>\n<li>Yes, I tried other pos weights but the final result is similar. 1 works on my solution.</li>\n<li>I select the best pf1 (non binariezed) weights.</li>\n<li>I think 10 epoch is enough, I also tried 20, but the total val score is similar.</li>\n</ol>",
          "rawMarkdown": "1. Yes, I tried other pos weights but the final result is similar. 1 works on my solution.\n2. I select the best pf1 (non binariezed) weights.\n3. I think 10 epoch is enough, I also tried 20, but the total val score is similar.",
          "votes": 1,
          "replies": [
            {
              "id": 2105064,
              "postDate": "2023-01-18T08:42:03.500Z",
              "content": "<p>As for 3, I detected a new finding.<br>\nAccording to my latest experiments on cropped images, it's easier for training set to overfit, and reduce epoch number from 10 to 5 is enough to get similar val score and also lower train score (less overfit).</p>",
              "rawMarkdown": "As for 3, I detected a new finding.\nAccording to my latest experiments on cropped images, it's easier for training set to overfit, and reduce epoch number from 10 to 5 is enough to get similar val score and also lower train score (less overfit).",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2097286,
      "postDate": "2023-01-12T14:57:56.947Z",
      "content": "<p>Thank you for sharing your work! As you already said, you use StratifiedGroupKFold with k=4 and binarize f1 to get the final prediction. How do you find the threshold to use in the submissions? As you compute the mean along all the folds, do you compute the mean of all the thresholds?</p>",
      "rawMarkdown": "Thank you for sharing your work! As you already said, you use StratifiedGroupKFold with k=4 and binarize f1 to get the final prediction. How do you find the threshold to use in the submissions? As you compute the mean along all the folds, do you compute the mean of all the thresholds?",
      "replies": [
        {
          "id": 2104975,
          "postDate": "2023-01-18T07:20:29.123Z",
          "content": "<p>I search the threshold based on all predictions (from all folds models).</p>",
          "rawMarkdown": "I search the threshold based on all predictions (from all folds models)."
        }
      ]
    },
    {
      "id": 2095373,
      "postDate": "2023-01-11T10:08:45.247Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yiheng\" target=\"_blank\">@yiheng</a> . Are you also tracking AUC? If so, could you report it here? Thanks. </p>",
      "rawMarkdown": "Hi @yiheng . Are you also tracking AUC? If so, could you report it here? Thanks. ",
      "replies": [
        {
          "id": 2095389,
          "postDate": "2023-01-11T10:18:57.003Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> , AUC is printed in my training pipeline however I did not record it in my sheet : (<br>\nI will add this part in my following experiments, and post my updates after getting a better result.</p>",
          "rawMarkdown": "Hi @pheadrus , AUC is printed in my training pipeline however I did not record it in my sheet : (\nI will add this part in my following experiments, and post my updates after getting a better result.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2095341,
      "postDate": "2023-01-11T09:44:33.350Z",
      "content": "<p><a href=\"https://www.kaggle.com/yiheng\" target=\"_blank\">@yiheng</a> <br>\nDid you make any resampling (under or over)? And how much your fbeta fluctuated during training time on validation set? Did you get 0.10475 in a consistent progress? </p>",
      "rawMarkdown": "@yiheng \nDid you make any resampling (under or over)? And how much your fbeta fluctuated during training time on validation set? Did you get 0.10475 in a consistent progress? ",
      "replies": [
        {
          "id": 2095390,
          "postDate": "2023-01-11T10:20:21.073Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/simonalerdic\" target=\"_blank\">@simonalerdic</a> , in 512 * 512 experiments, val score increases consistently.<br>\nSo far I did not try resampling but just change the pos weight in bce loss.</p>",
          "rawMarkdown": "Hi @simonalerdic , in 512 * 512 experiments, val score increases consistently.\nSo far I did not try resampling but just change the pos weight in bce loss.",
          "votes": 1,
          "replies": [
            {
              "id": 2096561,
              "postDate": "2023-01-12T06:17:58.130Z",
              "content": "<p>Is it possible to reproduce your training notebook score on kaggle env? Have you tried? In here, model size and image size, batch size gets limited due to resource and adjusting them from yours gave poor result. Could you please confirm?</p>",
              "rawMarkdown": "Is it possible to reproduce your training notebook score on kaggle env? Have you tried? In here, model size and image size, batch size gets limited due to resource and adjusting them from yours gave poor result. Could you please confirm?"
            },
            {
              "id": 2105047,
              "postDate": "2023-01-18T08:27:34.160Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/simonalerdic\" target=\"_blank\">@simonalerdic</a> , I'm afraid I cannot do this kind of works recently since we have limited time for this challenge. I prefer finishing my current experiments first:</p>\n<ol>\n<li>try training with cropped images</li>\n<li>try different networks</li>\n<li>try different input size</li>\n</ol>\n<p>If having 512*512 results, I will post my log here if it is helpful.</p>",
              "rawMarkdown": "Hi @simonalerdic , I'm afraid I cannot do this kind of works recently since we have limited time for this challenge. I prefer finishing my current experiments first:\n1. try training with cropped images\n2. try different networks\n3. try different input size\n\nIf having 512*512 results, I will post my log here if it is helpful."
            }
          ]
        }
      ]
    },
    {
      "id": 2095096,
      "postDate": "2023-01-11T07:43:59.407Z",
      "content": "<p>Is the batch size an important parameter? I reduced batch size to avoid out of memory, but the validation score will be very low.</p>",
      "rawMarkdown": "Is the batch size an important parameter? I reduced batch size to avoid out of memory, but the validation score will be very low.",
      "replies": [
        {
          "id": 2095376,
          "postDate": "2023-01-11T10:10:02.600Z",
          "content": "<p>I think it has some impacts, since the percentage of pos samples is low. Maybe you can try to accumulate gradients if not having enough gpu memory.</p>",
          "rawMarkdown": "I think it has some impacts, since the percentage of pos samples is low. Maybe you can try to accumulate gradients if not having enough gpu memory.",
          "votes": 3,
          "replies": [
            {
              "id": 2103574,
              "postDate": "2023-01-17T08:43:27.883Z",
              "content": "<p>Thank you!</p>",
              "rawMarkdown": "Thank you!"
            }
          ]
        },
        {
          "id": 2095496,
          "postDate": "2023-01-11T12:02:35.327Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/juntaojiang\" target=\"_blank\">@juntaojiang</a> , batch size is really trick, while the relationship between BS and LR is also very important.</p>\n<ol>\n<li>When you reduce BS, try to scale LR at the same time. For example, the original BS is 32 and the LR is 1e-4, and you want to reduce the BS to 16, try to reduce LR to 5e-5. If you have more time, you can search the optimal LR around 5e-5.</li>\n<li>If you use multiple GPU to train, never put only one image per gpus. For example, if you use 8 cards, your minimum batch size should be 16, while there are 2 images per gpus. If you set BS=8, then there is 1 image per gpu, which will cause the model to perform much worse. It's caused by  the BatchNorm in most CNN models.</li>\n</ol>",
          "rawMarkdown": "Hi @juntaojiang , batch size is really trick, while the relationship between BS and LR is also very important.\n\n1. When you reduce BS, try to scale LR at the same time. For example, the original BS is 32 and the LR is 1e-4, and you want to reduce the BS to 16, try to reduce LR to 5e-5. If you have more time, you can search the optimal LR around 5e-5.\n2. If you use multiple GPU to train, never put only one image per gpus. For example, if you use 8 cards, your minimum batch size should be 16, while there are 2 images per gpus. If you set BS=8, then there is 1 image per gpu, which will cause the model to perform much worse. It's caused by  the BatchNorm in most CNN models.",
          "votes": 10,
          "replies": [
            {
              "id": 2103572,
              "postDate": "2023-01-17T08:43:15.090Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>  Thank you for your kind suggestion! I will have a try.</p>",
              "rawMarkdown": "Hi @forcewithme  Thank you for your kind suggestion! I will have a try."
            }
          ]
        }
      ]
    },
    {
      "id": 2095041,
      "postDate": "2023-01-11T07:04:35.680Z",
      "content": "<p>good information…</p>",
      "rawMarkdown": "good information..."
    },
    {
      "id": 2095019,
      "postDate": "2023-01-11T06:44:23.623Z",
      "content": "<p>Nice work! btw, how do u determine <code>clf_threshold = 0.1</code>?</p>",
      "rawMarkdown": "Nice work! btw, how do u determine `clf_threshold = 0.1`?",
      "replies": [
        {
          "id": 2095347,
          "postDate": "2023-01-11T09:50:54.737Z",
          "content": "<p>In the training step, I just manually set it to 0.1 to take a look (avoid doing search on every epoch). <br>\nThe final threshold I used is searched on all 4 folds predictions.</p>",
          "rawMarkdown": "In the training step, I just manually set it to 0.1 to take a look (avoid doing search on every epoch). \nThe final threshold I used is searched on all 4 folds predictions.",
          "votes": 1,
          "replies": [
            {
              "id": 2095377,
              "postDate": "2023-01-11T10:10:28.667Z",
              "content": "<p>Got it, thx!</p>",
              "rawMarkdown": "Got it, thx!"
            }
          ]
        }
      ]
    },
    {
      "id": 2094563,
      "postDate": "2023-01-10T20:44:22.340Z",
      "content": "<p>Excellent post, very helpful and concise. Thanks for sharing!</p>",
      "rawMarkdown": "Excellent post, very helpful and concise. Thanks for sharing!"
    },
    {
      "id": 2092574,
      "postDate": "2023-01-09T12:36:00.047Z",
      "content": "<p>Great findings, thanks for sharing!<br>\nHow much VRAM do you have for 1024x1024 input images with EffNetV2S? 👀</p>",
      "rawMarkdown": "Great findings, thanks for sharing!\nHow much VRAM do you have for 1024x1024 input images with EffNetV2S? 👀",
      "replies": [
        {
          "id": 2092592,
          "postDate": "2023-01-09T12:44:38.877Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/davidlandup\" target=\"_blank\">@davidlandup</a> , for training, I used batch size = 64 on 4 * 32G gpus.</p>",
          "rawMarkdown": "Hi @davidlandup , for training, I used batch size = 64 on 4 * 32G gpus.",
          "votes": 1,
          "replies": [
            {
              "id": 2092606,
              "postDate": "2023-01-09T12:50:00.063Z",
              "content": "<p>Thanks for the quick response! Packing some serious compute 😨</p>",
              "rawMarkdown": "Thanks for the quick response! Packing some serious compute 😨",
              "votes": 1
            },
            {
              "id": 2092648,
              "postDate": "2023-01-09T13:24:27.417Z",
              "content": "<p>today I tried to train a 2048 * 2048 model but failed. My GPUs are also not enough, but I think it's interesting to find better ways to optimize our solution with limited resources. I think crop foreground is a good way to avoid using large size, maybe you can try it, and I will also have a try and share my findings later.</p>",
              "rawMarkdown": "today I tried to train a 2048 * 2048 model but failed. My GPUs are also not enough, but I think it's interesting to find better ways to optimize our solution with limited resources. I think crop foreground is a good way to avoid using large size, maybe you can try it, and I will also have a try and share my findings later."
            },
            {
              "id": 2093767,
              "postDate": "2023-01-10T09:58:14.953Z",
              "content": "<p>I wanted to create a baseline before cropping. Locally, I had an unreasonably high validation AUC (97%) which is a clear indication of a messed up data pipeline.<br>\nI'm currently getting ~0.06 pF1 (75% val AUC)  locally on small images (224x224).</p>\n<p>Not sure if the same pipeline on larger images will translate to a much higher pF1, but I'd like to get it up to ~0.1 before kicking off a 1024x1024 or 2048x2048 run, at which point ROI extraction will likely play a much larger role. I'll probably experiment with cropping in the next couple of days to see if that helps get a higher base pF1…</p>\n<p>Posting updates if there are any :)</p>",
              "rawMarkdown": "I wanted to create a baseline before cropping. Locally, I had an unreasonably high validation AUC (97%) which is a clear indication of a messed up data pipeline.\nI'm currently getting ~0.06 pF1 (75% val AUC)  locally on small images (224x224).\n\nNot sure if the same pipeline on larger images will translate to a much higher pF1, but I'd like to get it up to ~0.1 before kicking off a 1024x1024 or 2048x2048 run, at which point ROI extraction will likely play a much larger role. I'll probably experiment with cropping in the next couple of days to see if that helps get a higher base pF1...\n\nPosting updates if there are any :)",
              "votes": 2
            },
            {
              "id": 2094047,
              "postDate": "2023-01-10T15:07:42.807Z",
              "content": "<p>In my experiments,  I found that increasing the image size improved by quite a lot the results. But as you said, I would suggest having a strong pipeline and playing with it. ROI extraction also played for me a huge improvement.</p>",
              "rawMarkdown": "In my experiments,  I found that increasing the image size improved by quite a lot the results. But as you said, I would suggest having a strong pipeline and playing with it. ROI extraction also played for me a huge improvement.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2092484,
      "postDate": "2023-01-09T11:17:29.120Z",
      "content": "<p>What's the inference time on the 4 fold 1024 * 1024 setup?</p>",
      "rawMarkdown": "What's the inference time on the 4 fold 1024 * 1024 setup?",
      "replies": [
        {
          "id": 2092591,
          "postDate": "2023-01-09T12:43:53.213Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kiranik\" target=\"_blank\">@kiranik</a> , sorry I did not record the time as I submitted before going to bed : ) Is there any easy way to check the inference time?</p>",
          "rawMarkdown": "Hi @kiranik , sorry I did not record the time as I submitted before going to bed : ) Is there any easy way to check the inference time?",
          "votes": 1,
          "replies": [
            {
              "id": 2092624,
              "postDate": "2023-01-09T13:03:21.177Z",
              "content": "<p>Alright, yes it's not easy, only way is to manually check if the personally schedule allow at the time and get some ~time intervall for a new setup,  some also use the train set for time and memory check. </p>",
              "rawMarkdown": "Alright, yes it's not easy, only way is to manually check if the personally schedule allow at the time and get some ~time intervall for a new setup,  some also use the train set for time and memory check. ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2116345,
      "postDate": "2023-01-26T12:38:28.017Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2115421,
      "postDate": "2023-01-25T18:50:37.333Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2110815,
      "postDate": "2023-01-22T12:52:09.127Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2095271,
      "postDate": "2023-01-11T09:04:37.840Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2093668,
      "postDate": "2023-01-10T07:53:42.943Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2093522,
      "postDate": "2023-01-10T04:10:12.667Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2097241,
      "postDate": "2023-01-12T14:32:09.333Z",
      "content": "<p>Thanks for sharing your insights</p>",
      "rawMarkdown": "Thanks for sharing your insights\n"
    },
    {
      "id": 2095106,
      "postDate": "2023-01-11T07:51:37.493Z",
      "content": "<p>thank you!</p>",
      "rawMarkdown": "thank you!"
    }
  ],
  "comments": [
    {
      "id": 2092074,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2023-01-09T03:46:48.203000",
      "content": "<p>[Update 1] My purpose is to make use of MONAI, as well as other libs (such as timm for network, dali for acceleration) to build a decent model.<br>\nFirst of all, I wish to thank <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">@theoviel</a> , I followed <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369769\" target=\"_blank\">your topic</a> and <a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-inference\" target=\"_blank\">your code</a> to design my training pipeline, which can be summarized as the following settings:</p>\n<ul>\n<li>group stratified 4 fold data split</li>\n<li>512 * 512 input size</li>\n<li>h flip augmentation</li>\n<li>4 epoch, lr decay from 3e-4</li>\n<li>timm, efnv2s network</li>\n</ul>\n<p>Based on these settings, I started to do experiments, and my initial purpose is to reach pf1 ~0.083 and auc 0.755 (these two values are from theoviel's open sourced notebook I mentioned above). What I tuned include:</p>\n<ol>\n<li>number of epochs</li>\n<li>pos sample weights on bce loss</li>\n<li>drop rate of efficientnet (this idea is achieved from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> in <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333\" target=\"_blank\">his topic</a><br>\nafter ~ 8 experiments, I got pf1 0.10475 on one fold locally. Then I switched to use larger input size: 1024 * 1024 and do all 4 folds experiments. In this step, I did not tune any parameters, but just try to finish training, and then submit to LB.</li>\n</ol>",
      "votes": 13,
      "replies": [
        {
          "id": 2092075,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-09T03:49:50.580000",
          "content": "<p>My next plan:</p>\n<ol>\n<li>try to train a 2048 * 2048 model, which I expect a ~0.51 LB result as <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> shared : )</li>\n<li>adjust the input, since there are many black area in each figure, which may be useless.</li>\n</ol>",
          "votes": 6,
          "replies": [
            {
              "id": 2092470,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-09T11:03:51.207000",
              "content": "<p>\"adjust the input, since there are many black area in each figure\"<br>\n<a href=\"https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region\" target=\"_blank\">https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region</a></p>\n<p>i think this will take about 10 min to process the all kaggle test images</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2095384,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2023-01-11T10:14:51.273000",
              "content": "<p>Now I give up to train 2048 size model because of limited resources, and I will focus on roi based method as others suggested.<br>\nTo get a quick experiment, I will try to use the dataset provided by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> in his topic: <a href=\"https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369754\" target=\"_blank\">https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369754</a> (Thanks Remek!) locally, and if getting any progresses I will enhance my inference notebook and share my results.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2096719,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-01-12T08:19:59.647000",
              "content": "<p>if you are using crop strategy, do remember to take care of the case if no box can be detected. i spend a few days to debug this submission error.</p>\n<pre><code>box = run_box_detection_model(image)\nif box is None:\n      use whole image\nelse:\n       crop image\n</code></pre>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2097348,
              "author_name": "Remek Kinas",
              "author_url": "",
              "post_date": "2023-01-12T15:57:13.713000",
              "content": "<p>and … use </p>\n<pre><code>try:\n....\nexcept Exception as e \n... rollback operations in try\n</code></pre>\n<p>I have one place in code which is unstable - probably in test DS is dicom file which break rules from train. I have spent a lot of time debugging - have not found any problem. Then decided to wrap code in try except and for this image rollback all operations - use oryginal image instead of cropped and postprocessed.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2101521,
              "author_name": "Phaedrus",
              "author_url": "",
              "post_date": "2023-01-16T01:11:18.770000",
              "content": "<p>Which part of Heng's code did you wrap? I tried it but sub still fails. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2092868,
          "author_name": "Balaji Selvaraj",
          "author_url": "",
          "post_date": "2023-01-09T17:01:04.390000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yiheng\" target=\"_blank\">@yiheng</a> </p>\n<p>Thanks for sharing your insights.. Can you mention the basis for your CV split.. </p>\n<p>I am currently using StratifiedGroupKFold (5 folds) with patient_id for grouping</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2093695,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2023-01-10T08:19:05.527000",
              "content": "<p>Similarly, I also used StratifiedGroupKFold with k=4 (the same as Theo does in his notebook)</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2100708,
          "author_name": "i_love_huyen_tran",
          "author_url": "",
          "post_date": "2023-01-15T11:27:02.840000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/yiheng\" target=\"_blank\">@yiheng</a> I want to know when ensembling 4 folds, which validation set do you use to choose threshold to binarize ?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2101783,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2023-01-16T06:54:52.707000",
              "content": "<p>Hi, the whole dataset. I used each fold's model to do predictions on each validation set, and then combine all predictions together to search the threshold.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2100083,
      "author_name": "AbanoubSamir004",
      "author_url": "",
      "post_date": "2023-01-14T23:32:17.640000",
      "content": "<p>Great Job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2092163,
      "author_name": "Aadam",
      "author_url": "",
      "post_date": "2023-01-09T06:42:59.137000",
      "content": "<p>Are you using Kaggle for training, some other cloud provider, or local GPU setup?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2092179,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-09T07:17:24.047000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/aadimator\" target=\"_blank\">@aadimator</a> , I use local GPUs</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2104971,
      "author_name": "Yiheng Wang",
      "author_url": "",
      "post_date": "2023-01-18T07:17:55.217000",
      "content": "<p>[Update 2] I used <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's trained model (in: <a href=\"https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region\" target=\"_blank\">https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region</a>) to infer the bounding boxes for all training images and then do the following preprocessing:<br>\ndicom ---&gt; crop by bounding box ---&gt; resize to 1024 * 1024<br>\nthe resized dataset is uploaded to: <a href=\"https://www.kaggle.com/datasets/yiheng/rsna-breast-1024\" target=\"_blank\">https://www.kaggle.com/datasets/yiheng/rsna-breast-1024</a></p>\n<p>Now I'm training with this dataset, if having any progresses, I will update the topic and the notebook.<br>\nThe next plans are:</p>\n<ol>\n<li>train with different backbone</li>\n<li>train with different input size</li>\n<li>optimize the inference time</li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 2110936,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-01-22T14:39:57.640000",
          "content": "<p>remember to use try except to capture dengerate case at submissiom<br>\n<a href=\"https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example/comments#2104692\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3hr-tensorrt-nextvit-example/comments#2104692</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2102690,
      "author_name": "nicehzj",
      "author_url": "",
      "post_date": "2023-01-16T19:37:33.720000",
      "content": "<p>Hi! Thanks for your work!<br>\nI tried to reproduce and set batch_size to 16 due to my single GPU… and I find that the eval loss is increasing instead of decreasing…<br>\nShould I decreas the lr for my smaller batch_size?</p>\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2102835,
          "author_name": "slime",
          "author_url": "",
          "post_date": "2023-01-16T21:29:17.630000",
          "content": "<p>I tried this as well and failed, the model didn't converge, to make effnetv2 train I had to enable gradient checkpointing and keep batch size at 64, and the easier thing that I didn't test is to lower the learning rate, maybe it'll work comparatively to keeping batch size large</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2102910,
              "author_name": "nicehzj",
              "author_url": "",
              "post_date": "2023-01-16T22:06:47.737000",
              "content": "<p>I tried many models and I found effnetv2 is really hard to train… The eval loss often didn't converge…<br>\nI will try to decrease the lr from 3e-4 to 1e-4, then maybe 5e-5 to have a look.<br>\nBut as you said, i don't know if this only works for large batch size</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2103174,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-17T04:00:31.857000",
          "content": "<p>could you try large batch size with smaller image size first? Or maybe you can upsample the positive samples, I guess a small batch size may make the task hard since there may have many batches that have no pos samples.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2103803,
              "author_name": "nicehzj",
              "author_url": "",
              "post_date": "2023-01-17T11:58:52.540000",
              "content": "<p>Thanks! now I'm trying to use upsampling and also gradient accumulation to have a try!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2104315,
              "author_name": "nicehzj",
              "author_url": "",
              "post_date": "2023-01-17T18:40:59.550000",
              "content": "<p>I tried with upsampling the pos to 1/8 of each batch and gradient accumulation to get 64 a batch.<br>\nBut I still get the overfitting problems and cannot even close to your score…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2104979,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2023-01-18T07:22:45.210000",
              "content": "<p>what is your val score? Have you tried to train on 512 * 512 size? It should get a ~0.1 val score.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2099092,
      "author_name": "3rd Company 3rd Platoon Leader",
      "author_url": "",
      "post_date": "2023-01-14T07:01:16.593000",
      "content": "<p>Great Job!<br>\nSimple &amp; Impressive for your sharing.</p>\n<p>Could you get my 3 simple questions?</p>\n<ol>\n<li><p>As following your training config files, the pos_weight in loss function set to 1.<br>\nDid your 0.47 LB score come from it? or other value?<br>\nI'm just wondering because the pos_weight in most of other codes set to range of 10~50.</p></li>\n<li><p>In validation for each folds, how did you choose the best epoch weight?<br>\n2 scores(bin pfbeta1(thres 0.1) &amp; auc) can be printed.<br>\nwhich scores did you follow?</p></li>\n<li><p>In your experiment, Was just 10 epoch training enough to get the result?</p></li>\n</ol>\n<p>Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2101788,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-16T06:56:48.293000",
          "content": "<ol>\n<li>Yes, I tried other pos weights but the final result is similar. 1 works on my solution.</li>\n<li>I select the best pf1 (non binariezed) weights.</li>\n<li>I think 10 epoch is enough, I also tried 20, but the total val score is similar.</li>\n</ol>",
          "votes": 1,
          "replies": [
            {
              "id": 2105064,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2023-01-18T08:42:03.500000",
              "content": "<p>As for 3, I detected a new finding.<br>\nAccording to my latest experiments on cropped images, it's easier for training set to overfit, and reduce epoch number from 10 to 5 is enough to get similar val score and also lower train score (less overfit).</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2097286,
      "author_name": "David Serrano",
      "author_url": "",
      "post_date": "2023-01-12T14:57:56.947000",
      "content": "<p>Thank you for sharing your work! As you already said, you use StratifiedGroupKFold with k=4 and binarize f1 to get the final prediction. How do you find the threshold to use in the submissions? As you compute the mean along all the folds, do you compute the mean of all the thresholds?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2104975,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-18T07:20:29.123000",
          "content": "<p>I search the threshold based on all predictions (from all folds models).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2095373,
      "author_name": "Phaedrus",
      "author_url": "",
      "post_date": "2023-01-11T10:08:45.247000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/yiheng\" target=\"_blank\">@yiheng</a> . Are you also tracking AUC? If so, could you report it here? Thanks. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2095389,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-11T10:18:57.003000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/pheadrus\" target=\"_blank\">@pheadrus</a> , AUC is printed in my training pipeline however I did not record it in my sheet : (<br>\nI will add this part in my following experiments, and post my updates after getting a better result.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2095341,
      "author_name": "Simon Alerdic",
      "author_url": "",
      "post_date": "2023-01-11T09:44:33.350000",
      "content": "<p><a href=\"https://www.kaggle.com/yiheng\" target=\"_blank\">@yiheng</a> <br>\nDid you make any resampling (under or over)? And how much your fbeta fluctuated during training time on validation set? Did you get 0.10475 in a consistent progress? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2095390,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-11T10:20:21.073000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/simonalerdic\" target=\"_blank\">@simonalerdic</a> , in 512 * 512 experiments, val score increases consistently.<br>\nSo far I did not try resampling but just change the pos weight in bce loss.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2096561,
              "author_name": "Simon Alerdic",
              "author_url": "",
              "post_date": "2023-01-12T06:17:58.130000",
              "content": "<p>Is it possible to reproduce your training notebook score on kaggle env? Have you tried? In here, model size and image size, batch size gets limited due to resource and adjusting them from yours gave poor result. Could you please confirm?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2105047,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2023-01-18T08:27:34.160000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/simonalerdic\" target=\"_blank\">@simonalerdic</a> , I'm afraid I cannot do this kind of works recently since we have limited time for this challenge. I prefer finishing my current experiments first:</p>\n<ol>\n<li>try training with cropped images</li>\n<li>try different networks</li>\n<li>try different input size</li>\n</ol>\n<p>If having 512*512 results, I will post my log here if it is helpful.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2095096,
      "author_name": "Juntao Jiang",
      "author_url": "",
      "post_date": "2023-01-11T07:43:59.407000",
      "content": "<p>Is the batch size an important parameter? I reduced batch size to avoid out of memory, but the validation score will be very low.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2095376,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-11T10:10:02.600000",
          "content": "<p>I think it has some impacts, since the percentage of pos samples is low. Maybe you can try to accumulate gradients if not having enough gpu memory.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2103574,
              "author_name": "Juntao Jiang",
              "author_url": "",
              "post_date": "2023-01-17T08:43:27.883000",
              "content": "<p>Thank you!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2095496,
          "author_name": "ForcewithMe",
          "author_url": "",
          "post_date": "2023-01-11T12:02:35.327000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/juntaojiang\" target=\"_blank\">@juntaojiang</a> , batch size is really trick, while the relationship between BS and LR is also very important.</p>\n<ol>\n<li>When you reduce BS, try to scale LR at the same time. For example, the original BS is 32 and the LR is 1e-4, and you want to reduce the BS to 16, try to reduce LR to 5e-5. If you have more time, you can search the optimal LR around 5e-5.</li>\n<li>If you use multiple GPU to train, never put only one image per gpus. For example, if you use 8 cards, your minimum batch size should be 16, while there are 2 images per gpus. If you set BS=8, then there is 1 image per gpu, which will cause the model to perform much worse. It's caused by  the BatchNorm in most CNN models.</li>\n</ol>",
          "votes": 10,
          "replies": [
            {
              "id": 2103572,
              "author_name": "Juntao Jiang",
              "author_url": "",
              "post_date": "2023-01-17T08:43:15.090000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/forcewithme\" target=\"_blank\">@forcewithme</a>  Thank you for your kind suggestion! I will have a try.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2095041,
      "author_name": "Vaidehi Savaliya",
      "author_url": "",
      "post_date": "2023-01-11T07:04:35.680000",
      "content": "<p>good information…</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2095019,
      "author_name": "阳光开朗大男孩",
      "author_url": "",
      "post_date": "2023-01-11T06:44:23.623000",
      "content": "<p>Nice work! btw, how do u determine <code>clf_threshold = 0.1</code>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2095347,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-11T09:50:54.737000",
          "content": "<p>In the training step, I just manually set it to 0.1 to take a look (avoid doing search on every epoch). <br>\nThe final threshold I used is searched on all 4 folds predictions.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2095377,
              "author_name": "阳光开朗大男孩",
              "author_url": "",
              "post_date": "2023-01-11T10:10:28.667000",
              "content": "<p>Got it, thx!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2094563,
      "author_name": "Gustavo Ruiz",
      "author_url": "",
      "post_date": "2023-01-10T20:44:22.340000",
      "content": "<p>Excellent post, very helpful and concise. Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2092574,
      "author_name": "David Landup",
      "author_url": "",
      "post_date": "2023-01-09T12:36:00.047000",
      "content": "<p>Great findings, thanks for sharing!<br>\nHow much VRAM do you have for 1024x1024 input images with EffNetV2S? 👀</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2092592,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-09T12:44:38.877000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/davidlandup\" target=\"_blank\">@davidlandup</a> , for training, I used batch size = 64 on 4 * 32G gpus.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2092606,
              "author_name": "David Landup",
              "author_url": "",
              "post_date": "2023-01-09T12:50:00.063000",
              "content": "<p>Thanks for the quick response! Packing some serious compute 😨</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2092648,
              "author_name": "Yiheng Wang",
              "author_url": "",
              "post_date": "2023-01-09T13:24:27.417000",
              "content": "<p>today I tried to train a 2048 * 2048 model but failed. My GPUs are also not enough, but I think it's interesting to find better ways to optimize our solution with limited resources. I think crop foreground is a good way to avoid using large size, maybe you can try it, and I will also have a try and share my findings later.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2093767,
              "author_name": "David Landup",
              "author_url": "",
              "post_date": "2023-01-10T09:58:14.953000",
              "content": "<p>I wanted to create a baseline before cropping. Locally, I had an unreasonably high validation AUC (97%) which is a clear indication of a messed up data pipeline.<br>\nI'm currently getting ~0.06 pF1 (75% val AUC)  locally on small images (224x224).</p>\n<p>Not sure if the same pipeline on larger images will translate to a much higher pF1, but I'd like to get it up to ~0.1 before kicking off a 1024x1024 or 2048x2048 run, at which point ROI extraction will likely play a much larger role. I'll probably experiment with cropping in the next couple of days to see if that helps get a higher base pF1…</p>\n<p>Posting updates if there are any :)</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2094047,
              "author_name": "David Serrano",
              "author_url": "",
              "post_date": "2023-01-10T15:07:42.807000",
              "content": "<p>In my experiments,  I found that increasing the image size improved by quite a lot the results. But as you said, I would suggest having a strong pipeline and playing with it. ROI extraction also played for me a huge improvement.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2092484,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2023-01-09T11:17:29.120000",
      "content": "<p>What's the inference time on the 4 fold 1024 * 1024 setup?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2092591,
          "author_name": "Yiheng Wang",
          "author_url": "",
          "post_date": "2023-01-09T12:43:53.213000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/kiranik\" target=\"_blank\">@kiranik</a> , sorry I did not record the time as I submitted before going to bed : ) Is there any easy way to check the inference time?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2092624,
              "author_name": "Kirderf",
              "author_url": "",
              "post_date": "2023-01-09T13:03:21.177000",
              "content": "<p>Alright, yes it's not easy, only way is to manually check if the personally schedule allow at the time and get some ~time intervall for a new setup,  some also use the train set for time and memory check. </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2116345,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-26T12:38:28.017000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2115421,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-25T18:50:37.333000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2110815,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-22T12:52:09.127000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2095271,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-11T09:04:37.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2093668,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-10T07:53:42.943000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2093522,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-01-10T04:10:12.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2097241,
      "author_name": "LIUZIWEI0827",
      "author_url": "",
      "post_date": "2023-01-12T14:32:09.333000",
      "content": "<p>Thanks for sharing your insights</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2095106,
      "author_name": "yang zhao",
      "author_url": "",
      "post_date": "2023-01-11T07:51:37.493000",
      "content": "<p>thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2092068": "In this topic, I will share my experiences in this challenge, and open source all of my works on how to get a 0.47 (the number may change) model.\n\nInference code link:\nhttps://www.kaggle.com/code/yiheng/monai-baseline\nTraining code link:\nhttps://www.kaggle.com/code/yiheng/monai-pipeline-training/notebook\n\nyou can also clone the training pipeline from github:\n```\ngit clone git@github.com:yiheng-wang-nv/rsna_breast_monai_solution.git\n```\n\nThe following is my local and lb results (will update if having any progresses):\n\n| | pf1 local | bin pf1 local | LB |\n| --- | --- | --- | --- |\n| fold 0, agg by max | 0.2753 | 0.3352 | 0.38 |\n| fold 0, agg by mean | 0.2042 | 0.3392 | 0.42 |\n| fold 0, agg by mean, 2 tta | 0.2045 | 0.3879 | 0.44 |\n| 4 folds, agg by mean, 2 tta | 0.2234 | 0.3129 | 0.47 |",
    "2092074": "[Update 1] My purpose is to make use of MONAI, as well as other libs (such as timm for network, dali for acceleration) to build a decent model.\nFirst of all, I wish to thank @theoviel , I followed [your topic](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/369769) and [your code](https://www.kaggle.com/code/theoviel/rsna-breast-baseline-inference) to design my training pipeline, which can be summarized as the following settings:\n\n- group stratified 4 fold data split\n- 512 * 512 input size\n- h flip augmentation\n- 4 epoch, lr decay from 3e-4\n- timm, efnv2s network\n\nBased on these settings, I started to do experiments, and my initial purpose is to reach pf1 ~0.083 and auc 0.755 (these two values are from theoviel's open sourced notebook I mentioned above). What I tuned include:\n\n1. number of epochs\n2. pos sample weights on bce loss\n3. drop rate of efficientnet (this idea is achieved from @hengck23 in [his topic](https://www.kaggle.com/competitions/rsna-breast-cancer-detection/discussion/370333)\nafter ~ 8 experiments, I got pf1 0.10475 on one fold locally. Then I switched to use larger input size: 1024 * 1024 and do all 4 folds experiments. In this step, I did not tune any parameters, but just try to finish training, and then submit to LB.",
    "2100083": "Great Job!",
    "2092163": "Are you using Kaggle for training, some other cloud provider, or local GPU setup?",
    "2104971": "[Update 2] I used @hengck23 's trained model (in: https://www.kaggle.com/code/hengck23/proprocess-function-e-g-crop-breast-region) to infer the bounding boxes for all training images and then do the following preprocessing:\ndicom ---> crop by bounding box ---> resize to 1024 * 1024\nthe resized dataset is uploaded to: https://www.kaggle.com/datasets/yiheng/rsna-breast-1024\n\nNow I'm training with this dataset, if having any progresses, I will update the topic and the notebook.\nThe next plans are:\n1. train with different backbone\n2. train with different input size\n3. optimize the inference time",
    "2102690": "Hi! Thanks for your work!\nI tried to reproduce and set batch_size to 16 due to my single GPU... and I find that the eval loss is increasing instead of decreasing...\nShould I decreas the lr for my smaller batch_size?\n\nThanks!",
    "2099092": "Great Job!\nSimple & Impressive for your sharing.\n\nCould you get my 3 simple questions?\n1. As following your training config files, the pos_weight in loss function set to 1.\nDid your 0.47 LB score come from it? or other value?\nI'm just wondering because the pos_weight in most of other codes set to range of 10~50.\n\n2. In validation for each folds, how did you choose the best epoch weight?\n2 scores(bin pfbeta1(thres 0.1) & auc) can be printed.\nwhich scores did you follow?\n\n3. In your experiment, Was just 10 epoch training enough to get the result?\n\nThank you!",
    "2097286": "Thank you for sharing your work! As you already said, you use StratifiedGroupKFold with k=4 and binarize f1 to get the final prediction. How do you find the threshold to use in the submissions? As you compute the mean along all the folds, do you compute the mean of all the thresholds?",
    "2095373": "Hi @yiheng . Are you also tracking AUC? If so, could you report it here? Thanks. ",
    "2095341": "@yiheng \nDid you make any resampling (under or over)? And how much your fbeta fluctuated during training time on validation set? Did you get 0.10475 in a consistent progress? ",
    "2095096": "Is the batch size an important parameter? I reduced batch size to avoid out of memory, but the validation score will be very low.",
    "2095041": "good information...",
    "2095019": "Nice work! btw, how do u determine `clf_threshold = 0.1`?",
    "2094563": "Excellent post, very helpful and concise. Thanks for sharing!",
    "2092574": "Great findings, thanks for sharing!\nHow much VRAM do you have for 1024x1024 input images with EffNetV2S? 👀",
    "2092484": "What's the inference time on the 4 fold 1024 * 1024 setup?",
    "2116345": "",
    "2115421": "",
    "2110815": "",
    "2095271": "",
    "2093668": "",
    "2093522": "",
    "2097241": "Thanks for sharing your insights\n",
    "2095106": "thank you!"
  }
}