{
  "id": 372592,
  "title": "Training with size 512 vs 1024 experience",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372592",
  "author_name": "Martin Kovacevic Buvinic",
  "post_date": "2022-12-16T19:13:08.016000",
  "votes": 17,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I wanted to share that I have been training models with image size 512 and 1024, what I found interesting is the following:</p>\n<p>For example if we have backbone A and backbone B and train each backbone with the same hyperparammeters for image size 512, backbone A has a better CV compared to B.</p>\n<p>What is interesting is that when I do the same for image size 1024, backbone B has a much better CV compared to A.</p>\n<p>The main idea was to train with smaller image size and then if I have good results I can use bigger image size, nevertheless this experiment suggest that results from using image size 512 will not always have the same behaviour for image size 1024.</p>\n<p>For experts in computer vision this could not be an amazing discovery but for the rest that are learning I found it was interesting to share this</p>",
  "messages": [
    {
      "id": 2067491,
      "postDate": "2022-12-16T19:13:08.017Z",
      "content": "<p>I wanted to share that I have been training models with image size 512 and 1024, what I found interesting is the following:</p>\n<p>For example if we have backbone A and backbone B and train each backbone with the same hyperparammeters for image size 512, backbone A has a better CV compared to B.</p>\n<p>What is interesting is that when I do the same for image size 1024, backbone B has a much better CV compared to A.</p>\n<p>The main idea was to train with smaller image size and then if I have good results I can use bigger image size, nevertheless this experiment suggest that results from using image size 512 will not always have the same behaviour for image size 1024.</p>\n<p>For experts in computer vision this could not be an amazing discovery but for the rest that are learning I found it was interesting to share this</p>",
      "rawMarkdown": "I wanted to share that I have been training models with image size 512 and 1024, what I found interesting is the following:\n\nFor example if we have backbone A and backbone B and train each backbone with the same hyperparammeters for image size 512, backbone A has a better CV compared to B.\n\nWhat is interesting is that when I do the same for image size 1024, backbone B has a much better CV compared to A.\n\nThe main idea was to train with smaller image size and then if I have good results I can use bigger image size, nevertheless this experiment suggest that results from using image size 512 will not always have the same behaviour for image size 1024.\n\nFor experts in computer vision this could not be an amazing discovery but for the rest that are learning I found it was interesting to share this",
      "votes": 17
    },
    {
      "id": 2067508,
      "postDate": "2022-12-16T19:30:34.720Z",
      "content": "<p>this is a common observation in this competition data. Even for one single backbone, different oversampling strategy gives non-similar results for different size. </p>\n<p>I think the reasons could be:</p>\n<ol>\n<li>too little pos data to make the cv metric (and lb) reliable in the first place </li>\n<li>you may think we are having a 2 class binary problem, but it is not true. it is one class vs background (not a cat  versus dog problem but cat versus random images problem)<br>\nwe only have 54705*0.02=1094 pos images.</li>\n</ol>\n<p>oversampling, pos weight can easily leads to over fitting to a particular one pos sample.</p>\n<hr>\n<p>if you want to visualise learned feature space you can write code to do playground like this <a href=\"https://playground.tensorflow.org\" target=\"_blank\">https://playground.tensorflow.org</a>.<br>\nThen change the loss function to focal loss, (heavily) weighted BCE, etc for different data of pos to background ratio. you can see That the  learned feature space can changed drastically for different samples</p>\n<hr>\n<p>on a side note, i think some kagglers may observed that some models are stuck are learning to predict only background class. This is when you have only cv in the range of 0.05</p>",
      "rawMarkdown": "this is a common observation in this competition data. Even for one single backbone, different oversampling strategy gives non-similar results for different size. \n\nI think the reasons could be:\n1. too little pos data to make the cv metric (and lb) reliable in the first place \n2. you may think we are having a 2 class binary problem, but it is not true. it is one class vs background (not a cat  versus dog problem but cat versus random images problem)\nwe only have 54705*0.02=1094 pos images.\n\noversampling, pos weight can easily leads to over fitting to a particular one pos sample.\n\n---\n\nif you want to visualise learned feature space you can write code to do playground like this https://playground.tensorflow.org.\nThen change the loss function to focal loss, (heavily) weighted BCE, etc for different data of pos to background ratio. you can see That the  learned feature space can changed drastically for different samples\n\n--- \non a side note, i think some kagglers may observed that some models are stuck are learning to predict only background class. This is when you have only cv in the range of 0.05",
      "votes": 7,
      "replies": [
        {
          "id": 2068223,
          "postDate": "2022-12-17T16:19:17.247Z",
          "content": "<blockquote>\n  <p>i think some kagglers may observed that some models are stuck are learning to predict only background class. This is when you have only cv in the range of 0.05</p>\n</blockquote>\n<p>😂 This is exactly me. I have accounted for the imbalance in my loss function, but am still stuck in the 0.05 range. Do you have any suggestions?</p>\n<p>Thanks!</p>",
          "rawMarkdown": ">i think some kagglers may observed that some models are stuck are learning to predict only background class. This is when you have only cv in the range of 0.05\n\n😂 This is exactly me. I have accounted for the imbalance in my loss function, but am still stuck in the 0.05 range. Do you have any suggestions?\n\nThanks!",
          "replies": [
            {
              "id": 2068377,
              "postDate": "2022-12-17T20:06:12.637Z",
              "content": "<p>efficentnetb2/3/4 is one that not sensitive to the kaggle dataset here.<br>\nuse batch with pos/to neg in the range of 1/32 to 1/16. and rate from 1e-3, deccrease to 1-5</p>\n<p>for larger model (e.g. &gt;0.85 top-1 in imagenet), you still can get it to train by lowering the learning rate (initial rate = 1e-4) or use  pos/to neg in range of 1/2 to 1/8. But even if these can achieve good CV (say in the range of 0.40), they may perform poorly in LB(only in range of  0.35). A better solution for large model is the use of segmentation as aux loss (from external data or pesudo mask from small model)</p>",
              "rawMarkdown": "efficentnetb2/3/4 is one that not sensitive to the kaggle dataset here.\nuse batch with pos/to neg in the range of 1/32 to 1/16. and rate from 1e-3, deccrease to 1-5\n\nfor larger model (e.g. >0.85 top-1 in imagenet), you still can get it to train by lowering the learning rate (initial rate = 1e-4) or use  pos/to neg in range of 1/2 to 1/8. But even if these can achieve good CV (say in the range of 0.40), they may perform poorly in LB(only in range of  0.35). A better solution for large model is the use of segmentation as aux loss (from external data or pesudo mask from small model)\n",
              "votes": 7
            },
            {
              "id": 2068378,
              "postDate": "2022-12-17T20:09:53.953Z",
              "content": "<p>That's very helpful, thank you for sharing your thoughts! 😄</p>",
              "rawMarkdown": "That's very helpful, thank you for sharing your thoughts! 😄"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2067508,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-12-16T19:30:34.720000",
      "content": "<p>this is a common observation in this competition data. Even for one single backbone, different oversampling strategy gives non-similar results for different size. </p>\n<p>I think the reasons could be:</p>\n<ol>\n<li>too little pos data to make the cv metric (and lb) reliable in the first place </li>\n<li>you may think we are having a 2 class binary problem, but it is not true. it is one class vs background (not a cat  versus dog problem but cat versus random images problem)<br>\nwe only have 54705*0.02=1094 pos images.</li>\n</ol>\n<p>oversampling, pos weight can easily leads to over fitting to a particular one pos sample.</p>\n<hr>\n<p>if you want to visualise learned feature space you can write code to do playground like this <a href=\"https://playground.tensorflow.org\" target=\"_blank\">https://playground.tensorflow.org</a>.<br>\nThen change the loss function to focal loss, (heavily) weighted BCE, etc for different data of pos to background ratio. you can see That the  learned feature space can changed drastically for different samples</p>\n<hr>\n<p>on a side note, i think some kagglers may observed that some models are stuck are learning to predict only background class. This is when you have only cv in the range of 0.05</p>",
      "votes": 7,
      "replies": [
        {
          "id": 2068223,
          "author_name": "Julia Nickerson",
          "author_url": "",
          "post_date": "2022-12-17T16:19:17.247000",
          "content": "<blockquote>\n  <p>i think some kagglers may observed that some models are stuck are learning to predict only background class. This is when you have only cv in the range of 0.05</p>\n</blockquote>\n<p>😂 This is exactly me. I have accounted for the imbalance in my loss function, but am still stuck in the 0.05 range. Do you have any suggestions?</p>\n<p>Thanks!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2068377,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2022-12-17T20:06:12.637000",
              "content": "<p>efficentnetb2/3/4 is one that not sensitive to the kaggle dataset here.<br>\nuse batch with pos/to neg in the range of 1/32 to 1/16. and rate from 1e-3, deccrease to 1-5</p>\n<p>for larger model (e.g. &gt;0.85 top-1 in imagenet), you still can get it to train by lowering the learning rate (initial rate = 1e-4) or use  pos/to neg in range of 1/2 to 1/8. But even if these can achieve good CV (say in the range of 0.40), they may perform poorly in LB(only in range of  0.35). A better solution for large model is the use of segmentation as aux loss (from external data or pesudo mask from small model)</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2068378,
              "author_name": "Julia Nickerson",
              "author_url": "",
              "post_date": "2022-12-17T20:09:53.953000",
              "content": "<p>That's very helpful, thank you for sharing your thoughts! 😄</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2067491": "I wanted to share that I have been training models with image size 512 and 1024, what I found interesting is the following:\n\nFor example if we have backbone A and backbone B and train each backbone with the same hyperparammeters for image size 512, backbone A has a better CV compared to B.\n\nWhat is interesting is that when I do the same for image size 1024, backbone B has a much better CV compared to A.\n\nThe main idea was to train with smaller image size and then if I have good results I can use bigger image size, nevertheless this experiment suggest that results from using image size 512 will not always have the same behaviour for image size 1024.\n\nFor experts in computer vision this could not be an amazing discovery but for the rest that are learning I found it was interesting to share this",
    "2067508": "this is a common observation in this competition data. Even for one single backbone, different oversampling strategy gives non-similar results for different size. \n\nI think the reasons could be:\n1. too little pos data to make the cv metric (and lb) reliable in the first place \n2. you may think we are having a 2 class binary problem, but it is not true. it is one class vs background (not a cat  versus dog problem but cat versus random images problem)\nwe only have 54705*0.02=1094 pos images.\n\noversampling, pos weight can easily leads to over fitting to a particular one pos sample.\n\n---\n\nif you want to visualise learned feature space you can write code to do playground like this https://playground.tensorflow.org.\nThen change the loss function to focal loss, (heavily) weighted BCE, etc for different data of pos to background ratio. you can see That the  learned feature space can changed drastically for different samples\n\n--- \non a side note, i think some kagglers may observed that some models are stuck are learning to predict only background class. This is when you have only cv in the range of 0.05"
  }
}