{
  "id": 413767,
  "title": "U-Net is missing something",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/413767",
  "author_name": "JEANMPIA",
  "post_date": "2023-05-30T03:38:06.822000",
  "votes": 40,
  "comment_count": 7,
  "views": 0,
  "content": "<h2>Spoiler: It doesn't use BatchNorm</h2>\n<p>It's very likely that if you are using U-Net for this competition, you implemented it yourself. After reading the paper (or any documentation that refer to U-Net), you realised that there was this <strong>Conv -&gt; Relu -&gt; Conv -&gt; Relu</strong> that was being repeated at every downsampling and upsampling stage, so naturally, you made it into a function (or most likely a Class):</p>\n<pre><code> ():\n     nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, , padding=, bias=),\n        nn.ReLU(inplace=),\n        nn.Conv2d(out_channels, out_channels, , padding=, bias=),\n        nn.ReLU(inplace=)\n    )   \n</code></pre>\n<p>Even if this implementation is reasonable and produces good results in itself, kaggle is all about improving the score at all cost, and sometimes, at <strong>a very low cost</strong>. Indeed, by simply adding BatchNorm, I was able to increase my CV and LB at the same time while having a lower number of epochs:</p>\n<pre><code> ():\n     nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, , padding=, bias=), \n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(inplace=),\n        nn.Conv2d(out_channels, out_channels, , padding=, bias=), \n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(inplace=)\n    )   \n</code></pre>\n<p><em>If you are wondering why they didn't use BatchNorm, it's because the two papers came out at the same time:</em><br>\n<a href=\"https://arxiv.org/abs/1502.03167\" target=\"_blank\">BatchNorm paper</a> - <a href=\"https://arxiv.org/abs/1505.04597\" target=\"_blank\">U-Net paper</a></p>",
  "messages": [
    {
      "id": 2280300,
      "postDate": "2023-05-30T03:38:06.823Z",
      "content": "<h2>Spoiler: It doesn't use BatchNorm</h2>\n<p>It's very likely that if you are using U-Net for this competition, you implemented it yourself. After reading the paper (or any documentation that refer to U-Net), you realised that there was this <strong>Conv -&gt; Relu -&gt; Conv -&gt; Relu</strong> that was being repeated at every downsampling and upsampling stage, so naturally, you made it into a function (or most likely a Class):</p>\n<pre><code> ():\n     nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, , padding=, bias=),\n        nn.ReLU(inplace=),\n        nn.Conv2d(out_channels, out_channels, , padding=, bias=),\n        nn.ReLU(inplace=)\n    )   \n</code></pre>\n<p>Even if this implementation is reasonable and produces good results in itself, kaggle is all about improving the score at all cost, and sometimes, at <strong>a very low cost</strong>. Indeed, by simply adding BatchNorm, I was able to increase my CV and LB at the same time while having a lower number of epochs:</p>\n<pre><code> ():\n     nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, , padding=, bias=), \n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(inplace=),\n        nn.Conv2d(out_channels, out_channels, , padding=, bias=), \n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(inplace=)\n    )   \n</code></pre>\n<p><em>If you are wondering why they didn't use BatchNorm, it's because the two papers came out at the same time:</em><br>\n<a href=\"https://arxiv.org/abs/1502.03167\" target=\"_blank\">BatchNorm paper</a> - <a href=\"https://arxiv.org/abs/1505.04597\" target=\"_blank\">U-Net paper</a></p>",
      "rawMarkdown": "## Spoiler: It doesn't use BatchNorm\nIt's very likely that if you are using U-Net for this competition, you implemented it yourself. After reading the paper (or any documentation that refer to U-Net), you realised that there was this **Conv -> Relu -> Conv -> Relu** that was being repeated at every downsampling and upsampling stage, so naturally, you made it into a function (or most likely a Class):\n```python\ndef double_conv(in_channels, out_channels):\n    return nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, 3, padding=1, bias=True),\n        nn.ReLU(inplace=True),\n        nn.Conv2d(out_channels, out_channels, 3, padding=1, bias=True),\n        nn.ReLU(inplace=True)\n    )   \n```\n\nEven if this implementation is reasonable and produces good results in itself, kaggle is all about improving the score at all cost, and sometimes, at **a very low cost**. Indeed, by simply adding BatchNorm, I was able to increase my CV and LB at the same time while having a lower number of epochs:\n\n```python\ndef double_conv(in_channels, out_channels):\n    return nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, 3, padding=1, bias=False), \n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(inplace=True),\n        nn.Conv2d(out_channels, out_channels, 3, padding=1, bias=False), \n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(inplace=True)\n    )   \n```\n\n*If you are wondering why they didn't use BatchNorm, it's because the two papers came out at the same time:*\n[BatchNorm paper](https://arxiv.org/abs/1502.03167) - [U-Net paper](https://arxiv.org/abs/1505.04597)",
      "votes": 39
    },
    {
      "id": 2280303,
      "postDate": "2023-05-30T03:47:37.727Z",
      "content": "<p>When asked ChatGPT why U-Net doesn't use BatchNorm:</p>\n<pre><code>The original U-Net architecture, proposed by Ronneberger et al. in 2015, does not use Batch Normalization (BatchNorm). There could be several reasons for this:\n\nHistorical Context: BatchNorm was proposed in a paper by Ioffe and Szegedy in 2015, around the same time that the U-Net paper was published. It's possible that the U-Net authors were not aware of BatchNorm, or they didn't have enough time to incorporate it into their model and test it thoroughly.\n\nModel Simplicity: The U-Net was designed to be a relatively simple and straightforward architecture for biomedical image segmentation. The authors may have preferred to keep the architecture as simple as possible to make it easier for other researchers to understand and implement.\n\nDifferent Objectives: BatchNorm is primarily used to address the problem of internal covariate shift, where the distribution of each layer's inputs changes during training. This can make training deep networks difficult. However, the U-Net was designed for a different purpose: to perform precise localization in biomedical image segmentation tasks. The authors may have felt that addressing internal covariate shift was not a major concern for this task.\n\nComputational Resources: BatchNorm requires additional computations during both the forward and backward passes of training, which increases the computational resources required. If the authors were working with limited computational resources, they may have chosen to omit BatchNorm to reduce the computational requirements of the model.\n\nThat being said, variations of U-Net that do include BatchNorm have been proposed since the original paper, and these variations have shown improved performance in some cases. The decision to use BatchNorm or not often depends on the specific task and dataset, as well as the computational resources available.\n</code></pre>",
      "rawMarkdown": "When asked ChatGPT why U-Net doesn't use BatchNorm:\n\n```\nThe original U-Net architecture, proposed by Ronneberger et al. in 2015, does not use Batch Normalization (BatchNorm). There could be several reasons for this:\n\nHistorical Context: BatchNorm was proposed in a paper by Ioffe and Szegedy in 2015, around the same time that the U-Net paper was published. It's possible that the U-Net authors were not aware of BatchNorm, or they didn't have enough time to incorporate it into their model and test it thoroughly.\n\nModel Simplicity: The U-Net was designed to be a relatively simple and straightforward architecture for biomedical image segmentation. The authors may have preferred to keep the architecture as simple as possible to make it easier for other researchers to understand and implement.\n\nDifferent Objectives: BatchNorm is primarily used to address the problem of internal covariate shift, where the distribution of each layer's inputs changes during training. This can make training deep networks difficult. However, the U-Net was designed for a different purpose: to perform precise localization in biomedical image segmentation tasks. The authors may have felt that addressing internal covariate shift was not a major concern for this task.\n\nComputational Resources: BatchNorm requires additional computations during both the forward and backward passes of training, which increases the computational resources required. If the authors were working with limited computational resources, they may have chosen to omit BatchNorm to reduce the computational requirements of the model.\n\nThat being said, variations of U-Net that do include BatchNorm have been proposed since the original paper, and these variations have shown improved performance in some cases. The decision to use BatchNorm or not often depends on the specific task and dataset, as well as the computational resources available.\n```",
      "votes": 2
    },
    {
      "id": 2280525,
      "postDate": "2023-05-30T06:43:56.280Z",
      "content": "<p>Thank you for sharing your great comparison and explanation! </p>",
      "rawMarkdown": "Thank you for sharing your great comparison and explanation! ",
      "votes": 1
    },
    {
      "id": 2324201,
      "postDate": "2023-06-30T13:17:51.993Z",
      "content": "<p>Thanks for sharing. What's the difference in CV and LB scores? How much boost are you getting?</p>",
      "rawMarkdown": "Thanks for sharing. What's the difference in CV and LB scores? How much boost are you getting?",
      "replies": [
        {
          "id": 2324276,
          "postDate": "2023-06-30T14:02:17.893Z",
          "content": "<p>I don't remember since i use smp implementation now which has a parameter for enabling BatchNorm. I recall it did improve the CV but also sped up training by lowering the number of epochs needed.</p>",
          "rawMarkdown": "I don't remember since i use smp implementation now which has a parameter for enabling BatchNorm. I recall it did improve the CV but also sped up training by lowering the number of epochs needed.",
          "votes": 1,
          "replies": [
            {
              "id": 2339677,
              "postDate": "2023-07-11T02:35:21.043Z",
              "content": "<p><a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> You mean <code>smp.Unet(decoder_use_batchnorm=True)</code>?</p>",
              "rawMarkdown": "@janmpia You mean `smp.Unet(decoder_use_batchnorm=True)`?",
              "votes": 1
            },
            {
              "id": 2340339,
              "postDate": "2023-07-11T12:30:59.913Z",
              "content": "<p>hello <a href=\"https://www.kaggle.com/riow1983\" target=\"_blank\">@riow1983</a>,<br>\nyes thats what I mean.</p>",
              "rawMarkdown": "hello @riow1983,\nyes thats what I mean.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2282848,
      "postDate": "2023-05-31T21:20:01.990Z",
      "content": "<p>Very interesting finding! Thank you!</p>",
      "rawMarkdown": "Very interesting finding! Thank you!"
    }
  ],
  "comments": [
    {
      "id": 2280303,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-05-30T03:47:37.727000",
      "content": "<p>When asked ChatGPT why U-Net doesn't use BatchNorm:</p>\n<pre><code>The original U-Net architecture, proposed by Ronneberger et al. in 2015, does not use Batch Normalization (BatchNorm). There could be several reasons for this:\n\nHistorical Context: BatchNorm was proposed in a paper by Ioffe and Szegedy in 2015, around the same time that the U-Net paper was published. It's possible that the U-Net authors were not aware of BatchNorm, or they didn't have enough time to incorporate it into their model and test it thoroughly.\n\nModel Simplicity: The U-Net was designed to be a relatively simple and straightforward architecture for biomedical image segmentation. The authors may have preferred to keep the architecture as simple as possible to make it easier for other researchers to understand and implement.\n\nDifferent Objectives: BatchNorm is primarily used to address the problem of internal covariate shift, where the distribution of each layer's inputs changes during training. This can make training deep networks difficult. However, the U-Net was designed for a different purpose: to perform precise localization in biomedical image segmentation tasks. The authors may have felt that addressing internal covariate shift was not a major concern for this task.\n\nComputational Resources: BatchNorm requires additional computations during both the forward and backward passes of training, which increases the computational resources required. If the authors were working with limited computational resources, they may have chosen to omit BatchNorm to reduce the computational requirements of the model.\n\nThat being said, variations of U-Net that do include BatchNorm have been proposed since the original paper, and these variations have shown improved performance in some cases. The decision to use BatchNorm or not often depends on the specific task and dataset, as well as the computational resources available.\n</code></pre>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2280525,
      "author_name": "LUPIN11",
      "author_url": "",
      "post_date": "2023-05-30T06:43:56.280000",
      "content": "<p>Thank you for sharing your great comparison and explanation! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2324201,
      "author_name": "CroDoc",
      "author_url": "",
      "post_date": "2023-06-30T13:17:51.993000",
      "content": "<p>Thanks for sharing. What's the difference in CV and LB scores? How much boost are you getting?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2324276,
          "author_name": "JEANMPIA",
          "author_url": "",
          "post_date": "2023-06-30T14:02:17.893000",
          "content": "<p>I don't remember since i use smp implementation now which has a parameter for enabling BatchNorm. I recall it did improve the CV but also sped up training by lowering the number of epochs needed.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2339677,
              "author_name": "Ryosuke Horiuchi",
              "author_url": "",
              "post_date": "2023-07-11T02:35:21.043000",
              "content": "<p><a href=\"https://www.kaggle.com/janmpia\" target=\"_blank\">@janmpia</a> You mean <code>smp.Unet(decoder_use_batchnorm=True)</code>?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2340339,
              "author_name": "JEANMPIA",
              "author_url": "",
              "post_date": "2023-07-11T12:30:59.913000",
              "content": "<p>hello <a href=\"https://www.kaggle.com/riow1983\" target=\"_blank\">@riow1983</a>,<br>\nyes thats what I mean.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2282848,
      "author_name": "Marla E.",
      "author_url": "",
      "post_date": "2023-05-31T21:20:01.990000",
      "content": "<p>Very interesting finding! Thank you!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2280300": "## Spoiler: It doesn't use BatchNorm\nIt's very likely that if you are using U-Net for this competition, you implemented it yourself. After reading the paper (or any documentation that refer to U-Net), you realised that there was this **Conv -> Relu -> Conv -> Relu** that was being repeated at every downsampling and upsampling stage, so naturally, you made it into a function (or most likely a Class):\n```python\ndef double_conv(in_channels, out_channels):\n    return nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, 3, padding=1, bias=True),\n        nn.ReLU(inplace=True),\n        nn.Conv2d(out_channels, out_channels, 3, padding=1, bias=True),\n        nn.ReLU(inplace=True)\n    )   \n```\n\nEven if this implementation is reasonable and produces good results in itself, kaggle is all about improving the score at all cost, and sometimes, at **a very low cost**. Indeed, by simply adding BatchNorm, I was able to increase my CV and LB at the same time while having a lower number of epochs:\n\n```python\ndef double_conv(in_channels, out_channels):\n    return nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, 3, padding=1, bias=False), \n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(inplace=True),\n        nn.Conv2d(out_channels, out_channels, 3, padding=1, bias=False), \n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(inplace=True)\n    )   \n```\n\n*If you are wondering why they didn't use BatchNorm, it's because the two papers came out at the same time:*\n[BatchNorm paper](https://arxiv.org/abs/1502.03167) - [U-Net paper](https://arxiv.org/abs/1505.04597)",
    "2280303": "When asked ChatGPT why U-Net doesn't use BatchNorm:\n\n```\nThe original U-Net architecture, proposed by Ronneberger et al. in 2015, does not use Batch Normalization (BatchNorm). There could be several reasons for this:\n\nHistorical Context: BatchNorm was proposed in a paper by Ioffe and Szegedy in 2015, around the same time that the U-Net paper was published. It's possible that the U-Net authors were not aware of BatchNorm, or they didn't have enough time to incorporate it into their model and test it thoroughly.\n\nModel Simplicity: The U-Net was designed to be a relatively simple and straightforward architecture for biomedical image segmentation. The authors may have preferred to keep the architecture as simple as possible to make it easier for other researchers to understand and implement.\n\nDifferent Objectives: BatchNorm is primarily used to address the problem of internal covariate shift, where the distribution of each layer's inputs changes during training. This can make training deep networks difficult. However, the U-Net was designed for a different purpose: to perform precise localization in biomedical image segmentation tasks. The authors may have felt that addressing internal covariate shift was not a major concern for this task.\n\nComputational Resources: BatchNorm requires additional computations during both the forward and backward passes of training, which increases the computational resources required. If the authors were working with limited computational resources, they may have chosen to omit BatchNorm to reduce the computational requirements of the model.\n\nThat being said, variations of U-Net that do include BatchNorm have been proposed since the original paper, and these variations have shown improved performance in some cases. The decision to use BatchNorm or not often depends on the specific task and dataset, as well as the computational resources available.\n```",
    "2280525": "Thank you for sharing your great comparison and explanation! ",
    "2324201": "Thanks for sharing. What's the difference in CV and LB scores? How much boost are you getting?",
    "2282848": "Very interesting finding! Thank you!"
  }
}