{
  "id": 410111,
  "title": "Basic Question regarding Binary Image Classification",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/410111",
  "author_name": "Jan H",
  "post_date": "2023-05-14T01:01:18.972000",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey guys, <br>\ni am pretty new to the field of classification using neural networks. For this competition i though about using a cnn classifier to classify each pixel either as trail or as \"background\" (non-trail). When i did some research, i came accross the approach to use the sigmoid activation function for the output, such that each pixel in the cnn output gets a value between 0 and 1 (which then translates to the class), resulting in an output shape of [Image_hight, Image_Width, 1] where the last channel translates to the class label (which is between 0 and 1 after the sigmoid activation). The \"true mask\" values are either 0 or 1, depending on whether the pixel is a trail or not.</p>\n<p><strong>My question now is</strong>: How does the neural network use this output to assess the training results? How are the output values, which are between 0 and 1 (after my Sigmoid) translated / converted to the class labels, which are then used to assess the metrics and loss? This might be a pretty simple question, but after some research, i couldnt figure it out exactly. </p>\n<p>Thanks for your help guys!</p>",
  "messages": [
    {
      "id": 2258475,
      "postDate": "2023-05-14T08:32:02.917Z",
      "content": "<p>Let's say your model's output is <code>0.67</code> (for a specific pixel) and the true (ground truth or GT) pixel value is <code>1</code> then your loss should be lower than the case when your model outputs <code>0.4</code> for the same pixel.  </p>\n<p>To make things a bit more technical, consider the Binary Cross Entropy loss applied on all pixels and the value summed up for all samples in your dataset:</p>\n<p>$$BCE(x=output, y=GT) = \\sum\\limits_{0}^{n} \\sum\\limits_{i=0}^{N} \\sum\\limits_{j=0}^{M} y_{ij} log(x_{ij}) + (1-y_{ij})log(1-x_{ij})$$</p>\n<p>where,<br>\n<em>n -&gt; Total number of examples</em><br>\n<em>N, M -&gt; Mask width and height</em></p>\n<p><strong>Note 1:</strong><br>\nOther losses and metrics might threshold your model's output and then treat it like a discrete problem. <br>\nOnly counting of True Positives, True Negatives, False Positives  and False Negatives remains hereon. That's how losses or metrics like F1 Score, IoU are calculated most times.</p>\n<p><strong>Note 2:</strong><br>\nThe threshold is usually kept at <code>0.5</code> in practice.</p>\n<p>Hope this helps a bit :)</p>",
      "rawMarkdown": "Let's say your model's output is `0.67` (for a specific pixel) and the true (ground truth or GT) pixel value is `1` then your loss should be lower than the case when your model outputs `0.4` for the same pixel.  \n\nTo make things a bit more technical, consider the Binary Cross Entropy loss applied on all pixels and the value summed up for all samples in your dataset:\n\n$$BCE(x=output, y=GT) = \\sum\\limits_{0}^{n} \\sum\\limits_{i=0}^{N} \\sum\\limits_{j=0}^{M} y_{ij} log(x_{ij}) + (1-y_{ij})log(1-x_{ij})$$\n\nwhere,\n*n -> Total number of examples*\n*N, M -> Mask width and height*\n\n**Note 1:**\nOther losses and metrics might threshold your model's output and then treat it like a discrete problem. \nOnly counting of True Positives, True Negatives, False Positives  and False Negatives remains hereon. That's how losses or metrics like F1 Score, IoU are calculated most times.\n\n**Note 2:**\nThe threshold is usually kept at `0.5` in practice.\n \n\nHope this helps a bit :)\n",
      "votes": 2,
      "replies": [
        {
          "id": 2258979,
          "postDate": "2023-05-14T15:37:10.920Z",
          "content": "<p>Thank you! That helped a lot! <br>\nFrom this i just have another subsequent question: for the model evaluation / inference, we would then have to transfer every output values into discrete values between 0 and 1 manually (e.g using a threshold of 0.5) right?</p>",
          "rawMarkdown": "Thank you! That helped a lot! \nFrom this i just have another subsequent question: for the model evaluation / inference, we would then have to transfer every output values into discrete values between 0 and 1 manually (e.g using a threshold of 0.5) right?",
          "votes": 1,
          "replies": [
            {
              "id": 2259063,
              "postDate": "2023-05-14T17:00:09.417Z",
              "content": "<p>Exactly. The threshold is a hyperparameter which you can optimize later, but usually you just use 0.5</p>",
              "rawMarkdown": "Exactly. The threshold is a hyperparameter which you can optimize later, but usually you just use 0.5",
              "votes": 1
            }
          ]
        },
        {
          "id": 2259190,
          "postDate": "2023-05-14T18:31:36.707Z",
          "content": "<p>Great explanation! Just to add to your answer, </p>\n<p>Sigmoid functions are often used as activation functions in the output layer for binary classification problems. The output of the sigmoid function can be interpreted as the estimated probability that a given input belongs to the positive class.</p>\n<p>The sigmoid function, also known as the logistic function, has the mathematical form:<br>\nσ(x) = 1 / (1 + e^(-x))<br>\nIt takes real number as input and produces an output between 0 and 1. When the input is positive, the sigmoid function approaches 1, and when the input is negative, it approaches 0. The value of the sigmoid function at 0 is exactly 0.5.</p>",
          "rawMarkdown": "Great explanation! Just to add to your answer, \n\nSigmoid functions are often used as activation functions in the output layer for binary classification problems. The output of the sigmoid function can be interpreted as the estimated probability that a given input belongs to the positive class.\n\nThe sigmoid function, also known as the logistic function, has the mathematical form:\nσ(x) = 1 / (1 + e^(-x))\nIt takes real number as input and produces an output between 0 and 1. When the input is positive, the sigmoid function approaches 1, and when the input is negative, it approaches 0. The value of the sigmoid function at 0 is exactly 0.5.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2258178,
      "postDate": "2023-05-14T01:01:18.973Z",
      "content": "<p>Hey guys, <br>\ni am pretty new to the field of classification using neural networks. For this competition i though about using a cnn classifier to classify each pixel either as trail or as \"background\" (non-trail). When i did some research, i came accross the approach to use the sigmoid activation function for the output, such that each pixel in the cnn output gets a value between 0 and 1 (which then translates to the class), resulting in an output shape of [Image_hight, Image_Width, 1] where the last channel translates to the class label (which is between 0 and 1 after the sigmoid activation). The \"true mask\" values are either 0 or 1, depending on whether the pixel is a trail or not.</p>\n<p><strong>My question now is</strong>: How does the neural network use this output to assess the training results? How are the output values, which are between 0 and 1 (after my Sigmoid) translated / converted to the class labels, which are then used to assess the metrics and loss? This might be a pretty simple question, but after some research, i couldnt figure it out exactly. </p>\n<p>Thanks for your help guys!</p>",
      "rawMarkdown": "Hey guys, \ni am pretty new to the field of classification using neural networks. For this competition i though about using a cnn classifier to classify each pixel either as trail or as \"background\" (non-trail). When i did some research, i came accross the approach to use the sigmoid activation function for the output, such that each pixel in the cnn output gets a value between 0 and 1 (which then translates to the class), resulting in an output shape of [Image_hight, Image_Width, 1] where the last channel translates to the class label (which is between 0 and 1 after the sigmoid activation). The \"true mask\" values are either 0 or 1, depending on whether the pixel is a trail or not.\n\n**My question now is**: How does the neural network use this output to assess the training results? How are the output values, which are between 0 and 1 (after my Sigmoid) translated / converted to the class labels, which are then used to assess the metrics and loss? This might be a pretty simple question, but after some research, i couldnt figure it out exactly. \n\nThanks for your help guys!",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2258475,
      "author_name": "Aryan Garg",
      "author_url": "",
      "post_date": "2023-05-14T08:32:02.917000",
      "content": "<p>Let's say your model's output is <code>0.67</code> (for a specific pixel) and the true (ground truth or GT) pixel value is <code>1</code> then your loss should be lower than the case when your model outputs <code>0.4</code> for the same pixel.  </p>\n<p>To make things a bit more technical, consider the Binary Cross Entropy loss applied on all pixels and the value summed up for all samples in your dataset:</p>\n<p>$$BCE(x=output, y=GT) = \\sum\\limits_{0}^{n} \\sum\\limits_{i=0}^{N} \\sum\\limits_{j=0}^{M} y_{ij} log(x_{ij}) + (1-y_{ij})log(1-x_{ij})$$</p>\n<p>where,<br>\n<em>n -&gt; Total number of examples</em><br>\n<em>N, M -&gt; Mask width and height</em></p>\n<p><strong>Note 1:</strong><br>\nOther losses and metrics might threshold your model's output and then treat it like a discrete problem. <br>\nOnly counting of True Positives, True Negatives, False Positives  and False Negatives remains hereon. That's how losses or metrics like F1 Score, IoU are calculated most times.</p>\n<p><strong>Note 2:</strong><br>\nThe threshold is usually kept at <code>0.5</code> in practice.</p>\n<p>Hope this helps a bit :)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2258979,
          "author_name": "Jan H",
          "author_url": "",
          "post_date": "2023-05-14T15:37:10.920000",
          "content": "<p>Thank you! That helped a lot! <br>\nFrom this i just have another subsequent question: for the model evaluation / inference, we would then have to transfer every output values into discrete values between 0 and 1 manually (e.g using a threshold of 0.5) right?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2259063,
              "author_name": "Patchef",
              "author_url": "",
              "post_date": "2023-05-14T17:00:09.417000",
              "content": "<p>Exactly. The threshold is a hyperparameter which you can optimize later, but usually you just use 0.5</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2259190,
          "author_name": "Soumyadeep Khandual",
          "author_url": "",
          "post_date": "2023-05-14T18:31:36.707000",
          "content": "<p>Great explanation! Just to add to your answer, </p>\n<p>Sigmoid functions are often used as activation functions in the output layer for binary classification problems. The output of the sigmoid function can be interpreted as the estimated probability that a given input belongs to the positive class.</p>\n<p>The sigmoid function, also known as the logistic function, has the mathematical form:<br>\nσ(x) = 1 / (1 + e^(-x))<br>\nIt takes real number as input and produces an output between 0 and 1. When the input is positive, the sigmoid function approaches 1, and when the input is negative, it approaches 0. The value of the sigmoid function at 0 is exactly 0.5.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2258475": "Let's say your model's output is `0.67` (for a specific pixel) and the true (ground truth or GT) pixel value is `1` then your loss should be lower than the case when your model outputs `0.4` for the same pixel.  \n\nTo make things a bit more technical, consider the Binary Cross Entropy loss applied on all pixels and the value summed up for all samples in your dataset:\n\n$$BCE(x=output, y=GT) = \\sum\\limits_{0}^{n} \\sum\\limits_{i=0}^{N} \\sum\\limits_{j=0}^{M} y_{ij} log(x_{ij}) + (1-y_{ij})log(1-x_{ij})$$\n\nwhere,\n*n -> Total number of examples*\n*N, M -> Mask width and height*\n\n**Note 1:**\nOther losses and metrics might threshold your model's output and then treat it like a discrete problem. \nOnly counting of True Positives, True Negatives, False Positives  and False Negatives remains hereon. That's how losses or metrics like F1 Score, IoU are calculated most times.\n\n**Note 2:**\nThe threshold is usually kept at `0.5` in practice.\n \n\nHope this helps a bit :)\n",
    "2258178": "Hey guys, \ni am pretty new to the field of classification using neural networks. For this competition i though about using a cnn classifier to classify each pixel either as trail or as \"background\" (non-trail). When i did some research, i came accross the approach to use the sigmoid activation function for the output, such that each pixel in the cnn output gets a value between 0 and 1 (which then translates to the class), resulting in an output shape of [Image_hight, Image_Width, 1] where the last channel translates to the class label (which is between 0 and 1 after the sigmoid activation). The \"true mask\" values are either 0 or 1, depending on whether the pixel is a trail or not.\n\n**My question now is**: How does the neural network use this output to assess the training results? How are the output values, which are between 0 and 1 (after my Sigmoid) translated / converted to the class labels, which are then used to assess the metrics and loss? This might be a pretty simple question, but after some research, i couldnt figure it out exactly. \n\nThanks for your help guys!"
  }
}