{
  "id": 341941,
  "title": "What should be metric / loss function?",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/341941",
  "author_name": "Anurag Dhadse",
  "post_date": "2022-08-05T01:09:40.787000",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The competition describes the evaluation to be done using <em>weighted multi-class log loss</em> (weighted categorical crossentropy). Is this supposed to be <code>metric</code> while training the model or <code>loss</code>?</p>\n<p>Many Notebooks in the code sections are using <code>crossentropy</code> as the loss function, and <code>accuracy</code> as the evaluation metric while the output layer is a Dense layer with 2 nodes instead of one. Isn't this incorrect?</p>\n<p>My intuition so far is to use a weighted categorical crossentropy as metric both during evaluation and training and <code>binarycrossentropy</code> as loss function with weights:</p>\n<pre><code># github.com/wassname\n# https://gist.github.com/wassname/ce364fddfc8a025bfab4348cf5de852d\nfrom keras import backend as K\ndef weighted_categorical_crossentropy(weights):\n    \"\"\"python\n    A weighted version of keras.objectives.categorical_crossentropy\n\n    Variables:\n        weights: numpy array of shape (C,) where C is the number of classes\n\n    Usage:\n        weights = np.array([0.5,2,10]) # Class one at 0.5, class 2 twice the normal weights, class 3 10x.\n        loss = weighted_categorical_crossentropy(weights)\n        model.compile(loss=loss,optimizer='adam')\n    \"\"\"\n\n    weights = K.variable(weights)\n\n    def loss(y_true, y_pred):\n        # scale predictions so that the class probas of each sample sum to 1\n        y_pred /= K.sum(y_pred, axis=-1, keepdims=True)\n        # clip to prevent NaN's and Inf's\n        y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon())\n        # calc\n        loss = y_true * K.log(y_pred) * weights\n        loss = -K.sum(loss, -1)\n        return loss\n\n    return loss\n</code></pre>\n<p>Is this correct or is there better way?<br>\nTagging <a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> </p>",
  "messages": [
    {
      "id": 1885535,
      "postDate": "2022-08-05T08:44:12.260Z",
      "content": "<p></p>",
      "rawMarkdown": "~~If it works it works~~",
      "votes": 1
    },
    {
      "id": 1885166,
      "postDate": "2022-08-05T01:09:40.787Z",
      "content": "<p>The competition describes the evaluation to be done using <em>weighted multi-class log loss</em> (weighted categorical crossentropy). Is this supposed to be <code>metric</code> while training the model or <code>loss</code>?</p>\n<p>Many Notebooks in the code sections are using <code>crossentropy</code> as the loss function, and <code>accuracy</code> as the evaluation metric while the output layer is a Dense layer with 2 nodes instead of one. Isn't this incorrect?</p>\n<p>My intuition so far is to use a weighted categorical crossentropy as metric both during evaluation and training and <code>binarycrossentropy</code> as loss function with weights:</p>\n<pre><code># github.com/wassname\n# https://gist.github.com/wassname/ce364fddfc8a025bfab4348cf5de852d\nfrom keras import backend as K\ndef weighted_categorical_crossentropy(weights):\n    \"\"\"python\n    A weighted version of keras.objectives.categorical_crossentropy\n\n    Variables:\n        weights: numpy array of shape (C,) where C is the number of classes\n\n    Usage:\n        weights = np.array([0.5,2,10]) # Class one at 0.5, class 2 twice the normal weights, class 3 10x.\n        loss = weighted_categorical_crossentropy(weights)\n        model.compile(loss=loss,optimizer='adam')\n    \"\"\"\n\n    weights = K.variable(weights)\n\n    def loss(y_true, y_pred):\n        # scale predictions so that the class probas of each sample sum to 1\n        y_pred /= K.sum(y_pred, axis=-1, keepdims=True)\n        # clip to prevent NaN's and Inf's\n        y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon())\n        # calc\n        loss = y_true * K.log(y_pred) * weights\n        loss = -K.sum(loss, -1)\n        return loss\n\n    return loss\n</code></pre>\n<p>Is this correct or is there better way?<br>\nTagging <a href=\"https://www.kaggle.com/barbaroserdal\" target=\"_blank\">@barbaroserdal</a> </p>",
      "rawMarkdown": "The competition describes the evaluation to be done using *weighted multi-class log loss* (weighted categorical crossentropy). Is this supposed to be `metric` while training the model or `loss`?\n\nMany Notebooks in the code sections are using `crossentropy` as the loss function, and `accuracy` as the evaluation metric while the output layer is a Dense layer with 2 nodes instead of one. Isn't this incorrect?\n\nMy intuition so far is to use a weighted categorical crossentropy as metric both during evaluation and training and `binarycrossentropy` as loss function with weights:\n\n```\n# github.com/wassname\n# https://gist.github.com/wassname/ce364fddfc8a025bfab4348cf5de852d\nfrom keras import backend as K\ndef weighted_categorical_crossentropy(weights):\n    \"\"\"python\n    A weighted version of keras.objectives.categorical_crossentropy\n    \n    Variables:\n        weights: numpy array of shape (C,) where C is the number of classes\n    \n    Usage:\n        weights = np.array([0.5,2,10]) # Class one at 0.5, class 2 twice the normal weights, class 3 10x.\n        loss = weighted_categorical_crossentropy(weights)\n        model.compile(loss=loss,optimizer='adam')\n    \"\"\"\n    \n    weights = K.variable(weights)\n        \n    def loss(y_true, y_pred):\n        # scale predictions so that the class probas of each sample sum to 1\n        y_pred /= K.sum(y_pred, axis=-1, keepdims=True)\n        # clip to prevent NaN's and Inf's\n        y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon())\n        # calc\n        loss = y_true * K.log(y_pred) * weights\n        loss = -K.sum(loss, -1)\n        return loss\n    \n    return loss\n```\n\nIs this correct or is there better way?\nTagging @barbaroserdal ",
      "votes": 1
    },
    {
      "id": 1950128,
      "postDate": "2022-09-22T06:37:17.420Z",
      "content": "<p>There is no correct or incorrect way in data science. You have to find what works for you by trial and error.</p>\n<blockquote>\n  <p>Submissions are evaluated using a weighted multi-class logarithmic loss. The overall effect is such that each class is roughly equally important for the final score.</p>\n</blockquote>\n<p>I don't think we have to use weighted cross entropy loss function since each class is equally important. I'm currently using binary cross entropy for training and sklearn.metrics.log_loss for evaluation after sigmoiding logits.</p>",
      "rawMarkdown": "There is no correct or incorrect way in data science. You have to find what works for you by trial and error.\n\n> Submissions are evaluated using a weighted multi-class logarithmic loss. The overall effect is such that each class is roughly equally important for the final score.\n\nI don't think we have to use weighted cross entropy loss function since each class is equally important. I'm currently using binary cross entropy for training and sklearn.metrics.log_loss for evaluation after sigmoiding logits."
    },
    {
      "id": 1891665,
      "postDate": "2022-08-09T15:07:40.007Z",
      "content": "<blockquote>\n  <p>Is this supposed to be metric while training the model or loss?</p>\n</blockquote>\n<p>My intuition is that it doesn't have to be. We may train the model with a loss function of our choice and we may evaluate with whatever metric we want to. As a result we just submit a submission file as described in evaluation section and they apply a metric to this submission data. Still, that's my guess how it works.</p>\n<blockquote>\n  <p>My intuition so far is to use a weighted categorical crossentropy as metric both during evaluation and training and binarycrossentropy as loss function with weights:</p>\n</blockquote>\n<p>I also want to use binary cross entropy loss and evaluate the model using simple accuracy. Not sure if this will work well, but as for now I don't see a reason why it wouldn't work. </p>\n<p>Question for you: if you want to use weighted categorical crossentropy, how would you set weights for each class?</p>",
      "rawMarkdown": "> Is this supposed to be metric while training the model or loss?\n\nMy intuition is that it doesn't have to be. We may train the model with a loss function of our choice and we may evaluate with whatever metric we want to. As a result we just submit a submission file as described in evaluation section and they apply a metric to this submission data. Still, that's my guess how it works.\n\n> My intuition so far is to use a weighted categorical crossentropy as metric both during evaluation and training and binarycrossentropy as loss function with weights:\n\nI also want to use binary cross entropy loss and evaluate the model using simple accuracy. Not sure if this will work well, but as for now I don't see a reason why it wouldn't work. \n\nQuestion for you: if you want to use weighted categorical crossentropy, how would you set weights for each class?\n"
    },
    {
      "id": 1886327,
      "postDate": "2022-08-05T17:48:16.143Z",
      "content": "<p>If I understand correctly, binary classification can also be treated as multi-classification with 2 outputs. Your approach is also valid (binary-crossentropy with a single output). Accuracy may not be an optimal metric to estimate a model's performance, but we do not know the exact weights of each category.</p>",
      "rawMarkdown": "If I understand correctly, binary classification can also be treated as multi-classification with 2 outputs. Your approach is also valid (binary-crossentropy with a single output). Accuracy may not be an optimal metric to estimate a model's performance, but we do not know the exact weights of each category.",
      "replies": [
        {
          "id": 1886544,
          "postDate": "2022-08-06T00:22:07.397Z",
          "content": "<p>The weights are 0.5 and 0.5. <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335787#1848292\" target=\"_blank\">https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335787#1848292</a></p>",
          "rawMarkdown": "The weights are 0.5 and 0.5. https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335787#1848292",
          "votes": -1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1885535,
      "author_name": "majoraregalia",
      "author_url": "",
      "post_date": "2022-08-05T08:44:12.260000",
      "content": "<p></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1950128,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-09-22T06:37:17.420000",
      "content": "<p>There is no correct or incorrect way in data science. You have to find what works for you by trial and error.</p>\n<blockquote>\n  <p>Submissions are evaluated using a weighted multi-class logarithmic loss. The overall effect is such that each class is roughly equally important for the final score.</p>\n</blockquote>\n<p>I don't think we have to use weighted cross entropy loss function since each class is equally important. I'm currently using binary cross entropy for training and sklearn.metrics.log_loss for evaluation after sigmoiding logits.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1891665,
      "author_name": "AIGeekProgrammer",
      "author_url": "",
      "post_date": "2022-08-09T15:07:40.007000",
      "content": "<blockquote>\n  <p>Is this supposed to be metric while training the model or loss?</p>\n</blockquote>\n<p>My intuition is that it doesn't have to be. We may train the model with a loss function of our choice and we may evaluate with whatever metric we want to. As a result we just submit a submission file as described in evaluation section and they apply a metric to this submission data. Still, that's my guess how it works.</p>\n<blockquote>\n  <p>My intuition so far is to use a weighted categorical crossentropy as metric both during evaluation and training and binarycrossentropy as loss function with weights:</p>\n</blockquote>\n<p>I also want to use binary cross entropy loss and evaluate the model using simple accuracy. Not sure if this will work well, but as for now I don't see a reason why it wouldn't work. </p>\n<p>Question for you: if you want to use weighted categorical crossentropy, how would you set weights for each class?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1886327,
      "author_name": "hjunlee941",
      "author_url": "",
      "post_date": "2022-08-05T17:48:16.143000",
      "content": "<p>If I understand correctly, binary classification can also be treated as multi-classification with 2 outputs. Your approach is also valid (binary-crossentropy with a single output). Accuracy may not be an optimal metric to estimate a model's performance, but we do not know the exact weights of each category.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1886544,
          "author_name": "Anurag Dhadse",
          "author_url": "",
          "post_date": "2022-08-06T00:22:07.397000",
          "content": "<p>The weights are 0.5 and 0.5. <a href=\"https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335787#1848292\" target=\"_blank\">https://www.kaggle.com/competitions/mayo-clinic-strip-ai/discussion/335787#1848292</a></p>",
          "votes": -1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1885535": "~~If it works it works~~",
    "1885166": "The competition describes the evaluation to be done using *weighted multi-class log loss* (weighted categorical crossentropy). Is this supposed to be `metric` while training the model or `loss`?\n\nMany Notebooks in the code sections are using `crossentropy` as the loss function, and `accuracy` as the evaluation metric while the output layer is a Dense layer with 2 nodes instead of one. Isn't this incorrect?\n\nMy intuition so far is to use a weighted categorical crossentropy as metric both during evaluation and training and `binarycrossentropy` as loss function with weights:\n\n```\n# github.com/wassname\n# https://gist.github.com/wassname/ce364fddfc8a025bfab4348cf5de852d\nfrom keras import backend as K\ndef weighted_categorical_crossentropy(weights):\n    \"\"\"python\n    A weighted version of keras.objectives.categorical_crossentropy\n    \n    Variables:\n        weights: numpy array of shape (C,) where C is the number of classes\n    \n    Usage:\n        weights = np.array([0.5,2,10]) # Class one at 0.5, class 2 twice the normal weights, class 3 10x.\n        loss = weighted_categorical_crossentropy(weights)\n        model.compile(loss=loss,optimizer='adam')\n    \"\"\"\n    \n    weights = K.variable(weights)\n        \n    def loss(y_true, y_pred):\n        # scale predictions so that the class probas of each sample sum to 1\n        y_pred /= K.sum(y_pred, axis=-1, keepdims=True)\n        # clip to prevent NaN's and Inf's\n        y_pred = K.clip(y_pred, K.epsilon(), 1 - K.epsilon())\n        # calc\n        loss = y_true * K.log(y_pred) * weights\n        loss = -K.sum(loss, -1)\n        return loss\n    \n    return loss\n```\n\nIs this correct or is there better way?\nTagging @barbaroserdal ",
    "1950128": "There is no correct or incorrect way in data science. You have to find what works for you by trial and error.\n\n> Submissions are evaluated using a weighted multi-class logarithmic loss. The overall effect is such that each class is roughly equally important for the final score.\n\nI don't think we have to use weighted cross entropy loss function since each class is equally important. I'm currently using binary cross entropy for training and sklearn.metrics.log_loss for evaluation after sigmoiding logits.",
    "1891665": "> Is this supposed to be metric while training the model or loss?\n\nMy intuition is that it doesn't have to be. We may train the model with a loss function of our choice and we may evaluate with whatever metric we want to. As a result we just submit a submission file as described in evaluation section and they apply a metric to this submission data. Still, that's my guess how it works.\n\n> My intuition so far is to use a weighted categorical crossentropy as metric both during evaluation and training and binarycrossentropy as loss function with weights:\n\nI also want to use binary cross entropy loss and evaluate the model using simple accuracy. Not sure if this will work well, but as for now I don't see a reason why it wouldn't work. \n\nQuestion for you: if you want to use weighted categorical crossentropy, how would you set weights for each class?\n",
    "1886327": "If I understand correctly, binary classification can also be treated as multi-classification with 2 outputs. Your approach is also valid (binary-crossentropy with a single output). Accuracy may not be an optimal metric to estimate a model's performance, but we do not know the exact weights of each category."
  }
}