{
  "id": 369400,
  "title": "What's Best loss function in Breast Cancer Detection Task? [personal opinion]",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/369400",
  "author_name": "Wongi Park",
  "post_date": "2022-11-30T03:44:22.478000",
  "votes": 18,
  "comment_count": 5,
  "views": 0,
  "content": "<h2>F1 score</h2>\n<p>I think this Task requires f1 score, so we need to utilize the loss function related to f1 score.</p>\n<p>F1 score (also known as F-measure, or balanced F-score) is an error metric which measures model performance by calculating the harmonic mean of precision and recall for the minority positive class.</p>\n<h2>F1 loss function</h2>\n<pre><code>import torch\n\ndef f1_loss(y_true, y_pred):\n    tp = torch.sum(torch.tensor(y_true*y_pred).float(), 0)\n    tn = torch.sum(torch.tensor((1-y_true)*(1-y_pred)).float(), 0)\n    fp = torch.sum(torch.tensor((1-y_true)*y_pred).float(), 0)\n    fn = torch.sum(torch.tensor(y_true*(1-y_pred)).float(), 0)\n\n    p = tp/(tp+fp +1e-7)\n    r = tp/(tp+fn+ 1e-7)\n\n    f1 = 2*p*r / (p+r+1e-7)\n    f1 = torch.where(torch.isnan(f1), torch.zeros_like(f1), f1)\n    return 1-torch.mean(f1)\n</code></pre>\n<h2>F1 evaluate function</h2>\n<pre><code>def pfbeta(labels, predictions, beta):\n    y_true_count = 0\n    ctp = 0\n    cfp = 0\n\n    for idx in range(len(labels)):\n        prediction = min(max(predictions[idx], 0), 1)\n        if (labels[idx]):\n            y_true_count += 1\n            ctp += prediction\n            cfp += 1 - prediction\n        else:\n            cfp += prediction\n\n    beta_squared = beta * beta\n    c_precision = ctp / (ctp + cfp)\n    c_recall = ctp / y_true_count\n    if (c_precision &gt; 0 and c_recall &gt; 0):\n        result = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n        return result\n    else:\n        return 0\n</code></pre>",
  "messages": [
    {
      "id": 2049317,
      "postDate": "2022-11-30T03:44:22.480Z",
      "content": "<h2>F1 score</h2>\n<p>I think this Task requires f1 score, so we need to utilize the loss function related to f1 score.</p>\n<p>F1 score (also known as F-measure, or balanced F-score) is an error metric which measures model performance by calculating the harmonic mean of precision and recall for the minority positive class.</p>\n<h2>F1 loss function</h2>\n<pre><code>import torch\n\ndef f1_loss(y_true, y_pred):\n    tp = torch.sum(torch.tensor(y_true*y_pred).float(), 0)\n    tn = torch.sum(torch.tensor((1-y_true)*(1-y_pred)).float(), 0)\n    fp = torch.sum(torch.tensor((1-y_true)*y_pred).float(), 0)\n    fn = torch.sum(torch.tensor(y_true*(1-y_pred)).float(), 0)\n\n    p = tp/(tp+fp +1e-7)\n    r = tp/(tp+fn+ 1e-7)\n\n    f1 = 2*p*r / (p+r+1e-7)\n    f1 = torch.where(torch.isnan(f1), torch.zeros_like(f1), f1)\n    return 1-torch.mean(f1)\n</code></pre>\n<h2>F1 evaluate function</h2>\n<pre><code>def pfbeta(labels, predictions, beta):\n    y_true_count = 0\n    ctp = 0\n    cfp = 0\n\n    for idx in range(len(labels)):\n        prediction = min(max(predictions[idx], 0), 1)\n        if (labels[idx]):\n            y_true_count += 1\n            ctp += prediction\n            cfp += 1 - prediction\n        else:\n            cfp += prediction\n\n    beta_squared = beta * beta\n    c_precision = ctp / (ctp + cfp)\n    c_recall = ctp / y_true_count\n    if (c_precision &gt; 0 and c_recall &gt; 0):\n        result = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n        return result\n    else:\n        return 0\n</code></pre>",
      "rawMarkdown": "## F1 score\nI think this Task requires f1 score, so we need to utilize the loss function related to f1 score.\n\nF1 score (also known as F-measure, or balanced F-score) is an error metric which measures model performance by calculating the harmonic mean of precision and recall for the minority positive class.\n\n## F1 loss function\n```\nimport torch\n\ndef f1_loss(y_true, y_pred):\n    tp = torch.sum(torch.tensor(y_true*y_pred).float(), 0)\n    tn = torch.sum(torch.tensor((1-y_true)*(1-y_pred)).float(), 0)\n    fp = torch.sum(torch.tensor((1-y_true)*y_pred).float(), 0)\n    fn = torch.sum(torch.tensor(y_true*(1-y_pred)).float(), 0)\n\n    p = tp/(tp+fp +1e-7)\n    r = tp/(tp+fn+ 1e-7)\n\n    f1 = 2*p*r / (p+r+1e-7)\n    f1 = torch.where(torch.isnan(f1), torch.zeros_like(f1), f1)\n    return 1-torch.mean(f1)\n```\n\n## F1 evaluate function\n```\ndef pfbeta(labels, predictions, beta):\n    y_true_count = 0\n    ctp = 0\n    cfp = 0\n\n    for idx in range(len(labels)):\n        prediction = min(max(predictions[idx], 0), 1)\n        if (labels[idx]):\n            y_true_count += 1\n            ctp += prediction\n            cfp += 1 - prediction\n        else:\n            cfp += prediction\n\n    beta_squared = beta * beta\n    c_precision = ctp / (ctp + cfp)\n    c_recall = ctp / y_true_count\n    if (c_precision > 0 and c_recall > 0):\n        result = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n        return result\n    else:\n        return 0\n```",
      "votes": 18
    },
    {
      "id": 2049979,
      "postDate": "2022-11-30T13:13:34.353Z",
      "content": "<p>The problem of the F1-score is that it is not differentiable, thus it cannot be used as a loss function to compute gradients and update the weights when training the model. You can however use differentiable approximations that can be used as loss functions. Like <a href=\"https://datascience.stackexchange.com/questions/66581/is-it-possible-to-make-f1-score-differentiable-and-use-it-directly-as-a-loss-fun\" target=\"_blank\">Dice Loss</a> and <a href=\"https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d\" target=\"_blank\">soft F1 Loss</a></p>",
      "rawMarkdown": "The problem of the F1-score is that it is not differentiable, thus it cannot be used as a loss function to compute gradients and update the weights when training the model. You can however use differentiable approximations that can be used as loss functions. Like [Dice Loss](https://datascience.stackexchange.com/questions/66581/is-it-possible-to-make-f1-score-differentiable-and-use-it-directly-as-a-loss-fun) and [soft F1 Loss](https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d)",
      "votes": 11,
      "replies": [
        {
          "id": 2050138,
          "postDate": "2022-11-30T14:47:55.917Z",
          "content": "<p>I didn't even think about it. Thank's give me advice! but, I think this is useful for measurement.</p>",
          "rawMarkdown": "I didn't even think about it. Thank's give me advice! but, I think this is useful for measurement.",
          "votes": 1
        },
        {
          "id": 2051176,
          "postDate": "2022-12-01T09:08:03.580Z",
          "content": "<p>Good point, <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>! Will be interesting to see how well these losses will help with addressing the class imbalance.   </p>",
          "rawMarkdown": "Good point, @alejopaullier! Will be interesting to see how well these losses will help with addressing the class imbalance.   "
        },
        {
          "id": 2051177,
          "postDate": "2022-12-01T09:09:04.080Z",
          "content": "<p>This got me thinking -- one other thing that might work really well here could be label smoothing. Some of these training examples might be really hard for our model + label smoothing helps with calibration, might be worth  try.</p>",
          "rawMarkdown": "This got me thinking -- one other thing that might work really well here could be label smoothing. Some of these training examples might be really hard for our model + label smoothing helps with calibration, might be worth  try."
        }
      ]
    },
    {
      "id": 2052614,
      "postDate": "2022-12-02T10:52:09.717Z",
      "content": "<p>What is the value of beta in the competition?</p>",
      "rawMarkdown": "What is the value of beta in the competition?",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2049979,
      "author_name": "moth",
      "author_url": "",
      "post_date": "2022-11-30T13:13:34.353000",
      "content": "<p>The problem of the F1-score is that it is not differentiable, thus it cannot be used as a loss function to compute gradients and update the weights when training the model. You can however use differentiable approximations that can be used as loss functions. Like <a href=\"https://datascience.stackexchange.com/questions/66581/is-it-possible-to-make-f1-score-differentiable-and-use-it-directly-as-a-loss-fun\" target=\"_blank\">Dice Loss</a> and <a href=\"https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d\" target=\"_blank\">soft F1 Loss</a></p>",
      "votes": 11,
      "replies": [
        {
          "id": 2050138,
          "author_name": "Wongi Park",
          "author_url": "",
          "post_date": "2022-11-30T14:47:55.917000",
          "content": "<p>I didn't even think about it. Thank's give me advice! but, I think this is useful for measurement.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2051176,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-01T09:08:03.580000",
          "content": "<p>Good point, <a href=\"https://www.kaggle.com/alejopaullier\" target=\"_blank\">@alejopaullier</a>! Will be interesting to see how well these losses will help with addressing the class imbalance.   </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2051177,
          "author_name": "Radek Osmulski",
          "author_url": "",
          "post_date": "2022-12-01T09:09:04.080000",
          "content": "<p>This got me thinking -- one other thing that might work really well here could be label smoothing. Some of these training examples might be really hard for our model + label smoothing helps with calibration, might be worth  try.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2052614,
      "author_name": "Vladimir Slaykovskiy",
      "author_url": "",
      "post_date": "2022-12-02T10:52:09.717000",
      "content": "<p>What is the value of beta in the competition?</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2049317": "## F1 score\nI think this Task requires f1 score, so we need to utilize the loss function related to f1 score.\n\nF1 score (also known as F-measure, or balanced F-score) is an error metric which measures model performance by calculating the harmonic mean of precision and recall for the minority positive class.\n\n## F1 loss function\n```\nimport torch\n\ndef f1_loss(y_true, y_pred):\n    tp = torch.sum(torch.tensor(y_true*y_pred).float(), 0)\n    tn = torch.sum(torch.tensor((1-y_true)*(1-y_pred)).float(), 0)\n    fp = torch.sum(torch.tensor((1-y_true)*y_pred).float(), 0)\n    fn = torch.sum(torch.tensor(y_true*(1-y_pred)).float(), 0)\n\n    p = tp/(tp+fp +1e-7)\n    r = tp/(tp+fn+ 1e-7)\n\n    f1 = 2*p*r / (p+r+1e-7)\n    f1 = torch.where(torch.isnan(f1), torch.zeros_like(f1), f1)\n    return 1-torch.mean(f1)\n```\n\n## F1 evaluate function\n```\ndef pfbeta(labels, predictions, beta):\n    y_true_count = 0\n    ctp = 0\n    cfp = 0\n\n    for idx in range(len(labels)):\n        prediction = min(max(predictions[idx], 0), 1)\n        if (labels[idx]):\n            y_true_count += 1\n            ctp += prediction\n            cfp += 1 - prediction\n        else:\n            cfp += prediction\n\n    beta_squared = beta * beta\n    c_precision = ctp / (ctp + cfp)\n    c_recall = ctp / y_true_count\n    if (c_precision > 0 and c_recall > 0):\n        result = (1 + beta_squared) * (c_precision * c_recall) / (beta_squared * c_precision + c_recall)\n        return result\n    else:\n        return 0\n```",
    "2049979": "The problem of the F1-score is that it is not differentiable, thus it cannot be used as a loss function to compute gradients and update the weights when training the model. You can however use differentiable approximations that can be used as loss functions. Like [Dice Loss](https://datascience.stackexchange.com/questions/66581/is-it-possible-to-make-f1-score-differentiable-and-use-it-directly-as-a-loss-fun) and [soft F1 Loss](https://towardsdatascience.com/the-unknown-benefits-of-using-a-soft-f1-loss-in-classification-systems-753902c0105d)",
    "2052614": "What is the value of beta in the competition?"
  }
}