{
  "id": 193598,
  "title": "How To Beat Baselines Explained",
  "url": "/competitions/rsna-str-pulmonary-embolism-detection/discussion/193598",
  "author_name": "Chris Deotte",
  "post_date": "2020-10-27T19:33:59.640000",
  "votes": 39,
  "comment_count": 8,
  "views": 0,
  "content": "<p>For the first few months of this competition, most teams struggled to beat the baselines that used simple mean predictions such as OsciiArt's LB 0.325 <a href=\"https://www.kaggle.com/osciiart/baseline-with-no-image\" target=\"_blank\">here</a> and Paulo Pinto's LB 0.434 <a href=\"https://www.kaggle.com/paulorzp/mean-baseline\" target=\"_blank\">here</a>. Why were CNN struggling to beat simple means?</p>\n<p>The answer is the competition metric. When training, we need to use log loss with sample weights. With the correct loss, we can beat these baselines using a CNN training on only 32x32 images! And only using a portion of the train data!</p>\n<h1>Reduce Data</h1>\n<p>First we observe that images from patients without pulmonary embolism do not affect CV LB, therefore remove these 70% of images from train data</p>\n<pre><code>train['weight'] = train.groupby('StudyInstanceUID')\\\n                       .pe_present_on_image.transform('mean')\nX_train = train.loc[train.weight&gt;0]\nprint( X_train.shape[0] / train.shape[0] )\n# THIS PRINTS 30%\n</code></pre>\n<h1>The Trick</h1>\n<p>The metric for image level prediction <code>pe_present_on_image</code> is a sample weighted log loss, therefore you need to use <code>sample_weight</code> when training. </p>\n<pre><code>model.compile(loss='binary_crossentropy', optimizer = 'adam')\nmodel.fit(X_train['image'], X_train['pe_present_on_image'], \n          sample_weight = X_train['weight'])\n</code></pre>\n<p>Here <code>X_train</code> <code>weight</code> is the same that was computed above.</p>\n<h1>Understanding the Metric</h1>\n<p>The above two insights come from understanding this competition metric. The description page is very confusing. The bottom line is that the competition metric is a weighted average of 10 log losses. And the 10th log loss, the image log loss is a sample weighted log loss.</p>\n<p>They give us the 9 weights for exam level predictions and we need to compute the weight for the 10th image level prediction <code>pe_present_on_image</code>. Their 9 weights add up to 1, so the metric is simply:</p>\n<pre><code>metric = 0\nfor k in range(9): metric += exam_weight[k] * exam_logloss[k]\nmetric += image_weight * image_weighted_logloss\nmetric /= (1+image_weight)\n</code></pre>\n<p>There are many posts about computing the metric. Some are wrong. Many are inefficient. The metric can be computed below in a few seconds.</p>\n<pre><code># DATAFRAME TRUE IS THE ORIGINAL TRAIN.CSV FILE\n# DATAFRAME VALID IS TRAIN.CSV WITH YOUR PREDICTIONS\n\n# TEN WEIGHTS\nexam_weights = {\n    \"negative_exam_for_pe\": 0.0736196319,\n    \"indeterminate\": 0.09202453988,\n    \"chronic_pe\": 0.1042944785,\n    \"acute_and_chronic_pe\": 0.1042944785,\n    \"central_pe\": 0.1877300613,\n    \"leftsided_pe\": 0.06257668712,\n    \"rightsided_pe\": 0.06257668712,\n    \"rv_lv_ratio_gte_1\": 0.2346625767,\n    \"rv_lv_ratio_lt_1\": 0.0782208589\n}\nimage_weight = np.sum( true.groupby(\"StudyInstanceUID\")\\\n    ['pe_present_on_image'].transform(\"mean\") )\nimage_weight *= 0.07361963 / true.StudyInstanceUID.nunique()\n\n# EXAM PREDICTIONS\nTARGETS = list( exam_weights.keys() )\nexam_pred = valid.groupby('StudyInstanceUID')[TARGETS].mean()\nexam_true = true.groupby('StudyInstanceUID')[TARGETS].mean()\n\n# COMPUTE WEIGHTED AVERAGE OF TEN LOG LOSSES\nmetric = 0\n# NINE EXAM LOG LOSSES\nfor target, weight in exam_weights.items():\n    metric += weight * log_loss(exam_true[target], exam_pred[target])\n# ONE IMAGE WEIGHTED LOG LOSS\nimage_loss = log_loss(true['pe_present_on_image'], \n    valid['pe_present_on_image'],\n    sample_weight = true.groupby(\"StudyInstanceUID\")\\ \n     ['pe_present_on_image'].transform(\"mean\"))\n# WEIGHTED AVERAGE\nmetric += image_weight * image_loss\nmetric /= (1 + image_weight)\n</code></pre>",
  "messages": [
    {
      "id": 1062397,
      "postDate": "2020-10-27T19:33:59.640Z",
      "content": "<p>For the first few months of this competition, most teams struggled to beat the baselines that used simple mean predictions such as OsciiArt's LB 0.325 <a href=\"https://www.kaggle.com/osciiart/baseline-with-no-image\" target=\"_blank\">here</a> and Paulo Pinto's LB 0.434 <a href=\"https://www.kaggle.com/paulorzp/mean-baseline\" target=\"_blank\">here</a>. Why were CNN struggling to beat simple means?</p>\n<p>The answer is the competition metric. When training, we need to use log loss with sample weights. With the correct loss, we can beat these baselines using a CNN training on only 32x32 images! And only using a portion of the train data!</p>\n<h1>Reduce Data</h1>\n<p>First we observe that images from patients without pulmonary embolism do not affect CV LB, therefore remove these 70% of images from train data</p>\n<pre><code>train['weight'] = train.groupby('StudyInstanceUID')\\\n                       .pe_present_on_image.transform('mean')\nX_train = train.loc[train.weight&gt;0]\nprint( X_train.shape[0] / train.shape[0] )\n# THIS PRINTS 30%\n</code></pre>\n<h1>The Trick</h1>\n<p>The metric for image level prediction <code>pe_present_on_image</code> is a sample weighted log loss, therefore you need to use <code>sample_weight</code> when training. </p>\n<pre><code>model.compile(loss='binary_crossentropy', optimizer = 'adam')\nmodel.fit(X_train['image'], X_train['pe_present_on_image'], \n          sample_weight = X_train['weight'])\n</code></pre>\n<p>Here <code>X_train</code> <code>weight</code> is the same that was computed above.</p>\n<h1>Understanding the Metric</h1>\n<p>The above two insights come from understanding this competition metric. The description page is very confusing. The bottom line is that the competition metric is a weighted average of 10 log losses. And the 10th log loss, the image log loss is a sample weighted log loss.</p>\n<p>They give us the 9 weights for exam level predictions and we need to compute the weight for the 10th image level prediction <code>pe_present_on_image</code>. Their 9 weights add up to 1, so the metric is simply:</p>\n<pre><code>metric = 0\nfor k in range(9): metric += exam_weight[k] * exam_logloss[k]\nmetric += image_weight * image_weighted_logloss\nmetric /= (1+image_weight)\n</code></pre>\n<p>There are many posts about computing the metric. Some are wrong. Many are inefficient. The metric can be computed below in a few seconds.</p>\n<pre><code># DATAFRAME TRUE IS THE ORIGINAL TRAIN.CSV FILE\n# DATAFRAME VALID IS TRAIN.CSV WITH YOUR PREDICTIONS\n\n# TEN WEIGHTS\nexam_weights = {\n    \"negative_exam_for_pe\": 0.0736196319,\n    \"indeterminate\": 0.09202453988,\n    \"chronic_pe\": 0.1042944785,\n    \"acute_and_chronic_pe\": 0.1042944785,\n    \"central_pe\": 0.1877300613,\n    \"leftsided_pe\": 0.06257668712,\n    \"rightsided_pe\": 0.06257668712,\n    \"rv_lv_ratio_gte_1\": 0.2346625767,\n    \"rv_lv_ratio_lt_1\": 0.0782208589\n}\nimage_weight = np.sum( true.groupby(\"StudyInstanceUID\")\\\n    ['pe_present_on_image'].transform(\"mean\") )\nimage_weight *= 0.07361963 / true.StudyInstanceUID.nunique()\n\n# EXAM PREDICTIONS\nTARGETS = list( exam_weights.keys() )\nexam_pred = valid.groupby('StudyInstanceUID')[TARGETS].mean()\nexam_true = true.groupby('StudyInstanceUID')[TARGETS].mean()\n\n# COMPUTE WEIGHTED AVERAGE OF TEN LOG LOSSES\nmetric = 0\n# NINE EXAM LOG LOSSES\nfor target, weight in exam_weights.items():\n    metric += weight * log_loss(exam_true[target], exam_pred[target])\n# ONE IMAGE WEIGHTED LOG LOSS\nimage_loss = log_loss(true['pe_present_on_image'], \n    valid['pe_present_on_image'],\n    sample_weight = true.groupby(\"StudyInstanceUID\")\\ \n     ['pe_present_on_image'].transform(\"mean\"))\n# WEIGHTED AVERAGE\nmetric += image_weight * image_loss\nmetric /= (1 + image_weight)\n</code></pre>",
      "rawMarkdown": "For the first few months of this competition, most teams struggled to beat the baselines that used simple mean predictions such as OsciiArt's LB 0.325 [here][1] and Paulo Pinto's LB 0.434 [here][2]. Why were CNN struggling to beat simple means?\n\nThe answer is the competition metric. When training, we need to use log loss with sample weights. With the correct loss, we can beat these baselines using a CNN training on only 32x32 images! And only using a portion of the train data!\n\n# Reduce Data\nFirst we observe that images from patients without pulmonary embolism do not affect CV LB, therefore remove these 70% of images from train data\n\n    train['weight'] = train.groupby('StudyInstanceUID')\\\n                           .pe_present_on_image.transform('mean')\n    X_train = train.loc[train.weight>0]\n    print( X_train.shape[0] / train.shape[0] )\n    # THIS PRINTS 30%\n\n# The Trick\nThe metric for image level prediction `pe_present_on_image` is a sample weighted log loss, therefore you need to use `sample_weight` when training. \n\n    model.compile(loss='binary_crossentropy', optimizer = 'adam')\n    model.fit(X_train['image'], X_train['pe_present_on_image'], \n              sample_weight = X_train['weight'])\n\nHere `X_train` `weight` is the same that was computed above.\n\n# Understanding the Metric\nThe above two insights come from understanding this competition metric. The description page is very confusing. The bottom line is that the competition metric is a weighted average of 10 log losses. And the 10th log loss, the image log loss is a sample weighted log loss.\n\nThey give us the 9 weights for exam level predictions and we need to compute the weight for the 10th image level prediction `pe_present_on_image`. Their 9 weights add up to 1, so the metric is simply:\n\n    metric = 0\n    for k in range(9): metric += exam_weight[k] * exam_logloss[k]\n    metric += image_weight * image_weighted_logloss\n    metric /= (1+image_weight)\n\nThere are many posts about computing the metric. Some are wrong. Many are inefficient. The metric can be computed below in a few seconds.\n\n    # DATAFRAME TRUE IS THE ORIGINAL TRAIN.CSV FILE\n    # DATAFRAME VALID IS TRAIN.CSV WITH YOUR PREDICTIONS\n\n    # TEN WEIGHTS\n    exam_weights = {\n        \"negative_exam_for_pe\": 0.0736196319,\n        \"indeterminate\": 0.09202453988,\n        \"chronic_pe\": 0.1042944785,\n        \"acute_and_chronic_pe\": 0.1042944785,\n        \"central_pe\": 0.1877300613,\n        \"leftsided_pe\": 0.06257668712,\n        \"rightsided_pe\": 0.06257668712,\n        \"rv_lv_ratio_gte_1\": 0.2346625767,\n        \"rv_lv_ratio_lt_1\": 0.0782208589\n    }\n    image_weight = np.sum( true.groupby(\"StudyInstanceUID\")\\\n        ['pe_present_on_image'].transform(\"mean\") )\n    image_weight *= 0.07361963 / true.StudyInstanceUID.nunique()\n\n    # EXAM PREDICTIONS\n    TARGETS = list( exam_weights.keys() )\n    exam_pred = valid.groupby('StudyInstanceUID')[TARGETS].mean()\n    exam_true = true.groupby('StudyInstanceUID')[TARGETS].mean()\n\n    # COMPUTE WEIGHTED AVERAGE OF TEN LOG LOSSES\n    metric = 0\n    # NINE EXAM LOG LOSSES\n    for target, weight in exam_weights.items():\n        metric += weight * log_loss(exam_true[target], exam_pred[target])\n    # ONE IMAGE WEIGHTED LOG LOSS\n    image_loss = log_loss(true['pe_present_on_image'], \n        valid['pe_present_on_image'],\n        sample_weight = true.groupby(\"StudyInstanceUID\")\\ \n         ['pe_present_on_image'].transform(\"mean\"))\n    # WEIGHTED AVERAGE\n    metric += image_weight * image_loss\n    metric /= (1 + image_weight)\n\n[1]: https://www.kaggle.com/osciiart/baseline-with-no-image\n[2]: https://www.kaggle.com/paulorzp/mean-baseline",
      "votes": 39
    },
    {
      "id": 1062406,
      "postDate": "2020-10-27T19:44:11.327Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> sir thank you for this awesome writeup and also for this video : <a href=\"https://www.youtube.com/watch?v=L1QKTPb6V_I\" target=\"_blank\">Grandmaster Series – How to Build a World-Class ML Model for Melanoma Detection</a></p>",
      "rawMarkdown": "@cdeotte sir thank you for this awesome writeup and also for this video : [Grandmaster Series – How to Build a World-Class ML Model for Melanoma Detection](https://www.youtube.com/watch?v=L1QKTPb6V_I)",
      "votes": 4
    },
    {
      "id": 1067756,
      "postDate": "2020-11-02T17:29:09.373Z",
      "content": "<p>wow!! Thanks! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>\n<p>clearly shows how much all of us can learn… and thanks for making this community a great place to share and learn from great minds!!</p>",
      "rawMarkdown": "wow!! Thanks! @cdeotte\n\nclearly shows how much all of us can learn... and thanks for making this community a great place to share and learn from great minds!!",
      "votes": 1
    },
    {
      "id": 1063040,
      "postDate": "2020-10-28T13:01:54.647Z",
      "content": "<p>Wow. It is really helpful for me. Thanks! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "Wow. It is really helpful for me. Thanks! @cdeotte ",
      "votes": 1
    },
    {
      "id": 1062958,
      "postDate": "2020-10-28T10:58:45.977Z",
      "content": "<p>Nice explanations. Interesting. We have the same observation with the competition metric. With that revelation on the competition metric, our team approached it slightly differently. </p>\n<p>My (maybe flawed) principle is try not to throw away any training data.<br>\nWe know 70% of studies are negative, 30% of studies are positive, and ~5% of all images are positive.<br>\nKnowing that negative studies don't count towards the LB, the distribution of positive images that will be counted towards the LB is about 5% of all images divided by 30% of all studies = ~15%</p>\n<p>To ensure our trained model learn this distribution, we train our models with positive weighting of 3.0 (3 times of 5% to get 15%), and further tuned it to 2.0-2.5.</p>",
      "rawMarkdown": "Nice explanations. Interesting. We have the same observation with the competition metric. With that revelation on the competition metric, our team approached it slightly differently. \n\nMy (maybe flawed) principle is try not to throw away any training data.\nWe know 70% of studies are negative, 30% of studies are positive, and ~5% of all images are positive.\nKnowing that negative studies don't count towards the LB, the distribution of positive images that will be counted towards the LB is about 5% of all images divided by 30% of all studies = ~15%\n\nTo ensure our trained model learn this distribution, we train our models with positive weighting of 3.0 (3 times of 5% to get 15%), and further tuned it to 2.0-2.5.",
      "votes": 1
    },
    {
      "id": 1062455,
      "postDate": "2020-10-27T21:06:41.917Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the helpful information!</p>",
      "rawMarkdown": "@cdeotte Thanks for the helpful information!",
      "votes": 1
    },
    {
      "id": 1062616,
      "postDate": "2020-10-28T03:24:54.780Z",
      "content": "<p>You can throw it out 70% of the data….omg, so much frustration in the past few weeks training 1TB of data for days and days to no avail. Thank you for this precious lesson.</p>",
      "rawMarkdown": "You can throw it out 70% of the data....omg, so much frustration in the past few weeks training 1TB of data for days and days to no avail. Thank you for this precious lesson.",
      "votes": 2
    },
    {
      "id": 1063291,
      "postDate": "2020-10-28T17:49:37.823Z",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> …</p>",
      "rawMarkdown": "Thanks a lot @cdeotte ...",
      "votes": 1
    },
    {
      "id": 1062414,
      "postDate": "2020-10-27T19:51:24.857Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks sir.</p>",
      "rawMarkdown": "@cdeotte thanks sir.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1062406,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2020-10-27T19:44:11.327000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> sir thank you for this awesome writeup and also for this video : <a href=\"https://www.youtube.com/watch?v=L1QKTPb6V_I\" target=\"_blank\">Grandmaster Series – How to Build a World-Class ML Model for Melanoma Detection</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1067756,
      "author_name": "Kamal Das",
      "author_url": "",
      "post_date": "2020-11-02T17:29:09.373000",
      "content": "<p>wow!! Thanks! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a></p>\n<p>clearly shows how much all of us can learn… and thanks for making this community a great place to share and learn from great minds!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1063040,
      "author_name": "Heroseo",
      "author_url": "",
      "post_date": "2020-10-28T13:01:54.647000",
      "content": "<p>Wow. It is really helpful for me. Thanks! <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062958,
      "author_name": "Yee Ng",
      "author_url": "",
      "post_date": "2020-10-28T10:58:45.977000",
      "content": "<p>Nice explanations. Interesting. We have the same observation with the competition metric. With that revelation on the competition metric, our team approached it slightly differently. </p>\n<p>My (maybe flawed) principle is try not to throw away any training data.<br>\nWe know 70% of studies are negative, 30% of studies are positive, and ~5% of all images are positive.<br>\nKnowing that negative studies don't count towards the LB, the distribution of positive images that will be counted towards the LB is about 5% of all images divided by 30% of all studies = ~15%</p>\n<p>To ensure our trained model learn this distribution, we train our models with positive weighting of 3.0 (3 times of 5% to get 15%), and further tuned it to 2.0-2.5.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062455,
      "author_name": "Brenda N",
      "author_url": "",
      "post_date": "2020-10-27T21:06:41.917000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Thanks for the helpful information!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062616,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-10-28T03:24:54.780000",
      "content": "<p>You can throw it out 70% of the data….omg, so much frustration in the past few weeks training 1TB of data for days and days to no avail. Thank you for this precious lesson.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1063291,
      "author_name": "Manoj Prabhakar",
      "author_url": "",
      "post_date": "2020-10-28T17:49:37.823000",
      "content": "<p>Thanks a lot <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> …</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1062414,
      "author_name": "ask9",
      "author_url": "",
      "post_date": "2020-10-27T19:51:24.857000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thanks sir.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1062397": "For the first few months of this competition, most teams struggled to beat the baselines that used simple mean predictions such as OsciiArt's LB 0.325 [here][1] and Paulo Pinto's LB 0.434 [here][2]. Why were CNN struggling to beat simple means?\n\nThe answer is the competition metric. When training, we need to use log loss with sample weights. With the correct loss, we can beat these baselines using a CNN training on only 32x32 images! And only using a portion of the train data!\n\n# Reduce Data\nFirst we observe that images from patients without pulmonary embolism do not affect CV LB, therefore remove these 70% of images from train data\n\n    train['weight'] = train.groupby('StudyInstanceUID')\\\n                           .pe_present_on_image.transform('mean')\n    X_train = train.loc[train.weight>0]\n    print( X_train.shape[0] / train.shape[0] )\n    # THIS PRINTS 30%\n\n# The Trick\nThe metric for image level prediction `pe_present_on_image` is a sample weighted log loss, therefore you need to use `sample_weight` when training. \n\n    model.compile(loss='binary_crossentropy', optimizer = 'adam')\n    model.fit(X_train['image'], X_train['pe_present_on_image'], \n              sample_weight = X_train['weight'])\n\nHere `X_train` `weight` is the same that was computed above.\n\n# Understanding the Metric\nThe above two insights come from understanding this competition metric. The description page is very confusing. The bottom line is that the competition metric is a weighted average of 10 log losses. And the 10th log loss, the image log loss is a sample weighted log loss.\n\nThey give us the 9 weights for exam level predictions and we need to compute the weight for the 10th image level prediction `pe_present_on_image`. Their 9 weights add up to 1, so the metric is simply:\n\n    metric = 0\n    for k in range(9): metric += exam_weight[k] * exam_logloss[k]\n    metric += image_weight * image_weighted_logloss\n    metric /= (1+image_weight)\n\nThere are many posts about computing the metric. Some are wrong. Many are inefficient. The metric can be computed below in a few seconds.\n\n    # DATAFRAME TRUE IS THE ORIGINAL TRAIN.CSV FILE\n    # DATAFRAME VALID IS TRAIN.CSV WITH YOUR PREDICTIONS\n\n    # TEN WEIGHTS\n    exam_weights = {\n        \"negative_exam_for_pe\": 0.0736196319,\n        \"indeterminate\": 0.09202453988,\n        \"chronic_pe\": 0.1042944785,\n        \"acute_and_chronic_pe\": 0.1042944785,\n        \"central_pe\": 0.1877300613,\n        \"leftsided_pe\": 0.06257668712,\n        \"rightsided_pe\": 0.06257668712,\n        \"rv_lv_ratio_gte_1\": 0.2346625767,\n        \"rv_lv_ratio_lt_1\": 0.0782208589\n    }\n    image_weight = np.sum( true.groupby(\"StudyInstanceUID\")\\\n        ['pe_present_on_image'].transform(\"mean\") )\n    image_weight *= 0.07361963 / true.StudyInstanceUID.nunique()\n\n    # EXAM PREDICTIONS\n    TARGETS = list( exam_weights.keys() )\n    exam_pred = valid.groupby('StudyInstanceUID')[TARGETS].mean()\n    exam_true = true.groupby('StudyInstanceUID')[TARGETS].mean()\n\n    # COMPUTE WEIGHTED AVERAGE OF TEN LOG LOSSES\n    metric = 0\n    # NINE EXAM LOG LOSSES\n    for target, weight in exam_weights.items():\n        metric += weight * log_loss(exam_true[target], exam_pred[target])\n    # ONE IMAGE WEIGHTED LOG LOSS\n    image_loss = log_loss(true['pe_present_on_image'], \n        valid['pe_present_on_image'],\n        sample_weight = true.groupby(\"StudyInstanceUID\")\\ \n         ['pe_present_on_image'].transform(\"mean\"))\n    # WEIGHTED AVERAGE\n    metric += image_weight * image_loss\n    metric /= (1 + image_weight)\n\n[1]: https://www.kaggle.com/osciiart/baseline-with-no-image\n[2]: https://www.kaggle.com/paulorzp/mean-baseline",
    "1062406": "@cdeotte sir thank you for this awesome writeup and also for this video : [Grandmaster Series – How to Build a World-Class ML Model for Melanoma Detection](https://www.youtube.com/watch?v=L1QKTPb6V_I)",
    "1067756": "wow!! Thanks! @cdeotte\n\nclearly shows how much all of us can learn... and thanks for making this community a great place to share and learn from great minds!!",
    "1063040": "Wow. It is really helpful for me. Thanks! @cdeotte ",
    "1062958": "Nice explanations. Interesting. We have the same observation with the competition metric. With that revelation on the competition metric, our team approached it slightly differently. \n\nMy (maybe flawed) principle is try not to throw away any training data.\nWe know 70% of studies are negative, 30% of studies are positive, and ~5% of all images are positive.\nKnowing that negative studies don't count towards the LB, the distribution of positive images that will be counted towards the LB is about 5% of all images divided by 30% of all studies = ~15%\n\nTo ensure our trained model learn this distribution, we train our models with positive weighting of 3.0 (3 times of 5% to get 15%), and further tuned it to 2.0-2.5.",
    "1062455": "@cdeotte Thanks for the helpful information!",
    "1062616": "You can throw it out 70% of the data....omg, so much frustration in the past few weeks training 1TB of data for days and days to no avail. Thank you for this precious lesson.",
    "1063291": "Thanks a lot @cdeotte ...",
    "1062414": "@cdeotte thanks sir."
  }
}