{
  "id": 145574,
  "title": "Watch Out For Suspicious Masks",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/145574",
  "author_name": "Leonie",
  "post_date": "2020-04-23T17:53:07.520000",
  "votes": 33,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Many of you already know that there are 100 images without a mask. \nDuring my EDA I found two more suspicious things:\n* There are 4 masks that only contain background values\n* There are 85 masks that do not contain pixels marked as cancerous although the image is labeled with an ISUP grade &gt; 0</p>\n\n<p>It might be a good idea to avoid using these test cases for training you model but instead use them for validation. I am not medically trained but for know I am assuming the <code>isup_grade</code> is still correct for the concering <code>image_id</code>s.</p>\n\n<p>Details and a .csv containing all suspicious <code>image_id</code>s can be found in <a href=\"https://www.kaggle.com/iamleonie/panda-eda-visualizations-suspicious-data\">my EDA</a>.</p>",
  "messages": [
    {
      "id": 818226,
      "postDate": "2020-04-23T17:53:07.520Z",
      "content": "<p>Many of you already know that there are 100 images without a mask. \nDuring my EDA I found two more suspicious things:\n* There are 4 masks that only contain background values\n* There are 85 masks that do not contain pixels marked as cancerous although the image is labeled with an ISUP grade &gt; 0</p>\n\n<p>It might be a good idea to avoid using these test cases for training you model but instead use them for validation. I am not medically trained but for know I am assuming the <code>isup_grade</code> is still correct for the concering <code>image_id</code>s.</p>\n\n<p>Details and a .csv containing all suspicious <code>image_id</code>s can be found in <a href=\"https://www.kaggle.com/iamleonie/panda-eda-visualizations-suspicious-data\">my EDA</a>.</p>",
      "rawMarkdown": "Many of you already know that there are 100 images without a mask. \nDuring my EDA I found two more suspicious things:\n* There are 4 masks that only contain background values\n* There are 85 masks that do not contain pixels marked as cancerous although the image is labeled with an ISUP grade &gt; 0\n\nIt might be a good idea to avoid using these test cases for training you model but instead use them for validation. I am not medically trained but for know I am assuming the `isup_grade` is still correct for the concering `image_id`s.\n\nDetails and a .csv containing all suspicious `image_id`s can be found in [my EDA](https://www.kaggle.com/iamleonie/panda-eda-visualizations-suspicious-data).",
      "votes": 32
    },
    {
      "id": 818246,
      "postDate": "2020-04-23T18:09:23.710Z",
      "content": "<p>From <a href=\"https://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset\">https://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset</a>:</p>\n\n<blockquote>\n  <p>The label masks of Radboudumc were semi-automatically generated by several deep learning algorithms, contain noise, and can be considered as weakly-supervised labels. The label masks of Karolinska were semi-autotomatically generated based on annotations by a pathologist.</p>\n</blockquote>\n\n<p>I noticed the same thing you did. The 85 masks that do not contain pixels labeled as cancerous even though ISUP grade &gt; 0 are from Radboudumc, so the DL model probably just missed them. These are probably tougher cases so might be a good set to test your model against harder cases. </p>\n\n<p>The Karolinska labels were generated based on pathologist annotations, so likely are more reliable. However, they are not as granular as the DL-generated labels.</p>\n\n<p>In any case, these seem to be weak labels, so we need to be careful about using them.</p>",
      "rawMarkdown": "From https://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset:\n\n&gt; The label masks of Radboudumc were semi-automatically generated by several deep learning algorithms, contain noise, and can be considered as weakly-supervised labels. The label masks of Karolinska were semi-autotomatically generated based on annotations by a pathologist.\n\nI noticed the same thing you did. The 85 masks that do not contain pixels labeled as cancerous even though ISUP grade &gt; 0 are from Radboudumc, so the DL model probably just missed them. These are probably tougher cases so might be a good set to test your model against harder cases. \n\nThe Karolinska labels were generated based on pathologist annotations, so likely are more reliable. However, they are not as granular as the DL-generated labels.\n\nIn any case, these seem to be weak labels, so we need to be careful about using them.",
      "votes": 12,
      "replies": [
        {
          "id": 818277,
          "postDate": "2020-04-23T18:40:43.293Z",
          "content": "<p>Thank you for explanation. This helped a lot!</p>",
          "rawMarkdown": "Thank you for explanation. This helped a lot!"
        },
        {
          "id": 820897,
          "postDate": "2020-04-25T19:40:09.933Z",
          "rawMarkdown": ""
        },
        {
          "id": 820899,
          "postDate": "2020-04-25T19:42:30.477Z",
          "content": "<p><a href=\"/iamleonie\">@iamleonie</a>,</p>\n\n<p>when doing augmentation on the images we have to apply the same augmentation to the mask too?</p>",
          "rawMarkdown": "@iamleonie,\n\nwhen doing augmentation on the images we have to apply the same augmentation to the mask too?"
        },
        {
          "id": 823772,
          "postDate": "2020-04-27T21:55:40.080Z",
          "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a>  you only need to apply geometric transformations to both (e.g. rotation, scale, shift, ...), other non geometric transformations (e.g. brightness, contrast, ..) only need to be applied to the image.</p>",
          "rawMarkdown": "@oscarrangel  you only need to apply geometric transformations to both (e.g. rotation, scale, shift, ...), other non geometric transformations (e.g. brightness, contrast, ..) only need to be applied to the image.",
          "votes": 1
        },
        {
          "id": 823812,
          "postDate": "2020-04-27T23:07:59.657Z",
          "content": "<p><a href=\"/ngcferreira\">@ngcferreira</a> thanks!</p>",
          "rawMarkdown": "@ngcferreira thanks!"
        }
      ]
    },
    {
      "id": 824691,
      "postDate": "2020-04-28T14:39:02.333Z",
      "content": "<p>Good work!</p>",
      "rawMarkdown": "Good work!"
    },
    {
      "id": 820851,
      "postDate": "2020-04-25T19:00:03.323Z",
      "content": "<p>Can you Please tell me what are the masks exactly? And when I try to open a mask using openslide why it shows a black canvas . Also why after adding cmap the pattern appears and how people got to know that this will work?\nI am trying to learn here</p>",
      "rawMarkdown": "Can you Please tell me what are the masks exactly? And when I try to open a mask using openslide why it shows a black canvas . Also why after adding cmap the pattern appears and how people got to know that this will work?\nI am trying to learn here",
      "replies": [
        {
          "id": 820886,
          "postDate": "2020-04-25T19:29:28.583Z",
          "content": "<p>The details and a list of the images are found in their <a href=\"https://www.kaggle.com/iamleonie/panda-eda-visualizations-suspicious-data\">EDA</a>.</p>\n\n<p>And for your question about viewing the masks: the masks are not image data like the WSIs. Instead of containing a range of values from 0 to 255, they only go up to a maximum of 6, representing the different class labels (check the dataset description for details on mask labels). Therefor when you try to visualize the mask, it will appear very dark as every value is close to 0. Applying the color map fixes the problem by assigning each label between 0 and 6 a distinct color.</p>",
          "rawMarkdown": "The details and a list of the images are found in their [EDA](https://www.kaggle.com/iamleonie/panda-eda-visualizations-suspicious-data).\n\nAnd for your question about viewing the masks: the masks are not image data like the WSIs. Instead of containing a range of values from 0 to 255, they only go up to a maximum of 6, representing the different class labels (check the dataset description for details on mask labels). Therefor when you try to visualize the mask, it will appear very dark as every value is close to 0. Applying the color map fixes the problem by assigning each label between 0 and 6 a distinct color.",
          "votes": 4
        },
        {
          "id": 820941,
          "postDate": "2020-04-25T19:59:17.803Z",
          "content": "<p>Thanks a Lot for the help, I now understand everything</p>",
          "rawMarkdown": "Thanks a Lot for the help, I now understand everything"
        },
        {
          "id": 867379,
          "postDate": "2020-05-30T08:24:20.213Z",
          "content": "<p>Thanks for this reply <a href=\"/matthewmasters\">@matthewmasters</a> </p>",
          "rawMarkdown": "Thanks for this reply @matthewmasters "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 818246,
      "author_name": "Ian Pan",
      "author_url": "",
      "post_date": "2020-04-23T18:09:23.710000",
      "content": "<p>From <a href=\"https://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset\">https://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset</a>:</p>\n\n<blockquote>\n  <p>The label masks of Radboudumc were semi-automatically generated by several deep learning algorithms, contain noise, and can be considered as weakly-supervised labels. The label masks of Karolinska were semi-autotomatically generated based on annotations by a pathologist.</p>\n</blockquote>\n\n<p>I noticed the same thing you did. The 85 masks that do not contain pixels labeled as cancerous even though ISUP grade &gt; 0 are from Radboudumc, so the DL model probably just missed them. These are probably tougher cases so might be a good set to test your model against harder cases. </p>\n\n<p>The Karolinska labels were generated based on pathologist annotations, so likely are more reliable. However, they are not as granular as the DL-generated labels.</p>\n\n<p>In any case, these seem to be weak labels, so we need to be careful about using them.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 818277,
          "author_name": "Leonie",
          "author_url": "",
          "post_date": "2020-04-23T18:40:43.293000",
          "content": "<p>Thank you for explanation. This helped a lot!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 820897,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-04-25T19:40:09.933000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 820899,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-04-25T19:42:30.477000",
          "content": "<p><a href=\"/iamleonie\">@iamleonie</a>,</p>\n\n<p>when doing augmentation on the images we have to apply the same augmentation to the mask too?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 823772,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-04-27T21:55:40.080000",
          "content": "<p><a href=\"/oscarrangel\">@oscarrangel</a>  you only need to apply geometric transformations to both (e.g. rotation, scale, shift, ...), other non geometric transformations (e.g. brightness, contrast, ..) only need to be applied to the image.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 823812,
          "author_name": "TheStoneMX",
          "author_url": "",
          "post_date": "2020-04-27T23:07:59.657000",
          "content": "<p><a href=\"/ngcferreira\">@ngcferreira</a> thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 824691,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-28T14:39:02.333000",
      "content": "<p>Good work!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820851,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2020-04-25T19:00:03.323000",
      "content": "<p>Can you Please tell me what are the masks exactly? And when I try to open a mask using openslide why it shows a black canvas . Also why after adding cmap the pattern appears and how people got to know that this will work?\nI am trying to learn here</p>",
      "votes": 0,
      "replies": [
        {
          "id": 820886,
          "author_name": "Matt",
          "author_url": "",
          "post_date": "2020-04-25T19:29:28.583000",
          "content": "<p>The details and a list of the images are found in their <a href=\"https://www.kaggle.com/iamleonie/panda-eda-visualizations-suspicious-data\">EDA</a>.</p>\n\n<p>And for your question about viewing the masks: the masks are not image data like the WSIs. Instead of containing a range of values from 0 to 255, they only go up to a maximum of 6, representing the different class labels (check the dataset description for details on mask labels). Therefor when you try to visualize the mask, it will appear very dark as every value is close to 0. Applying the color map fixes the problem by assigning each label between 0 and 6 a distinct color.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 820941,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2020-04-25T19:59:17.803000",
          "content": "<p>Thanks a Lot for the help, I now understand everything</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 867379,
          "author_name": "Kurian Benoy",
          "author_url": "",
          "post_date": "2020-05-30T08:24:20.213000",
          "content": "<p>Thanks for this reply <a href=\"/matthewmasters\">@matthewmasters</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "818226": "Many of you already know that there are 100 images without a mask. \nDuring my EDA I found two more suspicious things:\n* There are 4 masks that only contain background values\n* There are 85 masks that do not contain pixels marked as cancerous although the image is labeled with an ISUP grade &gt; 0\n\nIt might be a good idea to avoid using these test cases for training you model but instead use them for validation. I am not medically trained but for know I am assuming the `isup_grade` is still correct for the concering `image_id`s.\n\nDetails and a .csv containing all suspicious `image_id`s can be found in [my EDA](https://www.kaggle.com/iamleonie/panda-eda-visualizations-suspicious-data).",
    "818246": "From https://www.kaggle.com/wouterbulten/getting-started-with-the-panda-dataset:\n\n&gt; The label masks of Radboudumc were semi-automatically generated by several deep learning algorithms, contain noise, and can be considered as weakly-supervised labels. The label masks of Karolinska were semi-autotomatically generated based on annotations by a pathologist.\n\nI noticed the same thing you did. The 85 masks that do not contain pixels labeled as cancerous even though ISUP grade &gt; 0 are from Radboudumc, so the DL model probably just missed them. These are probably tougher cases so might be a good set to test your model against harder cases. \n\nThe Karolinska labels were generated based on pathologist annotations, so likely are more reliable. However, they are not as granular as the DL-generated labels.\n\nIn any case, these seem to be weak labels, so we need to be careful about using them.",
    "824691": "Good work!",
    "820851": "Can you Please tell me what are the masks exactly? And when I try to open a mask using openslide why it shows a black canvas . Also why after adding cmap the pattern appears and how people got to know that this will work?\nI am trying to learn here"
  }
}