{
  "id": 145742,
  "title": "Ideas on how to standardize dataset across data providers",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/145742",
  "author_name": "Matt",
  "post_date": "2020-04-24T10:42:11.518000",
  "votes": 18,
  "comment_count": 4,
  "views": 0,
  "content": "<p>There's several major differences between the datasets coming from each of the two centers.</p>\n\n<p>Here's a list of my observances so far:\n1. <strong>Segmentation masks</strong>\n    a. <strong>Labels</strong> - As described in the <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/data\">data description</a>, each center has a different method of labeling the masks. Radboud has 6 more descriptive labels while Karolinska only has 3. Therefor, in order to standardize the mask values, you must either simplify from 6 labels to 3, or you must extrapolate from the known Gleason score to go from 3 to 6. For example, if a slide from Karolinska is classified as 3+3, you can assume (though not always right), that the mask region classified as 'cancerous' can be described in terms of the 6 classes as a 'Gleason 3'.\n    b. <strong>Method</strong> - In addition to the difference in labeling, there is also a clear difference to the Method of each data providers' masks. While Karolinska labels broad regions, Radboud labels individual prostate glands. Below are two masks from each center illustrating the difference:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F846065%2F7cdc0ca71cbaa116852bada67ae2934d%2Fexample.png?generation=1587722849050621&amp;alt=media\" alt=\"\">\n    The only idea I have to harmonize these is to radially expand or blur the Radboud labels so it contains more of the surrounding region. I'd be interested to hear what kind of ideas you all have, I'm sure there are better ways.\n2. <strong>Equipment/Methodology</strong> - The equipment used at each center is different and from visual inspection of the slides, there are some noticeable differences. For example, the Karolinska images usually appear much more vibrant than Radboud. This can be standardized by simply normalizing the colors of all images. I'm sure there are many other ways the different equipment and methods have an influence on the slide's appearance.\n3. <strong>Pen marks?</strong> - The organizers noted that the pathologist at Karolinska used pen on some of the training slides (however not in the test). I haven't seen any examples of pen marks so I'm not sure about this one. Perhaps if the color strongly differs from the H&amp;E stain, it can be removed easily with some pre-processing.</p>\n\n<p>These are just my initial thoughts as about the problem, I'm sure there are many more things to be noted and ways these approaches could be improved. I'd love to hear your thoughts and feedback.</p>",
  "messages": [
    {
      "id": 819076,
      "postDate": "2020-04-24T10:42:11.517Z",
      "content": "<p>There's several major differences between the datasets coming from each of the two centers.</p>\n\n<p>Here's a list of my observances so far:\n1. <strong>Segmentation masks</strong>\n    a. <strong>Labels</strong> - As described in the <a href=\"https://www.kaggle.com/c/prostate-cancer-grade-assessment/data\">data description</a>, each center has a different method of labeling the masks. Radboud has 6 more descriptive labels while Karolinska only has 3. Therefor, in order to standardize the mask values, you must either simplify from 6 labels to 3, or you must extrapolate from the known Gleason score to go from 3 to 6. For example, if a slide from Karolinska is classified as 3+3, you can assume (though not always right), that the mask region classified as 'cancerous' can be described in terms of the 6 classes as a 'Gleason 3'.\n    b. <strong>Method</strong> - In addition to the difference in labeling, there is also a clear difference to the Method of each data providers' masks. While Karolinska labels broad regions, Radboud labels individual prostate glands. Below are two masks from each center illustrating the difference:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F846065%2F7cdc0ca71cbaa116852bada67ae2934d%2Fexample.png?generation=1587722849050621&amp;alt=media\" alt=\"\">\n    The only idea I have to harmonize these is to radially expand or blur the Radboud labels so it contains more of the surrounding region. I'd be interested to hear what kind of ideas you all have, I'm sure there are better ways.\n2. <strong>Equipment/Methodology</strong> - The equipment used at each center is different and from visual inspection of the slides, there are some noticeable differences. For example, the Karolinska images usually appear much more vibrant than Radboud. This can be standardized by simply normalizing the colors of all images. I'm sure there are many other ways the different equipment and methods have an influence on the slide's appearance.\n3. <strong>Pen marks?</strong> - The organizers noted that the pathologist at Karolinska used pen on some of the training slides (however not in the test). I haven't seen any examples of pen marks so I'm not sure about this one. Perhaps if the color strongly differs from the H&amp;E stain, it can be removed easily with some pre-processing.</p>\n\n<p>These are just my initial thoughts as about the problem, I'm sure there are many more things to be noted and ways these approaches could be improved. I'd love to hear your thoughts and feedback.</p>",
      "rawMarkdown": "There's several major differences between the datasets coming from each of the two centers.\n\nHere's a list of my observances so far:\n1. **Segmentation masks**\n    a. **Labels** - As described in the [data description](https://www.kaggle.com/c/prostate-cancer-grade-assessment/data), each center has a different method of labeling the masks. Radboud has 6 more descriptive labels while Karolinska only has 3. Therefor, in order to standardize the mask values, you must either simplify from 6 labels to 3, or you must extrapolate from the known Gleason score to go from 3 to 6. For example, if a slide from Karolinska is classified as 3+3, you can assume (though not always right), that the mask region classified as 'cancerous' can be described in terms of the 6 classes as a 'Gleason 3'.\n    b. **Method** - In addition to the difference in labeling, there is also a clear difference to the Method of each data providers' masks. While Karolinska labels broad regions, Radboud labels individual prostate glands. Below are two masks from each center illustrating the difference:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F846065%2F7cdc0ca71cbaa116852bada67ae2934d%2Fexample.png?generation=1587722849050621&amp;alt=media)\n    The only idea I have to harmonize these is to radially expand or blur the Radboud labels so it contains more of the surrounding region. I'd be interested to hear what kind of ideas you all have, I'm sure there are better ways.\n2. **Equipment/Methodology** - The equipment used at each center is different and from visual inspection of the slides, there are some noticeable differences. For example, the Karolinska images usually appear much more vibrant than Radboud. This can be standardized by simply normalizing the colors of all images. I'm sure there are many other ways the different equipment and methods have an influence on the slide's appearance.\n3. **Pen marks?** - The organizers noted that the pathologist at Karolinska used pen on some of the training slides (however not in the test). I haven't seen any examples of pen marks so I'm not sure about this one. Perhaps if the color strongly differs from the H&amp;E stain, it can be removed easily with some pre-processing.\n\nThese are just my initial thoughts as about the problem, I'm sure there are many more things to be noted and ways these approaches could be improved. I'd love to hear your thoughts and feedback.\n\n\n\n",
      "votes": 18
    },
    {
      "id": 961320,
      "postDate": "2020-08-07T05:10:27.827Z",
      "content": "<p>I recently found this competetion and am excited.   Here is my two cents:  1) Prostate cancers are mostly originated cells in glands and tend to remain in their original glandular formations. 2) When they become more malignant, they tend to invade the surrounding stromal tissue (through losing basal layers), meaning losing some degrees of their original glandular structures.  3) Radboud experts marked only glandular cells (or at least tried to do it).  4) Karolinska experts marked the  whole intact tissue that may contain non-glandular cells (i.e. stromal cells etc).  5) In highly malignant cancers, we may not be able to find any glandular structures.  I am getting some ideas on why the competetion organizers presented mask images in this mixed format so I will scan through comments right now.</p>",
      "rawMarkdown": "I recently found this competetion and am excited.   Here is my two cents:  1) Prostate cancers are mostly originated cells in glands and tend to remain in their original glandular formations. 2) When they become more malignant, they tend to invade the surrounding stromal tissue (through losing basal layers), meaning losing some degrees of their original glandular structures.  3) Radboud experts marked only glandular cells (or at least tried to do it).  4) Karolinska experts marked the  whole intact tissue that may contain non-glandular cells (i.e. stromal cells etc).  5) In highly malignant cancers, we may not be able to find any glandular structures.  I am getting some ideas on why the competetion organizers presented mask images in this mixed format so I will scan through comments right now."
    },
    {
      "id": 821164,
      "postDate": "2020-04-26T00:43:15.537Z",
      "content": "<p>\" simply normalizing the colors of all images\" it's much more complicated than that. This is what is referred as stain normalization or H&amp;E normalization. There are many techniques in the literature available for solving this kind of problem, including some deep learning-based methods.</p>",
      "rawMarkdown": "\" simply normalizing the colors of all images\" it's much more complicated than that. This is what is referred as stain normalization or H&amp;E normalization. There are many techniques in the literature available for solving this kind of problem, including some deep learning-based methods.",
      "replies": [
        {
          "id": 822410,
          "postDate": "2020-04-26T22:07:50.130Z",
          "content": "<p>Yes. In addition to that, I found this nice library to normalize stain <a href=\"https://github.com/Peter554/StainTools\">https://github.com/Peter554/StainTools</a></p>",
          "rawMarkdown": "Yes. In addition to that, I found this nice library to normalize stain https://github.com/Peter554/StainTools"
        }
      ]
    },
    {
      "id": 819085,
      "postDate": "2020-04-24T10:50:11.870Z",
      "content": "<p>thanks for sharing :)</p>",
      "rawMarkdown": "thanks for sharing :)"
    }
  ],
  "comments": [
    {
      "id": 961320,
      "author_name": "kimdesok",
      "author_url": "",
      "post_date": "2020-08-07T05:10:27.827000",
      "content": "<p>I recently found this competetion and am excited.   Here is my two cents:  1) Prostate cancers are mostly originated cells in glands and tend to remain in their original glandular formations. 2) When they become more malignant, they tend to invade the surrounding stromal tissue (through losing basal layers), meaning losing some degrees of their original glandular structures.  3) Radboud experts marked only glandular cells (or at least tried to do it).  4) Karolinska experts marked the  whole intact tissue that may contain non-glandular cells (i.e. stromal cells etc).  5) In highly malignant cancers, we may not be able to find any glandular structures.  I am getting some ideas on why the competetion organizers presented mask images in this mixed format so I will scan through comments right now.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 821164,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2020-04-26T00:43:15.537000",
      "content": "<p>\" simply normalizing the colors of all images\" it's much more complicated than that. This is what is referred as stain normalization or H&amp;E normalization. There are many techniques in the literature available for solving this kind of problem, including some deep learning-based methods.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 822410,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-26T22:07:50.130000",
          "content": "<p>Yes. In addition to that, I found this nice library to normalize stain <a href=\"https://github.com/Peter554/StainTools\">https://github.com/Peter554/StainTools</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 819085,
      "author_name": "Alberto Maria Falletta",
      "author_url": "",
      "post_date": "2020-04-24T10:50:11.870000",
      "content": "<p>thanks for sharing :)</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "819076": "There's several major differences between the datasets coming from each of the two centers.\n\nHere's a list of my observances so far:\n1. **Segmentation masks**\n    a. **Labels** - As described in the [data description](https://www.kaggle.com/c/prostate-cancer-grade-assessment/data), each center has a different method of labeling the masks. Radboud has 6 more descriptive labels while Karolinska only has 3. Therefor, in order to standardize the mask values, you must either simplify from 6 labels to 3, or you must extrapolate from the known Gleason score to go from 3 to 6. For example, if a slide from Karolinska is classified as 3+3, you can assume (though not always right), that the mask region classified as 'cancerous' can be described in terms of the 6 classes as a 'Gleason 3'.\n    b. **Method** - In addition to the difference in labeling, there is also a clear difference to the Method of each data providers' masks. While Karolinska labels broad regions, Radboud labels individual prostate glands. Below are two masks from each center illustrating the difference:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F846065%2F7cdc0ca71cbaa116852bada67ae2934d%2Fexample.png?generation=1587722849050621&amp;alt=media)\n    The only idea I have to harmonize these is to radially expand or blur the Radboud labels so it contains more of the surrounding region. I'd be interested to hear what kind of ideas you all have, I'm sure there are better ways.\n2. **Equipment/Methodology** - The equipment used at each center is different and from visual inspection of the slides, there are some noticeable differences. For example, the Karolinska images usually appear much more vibrant than Radboud. This can be standardized by simply normalizing the colors of all images. I'm sure there are many other ways the different equipment and methods have an influence on the slide's appearance.\n3. **Pen marks?** - The organizers noted that the pathologist at Karolinska used pen on some of the training slides (however not in the test). I haven't seen any examples of pen marks so I'm not sure about this one. Perhaps if the color strongly differs from the H&amp;E stain, it can be removed easily with some pre-processing.\n\nThese are just my initial thoughts as about the problem, I'm sure there are many more things to be noted and ways these approaches could be improved. I'd love to hear your thoughts and feedback.\n\n\n\n",
    "961320": "I recently found this competetion and am excited.   Here is my two cents:  1) Prostate cancers are mostly originated cells in glands and tend to remain in their original glandular formations. 2) When they become more malignant, they tend to invade the surrounding stromal tissue (through losing basal layers), meaning losing some degrees of their original glandular structures.  3) Radboud experts marked only glandular cells (or at least tried to do it).  4) Karolinska experts marked the  whole intact tissue that may contain non-glandular cells (i.e. stromal cells etc).  5) In highly malignant cancers, we may not be able to find any glandular structures.  I am getting some ideas on why the competetion organizers presented mask images in this mixed format so I will scan through comments right now.",
    "821164": "\" simply normalizing the colors of all images\" it's much more complicated than that. This is what is referred as stain normalization or H&amp;E normalization. There are many techniques in the literature available for solving this kind of problem, including some deep learning-based methods.",
    "819085": "thanks for sharing :)"
  }
}