{
  "id": 147378,
  "title": "Another institution in the hidden test set ?",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/147378",
  "author_name": "Benjamin Dubreu",
  "post_date": "2020-04-30T13:06:43.825000",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi ! </p>\n\n<p>should we expect data from another institution in the hidden part of the test set ? Right now I'd like to try normalizing the picture depending on which institution they come from. But i can only do that if the test images are from Radboud OR Karolinksa.</p>\n\n<p>Thanks for your time </p>",
  "messages": [
    {
      "id": 827677,
      "postDate": "2020-04-30T13:06:43.827Z",
      "content": "<p>Hi ! </p>\n\n<p>should we expect data from another institution in the hidden part of the test set ? Right now I'd like to try normalizing the picture depending on which institution they come from. But i can only do that if the test images are from Radboud OR Karolinksa.</p>\n\n<p>Thanks for your time </p>",
      "rawMarkdown": "Hi ! \n\nshould we expect data from another institution in the hidden part of the test set ? Right now I'd like to try normalizing the picture depending on which institution they come from. But i can only do that if the test images are from Radboud OR Karolinksa.\n\nThanks for your time \n",
      "votes": 4
    },
    {
      "id": 827805,
      "postDate": "2020-04-30T14:55:34.973Z",
      "content": "<p>Does it really matter? I guess the difference between the institutes is in the way the images are masked. Test images do not need to be masked right? So, How would it make a difference?</p>",
      "rawMarkdown": "Does it really matter? I guess the difference between the institutes is in the way the images are masked. Test images do not need to be masked right? So, How would it make a difference?",
      "votes": 1,
      "replies": [
        {
          "id": 827815,
          "postDate": "2020-04-30T15:00:42.430Z",
          "content": "<p>The hidden test set contains images from the same institutions as the rest of the data.</p>\n\n<p>The appearance of the images between (and within) institutions may vary due to different procedures, reagents and personnel in the lab, plus different imaging hardware. This is a key challenge in digital pathology. </p>",
          "rawMarkdown": "The hidden test set contains images from the same institutions as the rest of the data.\n\nThe appearance of the images between (and within) institutions may vary due to different procedures, reagents and personnel in the lab, plus different imaging hardware. This is a key challenge in digital pathology. ",
          "votes": 8
        },
        {
          "id": 827819,
          "postDate": "2020-04-30T15:03:27.410Z",
          "content": "<p>Thanks for enlightening me.</p>",
          "rawMarkdown": "Thanks for enlightening me."
        }
      ]
    },
    {
      "id": 828601,
      "postDate": "2020-05-01T07:16:40.733Z",
      "content": "<p><a href=\"/venky2506\">@venky2506</a> : I believe it matters. Since you can predict with close to 100% accuracy if a pic is from radboud or karolinska, you could use a model to detect that in the test set and then act consequently with the picture. But if there was a third institution involved, it wouldn't make sense. (note that I don't think It's the main idea I want to consider, but if I've got time left towards the end, I might give it a try)</p>\n\n<p>@Kimmo Kartasio: thanks for your answer ! In a future competition, you might want to add a third institution in the test set, to ensure that competitors solutions generalize well to totally new data ;) Thanks for hosting this comp' by the way !</p>",
      "rawMarkdown": "@venky2506 : I believe it matters. Since you can predict with close to 100% accuracy if a pic is from radboud or karolinska, you could use a model to detect that in the test set and then act consequently with the picture. But if there was a third institution involved, it wouldn't make sense. (note that I don't think It's the main idea I want to consider, but if I've got time left towards the end, I might give it a try)\n\n@Kimmo Kartasio: thanks for your answer ! In a future competition, you might want to add a third institution in the test set, to ensure that competitors solutions generalize well to totally new data ;) Thanks for hosting this comp' by the way !",
      "votes": 2,
      "replies": [
        {
          "id": 829334,
          "postDate": "2020-05-01T17:39:58.833Z",
          "content": "<p><a href=\"/bdubreu\">@bdubreu</a> you're not wrong that more data would be nice, but it's important to understand that acquiring the data for this competition (and most of the other competitions) involved a <em>ton</em> of work over several years by the host teams and medical professionals at Radboud and Karolinska. As the overview tab notes, this is the largest public whole-slide image dataset available. That's not because Radboud and Karolinska do more prostate biopsies than anywhere else, it's because publishing this kind of data is a very hard undertaking.</p>",
          "rawMarkdown": "@bdubreu you're not wrong that more data would be nice, but it's important to understand that acquiring the data for this competition (and most of the other competitions) involved a _ton_ of work over several years by the host teams and medical professionals at Radboud and Karolinska. As the overview tab notes, this is the largest public whole-slide image dataset available. That's not because Radboud and Karolinska do more prostate biopsies than anywhere else, it's because publishing this kind of data is a very hard undertaking.",
          "votes": 3
        },
        {
          "id": 829963,
          "postDate": "2020-05-02T07:52:29.213Z",
          "content": "<p>Oh yeah that's totally not a criticism, gathering all this data must have been quite the hassle ^^ I said that in your interest: the final goal of machine learning is taking the model to a new hospital and hope it will generalize there. Here, if there is no other institution, I could probably better my kaggle score by adapting my pipeline to the institution (since we have the info in the test df I believe, but even without it a model can predict with a huge degree of accuracy which institution a pic is from). However, such a solution wouldn't provide you guys with anything that would be useful in real world... Let's hope the final pipelines are not \"institution specific\" ;)</p>",
          "rawMarkdown": "Oh yeah that's totally not a criticism, gathering all this data must have been quite the hassle ^^ I said that in your interest: the final goal of machine learning is taking the model to a new hospital and hope it will generalize there. Here, if there is no other institution, I could probably better my kaggle score by adapting my pipeline to the institution (since we have the info in the test df I believe, but even without it a model can predict with a huge degree of accuracy which institution a pic is from). However, such a solution wouldn't provide you guys with anything that would be useful in real world... Let's hope the final pipelines are not \"institution specific\" ;)"
        }
      ]
    },
    {
      "id": 829826,
      "postDate": "2020-05-02T05:36:03.120Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 827805,
      "author_name": "Son of Anton v3.0",
      "author_url": "",
      "post_date": "2020-04-30T14:55:34.973000",
      "content": "<p>Does it really matter? I guess the difference between the institutes is in the way the images are masked. Test images do not need to be masked right? So, How would it make a difference?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 827815,
          "author_name": "Kimmo Kartasalo",
          "author_url": "",
          "post_date": "2020-04-30T15:00:42.430000",
          "content": "<p>The hidden test set contains images from the same institutions as the rest of the data.</p>\n\n<p>The appearance of the images between (and within) institutions may vary due to different procedures, reagents and personnel in the lab, plus different imaging hardware. This is a key challenge in digital pathology. </p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 827819,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-04-30T15:03:27.410000",
          "content": "<p>Thanks for enlightening me.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 828601,
      "author_name": "Benjamin Dubreu",
      "author_url": "",
      "post_date": "2020-05-01T07:16:40.733000",
      "content": "<p><a href=\"/venky2506\">@venky2506</a> : I believe it matters. Since you can predict with close to 100% accuracy if a pic is from radboud or karolinska, you could use a model to detect that in the test set and then act consequently with the picture. But if there was a third institution involved, it wouldn't make sense. (note that I don't think It's the main idea I want to consider, but if I've got time left towards the end, I might give it a try)</p>\n\n<p>@Kimmo Kartasio: thanks for your answer ! In a future competition, you might want to add a third institution in the test set, to ensure that competitors solutions generalize well to totally new data ;) Thanks for hosting this comp' by the way !</p>",
      "votes": 2,
      "replies": [
        {
          "id": 829334,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2020-05-01T17:39:58.833000",
          "content": "<p><a href=\"/bdubreu\">@bdubreu</a> you're not wrong that more data would be nice, but it's important to understand that acquiring the data for this competition (and most of the other competitions) involved a <em>ton</em> of work over several years by the host teams and medical professionals at Radboud and Karolinska. As the overview tab notes, this is the largest public whole-slide image dataset available. That's not because Radboud and Karolinska do more prostate biopsies than anywhere else, it's because publishing this kind of data is a very hard undertaking.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 829963,
          "author_name": "Benjamin Dubreu",
          "author_url": "",
          "post_date": "2020-05-02T07:52:29.213000",
          "content": "<p>Oh yeah that's totally not a criticism, gathering all this data must have been quite the hassle ^^ I said that in your interest: the final goal of machine learning is taking the model to a new hospital and hope it will generalize there. Here, if there is no other institution, I could probably better my kaggle score by adapting my pipeline to the institution (since we have the info in the test df I believe, but even without it a model can predict with a huge degree of accuracy which institution a pic is from). However, such a solution wouldn't provide you guys with anything that would be useful in real world... Let's hope the final pipelines are not \"institution specific\" ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 829826,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-02T05:36:03.120000",
      "content": "",
      "votes": -3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "827677": "Hi ! \n\nshould we expect data from another institution in the hidden part of the test set ? Right now I'd like to try normalizing the picture depending on which institution they come from. But i can only do that if the test images are from Radboud OR Karolinksa.\n\nThanks for your time \n",
    "827805": "Does it really matter? I guess the difference between the institutes is in the way the images are masked. Test images do not need to be masked right? So, How would it make a difference?",
    "828601": "@venky2506 : I believe it matters. Since you can predict with close to 100% accuracy if a pic is from radboud or karolinska, you could use a model to detect that in the test set and then act consequently with the picture. But if there was a third institution involved, it wouldn't make sense. (note that I don't think It's the main idea I want to consider, but if I've got time left towards the end, I might give it a try)\n\n@Kimmo Kartasio: thanks for your answer ! In a future competition, you might want to add a third institution in the test set, to ensure that competitors solutions generalize well to totally new data ;) Thanks for hosting this comp' by the way !",
    "829826": ""
  }
}