{
  "id": 356054,
  "title": "[Important Resource for competition] Normalized Tiled Dataset",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/356054",
  "author_name": "Mrinal Tyagi",
  "post_date": "2022-09-29T06:09:30.238000",
  "votes": 6,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Hello everyone, do check out my new dataset regarding MAYO Clinic - STRIP AI competition.<br>\nWould love to encourage everyone to use the dataset for solving the problem if they find it useful. </p>\n<p>Dataset Link -&gt; <a href=\"https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic-strip-ai-normalized-dataset\" target=\"_blank\">Link</a></p>\n<p>Do upvote if the dataset helped you in any form.</p>",
  "messages": [
    {
      "id": 1961372,
      "postDate": "2022-09-29T06:09:30.237Z",
      "content": "<p>Hello everyone, do check out my new dataset regarding MAYO Clinic - STRIP AI competition.<br>\nWould love to encourage everyone to use the dataset for solving the problem if they find it useful. </p>\n<p>Dataset Link -&gt; <a href=\"https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic-strip-ai-normalized-dataset\" target=\"_blank\">Link</a></p>\n<p>Do upvote if the dataset helped you in any form.</p>",
      "rawMarkdown": "Hello everyone, do check out my new dataset regarding MAYO Clinic - STRIP AI competition.\nWould love to encourage everyone to use the dataset for solving the problem if they find it useful. \n\nDataset Link -> [Link](https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic-strip-ai-normalized-dataset)\n\nDo upvote if the dataset helped you in any form.",
      "votes": 6
    },
    {
      "id": 1965598,
      "postDate": "2022-10-01T12:50:20.020Z",
      "content": "<p>I use your first dataset to train a neural network, then extract some tile using threshold methods and found good results in validation sets ( weighted log loss of 0.51), but submission gave me 0.8. Have you an idea about difference between the public, private set and the train set at the high resolution? A coloring problem?</p>",
      "rawMarkdown": "I use your first dataset to train a neural network, then extract some tile using threshold methods and found good results in validation sets ( weighted log loss of 0.51), but submission gave me 0.8. Have you an idea about difference between the public, private set and the train set at the high resolution? A coloring problem?",
      "votes": 1,
      "replies": [
        {
          "id": 1965765,
          "postDate": "2022-10-01T14:17:49.823Z",
          "content": "<p>I feel like the problem is in the model as i have come to infer from my experiments as well. The point is due to class imbalance, the model that is trained might be biased towards a specific class, and then during evaluation, they are using a weighted loss, and hence these biases might be affected. According to my opinion, the task is to make an unbiased model which can be given for evaluation to get peak performance when using weighted loss. Still, I am not an expert, would love to know others' opinions. This was my take on this as I have encountered the same problem. </p>",
          "rawMarkdown": "I feel like the problem is in the model as i have come to infer from my experiments as well. The point is due to class imbalance, the model that is trained might be biased towards a specific class, and then during evaluation, they are using a weighted loss, and hence these biases might be affected. According to my opinion, the task is to make an unbiased model which can be given for evaluation to get peak performance when using weighted loss. Still, I am not an expert, would love to know others' opinions. This was my take on this as I have encountered the same problem. "
        }
      ]
    },
    {
      "id": 1962705,
      "postDate": "2022-09-29T22:15:24.537Z",
      "content": "<p>Hi, are the tiles of this dataset for only one image?.</p>\n<p>Thanks a lot!</p>",
      "rawMarkdown": "Hi, are the tiles of this dataset for only one image?.\n\nThanks a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 1962886,
          "postDate": "2022-09-30T03:27:51.270Z",
          "content": "<p>No it is for the complete training set. </p>",
          "rawMarkdown": "No it is for the complete training set. "
        },
        {
          "id": 1963928,
          "postDate": "2022-09-30T14:06:08.020Z",
          "content": "<p>Hi, how do I do to map each image? the image files start with the same id.</p>\n<p>Thanks!</p>",
          "rawMarkdown": "Hi, how do I do to map each image? the image files start with the same id.\n\nThanks!"
        },
        {
          "id": 1963959,
          "postDate": "2022-09-30T14:17:38.620Z",
          "content": "<p>Please refer this notebook for tiling. <a href=\"https://www.kaggle.com/code/tr1gg3rtrash/mayo-clinic-best-preprocessing-notebook\" target=\"_blank\">https://www.kaggle.com/code/tr1gg3rtrash/mayo-clinic-best-preprocessing-notebook</a></p>\n<p>Yes the image file starts with same id but has row level and col information as well. </p>",
          "rawMarkdown": "Please refer this notebook for tiling. https://www.kaggle.com/code/tr1gg3rtrash/mayo-clinic-best-preprocessing-notebook\n\nYes the image file starts with same id but has row level and col information as well. "
        },
        {
          "id": 1964123,
          "postDate": "2022-09-30T15:27:52.970Z",
          "content": "<p>I couldn't find this image: train/CE/4ded24_0*.png</p>",
          "rawMarkdown": "I couldn't find this image: train/CE/4ded24_0*.png"
        }
      ]
    },
    {
      "id": 1961844,
      "postDate": "2022-09-29T11:40:04.747Z",
      "content": "<p>thanks. I am one of the users of the other dataset, which I found useful to avoid downloading all th GBs of the original slides :) . Thanks for it and also for the effort put in this one.<br>\nI was interested in this one, however I see the yellow color is lost. From the Mayo paper, the staining used for these slides should be MSB, which is a trichrome staining . The Macenko paper tells that \"When three or more stains are present in a slide, results are sometimes inconsistent\". Normalized images in fact look as H&amp;E… yellow areas are red blood cells, which in this case are likely important. </p>",
      "rawMarkdown": "thanks. I am one of the users of the other dataset, which I found useful to avoid downloading all th GBs of the original slides :) . Thanks for it and also for the effort put in this one.\nI was interested in this one, however I see the yellow color is lost. From the Mayo paper, the staining used for these slides should be MSB, which is a trichrome staining . The Macenko paper tells that \"When three or more stains are present in a slide, results are sometimes inconsistent\". Normalized images in fact look as H&E... yellow areas are red blood cells, which in this case are likely important. ",
      "votes": 2,
      "replies": [
        {
          "id": 1961871,
          "postDate": "2022-09-29T11:51:12.837Z",
          "content": "<p>Thank you for the great feedback. Really helps in keeping the motivation high. I did not consider this point while creating the dataset. But will surely try to work in this direction as well. In case you have any specific type of technique that you would want this dataset to be on, do let me know. Would love to help out in any way possible. And really appreciated the feedback. Always good to get some to grow. </p>",
          "rawMarkdown": "Thank you for the great feedback. Really helps in keeping the motivation high. I did not consider this point while creating the dataset. But will surely try to work in this direction as well. In case you have any specific type of technique that you would want this dataset to be on, do let me know. Would love to help out in any way possible. And really appreciated the feedback. Always good to get some to grow. ",
          "votes": 1
        },
        {
          "id": 1961984,
          "postDate": "2022-09-29T13:06:38.737Z",
          "content": "<p>It is worth trying to use this dataset. Sometimes 3 stains give results inconsistent, maybe this true for human and false for neural networks.</p>",
          "rawMarkdown": "It is worth trying to use this dataset. Sometimes 3 stains give results inconsistent, maybe this true for human and false for neural networks.",
          "votes": -1
        },
        {
          "id": 1961992,
          "postDate": "2022-09-29T13:09:05.133Z",
          "content": "<p>This dataset is balanced.</p>",
          "rawMarkdown": "This dataset is balanced.",
          "votes": -1
        },
        {
          "id": 1962022,
          "postDate": "2022-09-29T13:22:05.487Z",
          "content": "<p>I entered late in this challenge, so I do not have many bullets to shoot ☹️ (and also spent the first days only to be able to submit - another annoyance). The underlying issue of this challenge is that there is an interesting scientific problem to solve, but there are also resource constraints surrounding it that make it difficult to find the best solution. If you have to tile the slides and also normalize (both sensible options), likely 9 hours with 2 CPUs are not sufficient, even if normalization, done the right way, would be the way to go. <br>\nSaid that, what I was going to do, and eventually will do if I see some signal in my experiments, is at least to do white balancing. <br>\nThese slides are of very variable quality; in a regular digital pathology workflow a number of them would have been discarded because of bad quality, sometimes due to badly set scanner. The most visible issue is white balancing, that a well set scanner does by itself (and since from their paper it seems all was acquired with the same scanner or at least the same model, there should not be so much variablity). Here we have pink backgrounds, yellow backgrounds, green backgrounds, (Not sure if these artifacts come from conversion to TIFF), white of course, … and some very dirty slides. Forgetting the latter, I checked whether the top left or the bottom right tiles were always background, and it seems so. Basing on that, I would white-balance the rest of the tiles. This could be adequate for most but not all the slides, because some of them show a sort of dually colored background, which would make balancing impossible. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F901dbda00f5f6f04d4b5eac98bd5258b%2F6baf51_0.tif-150-150.jpg?generation=1664457416896428&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F1d8f0cef20e546a588be7331e1a50923%2Fa59c0d_0.tif-150-150.jpg?generation=1664457450042023&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F5238b382e791db5e4524b44699c5ee62%2F5bfaf8_0.tif-150-150.jpg?generation=1664457534986137&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2Fe4a70d73321d1246e41261fa2190695a%2F4094c4_0.tif-150-150.jpg?generation=1664457658540348&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "I entered late in this challenge, so I do not have many bullets to shoot ☹️ (and also spent the first days only to be able to submit - another annoyance). The underlying issue of this challenge is that there is an interesting scientific problem to solve, but there are also resource constraints surrounding it that make it difficult to find the best solution. If you have to tile the slides and also normalize (both sensible options), likely 9 hours with 2 CPUs are not sufficient, even if normalization, done the right way, would be the way to go. \nSaid that, what I was going to do, and eventually will do if I see some signal in my experiments, is at least to do white balancing. \nThese slides are of very variable quality; in a regular digital pathology workflow a number of them would have been discarded because of bad quality, sometimes due to badly set scanner. The most visible issue is white balancing, that a well set scanner does by itself (and since from their paper it seems all was acquired with the same scanner or at least the same model, there should not be so much variablity). Here we have pink backgrounds, yellow backgrounds, green backgrounds, (Not sure if these artifacts come from conversion to TIFF), white of course, ... and some very dirty slides. Forgetting the latter, I checked whether the top left or the bottom right tiles were always background, and it seems so. Basing on that, I would white-balance the rest of the tiles. This could be adequate for most but not all the slides, because some of them show a sort of dually colored background, which would make balancing impossible. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F901dbda00f5f6f04d4b5eac98bd5258b%2F6baf51_0.tif-150-150.jpg?generation=1664457416896428&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F1d8f0cef20e546a588be7331e1a50923%2Fa59c0d_0.tif-150-150.jpg?generation=1664457450042023&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F5238b382e791db5e4524b44699c5ee62%2F5bfaf8_0.tif-150-150.jpg?generation=1664457534986137&alt=media) ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2Fe4a70d73321d1246e41261fa2190695a%2F4094c4_0.tif-150-150.jpg?generation=1664457658540348&alt=media)",
          "votes": 1
        },
        {
          "id": 1962097,
          "postDate": "2022-09-29T13:47:09.500Z",
          "content": "<p><a href=\"https://www.kaggle.com/vdellamea\" target=\"_blank\">@vdellamea</a> I completely agree with you and I have sort of tried doing the same. I have locally created a dataset that first tiles the dataset. Then I perform h and e normalization on the dataset. Due to that most of the white area images are removed as in normalization formulae it is not able to find the eigen vectors. Then I perform otsu binarization and find an area in the image. If that area is above a specific threshold say 20 percent (as an area less than that would I think be useless), I remove that image. Also, another task that I did was separate out all these types of images that you have shared in the above post manually in the whole dataset (talking about the tiled one), and then calculate the statistical properties of the discarded images and hence use it as another way to separate out good and plain images. The outcome of this technique is surprisingly great and can share the dataset if you wanna have a look. Hopefully, I was able to understand some of the concepts you talked about above. </p>",
          "rawMarkdown": "@vdellamea I completely agree with you and I have sort of tried doing the same. I have locally created a dataset that first tiles the dataset. Then I perform h and e normalization on the dataset. Due to that most of the white area images are removed as in normalization formulae it is not able to find the eigen vectors. Then I perform otsu binarization and find an area in the image. If that area is above a specific threshold say 20 percent (as an area less than that would I think be useless), I remove that image. Also, another task that I did was separate out all these types of images that you have shared in the above post manually in the whole dataset (talking about the tiled one), and then calculate the statistical properties of the discarded images and hence use it as another way to separate out good and plain images. The outcome of this technique is surprisingly great and can share the dataset if you wanna have a look. Hopefully, I was able to understand some of the concepts you talked about above. ",
          "votes": 1
        },
        {
          "id": 1962107,
          "postDate": "2022-09-29T13:50:37.933Z",
          "content": "<p>Imbalance was surely one of the major issue of the main dataset and idts there is time left in the competition to calculate class weights based on statistical properties of the images which I performed in one of my research works and worked out great and that was too in medical imaging. Then ultimately have to resort to using conventional class weights by wcce which would ultimately give the majority class very less weightage as the imbalance is in ratio 1: 2 between classes. </p>",
          "rawMarkdown": "Imbalance was surely one of the major issue of the main dataset and idts there is time left in the competition to calculate class weights based on statistical properties of the images which I performed in one of my research works and worked out great and that was too in medical imaging. Then ultimately have to resort to using conventional class weights by wcce which would ultimately give the majority class very less weightage as the imbalance is in ratio 1: 2 between classes. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 1961603,
      "postDate": "2022-09-29T09:01:14.807Z",
      "content": "<p>May you share what normalization was used?</p>",
      "rawMarkdown": "May you share what normalization was used?",
      "votes": 2,
      "replies": [
        {
          "id": 1961623,
          "postDate": "2022-09-29T09:18:55.817Z",
          "content": "<p>Please refer this tutorial for the complete information. <a href=\"https://www.youtube.com/watch?v=yUrwEYgZUsA\" target=\"_blank\">https://www.youtube.com/watch?v=yUrwEYgZUsA</a></p>\n<p>The paper for the following technique is available here: <a href=\"http://wwwx.cs.unc.edu/~mn/sites/default/files/macenko2009.pdf\" target=\"_blank\">http://wwwx.cs.unc.edu/~mn/sites/default/files/macenko2009.pdf</a></p>",
          "rawMarkdown": "Please refer this tutorial for the complete information. https://www.youtube.com/watch?v=yUrwEYgZUsA\n\nThe paper for the following technique is available here: http://wwwx.cs.unc.edu/~mn/sites/default/files/macenko2009.pdf",
          "votes": 2
        }
      ]
    },
    {
      "id": 1969108,
      "postDate": "2022-10-03T10:41:29.053Z",
      "content": "<p>This picture does not seem to be a direct screenshot of the original picture, what kind of processing have they experienced?</p>",
      "rawMarkdown": "This picture does not seem to be a direct screenshot of the original picture, what kind of processing have they experienced?"
    },
    {
      "id": 1967396,
      "postDate": "2022-10-02T13:31:36.373Z",
      "content": "<p>Do checkout the inferencing pipeline for the dataset: <a href=\"https://www.kaggle.com/tr1gg3rtrash/mayo-clinic-tiled-inference\" target=\"_blank\">https://www.kaggle.com/tr1gg3rtrash/mayo-clinic-tiled-inference</a></p>",
      "rawMarkdown": "Do checkout the inferencing pipeline for the dataset: https://www.kaggle.com/tr1gg3rtrash/mayo-clinic-tiled-inference",
      "replies": [
        {
          "id": 1967398,
          "postDate": "2022-10-02T13:32:09.810Z",
          "content": "<p><a href=\"https://www.kaggle.com/vdellamea\" target=\"_blank\">@vdellamea</a> </p>",
          "rawMarkdown": "@vdellamea "
        }
      ]
    },
    {
      "id": 1964643,
      "postDate": "2022-09-30T20:21:00.837Z",
      "content": "<p>Hi, the issue here is What happens when the submission is made?, I suppose that the final test will be around 280 images, which will be a complex task to generate slices</p>",
      "rawMarkdown": "Hi, the issue here is What happens when the submission is made?, I suppose that the final test will be around 280 images, which will be a complex task to generate slices",
      "replies": [
        {
          "id": 1964736,
          "postDate": "2022-09-30T21:12:12.500Z",
          "content": "<p>I have made a submission with tiling of images and it takes around 5-6 hours for submission. It depends on how memory efficient your testing pipeline. </p>",
          "rawMarkdown": "I have made a submission with tiling of images and it takes around 5-6 hours for submission. It depends on how memory efficient your testing pipeline. ",
          "votes": 1
        },
        {
          "id": 1964846,
          "postDate": "2022-09-30T23:36:11.327Z",
          "content": "<p>Did you see differences respect to accuracy or loss when you use this tiles?</p>\n<p>Thanks for your contribution and your comments!</p>",
          "rawMarkdown": "Did you see differences respect to accuracy or loss when you use this tiles?\n\nThanks for your contribution and your comments!"
        },
        {
          "id": 1964854,
          "postDate": "2022-09-30T23:52:11.403Z",
          "content": "<p>I tried binary cross entropy with resnet and was able to reach 0.49 loss ig. But I feel the competition is not just with a model with the lowest loss. </p>",
          "rawMarkdown": "I tried binary cross entropy with resnet and was able to reach 0.49 loss ig. But I feel the competition is not just with a model with the lowest loss. "
        },
        {
          "id": 1964902,
          "postDate": "2022-10-01T01:38:23.747Z",
          "content": "<p>In order to process this different number of tiles by image did you use a MIL model?</p>",
          "rawMarkdown": "In order to process this different number of tiles by image did you use a MIL model?"
        },
        {
          "id": 1964909,
          "postDate": "2022-10-01T01:59:51.020Z",
          "content": "<p>Basically, i used a deepzoomgenerator class of openslide library. It divides the images into slides, then performed a normalization function to obtain normalized image of the slide. </p>",
          "rawMarkdown": "Basically, i used a deepzoomgenerator class of openslide library. It divides the images into slides, then performed a normalization function to obtain normalized image of the slide. "
        },
        {
          "id": 1964917,
          "postDate": "2022-10-01T02:08:24.373Z",
          "content": "<p>Sorry I wanted to say in order to train </p>",
          "rawMarkdown": "Sorry I wanted to say in order to train "
        },
        {
          "id": 1964923,
          "postDate": "2022-10-01T02:19:35.513Z",
          "content": "<p>yupp same process is used for train as wekk. in case of test, tiles are extracted and then the mean prediction is extracted for an image using the tiles of the image. </p>",
          "rawMarkdown": "yupp same process is used for train as wekk. in case of test, tiles are extracted and then the mean prediction is extracted for an image using the tiles of the image. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1965598,
      "author_name": "Pierre Tisseur",
      "author_url": "",
      "post_date": "2022-10-01T12:50:20.020000",
      "content": "<p>I use your first dataset to train a neural network, then extract some tile using threshold methods and found good results in validation sets ( weighted log loss of 0.51), but submission gave me 0.8. Have you an idea about difference between the public, private set and the train set at the high resolution? A coloring problem?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1965765,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-10-01T14:17:49.823000",
          "content": "<p>I feel like the problem is in the model as i have come to infer from my experiments as well. The point is due to class imbalance, the model that is trained might be biased towards a specific class, and then during evaluation, they are using a weighted loss, and hence these biases might be affected. According to my opinion, the task is to make an unbiased model which can be given for evaluation to get peak performance when using weighted loss. Still, I am not an expert, would love to know others' opinions. This was my take on this as I have encountered the same problem. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1962705,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-09-29T22:15:24.537000",
      "content": "<p>Hi, are the tiles of this dataset for only one image?.</p>\n<p>Thanks a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1962886,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-30T03:27:51.270000",
          "content": "<p>No it is for the complete training set. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1963928,
          "author_name": "Pablo Larrosa",
          "author_url": "",
          "post_date": "2022-09-30T14:06:08.020000",
          "content": "<p>Hi, how do I do to map each image? the image files start with the same id.</p>\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1963959,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-30T14:17:38.620000",
          "content": "<p>Please refer this notebook for tiling. <a href=\"https://www.kaggle.com/code/tr1gg3rtrash/mayo-clinic-best-preprocessing-notebook\" target=\"_blank\">https://www.kaggle.com/code/tr1gg3rtrash/mayo-clinic-best-preprocessing-notebook</a></p>\n<p>Yes the image file starts with same id but has row level and col information as well. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1964123,
          "author_name": "Pablo Larrosa",
          "author_url": "",
          "post_date": "2022-09-30T15:27:52.970000",
          "content": "<p>I couldn't find this image: train/CE/4ded24_0*.png</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1961844,
      "author_name": "MITEL-UNIUD",
      "author_url": "",
      "post_date": "2022-09-29T11:40:04.747000",
      "content": "<p>thanks. I am one of the users of the other dataset, which I found useful to avoid downloading all th GBs of the original slides :) . Thanks for it and also for the effort put in this one.<br>\nI was interested in this one, however I see the yellow color is lost. From the Mayo paper, the staining used for these slides should be MSB, which is a trichrome staining . The Macenko paper tells that \"When three or more stains are present in a slide, results are sometimes inconsistent\". Normalized images in fact look as H&amp;E… yellow areas are red blood cells, which in this case are likely important. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1961871,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-29T11:51:12.837000",
          "content": "<p>Thank you for the great feedback. Really helps in keeping the motivation high. I did not consider this point while creating the dataset. But will surely try to work in this direction as well. In case you have any specific type of technique that you would want this dataset to be on, do let me know. Would love to help out in any way possible. And really appreciated the feedback. Always good to get some to grow. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1961984,
          "author_name": "Pierre Tisseur",
          "author_url": "",
          "post_date": "2022-09-29T13:06:38.737000",
          "content": "<p>It is worth trying to use this dataset. Sometimes 3 stains give results inconsistent, maybe this true for human and false for neural networks.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1961992,
          "author_name": "Pierre Tisseur",
          "author_url": "",
          "post_date": "2022-09-29T13:09:05.133000",
          "content": "<p>This dataset is balanced.</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 1962022,
          "author_name": "MITEL-UNIUD",
          "author_url": "",
          "post_date": "2022-09-29T13:22:05.487000",
          "content": "<p>I entered late in this challenge, so I do not have many bullets to shoot ☹️ (and also spent the first days only to be able to submit - another annoyance). The underlying issue of this challenge is that there is an interesting scientific problem to solve, but there are also resource constraints surrounding it that make it difficult to find the best solution. If you have to tile the slides and also normalize (both sensible options), likely 9 hours with 2 CPUs are not sufficient, even if normalization, done the right way, would be the way to go. <br>\nSaid that, what I was going to do, and eventually will do if I see some signal in my experiments, is at least to do white balancing. <br>\nThese slides are of very variable quality; in a regular digital pathology workflow a number of them would have been discarded because of bad quality, sometimes due to badly set scanner. The most visible issue is white balancing, that a well set scanner does by itself (and since from their paper it seems all was acquired with the same scanner or at least the same model, there should not be so much variablity). Here we have pink backgrounds, yellow backgrounds, green backgrounds, (Not sure if these artifacts come from conversion to TIFF), white of course, … and some very dirty slides. Forgetting the latter, I checked whether the top left or the bottom right tiles were always background, and it seems so. Basing on that, I would white-balance the rest of the tiles. This could be adequate for most but not all the slides, because some of them show a sort of dually colored background, which would make balancing impossible. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F901dbda00f5f6f04d4b5eac98bd5258b%2F6baf51_0.tif-150-150.jpg?generation=1664457416896428&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F1d8f0cef20e546a588be7331e1a50923%2Fa59c0d_0.tif-150-150.jpg?generation=1664457450042023&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2F5238b382e791db5e4524b44699c5ee62%2F5bfaf8_0.tif-150-150.jpg?generation=1664457534986137&amp;alt=media\" alt=\"\"> <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3526812%2Fe4a70d73321d1246e41261fa2190695a%2F4094c4_0.tif-150-150.jpg?generation=1664457658540348&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1962097,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-29T13:47:09.500000",
          "content": "<p><a href=\"https://www.kaggle.com/vdellamea\" target=\"_blank\">@vdellamea</a> I completely agree with you and I have sort of tried doing the same. I have locally created a dataset that first tiles the dataset. Then I perform h and e normalization on the dataset. Due to that most of the white area images are removed as in normalization formulae it is not able to find the eigen vectors. Then I perform otsu binarization and find an area in the image. If that area is above a specific threshold say 20 percent (as an area less than that would I think be useless), I remove that image. Also, another task that I did was separate out all these types of images that you have shared in the above post manually in the whole dataset (talking about the tiled one), and then calculate the statistical properties of the discarded images and hence use it as another way to separate out good and plain images. The outcome of this technique is surprisingly great and can share the dataset if you wanna have a look. Hopefully, I was able to understand some of the concepts you talked about above. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1962107,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-29T13:50:37.933000",
          "content": "<p>Imbalance was surely one of the major issue of the main dataset and idts there is time left in the competition to calculate class weights based on statistical properties of the images which I performed in one of my research works and worked out great and that was too in medical imaging. Then ultimately have to resort to using conventional class weights by wcce which would ultimately give the majority class very less weightage as the imbalance is in ratio 1: 2 between classes. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1961603,
      "author_name": "anthony",
      "author_url": "",
      "post_date": "2022-09-29T09:01:14.807000",
      "content": "<p>May you share what normalization was used?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1961623,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-29T09:18:55.817000",
          "content": "<p>Please refer this tutorial for the complete information. <a href=\"https://www.youtube.com/watch?v=yUrwEYgZUsA\" target=\"_blank\">https://www.youtube.com/watch?v=yUrwEYgZUsA</a></p>\n<p>The paper for the following technique is available here: <a href=\"http://wwwx.cs.unc.edu/~mn/sites/default/files/macenko2009.pdf\" target=\"_blank\">http://wwwx.cs.unc.edu/~mn/sites/default/files/macenko2009.pdf</a></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1969108,
      "author_name": "qiucen",
      "author_url": "",
      "post_date": "2022-10-03T10:41:29.053000",
      "content": "<p>This picture does not seem to be a direct screenshot of the original picture, what kind of processing have they experienced?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1967396,
      "author_name": "Mrinal Tyagi",
      "author_url": "",
      "post_date": "2022-10-02T13:31:36.373000",
      "content": "<p>Do checkout the inferencing pipeline for the dataset: <a href=\"https://www.kaggle.com/tr1gg3rtrash/mayo-clinic-tiled-inference\" target=\"_blank\">https://www.kaggle.com/tr1gg3rtrash/mayo-clinic-tiled-inference</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 1967398,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-10-02T13:32:09.810000",
          "content": "<p><a href=\"https://www.kaggle.com/vdellamea\" target=\"_blank\">@vdellamea</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1964643,
      "author_name": "Pablo Larrosa",
      "author_url": "",
      "post_date": "2022-09-30T20:21:00.837000",
      "content": "<p>Hi, the issue here is What happens when the submission is made?, I suppose that the final test will be around 280 images, which will be a complex task to generate slices</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1964736,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-30T21:12:12.500000",
          "content": "<p>I have made a submission with tiling of images and it takes around 5-6 hours for submission. It depends on how memory efficient your testing pipeline. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1964846,
          "author_name": "Pablo Larrosa",
          "author_url": "",
          "post_date": "2022-09-30T23:36:11.327000",
          "content": "<p>Did you see differences respect to accuracy or loss when you use this tiles?</p>\n<p>Thanks for your contribution and your comments!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1964854,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-09-30T23:52:11.403000",
          "content": "<p>I tried binary cross entropy with resnet and was able to reach 0.49 loss ig. But I feel the competition is not just with a model with the lowest loss. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1964902,
          "author_name": "Pablo Larrosa",
          "author_url": "",
          "post_date": "2022-10-01T01:38:23.747000",
          "content": "<p>In order to process this different number of tiles by image did you use a MIL model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1964909,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-10-01T01:59:51.020000",
          "content": "<p>Basically, i used a deepzoomgenerator class of openslide library. It divides the images into slides, then performed a normalization function to obtain normalized image of the slide. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1964917,
          "author_name": "Pablo Larrosa",
          "author_url": "",
          "post_date": "2022-10-01T02:08:24.373000",
          "content": "<p>Sorry I wanted to say in order to train </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1964923,
          "author_name": "Mrinal Tyagi",
          "author_url": "",
          "post_date": "2022-10-01T02:19:35.513000",
          "content": "<p>yupp same process is used for train as wekk. in case of test, tiles are extracted and then the mean prediction is extracted for an image using the tiles of the image. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1961372": "Hello everyone, do check out my new dataset regarding MAYO Clinic - STRIP AI competition.\nWould love to encourage everyone to use the dataset for solving the problem if they find it useful. \n\nDataset Link -> [Link](https://www.kaggle.com/datasets/tr1gg3rtrash/mayo-clinic-strip-ai-normalized-dataset)\n\nDo upvote if the dataset helped you in any form.",
    "1965598": "I use your first dataset to train a neural network, then extract some tile using threshold methods and found good results in validation sets ( weighted log loss of 0.51), but submission gave me 0.8. Have you an idea about difference between the public, private set and the train set at the high resolution? A coloring problem?",
    "1962705": "Hi, are the tiles of this dataset for only one image?.\n\nThanks a lot!",
    "1961844": "thanks. I am one of the users of the other dataset, which I found useful to avoid downloading all th GBs of the original slides :) . Thanks for it and also for the effort put in this one.\nI was interested in this one, however I see the yellow color is lost. From the Mayo paper, the staining used for these slides should be MSB, which is a trichrome staining . The Macenko paper tells that \"When three or more stains are present in a slide, results are sometimes inconsistent\". Normalized images in fact look as H&E... yellow areas are red blood cells, which in this case are likely important. ",
    "1961603": "May you share what normalization was used?",
    "1969108": "This picture does not seem to be a direct screenshot of the original picture, what kind of processing have they experienced?",
    "1967396": "Do checkout the inferencing pipeline for the dataset: https://www.kaggle.com/tr1gg3rtrash/mayo-clinic-tiled-inference",
    "1964643": "Hi, the issue here is What happens when the submission is made?, I suppose that the final test will be around 280 images, which will be a complex task to generate slices"
  }
}