{
  "id": 335755,
  "title": "Mayo Clinic Sliced 1024x1024 JPG Datasets",
  "url": "/competitions/mayo-clinic-strip-ai/discussion/335755",
  "author_name": "Rob Mulla",
  "post_date": "2022-07-07T17:50:49.362000",
  "votes": 47,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I adapted some code that <a href=\"https://www.kaggle.com/simsonimus\" target=\"_blank\">@simsonimus</a> created for last years <a href=\"https://www.kaggle.com/code/simsonimus/hubmap-training\" target=\"_blank\">HuBMap competition</a> to slice the tif files into more manageable 1024x1024 jpg images.</p>\n<p>The notebook to create the datasets is here: <a href=\"https://www.kaggle.com/code/robikscube/mayo-clinic-image-dataset-1024-jpg\" target=\"_blank\">https://www.kaggle.com/code/robikscube/mayo-clinic-image-dataset-1024-jpg</a></p>\n<p>Since this dataset is so large I had to run it in chunks based on the filesizes or else the notebooks would fail.  I'll update the links here as they complete. Hopefully some will find this useful to make the data more managable.</p>\n<ul>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part1\" target=\"_blank\">Part 1</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part2-1\" target=\"_blank\">Part 2</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part3\" target=\"_blank\">Part 3</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part4-1\" target=\"_blank\">Part 4</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part5-1\" target=\"_blank\">Part 5</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part6\" target=\"_blank\">Part 6</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part7-1\" target=\"_blank\">Part 7</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part8\" target=\"_blank\">Part 8</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part9\" target=\"_blank\">Part 9</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part10\" target=\"_blank\">Part 10</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024x1024-sliced-jpg-test-set\" target=\"_blank\">Mayo Clinic 1024x1024 - Test Set</a></p></li>\n</ul>",
  "messages": [
    {
      "id": 1847210,
      "postDate": "2022-07-07T17:50:49.363Z",
      "content": "<p>I adapted some code that <a href=\"https://www.kaggle.com/simsonimus\" target=\"_blank\">@simsonimus</a> created for last years <a href=\"https://www.kaggle.com/code/simsonimus/hubmap-training\" target=\"_blank\">HuBMap competition</a> to slice the tif files into more manageable 1024x1024 jpg images.</p>\n<p>The notebook to create the datasets is here: <a href=\"https://www.kaggle.com/code/robikscube/mayo-clinic-image-dataset-1024-jpg\" target=\"_blank\">https://www.kaggle.com/code/robikscube/mayo-clinic-image-dataset-1024-jpg</a></p>\n<p>Since this dataset is so large I had to run it in chunks based on the filesizes or else the notebooks would fail.  I'll update the links here as they complete. Hopefully some will find this useful to make the data more managable.</p>\n<ul>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part1\" target=\"_blank\">Part 1</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part2-1\" target=\"_blank\">Part 2</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part3\" target=\"_blank\">Part 3</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part4-1\" target=\"_blank\">Part 4</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part5-1\" target=\"_blank\">Part 5</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part6\" target=\"_blank\">Part 6</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part7-1\" target=\"_blank\">Part 7</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part8\" target=\"_blank\">Part 8</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part9\" target=\"_blank\">Part 9</a></p></li>\n<li><p>Mayo Clinic 1024x1024 - <a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part10\" target=\"_blank\">Part 10</a></p></li>\n<li><p><a href=\"https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024x1024-sliced-jpg-test-set\" target=\"_blank\">Mayo Clinic 1024x1024 - Test Set</a></p></li>\n</ul>",
      "rawMarkdown": "I adapted some code that @simsonimus created for last years [HuBMap competition](https://www.kaggle.com/code/simsonimus/hubmap-training) to slice the tif files into more manageable 1024x1024 jpg images.\n\nThe notebook to create the datasets is here: https://www.kaggle.com/code/robikscube/mayo-clinic-image-dataset-1024-jpg\n\nSince this dataset is so large I had to run it in chunks based on the filesizes or else the notebooks would fail.  I'll update the links here as they complete. Hopefully some will find this useful to make the data more managable.\n\n- Mayo Clinic 1024x1024 - [Part 1](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part1)\n- Mayo Clinic 1024x1024 - [Part 2](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part2-1)\n- Mayo Clinic 1024x1024 - [Part 3](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part3)\n- Mayo Clinic 1024x1024 - [Part 4](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part4-1)\n- Mayo Clinic 1024x1024 - [Part 5](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part5-1)\n- Mayo Clinic 1024x1024 - [Part 6](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part6)\n- Mayo Clinic 1024x1024 - [Part 7](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part7-1)\n- Mayo Clinic 1024x1024 - [Part 8](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part8)\n- Mayo Clinic 1024x1024 - [Part 9](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part9)\n- Mayo Clinic 1024x1024 - [Part 10](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part10)\n\n- [Mayo Clinic 1024x1024 - Test Set](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024x1024-sliced-jpg-test-set)",
      "votes": 47
    },
    {
      "id": 1847497,
      "postDate": "2022-07-08T00:59:29.873Z",
      "content": "<p>I've been working with this dataset and the original TIF files are really hard to work with. This will be a great help!</p>",
      "rawMarkdown": "I've been working with this dataset and the original TIF files are really hard to work with. This will be a great help!\n",
      "votes": 3
    },
    {
      "id": 1849441,
      "postDate": "2022-07-09T14:15:05.153Z",
      "content": "<p>It seems like every group of your dataset misses the first slide.</p>",
      "rawMarkdown": "It seems like every group of your dataset misses the first slide.",
      "votes": 1
    },
    {
      "id": 1849414,
      "postDate": "2022-07-09T13:39:39.433Z",
      "content": "<p>I was having a hard time downloading the data, so thank you!</p>",
      "rawMarkdown": "I was having a hard time downloading the data, so thank you!",
      "votes": 1
    },
    {
      "id": 1847220,
      "postDate": "2022-07-07T18:09:00.523Z",
      "content": "<p>Thank you, now I understand why you sliced the data.<br>\nBut why did use mask on the sliced data knowing that this not a segmentation problem.</p>",
      "rawMarkdown": "Thank you, now I understand why you sliced the data.\nBut why did use mask on the sliced data knowing that this not a segmentation problem.",
      "votes": 1,
      "replies": [
        {
          "id": 1847241,
          "postDate": "2022-07-07T18:28:33.027Z",
          "content": "<p>The mask code is just caryover from the source notebook used in the previous competition. I did not use it in creating these jpg files.</p>",
          "rawMarkdown": "The mask code is just caryover from the source notebook used in the previous competition. I did not use it in creating these jpg files.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1847900,
      "postDate": "2022-07-08T08:33:49.383Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> , if we create slices of the actual image, what happens to the label? Unlike Hubmap we don't have pixel-wise label so there is a good chance some slices may not have the same label as the image. Won't it create problem?</p>",
      "rawMarkdown": "Hi @robikscube , if we create slices of the actual image, what happens to the label? Unlike Hubmap we don't have pixel-wise label so there is a good chance some slices may not have the same label as the image. Won't it create problem?",
      "votes": 2,
      "replies": [
        {
          "id": 1848367,
          "postDate": "2022-07-08T14:52:05.577Z",
          "content": "<p>That's a good point. I'm not 100% sure that these will be useful yet. But I'm assuming that we will need to somehow reduce the tif files to make them more managable. One thought would be to remove any of the blank slices and then train a model with examples from each slice with the given labels, and then average on predictions of the slices for the test set.</p>\n<p>If you have ideas for different data formatting for this competition let me know- I'm happy to create more datasets :D</p>",
          "rawMarkdown": "That's a good point. I'm not 100% sure that these will be useful yet. But I'm assuming that we will need to somehow reduce the tif files to make them more managable. One thought would be to remove any of the blank slices and then train a model with examples from each slice with the given labels, and then average on predictions of the slices for the test set.\n\nIf you have ideas for different data formatting for this competition let me know- I'm happy to create more datasets :D",
          "votes": 4
        }
      ]
    },
    {
      "id": 1956507,
      "postDate": "2022-09-26T13:37:24.290Z",
      "content": "<p>the test dataset has missing slices for<code>006388_0.tif</code> image</p>",
      "rawMarkdown": "the test dataset has missing slices for` 006388_0.tif` image"
    }
  ],
  "comments": [
    {
      "id": 1847497,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-07-08T00:59:29.873000",
      "content": "<p>I've been working with this dataset and the original TIF files are really hard to work with. This will be a great help!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1849441,
      "author_name": "Leon",
      "author_url": "",
      "post_date": "2022-07-09T14:15:05.153000",
      "content": "<p>It seems like every group of your dataset misses the first slide.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1849414,
      "author_name": "akahane",
      "author_url": "",
      "post_date": "2022-07-09T13:39:39.433000",
      "content": "<p>I was having a hard time downloading the data, so thank you!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1847220,
      "author_name": "Med Ali Bouchhioua",
      "author_url": "",
      "post_date": "2022-07-07T18:09:00.523000",
      "content": "<p>Thank you, now I understand why you sliced the data.<br>\nBut why did use mask on the sliced data knowing that this not a segmentation problem.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1847241,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2022-07-07T18:28:33.027000",
          "content": "<p>The mask code is just caryover from the source notebook used in the previous competition. I did not use it in creating these jpg files.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1847900,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2022-07-08T08:33:49.383000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/robikscube\" target=\"_blank\">@robikscube</a> , if we create slices of the actual image, what happens to the label? Unlike Hubmap we don't have pixel-wise label so there is a good chance some slices may not have the same label as the image. Won't it create problem?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1848367,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2022-07-08T14:52:05.577000",
          "content": "<p>That's a good point. I'm not 100% sure that these will be useful yet. But I'm assuming that we will need to somehow reduce the tif files to make them more managable. One thought would be to remove any of the blank slices and then train a model with examples from each slice with the given labels, and then average on predictions of the slices for the test set.</p>\n<p>If you have ideas for different data formatting for this competition let me know- I'm happy to create more datasets :D</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1956507,
      "author_name": "Antriksh jain",
      "author_url": "",
      "post_date": "2022-09-26T13:37:24.290000",
      "content": "<p>the test dataset has missing slices for<code>006388_0.tif</code> image</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1847210": "I adapted some code that @simsonimus created for last years [HuBMap competition](https://www.kaggle.com/code/simsonimus/hubmap-training) to slice the tif files into more manageable 1024x1024 jpg images.\n\nThe notebook to create the datasets is here: https://www.kaggle.com/code/robikscube/mayo-clinic-image-dataset-1024-jpg\n\nSince this dataset is so large I had to run it in chunks based on the filesizes or else the notebooks would fail.  I'll update the links here as they complete. Hopefully some will find this useful to make the data more managable.\n\n- Mayo Clinic 1024x1024 - [Part 1](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part1)\n- Mayo Clinic 1024x1024 - [Part 2](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part2-1)\n- Mayo Clinic 1024x1024 - [Part 3](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part3)\n- Mayo Clinic 1024x1024 - [Part 4](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part4-1)\n- Mayo Clinic 1024x1024 - [Part 5](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part5-1)\n- Mayo Clinic 1024x1024 - [Part 6](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part6)\n- Mayo Clinic 1024x1024 - [Part 7](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part7-1)\n- Mayo Clinic 1024x1024 - [Part 8](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part8)\n- Mayo Clinic 1024x1024 - [Part 9](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part9)\n- Mayo Clinic 1024x1024 - [Part 10](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024-jpg-part10)\n\n- [Mayo Clinic 1024x1024 - Test Set](https://www.kaggle.com/datasets/robikscube/mayo-clinic-1024x1024-sliced-jpg-test-set)",
    "1847497": "I've been working with this dataset and the original TIF files are really hard to work with. This will be a great help!\n",
    "1849441": "It seems like every group of your dataset misses the first slide.",
    "1849414": "I was having a hard time downloading the data, so thank you!",
    "1847220": "Thank you, now I understand why you sliced the data.\nBut why did use mask on the sliced data knowing that this not a segmentation problem.",
    "1847900": "Hi @robikscube , if we create slices of the actual image, what happens to the label? Unlike Hubmap we don't have pixel-wise label so there is a good chance some slices may not have the same label as the image. Won't it create problem?",
    "1956507": "the test dataset has missing slices for` 006388_0.tif` image"
  }
}