{
  "id": 148411,
  "title": "Level 1 + 2 in jpeg (images only) dataset: 11 GB",
  "url": "/competitions/prostate-cancer-grade-assessment/discussion/148411",
  "author_name": "Konstantin Lopukhin",
  "post_date": "2020-05-04T10:00:04.075000",
  "votes": 36,
  "comment_count": 16,
  "views": 0,
  "content": "<p>Here is an image-only dataset with level 1 and level 2 images (cropped), no masks. It's size is 9 GB to download, around 11 GB uncompressed. Link: <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">https://www.kaggle.com/lopuhin/panda-2020-level-1-2</a></p>",
  "messages": [
    {
      "id": 832632,
      "postDate": "2020-05-04T10:00:04.077Z",
      "content": "<p>Here is an image-only dataset with level 1 and level 2 images (cropped), no masks. It's size is 9 GB to download, around 11 GB uncompressed. Link: <a href=\"https://www.kaggle.com/lopuhin/panda-2020-level-1-2\">https://www.kaggle.com/lopuhin/panda-2020-level-1-2</a></p>",
      "rawMarkdown": "Here is an image-only dataset with level 1 and level 2 images (cropped), no masks. It's size is 9 GB to download, around 11 GB uncompressed. Link: https://www.kaggle.com/lopuhin/panda-2020-level-1-2",
      "votes": 36
    },
    {
      "id": 871698,
      "postDate": "2020-06-02T15:22:47.617Z",
      "content": "<p>thanks for sharing! have you by chance encountered issues related to <code>PIL.Image.DecompressionBombError</code>? I get it when training on 25x256x256 images made from tiles</p>",
      "rawMarkdown": "thanks for sharing! have you by chance encountered issues related to `PIL.Image.DecompressionBombError`? I get it when training on 25x256x256 images made from tiles",
      "votes": 1,
      "replies": [
        {
          "id": 871743,
          "postDate": "2020-06-02T15:55:11.317Z",
          "content": "<p>ok, solved it, for whoever might be interested, you can increase the maximum allowed size: <a href=\"https://stackoverflow.com/questions/51152059/pillow-in-python-wont-let-me-open-image-exceeds-limit\">https://stackoverflow.com/questions/51152059/pillow-in-python-wont-let-me-open-image-exceeds-limit</a></p>",
          "rawMarkdown": "ok, solved it, for whoever might be interested, you can increase the maximum allowed size: https://stackoverflow.com/questions/51152059/pillow-in-python-wont-let-me-open-image-exceeds-limit",
          "votes": 1
        },
        {
          "id": 874356,
          "postDate": "2020-06-04T21:40:15.083Z",
          "content": "<p><code>PIL.Image.MAX_IMAGE_PIXELS = None</code> instead of maximum size to skip that DOS verifier.</p>",
          "rawMarkdown": "`PIL.Image.MAX_IMAGE_PIXELS = None` instead of maximum size to skip that DOS verifier."
        }
      ]
    },
    {
      "id": 857061,
      "postDate": "2020-05-22T09:12:33.710Z",
      "content": "<p>Hey, thanks for sharing :) Did you get your best results out of this dataset or did you have to use original tiffs? </p>",
      "rawMarkdown": "Hey, thanks for sharing :) Did you get your best results out of this dataset or did you have to use original tiffs? ",
      "replies": [
        {
          "id": 857068,
          "postDate": "2020-05-22T09:21:59.350Z",
          "content": "<p>My best submission is trained on this dataset on level 1, I use original tiffs only for submission (the only trick is to re-apply jpeg compression/decompression during submission). I'm yet to train a model on higher resolution.</p>",
          "rawMarkdown": "My best submission is trained on this dataset on level 1, I use original tiffs only for submission (the only trick is to re-apply jpeg compression/decompression during submission). I'm yet to train a model on higher resolution.",
          "votes": 3
        },
        {
          "id": 857094,
          "postDate": "2020-05-22T09:44:10.863Z",
          "content": "<p>Thanks. As I understand, to submit one has to repeat <code>image_to_jpeg</code> method partially (without saving). And do compression with the same quality (90 in your case). Am I understand that correctly? Can you provide some pseudo code for submission part? The part with compression/decompression is unclear for me ((</p>",
          "rawMarkdown": "Thanks. As I understand, to submit one has to repeat `image_to_jpeg` method partially (without saving). And do compression with the same quality (90 in your case). Am I understand that correctly? Can you provide some pseudo code for submission part? The part with compression/decompression is unclear for me (("
        },
        {
          "id": 857143,
          "postDate": "2020-05-22T10:38:42.610Z",
          "content": "<p>Yes, here is the actual code which I use (it's not optimal, I plan to move away from PIL here):</p>\n\n<p><code>\n            image = crop_white(skimage.io.MultiImage(\n                str(self.root / f'{item.image_id}.tiff'))[self.level])\n            if self.level != 0:\n                # use PIL as jpeg4py is not available on kaggle\n                buffer = io.BytesIO()\n                Image.fromarray(image).save(buffer, format='jpeg', quality=90)\n                image = np.array(Image.open(buffer))\n</code></p>",
          "rawMarkdown": "Yes, here is the actual code which I use (it's not optimal, I plan to move away from PIL here):\n\n```\n            image = crop_white(skimage.io.MultiImage(\n                str(self.root / f'{item.image_id}.tiff'))[self.level])\n            if self.level != 0:\n                # use PIL as jpeg4py is not available on kaggle\n                buffer = io.BytesIO()\n                Image.fromarray(image).save(buffer, format='jpeg', quality=90)\n                image = np.array(Image.open(buffer))\n```"
        },
        {
          "id": 857152,
          "postDate": "2020-05-22T10:55:36.947Z",
          "content": "<p><code>image = skimage.io.MultiImage(str(path))[1]</code>\n<code>image = crop_white(image)</code>\n<code>image = Image.fromarray(image)</code>\n<code>inmem = io.BytesIO()</code>\n<code>image.save(inmem, format=\"jpeg\", quality=90)</code>\n<code>img = Image.open(inmem)</code>\n<code># continue with preprocessing for your network here</code>\nShould be something like this?</p>",
          "rawMarkdown": "`image = skimage.io.MultiImage(str(path))[1]`\n`image = crop_white(image)`\n`image = Image.fromarray(image)`\n`inmem = io.BytesIO()`\n`image.save(inmem, format=\"jpeg\", quality=90)`\n`img = Image.open(inmem)`\n`# continue with preprocessing for your network here`\nShould be something like this?",
          "votes": 1
        },
        {
          "id": 857156,
          "postDate": "2020-05-22T10:57:38.440Z",
          "content": "<blockquote>\n  <p>Should be something like this?</p>\n</blockquote>\n\n<p>yes looks look to me</p>",
          "rawMarkdown": "&gt; Should be something like this?\n\nyes looks look to me"
        },
        {
          "id": 857157,
          "postDate": "2020-05-22T10:57:46.830Z",
          "content": "<p>Thanks. Get into the same code, while this page was not updated! )))</p>",
          "rawMarkdown": "Thanks. Get into the same code, while this page was not updated! )))",
          "votes": 1
        },
        {
          "id": 857209,
          "postDate": "2020-05-22T12:17:03.810Z",
          "content": "<p><a href=\"/lopuhin\">@lopuhin</a> Did you compare how much faster is to train with jpeg instead of tiff?</p>",
          "rawMarkdown": "@lopuhin Did you compare how much faster is to train with jpeg instead of tiff?"
        },
        {
          "id": 857240,
          "postDate": "2020-05-22T12:52:42.263Z",
          "content": "<p>Quick comparison load tiffs vs jpeg on the kaggle kernel gave me roughly the same time<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F789366%2F02c2b9238524242279f2504ef8821734%2FCapture.PNG?generation=1590151952480407&amp;alt=media\" alt=\"\">\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F789366%2F7b9f01aac3c99f0c43d8e5e6c55c2da4%2FCapture1.PNG?generation=1590151953238420&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Quick comparison load tiffs vs jpeg on the kaggle kernel gave me roughly the same time![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F789366%2F02c2b9238524242279f2504ef8821734%2FCapture.PNG?generation=1590151952480407&amp;alt=media)\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F789366%2F7b9f01aac3c99f0c43d8e5e6c55c2da4%2FCapture1.PNG?generation=1590151953238420&amp;alt=media)\n",
          "votes": 1
        },
        {
          "id": 857279,
          "postDate": "2020-05-22T13:24:22.047Z",
          "content": "<p>Thanks for the dataset <a href=\"/lopuhin\">@lopuhin</a>, I use it when running on servers, really helpful! Interestingly, what I've noticed is that, whether or not I compress the .tiff images to jpeg on submission have no (or insignificant) effect on LB.. Note that I use albumentations JpegCompression (I set it to always_apply=True, and the quality range to be 90-90). However, I guess it's such an easy thing to include in the submission code anyways so it's unimportant in that sense.. :-)</p>",
          "rawMarkdown": "Thanks for the dataset @lopuhin, I use it when running on servers, really helpful! Interestingly, what I've noticed is that, whether or not I compress the .tiff images to jpeg on submission have no (or insignificant) effect on LB.. Note that I use albumentations JpegCompression (I set it to always_apply=True, and the quality range to be 90-90). However, I guess it's such an easy thing to include in the submission code anyways so it's unimportant in that sense.. :-)",
          "votes": 1
        },
        {
          "id": 858060,
          "postDate": "2020-05-23T07:33:48.317Z",
          "content": "<blockquote>\n  <p>I use it when running on servers, really helpful! </p>\n</blockquote>\n\n<p>Yep that was my motivation as well, to quickly copy it.</p>\n\n<blockquote>\n  <p>whether or not I compress the .tiff images to jpeg on submission have no (or insignificant) effect on LB.. Note that I use albumentations JpegCompression</p>\n</blockquote>\n\n<p>For me the difference was around 0.02 both on LB and validation, but I didn't use this augmentation - probably that's the reason.</p>\n\n<blockquote>\n  <p>Quick comparison load tiffs vs jpeg on the kaggle kernel gave me roughly the same time</p>\n</blockquote>\n\n<p>Good point - I don't recall exact numbers, the difference may be not so large. Two notes here: jpeg4py was around 2x faster than cv2 IIRC (I'm not sure if proper turbojpeg is available on kaggle though), and also what matters for dataloader speed is not \"wall time\" but \"total time\", because you're launching multiple workers in any case. Smaller wall time but larger total time means multiple cores are utilized.</p>",
          "rawMarkdown": "&gt;  I use it when running on servers, really helpful! \n\nYep that was my motivation as well, to quickly copy it.\n\n&gt; whether or not I compress the .tiff images to jpeg on submission have no (or insignificant) effect on LB.. Note that I use albumentations JpegCompression\n\nFor me the difference was around 0.02 both on LB and validation, but I didn't use this augmentation - probably that's the reason.\n\n&gt; Quick comparison load tiffs vs jpeg on the kaggle kernel gave me roughly the same time\n\nGood point - I don't recall exact numbers, the difference may be not so large. Two notes here: jpeg4py was around 2x faster than cv2 IIRC (I'm not sure if proper turbojpeg is available on kaggle though), and also what matters for dataloader speed is not \"wall time\" but \"total time\", because you're launching multiple workers in any case. Smaller wall time but larger total time means multiple cores are utilized.",
          "votes": 1
        },
        {
          "id": 858816,
          "postDate": "2020-05-23T20:35:12.477Z",
          "content": "<p>Yep, compared both locally  as a part of dataloaders with the same number of workers, still no significant difference.\nAnyways, good to have if need to copy 11GB is better than 300GB )))</p>",
          "rawMarkdown": "Yep, compared both locally  as a part of dataloaders with the same number of workers, still no significant difference.\nAnyways, good to have if need to copy 11GB is better than 300GB )))",
          "votes": 1
        }
      ]
    },
    {
      "id": 833112,
      "postDate": "2020-05-04T15:57:48.807Z",
      "content": "<p>Thanks 🙏Will be great help.</p>",
      "rawMarkdown": "Thanks 🙏Will be great help."
    }
  ],
  "comments": [
    {
      "id": 871698,
      "author_name": "Marco Perini",
      "author_url": "",
      "post_date": "2020-06-02T15:22:47.617000",
      "content": "<p>thanks for sharing! have you by chance encountered issues related to <code>PIL.Image.DecompressionBombError</code>? I get it when training on 25x256x256 images made from tiles</p>",
      "votes": 1,
      "replies": [
        {
          "id": 871743,
          "author_name": "Marco Perini",
          "author_url": "",
          "post_date": "2020-06-02T15:55:11.317000",
          "content": "<p>ok, solved it, for whoever might be interested, you can increase the maximum allowed size: <a href=\"https://stackoverflow.com/questions/51152059/pillow-in-python-wont-let-me-open-image-exceeds-limit\">https://stackoverflow.com/questions/51152059/pillow-in-python-wont-let-me-open-image-exceeds-limit</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 874356,
          "author_name": "BachT",
          "author_url": "",
          "post_date": "2020-06-04T21:40:15.083000",
          "content": "<p><code>PIL.Image.MAX_IMAGE_PIXELS = None</code> instead of maximum size to skip that DOS verifier.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 857061,
      "author_name": "A.Demyanchuk",
      "author_url": "",
      "post_date": "2020-05-22T09:12:33.710000",
      "content": "<p>Hey, thanks for sharing :) Did you get your best results out of this dataset or did you have to use original tiffs? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 857068,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-22T09:21:59.350000",
          "content": "<p>My best submission is trained on this dataset on level 1, I use original tiffs only for submission (the only trick is to re-apply jpeg compression/decompression during submission). I'm yet to train a model on higher resolution.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 857094,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-22T09:44:10.863000",
          "content": "<p>Thanks. As I understand, to submit one has to repeat <code>image_to_jpeg</code> method partially (without saving). And do compression with the same quality (90 in your case). Am I understand that correctly? Can you provide some pseudo code for submission part? The part with compression/decompression is unclear for me ((</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 857143,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-22T10:38:42.610000",
          "content": "<p>Yes, here is the actual code which I use (it's not optimal, I plan to move away from PIL here):</p>\n\n<p><code>\n            image = crop_white(skimage.io.MultiImage(\n                str(self.root / f'{item.image_id}.tiff'))[self.level])\n            if self.level != 0:\n                # use PIL as jpeg4py is not available on kaggle\n                buffer = io.BytesIO()\n                Image.fromarray(image).save(buffer, format='jpeg', quality=90)\n                image = np.array(Image.open(buffer))\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 857152,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-22T10:55:36.947000",
          "content": "<p><code>image = skimage.io.MultiImage(str(path))[1]</code>\n<code>image = crop_white(image)</code>\n<code>image = Image.fromarray(image)</code>\n<code>inmem = io.BytesIO()</code>\n<code>image.save(inmem, format=\"jpeg\", quality=90)</code>\n<code>img = Image.open(inmem)</code>\n<code># continue with preprocessing for your network here</code>\nShould be something like this?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 857156,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-22T10:57:38.440000",
          "content": "<blockquote>\n  <p>Should be something like this?</p>\n</blockquote>\n\n<p>yes looks look to me</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 857157,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-22T10:57:46.830000",
          "content": "<p>Thanks. Get into the same code, while this page was not updated! )))</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 857209,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-22T12:17:03.810000",
          "content": "<p><a href=\"/lopuhin\">@lopuhin</a> Did you compare how much faster is to train with jpeg instead of tiff?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 857240,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-22T12:52:42.263000",
          "content": "<p>Quick comparison load tiffs vs jpeg on the kaggle kernel gave me roughly the same time<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F789366%2F02c2b9238524242279f2504ef8821734%2FCapture.PNG?generation=1590151952480407&amp;alt=media\" alt=\"\">\n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F789366%2F7b9f01aac3c99f0c43d8e5e6c55c2da4%2FCapture1.PNG?generation=1590151953238420&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 857279,
          "author_name": "Alex",
          "author_url": "",
          "post_date": "2020-05-22T13:24:22.047000",
          "content": "<p>Thanks for the dataset <a href=\"/lopuhin\">@lopuhin</a>, I use it when running on servers, really helpful! Interestingly, what I've noticed is that, whether or not I compress the .tiff images to jpeg on submission have no (or insignificant) effect on LB.. Note that I use albumentations JpegCompression (I set it to always_apply=True, and the quality range to be 90-90). However, I guess it's such an easy thing to include in the submission code anyways so it's unimportant in that sense.. :-)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 858060,
          "author_name": "Konstantin Lopukhin",
          "author_url": "",
          "post_date": "2020-05-23T07:33:48.317000",
          "content": "<blockquote>\n  <p>I use it when running on servers, really helpful! </p>\n</blockquote>\n\n<p>Yep that was my motivation as well, to quickly copy it.</p>\n\n<blockquote>\n  <p>whether or not I compress the .tiff images to jpeg on submission have no (or insignificant) effect on LB.. Note that I use albumentations JpegCompression</p>\n</blockquote>\n\n<p>For me the difference was around 0.02 both on LB and validation, but I didn't use this augmentation - probably that's the reason.</p>\n\n<blockquote>\n  <p>Quick comparison load tiffs vs jpeg on the kaggle kernel gave me roughly the same time</p>\n</blockquote>\n\n<p>Good point - I don't recall exact numbers, the difference may be not so large. Two notes here: jpeg4py was around 2x faster than cv2 IIRC (I'm not sure if proper turbojpeg is available on kaggle though), and also what matters for dataloader speed is not \"wall time\" but \"total time\", because you're launching multiple workers in any case. Smaller wall time but larger total time means multiple cores are utilized.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 858816,
          "author_name": "A.Demyanchuk",
          "author_url": "",
          "post_date": "2020-05-23T20:35:12.477000",
          "content": "<p>Yep, compared both locally  as a part of dataloaders with the same number of workers, still no significant difference.\nAnyways, good to have if need to copy 11GB is better than 300GB )))</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 833112,
      "author_name": "Debanga Raj Neog",
      "author_url": "",
      "post_date": "2020-05-04T15:57:48.807000",
      "content": "<p>Thanks 🙏Will be great help.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "832632": "Here is an image-only dataset with level 1 and level 2 images (cropped), no masks. It's size is 9 GB to download, around 11 GB uncompressed. Link: https://www.kaggle.com/lopuhin/panda-2020-level-1-2",
    "871698": "thanks for sharing! have you by chance encountered issues related to `PIL.Image.DecompressionBombError`? I get it when training on 25x256x256 images made from tiles",
    "857061": "Hey, thanks for sharing :) Did you get your best results out of this dataset or did you have to use original tiffs? ",
    "833112": "Thanks 🙏Will be great help."
  }
}