{
  "id": 419581,
  "title": "What is your environment ? (Data size looks really big. I'm hesitating...)",
  "url": "/competitions/google-research-identify-contrails-reduce-global-warming/discussion/419581",
  "author_name": "k_tomo",
  "post_date": "2023-06-26T16:09:27.103000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I'm begginer of image competition. Therefore, I'd like to know good environment to deal with a lot of images.</p>\n<p>This competition looks very interesting. However, the data size is huge!(over 400GB!!).<br>\nHow do you deal with such a big data? Please share your way.</p>",
  "messages": [
    {
      "id": 2319143,
      "postDate": "2023-06-26T23:12:55.807Z",
      "content": "<p>The original data size is huge, but by preprocessing (selecting time stamp and band width) it results in around 20GB. So you can conduct model training in kaggle notebooks.</p>",
      "rawMarkdown": "The original data size is huge, but by preprocessing (selecting time stamp and band width) it results in around 20GB. So you can conduct model training in kaggle notebooks.",
      "votes": 3
    },
    {
      "id": 2318932,
      "postDate": "2023-06-26T17:25:34.670Z",
      "content": "<p>I use my local setup which has 2x 4090 ,7950x , 2x2TB NVME PCIE gen 4 (WD SN850X really good SSD when you are reading a lot of npy files)</p>",
      "rawMarkdown": "I use my local setup which has 2x 4090 ,7950x , 2x2TB NVME PCIE gen 4 (WD SN850X really good SSD when you are reading a lot of npy files)",
      "votes": 3
    },
    {
      "id": 2321092,
      "postDate": "2023-06-28T09:24:51.840Z",
      "content": "<p>I don't intend to promote my datasets but please look at previous discussions and preprocessing done.</p>\n<p>You are more then welcome to read this discussion thread and hope this helps you.</p>\n<p><a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/411713\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/411713</a></p>",
      "rawMarkdown": "I don't intend to promote my datasets but please look at previous discussions and preprocessing done.\n\nYou are more then welcome to read this discussion thread and hope this helps you.\n\nhttps://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/411713",
      "votes": 1
    },
    {
      "id": 2318850,
      "postDate": "2023-06-26T16:09:27.103Z",
      "content": "<p>I'm begginer of image competition. Therefore, I'd like to know good environment to deal with a lot of images.</p>\n<p>This competition looks very interesting. However, the data size is huge!(over 400GB!!).<br>\nHow do you deal with such a big data? Please share your way.</p>",
      "rawMarkdown": "I'm begginer of image competition. Therefore, I'd like to know good environment to deal with a lot of images.\n\nThis competition looks very interesting. However, the data size is huge!(over 400GB!!).\nHow do you deal with such a big data? Please share your way.\n\n",
      "votes": 1
    },
    {
      "id": 2323141,
      "postDate": "2023-06-29T17:30:10.167Z",
      "content": "<p>I use following techniques to fit all the data to memory:</p>\n<ol>\n<li>As mentioned before, converting to \"ash\" RGB color-scheme and selecting single image with index 4 of provided 8 frames reduces dataset 3 x 8 = 24 times to below 20 Gb</li>\n<li>float32 images could be clipped to [0, 1] and multiplied by 255 and then quantized to uint8 to reduce the size of \"ash\" dataset to merely 5 Gb. While losing information by quantization, it allows to e. g. use all the 8 frames on 64 Gb RAM machine if your solution needs that or use all raw band images. You should check if quantization harms performance of your solution of course.</li>\n</ol>\n<p>Also I cache the pre-processed images in RAM within single run and dump the cache to disk, it is optional so as you could simply pre-process the dataset once and save it, but allows to try different pre-processing schemes without much versioning pain.</p>\n<p>In my experience, all the data should be in memory because it greatly reduces epoch time.</p>",
      "rawMarkdown": "I use following techniques to fit all the data to memory:\n\n1. As mentioned before, converting to \"ash\" RGB color-scheme and selecting single image with index 4 of provided 8 frames reduces dataset 3 x 8 = 24 times to below 20 Gb\n2. float32 images could be clipped to [0, 1] and multiplied by 255 and then quantized to uint8 to reduce the size of \"ash\" dataset to merely 5 Gb. While losing information by quantization, it allows to e. g. use all the 8 frames on 64 Gb RAM machine if your solution needs that or use all raw band images. You should check if quantization harms performance of your solution of course.\n\nAlso I cache the pre-processed images in RAM within single run and dump the cache to disk, it is optional so as you could simply pre-process the dataset once and save it, but allows to try different pre-processing schemes without much versioning pain.\n\nIn my experience, all the data should be in memory because it greatly reduces epoch time.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2319143,
      "author_name": "luddite^",
      "author_url": "",
      "post_date": "2023-06-26T23:12:55.807000",
      "content": "<p>The original data size is huge, but by preprocessing (selecting time stamp and band width) it results in around 20GB. So you can conduct model training in kaggle notebooks.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2318932,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2023-06-26T17:25:34.670000",
      "content": "<p>I use my local setup which has 2x 4090 ,7950x , 2x2TB NVME PCIE gen 4 (WD SN850X really good SSD when you are reading a lot of npy files)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2321092,
      "author_name": "Kenni",
      "author_url": "",
      "post_date": "2023-06-28T09:24:51.840000",
      "content": "<p>I don't intend to promote my datasets but please look at previous discussions and preprocessing done.</p>\n<p>You are more then welcome to read this discussion thread and hope this helps you.</p>\n<p><a href=\"https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/411713\" target=\"_blank\">https://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/411713</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2323141,
      "author_name": "Mikhail Kotyushev",
      "author_url": "",
      "post_date": "2023-06-29T17:30:10.167000",
      "content": "<p>I use following techniques to fit all the data to memory:</p>\n<ol>\n<li>As mentioned before, converting to \"ash\" RGB color-scheme and selecting single image with index 4 of provided 8 frames reduces dataset 3 x 8 = 24 times to below 20 Gb</li>\n<li>float32 images could be clipped to [0, 1] and multiplied by 255 and then quantized to uint8 to reduce the size of \"ash\" dataset to merely 5 Gb. While losing information by quantization, it allows to e. g. use all the 8 frames on 64 Gb RAM machine if your solution needs that or use all raw band images. You should check if quantization harms performance of your solution of course.</li>\n</ol>\n<p>Also I cache the pre-processed images in RAM within single run and dump the cache to disk, it is optional so as you could simply pre-process the dataset once and save it, but allows to try different pre-processing schemes without much versioning pain.</p>\n<p>In my experience, all the data should be in memory because it greatly reduces epoch time.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2319143": "The original data size is huge, but by preprocessing (selecting time stamp and band width) it results in around 20GB. So you can conduct model training in kaggle notebooks.",
    "2318932": "I use my local setup which has 2x 4090 ,7950x , 2x2TB NVME PCIE gen 4 (WD SN850X really good SSD when you are reading a lot of npy files)",
    "2321092": "I don't intend to promote my datasets but please look at previous discussions and preprocessing done.\n\nYou are more then welcome to read this discussion thread and hope this helps you.\n\nhttps://www.kaggle.com/competitions/google-research-identify-contrails-reduce-global-warming/discussion/411713",
    "2318850": "I'm begginer of image competition. Therefore, I'd like to know good environment to deal with a lot of images.\n\nThis competition looks very interesting. However, the data size is huge!(over 400GB!!).\nHow do you deal with such a big data? Please share your way.\n\n",
    "2323141": "I use following techniques to fit all the data to memory:\n\n1. As mentioned before, converting to \"ash\" RGB color-scheme and selecting single image with index 4 of provided 8 frames reduces dataset 3 x 8 = 24 times to below 20 Gb\n2. float32 images could be clipped to [0, 1] and multiplied by 255 and then quantized to uint8 to reduce the size of \"ash\" dataset to merely 5 Gb. While losing information by quantization, it allows to e. g. use all the 8 frames on 64 Gb RAM machine if your solution needs that or use all raw band images. You should check if quantization harms performance of your solution of course.\n\nAlso I cache the pre-processed images in RAM within single run and dump the cache to disk, it is optional so as you could simply pre-process the dataset once and save it, but allows to try different pre-processing schemes without much versioning pain.\n\nIn my experience, all the data should be in memory because it greatly reduces epoch time."
  }
}