{
  "id": 495109,
  "title": "Does the low-resolution data amount to 744GB?",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/495109",
  "author_name": "chumajin",
  "post_date": "2024-04-19T17:18:08.295000",
  "votes": 14,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I tried recreating the training data from the low-resolution data described in the Data explanation <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">page</a>.</p>\n<pre><code>from datasets import load_dataset\n\n  load_dataset()\n</code></pre>\n<p>But the data download from Hugging Face doesn't seem to be finishing any time soon.</p>\n<p>Therefore, when I checked <a href=\"https://leap-stc.github.io/ClimSim/dataset.html\" target=\"_blank\">GitHub</a>, there is a description here that low-resolution real geography is as follows.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F6c0d94c9ea3aac387b134492e6e95b4a%2FClipboard01.jpg?generation=1713546432746325&amp;alt=media\"></p>\n<p>From this description, the low-resolution data is 744GB. So, I understood this explanation in data explanation <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">page</a>.</p>\n<pre><code>Given  size   data, we recommend downloading       converting them   more manageable  (like parquet  npy). If you have access   High-Performance Computing (HPC) cluster  sufficient storage availability, you can also download  data directly  your remote cluster   Kaggle API   following :\n\nkaggle competitions download -c ClimSim\n</code></pre>\n<p>Maybe we should download each file, for example from <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res/tree/main/train\" target=\"_blank\">low-resolution hugging face site</a> and remake train data with <a href=\"https://github.com/leap-stc/ClimSim/blob/main/for_kaggle_users.py\" target=\"_blank\">this script</a>.</p>\n<p>Or, using kaggle dataset ( 192.3 Gb ).</p>\n<p>Anyway, this competition seems to be testing the ability to process a lot of data.<br>\nPlease let me know if I have misunderstood anything.</p>\n<p>Enjoy !!</p>",
  "messages": [
    {
      "id": 2761154,
      "postDate": "2024-04-19T17:18:08.297Z",
      "content": "<p>I tried recreating the training data from the low-resolution data described in the Data explanation <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">page</a>.</p>\n<pre><code>from datasets import load_dataset\n\n  load_dataset()\n</code></pre>\n<p>But the data download from Hugging Face doesn't seem to be finishing any time soon.</p>\n<p>Therefore, when I checked <a href=\"https://leap-stc.github.io/ClimSim/dataset.html\" target=\"_blank\">GitHub</a>, there is a description here that low-resolution real geography is as follows.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F6c0d94c9ea3aac387b134492e6e95b4a%2FClipboard01.jpg?generation=1713546432746325&amp;alt=media\"></p>\n<p>From this description, the low-resolution data is 744GB. So, I understood this explanation in data explanation <a href=\"https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data\" target=\"_blank\">page</a>.</p>\n<pre><code>Given  size   data, we recommend downloading       converting them   more manageable  (like parquet  npy). If you have access   High-Performance Computing (HPC) cluster  sufficient storage availability, you can also download  data directly  your remote cluster   Kaggle API   following :\n\nkaggle competitions download -c ClimSim\n</code></pre>\n<p>Maybe we should download each file, for example from <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res/tree/main/train\" target=\"_blank\">low-resolution hugging face site</a> and remake train data with <a href=\"https://github.com/leap-stc/ClimSim/blob/main/for_kaggle_users.py\" target=\"_blank\">this script</a>.</p>\n<p>Or, using kaggle dataset ( 192.3 Gb ).</p>\n<p>Anyway, this competition seems to be testing the ability to process a lot of data.<br>\nPlease let me know if I have misunderstood anything.</p>\n<p>Enjoy !!</p>",
      "rawMarkdown": "I tried recreating the training data from the low-resolution data described in the Data explanation [page](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data).\n~~~\nfrom datasets import load_dataset\n\ndataset = load_dataset(\"LEAP/ClimSim_low-res\")\n~~~\n\nBut the data download from Hugging Face doesn't seem to be finishing any time soon.\n\nTherefore, when I checked [GitHub](https://leap-stc.github.io/ClimSim/dataset.html), there is a description here that low-resolution real geography is as follows.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F6c0d94c9ea3aac387b134492e6e95b4a%2FClipboard01.jpg?generation=1713546432746325&alt=media)\n\nFrom this description, the low-resolution data is 744GB. So, I understood this explanation in data explanation [page](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data).\n\n~~~\nGiven the size of the data, we recommend downloading one file at a time and converting them to a more manageable format (like parquet or npy). If you have access to a High-Performance Computing (HPC) cluster with sufficient storage availability, you can also download the data directly to your remote cluster using the Kaggle API with the following command:\n\nkaggle competitions download -c ClimSim\n~~~\n\nMaybe we should download each file, for example from [low-resolution hugging face site](https://huggingface.co/datasets/LEAP/ClimSim_low-res/tree/main/train) and remake train data with [this script](https://github.com/leap-stc/ClimSim/blob/main/for_kaggle_users.py).\n\nOr, using kaggle dataset ( 192.3 Gb ).\n\nAnyway, this competition seems to be testing the ability to process a lot of data.\nPlease let me know if I have misunderstood anything.\n\nEnjoy !!",
      "votes": 13
    },
    {
      "id": 2761161,
      "postDate": "2024-04-19T17:20:50.917Z",
      "content": "<p><a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> you are absolutely right. This data amounts to 740 Gb! I tried to work on this earlier today and was warded off due to the size of the data! <br>\nI am unsure how one may process this data at this size with the resources available at an individual's disposal. Perhaps a cap on the size of data considering the nature of the participant base should be considered. </p>",
      "rawMarkdown": "@chumajin you are absolutely right. This data amounts to 740 Gb! I tried to work on this earlier today and was warded off due to the size of the data! \nI am unsure how one may process this data at this size with the resources available at an individual's disposal. Perhaps a cap on the size of data considering the nature of the participant base should be considered. ",
      "votes": 2,
      "replies": [
        {
          "id": 2761169,
          "postDate": "2024-04-19T17:24:12.163Z",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Thank you for comment! I also ended up not being able to download it. Yes, using all the data might be difficult on a typical machine. It seems that the key will be which data to use and how much of it to use.</p>",
          "rawMarkdown": "@ravi20076 Thank you for comment! I also ended up not being able to download it. Yes, using all the data might be difficult on a typical machine. It seems that the key will be which data to use and how much of it to use.",
          "votes": 3
        },
        {
          "id": 2766228,
          "postDate": "2024-04-21T15:01:08.917Z",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> i am doing batch processing for neural nets although i agree with some of your points i have tried some iterations with spark XGB too and it took more that 8 hours <br>\nbut in my opinion working with npy files might do the trick </p>",
          "rawMarkdown": "Hey @ravi20076 i am doing batch processing for neural nets although i agree with some of your points i have tried some iterations with spark XGB too and it took more that 8 hours \nbut in my opinion working with npy files might do the trick \n",
          "votes": 2,
          "replies": [
            {
              "id": 2766231,
              "postDate": "2024-04-21T15:02:40.680Z",
              "content": "<p>Please rate my solution if helpful <br>\n<a href=\"https://www.kaggle.com/code/starcs2001/neural-nets-starter-with-batch-processing\" target=\"_blank\">https://www.kaggle.com/code/starcs2001/neural-nets-starter-with-batch-processing</a></p>",
              "rawMarkdown": "Please rate my solution if helpful \nhttps://www.kaggle.com/code/starcs2001/neural-nets-starter-with-batch-processing\n",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2761161,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2024-04-19T17:20:50.917000",
      "content": "<p><a href=\"https://www.kaggle.com/chumajin\" target=\"_blank\">@chumajin</a> you are absolutely right. This data amounts to 740 Gb! I tried to work on this earlier today and was warded off due to the size of the data! <br>\nI am unsure how one may process this data at this size with the resources available at an individual's disposal. Perhaps a cap on the size of data considering the nature of the participant base should be considered. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2761169,
          "author_name": "chumajin",
          "author_url": "",
          "post_date": "2024-04-19T17:24:12.163000",
          "content": "<p><a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> Thank you for comment! I also ended up not being able to download it. Yes, using all the data might be difficult on a typical machine. It seems that the key will be which data to use and how much of it to use.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 2766228,
          "author_name": "Satej Raste",
          "author_url": "",
          "post_date": "2024-04-21T15:01:08.917000",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/ravi20076\" target=\"_blank\">@ravi20076</a> i am doing batch processing for neural nets although i agree with some of your points i have tried some iterations with spark XGB too and it took more that 8 hours <br>\nbut in my opinion working with npy files might do the trick </p>",
          "votes": 2,
          "replies": [
            {
              "id": 2766231,
              "author_name": "Satej Raste",
              "author_url": "",
              "post_date": "2024-04-21T15:02:40.680000",
              "content": "<p>Please rate my solution if helpful <br>\n<a href=\"https://www.kaggle.com/code/starcs2001/neural-nets-starter-with-batch-processing\" target=\"_blank\">https://www.kaggle.com/code/starcs2001/neural-nets-starter-with-batch-processing</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2761154": "I tried recreating the training data from the low-resolution data described in the Data explanation [page](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data).\n~~~\nfrom datasets import load_dataset\n\ndataset = load_dataset(\"LEAP/ClimSim_low-res\")\n~~~\n\nBut the data download from Hugging Face doesn't seem to be finishing any time soon.\n\nTherefore, when I checked [GitHub](https://leap-stc.github.io/ClimSim/dataset.html), there is a description here that low-resolution real geography is as follows.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4001300%2F6c0d94c9ea3aac387b134492e6e95b4a%2FClipboard01.jpg?generation=1713546432746325&alt=media)\n\nFrom this description, the low-resolution data is 744GB. So, I understood this explanation in data explanation [page](https://www.kaggle.com/competitions/leap-atmospheric-physics-ai-climsim/data).\n\n~~~\nGiven the size of the data, we recommend downloading one file at a time and converting them to a more manageable format (like parquet or npy). If you have access to a High-Performance Computing (HPC) cluster with sufficient storage availability, you can also download the data directly to your remote cluster using the Kaggle API with the following command:\n\nkaggle competitions download -c ClimSim\n~~~\n\nMaybe we should download each file, for example from [low-resolution hugging face site](https://huggingface.co/datasets/LEAP/ClimSim_low-res/tree/main/train) and remake train data with [this script](https://github.com/leap-stc/ClimSim/blob/main/for_kaggle_users.py).\n\nOr, using kaggle dataset ( 192.3 Gb ).\n\nAnyway, this competition seems to be testing the ability to process a lot of data.\nPlease let me know if I have misunderstood anything.\n\nEnjoy !!",
    "2761161": "@chumajin you are absolutely right. This data amounts to 740 Gb! I tried to work on this earlier today and was warded off due to the size of the data! \nI am unsure how one may process this data at this size with the resources available at an individual's disposal. Perhaps a cap on the size of data considering the nature of the participant base should be considered. "
  }
}