{
  "id": 369742,
  "title": "200+GB dataset is crazy",
  "url": "/competitions/g2net-detecting-continuous-gravitational-waves/discussion/369742",
  "author_name": "Apa",
  "post_date": "2022-12-01T09:12:18.797000",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I just signed up for the competition. Generally speaking, I will download the data set to the local training, and then upload the model and inference code for submission. Then I was stunned to see the data volume of 200G, it would be too inefficient to do so. Overwhelmed!!! Can you share how you work?😑😑😑</p>",
  "messages": [
    {
      "id": 2051184,
      "postDate": "2022-12-01T09:12:18.797Z",
      "content": "<p>I just signed up for the competition. Generally speaking, I will download the data set to the local training, and then upload the model and inference code for submission. Then I was stunned to see the data volume of 200G, it would be too inefficient to do so. Overwhelmed!!! Can you share how you work?😑😑😑</p>",
      "rawMarkdown": "I just signed up for the competition. Generally speaking, I will download the data set to the local training, and then upload the model and inference code for submission. Then I was stunned to see the data volume of 200G, it would be too inefficient to do so. Overwhelmed!!! Can you share how you work?😑😑😑",
      "votes": 4
    },
    {
      "id": 2052896,
      "postDate": "2022-12-02T16:14:20.033Z",
      "content": "<p>Download the database from web page directly is easy to interrupt and fail. You can copy the download link, put it in the Thunder cloud disk, and then download it from the cloud disk with a high speed. That's what I do.</p>",
      "rawMarkdown": "Download the database from web page directly is easy to interrupt and fail. You can copy the download link, put it in the Thunder cloud disk, and then download it from the cloud disk with a high speed. That's what I do.",
      "votes": 1,
      "replies": [
        {
          "id": 2055340,
          "postDate": "2022-12-05T02:52:24.310Z",
          "content": "<p>Thanks, I successfully downloaded the dataset</p>",
          "rawMarkdown": "Thanks, I successfully downloaded the dataset"
        }
      ]
    },
    {
      "id": 2051290,
      "postDate": "2022-12-01T10:23:03.707Z",
      "content": "<p>Most of that is test data. You need to generate train data yourself. And it needs to be much more than 200GB.</p>",
      "rawMarkdown": "Most of that is test data. You need to generate train data yourself. And it needs to be much more than 200GB.",
      "votes": 1,
      "replies": [
        {
          "id": 2052193,
          "postDate": "2022-12-02T02:00:21.980Z",
          "content": "<p>😂😂😂😂😂😂😂</p>",
          "rawMarkdown": "😂😂😂😂😂😂😂",
          "votes": 2
        }
      ]
    },
    {
      "id": 2052092,
      "postDate": "2022-12-01T22:54:57.573Z",
      "content": "<p>I work on a local setup for iteration speed. I post-processed the 200G into smaller .npy files. One example is Jun's notebook (<a href=\"https://www.kaggle.com/code/junkoda/basic-spectrogram-image-classification)\" target=\"_blank\">https://www.kaggle.com/code/junkoda/basic-spectrogram-image-classification)</a>, I just saved his post-processed data (2x360x128) into .npy. He/she does a 32 point average on the data (without gaps). This makes the 197GB test set -&gt; 2.8GB test set. An added bonus is that you can load it all into memory, and increase the dataloader speed significantly. You do all the .hdf5 preprocessing one time up front.<br>\nI did download the whole dataset.</p>",
      "rawMarkdown": "I work on a local setup for iteration speed. I post-processed the 200G into smaller .npy files. One example is Jun's notebook (https://www.kaggle.com/code/junkoda/basic-spectrogram-image-classification), I just saved his post-processed data (2x360x128) into .npy. He/she does a 32 point average on the data (without gaps). This makes the 197GB test set -> 2.8GB test set. An added bonus is that you can load it all into memory, and increase the dataloader speed significantly. You do all the .hdf5 preprocessing one time up front.\nI did download the whole dataset.",
      "votes": 2,
      "replies": [
        {
          "id": 2052192,
          "postDate": "2022-12-02T01:59:48.893Z",
          "content": "<p>Thanks for your reply I will try it</p>",
          "rawMarkdown": "Thanks for your reply I will try it"
        }
      ]
    },
    {
      "id": 2074748,
      "postDate": "2022-12-24T14:59:53.987Z",
      "content": "<p>I have encountered similar problems, so the data itself is large, and there are not sth wrong with my operation, is that right?</p>",
      "rawMarkdown": "I have encountered similar problems, so the data itself is large, and there are not sth wrong with my operation, is that right?"
    },
    {
      "id": 2053375,
      "postDate": "2022-12-03T07:44:25.050Z",
      "content": "<p>RSNA has a larger dataset. . .</p>",
      "rawMarkdown": "RSNA has a larger dataset. . ."
    },
    {
      "id": 2051200,
      "postDate": "2022-12-01T09:24:32.703Z",
      "content": "<p>Hah, \"nice\". I will pass competition <a href=\"https://www.kaggle.com/yasso1\" target=\"_blank\">@yasso1</a> </p>",
      "rawMarkdown": "Hah, \"nice\". I will pass competition @yasso1 ",
      "replies": [
        {
          "id": 2052194,
          "postDate": "2022-12-02T02:01:14.117Z",
          "content": "<p>I decided to give it a try, even though the first</p>",
          "rawMarkdown": "I decided to give it a try, even though the first"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2052896,
      "author_name": "E-Max AI",
      "author_url": "",
      "post_date": "2022-12-02T16:14:20.033000",
      "content": "<p>Download the database from web page directly is easy to interrupt and fail. You can copy the download link, put it in the Thunder cloud disk, and then download it from the cloud disk with a high speed. That's what I do.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2055340,
          "author_name": "Apa",
          "author_url": "",
          "post_date": "2022-12-05T02:52:24.310000",
          "content": "<p>Thanks, I successfully downloaded the dataset</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2051290,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2022-12-01T10:23:03.707000",
      "content": "<p>Most of that is test data. You need to generate train data yourself. And it needs to be much more than 200GB.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2052193,
          "author_name": "Apa",
          "author_url": "",
          "post_date": "2022-12-02T02:00:21.980000",
          "content": "<p>😂😂😂😂😂😂😂</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2052092,
      "author_name": "datadote",
      "author_url": "",
      "post_date": "2022-12-01T22:54:57.573000",
      "content": "<p>I work on a local setup for iteration speed. I post-processed the 200G into smaller .npy files. One example is Jun's notebook (<a href=\"https://www.kaggle.com/code/junkoda/basic-spectrogram-image-classification)\" target=\"_blank\">https://www.kaggle.com/code/junkoda/basic-spectrogram-image-classification)</a>, I just saved his post-processed data (2x360x128) into .npy. He/she does a 32 point average on the data (without gaps). This makes the 197GB test set -&gt; 2.8GB test set. An added bonus is that you can load it all into memory, and increase the dataloader speed significantly. You do all the .hdf5 preprocessing one time up front.<br>\nI did download the whole dataset.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2052192,
          "author_name": "Apa",
          "author_url": "",
          "post_date": "2022-12-02T01:59:48.893000",
          "content": "<p>Thanks for your reply I will try it</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2074748,
      "author_name": "Niu Meiqi",
      "author_url": "",
      "post_date": "2022-12-24T14:59:53.987000",
      "content": "<p>I have encountered similar problems, so the data itself is large, and there are not sth wrong with my operation, is that right?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2053375,
      "author_name": "E-Max AI",
      "author_url": "",
      "post_date": "2022-12-03T07:44:25.050000",
      "content": "<p>RSNA has a larger dataset. . .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2051200,
      "author_name": "Olqa_842",
      "author_url": "",
      "post_date": "2022-12-01T09:24:32.703000",
      "content": "<p>Hah, \"nice\". I will pass competition <a href=\"https://www.kaggle.com/yasso1\" target=\"_blank\">@yasso1</a> </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2052194,
          "author_name": "Apa",
          "author_url": "",
          "post_date": "2022-12-02T02:01:14.117000",
          "content": "<p>I decided to give it a try, even though the first</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2051184": "I just signed up for the competition. Generally speaking, I will download the data set to the local training, and then upload the model and inference code for submission. Then I was stunned to see the data volume of 200G, it would be too inefficient to do so. Overwhelmed!!! Can you share how you work?😑😑😑",
    "2052896": "Download the database from web page directly is easy to interrupt and fail. You can copy the download link, put it in the Thunder cloud disk, and then download it from the cloud disk with a high speed. That's what I do.",
    "2051290": "Most of that is test data. You need to generate train data yourself. And it needs to be much more than 200GB.",
    "2052092": "I work on a local setup for iteration speed. I post-processed the 200G into smaller .npy files. One example is Jun's notebook (https://www.kaggle.com/code/junkoda/basic-spectrogram-image-classification), I just saved his post-processed data (2x360x128) into .npy. He/she does a 32 point average on the data (without gaps). This makes the 197GB test set -> 2.8GB test set. An added bonus is that you can load it all into memory, and increase the dataloader speed significantly. You do all the .hdf5 preprocessing one time up front.\nI did download the whole dataset.",
    "2074748": "I have encountered similar problems, so the data itself is large, and there are not sth wrong with my operation, is that right?",
    "2053375": "RSNA has a larger dataset. . .",
    "2051200": "Hah, \"nice\". I will pass competition @yasso1 "
  }
}