{
  "id": 506278,
  "title": "Full train dataset generation",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/506278",
  "author_name": "Federico Peccia",
  "post_date": "2024-05-21T08:49:59.710000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As we know, the train csv file was generated using <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">this data</a> as input, transformed using <a href=\"https://github.com/leap-stc/ClimSim/blob/main/for_kaggle_users.py\" target=\"_blank\">this script</a>. The script samples the training data to generate a reduced sample.</p>\n<p>I am looking at generating a new training csv using the full dataset, but I am not sure if this is really worth the effort. Has someone already tried this? Or do you think that the data we have is already enough?</p>",
  "messages": [
    {
      "id": 2827114,
      "postDate": "2024-05-21T08:49:59.710Z",
      "content": "<p>As we know, the train csv file was generated using <a href=\"https://huggingface.co/datasets/LEAP/ClimSim_low-res\" target=\"_blank\">this data</a> as input, transformed using <a href=\"https://github.com/leap-stc/ClimSim/blob/main/for_kaggle_users.py\" target=\"_blank\">this script</a>. The script samples the training data to generate a reduced sample.</p>\n<p>I am looking at generating a new training csv using the full dataset, but I am not sure if this is really worth the effort. Has someone already tried this? Or do you think that the data we have is already enough?</p>",
      "rawMarkdown": "As we know, the train csv file was generated using [this data](https://huggingface.co/datasets/LEAP/ClimSim_low-res) as input, transformed using [this script](https://github.com/leap-stc/ClimSim/blob/main/for_kaggle_users.py). The script samples the training data to generate a reduced sample.\n\nI am looking at generating a new training csv using the full dataset, but I am not sure if this is really worth the effort. Has someone already tried this? Or do you think that the data we have is already enough?",
      "votes": 4
    },
    {
      "id": 2829533,
      "postDate": "2024-05-22T16:12:15.373Z",
      "content": "<p><a href=\"https://www.kaggle.com/fpeccia\" target=\"_blank\">@fpeccia</a> I uploaded the full dataset for the first year: <a href=\"https://www.kaggle.com/datasets/abiolatti/leap-complete-training-data-0001/data\" target=\"_blank\">https://www.kaggle.com/datasets/abiolatti/leap-complete-training-data-0001/data</a> </p>",
      "rawMarkdown": "@fpeccia I uploaded the full dataset for the first year: https://www.kaggle.com/datasets/abiolatti/leap-complete-training-data-0001/data ",
      "votes": 2
    },
    {
      "id": 2827533,
      "postDate": "2024-05-21T14:37:13.330Z",
      "content": "<p>I'm currently trying, but it takes a while 😅</p>\n<p>I would suggest using a better data format than CSV, for example parquet or tfrecord</p>",
      "rawMarkdown": "I'm currently trying, but it takes a while 😅\n\nI would suggest using a better data format than CSV, for example parquet or tfrecord",
      "replies": [
        {
          "id": 2827996,
          "postDate": "2024-05-21T19:15:28.730Z",
          "content": "<p>The first thing I did with the original training and test data was to convert it to parquet, it was impossible to work otherwise hehe</p>",
          "rawMarkdown": "The first thing I did with the original training and test data was to convert it to parquet, it was impossible to work otherwise hehe",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2829533,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2024-05-22T16:12:15.373000",
      "content": "<p><a href=\"https://www.kaggle.com/fpeccia\" target=\"_blank\">@fpeccia</a> I uploaded the full dataset for the first year: <a href=\"https://www.kaggle.com/datasets/abiolatti/leap-complete-training-data-0001/data\" target=\"_blank\">https://www.kaggle.com/datasets/abiolatti/leap-complete-training-data-0001/data</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2827533,
      "author_name": "Amedeo Biolatti",
      "author_url": "",
      "post_date": "2024-05-21T14:37:13.330000",
      "content": "<p>I'm currently trying, but it takes a while 😅</p>\n<p>I would suggest using a better data format than CSV, for example parquet or tfrecord</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2827996,
          "author_name": "Federico Peccia",
          "author_url": "",
          "post_date": "2024-05-21T19:15:28.730000",
          "content": "<p>The first thing I did with the original training and test data was to convert it to parquet, it was impossible to work otherwise hehe</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2827114": "As we know, the train csv file was generated using [this data](https://huggingface.co/datasets/LEAP/ClimSim_low-res) as input, transformed using [this script](https://github.com/leap-stc/ClimSim/blob/main/for_kaggle_users.py). The script samples the training data to generate a reduced sample.\n\nI am looking at generating a new training csv using the full dataset, but I am not sure if this is really worth the effort. Has someone already tried this? Or do you think that the data we have is already enough?",
    "2829533": "@fpeccia I uploaded the full dataset for the first year: https://www.kaggle.com/datasets/abiolatti/leap-complete-training-data-0001/data ",
    "2827533": "I'm currently trying, but it takes a while 😅\n\nI would suggest using a better data format than CSV, for example parquet or tfrecord"
  }
}