{
  "id": 508922,
  "title": "Whether increasing the amount of data can improve the score",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/508922",
  "author_name": "Zhuoqun Li",
  "post_date": "2024-05-31T13:44:15.630000",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Now that I'm only using about 30% of data for training, I wonder if using all data for training will help improve my score? Can anyone give me some advice?</p>",
  "messages": [
    {
      "id": 2847271,
      "postDate": "2024-05-31T13:44:15.630Z",
      "content": "<p>Now that I'm only using about 30% of data for training, I wonder if using all data for training will help improve my score? Can anyone give me some advice?</p>",
      "rawMarkdown": "Now that I'm only using about 30% of data for training, I wonder if using all data for training will help improve my score? Can anyone give me some advice?",
      "votes": 3
    },
    {
      "id": 2847599,
      "postDate": "2024-05-31T15:37:47.663Z",
      "content": "<p>Going from ~0.9M to ~9M samples in training (*10) give me between +0.06 (bad models) to +0.03 (best models) boost in score.<br>\nSo yea, going from 30% to 100% would give you a boost. But not a HUGE one.</p>",
      "rawMarkdown": "Going from ~0.9M to ~9M samples in training (*10) give me between +0.06 (bad models) to +0.03 (best models) boost in score.\nSo yea, going from 30% to 100% would give you a boost. But not a HUGE one.",
      "votes": 2,
      "replies": [
        {
          "id": 2848270,
          "postDate": "2024-06-01T00:11:19.923Z",
          "content": "<p>what if using  all data from <a href=\"https://github.com/leap-stc/ClimSim\" target=\"_blank\">https://github.com/leap-stc/ClimSim</a>, which has 100M samples, have you tried that?</p>",
          "rawMarkdown": "what if using  all data from https://github.com/leap-stc/ClimSim, which has 100M samples, have you tried that?",
          "replies": [
            {
              "id": 2848434,
              "postDate": "2024-06-01T03:12:48.940Z",
              "content": "<p>I think this might require a lot of computing resources</p>",
              "rawMarkdown": "I think this might require a lot of computing resources"
            },
            {
              "id": 2848746,
              "postDate": "2024-06-01T06:44:26.057Z",
              "content": "<p>Not yet, my score is Kaggle data only</p>",
              "rawMarkdown": "Not yet, my score is Kaggle data only",
              "votes": 6
            },
            {
              "id": 2848851,
              "postDate": "2024-06-01T08:50:34.517Z",
              "content": "<p>Not <em>THAT</em> much, I'm currently using all the data from HF and I can train for one epoch in less than one hour on my laptop (RTX4080) if I keep the model small, 2/3 hours with bigger models. I released about half of the data in tfrecord format as a dataset if you want to try</p>",
              "rawMarkdown": "Not *THAT* much, I'm currently using all the data from HF and I can train for one epoch in less than one hour on my laptop (RTX4080) if I keep the model small, 2/3 hours with bigger models. I released about half of the data in tfrecord format as a dataset if you want to try",
              "votes": 2
            },
            {
              "id": 2855139,
              "postDate": "2024-06-04T16:07:44.100Z",
              "content": "<p>Amadeo, I am curious, how do you prepare you data? do you use your laptop or do you use cloud computing? I have been using Google Cloud, but most of the cost is in accessing the data from object store (\"buckets\").</p>",
              "rawMarkdown": "Amadeo, I am curious, how do you prepare you data? do you use your laptop or do you use cloud computing? I have been using Google Cloud, but most of the cost is in accessing the data from object store (\"buckets\")."
            },
            {
              "id": 2855236,
              "postDate": "2024-06-04T17:14:38.593Z",
              "content": "<p>Everything is done locally (takes about 48h for the full dataset using a single thread). I tried to use a Kaggle notebook but the APIs don't allow to update a dataset without deleting the current files.</p>\n<p>I uploaded the first 4 year of the dataset here: <a href=\"https://www.kaggle.com/datasets/abiolatti/climsim-lowres-complete-normalized\" target=\"_blank\">https://www.kaggle.com/datasets/abiolatti/climsim-lowres-complete-normalized</a></p>",
              "rawMarkdown": "Everything is done locally (takes about 48h for the full dataset using a single thread). I tried to use a Kaggle notebook but the APIs don't allow to update a dataset without deleting the current files.\n\nI uploaded the first 4 year of the dataset here: https://www.kaggle.com/datasets/abiolatti/climsim-lowres-complete-normalized",
              "votes": 2
            },
            {
              "id": 2855295,
              "postDate": "2024-06-04T17:44:43.343Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2857125,
              "postDate": "2024-06-05T17:52:39.107Z",
              "content": "<p>Great, thanks as always for the insight. </p>",
              "rawMarkdown": "Great, thanks as always for the insight. "
            }
          ]
        },
        {
          "id": 2848432,
          "postDate": "2024-06-01T03:12:12.733Z",
          "content": "<p>Thank you, it's very helpful.</p>",
          "rawMarkdown": "Thank you, it's very helpful."
        }
      ]
    },
    {
      "id": 2847327,
      "postDate": "2024-05-31T14:00:30.923Z",
      "content": "<p>I used 95% and 80% of the competition data to train my models. When using 80%, the LB score loses 0.0024 compared to 95%. I suggest using at least 70% of the competition data to train your models.</p>\n<p>To handle this size of data, I recommend using Kaggle TPUv3-8 or TPUv2-8 instead of GPU.</p>",
      "rawMarkdown": "I used 95% and 80% of the competition data to train my models. When using 80%, the LB score loses 0.0024 compared to 95%. I suggest using at least 70% of the competition data to train your models.\n\nTo handle this size of data, I recommend using Kaggle TPUv3-8 or TPUv2-8 instead of GPU.",
      "replies": [
        {
          "id": 2847374,
          "postDate": "2024-05-31T14:14:00.880Z",
          "content": "<p>thank you for your sharing</p>",
          "rawMarkdown": "thank you for your sharing"
        }
      ]
    },
    {
      "id": 2878353,
      "postDate": "2024-06-18T22:23:38.820Z",
      "content": "<p>Yeah, had the same doubt, especially as I am trying to run the model on my system, so can only work on a subset of data.</p>",
      "rawMarkdown": "Yeah, had the same doubt, especially as I am trying to run the model on my system, so can only work on a subset of data.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2847599,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-05-31T15:37:47.663000",
      "content": "<p>Going from ~0.9M to ~9M samples in training (*10) give me between +0.06 (bad models) to +0.03 (best models) boost in score.<br>\nSo yea, going from 30% to 100% would give you a boost. But not a HUGE one.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2848270,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-06-01T00:11:19.923000",
          "content": "<p>what if using  all data from <a href=\"https://github.com/leap-stc/ClimSim\" target=\"_blank\">https://github.com/leap-stc/ClimSim</a>, which has 100M samples, have you tried that?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2848434,
              "author_name": "Zhuoqun Li",
              "author_url": "",
              "post_date": "2024-06-01T03:12:48.940000",
              "content": "<p>I think this might require a lot of computing resources</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2848746,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-06-01T06:44:26.057000",
              "content": "<p>Not yet, my score is Kaggle data only</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2848851,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-06-01T08:50:34.517000",
              "content": "<p>Not <em>THAT</em> much, I'm currently using all the data from HF and I can train for one epoch in less than one hour on my laptop (RTX4080) if I keep the model small, 2/3 hours with bigger models. I released about half of the data in tfrecord format as a dataset if you want to try</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2855139,
              "author_name": "Juan D C F",
              "author_url": "",
              "post_date": "2024-06-04T16:07:44.100000",
              "content": "<p>Amadeo, I am curious, how do you prepare you data? do you use your laptop or do you use cloud computing? I have been using Google Cloud, but most of the cost is in accessing the data from object store (\"buckets\").</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2855236,
              "author_name": "Amedeo Biolatti",
              "author_url": "",
              "post_date": "2024-06-04T17:14:38.593000",
              "content": "<p>Everything is done locally (takes about 48h for the full dataset using a single thread). I tried to use a Kaggle notebook but the APIs don't allow to update a dataset without deleting the current files.</p>\n<p>I uploaded the first 4 year of the dataset here: <a href=\"https://www.kaggle.com/datasets/abiolatti/climsim-lowres-complete-normalized\" target=\"_blank\">https://www.kaggle.com/datasets/abiolatti/climsim-lowres-complete-normalized</a></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2855295,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-06-04T17:44:43.343000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2857125,
              "author_name": "Juan D C F",
              "author_url": "",
              "post_date": "2024-06-05T17:52:39.107000",
              "content": "<p>Great, thanks as always for the insight. </p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2848432,
          "author_name": "Zhuoqun Li",
          "author_url": "",
          "post_date": "2024-06-01T03:12:12.733000",
          "content": "<p>Thank you, it's very helpful.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2847327,
      "author_name": "Urazalinov Baurzhan",
      "author_url": "",
      "post_date": "2024-05-31T14:00:30.923000",
      "content": "<p>I used 95% and 80% of the competition data to train my models. When using 80%, the LB score loses 0.0024 compared to 95%. I suggest using at least 70% of the competition data to train your models.</p>\n<p>To handle this size of data, I recommend using Kaggle TPUv3-8 or TPUv2-8 instead of GPU.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2847374,
          "author_name": "Zhuoqun Li",
          "author_url": "",
          "post_date": "2024-05-31T14:14:00.880000",
          "content": "<p>thank you for your sharing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2878353,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-06-18T22:23:38.820000",
      "content": "<p>Yeah, had the same doubt, especially as I am trying to run the model on my system, so can only work on a subset of data.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2847271": "Now that I'm only using about 30% of data for training, I wonder if using all data for training will help improve my score? Can anyone give me some advice?",
    "2847599": "Going from ~0.9M to ~9M samples in training (*10) give me between +0.06 (bad models) to +0.03 (best models) boost in score.\nSo yea, going from 30% to 100% would give you a boost. But not a HUGE one.",
    "2847327": "I used 95% and 80% of the competition data to train my models. When using 80%, the LB score loses 0.0024 compared to 95%. I suggest using at least 70% of the competition data to train your models.\n\nTo handle this size of data, I recommend using Kaggle TPUv3-8 or TPUv2-8 instead of GPU.",
    "2878353": "Yeah, had the same doubt, especially as I am trying to run the model on my system, so can only work on a subset of data."
  }
}