{
  "id": 523904,
  "title": "is Test set size : 200GB for 800 exoplanets ? [Host conformed]",
  "url": "/competitions/ariel-data-challenge-2024/discussion/523904",
  "author_name": "SeshuRaju 🧘‍♂️",
  "post_date": "2024-08-03T13:45:03.855000",
  "votes": 2,
  "comment_count": 10,
  "views": 0,
  "content": "<blockquote>\n  <p><strong>Each exoplanet size in training set is around ~ 256MB =&gt; for 673 exoplanets ~ 168GB</strong> !</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>is Test set is <strong>~ 200GB for 800 exoplanets</strong> ? <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> </p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>How to handle inference if it is 200GB</strong> -. This is the largest hidden test set size I've seen so far. (<strong>processing data one by one exoplanet took ~40mins to 60mins on submission</strong>)</p>\n</blockquote>",
  "messages": [
    {
      "id": 2947627,
      "postDate": "2024-08-05T10:30:49.047Z",
      "content": "<p>Yes I can confirm that the test data is huge. with only hundreds of examples we are already talking about 200GB of disk space. However, this is what the raw data looks like for many astronomical observations :/</p>",
      "rawMarkdown": "Yes I can confirm that the test data is huge. with only hundreds of examples we are already talking about 200GB of disk space. However, this is what the raw data looks like for many astronomical observations :/",
      "votes": 3,
      "replies": [
        {
          "id": 2947631,
          "postDate": "2024-08-05T10:32:25.820Z",
          "content": "<p>Thanks for the conformation about test set size <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> </p>",
          "rawMarkdown": "Thanks for the conformation about test set size @gordonyip ",
          "replies": [
            {
              "id": 2969747,
              "postDate": "2024-08-25T10:36:41.270Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2976779,
              "postDate": "2024-09-02T10:36:12.950Z",
              "content": "<p><a href=\"https://www.kaggle.com/edotirapuka\" target=\"_blank\">@edotirapuka</a> My library only takes about 1.2 hours to process all 800 samples<br>\n<a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/531453\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/531453</a></p>",
              "rawMarkdown": "@edotirapuka My library only takes about 1.2 hours to process all 800 samples\nhttps://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/531453\n",
              "votes": 3
            },
            {
              "id": 2976914,
              "postDate": "2024-09-02T12:35:57.853Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            },
            {
              "id": 2976989,
              "postDate": "2024-09-02T13:42:59.957Z",
              "content": "<p><a href=\"https://www.kaggle.com/edotirapuka\" target=\"_blank\">@edotirapuka</a> My library simply follow all the steps from the host calibration notebook </p>",
              "rawMarkdown": "@edotirapuka My library simply follow all the steps from the host calibration notebook ",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2945509,
      "postDate": "2024-08-03T13:45:03.857Z",
      "content": "<blockquote>\n  <p><strong>Each exoplanet size in training set is around ~ 256MB =&gt; for 673 exoplanets ~ 168GB</strong> !</p>\n</blockquote>\n<hr>\n<blockquote>\n  <p>is Test set is <strong>~ 200GB for 800 exoplanets</strong> ? <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> </p>\n</blockquote>\n<hr>\n<blockquote>\n  <p><strong>How to handle inference if it is 200GB</strong> -. This is the largest hidden test set size I've seen so far. (<strong>processing data one by one exoplanet took ~40mins to 60mins on submission</strong>)</p>\n</blockquote>",
      "rawMarkdown": "> **Each exoplanet size in training set is around ~ 256MB => for 673 exoplanets ~ 168GB** !\n\n---\n\n> is Test set is **~ 200GB for 800 exoplanets** ? @gordonyip \n\n---\n\n> **How to handle inference if it is 200GB** -~~ load each exoplanet one by one or in batches, then predict the targets within 9 hours~~. This is the largest hidden test set size I've seen so far. (**processing data one by one exoplanet took ~40mins to 60mins on submission**)",
      "votes": 2
    },
    {
      "id": 2968996,
      "postDate": "2024-08-24T14:13:23.940Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2968997,
          "postDate": "2024-08-24T14:13:53.027Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2978975,
              "postDate": "2024-09-04T12:36:26.507Z",
              "content": "<p><a href=\"https://www.kaggle.com/edotirapuka\" target=\"_blank\">@edotirapuka</a> , sorry for the late reply. after running the notebook you should produce something similar to the sample_submission.csv, so it should have 1+ 283*2 columns, the first 283 columns are the mean prediction, and the next 283 are the corresponding uncertainty for each wavelength, the wavelength order follows train_label.csv</p>",
              "rawMarkdown": "@edotirapuka , sorry for the late reply. after running the notebook you should produce something similar to the sample_submission.csv, so it should have 1+ 283*2 columns, the first 283 columns are the mean prediction, and the next 283 are the corresponding uncertainty for each wavelength, the wavelength order follows train_label.csv",
              "votes": 2
            },
            {
              "id": 2979193,
              "postDate": "2024-09-04T15:53:49.197Z",
              "rawMarkdown": "",
              "votes": 1,
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2947627,
      "author_name": "Gordon Yip",
      "author_url": "",
      "post_date": "2024-08-05T10:30:49.047000",
      "content": "<p>Yes I can confirm that the test data is huge. with only hundreds of examples we are already talking about 200GB of disk space. However, this is what the raw data looks like for many astronomical observations :/</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2947631,
          "author_name": "SeshuRaju 🧘‍♂️",
          "author_url": "",
          "post_date": "2024-08-05T10:32:25.820000",
          "content": "<p>Thanks for the conformation about test set size <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2969747,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-08-25T10:36:41.270000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2976779,
              "author_name": "ChingYinNg",
              "author_url": "",
              "post_date": "2024-09-02T10:36:12.950000",
              "content": "<p><a href=\"https://www.kaggle.com/edotirapuka\" target=\"_blank\">@edotirapuka</a> My library only takes about 1.2 hours to process all 800 samples<br>\n<a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/531453\" target=\"_blank\">https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/531453</a></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2976914,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-09-02T12:35:57.853000",
              "content": "",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2976989,
              "author_name": "ChingYinNg",
              "author_url": "",
              "post_date": "2024-09-02T13:42:59.957000",
              "content": "<p><a href=\"https://www.kaggle.com/edotirapuka\" target=\"_blank\">@edotirapuka</a> My library simply follow all the steps from the host calibration notebook </p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2968996,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-08-24T14:13:23.940000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2968997,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-08-24T14:13:53.027000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2978975,
              "author_name": "Gordon Yip",
              "author_url": "",
              "post_date": "2024-09-04T12:36:26.507000",
              "content": "<p><a href=\"https://www.kaggle.com/edotirapuka\" target=\"_blank\">@edotirapuka</a> , sorry for the late reply. after running the notebook you should produce something similar to the sample_submission.csv, so it should have 1+ 283*2 columns, the first 283 columns are the mean prediction, and the next 283 are the corresponding uncertainty for each wavelength, the wavelength order follows train_label.csv</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2979193,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-09-04T15:53:49.197000",
              "content": "",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2947627": "Yes I can confirm that the test data is huge. with only hundreds of examples we are already talking about 200GB of disk space. However, this is what the raw data looks like for many astronomical observations :/",
    "2945509": "> **Each exoplanet size in training set is around ~ 256MB => for 673 exoplanets ~ 168GB** !\n\n---\n\n> is Test set is **~ 200GB for 800 exoplanets** ? @gordonyip \n\n---\n\n> **How to handle inference if it is 200GB** -~~ load each exoplanet one by one or in batches, then predict the targets within 9 hours~~. This is the largest hidden test set size I've seen so far. (**processing data one by one exoplanet took ~40mins to 60mins on submission**)",
    "2968996": ""
  }
}