{
  "id": 529663,
  "title": "Managing with data sixe",
  "url": "/competitions/ariel-data-challenge-2024/discussion/529663",
  "author_name": "v_parth7",
  "post_date": "2024-08-22T07:04:46.187000",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hey I am stuck on a problem that the dataset size is too large, how can I run it on my PC or even of Kaggle Notebooks or Google Collab.<br>\nCan anyone please suggest me how to deal with such a problem?<br>\nThanks</p>",
  "messages": [
    {
      "id": 2966772,
      "postDate": "2024-08-22T07:39:59.993Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vparh7\" target=\"_blank\">@vparh7</a> Many public notebooks show how to do it. It all boils down to two principles:</p>\n<ol>\n<li>Assuming that you are interested in the star's brightness rather than its shape, you can aggregate every 32*32 pixel image to a single float32 number.</li>\n<li>If the planet's transit takes two hours, you don't need data at a 0.1 second resolution. You can downsample the FGS1 time series by a factor of 300 and the AIRS-CH0 time series by a factor of 25.</li>\n</ol>\n<p>With these reductions I get a 673*225 array for FGS1 and a 673*225*356 array for AIRS-CH0. These arrays can be processed on every PC.</p>",
      "rawMarkdown": "Hi @vparh7 Many public notebooks show how to do it. It all boils down to two principles:\n1. Assuming that you are interested in the star's brightness rather than its shape, you can aggregate every 32\\*32 pixel image to a single float32 number.\n2. If the planet's transit takes two hours, you don't need data at a 0.1 second resolution. You can downsample the FGS1 time series by a factor of 300 and the AIRS-CH0 time series by a factor of 25.\n\nWith these reductions I get a 673\\*225 array for FGS1 and a 673\\*225\\*356 array for AIRS-CH0. These arrays can be processed on every PC.",
      "votes": 1
    },
    {
      "id": 2966742,
      "postDate": "2024-08-22T07:04:46.187Z",
      "content": "<p>Hey I am stuck on a problem that the dataset size is too large, how can I run it on my PC or even of Kaggle Notebooks or Google Collab.<br>\nCan anyone please suggest me how to deal with such a problem?<br>\nThanks</p>",
      "rawMarkdown": "Hey I am stuck on a problem that the dataset size is too large, how can I run it on my PC or even of Kaggle Notebooks or Google Collab.\nCan anyone please suggest me how to deal with such a problem?\nThanks",
      "votes": 1
    },
    {
      "id": 2966756,
      "postDate": "2024-08-22T07:27:33.283Z",
      "content": "<p>Use binning as shown in the pre-processing script by organizers</p>",
      "rawMarkdown": "Use binning as shown in the pre-processing script by organizers"
    },
    {
      "id": 2966748,
      "postDate": "2024-08-22T07:12:23.737Z",
      "content": "<p>The large dataset can be intimidating, where exactly are you experiencing issues, so we can try to help?</p>\n<p>Other kagglers have demonstrated how to successfully work with that data in kaggle notebooks, check out the <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/code\" target=\"_blank\">Code</a> Section. </p>",
      "rawMarkdown": "The large dataset can be intimidating, where exactly are you experiencing issues, so we can try to help?\n\nOther kagglers have demonstrated how to successfully work with that data in kaggle notebooks, check out the [Code](https://www.kaggle.com/competitions/ariel-data-challenge-2024/code) Section. "
    }
  ],
  "comments": [
    {
      "id": 2966772,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2024-08-22T07:39:59.993000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vparh7\" target=\"_blank\">@vparh7</a> Many public notebooks show how to do it. It all boils down to two principles:</p>\n<ol>\n<li>Assuming that you are interested in the star's brightness rather than its shape, you can aggregate every 32*32 pixel image to a single float32 number.</li>\n<li>If the planet's transit takes two hours, you don't need data at a 0.1 second resolution. You can downsample the FGS1 time series by a factor of 300 and the AIRS-CH0 time series by a factor of 25.</li>\n</ol>\n<p>With these reductions I get a 673*225 array for FGS1 and a 673*225*356 array for AIRS-CH0. These arrays can be processed on every PC.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2966756,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-08-22T07:27:33.283000",
      "content": "<p>Use binning as shown in the pre-processing script by organizers</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2966748,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2024-08-22T07:12:23.737000",
      "content": "<p>The large dataset can be intimidating, where exactly are you experiencing issues, so we can try to help?</p>\n<p>Other kagglers have demonstrated how to successfully work with that data in kaggle notebooks, check out the <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/code\" target=\"_blank\">Code</a> Section. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2966772": "Hi @vparh7 Many public notebooks show how to do it. It all boils down to two principles:\n1. Assuming that you are interested in the star's brightness rather than its shape, you can aggregate every 32\\*32 pixel image to a single float32 number.\n2. If the planet's transit takes two hours, you don't need data at a 0.1 second resolution. You can downsample the FGS1 time series by a factor of 300 and the AIRS-CH0 time series by a factor of 25.\n\nWith these reductions I get a 673\\*225 array for FGS1 and a 673\\*225\\*356 array for AIRS-CH0. These arrays can be processed on every PC.",
    "2966742": "Hey I am stuck on a problem that the dataset size is too large, how can I run it on my PC or even of Kaggle Notebooks or Google Collab.\nCan anyone please suggest me how to deal with such a problem?\nThanks",
    "2966756": "Use binning as shown in the pre-processing script by organizers",
    "2966748": "The large dataset can be intimidating, where exactly are you experiencing issues, so we can try to help?\n\nOther kagglers have demonstrated how to successfully work with that data in kaggle notebooks, check out the [Code](https://www.kaggle.com/competitions/ariel-data-challenge-2024/code) Section. "
  }
}