{
  "id": 528247,
  "title": "Do Efficient Data Pre-Processing!",
  "url": "/competitions/ariel-data-challenge-2024/discussion/528247",
  "author_name": "Pascal Pfeiffer",
  "post_date": "2024-08-15T11:44:31.809000",
  "votes": 38,
  "comment_count": 10,
  "views": 0,
  "content": "<p>In this competition, the ability to efficiently process and analyze the large amount of data is crucial. We deal with noisy data and need to pre-process the data for around 800 planets. Domain knowledge plays a central role here, as it allows us to identify and mitigate noise and enhance the quality of the data that we feed into our models. Thankfully, the hosts are kind to provide a starter notebook with that domain knowledge: <a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\">https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data</a></p>\n<p>For training, we can do all of this pre-processing once and the just save and reuse the engineered features, but for submissions, we also need to process the 800 planets in the submission kernel. Our submission time is capped at 9 hours (GPU and CPU kernels for test submissions!). This means we must ensure that our pre-processing pipeline is both effective and time-efficient. Even the seemingly simple task of loading this data from disk into RAM takes considerable time. For instance, loading only AIRS data for all 800 planets takes approximately 26 minutes on a single core. And further processing steps add up quickly.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F66ca9dc6449179b3df3ebffa86f96817%2Fpreprocessing_time.png?generation=1723726459932617&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>Task</th>\n<th>Time for AIRS [s]</th>\n<th>Time for FGS [s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Loading Data</td>\n<td>1.28</td>\n<td>0.45</td>\n</tr>\n<tr>\n<td>ADC Conversion</td>\n<td>0.34</td>\n<td>0.36</td>\n</tr>\n<tr>\n<td>Masking Hot/Dead Pixels</td>\n<td>3.13</td>\n<td>3.35</td>\n</tr>\n<tr>\n<td>Linear Correction</td>\n<td>19.4</td>\n<td>25.8</td>\n</tr>\n<tr>\n<td>Dark Cleaning</td>\n<td>1.44</td>\n<td>1.68</td>\n</tr>\n<tr>\n<td>Correlated Double Sampling (CDS)</td>\n<td>0.21</td>\n<td>0.28</td>\n</tr>\n<tr>\n<td>Transposing</td>\n<td>0.00 (negligible)</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>Flat Field Correction</td>\n<td>1.39</td>\n<td>1.47</td>\n</tr>\n<tr>\n<td>Feature Engineering</td>\n<td>0.27</td>\n<td>0.28</td>\n</tr>\n</tbody>\n</table>\n<p>In total, the preprocessing pipeline for AIRS data alone takes approximately 27.46 seconds per planet and per core, which translates to over 6 hours when processing all 800 planets on a single core. Adding FGS data, which takes even longer (33.67s) , this is already way above the 9h limit and we need to use multiple threads (can use 2 in the kernel without memory issues) or batching. Overall, with the naive approach, little runtime is left for actual modeling, so speeding up these pieces can bring large benefits depending on the modelling approach.</p>",
  "messages": [
    {
      "id": 2959839,
      "postDate": "2024-08-15T11:44:31.810Z",
      "content": "<p>In this competition, the ability to efficiently process and analyze the large amount of data is crucial. We deal with noisy data and need to pre-process the data for around 800 planets. Domain knowledge plays a central role here, as it allows us to identify and mitigate noise and enhance the quality of the data that we feed into our models. Thankfully, the hosts are kind to provide a starter notebook with that domain knowledge: <a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\">https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data</a></p>\n<p>For training, we can do all of this pre-processing once and the just save and reuse the engineered features, but for submissions, we also need to process the 800 planets in the submission kernel. Our submission time is capped at 9 hours (GPU and CPU kernels for test submissions!). This means we must ensure that our pre-processing pipeline is both effective and time-efficient. Even the seemingly simple task of loading this data from disk into RAM takes considerable time. For instance, loading only AIRS data for all 800 planets takes approximately 26 minutes on a single core. And further processing steps add up quickly.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F66ca9dc6449179b3df3ebffa86f96817%2Fpreprocessing_time.png?generation=1723726459932617&amp;alt=media\" alt=\"\"></p>\n<table>\n<thead>\n<tr>\n<th>Task</th>\n<th>Time for AIRS [s]</th>\n<th>Time for FGS [s]</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Loading Data</td>\n<td>1.28</td>\n<td>0.45</td>\n</tr>\n<tr>\n<td>ADC Conversion</td>\n<td>0.34</td>\n<td>0.36</td>\n</tr>\n<tr>\n<td>Masking Hot/Dead Pixels</td>\n<td>3.13</td>\n<td>3.35</td>\n</tr>\n<tr>\n<td>Linear Correction</td>\n<td>19.4</td>\n<td>25.8</td>\n</tr>\n<tr>\n<td>Dark Cleaning</td>\n<td>1.44</td>\n<td>1.68</td>\n</tr>\n<tr>\n<td>Correlated Double Sampling (CDS)</td>\n<td>0.21</td>\n<td>0.28</td>\n</tr>\n<tr>\n<td>Transposing</td>\n<td>0.00 (negligible)</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>Flat Field Correction</td>\n<td>1.39</td>\n<td>1.47</td>\n</tr>\n<tr>\n<td>Feature Engineering</td>\n<td>0.27</td>\n<td>0.28</td>\n</tr>\n</tbody>\n</table>\n<p>In total, the preprocessing pipeline for AIRS data alone takes approximately 27.46 seconds per planet and per core, which translates to over 6 hours when processing all 800 planets on a single core. Adding FGS data, which takes even longer (33.67s) , this is already way above the 9h limit and we need to use multiple threads (can use 2 in the kernel without memory issues) or batching. Overall, with the naive approach, little runtime is left for actual modeling, so speeding up these pieces can bring large benefits depending on the modelling approach.</p>",
      "rawMarkdown": "In this competition, the ability to efficiently process and analyze the large amount of data is crucial. We deal with noisy data and need to pre-process the data for around 800 planets. Domain knowledge plays a central role here, as it allows us to identify and mitigate noise and enhance the quality of the data that we feed into our models. Thankfully, the hosts are kind to provide a starter notebook with that domain knowledge: https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\n\nFor training, we can do all of this pre-processing once and the just save and reuse the engineered features, but for submissions, we also need to process the 800 planets in the submission kernel. Our submission time is capped at 9 hours (GPU and CPU kernels for test submissions!). This means we must ensure that our pre-processing pipeline is both effective and time-efficient. Even the seemingly simple task of loading this data from disk into RAM takes considerable time. For instance, loading only AIRS data for all 800 planets takes approximately 26 minutes on a single core. And further processing steps add up quickly.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F66ca9dc6449179b3df3ebffa86f96817%2Fpreprocessing_time.png?generation=1723726459932617&alt=media)\n\n| Task | Time for AIRS [s] | Time for FGS [s]\n| --- | --- | --- |\n| Loading Data | 1.28 | 0.45 |\n| ADC Conversion | 0.34 | 0.36 |\n| Masking Hot/Dead Pixels | 3.13 | 3.35 |\n| Linear Correction | 19.4 | 25.8 |\n| Dark Cleaning | 1.44 | 1.68 |\n| Correlated Double Sampling (CDS) | 0.21 | 0.28 |\n| Transposing | 0.00 (negligible) | 0.00 |\n| Flat Field Correction | 1.39 | 1.47 |\n| Feature Engineering | 0.27 | 0.28 |\n\n\nIn total, the preprocessing pipeline for AIRS data alone takes approximately 27.46 seconds per planet and per core, which translates to over 6 hours when processing all 800 planets on a single core. Adding FGS data, which takes even longer (33.67s) , this is already way above the 9h limit and we need to use multiple threads (can use 2 in the kernel without memory issues) or batching. Overall, with the naive approach, little runtime is left for actual modeling, so speeding up these pieces can bring large benefits depending on the modelling approach.",
      "votes": 38
    },
    {
      "id": 2967011,
      "postDate": "2024-08-22T12:56:05.113Z",
      "content": "<p>Update for preprocessing:</p>\n<table>\n<thead>\n<tr>\n<th>Task</th>\n<th>Time for AIRS [s]</th>\n<th>Time for FGS [s]</th>\n<th>Time for AIRS [s] UPDATE</th>\n<th>Time for FGS [s] UPDATE</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Loading Data</td>\n<td>1.28</td>\n<td>0.45</td>\n<td>1.27</td>\n<td>0.67</td>\n</tr>\n<tr>\n<td>ADC Conversion</td>\n<td>0.34</td>\n<td>0.36</td>\n<td>0.20</td>\n<td>0.22</td>\n</tr>\n<tr>\n<td>Masking Hot/Dead Pixels</td>\n<td>3.13</td>\n<td>3.35</td>\n<td>0.125</td>\n<td>0.134</td>\n</tr>\n<tr>\n<td>Linear Correction</td>\n<td>19.4</td>\n<td>25.8</td>\n<td>3.49 (1.96 GPU)</td>\n<td>5.61 (1.98 GPU)</td>\n</tr>\n<tr>\n<td>Dark Cleaning</td>\n<td>1.44</td>\n<td>1.68</td>\n<td>0.898</td>\n<td>0.957</td>\n</tr>\n<tr>\n<td>Correlated Double Sampling (CDS)</td>\n<td>0.21</td>\n<td>0.28</td>\n<td>0.189</td>\n<td>0.246</td>\n</tr>\n<tr>\n<td>Transposing</td>\n<td>0.00 (negligible)</td>\n<td>0.00</td>\n<td>0.00</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>Flat Field Correction</td>\n<td>1.39</td>\n<td>1.47</td>\n<td>1.38</td>\n<td>1.5</td>\n</tr>\n<tr>\n<td>Feature Engineering</td>\n<td>0.27</td>\n<td>0.28</td>\n<td>0.31</td>\n<td>0.34</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Ff954cfbb7f4e74d10d1bb3dc8e3c5a21%2Fpreprocessing_time2.png?generation=1724331293644455&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Update for preprocessing:\n\n| Task | Time for AIRS [s] | Time for FGS [s] | Time for AIRS [s] UPDATE | Time for FGS [s] UPDATE |\n| --- | --- | --- | --- | --- |\n| Loading Data | 1.28 | 0.45 | 1.27 | 0.67 |\n| ADC Conversion | 0.34 | 0.36 | 0.20 | 0.22 |\n| Masking Hot/Dead Pixels | 3.13 | 3.35 | 0.125 | 0.134 |\n| Linear Correction | 19.4 | 25.8 | 3.49 (1.96 GPU) | 5.61 (1.98 GPU) |\n| Dark Cleaning | 1.44 | 1.68 | 0.898 | 0.957 |\n| Correlated Double Sampling (CDS) | 0.21 | 0.28 | 0.189 | 0.246 |\n| Transposing | 0.00 (negligible) | 0.00 | 0.00 | 0.00 |\n| Flat Field Correction | 1.39 | 1.47 | 1.38 | 1.5 |\n| Feature Engineering | 0.27 | 0.28 | 0.31 | 0.34 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Ff954cfbb7f4e74d10d1bb3dc8e3c5a21%2Fpreprocessing_time2.png?generation=1724331293644455&alt=media)",
      "votes": 5
    },
    {
      "id": 2960003,
      "postDate": "2024-08-15T13:52:16.220Z",
      "content": "<p>Local test results for using linear correction vs not using it:</p>\n<pre><code>  (insample) score  linear correction\n  (insample) score  linear correction\n\n:\n   star  CV score  linear correction\n   star  CV score  linear correction\n</code></pre>\n<p>So, only a tiny drop. Though, again, might be different behavior for the unseen test solar system. Guess, I need to start subbing</p>",
      "rawMarkdown": "Local test results for using linear correction vs not using it:\n\n```\n0.486 - (insample) score **with** linear correction\n0.481 - (insample) score **without** linear correction\n\nupdate:\n0.455 - one star out CV score **with** linear correction\n0.449 - one star out CV score **without** linear correction\n```\n\nSo, only a tiny drop. Though, again, might be different behavior for the unseen test solar system. Guess, I need to start subbing",
      "votes": 6
    },
    {
      "id": 2966774,
      "postDate": "2024-08-22T07:40:58.610Z",
      "content": "<p>Your data cleaning process is meticulous and thorough—fantastic work!</p>",
      "rawMarkdown": "Your data cleaning process is meticulous and thorough—fantastic work!",
      "votes": 1
    },
    {
      "id": 2963261,
      "postDate": "2024-08-18T14:43:29.023Z",
      "content": "<p>I appreciate you providing such valuable insights. 😄<br>\nI'm also curious if there might be a way to reduce the linear correction time. </p>",
      "rawMarkdown": "I appreciate you providing such valuable insights. 😄\nI'm also curious if there might be a way to reduce the linear correction time. ",
      "votes": 1
    },
    {
      "id": 2961744,
      "postDate": "2024-08-16T21:21:56.497Z",
      "content": "<p>I did find that reducing the precision of the data files (float64 -&gt; float32) allowed me to use 3 cores (All of the preprocessing was done under 4 hours) without any memory issues. However, I have yet to run my model to check the loss in performance.  </p>",
      "rawMarkdown": "I did find that reducing the precision of the data files (float64 -> float32) allowed me to use 3 cores (All of the preprocessing was done under 4 hours) without any memory issues. However, I have yet to run my model to check the loss in performance.  ",
      "votes": 2,
      "replies": [
        {
          "id": 2961756,
          "postDate": "2024-08-16T21:32:09.827Z",
          "content": "<p>It actually even works with all four threads:<br>\n<a href=\"https://www.kaggle.com/code/ilu000/ariel24-data-prep\" target=\"_blank\">https://www.kaggle.com/code/ilu000/ariel24-data-prep</a></p>\n<p>Locally, the difference between float64 and float32 was tiny for me, but could become more important once we approach the limits of the signal extractions. Same as for linear correction, which might become more important. I assume that wasn't included in your 4h?</p>",
          "rawMarkdown": "It actually even works with all four threads:\nhttps://www.kaggle.com/code/ilu000/ariel24-data-prep\n\nLocally, the difference between float64 and float32 was tiny for me, but could become more important once we approach the limits of the signal extractions. Same as for linear correction, which might become more important. I assume that wasn't included in your 4h?",
          "votes": 2
        }
      ]
    },
    {
      "id": 2966773,
      "postDate": "2024-08-22T07:40:00.127Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 3024237,
      "postDate": "2024-10-21T12:56:48.177Z",
      "content": "<p>thanks for sharing!</p>",
      "rawMarkdown": "thanks for sharing!"
    },
    {
      "id": 2959991,
      "postDate": "2024-08-15T13:43:15.913Z",
      "content": "<p>thanks for the great tips</p>",
      "rawMarkdown": "thanks for the great tips\n"
    },
    {
      "id": 2959853,
      "postDate": "2024-08-15T11:56:20.760Z",
      "content": "<p>Useful tips! Thanks.</p>",
      "rawMarkdown": "Useful tips! Thanks."
    }
  ],
  "comments": [
    {
      "id": 2967011,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2024-08-22T12:56:05.113000",
      "content": "<p>Update for preprocessing:</p>\n<table>\n<thead>\n<tr>\n<th>Task</th>\n<th>Time for AIRS [s]</th>\n<th>Time for FGS [s]</th>\n<th>Time for AIRS [s] UPDATE</th>\n<th>Time for FGS [s] UPDATE</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Loading Data</td>\n<td>1.28</td>\n<td>0.45</td>\n<td>1.27</td>\n<td>0.67</td>\n</tr>\n<tr>\n<td>ADC Conversion</td>\n<td>0.34</td>\n<td>0.36</td>\n<td>0.20</td>\n<td>0.22</td>\n</tr>\n<tr>\n<td>Masking Hot/Dead Pixels</td>\n<td>3.13</td>\n<td>3.35</td>\n<td>0.125</td>\n<td>0.134</td>\n</tr>\n<tr>\n<td>Linear Correction</td>\n<td>19.4</td>\n<td>25.8</td>\n<td>3.49 (1.96 GPU)</td>\n<td>5.61 (1.98 GPU)</td>\n</tr>\n<tr>\n<td>Dark Cleaning</td>\n<td>1.44</td>\n<td>1.68</td>\n<td>0.898</td>\n<td>0.957</td>\n</tr>\n<tr>\n<td>Correlated Double Sampling (CDS)</td>\n<td>0.21</td>\n<td>0.28</td>\n<td>0.189</td>\n<td>0.246</td>\n</tr>\n<tr>\n<td>Transposing</td>\n<td>0.00 (negligible)</td>\n<td>0.00</td>\n<td>0.00</td>\n<td>0.00</td>\n</tr>\n<tr>\n<td>Flat Field Correction</td>\n<td>1.39</td>\n<td>1.47</td>\n<td>1.38</td>\n<td>1.5</td>\n</tr>\n<tr>\n<td>Feature Engineering</td>\n<td>0.27</td>\n<td>0.28</td>\n<td>0.31</td>\n<td>0.34</td>\n</tr>\n</tbody>\n</table>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Ff954cfbb7f4e74d10d1bb3dc8e3c5a21%2Fpreprocessing_time2.png?generation=1724331293644455&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 2960003,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2024-08-15T13:52:16.220000",
      "content": "<p>Local test results for using linear correction vs not using it:</p>\n<pre><code>  (insample) score  linear correction\n  (insample) score  linear correction\n\n:\n   star  CV score  linear correction\n   star  CV score  linear correction\n</code></pre>\n<p>So, only a tiny drop. Though, again, might be different behavior for the unseen test solar system. Guess, I need to start subbing</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 2966774,
      "author_name": "费文轩",
      "author_url": "",
      "post_date": "2024-08-22T07:40:58.610000",
      "content": "<p>Your data cleaning process is meticulous and thorough—fantastic work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2963261,
      "author_name": "WOOSUNG YOON",
      "author_url": "",
      "post_date": "2024-08-18T14:43:29.023000",
      "content": "<p>I appreciate you providing such valuable insights. 😄<br>\nI'm also curious if there might be a way to reduce the linear correction time. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2961744,
      "author_name": "Rahul Selvakumar",
      "author_url": "",
      "post_date": "2024-08-16T21:21:56.497000",
      "content": "<p>I did find that reducing the precision of the data files (float64 -&gt; float32) allowed me to use 3 cores (All of the preprocessing was done under 4 hours) without any memory issues. However, I have yet to run my model to check the loss in performance.  </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2961756,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2024-08-16T21:32:09.827000",
          "content": "<p>It actually even works with all four threads:<br>\n<a href=\"https://www.kaggle.com/code/ilu000/ariel24-data-prep\" target=\"_blank\">https://www.kaggle.com/code/ilu000/ariel24-data-prep</a></p>\n<p>Locally, the difference between float64 and float32 was tiny for me, but could become more important once we approach the limits of the signal extractions. Same as for linear correction, which might become more important. I assume that wasn't included in your 4h?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2966773,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-08-22T07:40:00.127000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3024237,
      "author_name": "Qixuan Sen",
      "author_url": "",
      "post_date": "2024-10-21T12:56:48.177000",
      "content": "<p>thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2959991,
      "author_name": "zj",
      "author_url": "",
      "post_date": "2024-08-15T13:43:15.913000",
      "content": "<p>thanks for the great tips</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2959853,
      "author_name": "KKY",
      "author_url": "",
      "post_date": "2024-08-15T11:56:20.760000",
      "content": "<p>Useful tips! Thanks.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2959839": "In this competition, the ability to efficiently process and analyze the large amount of data is crucial. We deal with noisy data and need to pre-process the data for around 800 planets. Domain knowledge plays a central role here, as it allows us to identify and mitigate noise and enhance the quality of the data that we feed into our models. Thankfully, the hosts are kind to provide a starter notebook with that domain knowledge: https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\n\nFor training, we can do all of this pre-processing once and the just save and reuse the engineered features, but for submissions, we also need to process the 800 planets in the submission kernel. Our submission time is capped at 9 hours (GPU and CPU kernels for test submissions!). This means we must ensure that our pre-processing pipeline is both effective and time-efficient. Even the seemingly simple task of loading this data from disk into RAM takes considerable time. For instance, loading only AIRS data for all 800 planets takes approximately 26 minutes on a single core. And further processing steps add up quickly.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F66ca9dc6449179b3df3ebffa86f96817%2Fpreprocessing_time.png?generation=1723726459932617&alt=media)\n\n| Task | Time for AIRS [s] | Time for FGS [s]\n| --- | --- | --- |\n| Loading Data | 1.28 | 0.45 |\n| ADC Conversion | 0.34 | 0.36 |\n| Masking Hot/Dead Pixels | 3.13 | 3.35 |\n| Linear Correction | 19.4 | 25.8 |\n| Dark Cleaning | 1.44 | 1.68 |\n| Correlated Double Sampling (CDS) | 0.21 | 0.28 |\n| Transposing | 0.00 (negligible) | 0.00 |\n| Flat Field Correction | 1.39 | 1.47 |\n| Feature Engineering | 0.27 | 0.28 |\n\n\nIn total, the preprocessing pipeline for AIRS data alone takes approximately 27.46 seconds per planet and per core, which translates to over 6 hours when processing all 800 planets on a single core. Adding FGS data, which takes even longer (33.67s) , this is already way above the 9h limit and we need to use multiple threads (can use 2 in the kernel without memory issues) or batching. Overall, with the naive approach, little runtime is left for actual modeling, so speeding up these pieces can bring large benefits depending on the modelling approach.",
    "2967011": "Update for preprocessing:\n\n| Task | Time for AIRS [s] | Time for FGS [s] | Time for AIRS [s] UPDATE | Time for FGS [s] UPDATE |\n| --- | --- | --- | --- | --- |\n| Loading Data | 1.28 | 0.45 | 1.27 | 0.67 |\n| ADC Conversion | 0.34 | 0.36 | 0.20 | 0.22 |\n| Masking Hot/Dead Pixels | 3.13 | 3.35 | 0.125 | 0.134 |\n| Linear Correction | 19.4 | 25.8 | 3.49 (1.96 GPU) | 5.61 (1.98 GPU) |\n| Dark Cleaning | 1.44 | 1.68 | 0.898 | 0.957 |\n| Correlated Double Sampling (CDS) | 0.21 | 0.28 | 0.189 | 0.246 |\n| Transposing | 0.00 (negligible) | 0.00 | 0.00 | 0.00 |\n| Flat Field Correction | 1.39 | 1.47 | 1.38 | 1.5 |\n| Feature Engineering | 0.27 | 0.28 | 0.31 | 0.34 |\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Ff954cfbb7f4e74d10d1bb3dc8e3c5a21%2Fpreprocessing_time2.png?generation=1724331293644455&alt=media)",
    "2960003": "Local test results for using linear correction vs not using it:\n\n```\n0.486 - (insample) score **with** linear correction\n0.481 - (insample) score **without** linear correction\n\nupdate:\n0.455 - one star out CV score **with** linear correction\n0.449 - one star out CV score **without** linear correction\n```\n\nSo, only a tiny drop. Though, again, might be different behavior for the unseen test solar system. Guess, I need to start subbing",
    "2966774": "Your data cleaning process is meticulous and thorough—fantastic work!",
    "2963261": "I appreciate you providing such valuable insights. 😄\nI'm also curious if there might be a way to reduce the linear correction time. ",
    "2961744": "I did find that reducing the precision of the data files (float64 -> float32) allowed me to use 3 cores (All of the preprocessing was done under 4 hours) without any memory issues. However, I have yet to run my model to check the loss in performance.  ",
    "2966773": "",
    "3024237": "thanks for sharing!",
    "2959991": "thanks for the great tips\n",
    "2959853": "Useful tips! Thanks."
  }
}