{
  "id": 531453,
  "title": "Fast calibration library with C and multiprocessing",
  "url": "/competitions/ariel-data-challenge-2024/discussion/531453",
  "author_name": "ChingYinNg",
  "post_date": "2024-09-01T11:58:38.145000",
  "votes": 17,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Here is my calibration library in C (Only CPU is needed):<br>\n<a href=\"https://www.kaggle.com/datasets/chingyinng/adc-calibration-library\" target=\"_blank\">https://www.kaggle.com/datasets/chingyinng/adc-calibration-library</a></p>\n<p>Benchmark for the training dataset: 4926 s</p>\n<p>Things I did to speed up the calibration:</p>\n<ul>\n<li>Used polars instead of pandas to speed up the <code>read_parquet</code> function</li>\n<li>Wrote the computational layer in C with -O3 flag</li>\n<li>Multiprocessing</li>\n</ul>\n<p>I followed the calibration notebook from host (<a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\">https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data</a>)</p>\n<ul>\n<li>All steps including linearity correction are enabled</li>\n<li>For dead or hot pixels, I simply did an averaging from the nearby four pixels</li>\n<li>I have clipped the negative values to zero</li>\n</ul>\n<p>If you would like to use this library, please give reference to this post or tag me, thanks.</p>\n<p>Sample usage:</p>\n<pre><code>import sys\n\nsys.path.append()\n calibrating import Calibrator\n\ncalibrator = Calibrator(\n    =,    # : training dataset; : test dataset\n    =,\n    =,\n    =,\n    =8,\n    =25,    #value\n    # =8,    # Use this option  you only want  calibrate a few files  testing\n)\ncalibrator.calibrate()\ncalibrator.concatenate_files()\ndel calibrator\n\n\ndata_test_airs = np.load(OUTPUT_DIR / )\ndata_test_fgs = np.load(OUTPUT_DIR / )\n</code></pre>",
  "messages": [
    {
      "id": 2975967,
      "postDate": "2024-09-01T11:58:38.147Z",
      "content": "<p>Here is my calibration library in C (Only CPU is needed):<br>\n<a href=\"https://www.kaggle.com/datasets/chingyinng/adc-calibration-library\" target=\"_blank\">https://www.kaggle.com/datasets/chingyinng/adc-calibration-library</a></p>\n<p>Benchmark for the training dataset: 4926 s</p>\n<p>Things I did to speed up the calibration:</p>\n<ul>\n<li>Used polars instead of pandas to speed up the <code>read_parquet</code> function</li>\n<li>Wrote the computational layer in C with -O3 flag</li>\n<li>Multiprocessing</li>\n</ul>\n<p>I followed the calibration notebook from host (<a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\">https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data</a>)</p>\n<ul>\n<li>All steps including linearity correction are enabled</li>\n<li>For dead or hot pixels, I simply did an averaging from the nearby four pixels</li>\n<li>I have clipped the negative values to zero</li>\n</ul>\n<p>If you would like to use this library, please give reference to this post or tag me, thanks.</p>\n<p>Sample usage:</p>\n<pre><code>import sys\n\nsys.path.append()\n calibrating import Calibrator\n\ncalibrator = Calibrator(\n    =,    # : training dataset; : test dataset\n    =,\n    =,\n    =,\n    =8,\n    =25,    #value\n    # =8,    # Use this option  you only want  calibrate a few files  testing\n)\ncalibrator.calibrate()\ncalibrator.concatenate_files()\ndel calibrator\n\n\ndata_test_airs = np.load(OUTPUT_DIR / )\ndata_test_fgs = np.load(OUTPUT_DIR / )\n</code></pre>",
      "rawMarkdown": "Here is my calibration library in C (Only CPU is needed):\nhttps://www.kaggle.com/datasets/chingyinng/adc-calibration-library\n\nBenchmark for the training dataset: 4926 s\n\nThings I did to speed up the calibration:\n* Used polars instead of pandas to speed up the `read_parquet` function\n* Wrote the computational layer in C with -O3 flag\n* Multiprocessing\n\nI followed the calibration notebook from host (https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data)\n* All steps including linearity correction are enabled\n* For dead or hot pixels, I simply did an averaging from the nearby four pixels\n* I have clipped the negative values to zero\n\nIf you would like to use this library, please give reference to this post or tag me, thanks.\n\nSample usage:\n```\nimport sys\n\nsys.path.append(\"/kaggle/input/adc-calibration-library\")\nfrom calibrating import Calibrator\n\ncalibrator = Calibrator(\n    is_test=True,    # False: training dataset; True: test dataset\n    data_dir=\"/kaggle/input/ariel-data-challenge-2024\",\n    output_dir=\"output\",\n    c_lib_path=\"/kaggle/input/adc-calibration-library/c_lib.so\",\n    num_workers=8,\n    time_binning_freq=25,    # Default value\n    # first_n_files=8,    # Use this option if you only want to calibrate a few files for testing\n)\ncalibrator.calibrate()\ncalibrator.concatenate_files()\ndel calibrator\n\n\ndata_test_airs = np.load(OUTPUT_DIR / \"data_test.npy\")\ndata_test_fgs = np.load(OUTPUT_DIR / \"data_test_FGS.npy\")\n```",
      "votes": 17
    },
    {
      "id": 2975978,
      "postDate": "2024-09-01T12:14:28.283Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chingyinng\" target=\"_blank\">@chingyinng</a>, can you publish the source code? It would help reproducibility.</p>",
      "rawMarkdown": "Hi @chingyinng, can you publish the source code? It would help reproducibility.",
      "votes": 1,
      "replies": [
        {
          "id": 2975982,
          "postDate": "2024-09-01T12:19:53.960Z",
          "content": "<p>Sure! I have updated the dataset with the source files.</p>",
          "rawMarkdown": "Sure! I have updated the dataset with the source files.",
          "votes": 4
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2975978,
      "author_name": "AmbrosM",
      "author_url": "",
      "post_date": "2024-09-01T12:14:28.283000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chingyinng\" target=\"_blank\">@chingyinng</a>, can you publish the source code? It would help reproducibility.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2975982,
          "author_name": "ChingYinNg",
          "author_url": "",
          "post_date": "2024-09-01T12:19:53.960000",
          "content": "<p>Sure! I have updated the dataset with the source files.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2975967": "Here is my calibration library in C (Only CPU is needed):\nhttps://www.kaggle.com/datasets/chingyinng/adc-calibration-library\n\nBenchmark for the training dataset: 4926 s\n\nThings I did to speed up the calibration:\n* Used polars instead of pandas to speed up the `read_parquet` function\n* Wrote the computational layer in C with -O3 flag\n* Multiprocessing\n\nI followed the calibration notebook from host (https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data)\n* All steps including linearity correction are enabled\n* For dead or hot pixels, I simply did an averaging from the nearby four pixels\n* I have clipped the negative values to zero\n\nIf you would like to use this library, please give reference to this post or tag me, thanks.\n\nSample usage:\n```\nimport sys\n\nsys.path.append(\"/kaggle/input/adc-calibration-library\")\nfrom calibrating import Calibrator\n\ncalibrator = Calibrator(\n    is_test=True,    # False: training dataset; True: test dataset\n    data_dir=\"/kaggle/input/ariel-data-challenge-2024\",\n    output_dir=\"output\",\n    c_lib_path=\"/kaggle/input/adc-calibration-library/c_lib.so\",\n    num_workers=8,\n    time_binning_freq=25,    # Default value\n    # first_n_files=8,    # Use this option if you only want to calibrate a few files for testing\n)\ncalibrator.calibrate()\ncalibrator.concatenate_files()\ndel calibrator\n\n\ndata_test_airs = np.load(OUTPUT_DIR / \"data_test.npy\")\ndata_test_fgs = np.load(OUTPUT_DIR / \"data_test_FGS.npy\")\n```",
    "2975978": "Hi @chingyinng, can you publish the source code? It would help reproducibility."
  }
}