{
  "id": 524153,
  "title": "ARIEL data processing pipeline",
  "url": "/competitions/ariel-data-challenge-2024/discussion/524153",
  "author_name": "DennisSakva",
  "post_date": "2024-08-04T21:47:50.455000",
  "votes": 16,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi everyone,<br>\nI'm trying to find out the correct procedure to produce calibrated images based on ARIEL red book (<a href=\"https://sci.esa.int/documents/34022/36216/Ariel_Definition_Study_Report_2020.pdf)\" target=\"_blank\">https://sci.esa.int/documents/34022/36216/Ariel_Definition_Study_Report_2020.pdf)</a>. Here are my thoughts. What do you think?</p>\n<ol>\n<li>We are given a signal in ADU. We need to convert it to electrons by multiplying by gain and adding offset</li>\n<li>Subtract readout noise (Subtract zeroth read?)</li>\n<li>Non-linearity correction. There are no details in the red book. But HST calibration (<a href=\"https://hst-docs.stsci.edu/wfc3dhb/chapter-3-wfc3-data-calibration/3-3-ir-data-calibration-steps#id-3.3IRDataCalibrationSteps-3.3.2IRZero-ReadSignalCorrection\" target=\"_blank\">https://hst-docs.stsci.edu/wfc3dhb/chapter-3-wfc3-data-calibration/3-3-ir-data-calibration-steps#id-3.3IRDataCalibrationSteps-3.3.2IRZero-ReadSignalCorrection</a>) is Fc=F<em>(1+C1+C2</em>F+C3<em>F^2+c4</em>F^3) but in our case there are 6 coefficients</li>\n<li>Subtract dark current</li>\n<li>Divide by flat field</li>\n</ol>\n<p>Does this pipeline looks reasonable to you?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F01f1f97ba14205a2a30c26e8a1e0ae4b%2FAriel1.png?generation=1722808009657800&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 2947006,
      "postDate": "2024-08-04T21:47:50.457Z",
      "content": "<p>Hi everyone,<br>\nI'm trying to find out the correct procedure to produce calibrated images based on ARIEL red book (<a href=\"https://sci.esa.int/documents/34022/36216/Ariel_Definition_Study_Report_2020.pdf)\" target=\"_blank\">https://sci.esa.int/documents/34022/36216/Ariel_Definition_Study_Report_2020.pdf)</a>. Here are my thoughts. What do you think?</p>\n<ol>\n<li>We are given a signal in ADU. We need to convert it to electrons by multiplying by gain and adding offset</li>\n<li>Subtract readout noise (Subtract zeroth read?)</li>\n<li>Non-linearity correction. There are no details in the red book. But HST calibration (<a href=\"https://hst-docs.stsci.edu/wfc3dhb/chapter-3-wfc3-data-calibration/3-3-ir-data-calibration-steps#id-3.3IRDataCalibrationSteps-3.3.2IRZero-ReadSignalCorrection\" target=\"_blank\">https://hst-docs.stsci.edu/wfc3dhb/chapter-3-wfc3-data-calibration/3-3-ir-data-calibration-steps#id-3.3IRDataCalibrationSteps-3.3.2IRZero-ReadSignalCorrection</a>) is Fc=F<em>(1+C1+C2</em>F+C3<em>F^2+c4</em>F^3) but in our case there are 6 coefficients</li>\n<li>Subtract dark current</li>\n<li>Divide by flat field</li>\n</ol>\n<p>Does this pipeline looks reasonable to you?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F01f1f97ba14205a2a30c26e8a1e0ae4b%2FAriel1.png?generation=1722808009657800&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi everyone,\nI'm trying to find out the correct procedure to produce calibrated images based on ARIEL red book (https://sci.esa.int/documents/34022/36216/Ariel_Definition_Study_Report_2020.pdf). Here are my thoughts. What do you think?\n1. We are given a signal in ADU. We need to convert it to electrons by multiplying by gain and adding offset\n2. Subtract readout noise (Subtract zeroth read?)\n3. Non-linearity correction. There are no details in the red book. But HST calibration (https://hst-docs.stsci.edu/wfc3dhb/chapter-3-wfc3-data-calibration/3-3-ir-data-calibration-steps#id-3.3IRDataCalibrationSteps-3.3.2IRZero-ReadSignalCorrection) is Fc=F*(1+C1+C2*F+C3*F^2+c4*F^3) but in our case there are 6 coefficients\n4. Subtract dark current\n5. Divide by flat field\n\nDoes this pipeline looks reasonable to you?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F01f1f97ba14205a2a30c26e8a1e0ae4b%2FAriel1.png?generation=1722808009657800&alt=media)\n",
      "votes": 16
    },
    {
      "id": 2947626,
      "postDate": "2024-08-05T10:22:18.643Z",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> for the heroic effort! We will be publishing the pipeline soon so that you can cross check your pipeline with ours, again, the calibration procedures are not perfect and they are prone to non-linear effects (jitter etc) in the data. so please do not take our pipeline as the ONLY pipeline. </p>",
      "rawMarkdown": "Thank you @sakvaua for the heroic effort! We will be publishing the pipeline soon so that you can cross check your pipeline with ours, again, the calibration procedures are not perfect and they are prone to non-linear effects (jitter etc) in the data. so please do not take our pipeline as the ONLY pipeline. ",
      "votes": 3
    },
    {
      "id": 2947155,
      "postDate": "2024-08-05T02:39:50.760Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/jnesbit6/sensor-noise-correction\" target=\"_blank\">Notebook - Sensor Noise Correction</a> <a href=\"https://www.kaggle.com/jnesbit6\" target=\"_blank\">@jnesbit6</a> shared public notebook</p>\n<p><a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> dead cells remove logic missing</p>\n<pre><code>    signal = signal.to_numpy().reshape(sensor_sizes_dict[sensor][])\n\n    \n    signal = signal - dark_frame.to_numpy().reshape(sensor_sizes_dict[sensor][])\n    signal = signal - read_frame.to_numpy().reshape(sensor_sizes_dict[sensor][])\n\n    \n    flat = flat_frame.to_numpy().reshape(sensor_sizes_dict[sensor][])\n    flat[dead_frame.to_numpy().reshape(sensor_sizes_dict[sensor][])] = \n    signal = signal/flat\n\n    \n    coefficients = linear_corr_frame.to_numpy().reshape([] + sensor_sizes_dict[sensor][])\n    coefficients = np.repeat(coefficients, sensor_sizes_dict[sensor][][], axis=)\n    corrected = coefficients[]\n     i  ():\n        corrected += np.multiply(np.power(signal, i+),coefficients[i+])\n\n    \n    corrected = corrected*planet_gain_offset[sensor + ].values + planet_gain_offset[sensor + ].values\n\n    \n    corrected[np.repeat(dead_frame.to_numpy().reshape(sensor_sizes_dict[sensor][]), sensor_sizes_dict[sensor][][], axis=)] = \n</code></pre>",
      "rawMarkdown": "[Notebook - Sensor Noise Correction](https://www.kaggle.com/code/jnesbit6/sensor-noise-correction) @jnesbit6 shared public notebook\n\n@sakvaua dead cells remove logic missing\n\n```python\n    signal = signal.to_numpy().reshape(sensor_sizes_dict[sensor][0])\n\n    # read and dark frame correction, just subtraction\n    signal = signal - dark_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1])\n    signal = signal - read_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1])\n\n    # flat frame correction + ensure dead pixels do not disrupt. We divide by the flat frame\n    flat = flat_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1])\n    flat[dead_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1])] = 1\n    signal = signal/flat\n\n    # linear correction - treat as sixth degree polynomial, first number is C, not coefficient?\n    coefficients = linear_corr_frame.to_numpy().reshape([6] + sensor_sizes_dict[sensor][1])\n    coefficients = np.repeat(coefficients, sensor_sizes_dict[sensor][0][0], axis=1)\n    corrected = coefficients[0]\n    for i in range(5):\n        corrected += np.multiply(np.power(signal, i+1),coefficients[i+1])\n        \n    # gain, offset\n    corrected = corrected*planet_gain_offset[sensor + \"_adc_gain\"].values + planet_gain_offset[sensor + \"_adc_offset\"].values\n        \n    # zero out malfunctioning pixels\n    corrected[np.repeat(dead_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1]), sensor_sizes_dict[sensor][0][0], axis=0)] = 0\n```",
      "votes": 3,
      "replies": [
        {
          "id": 2947384,
          "postDate": "2024-08-05T06:45:47.357Z",
          "content": "<p>Yeah, I forgot to mention dead pixels. I'm unsure if the order of calibration steps in this code snippet is correct.</p>",
          "rawMarkdown": "Yeah, I forgot to mention dead pixels. I'm unsure if the order of calibration steps in this code snippet is correct.",
          "votes": 1,
          "replies": [
            {
              "id": 2947395,
              "postDate": "2024-08-05T06:51:03.340Z",
              "content": "<p>Thanks, now i understood what you mean by order of calibrations.</p>",
              "rawMarkdown": "Thanks, now i understood what you mean by order of calibrations."
            }
          ]
        }
      ]
    },
    {
      "id": 2947042,
      "postDate": "2024-08-04T22:54:47.257Z",
      "content": "<p>Do hosts truly expect us to reverse-engineering SOC@ESA pre-processing pipeline?<br>\nI think more specification should be needed on how to handle calibration data (i.e. Level1 ~ Level1.5 product).</p>\n<p>E.g.</p>\n<ul>\n<li>what linear-correction coefficient mean?</li>\n</ul>",
      "rawMarkdown": "Do hosts truly expect us to reverse-engineering SOC@ESA pre-processing pipeline?\nI think more specification should be needed on how to handle calibration data (i.e. Level1 ~ Level1.5 product).\n\nE.g.\n\n* what linear-correction coefficient mean?",
      "votes": 2,
      "replies": [
        {
          "id": 2947422,
          "postDate": "2024-08-05T07:07:21.880Z",
          "content": "<p>Also, even and odd images from the AIRS-CH0 instrument have different exposures yet only one dark frame for calibration.</p>",
          "rawMarkdown": "Also, even and odd images from the AIRS-CH0 instrument have different exposures yet only one dark frame for calibration."
        },
        {
          "id": 2947623,
          "postDate": "2024-08-05T10:19:49.893Z",
          "content": "<p>No no no not at all! We are doing some final checks before we push out the calibration notebook to everyone :) </p>\n<p>We understand not everyone knows calibration, and it will be a BIG BIG ask for people to know</p>",
          "rawMarkdown": "No no no not at all! We are doing some final checks before we push out the calibration notebook to everyone :) \n\nWe understand not everyone knows calibration, and it will be a BIG BIG ask for people to know",
          "votes": 6
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2947626,
      "author_name": "Gordon Yip",
      "author_url": "",
      "post_date": "2024-08-05T10:22:18.643000",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> for the heroic effort! We will be publishing the pipeline soon so that you can cross check your pipeline with ours, again, the calibration procedures are not perfect and they are prone to non-linear effects (jitter etc) in the data. so please do not take our pipeline as the ONLY pipeline. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 2947155,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-08-05T02:39:50.760000",
      "content": "<p><a href=\"https://www.kaggle.com/code/jnesbit6/sensor-noise-correction\" target=\"_blank\">Notebook - Sensor Noise Correction</a> <a href=\"https://www.kaggle.com/jnesbit6\" target=\"_blank\">@jnesbit6</a> shared public notebook</p>\n<p><a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> dead cells remove logic missing</p>\n<pre><code>    signal = signal.to_numpy().reshape(sensor_sizes_dict[sensor][])\n\n    \n    signal = signal - dark_frame.to_numpy().reshape(sensor_sizes_dict[sensor][])\n    signal = signal - read_frame.to_numpy().reshape(sensor_sizes_dict[sensor][])\n\n    \n    flat = flat_frame.to_numpy().reshape(sensor_sizes_dict[sensor][])\n    flat[dead_frame.to_numpy().reshape(sensor_sizes_dict[sensor][])] = \n    signal = signal/flat\n\n    \n    coefficients = linear_corr_frame.to_numpy().reshape([] + sensor_sizes_dict[sensor][])\n    coefficients = np.repeat(coefficients, sensor_sizes_dict[sensor][][], axis=)\n    corrected = coefficients[]\n     i  ():\n        corrected += np.multiply(np.power(signal, i+),coefficients[i+])\n\n    \n    corrected = corrected*planet_gain_offset[sensor + ].values + planet_gain_offset[sensor + ].values\n\n    \n    corrected[np.repeat(dead_frame.to_numpy().reshape(sensor_sizes_dict[sensor][]), sensor_sizes_dict[sensor][][], axis=)] = \n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2947384,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-08-05T06:45:47.357000",
          "content": "<p>Yeah, I forgot to mention dead pixels. I'm unsure if the order of calibration steps in this code snippet is correct.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2947395,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2024-08-05T06:51:03.340000",
              "content": "<p>Thanks, now i understood what you mean by order of calibrations.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2947042,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-08-04T22:54:47.257000",
      "content": "<p>Do hosts truly expect us to reverse-engineering SOC@ESA pre-processing pipeline?<br>\nI think more specification should be needed on how to handle calibration data (i.e. Level1 ~ Level1.5 product).</p>\n<p>E.g.</p>\n<ul>\n<li>what linear-correction coefficient mean?</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 2947422,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-08-05T07:07:21.880000",
          "content": "<p>Also, even and odd images from the AIRS-CH0 instrument have different exposures yet only one dark frame for calibration.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2947623,
          "author_name": "Gordon Yip",
          "author_url": "",
          "post_date": "2024-08-05T10:19:49.893000",
          "content": "<p>No no no not at all! We are doing some final checks before we push out the calibration notebook to everyone :) </p>\n<p>We understand not everyone knows calibration, and it will be a BIG BIG ask for people to know</p>",
          "votes": 6,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2947006": "Hi everyone,\nI'm trying to find out the correct procedure to produce calibrated images based on ARIEL red book (https://sci.esa.int/documents/34022/36216/Ariel_Definition_Study_Report_2020.pdf). Here are my thoughts. What do you think?\n1. We are given a signal in ADU. We need to convert it to electrons by multiplying by gain and adding offset\n2. Subtract readout noise (Subtract zeroth read?)\n3. Non-linearity correction. There are no details in the red book. But HST calibration (https://hst-docs.stsci.edu/wfc3dhb/chapter-3-wfc3-data-calibration/3-3-ir-data-calibration-steps#id-3.3IRDataCalibrationSteps-3.3.2IRZero-ReadSignalCorrection) is Fc=F*(1+C1+C2*F+C3*F^2+c4*F^3) but in our case there are 6 coefficients\n4. Subtract dark current\n5. Divide by flat field\n\nDoes this pipeline looks reasonable to you?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F349155%2F01f1f97ba14205a2a30c26e8a1e0ae4b%2FAriel1.png?generation=1722808009657800&alt=media)\n",
    "2947626": "Thank you @sakvaua for the heroic effort! We will be publishing the pipeline soon so that you can cross check your pipeline with ours, again, the calibration procedures are not perfect and they are prone to non-linear effects (jitter etc) in the data. so please do not take our pipeline as the ONLY pipeline. ",
    "2947155": "[Notebook - Sensor Noise Correction](https://www.kaggle.com/code/jnesbit6/sensor-noise-correction) @jnesbit6 shared public notebook\n\n@sakvaua dead cells remove logic missing\n\n```python\n    signal = signal.to_numpy().reshape(sensor_sizes_dict[sensor][0])\n\n    # read and dark frame correction, just subtraction\n    signal = signal - dark_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1])\n    signal = signal - read_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1])\n\n    # flat frame correction + ensure dead pixels do not disrupt. We divide by the flat frame\n    flat = flat_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1])\n    flat[dead_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1])] = 1\n    signal = signal/flat\n\n    # linear correction - treat as sixth degree polynomial, first number is C, not coefficient?\n    coefficients = linear_corr_frame.to_numpy().reshape([6] + sensor_sizes_dict[sensor][1])\n    coefficients = np.repeat(coefficients, sensor_sizes_dict[sensor][0][0], axis=1)\n    corrected = coefficients[0]\n    for i in range(5):\n        corrected += np.multiply(np.power(signal, i+1),coefficients[i+1])\n        \n    # gain, offset\n    corrected = corrected*planet_gain_offset[sensor + \"_adc_gain\"].values + planet_gain_offset[sensor + \"_adc_offset\"].values\n        \n    # zero out malfunctioning pixels\n    corrected[np.repeat(dead_frame.to_numpy().reshape(sensor_sizes_dict[sensor][1]), sensor_sizes_dict[sensor][0][0], axis=0)] = 0\n```",
    "2947042": "Do hosts truly expect us to reverse-engineering SOC@ESA pre-processing pipeline?\nI think more specification should be needed on how to handle calibration data (i.e. Level1 ~ Level1.5 product).\n\nE.g.\n\n* what linear-correction coefficient mean?"
  }
}