{
  "id": 528066,
  "title": "Understanding Calibration Data in Our Simulations",
  "url": "/competitions/ariel-data-challenge-2024/discussion/528066",
  "author_name": "Lorenzo Mugnai",
  "post_date": "2024-08-14T17:30:51.908000",
  "votes": 44,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I wanted to take a moment to explain some key aspects of the data we're working on within our simulations, particularly around dark frames and flat fields. There have been some great questions raised, so I hope this post clarifies things for those who are working on similar projects.</p>\n<h2>Dark Frames vs Dark Current Calibration Products</h2>\n<p>In typical astronomical data processing, dark frames are images taken with the shutter closed, capturing only the inherent electronic noise of the detector, known as dark current. These dark frames are usually subtracted from the actual observation data to remove this unwanted noise.</p>\n<p>However, in our case, the \"dark frames\" we provide are not traditional dark frames collected by a telescope. Instead, they are <strong>estimates of the dark current</strong>. Because of this, these dark current products need to be scaled according to the integration time of each frame before being subtracted from the observational data. This ensures that the dark current is accurately accounted for in relation to the exposure duration.</p>\n<h2>Flat Fields in Our Simulations</h2>\n<p>Similarly, flat fields in typical observations are used to correct for variations in the pixel-to-pixel sensitivity of the detector. Observationally, these are obtained by imaging a uniformly illuminated field. The data we provide, however, are not observational flat fields. Instead, they are <strong>estimates of the relative efficiency</strong> of each pixel in detecting light, which is why these \"flats\" have a mean value of one. Their purpose is to correct for variations in quantum efficiency across the detector, but because they are simulated, they might differ statistically from observational flats.</p>\n<h2>Uniformity of Calibration Products</h2>\n<p>In our data, the calibration products (darks, flats, etc.) are identical across frames. This is because we’ve simulated observations using the same detector consistently, which reflects real-life scenarios where different observations would be made using the same telescope. As a result, the same calibration products would be available throughout all observations.</p>\n<h2>Should I calibrate?</h2>\n<p>Calibration is crucial in the study of planetary transits to ensure precise measurements of the brightness variations. Since transit signals are very weak, with a required accuracy of about 100 ppm, even small instrumental errors can introduce significant distortions.</p>\n<p>For example, if there is a constant additive signal ( C ) in the data, the transit depth, calculated as the difference of the out of transit signal and the in transit signal divided by the signal out of transit <br>\n$$\\frac{\\Delta S}{S} = \\frac{S_{\\text{oot}} - S_{\\text{it}}}{S_{\\text{oot}}}$$ <br>\nwill be distorted:</p>\n<p>$$\\frac{\\Delta S}{S} = \\frac{(S_{\\text{oot}}+C) - (S_{\\text{it}}+C)}{S_{\\text{oot}} + C} = \\frac{S_{\\text{oot}} - S_{\\text{it}}}{S_{\\text{oot}} + C}$$</p>\n<p>This causes a bias in the estimation, leading to inaccurate conclusions about the planet's characteristics. Calibration is essential to remove these errors, ensuring that the measurements accurately reflect the transit signal.</p>\n<h2><strong>Notes on the Provided Notebooks</strong></h2>\n<p>We’ve noticed a few issues in the notebook that we're currently addressing with you:</p>\n<ul>\n<li><strong>Dark Current</strong>: As mentioned earlier, the dark current should be scaled according to the integration time of each pixel. This scaling was not applied correctly in previous iterations of the notebook, but we included it in the updated ones (see Notebooks listed later).</li>\n<li><strong>Subtracting Dark Current from Flat Fields</strong>: This subtraction is incorrect based on the earlier discussion. The dark current should not be subtracted from the flat fields as they serve a different purpose.</li>\n<li><strong>Multiplication of Gain</strong>: To correctly invert the analogue-to-digital conversion equation, the data should be divided by the gain, not multiplied. However, this doesn’t change the final result since we’re working with relative measurements.</li>\n</ul>\n<p>We have updated our notebook to include these fixes, which you can find here:</p>\n<p><a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\"><strong>Updated Notebook</strong></a></p>\n<p>Additionally, we recommend reviewing our new notebook for single observation calibrations:</p>\n<p><a href=\"https://www.kaggle.com/code/gordonyip/calibrating-a-single-observation/notebook\" target=\"_blank\"><strong>Calibrating a Single Observation</strong></a></p>\n<p>I hope this explanation helps you better understand how to work with the data in our simulations. Feel free to reach out if you have any further questions or need clarification on any points!</p>\n<p>Best regards,</p>",
  "messages": [
    {
      "id": 2959214,
      "postDate": "2024-08-14T17:30:51.907Z",
      "content": "<p>Hi everyone,</p>\n<p>I wanted to take a moment to explain some key aspects of the data we're working on within our simulations, particularly around dark frames and flat fields. There have been some great questions raised, so I hope this post clarifies things for those who are working on similar projects.</p>\n<h2>Dark Frames vs Dark Current Calibration Products</h2>\n<p>In typical astronomical data processing, dark frames are images taken with the shutter closed, capturing only the inherent electronic noise of the detector, known as dark current. These dark frames are usually subtracted from the actual observation data to remove this unwanted noise.</p>\n<p>However, in our case, the \"dark frames\" we provide are not traditional dark frames collected by a telescope. Instead, they are <strong>estimates of the dark current</strong>. Because of this, these dark current products need to be scaled according to the integration time of each frame before being subtracted from the observational data. This ensures that the dark current is accurately accounted for in relation to the exposure duration.</p>\n<h2>Flat Fields in Our Simulations</h2>\n<p>Similarly, flat fields in typical observations are used to correct for variations in the pixel-to-pixel sensitivity of the detector. Observationally, these are obtained by imaging a uniformly illuminated field. The data we provide, however, are not observational flat fields. Instead, they are <strong>estimates of the relative efficiency</strong> of each pixel in detecting light, which is why these \"flats\" have a mean value of one. Their purpose is to correct for variations in quantum efficiency across the detector, but because they are simulated, they might differ statistically from observational flats.</p>\n<h2>Uniformity of Calibration Products</h2>\n<p>In our data, the calibration products (darks, flats, etc.) are identical across frames. This is because we’ve simulated observations using the same detector consistently, which reflects real-life scenarios where different observations would be made using the same telescope. As a result, the same calibration products would be available throughout all observations.</p>\n<h2>Should I calibrate?</h2>\n<p>Calibration is crucial in the study of planetary transits to ensure precise measurements of the brightness variations. Since transit signals are very weak, with a required accuracy of about 100 ppm, even small instrumental errors can introduce significant distortions.</p>\n<p>For example, if there is a constant additive signal ( C ) in the data, the transit depth, calculated as the difference of the out of transit signal and the in transit signal divided by the signal out of transit <br>\n$$\\frac{\\Delta S}{S} = \\frac{S_{\\text{oot}} - S_{\\text{it}}}{S_{\\text{oot}}}$$ <br>\nwill be distorted:</p>\n<p>$$\\frac{\\Delta S}{S} = \\frac{(S_{\\text{oot}}+C) - (S_{\\text{it}}+C)}{S_{\\text{oot}} + C} = \\frac{S_{\\text{oot}} - S_{\\text{it}}}{S_{\\text{oot}} + C}$$</p>\n<p>This causes a bias in the estimation, leading to inaccurate conclusions about the planet's characteristics. Calibration is essential to remove these errors, ensuring that the measurements accurately reflect the transit signal.</p>\n<h2><strong>Notes on the Provided Notebooks</strong></h2>\n<p>We’ve noticed a few issues in the notebook that we're currently addressing with you:</p>\n<ul>\n<li><strong>Dark Current</strong>: As mentioned earlier, the dark current should be scaled according to the integration time of each pixel. This scaling was not applied correctly in previous iterations of the notebook, but we included it in the updated ones (see Notebooks listed later).</li>\n<li><strong>Subtracting Dark Current from Flat Fields</strong>: This subtraction is incorrect based on the earlier discussion. The dark current should not be subtracted from the flat fields as they serve a different purpose.</li>\n<li><strong>Multiplication of Gain</strong>: To correctly invert the analogue-to-digital conversion equation, the data should be divided by the gain, not multiplied. However, this doesn’t change the final result since we’re working with relative measurements.</li>\n</ul>\n<p>We have updated our notebook to include these fixes, which you can find here:</p>\n<p><a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data\" target=\"_blank\"><strong>Updated Notebook</strong></a></p>\n<p>Additionally, we recommend reviewing our new notebook for single observation calibrations:</p>\n<p><a href=\"https://www.kaggle.com/code/gordonyip/calibrating-a-single-observation/notebook\" target=\"_blank\"><strong>Calibrating a Single Observation</strong></a></p>\n<p>I hope this explanation helps you better understand how to work with the data in our simulations. Feel free to reach out if you have any further questions or need clarification on any points!</p>\n<p>Best regards,</p>",
      "rawMarkdown": "Hi everyone,\n\nI wanted to take a moment to explain some key aspects of the data we're working on within our simulations, particularly around dark frames and flat fields. There have been some great questions raised, so I hope this post clarifies things for those who are working on similar projects.\n\n## Dark Frames vs Dark Current Calibration Products\n\nIn typical astronomical data processing, dark frames are images taken with the shutter closed, capturing only the inherent electronic noise of the detector, known as dark current. These dark frames are usually subtracted from the actual observation data to remove this unwanted noise.\n\nHowever, in our case, the \"dark frames\" we provide are not traditional dark frames collected by a telescope. Instead, they are **estimates of the dark current**. Because of this, these dark current products need to be scaled according to the integration time of each frame before being subtracted from the observational data. This ensures that the dark current is accurately accounted for in relation to the exposure duration.\n\n## Flat Fields in Our Simulations\n\nSimilarly, flat fields in typical observations are used to correct for variations in the pixel-to-pixel sensitivity of the detector. Observationally, these are obtained by imaging a uniformly illuminated field. The data we provide, however, are not observational flat fields. Instead, they are **estimates of the relative efficiency** of each pixel in detecting light, which is why these \"flats\" have a mean value of one. Their purpose is to correct for variations in quantum efficiency across the detector, but because they are simulated, they might differ statistically from observational flats.\n\n## Uniformity of Calibration Products\n\nIn our data, the calibration products (darks, flats, etc.) are identical across frames. This is because we’ve simulated observations using the same detector consistently, which reflects real-life scenarios where different observations would be made using the same telescope. As a result, the same calibration products would be available throughout all observations.\n\n## Should I calibrate?\n\nCalibration is crucial in the study of planetary transits to ensure precise measurements of the brightness variations. Since transit signals are very weak, with a required accuracy of about 100 ppm, even small instrumental errors can introduce significant distortions.\n\nFor example, if there is a constant additive signal \\( C \\) in the data, the transit depth, calculated as the difference of the out of transit signal and the in transit signal divided by the signal out of transit \n$$\\frac{\\Delta S}{S} = \\frac{S_{\\text{oot}} - S_{\\text{it}}}{S_{\\text{oot}}}$$ \nwill be distorted:\n\n$$\\frac{\\Delta S}{S} = \\frac{(S_{\\text{oot}}+C) - (S_{\\text{it}}+C)}{S_{\\text{oot}} + C} = \\frac{S_{\\text{oot}} - S_{\\text{it}}}{S_{\\text{oot}} + C}$$\n\nThis causes a bias in the estimation, leading to inaccurate conclusions about the planet's characteristics. Calibration is essential to remove these errors, ensuring that the measurements accurately reflect the transit signal.\n\n## **Notes on the Provided Notebooks**\n\nWe’ve noticed a few issues in the notebook that we're currently addressing with you:\n\n+ **Dark Current**: As mentioned earlier, the dark current should be scaled according to the integration time of each pixel. This scaling was not applied correctly in previous iterations of the notebook, but we included it in the updated ones (see Notebooks listed later).\n+ **Subtracting Dark Current from Flat Fields**: This subtraction is incorrect based on the earlier discussion. The dark current should not be subtracted from the flat fields as they serve a different purpose.\n+ **Multiplication of Gain**: To correctly invert the analogue-to-digital conversion equation, the data should be divided by the gain, not multiplied. However, this doesn’t change the final result since we’re working with relative measurements.\n\nWe have updated our notebook to include these fixes, which you can find here:\n\n[**Updated Notebook**](https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data)\n\nAdditionally, we recommend reviewing our new notebook for single observation calibrations:\n\n[**Calibrating a Single Observation**](https://www.kaggle.com/code/gordonyip/calibrating-a-single-observation/notebook)\n\nI hope this explanation helps you better understand how to work with the data in our simulations. Feel free to reach out if you have any further questions or need clarification on any points!\n\nBest regards,\n",
      "votes": 42
    },
    {
      "id": 2964080,
      "postDate": "2024-08-19T14:19:34.920Z",
      "content": "<p><a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> </p>\n<p>Can I ask the technical details of <code>read</code> product?</p>",
      "rawMarkdown": "@lorenzomugnai \n\nCan I ask the technical details of `read` product?",
      "votes": 4,
      "replies": [
        {
          "id": 2964133,
          "postDate": "2024-08-19T14:57:47.217Z",
          "content": "<p>Just in case. This is the exposure taken with the shortest possible exposure time with the shutter closed. You take several, average and it gives you the readout noise.<br>\nAfter that, you can subtract readout noise from the dark frame and get an estimate of dark current which can be scaled upwards or downwards depending on the exposure time. In our case, readout noise was already subtracted (supposedly) from the dark frame, it's not clear if it was subtracted from the light frames though.</p>",
          "rawMarkdown": "Just in case. This is the exposure taken with the shortest possible exposure time with the shutter closed. You take several, average and it gives you the readout noise.\nAfter that, you can subtract readout noise from the dark frame and get an estimate of dark current which can be scaled upwards or downwards depending on the exposure time. In our case, readout noise was already subtracted (supposedly) from the dark frame, it's not clear if it was subtracted from the light frames though.",
          "votes": 4,
          "replies": [
            {
              "id": 2964247,
              "postDate": "2024-08-19T16:49:17.997Z",
              "content": "<p>Yes, that’s correct. The \"read\" frames we provide are estimates of the read noise amplitude for each pixel. As mentioned earlier, in our case, the dark frames are estimations of the dark current and are not affected by read noise. However, the science (or light) frames are affected by read noise.</p>",
              "rawMarkdown": "Yes, that’s correct. The \"read\" frames we provide are estimates of the read noise amplitude for each pixel. As mentioned earlier, in our case, the dark frames are estimations of the dark current and are not affected by read noise. However, the science (or light) frames are affected by read noise.",
              "votes": 6
            },
            {
              "id": 2964378,
              "postDate": "2024-08-19T18:28:35.693Z",
              "content": "<p><a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> Does this picture correctly describes detector measurements?<br>\nThen, since read noise is constant on both two observations (i.e. at 0.1 and 4.5), we don't have to take care read noise because they cancels out by subtracting them?</p>\n<p>[EDIT] And if we should take care read noise, should we subtract them from actual measurements, or add to them?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fed1a38209bd2ababeb2757d3b66190c6%2FScreenshot%202024-08-20%20at%203.23.35.png?generation=1724091842365472&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "@lorenzomugnai Does this picture correctly describes detector measurements?\nThen, since read noise is constant on both two observations (i.e. at 0.1 and 4.5), we don't have to take care read noise because they cancels out by subtracting them?\n\n[EDIT] And if we should take care read noise, should we subtract them from actual measurements, or add to them?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fed1a38209bd2ababeb2757d3b66190c6%2FScreenshot%202024-08-20%20at%203.23.35.png?generation=1724091842365472&alt=media)",
              "votes": 5
            },
            {
              "id": 2964423,
              "postDate": "2024-08-19T18:55:40.280Z",
              "content": "<p>The read noise is an additional source of noise originating from the electronic components of the detector, such as the amplifiers and readout circuitry, and it is random in nature. Even though the read noise may have similar amplitudes for both images (at 0.1 s and 4.5 s in your diagram), it does not cancel out during subtraction.</p>\n<p>Regarding the way signal and dark current fill up the detector, your diagram looks representative. To clean up for dark current, you can refer to the <a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data?scriptVersionId=192915144&amp;cellId=20\" target=\"_blank\">example notebook</a> where we’ve implemented this process.</p>",
              "rawMarkdown": "The read noise is an additional source of noise originating from the electronic components of the detector, such as the amplifiers and readout circuitry, and it is random in nature. Even though the read noise may have similar amplitudes for both images (at 0.1 s and 4.5 s in your diagram), it does not cancel out during subtraction.\n\nRegarding the way signal and dark current fill up the detector, your diagram looks representative. To clean up for dark current, you can refer to the [example notebook](https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data?scriptVersionId=192915144&cellId=20) where we’ve implemented this process.",
              "votes": 4
            },
            {
              "id": 2964433,
              "postDate": "2024-08-19T19:09:19.980Z",
              "content": "<p>It's not cancelling in the same sense as subtracting one gaussian noise from another generated with the same parameters do to cancel each other.</p>",
              "rawMarkdown": "It's not cancelling in the same sense as subtracting one gaussian noise from another generated with the same parameters do to cancel each other.",
              "votes": 1
            },
            {
              "id": 2964509,
              "postDate": "2024-08-19T21:27:41.780Z",
              "content": "<p>That is correct. Because, as I said, read noise is random in nature. Thank you</p>",
              "rawMarkdown": "That is correct. Because, as I said, read noise is random in nature. Thank you",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2963610,
      "postDate": "2024-08-18T22:45:26.377Z",
      "content": "<p>So is the data provided uncalibrated or already calibrated to a certain degree? What units are it in; raw detector counts (dN?). Does the provided files properly flux calibrate everything to mJy? or just to Me/s?</p>",
      "rawMarkdown": "So is the data provided uncalibrated or already calibrated to a certain degree? What units are it in; raw detector counts (dN?). Does the provided files properly flux calibrate everything to mJy? or just to Me/s?",
      "votes": 1,
      "replies": [
        {
          "id": 2964254,
          "postDate": "2024-08-19T16:52:32.353Z",
          "content": "<p>The data we provide are raw and in units of ADU (Analog-to-Digital Units).</p>\n<p>ADU refers to the digital values generated by the detector in response to the light it collects. When light hits the detector, it creates a signal that is converted from an analogue signal (the voltage produced by photons hitting the detector) to a digital signal through an analogue-to-digital converter. These values are stored as integers because the analogue-to-digital conversion process rounds the signal to whole numbers for efficiency and simplicity in storage. The ADU value represents the strength of this signal but has not yet been calibrated into physical units like flux (e.g., mJy).</p>\n<p>The data you have, therefore, have not been flux-calibrated yet.</p>",
          "rawMarkdown": "The data we provide are raw and in units of ADU (Analog-to-Digital Units).\n\nADU refers to the digital values generated by the detector in response to the light it collects. When light hits the detector, it creates a signal that is converted from an analogue signal (the voltage produced by photons hitting the detector) to a digital signal through an analogue-to-digital converter. These values are stored as integers because the analogue-to-digital conversion process rounds the signal to whole numbers for efficiency and simplicity in storage. The ADU value represents the strength of this signal but has not yet been calibrated into physical units like flux (e.g., mJy).\n\nThe data you have, therefore, have not been flux-calibrated yet.",
          "votes": 8,
          "replies": [
            {
              "id": 2964653,
              "postDate": "2024-08-20T03:23:43.767Z",
              "content": "<p>Thanks for the clarification!</p>",
              "rawMarkdown": "Thanks for the clarification!"
            }
          ]
        }
      ]
    },
    {
      "id": 3028326,
      "postDate": "2024-10-25T21:36:57.240Z",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> !</p>\n<p>Thank you for awesome explanations. Regarding the order of transformations in notebook:</p>\n<p>Shouldn't the Flat Field Correction be done just after the linear corr? </p>\n<p>I am concerned with applying it after CDS. Does it mean that Flat was acquired using CDS too?</p>",
      "rawMarkdown": "Dear @lorenzomugnai !\n\nThank you for awesome explanations. Regarding the order of transformations in notebook:\n\nShouldn't the Flat Field Correction be done just after the linear corr? \n\nI am concerned with applying it after CDS. Does it mean that Flat was acquired using CDS too?",
      "replies": [
        {
          "id": 3028655,
          "postDate": "2024-10-26T11:51:09.290Z",
          "content": "<p>Hello,</p>\n<p>Thank you for your question. Applying the flat field correction before or after the CDS (Correlated Double Sampling) should not affect your results. The flat field is essentially a map of multiplicative factors applied pixel by pixel.</p>\n<p>Let’s assume your CDS is calculated using NDR_0 and NDR_1, where your pixel (XY) has a certain value in NDR_0 and NDR_1, and its corresponding multiplicative factor from the flat field is XY_F. If you apply the flat field correction before the CDS, you get:</p>\n<ul>\n<li>Corrected NDR_0: XY_0 / XY_F</li>\n<li>Corrected NDR_1: XY_1 / XY_F</li>\n</ul>\n<p>Then, the CDS would be calculated as:</p>\n<p>CDS = (XY_1 / XY_F) - (XY_0 / XY_F) = (XY_1 - XY_0) / XY_F</p>\n<p>As you can see, this is equivalent to calculating the CDS first and then applying the flat field correction afterward.</p>\n<p>I hope this clarifies your question.</p>",
          "rawMarkdown": "Hello,\n\nThank you for your question. Applying the flat field correction before or after the CDS (Correlated Double Sampling) should not affect your results. The flat field is essentially a map of multiplicative factors applied pixel by pixel.\n\nLet’s assume your CDS is calculated using NDR_0 and NDR_1, where your pixel (XY) has a certain value in NDR_0 and NDR_1, and its corresponding multiplicative factor from the flat field is XY_F. If you apply the flat field correction before the CDS, you get:\n\n- Corrected NDR_0: XY_0 / XY_F\n- Corrected NDR_1: XY_1 / XY_F\n\nThen, the CDS would be calculated as:\n\nCDS = (XY_1 / XY_F) - (XY_0 / XY_F) = (XY_1 - XY_0) / XY_F\n\nAs you can see, this is equivalent to calculating the CDS first and then applying the flat field correction afterward.\n\nI hope this clarifies your question.",
          "replies": [
            {
              "id": 3029124,
              "postDate": "2024-10-26T22:40:37.753Z",
              "content": "<p>Yeah. Makes perfect sens, omg, it was actually a silly question, but it was middle of night for me ;)</p>\n<p>On the other hand, I have another question, more of \"why\" or \"how\".</p>\n<p>There is the linear correction (which BTW could be better named nonlinear correction ;) ) step, why the flat correction is not part of that? In other words, I am wondering, how did you manage to obtain per pixels polynomial corrections and separately flat correction?</p>\n<p>I guess that there is a methodology details that motivates it. Would you be so kind to share that please to satisfy my curiosity?</p>",
              "rawMarkdown": "Yeah. Makes perfect sens, omg, it was actually a silly question, but it was middle of night for me ;)\n\nOn the other hand, I have another question, more of \"why\" or \"how\".\n\nThere is the linear correction (which BTW could be better named nonlinear correction ;) ) step, why the flat correction is not part of that? In other words, I am wondering, how did you manage to obtain per pixels polynomial corrections and separately flat correction?\n\nI guess that there is a methodology details that motivates it. Would you be so kind to share that please to satisfy my curiosity?"
            },
            {
              "id": 3030249,
              "postDate": "2024-10-28T10:39:35.633Z",
              "content": "<p>Thank you for your the question. <br>\nThe distinction between the nonlinear correction and the flat-field correction lies in the nature of the physical processes they aim to address, which is why they are treated separately.</p>\n<p>The nonlinear correction accounts for the deviation from a linear response as a pixel approaches its full-well capacity. This effect is especially important in high-precision photometric measurements, as pixels often exhibit a nonlinear behaviour when nearing saturation. To quantify this, we perform a detailed characterisation by exposing the detector to progressively increasing illumination levels until saturation. By analysing the response curve, we derive a polynomial correction that compensates for these deviations, allowing us to recover the true linear signal response of each pixel.</p>\n<p>The flat-field correction, however, targets a different source of variation: the intrinsic pixel-to-pixel variations in quantum efficiency across the detector. Even when uniformly illuminated, different pixels may register slightly different signal levels due to manufacturing inconsistencies or small variations in the detector's structure. The flat-field is typically derived by exposing the detector to a uniform light source, ensuring the entire field is evenly illuminated. The resulting map captures the relative sensitivity of each pixel, allowing us to correct for these differences and ensure a homogeneous response across the detector.</p>\n<p>In essence, while the nonlinear correction adjusts for how individual pixels behave under varying illumination levels near saturation, the flat-field correction ensures that each pixel’s response to a uniform flux is consistent. This separation allows us to more accurately address the specific characteristics of the detector at each stage of the calibration process.</p>\n<p>I hope this explanation satisfies your curiosity 😀</p>",
              "rawMarkdown": "Thank you for your the question. \nThe distinction between the nonlinear correction and the flat-field correction lies in the nature of the physical processes they aim to address, which is why they are treated separately.\n\nThe nonlinear correction accounts for the deviation from a linear response as a pixel approaches its full-well capacity. This effect is especially important in high-precision photometric measurements, as pixels often exhibit a nonlinear behaviour when nearing saturation. To quantify this, we perform a detailed characterisation by exposing the detector to progressively increasing illumination levels until saturation. By analysing the response curve, we derive a polynomial correction that compensates for these deviations, allowing us to recover the true linear signal response of each pixel.\n\nThe flat-field correction, however, targets a different source of variation: the intrinsic pixel-to-pixel variations in quantum efficiency across the detector. Even when uniformly illuminated, different pixels may register slightly different signal levels due to manufacturing inconsistencies or small variations in the detector's structure. The flat-field is typically derived by exposing the detector to a uniform light source, ensuring the entire field is evenly illuminated. The resulting map captures the relative sensitivity of each pixel, allowing us to correct for these differences and ensure a homogeneous response across the detector.\n\nIn essence, while the nonlinear correction adjusts for how individual pixels behave under varying illumination levels near saturation, the flat-field correction ensures that each pixel’s response to a uniform flux is consistent. This separation allows us to more accurately address the specific characteristics of the detector at each stage of the calibration process.\n\nI hope this explanation satisfies your curiosity 😀"
            }
          ]
        }
      ]
    },
    {
      "id": 3005967,
      "postDate": "2024-10-03T15:12:21.140Z",
      "content": "<p><a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> , thank you for this post! Could you please also explain the meaning of the axis_info.parquet columns?</p>",
      "rawMarkdown": "@lorenzomugnai , thank you for this post! Could you please also explain the meaning of the axis_info.parquet columns?"
    },
    {
      "id": 3004053,
      "postDate": "2024-10-01T14:26:43.417Z",
      "content": "<p>How might the changes in handling the dark current scaling and gain multiplication affect the accuracy of the final relative measurements + what implications might this have for future similar data analysis ? Thanks</p>",
      "rawMarkdown": "How might the changes in handling the dark current scaling and gain multiplication affect the accuracy of the final relative measurements + what implications might this have for future similar data analysis ? Thanks",
      "replies": [
        {
          "id": 3005646,
          "postDate": "2024-10-03T06:50:10.357Z",
          "content": "<p>Good questions! Allow me to use some math to answer this.</p>\n<h2>Dark Current</h2>\n<p>If dark current is removed improperly, it introduces an additive bias. Our signal is ( S + c ), where ( c ) is the dark current additive term. As seen before, the transit measurement is calculated as the flux out of transit minus the flux in transit, divided by the flux out of transit. In equations, this is:</p>\n<p>$$<br>\nf_1 = \\frac{S_{oot} - S_{it}}{S_{oot}}<br>\n$$</p>\n<p>As I shown in my post, if we have the additive term, it becomes:</p>\n<p>$$<br>\nf_2 = \\frac{S_{oot} + c - S_{it} - c}{S_{oot} + c} = \\frac{S_{oot} - S_{it}}{S_{oot} + c}<br>\n$$</p>\n<p>And we end up with a biased estimator because of the changed denominator. This means that we are now measuring an incorrect transit depth.</p>\n<p>Let's focus on the noise estimation now. If we use the \"simple\" formula:</p>\n<p>$$<br>\nf_1 = \\frac{S_{oot} - S_{it}}{S_{oot}}<br>\n$$</p>\n<p>assuming all the terms are independent (are they? 😉), we get from the noise propagation:</p>\n<p>$$<br>\n\\Delta f_1 = \\frac{S_{it}}{S_{oot}^2} \\Delta S_{oot} + \\frac{1}{S_{oot}} \\Delta S_{it}<br>\n$$</p>\n<p>However, if we compute the same for the biased equation:</p>\n<p>$$<br>\n\\Delta f_2 = \\frac{(c + S_{it}) \\Delta S_{oot} + (S_{oot} + c) \\Delta S_{it} + (S_{oot} - S_{it}) \\Delta c}{(S_{oot} + c)^2}<br>\n$$</p>\n<p>This is the noise equation you \"should\" use to properly account for the dark current term. But if you corrected for the dark current improperly, you don't really know that it's there. As a result, you end up using the previous equation, which is incorrect, leading to an uncertainty that is not representative and therefore does not account for the introduced bias.</p>\n<h2>Gain Multiplication</h2>\n<p>Multiplying by the gain is necessary to convert from integer values to floating-point values, obtaining the \"true\" electron counts. It is technically the same for all frames, so doing the math again, it disappears:</p>\n<p>$$<br>\nf_3 = \\frac{g \\cdot S_{oot} - g \\cdot S_{it}}{g \\cdot S_{oot}} = \\frac{g(S_{oot} - S_{it})}{g \\cdot S_{oot}} = \\frac{S_{oot} - S_{it}}{S_{oot}} = f_1<br>\n$$</p>\n<p>This doesn't really matter for relative measurements like ours. It will matter if you want or need to perform absolute calibrations to estimate the actual flux.</p>\n<h2>Applying the Equations in Model Fitting</h2>\n<p>These calculations are for a generic fit of the transit depth. How you apply these equations depends on \"how\" the quantity is estimated. Typically, a \"model\" is fitted to the data to determine the transit depth. If the model accounts for the bias term (e.g., includes the dark current ( c ) as a parameter or ensures proper scaling), the derived equations will accurately reflect the uncertainties and biases in the measurements.</p>\n<p>In other words, when performing a fit:</p>\n<ul>\n<li>If the fitting model includes the bias term correctly, the propagation of errors will account for it, leading to more accurate and unbiased estimates of the transit depth. Such a model may be referred to as an instrument model, systematic model, or detrending model, depending on the specific components it aims to correct or remove from the science signal.</li>\n<li>If the bias is not properly accounted for in the model, the uncertainties and the estimated transit depth will be biased, potentially leading to incorrect scientific conclusions.</li>\n</ul>\n<p>Additionally, care must be taken to avoid <strong>overfitting</strong> the model to the data. This occurs when the model becomes too complex, capturing not only the underlying signal but also the noise in such details that can result in poor generalization to new observations. However, the model should not be <strong>underfitting</strong> either by being too simplistic, as this may cause loss of important scientific signals by inadvertently incorporating them into the detrending process. We have learnt from experienced that this is hard to balance. </p>\n<h2>Conclusion</h2>\n<p>These are (or should be) standard practices and considerations in science. It's not easy and might even be \"less easy\" to generalize these processes, but science is what drives us forward, so it's not meant to be easy. As I tell my students,</p>\n<blockquote>\n  <p><strong>\"If this task were easy to do now, we would have done it before when it was difficult.\"</strong></p>\n</blockquote>",
          "rawMarkdown": "Good questions! Allow me to use some math to answer this.\n\n## Dark Current\n\nIf dark current is removed improperly, it introduces an additive bias. Our signal is \\( S + c \\), where \\( c \\) is the dark current additive term. As seen before, the transit measurement is calculated as the flux out of transit minus the flux in transit, divided by the flux out of transit. In equations, this is:\n\n$$\nf_1 = \\frac{S_{oot} - S_{it}}{S_{oot}}\n$$\n\nAs I shown in my post, if we have the additive term, it becomes:\n\n$$\nf_2 = \\frac{S_{oot} + c - S_{it} - c}{S_{oot} + c} = \\frac{S_{oot} - S_{it}}{S_{oot} + c}\n$$\n\nAnd we end up with a biased estimator because of the changed denominator. This means that we are now measuring an incorrect transit depth.\n\nLet's focus on the noise estimation now. If we use the \"simple\" formula:\n\n$$\nf_1 = \\frac{S_{oot} - S_{it}}{S_{oot}}\n$$\n\nassuming all the terms are independent (are they? 😉), we get from the noise propagation:\n\n$$\n\\Delta f_1 = \\frac{S_{it}}{S_{oot}^2} \\Delta S_{oot} + \\frac{1}{S_{oot}} \\Delta S_{it}\n$$\n\nHowever, if we compute the same for the biased equation:\n\n$$\n\\Delta f_2 = \\frac{(c + S_{it}) \\Delta S_{oot} + (S_{oot} + c) \\Delta S_{it} + (S_{oot} - S_{it}) \\Delta c}{(S_{oot} + c)^2}\n$$\n\nThis is the noise equation you \"should\" use to properly account for the dark current term. But if you corrected for the dark current improperly, you don't really know that it's there. As a result, you end up using the previous equation, which is incorrect, leading to an uncertainty that is not representative and therefore does not account for the introduced bias.\n\n## Gain Multiplication\n\nMultiplying by the gain is necessary to convert from integer values to floating-point values, obtaining the \"true\" electron counts. It is technically the same for all frames, so doing the math again, it disappears:\n\n$$\nf_3 = \\frac{g \\cdot S_{oot} - g \\cdot S_{it}}{g \\cdot S_{oot}} = \\frac{g(S_{oot} - S_{it})}{g \\cdot S_{oot}} = \\frac{S_{oot} - S_{it}}{S_{oot}} = f_1\n$$\n\nThis doesn't really matter for relative measurements like ours. It will matter if you want or need to perform absolute calibrations to estimate the actual flux.\n\n## Applying the Equations in Model Fitting\n\nThese calculations are for a generic fit of the transit depth. How you apply these equations depends on \"how\" the quantity is estimated. Typically, a \"model\" is fitted to the data to determine the transit depth. If the model accounts for the bias term (e.g., includes the dark current \\( c \\) as a parameter or ensures proper scaling), the derived equations will accurately reflect the uncertainties and biases in the measurements.\n\nIn other words, when performing a fit:\n- If the fitting model includes the bias term correctly, the propagation of errors will account for it, leading to more accurate and unbiased estimates of the transit depth. Such a model may be referred to as an instrument model, systematic model, or detrending model, depending on the specific components it aims to correct or remove from the science signal.\n- If the bias is not properly accounted for in the model, the uncertainties and the estimated transit depth will be biased, potentially leading to incorrect scientific conclusions.\n\nAdditionally, care must be taken to avoid **overfitting** the model to the data. This occurs when the model becomes too complex, capturing not only the underlying signal but also the noise in such details that can result in poor generalization to new observations. However, the model should not be **underfitting** either by being too simplistic, as this may cause loss of important scientific signals by inadvertently incorporating them into the detrending process. We have learnt from experienced that this is hard to balance. \n\n## Conclusion\n\nThese are (or should be) standard practices and considerations in science. It's not easy and might even be \"less easy\" to generalize these processes, but science is what drives us forward, so it's not meant to be easy. As I tell my students,\n\n> **\"If this task were easy to do now, we would have done it before when it was difficult.\"**",
          "votes": 4
        }
      ]
    },
    {
      "id": 2998751,
      "postDate": "2024-09-26T00:58:10.500Z",
      "content": "<p><a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a>  Thanks Lorenzo, i understand that we apply the linearity corrections prior to subtracting dark current - does it matter if it is done in this order rather than the other way round? (intuitively, I may have thought that we want to subtract the dark current prior to linearity corrections, as we can then first get the relevant measured current and figure out what the actual photos hitting the sensor is that generated these currents, but my understanding is likely flawed) </p>",
      "rawMarkdown": "@lorenzomugnai  Thanks Lorenzo, i understand that we apply the linearity corrections prior to subtracting dark current - does it matter if it is done in this order rather than the other way round? (intuitively, I may have thought that we want to subtract the dark current prior to linearity corrections, as we can then first get the relevant measured current and figure out what the actual photos hitting the sensor is that generated these currents, but my understanding is likely flawed) ",
      "replies": [
        {
          "id": 2999120,
          "postDate": "2024-09-26T11:45:17.040Z",
          "content": "<p>Thank you. Indeed, there are different approaches regarding the order of calibration steps, and sometimes the sequence can affect the results. In this specific case, the difference should be minimal; however, my opinion is that linearity correction should be applied before subtracting the dark current.</p>\n<p>Although the dark current is thermally generated and not in response to a light signal, it still contributes to filling the pixel well, thereby reducing the capacity to convert photons into electrons. It’s important to remember that the pixel well is filled by electrons, and the detector does not differentiate between photo-generated electrons and thermal electrons. Consequently, the observed deviation from linearity is due to both the filling of the well by incoming photons and the dark current. By applying the linearity correction first, we account for the intrinsic non-linearities of the sensor that arise from both of these factors.</p>",
          "rawMarkdown": "Thank you. Indeed, there are different approaches regarding the order of calibration steps, and sometimes the sequence can affect the results. In this specific case, the difference should be minimal; however, my opinion is that linearity correction should be applied before subtracting the dark current.\n\nAlthough the dark current is thermally generated and not in response to a light signal, it still contributes to filling the pixel well, thereby reducing the capacity to convert photons into electrons. It’s important to remember that the pixel well is filled by electrons, and the detector does not differentiate between photo-generated electrons and thermal electrons. Consequently, the observed deviation from linearity is due to both the filling of the well by incoming photons and the dark current. By applying the linearity correction first, we account for the intrinsic non-linearities of the sensor that arise from both of these factors.",
          "votes": 4
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2964080,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2024-08-19T14:19:34.920000",
      "content": "<p><a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> </p>\n<p>Can I ask the technical details of <code>read</code> product?</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2964133,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-08-19T14:57:47.217000",
          "content": "<p>Just in case. This is the exposure taken with the shortest possible exposure time with the shutter closed. You take several, average and it gives you the readout noise.<br>\nAfter that, you can subtract readout noise from the dark frame and get an estimate of dark current which can be scaled upwards or downwards depending on the exposure time. In our case, readout noise was already subtracted (supposedly) from the dark frame, it's not clear if it was subtracted from the light frames though.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2964247,
              "author_name": "Lorenzo Mugnai",
              "author_url": "",
              "post_date": "2024-08-19T16:49:17.997000",
              "content": "<p>Yes, that’s correct. The \"read\" frames we provide are estimates of the read noise amplitude for each pixel. As mentioned earlier, in our case, the dark frames are estimations of the dark current and are not affected by read noise. However, the science (or light) frames are affected by read noise.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2964378,
              "author_name": "Bilzard",
              "author_url": "",
              "post_date": "2024-08-19T18:28:35.693000",
              "content": "<p><a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> Does this picture correctly describes detector measurements?<br>\nThen, since read noise is constant on both two observations (i.e. at 0.1 and 4.5), we don't have to take care read noise because they cancels out by subtracting them?</p>\n<p>[EDIT] And if we should take care read noise, should we subtract them from actual measurements, or add to them?</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fed1a38209bd2ababeb2757d3b66190c6%2FScreenshot%202024-08-20%20at%203.23.35.png?generation=1724091842365472&amp;alt=media\" alt=\"\"></p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2964423,
              "author_name": "Lorenzo Mugnai",
              "author_url": "",
              "post_date": "2024-08-19T18:55:40.280000",
              "content": "<p>The read noise is an additional source of noise originating from the electronic components of the detector, such as the amplifiers and readout circuitry, and it is random in nature. Even though the read noise may have similar amplitudes for both images (at 0.1 s and 4.5 s in your diagram), it does not cancel out during subtraction.</p>\n<p>Regarding the way signal and dark current fill up the detector, your diagram looks representative. To clean up for dark current, you can refer to the <a href=\"https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data?scriptVersionId=192915144&amp;cellId=20\" target=\"_blank\">example notebook</a> where we’ve implemented this process.</p>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2964433,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-08-19T19:09:19.980000",
              "content": "<p>It's not cancelling in the same sense as subtracting one gaussian noise from another generated with the same parameters do to cancel each other.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2964509,
              "author_name": "Lorenzo Mugnai",
              "author_url": "",
              "post_date": "2024-08-19T21:27:41.780000",
              "content": "<p>That is correct. Because, as I said, read noise is random in nature. Thank you</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2963610,
      "author_name": "Gh0st",
      "author_url": "",
      "post_date": "2024-08-18T22:45:26.377000",
      "content": "<p>So is the data provided uncalibrated or already calibrated to a certain degree? What units are it in; raw detector counts (dN?). Does the provided files properly flux calibrate everything to mJy? or just to Me/s?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2964254,
          "author_name": "Lorenzo Mugnai",
          "author_url": "",
          "post_date": "2024-08-19T16:52:32.353000",
          "content": "<p>The data we provide are raw and in units of ADU (Analog-to-Digital Units).</p>\n<p>ADU refers to the digital values generated by the detector in response to the light it collects. When light hits the detector, it creates a signal that is converted from an analogue signal (the voltage produced by photons hitting the detector) to a digital signal through an analogue-to-digital converter. These values are stored as integers because the analogue-to-digital conversion process rounds the signal to whole numbers for efficiency and simplicity in storage. The ADU value represents the strength of this signal but has not yet been calibrated into physical units like flux (e.g., mJy).</p>\n<p>The data you have, therefore, have not been flux-calibrated yet.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2964653,
              "author_name": "Gh0st",
              "author_url": "",
              "post_date": "2024-08-20T03:23:43.767000",
              "content": "<p>Thanks for the clarification!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3028326,
      "author_name": "Maciej J Mikulski",
      "author_url": "",
      "post_date": "2024-10-25T21:36:57.240000",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> !</p>\n<p>Thank you for awesome explanations. Regarding the order of transformations in notebook:</p>\n<p>Shouldn't the Flat Field Correction be done just after the linear corr? </p>\n<p>I am concerned with applying it after CDS. Does it mean that Flat was acquired using CDS too?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3028655,
          "author_name": "Lorenzo Mugnai",
          "author_url": "",
          "post_date": "2024-10-26T11:51:09.290000",
          "content": "<p>Hello,</p>\n<p>Thank you for your question. Applying the flat field correction before or after the CDS (Correlated Double Sampling) should not affect your results. The flat field is essentially a map of multiplicative factors applied pixel by pixel.</p>\n<p>Let’s assume your CDS is calculated using NDR_0 and NDR_1, where your pixel (XY) has a certain value in NDR_0 and NDR_1, and its corresponding multiplicative factor from the flat field is XY_F. If you apply the flat field correction before the CDS, you get:</p>\n<ul>\n<li>Corrected NDR_0: XY_0 / XY_F</li>\n<li>Corrected NDR_1: XY_1 / XY_F</li>\n</ul>\n<p>Then, the CDS would be calculated as:</p>\n<p>CDS = (XY_1 / XY_F) - (XY_0 / XY_F) = (XY_1 - XY_0) / XY_F</p>\n<p>As you can see, this is equivalent to calculating the CDS first and then applying the flat field correction afterward.</p>\n<p>I hope this clarifies your question.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3029124,
              "author_name": "Maciej J Mikulski",
              "author_url": "",
              "post_date": "2024-10-26T22:40:37.753000",
              "content": "<p>Yeah. Makes perfect sens, omg, it was actually a silly question, but it was middle of night for me ;)</p>\n<p>On the other hand, I have another question, more of \"why\" or \"how\".</p>\n<p>There is the linear correction (which BTW could be better named nonlinear correction ;) ) step, why the flat correction is not part of that? In other words, I am wondering, how did you manage to obtain per pixels polynomial corrections and separately flat correction?</p>\n<p>I guess that there is a methodology details that motivates it. Would you be so kind to share that please to satisfy my curiosity?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3030249,
              "author_name": "Lorenzo Mugnai",
              "author_url": "",
              "post_date": "2024-10-28T10:39:35.633000",
              "content": "<p>Thank you for your the question. <br>\nThe distinction between the nonlinear correction and the flat-field correction lies in the nature of the physical processes they aim to address, which is why they are treated separately.</p>\n<p>The nonlinear correction accounts for the deviation from a linear response as a pixel approaches its full-well capacity. This effect is especially important in high-precision photometric measurements, as pixels often exhibit a nonlinear behaviour when nearing saturation. To quantify this, we perform a detailed characterisation by exposing the detector to progressively increasing illumination levels until saturation. By analysing the response curve, we derive a polynomial correction that compensates for these deviations, allowing us to recover the true linear signal response of each pixel.</p>\n<p>The flat-field correction, however, targets a different source of variation: the intrinsic pixel-to-pixel variations in quantum efficiency across the detector. Even when uniformly illuminated, different pixels may register slightly different signal levels due to manufacturing inconsistencies or small variations in the detector's structure. The flat-field is typically derived by exposing the detector to a uniform light source, ensuring the entire field is evenly illuminated. The resulting map captures the relative sensitivity of each pixel, allowing us to correct for these differences and ensure a homogeneous response across the detector.</p>\n<p>In essence, while the nonlinear correction adjusts for how individual pixels behave under varying illumination levels near saturation, the flat-field correction ensures that each pixel’s response to a uniform flux is consistent. This separation allows us to more accurately address the specific characteristics of the detector at each stage of the calibration process.</p>\n<p>I hope this explanation satisfies your curiosity 😀</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3005967,
      "author_name": "Andrey",
      "author_url": "",
      "post_date": "2024-10-03T15:12:21.140000",
      "content": "<p><a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a> , thank you for this post! Could you please also explain the meaning of the axis_info.parquet columns?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3004053,
      "author_name": "Seamus Mack",
      "author_url": "",
      "post_date": "2024-10-01T14:26:43.417000",
      "content": "<p>How might the changes in handling the dark current scaling and gain multiplication affect the accuracy of the final relative measurements + what implications might this have for future similar data analysis ? Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3005646,
          "author_name": "Lorenzo Mugnai",
          "author_url": "",
          "post_date": "2024-10-03T06:50:10.357000",
          "content": "<p>Good questions! Allow me to use some math to answer this.</p>\n<h2>Dark Current</h2>\n<p>If dark current is removed improperly, it introduces an additive bias. Our signal is ( S + c ), where ( c ) is the dark current additive term. As seen before, the transit measurement is calculated as the flux out of transit minus the flux in transit, divided by the flux out of transit. In equations, this is:</p>\n<p>$$<br>\nf_1 = \\frac{S_{oot} - S_{it}}{S_{oot}}<br>\n$$</p>\n<p>As I shown in my post, if we have the additive term, it becomes:</p>\n<p>$$<br>\nf_2 = \\frac{S_{oot} + c - S_{it} - c}{S_{oot} + c} = \\frac{S_{oot} - S_{it}}{S_{oot} + c}<br>\n$$</p>\n<p>And we end up with a biased estimator because of the changed denominator. This means that we are now measuring an incorrect transit depth.</p>\n<p>Let's focus on the noise estimation now. If we use the \"simple\" formula:</p>\n<p>$$<br>\nf_1 = \\frac{S_{oot} - S_{it}}{S_{oot}}<br>\n$$</p>\n<p>assuming all the terms are independent (are they? 😉), we get from the noise propagation:</p>\n<p>$$<br>\n\\Delta f_1 = \\frac{S_{it}}{S_{oot}^2} \\Delta S_{oot} + \\frac{1}{S_{oot}} \\Delta S_{it}<br>\n$$</p>\n<p>However, if we compute the same for the biased equation:</p>\n<p>$$<br>\n\\Delta f_2 = \\frac{(c + S_{it}) \\Delta S_{oot} + (S_{oot} + c) \\Delta S_{it} + (S_{oot} - S_{it}) \\Delta c}{(S_{oot} + c)^2}<br>\n$$</p>\n<p>This is the noise equation you \"should\" use to properly account for the dark current term. But if you corrected for the dark current improperly, you don't really know that it's there. As a result, you end up using the previous equation, which is incorrect, leading to an uncertainty that is not representative and therefore does not account for the introduced bias.</p>\n<h2>Gain Multiplication</h2>\n<p>Multiplying by the gain is necessary to convert from integer values to floating-point values, obtaining the \"true\" electron counts. It is technically the same for all frames, so doing the math again, it disappears:</p>\n<p>$$<br>\nf_3 = \\frac{g \\cdot S_{oot} - g \\cdot S_{it}}{g \\cdot S_{oot}} = \\frac{g(S_{oot} - S_{it})}{g \\cdot S_{oot}} = \\frac{S_{oot} - S_{it}}{S_{oot}} = f_1<br>\n$$</p>\n<p>This doesn't really matter for relative measurements like ours. It will matter if you want or need to perform absolute calibrations to estimate the actual flux.</p>\n<h2>Applying the Equations in Model Fitting</h2>\n<p>These calculations are for a generic fit of the transit depth. How you apply these equations depends on \"how\" the quantity is estimated. Typically, a \"model\" is fitted to the data to determine the transit depth. If the model accounts for the bias term (e.g., includes the dark current ( c ) as a parameter or ensures proper scaling), the derived equations will accurately reflect the uncertainties and biases in the measurements.</p>\n<p>In other words, when performing a fit:</p>\n<ul>\n<li>If the fitting model includes the bias term correctly, the propagation of errors will account for it, leading to more accurate and unbiased estimates of the transit depth. Such a model may be referred to as an instrument model, systematic model, or detrending model, depending on the specific components it aims to correct or remove from the science signal.</li>\n<li>If the bias is not properly accounted for in the model, the uncertainties and the estimated transit depth will be biased, potentially leading to incorrect scientific conclusions.</li>\n</ul>\n<p>Additionally, care must be taken to avoid <strong>overfitting</strong> the model to the data. This occurs when the model becomes too complex, capturing not only the underlying signal but also the noise in such details that can result in poor generalization to new observations. However, the model should not be <strong>underfitting</strong> either by being too simplistic, as this may cause loss of important scientific signals by inadvertently incorporating them into the detrending process. We have learnt from experienced that this is hard to balance. </p>\n<h2>Conclusion</h2>\n<p>These are (or should be) standard practices and considerations in science. It's not easy and might even be \"less easy\" to generalize these processes, but science is what drives us forward, so it's not meant to be easy. As I tell my students,</p>\n<blockquote>\n  <p><strong>\"If this task were easy to do now, we would have done it before when it was difficult.\"</strong></p>\n</blockquote>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2998751,
      "author_name": "Heisenger",
      "author_url": "",
      "post_date": "2024-09-26T00:58:10.500000",
      "content": "<p><a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a>  Thanks Lorenzo, i understand that we apply the linearity corrections prior to subtracting dark current - does it matter if it is done in this order rather than the other way round? (intuitively, I may have thought that we want to subtract the dark current prior to linearity corrections, as we can then first get the relevant measured current and figure out what the actual photos hitting the sensor is that generated these currents, but my understanding is likely flawed) </p>",
      "votes": 0,
      "replies": [
        {
          "id": 2999120,
          "author_name": "Lorenzo Mugnai",
          "author_url": "",
          "post_date": "2024-09-26T11:45:17.040000",
          "content": "<p>Thank you. Indeed, there are different approaches regarding the order of calibration steps, and sometimes the sequence can affect the results. In this specific case, the difference should be minimal; however, my opinion is that linearity correction should be applied before subtracting the dark current.</p>\n<p>Although the dark current is thermally generated and not in response to a light signal, it still contributes to filling the pixel well, thereby reducing the capacity to convert photons into electrons. It’s important to remember that the pixel well is filled by electrons, and the detector does not differentiate between photo-generated electrons and thermal electrons. Consequently, the observed deviation from linearity is due to both the filling of the well by incoming photons and the dark current. By applying the linearity correction first, we account for the intrinsic non-linearities of the sensor that arise from both of these factors.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2959214": "Hi everyone,\n\nI wanted to take a moment to explain some key aspects of the data we're working on within our simulations, particularly around dark frames and flat fields. There have been some great questions raised, so I hope this post clarifies things for those who are working on similar projects.\n\n## Dark Frames vs Dark Current Calibration Products\n\nIn typical astronomical data processing, dark frames are images taken with the shutter closed, capturing only the inherent electronic noise of the detector, known as dark current. These dark frames are usually subtracted from the actual observation data to remove this unwanted noise.\n\nHowever, in our case, the \"dark frames\" we provide are not traditional dark frames collected by a telescope. Instead, they are **estimates of the dark current**. Because of this, these dark current products need to be scaled according to the integration time of each frame before being subtracted from the observational data. This ensures that the dark current is accurately accounted for in relation to the exposure duration.\n\n## Flat Fields in Our Simulations\n\nSimilarly, flat fields in typical observations are used to correct for variations in the pixel-to-pixel sensitivity of the detector. Observationally, these are obtained by imaging a uniformly illuminated field. The data we provide, however, are not observational flat fields. Instead, they are **estimates of the relative efficiency** of each pixel in detecting light, which is why these \"flats\" have a mean value of one. Their purpose is to correct for variations in quantum efficiency across the detector, but because they are simulated, they might differ statistically from observational flats.\n\n## Uniformity of Calibration Products\n\nIn our data, the calibration products (darks, flats, etc.) are identical across frames. This is because we’ve simulated observations using the same detector consistently, which reflects real-life scenarios where different observations would be made using the same telescope. As a result, the same calibration products would be available throughout all observations.\n\n## Should I calibrate?\n\nCalibration is crucial in the study of planetary transits to ensure precise measurements of the brightness variations. Since transit signals are very weak, with a required accuracy of about 100 ppm, even small instrumental errors can introduce significant distortions.\n\nFor example, if there is a constant additive signal \\( C \\) in the data, the transit depth, calculated as the difference of the out of transit signal and the in transit signal divided by the signal out of transit \n$$\\frac{\\Delta S}{S} = \\frac{S_{\\text{oot}} - S_{\\text{it}}}{S_{\\text{oot}}}$$ \nwill be distorted:\n\n$$\\frac{\\Delta S}{S} = \\frac{(S_{\\text{oot}}+C) - (S_{\\text{it}}+C)}{S_{\\text{oot}} + C} = \\frac{S_{\\text{oot}} - S_{\\text{it}}}{S_{\\text{oot}} + C}$$\n\nThis causes a bias in the estimation, leading to inaccurate conclusions about the planet's characteristics. Calibration is essential to remove these errors, ensuring that the measurements accurately reflect the transit signal.\n\n## **Notes on the Provided Notebooks**\n\nWe’ve noticed a few issues in the notebook that we're currently addressing with you:\n\n+ **Dark Current**: As mentioned earlier, the dark current should be scaled according to the integration time of each pixel. This scaling was not applied correctly in previous iterations of the notebook, but we included it in the updated ones (see Notebooks listed later).\n+ **Subtracting Dark Current from Flat Fields**: This subtraction is incorrect based on the earlier discussion. The dark current should not be subtracted from the flat fields as they serve a different purpose.\n+ **Multiplication of Gain**: To correctly invert the analogue-to-digital conversion equation, the data should be divided by the gain, not multiplied. However, this doesn’t change the final result since we’re working with relative measurements.\n\nWe have updated our notebook to include these fixes, which you can find here:\n\n[**Updated Notebook**](https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data)\n\nAdditionally, we recommend reviewing our new notebook for single observation calibrations:\n\n[**Calibrating a Single Observation**](https://www.kaggle.com/code/gordonyip/calibrating-a-single-observation/notebook)\n\nI hope this explanation helps you better understand how to work with the data in our simulations. Feel free to reach out if you have any further questions or need clarification on any points!\n\nBest regards,\n",
    "2964080": "@lorenzomugnai \n\nCan I ask the technical details of `read` product?",
    "2963610": "So is the data provided uncalibrated or already calibrated to a certain degree? What units are it in; raw detector counts (dN?). Does the provided files properly flux calibrate everything to mJy? or just to Me/s?",
    "3028326": "Dear @lorenzomugnai !\n\nThank you for awesome explanations. Regarding the order of transformations in notebook:\n\nShouldn't the Flat Field Correction be done just after the linear corr? \n\nI am concerned with applying it after CDS. Does it mean that Flat was acquired using CDS too?",
    "3005967": "@lorenzomugnai , thank you for this post! Could you please also explain the meaning of the axis_info.parquet columns?",
    "3004053": "How might the changes in handling the dark current scaling and gain multiplication affect the accuracy of the final relative measurements + what implications might this have for future similar data analysis ? Thanks",
    "2998751": "@lorenzomugnai  Thanks Lorenzo, i understand that we apply the linearity corrections prior to subtracting dark current - does it matter if it is done in this order rather than the other way round? (intuitively, I may have thought that we want to subtract the dark current prior to linearity corrections, as we can then first get the relevant measured current and figure out what the actual photos hitting the sensor is that generated these currents, but my understanding is likely flawed) "
  }
}