{
  "id": 543855,
  "title": "19th Place Solution - Curve Fitting",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543855",
  "author_name": "kangourous",
  "post_date": "2024-11-01T19:56:38.687000",
  "votes": 11,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Many thanks to the Ariel team for organizing this competition ! This was my first kaggle competition and I enjoyed a lot working on this topic. <br>\nMy solution is mostly based on modelling the signal using a nonlinear function, and solving a least-squares problem using a gradient based algorithm such as the Levenberg-Marquardt algorithm.<br>\nI've just made a clean notebook of my solution here : <a href=\"https://www.kaggle.com/code/kangourous/ariel-19th-place-solution-curve-fitting\" target=\"_blank\">https://www.kaggle.com/code/kangourous/ariel-19th-place-solution-curve-fitting</a></p>\n<h2>Preprocessing</h2>\n<p>I used the organizer's notebook for calibration but accelerated the process by utilizing the GPU. For FGS, I calculated the mean of the 100 brightest pixels, and for AIRS-CH0, I took the 15 brightest pixels for each wavelength.</p>\n<h2>Transit zone</h2>\n<p>I've found that an accurate estimation of the transit boundaries for an observation was crucial to get a good estimation of the transit depth. As the transit zone is the same for all wavelength, I took the mean value for AIRS along frequency axis. Initially, I used the second derivative to spot the transit zone (thanks <a href=\"https://www.kaggle.com/code/rezanl/code-find-transition-zone-using-derivatives)\" target=\"_blank\">https://www.kaggle.com/code/rezanl/code-find-transition-zone-using-derivatives)</a>, but to do so I needed to smooth the signal which led to a loss of precision. So I tried a new approach : I've created a function \\(f(x)\\) to model the signal as the product of a polynomial function \\(p(x)\\) and a step function \\(s(x)\\) :</p>\n<p>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F0070712a06c42b1261872ae85c2e3e9d%2FScreenshot_20241101_200300.png?generation=1730487804551131&amp;alt=media\" alt=\"drawing\">\n</p>\n<p>\\(a_i\\) are the parameters of the polynomial function, \\(z_1, z_2, z_3, z_4\\) represent the key points for the transit event, and \\(d\\) represent the transit depth.<br>\nOur least square problem consists of finding the optimal value for \\(a_i\\), \\(z_i\\) and \\(d\\) so that \\(f(x)\\) fit best the data. To achieve this, I used the <code>dogbox</code> algorithm, for which <code>scipy.optimize</code> provides an implementation. This optimization process gives 2 useful information :</p>\n<ol>\n<li>the transit boundaries \\(z_i\\)</li>\n<li>the average background noise, modeled as a polynomial function of coefficient \\(a_i\\). I divided the signals by this curve to denoise globally the AIRS data.<br>\nWe will then apply the same optimization process for each wavelength for, this time, estimate the depth.</li>\n</ol>\n<p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Fbc3de9e636d5d906f1cd92f6355b55cd%2Fp(x).png?generation=1730488254786221&amp;alt=media\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F1e6f515897ca52a7d4003d9002175b0f%2Fs(x).png?generation=1730488267747307&amp;alt=media\"> \n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Ffb786d040264a1b5e19a0a222ca6d723%2Fall.png?generation=1730488281999400&amp;alt=media\">\n</p>\n<h2>Depth estimation</h2>\n<p>Now that we have the transit zone for each planet, I used \\(f(x)\\) again with the parameters \\(p_i\\) fixed, to model the signal for each wavelength. I solve the least square problem by finding the optimal \\(a_i\\) and \\(d\\) for each wavelength.  I then smoothed the result along the wavelength axis using a moving average window of size 73, weighted by a Taylor windows. This approach gave me 0.611 on LB. By analyzing sources of error, I found that my method underperformed for planets with spectra that have high standard deviation. For these planet, the moving average windows was to aggressive and flatten the spectrum. So I used a smaller window size if the standard deviation of the predicted sprectrum is higher than a threshold. I got an improvement of 0.03 on LB with this adjustment.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F6cd949a9ef3dc4bc6b538e1bf5333b67%2Fwin.png?generation=1730488878046794&amp;alt=media\" alt=\"\"></p>\n<h2>Sigma estimation</h2>\n<p>I only used the global root mean squared error for all the dataset.</p>\n<h2>Improvement</h2>\n<p>During the whole competition I tried to analyze each pixel individually, to see if it was possible to denoise, or find the transit depth at a pixel-level resolution. The thing is, it is very hard to understand what happen when you visualize the signal for a single pixel 😅. There is a lot of noise and this mysterious spike that happens everywhere. I noticed that the position of the spike was consistent across each pixels for a given planet, and there is positive and negative spikes so when we take the mean they compensate each other. I was able to model the spike using a  lorentzian function, but I did nothing with that because denoise each pixel individually was to heavy, and it seems that it didn't improve my results.</p>\n<p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F1f892ae27fd1b1da2d835bd908163c72%2Florentz.png?generation=1730489255229397&amp;alt=media\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Fd06389d503c15931e3aaf19cecce378a%2Fzoom.png?generation=1730489266853984&amp;alt=media\"> \n</p>",
  "messages": [
    {
      "id": 3034138,
      "postDate": "2024-11-01T19:56:38.687Z",
      "content": "<p>Many thanks to the Ariel team for organizing this competition ! This was my first kaggle competition and I enjoyed a lot working on this topic. <br>\nMy solution is mostly based on modelling the signal using a nonlinear function, and solving a least-squares problem using a gradient based algorithm such as the Levenberg-Marquardt algorithm.<br>\nI've just made a clean notebook of my solution here : <a href=\"https://www.kaggle.com/code/kangourous/ariel-19th-place-solution-curve-fitting\" target=\"_blank\">https://www.kaggle.com/code/kangourous/ariel-19th-place-solution-curve-fitting</a></p>\n<h2>Preprocessing</h2>\n<p>I used the organizer's notebook for calibration but accelerated the process by utilizing the GPU. For FGS, I calculated the mean of the 100 brightest pixels, and for AIRS-CH0, I took the 15 brightest pixels for each wavelength.</p>\n<h2>Transit zone</h2>\n<p>I've found that an accurate estimation of the transit boundaries for an observation was crucial to get a good estimation of the transit depth. As the transit zone is the same for all wavelength, I took the mean value for AIRS along frequency axis. Initially, I used the second derivative to spot the transit zone (thanks <a href=\"https://www.kaggle.com/code/rezanl/code-find-transition-zone-using-derivatives)\" target=\"_blank\">https://www.kaggle.com/code/rezanl/code-find-transition-zone-using-derivatives)</a>, but to do so I needed to smooth the signal which led to a loss of precision. So I tried a new approach : I've created a function \\(f(x)\\) to model the signal as the product of a polynomial function \\(p(x)\\) and a step function \\(s(x)\\) :</p>\n<p>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F0070712a06c42b1261872ae85c2e3e9d%2FScreenshot_20241101_200300.png?generation=1730487804551131&amp;alt=media\" alt=\"drawing\">\n</p>\n<p>\\(a_i\\) are the parameters of the polynomial function, \\(z_1, z_2, z_3, z_4\\) represent the key points for the transit event, and \\(d\\) represent the transit depth.<br>\nOur least square problem consists of finding the optimal value for \\(a_i\\), \\(z_i\\) and \\(d\\) so that \\(f(x)\\) fit best the data. To achieve this, I used the <code>dogbox</code> algorithm, for which <code>scipy.optimize</code> provides an implementation. This optimization process gives 2 useful information :</p>\n<ol>\n<li>the transit boundaries \\(z_i\\)</li>\n<li>the average background noise, modeled as a polynomial function of coefficient \\(a_i\\). I divided the signals by this curve to denoise globally the AIRS data.<br>\nWe will then apply the same optimization process for each wavelength for, this time, estimate the depth.</li>\n</ol>\n<p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Fbc3de9e636d5d906f1cd92f6355b55cd%2Fp(x).png?generation=1730488254786221&amp;alt=media\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F1e6f515897ca52a7d4003d9002175b0f%2Fs(x).png?generation=1730488267747307&amp;alt=media\"> \n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Ffb786d040264a1b5e19a0a222ca6d723%2Fall.png?generation=1730488281999400&amp;alt=media\">\n</p>\n<h2>Depth estimation</h2>\n<p>Now that we have the transit zone for each planet, I used \\(f(x)\\) again with the parameters \\(p_i\\) fixed, to model the signal for each wavelength. I solve the least square problem by finding the optimal \\(a_i\\) and \\(d\\) for each wavelength.  I then smoothed the result along the wavelength axis using a moving average window of size 73, weighted by a Taylor windows. This approach gave me 0.611 on LB. By analyzing sources of error, I found that my method underperformed for planets with spectra that have high standard deviation. For these planet, the moving average windows was to aggressive and flatten the spectrum. So I used a smaller window size if the standard deviation of the predicted sprectrum is higher than a threshold. I got an improvement of 0.03 on LB with this adjustment.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F6cd949a9ef3dc4bc6b538e1bf5333b67%2Fwin.png?generation=1730488878046794&amp;alt=media\" alt=\"\"></p>\n<h2>Sigma estimation</h2>\n<p>I only used the global root mean squared error for all the dataset.</p>\n<h2>Improvement</h2>\n<p>During the whole competition I tried to analyze each pixel individually, to see if it was possible to denoise, or find the transit depth at a pixel-level resolution. The thing is, it is very hard to understand what happen when you visualize the signal for a single pixel 😅. There is a lot of noise and this mysterious spike that happens everywhere. I noticed that the position of the spike was consistent across each pixels for a given planet, and there is positive and negative spikes so when we take the mean they compensate each other. I was able to model the spike using a  lorentzian function, but I did nothing with that because denoise each pixel individually was to heavy, and it seems that it didn't improve my results.</p>\n<p>\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F1f892ae27fd1b1da2d835bd908163c72%2Florentz.png?generation=1730489255229397&amp;alt=media\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Fd06389d503c15931e3aaf19cecce378a%2Fzoom.png?generation=1730489266853984&amp;alt=media\"> \n</p>",
      "rawMarkdown": "Many thanks to the Ariel team for organizing this competition ! This was my first kaggle competition and I enjoyed a lot working on this topic. \nMy solution is mostly based on modelling the signal using a nonlinear function, and solving a least-squares problem using a gradient based algorithm such as the Levenberg-Marquardt algorithm.\nI've just made a clean notebook of my solution here : https://www.kaggle.com/code/kangourous/ariel-19th-place-solution-curve-fitting\n\n\n## Preprocessing\nI used the organizer's notebook for calibration but accelerated the process by utilizing the GPU. For FGS, I calculated the mean of the 100 brightest pixels, and for AIRS-CH0, I took the 15 brightest pixels for each wavelength.\n\n## Transit zone\nI've found that an accurate estimation of the transit boundaries for an observation was crucial to get a good estimation of the transit depth. As the transit zone is the same for all wavelength, I took the mean value for AIRS along frequency axis. Initially, I used the second derivative to spot the transit zone (thanks https://www.kaggle.com/code/rezanl/code-find-transition-zone-using-derivatives), but to do so I needed to smooth the signal which led to a loss of precision. So I tried a new approach : I've created a function \\\\(f(x)\\\\) to model the signal as the product of a polynomial function \\\\(p(x)\\\\) and a step function \\\\(s(x)\\\\) :\n\n<p align=\"center\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F0070712a06c42b1261872ae85c2e3e9d%2FScreenshot_20241101_200300.png?generation=1730487804551131&alt=media\" alt=\"drawing\" width=\"500\" />\n</p>\n\n\\\\(a_i\\\\) are the parameters of the polynomial function, \\\\(z_1, z_2, z_3, z_4\\\\) represent the key points for the transit event, and \\\\(d\\\\) represent the transit depth.\nOur least square problem consists of finding the optimal value for \\\\(a_i\\\\), \\\\(z_i\\\\) and \\\\(d\\\\) so that \\\\(f(x)\\\\) fit best the data. To achieve this, I used the `dogbox` algorithm, for which `scipy.optimize` provides an implementation. This optimization process gives 2 useful information :\n1.  the transit boundaries \\\\(z_i\\\\)\n2. the average background noise, modeled as a polynomial function of coefficient \\\\(a_i\\\\). I divided the signals by this curve to denoise globally the AIRS data.\nWe will then apply the same optimization process for each wavelength for, this time, estimate the depth.\n\n<p float=\"left\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Fbc3de9e636d5d906f1cd92f6355b55cd%2Fp(x).png?generation=1730488254786221&alt=media\" width=\"450\" />\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F1e6f515897ca52a7d4003d9002175b0f%2Fs(x).png?generation=1730488267747307&alt=media\" width=\"450\" /> \n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Ffb786d040264a1b5e19a0a222ca6d723%2Fall.png?generation=1730488281999400&alt=media\" width=\"700\" />\n</p>\n\n## Depth estimation\nNow that we have the transit zone for each planet, I used \\\\(f(x)\\\\) again with the parameters \\\\(p_i\\\\) fixed, to model the signal for each wavelength. I solve the least square problem by finding the optimal \\\\(a_i\\\\) and \\\\(d\\\\) for each wavelength.  I then smoothed the result along the wavelength axis using a moving average window of size 73, weighted by a Taylor windows. This approach gave me 0.611 on LB. By analyzing sources of error, I found that my method underperformed for planets with spectra that have high standard deviation. For these planet, the moving average windows was to aggressive and flatten the spectrum. So I used a smaller window size if the standard deviation of the predicted sprectrum is higher than a threshold. I got an improvement of 0.03 on LB with this adjustment.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F6cd949a9ef3dc4bc6b538e1bf5333b67%2Fwin.png?generation=1730488878046794&alt=media)\n\n## Sigma estimation\nI only used the global root mean squared error for all the dataset.\n\n## Improvement\nDuring the whole competition I tried to analyze each pixel individually, to see if it was possible to denoise, or find the transit depth at a pixel-level resolution. The thing is, it is very hard to understand what happen when you visualize the signal for a single pixel 😅. There is a lot of noise and this mysterious spike that happens everywhere. I noticed that the position of the spike was consistent across each pixels for a given planet, and there is positive and negative spikes so when we take the mean they compensate each other. I was able to model the spike using a  lorentzian function, but I did nothing with that because denoise each pixel individually was to heavy, and it seems that it didn't improve my results.\n\n<p float=\"left\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F1f892ae27fd1b1da2d835bd908163c72%2Florentz.png?generation=1730489255229397&alt=media\" width=\"450\" />\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Fd06389d503c15931e3aaf19cecce378a%2Fzoom.png?generation=1730489266853984&alt=media\" width=\"450\" /> \n</p>",
      "votes": 11
    },
    {
      "id": 3034303,
      "postDate": "2024-11-02T01:59:19.607Z",
      "content": "<p>Thr 'spike' is not mysterious at all, it is the noise from the sensors getting hit and trying to refocus (or something along this line, the host explained it somewhere better) in practice, what happens is that the signal is no longer centered around the middle but is shifted slightly. This is why there are positive and negative spikes- the positive is on the side that the center of the signal moved to, and the negative is on the other side. But as you noted, when we sum on all the pixels for a certain wavelength, they cancel each other since we still capture all the signal. It's just that its mean is slightly shifted.</p>",
      "rawMarkdown": "Thr 'spike' is not mysterious at all, it is the noise from the sensors getting hit and trying to refocus (or something along this line, the host explained it somewhere better) in practice, what happens is that the signal is no longer centered around the middle but is shifted slightly. This is why there are positive and negative spikes- the positive is on the side that the center of the signal moved to, and the negative is on the other side. But as you noted, when we sum on all the pixels for a certain wavelength, they cancel each other since we still capture all the signal. It's just that its mean is slightly shifted.",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3034303,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-11-02T01:59:19.607000",
      "content": "<p>Thr 'spike' is not mysterious at all, it is the noise from the sensors getting hit and trying to refocus (or something along this line, the host explained it somewhere better) in practice, what happens is that the signal is no longer centered around the middle but is shifted slightly. This is why there are positive and negative spikes- the positive is on the side that the center of the signal moved to, and the negative is on the other side. But as you noted, when we sum on all the pixels for a certain wavelength, they cancel each other since we still capture all the signal. It's just that its mean is slightly shifted.</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3034138": "Many thanks to the Ariel team for organizing this competition ! This was my first kaggle competition and I enjoyed a lot working on this topic. \nMy solution is mostly based on modelling the signal using a nonlinear function, and solving a least-squares problem using a gradient based algorithm such as the Levenberg-Marquardt algorithm.\nI've just made a clean notebook of my solution here : https://www.kaggle.com/code/kangourous/ariel-19th-place-solution-curve-fitting\n\n\n## Preprocessing\nI used the organizer's notebook for calibration but accelerated the process by utilizing the GPU. For FGS, I calculated the mean of the 100 brightest pixels, and for AIRS-CH0, I took the 15 brightest pixels for each wavelength.\n\n## Transit zone\nI've found that an accurate estimation of the transit boundaries for an observation was crucial to get a good estimation of the transit depth. As the transit zone is the same for all wavelength, I took the mean value for AIRS along frequency axis. Initially, I used the second derivative to spot the transit zone (thanks https://www.kaggle.com/code/rezanl/code-find-transition-zone-using-derivatives), but to do so I needed to smooth the signal which led to a loss of precision. So I tried a new approach : I've created a function \\\\(f(x)\\\\) to model the signal as the product of a polynomial function \\\\(p(x)\\\\) and a step function \\\\(s(x)\\\\) :\n\n<p align=\"center\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F0070712a06c42b1261872ae85c2e3e9d%2FScreenshot_20241101_200300.png?generation=1730487804551131&alt=media\" alt=\"drawing\" width=\"500\" />\n</p>\n\n\\\\(a_i\\\\) are the parameters of the polynomial function, \\\\(z_1, z_2, z_3, z_4\\\\) represent the key points for the transit event, and \\\\(d\\\\) represent the transit depth.\nOur least square problem consists of finding the optimal value for \\\\(a_i\\\\), \\\\(z_i\\\\) and \\\\(d\\\\) so that \\\\(f(x)\\\\) fit best the data. To achieve this, I used the `dogbox` algorithm, for which `scipy.optimize` provides an implementation. This optimization process gives 2 useful information :\n1.  the transit boundaries \\\\(z_i\\\\)\n2. the average background noise, modeled as a polynomial function of coefficient \\\\(a_i\\\\). I divided the signals by this curve to denoise globally the AIRS data.\nWe will then apply the same optimization process for each wavelength for, this time, estimate the depth.\n\n<p float=\"left\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Fbc3de9e636d5d906f1cd92f6355b55cd%2Fp(x).png?generation=1730488254786221&alt=media\" width=\"450\" />\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F1e6f515897ca52a7d4003d9002175b0f%2Fs(x).png?generation=1730488267747307&alt=media\" width=\"450\" /> \n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Ffb786d040264a1b5e19a0a222ca6d723%2Fall.png?generation=1730488281999400&alt=media\" width=\"700\" />\n</p>\n\n## Depth estimation\nNow that we have the transit zone for each planet, I used \\\\(f(x)\\\\) again with the parameters \\\\(p_i\\\\) fixed, to model the signal for each wavelength. I solve the least square problem by finding the optimal \\\\(a_i\\\\) and \\\\(d\\\\) for each wavelength.  I then smoothed the result along the wavelength axis using a moving average window of size 73, weighted by a Taylor windows. This approach gave me 0.611 on LB. By analyzing sources of error, I found that my method underperformed for planets with spectra that have high standard deviation. For these planet, the moving average windows was to aggressive and flatten the spectrum. So I used a smaller window size if the standard deviation of the predicted sprectrum is higher than a threshold. I got an improvement of 0.03 on LB with this adjustment.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F6cd949a9ef3dc4bc6b538e1bf5333b67%2Fwin.png?generation=1730488878046794&alt=media)\n\n## Sigma estimation\nI only used the global root mean squared error for all the dataset.\n\n## Improvement\nDuring the whole competition I tried to analyze each pixel individually, to see if it was possible to denoise, or find the transit depth at a pixel-level resolution. The thing is, it is very hard to understand what happen when you visualize the signal for a single pixel 😅. There is a lot of noise and this mysterious spike that happens everywhere. I noticed that the position of the spike was consistent across each pixels for a given planet, and there is positive and negative spikes so when we take the mean they compensate each other. I was able to model the spike using a  lorentzian function, but I did nothing with that because denoise each pixel individually was to heavy, and it seems that it didn't improve my results.\n\n<p float=\"left\">\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2F1f892ae27fd1b1da2d835bd908163c72%2Florentz.png?generation=1730489255229397&alt=media\" width=\"450\" />\n  <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8810589%2Fd06389d503c15931e3aaf19cecce378a%2Fzoom.png?generation=1730489266853984&alt=media\" width=\"450\" /> \n</p>",
    "3034303": "Thr 'spike' is not mysterious at all, it is the noise from the sensors getting hit and trying to refocus (or something along this line, the host explained it somewhere better) in practice, what happens is that the signal is no longer centered around the middle but is shifted slightly. This is why there are positive and negative spikes- the positive is on the side that the center of the signal moved to, and the negative is on the other side. But as you noted, when we sum on all the pixels for a certain wavelength, they cancel each other since we still capture all the signal. It's just that its mean is slightly shifted."
  }
}