{
  "id": 544317,
  "title": "1st Place Solution",
  "url": "/competitions/ariel-data-challenge-2024/discussion/544317",
  "author_name": "daiwakun",
  "post_date": "2024-11-04T13:30:14.875000",
  "votes": 51,
  "comment_count": 16,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/code/cnumber/neurips-ariel-data-challenge-2024-final-submission\" target=\"_blank\">link to notebook</a></p>\n<h1>Introduction</h1>\n<p>We ( <a href=\"https://www.kaggle.com/daiwakun\" target=\"_blank\">@daiwakun</a> and <a href=\"https://www.kaggle.com/cnumber\" target=\"_blank\">@cnumber</a>) would like to thank the organizers for hosting this fascinating competition, especially <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> and <a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a>, who worked tirelessly so that competitors could focus on the more \"interesting\" parts.</p>\n<p>Also, a huge applause to <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> for maintaining first place for almost the entire duration of the competition. Although we managed to grab victory at the end of the race, we are confident that we would have lost had the competition period been one day shorter or one day longer.</p>\n<p>Many ideas in our solution were found during a deep examination of the ExoSim2 and TauREx3 code used for data generation, including gain drift fitting and foreground processing. Although we didn't (or couldn't) find any leaks, we have to admit that our solution somewhat \"hacks\" the simulator. Nevertheless, we hope the hosts can make use of some aspects of our solution, and hopefully, it would be useful for the actual data.</p>\n<p>For those who are curious about our significant score improvement during the last two days of the competition, we have decided to present what was going on in our team in <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/544316\" target=\"_blank\">another post</a></p>\n<h1>Solution</h1>\n<h2>Signal Preprocessing</h2>\n<p>Only the AIRS-CH0 channel was used because our solution heavily relies on the correlation between adjacent wavelengths, and also because we were unsure how to effectively utilize the FGS1 channel data due to issues with its distorted and fluctuating point spread function.</p>\n<p>The public notebook (we wanted to put a link, but we couldn't find the notebook we used) was used with the following changes:</p>\n<ol>\n<li><p><strong>Disabling Hot Pixel Processing</strong>: This led to a significant jump on the leaderboard. We assume that the hot pixel processing leads to unwanted information loss of the pixels in the middle of the sensor, making it almost impossible to correct the noise introduced by the time-dependent PSF distortion (see image below). Perhaps it is the sigma clip algorithm that is doing the wrong thing and that there exist algorithms that can handle hot pixels properly, but we weren't able to find them.</p></li>\n<li><p><strong>Foreground Processing</strong>: Some teams have noticed that multiplying the final spectrum by a coefficient around 1.006~1.008 gives a huge boost. This is because a foreground is added to the signal during the ExoSim2 simulation [link to GitHub]. The correct way to handle this is to estimate the wavelength-dependent foreground signal and to subtract it from the signal in the central region. In our solution, we decided to estimate the foreground signal with the regions <code>[0:8]</code> and <code>[24:32]</code>, and subtracted it from the central region <code>[8:24]</code>. The effect of considering the foreground properly can be seen in <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543853#3034987\" target=\"_blank\">this discussion post</a>.</p></li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2Fb07dae4e95a61af232cd678be36d516e%2Fbackground.png?generation=1730726806041614&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2F63105d03f4122d7967b25ca8a546853c%2Fpixel.png?generation=1730726844414945&amp;alt=media\" alt=\"\"></p>\n<h2>Initial Dip Estimation for Each Wavelength</h2>\n<h3>Identifying the Transit Interval</h3>\n<p>For identifying the time location of the transit, the wavelength-averaged signal was used. After estimating the approximate time position using a rule-based algorithm, the exact time was identified with a fitting algorithm.</p>\n<h3>Gain Drift</h3>\n<p>Gain drift refers to the variation in the detector's gain over time and across different wavelengths. These drifts can introduce systematic errors in the observed signal and must be corrected to accurately estimate the transit dip.</p>\n<p>ExoSim2 models the gain drift as:</p>\n<p>$$<br>\n(1 + f(t) \\cdot g(\\lambda))<br>\n$$</p>\n<p>where \\( f(t) \\) and \\( g(\\lambda) \\) are polynomial functions.</p>\n<p>It is notable that the final function form is different from a two-variable polynomial \\( h(t, \\lambda) \\), because while the wavelength and time components can be separated in the former, they cannot in the latter.</p>\n<p>By using this functional form directly in the fitting described below, we were able to achieve high performance in the dip estimation.</p>\n<h3>Dip Estimation with Gain Drift Fitting</h3>\n<p>The function used for fitting is as follows:</p>\n<p>$$<br>\ny_{\\text{pred}} = I(\\lambda) \\times \\text{Box}(\\lambda) \\times (1 + f(t) \\cdot g(\\lambda))<br>\n$$</p>\n<p>where \\( I(\\lambda) \\) is the spectrum of the star, and \\( \\text{Box}(\\lambda) \\) describes the dip; it is 1 outside the transit and \\( 1 - d_\\lambda \\) inside the transit.</p>\n<p>The number of fitting parameters are as follows:</p>\n<ul>\n<li>\\( f(t) \\): 5 parameters</li>\n<li>\\( g(\\lambda) \\): 5 parameters</li>\n</ul>\n<p>It is worth mentioning that the optimal \\( I(\\lambda) \\) and the optimal dip used in \\( \\text{Box}(\\lambda) \\), which minimize the mean squared error, can be found analytically if the other parameters are fixed, and they don't need to be considered by the fitting algorithm.</p>\n<p>To stabilize the fitting process, a two-stage fitting was used. In the first stage, \\( I(\\lambda) \\) was directly estimated from the raw signal with temporal averaging and fixed during the fitting. In the second stage, \\( I(\\lambda) \\) was not fixed and was set optimally for each step during the fitting, while for the fitting parameters, the result from the first stage was used as the initial parameter. The error for each data point used in the fitting was estimated from the variation of the signal at each wavelength.</p>\n<h3>Dip Error Estimation with Bootstrapping</h3>\n<p>The errors in the dip estimation at each wavelength differ due to the varying signal-to-noise ratios associated with each wavelength. <a href=\"https://en.wikipedia.org/wiki/Bootstrapping_(statistics)\" target=\"_blank\">Bootstrapping</a> was used to overcome this problem and to estimate the dip error of each wavelength.</p>\n<p>Further details are available in our code.</p>\n<h2>Dip Estimation Considering Wavelength Correlations</h2>\n<p>Three models were used: Gaussian Process Regression, AutoEncoder, and Non-negative Matrix Factorization (NMF). They were ensembled with the ratio 6:2:2.</p>\n<h3>Gaussian Process Regression</h3>\n<p>Nothing fancy, unlike the second-place solution.</p>\n<p>A simple kernel composed of RBF and Matern kernels was employed, and the errors calculated by bootstrapping were passed to <code>sklearn.gaussian_process.GaussianProcessRegressor</code> so that the model can consider the uncertainty of each data point.</p>\n<h3>AutoEncoder</h3>\n<p>We applied an autoencoder to capture relationships in the data.</p>\n<p>Unlike PCA, which captures linear relationships, autoencoders can model more complex and nonlinear patterns.</p>\n<p>Also, we expected that by training the model with MSE loss, the autoencoder model could take noise into account and recover the \"optimal\" spectrum.</p>\n<p>Important points were to normalize the data for each exoplanet and to take the moving median of the dip spectrum to smooth the input dip spectrum.</p>\n<p>The model we used is as below, with 4 nodes in the hidden layer:</p>\n<pre><code>input_data = Input(shape=(input_dim,))\nencoded = Dense(encoding_dim, activation=)(input_data)\ndecoded = Dense(input_dim, activation=)(encoded)\n\nautoencoder = Model(input_data, decoded)\n</code></pre>\n<h3>NMF</h3>\n<p>Very similar to the autoencoder, but we added it to enhance the diversity.</p>\n<p>Unlike with the autoencoder, the number of ranks was set to 5 for NMF.</p>\n<h3>Spectral Components Identified by NMF</h3>\n<p>The plot below shows the identified spectral components by NMF for the training data.</p>\n<p>It can be seen that the main contributing gases of each component are:</p>\n<ul>\n<li><strong>Component 1</strong>: CO₂</li>\n<li><strong>Component 2</strong>: CH₄</li>\n<li><strong>Component 3</strong>: H₂O</li>\n</ul>\n<p>This signifies NMF's ability to consider the correlation of the spectrum unsupervised.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2Fbc7f93b100cae6b768c0c17a04ff10a8%2Fnmf_spectral_components.png?generation=1730726867605917&amp;alt=media\" alt=\"\"></p>\n<h3>Sigma</h3>\n<p>We took the weighted average of the following components:</p>\n<ol>\n<li><strong>Constant value</strong> (planet and wavelength independent)</li>\n<li><strong>Standard deviation of the smoothed predicted dip spectrum</strong> (planet dependent, wavelength independent)</li>\n<li><strong>Uncertainty predicted by Gaussian Process Regression</strong> (planet and wavelength dependent)</li>\n</ol>\n<p>The constant value played an important role in cases where the dip spectrum was almost constant but had some bias.</p>\n<h2>The Contribution of Each Component</h2>\n<p>Since our solution incorporates a variety of ideas, we examined the contributions for those that seem important with late submissions.</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Public</th>\n<th>Private</th>\n<th>Public Loss</th>\n<th>Private Loss</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Final Submission</strong></td>\n<td>0.7330321</td>\n<td>0.7420624</td>\n<td>0.0000000</td>\n<td>0.0000000</td>\n</tr>\n<tr>\n<td>Only Gaussian Process Regression</td>\n<td>0.7221480</td>\n<td>0.7343485</td>\n<td>-0.0108841</td>\n<td>-0.0077139</td>\n</tr>\n<tr>\n<td>Only AutoEncoder</td>\n<td>0.7078056</td>\n<td>0.7181137</td>\n<td>-0.0252265</td>\n<td>-0.0239487</td>\n</tr>\n<tr>\n<td>Only NMF</td>\n<td>0.7017943</td>\n<td>0.7122631</td>\n<td>-0.0312378</td>\n<td>-0.0297993</td>\n</tr>\n<tr>\n<td>With Hot Pixel Processing</td>\n<td>0.7021653</td>\n<td>0.7224989</td>\n<td>-0.0308668</td>\n<td>-0.0195635</td>\n</tr>\n<tr>\n<td>Without Foreground Processing<br>with *1.008 to prediction</td>\n<td>0.7225193</td>\n<td>0.7298121</td>\n<td>-0.0105128</td>\n<td>-0.0122503</td>\n</tr>\n</tbody>\n</table>\n<h2>What Didn't Work</h2>\n<h3>Absorption Spectra of Known Gases</h3>\n<p>We tried to use TauREx3 to fit the dip spectrum.</p>\n<p>Although it worked really well on the training data, the leaderboard score was horrible, presumably because of the shift in the composition of the atmosphere.</p>\n<p>We did try to identify the gases in the test data, but without any success.</p>\n<h3>Denoising the Input Data with Machine Learning</h3>\n<p>It is quite hard to assume that taking the sum of the <code>[8:24]</code> channels is optimal.</p>\n<p>In some cases, the optimal range could be <code>[9:23]</code> or <code>[10:22]</code>, or even taking a weighted sum could be optimal.</p>\n<p>To resolve this problem, we tried several machine learning methods but didn't manage to beat <code>[8:24]</code>, probably due to the temporal spatial fluctuation of the signal.</p>\n<h2>What We Wanted to Do If We Had More Time</h2>\n<ul>\n<li>Utilize the FGS1 channel.</li>\n</ul>\n<h1>Final Comments</h1>\n<p>Deep analysis of the simulation code led to critical ideas needed for our victory. The function for the gain drift or the foreground processing method couldn't have been found without observing the ExoSim2 code. Though we had believed that this is the destiny of competitions that use simulation-generated data, we were very surprised that some of the top teams were able to achieve high scores without using them, so again, a big applause to them.</p>",
  "messages": [
    {
      "id": 3036336,
      "postDate": "2024-11-04T13:30:14.877Z",
      "content": "<p><a href=\"https://www.kaggle.com/code/cnumber/neurips-ariel-data-challenge-2024-final-submission\" target=\"_blank\">link to notebook</a></p>\n<h1>Introduction</h1>\n<p>We ( <a href=\"https://www.kaggle.com/daiwakun\" target=\"_blank\">@daiwakun</a> and <a href=\"https://www.kaggle.com/cnumber\" target=\"_blank\">@cnumber</a>) would like to thank the organizers for hosting this fascinating competition, especially <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> and <a href=\"https://www.kaggle.com/lorenzomugnai\" target=\"_blank\">@lorenzomugnai</a>, who worked tirelessly so that competitors could focus on the more \"interesting\" parts.</p>\n<p>Also, a huge applause to <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> for maintaining first place for almost the entire duration of the competition. Although we managed to grab victory at the end of the race, we are confident that we would have lost had the competition period been one day shorter or one day longer.</p>\n<p>Many ideas in our solution were found during a deep examination of the ExoSim2 and TauREx3 code used for data generation, including gain drift fitting and foreground processing. Although we didn't (or couldn't) find any leaks, we have to admit that our solution somewhat \"hacks\" the simulator. Nevertheless, we hope the hosts can make use of some aspects of our solution, and hopefully, it would be useful for the actual data.</p>\n<p>For those who are curious about our significant score improvement during the last two days of the competition, we have decided to present what was going on in our team in <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/544316\" target=\"_blank\">another post</a></p>\n<h1>Solution</h1>\n<h2>Signal Preprocessing</h2>\n<p>Only the AIRS-CH0 channel was used because our solution heavily relies on the correlation between adjacent wavelengths, and also because we were unsure how to effectively utilize the FGS1 channel data due to issues with its distorted and fluctuating point spread function.</p>\n<p>The public notebook (we wanted to put a link, but we couldn't find the notebook we used) was used with the following changes:</p>\n<ol>\n<li><p><strong>Disabling Hot Pixel Processing</strong>: This led to a significant jump on the leaderboard. We assume that the hot pixel processing leads to unwanted information loss of the pixels in the middle of the sensor, making it almost impossible to correct the noise introduced by the time-dependent PSF distortion (see image below). Perhaps it is the sigma clip algorithm that is doing the wrong thing and that there exist algorithms that can handle hot pixels properly, but we weren't able to find them.</p></li>\n<li><p><strong>Foreground Processing</strong>: Some teams have noticed that multiplying the final spectrum by a coefficient around 1.006~1.008 gives a huge boost. This is because a foreground is added to the signal during the ExoSim2 simulation [link to GitHub]. The correct way to handle this is to estimate the wavelength-dependent foreground signal and to subtract it from the signal in the central region. In our solution, we decided to estimate the foreground signal with the regions <code>[0:8]</code> and <code>[24:32]</code>, and subtracted it from the central region <code>[8:24]</code>. The effect of considering the foreground properly can be seen in <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543853#3034987\" target=\"_blank\">this discussion post</a>.</p></li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2Fb07dae4e95a61af232cd678be36d516e%2Fbackground.png?generation=1730726806041614&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2F63105d03f4122d7967b25ca8a546853c%2Fpixel.png?generation=1730726844414945&amp;alt=media\" alt=\"\"></p>\n<h2>Initial Dip Estimation for Each Wavelength</h2>\n<h3>Identifying the Transit Interval</h3>\n<p>For identifying the time location of the transit, the wavelength-averaged signal was used. After estimating the approximate time position using a rule-based algorithm, the exact time was identified with a fitting algorithm.</p>\n<h3>Gain Drift</h3>\n<p>Gain drift refers to the variation in the detector's gain over time and across different wavelengths. These drifts can introduce systematic errors in the observed signal and must be corrected to accurately estimate the transit dip.</p>\n<p>ExoSim2 models the gain drift as:</p>\n<p>$$<br>\n(1 + f(t) \\cdot g(\\lambda))<br>\n$$</p>\n<p>where \\( f(t) \\) and \\( g(\\lambda) \\) are polynomial functions.</p>\n<p>It is notable that the final function form is different from a two-variable polynomial \\( h(t, \\lambda) \\), because while the wavelength and time components can be separated in the former, they cannot in the latter.</p>\n<p>By using this functional form directly in the fitting described below, we were able to achieve high performance in the dip estimation.</p>\n<h3>Dip Estimation with Gain Drift Fitting</h3>\n<p>The function used for fitting is as follows:</p>\n<p>$$<br>\ny_{\\text{pred}} = I(\\lambda) \\times \\text{Box}(\\lambda) \\times (1 + f(t) \\cdot g(\\lambda))<br>\n$$</p>\n<p>where \\( I(\\lambda) \\) is the spectrum of the star, and \\( \\text{Box}(\\lambda) \\) describes the dip; it is 1 outside the transit and \\( 1 - d_\\lambda \\) inside the transit.</p>\n<p>The number of fitting parameters are as follows:</p>\n<ul>\n<li>\\( f(t) \\): 5 parameters</li>\n<li>\\( g(\\lambda) \\): 5 parameters</li>\n</ul>\n<p>It is worth mentioning that the optimal \\( I(\\lambda) \\) and the optimal dip used in \\( \\text{Box}(\\lambda) \\), which minimize the mean squared error, can be found analytically if the other parameters are fixed, and they don't need to be considered by the fitting algorithm.</p>\n<p>To stabilize the fitting process, a two-stage fitting was used. In the first stage, \\( I(\\lambda) \\) was directly estimated from the raw signal with temporal averaging and fixed during the fitting. In the second stage, \\( I(\\lambda) \\) was not fixed and was set optimally for each step during the fitting, while for the fitting parameters, the result from the first stage was used as the initial parameter. The error for each data point used in the fitting was estimated from the variation of the signal at each wavelength.</p>\n<h3>Dip Error Estimation with Bootstrapping</h3>\n<p>The errors in the dip estimation at each wavelength differ due to the varying signal-to-noise ratios associated with each wavelength. <a href=\"https://en.wikipedia.org/wiki/Bootstrapping_(statistics)\" target=\"_blank\">Bootstrapping</a> was used to overcome this problem and to estimate the dip error of each wavelength.</p>\n<p>Further details are available in our code.</p>\n<h2>Dip Estimation Considering Wavelength Correlations</h2>\n<p>Three models were used: Gaussian Process Regression, AutoEncoder, and Non-negative Matrix Factorization (NMF). They were ensembled with the ratio 6:2:2.</p>\n<h3>Gaussian Process Regression</h3>\n<p>Nothing fancy, unlike the second-place solution.</p>\n<p>A simple kernel composed of RBF and Matern kernels was employed, and the errors calculated by bootstrapping were passed to <code>sklearn.gaussian_process.GaussianProcessRegressor</code> so that the model can consider the uncertainty of each data point.</p>\n<h3>AutoEncoder</h3>\n<p>We applied an autoencoder to capture relationships in the data.</p>\n<p>Unlike PCA, which captures linear relationships, autoencoders can model more complex and nonlinear patterns.</p>\n<p>Also, we expected that by training the model with MSE loss, the autoencoder model could take noise into account and recover the \"optimal\" spectrum.</p>\n<p>Important points were to normalize the data for each exoplanet and to take the moving median of the dip spectrum to smooth the input dip spectrum.</p>\n<p>The model we used is as below, with 4 nodes in the hidden layer:</p>\n<pre><code>input_data = Input(shape=(input_dim,))\nencoded = Dense(encoding_dim, activation=)(input_data)\ndecoded = Dense(input_dim, activation=)(encoded)\n\nautoencoder = Model(input_data, decoded)\n</code></pre>\n<h3>NMF</h3>\n<p>Very similar to the autoencoder, but we added it to enhance the diversity.</p>\n<p>Unlike with the autoencoder, the number of ranks was set to 5 for NMF.</p>\n<h3>Spectral Components Identified by NMF</h3>\n<p>The plot below shows the identified spectral components by NMF for the training data.</p>\n<p>It can be seen that the main contributing gases of each component are:</p>\n<ul>\n<li><strong>Component 1</strong>: CO₂</li>\n<li><strong>Component 2</strong>: CH₄</li>\n<li><strong>Component 3</strong>: H₂O</li>\n</ul>\n<p>This signifies NMF's ability to consider the correlation of the spectrum unsupervised.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2Fbc7f93b100cae6b768c0c17a04ff10a8%2Fnmf_spectral_components.png?generation=1730726867605917&amp;alt=media\" alt=\"\"></p>\n<h3>Sigma</h3>\n<p>We took the weighted average of the following components:</p>\n<ol>\n<li><strong>Constant value</strong> (planet and wavelength independent)</li>\n<li><strong>Standard deviation of the smoothed predicted dip spectrum</strong> (planet dependent, wavelength independent)</li>\n<li><strong>Uncertainty predicted by Gaussian Process Regression</strong> (planet and wavelength dependent)</li>\n</ol>\n<p>The constant value played an important role in cases where the dip spectrum was almost constant but had some bias.</p>\n<h2>The Contribution of Each Component</h2>\n<p>Since our solution incorporates a variety of ideas, we examined the contributions for those that seem important with late submissions.</p>\n<table>\n<thead>\n<tr>\n<th>Method</th>\n<th>Public</th>\n<th>Private</th>\n<th>Public Loss</th>\n<th>Private Loss</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>Final Submission</strong></td>\n<td>0.7330321</td>\n<td>0.7420624</td>\n<td>0.0000000</td>\n<td>0.0000000</td>\n</tr>\n<tr>\n<td>Only Gaussian Process Regression</td>\n<td>0.7221480</td>\n<td>0.7343485</td>\n<td>-0.0108841</td>\n<td>-0.0077139</td>\n</tr>\n<tr>\n<td>Only AutoEncoder</td>\n<td>0.7078056</td>\n<td>0.7181137</td>\n<td>-0.0252265</td>\n<td>-0.0239487</td>\n</tr>\n<tr>\n<td>Only NMF</td>\n<td>0.7017943</td>\n<td>0.7122631</td>\n<td>-0.0312378</td>\n<td>-0.0297993</td>\n</tr>\n<tr>\n<td>With Hot Pixel Processing</td>\n<td>0.7021653</td>\n<td>0.7224989</td>\n<td>-0.0308668</td>\n<td>-0.0195635</td>\n</tr>\n<tr>\n<td>Without Foreground Processing<br>with *1.008 to prediction</td>\n<td>0.7225193</td>\n<td>0.7298121</td>\n<td>-0.0105128</td>\n<td>-0.0122503</td>\n</tr>\n</tbody>\n</table>\n<h2>What Didn't Work</h2>\n<h3>Absorption Spectra of Known Gases</h3>\n<p>We tried to use TauREx3 to fit the dip spectrum.</p>\n<p>Although it worked really well on the training data, the leaderboard score was horrible, presumably because of the shift in the composition of the atmosphere.</p>\n<p>We did try to identify the gases in the test data, but without any success.</p>\n<h3>Denoising the Input Data with Machine Learning</h3>\n<p>It is quite hard to assume that taking the sum of the <code>[8:24]</code> channels is optimal.</p>\n<p>In some cases, the optimal range could be <code>[9:23]</code> or <code>[10:22]</code>, or even taking a weighted sum could be optimal.</p>\n<p>To resolve this problem, we tried several machine learning methods but didn't manage to beat <code>[8:24]</code>, probably due to the temporal spatial fluctuation of the signal.</p>\n<h2>What We Wanted to Do If We Had More Time</h2>\n<ul>\n<li>Utilize the FGS1 channel.</li>\n</ul>\n<h1>Final Comments</h1>\n<p>Deep analysis of the simulation code led to critical ideas needed for our victory. The function for the gain drift or the foreground processing method couldn't have been found without observing the ExoSim2 code. Though we had believed that this is the destiny of competitions that use simulation-generated data, we were very surprised that some of the top teams were able to achieve high scores without using them, so again, a big applause to them.</p>",
      "rawMarkdown": "[link to notebook](https://www.kaggle.com/code/cnumber/neurips-ariel-data-challenge-2024-final-submission)\n\n# Introduction\n\nWe ( @daiwakun and @cnumber) would like to thank the organizers for hosting this fascinating competition, especially @gordonyip and @lorenzomugnai, who worked tirelessly so that competitors could focus on the more \"interesting\" parts.\n\nAlso, a huge applause to @jeroencottaar for maintaining first place for almost the entire duration of the competition. Although we managed to grab victory at the end of the race, we are confident that we would have lost had the competition period been one day shorter or one day longer.\n\nMany ideas in our solution were found during a deep examination of the ExoSim2 and TauREx3 code used for data generation, including gain drift fitting and foreground processing. Although we didn't (or couldn't) find any leaks, we have to admit that our solution somewhat \"hacks\" the simulator. Nevertheless, we hope the hosts can make use of some aspects of our solution, and hopefully, it would be useful for the actual data.\n\nFor those who are curious about our significant score improvement during the last two days of the competition, we have decided to present what was going on in our team in [another post](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/544316)\n\n# Solution\n\n## Signal Preprocessing\n\nOnly the AIRS-CH0 channel was used because our solution heavily relies on the correlation between adjacent wavelengths, and also because we were unsure how to effectively utilize the FGS1 channel data due to issues with its distorted and fluctuating point spread function.\n\nThe public notebook (we wanted to put a link, but we couldn't find the notebook we used) was used with the following changes:\n\n1. **Disabling Hot Pixel Processing**: This led to a significant jump on the leaderboard. We assume that the hot pixel processing leads to unwanted information loss of the pixels in the middle of the sensor, making it almost impossible to correct the noise introduced by the time-dependent PSF distortion (see image below). Perhaps it is the sigma clip algorithm that is doing the wrong thing and that there exist algorithms that can handle hot pixels properly, but we weren't able to find them.\n\n2. **Foreground Processing**: Some teams have noticed that multiplying the final spectrum by a coefficient around 1.006~1.008 gives a huge boost. This is because a foreground is added to the signal during the ExoSim2 simulation [link to GitHub]. The correct way to handle this is to estimate the wavelength-dependent foreground signal and to subtract it from the signal in the central region. In our solution, we decided to estimate the foreground signal with the regions `[0:8]` and `[24:32]`, and subtracted it from the central region `[8:24]`. The effect of considering the foreground properly can be seen in [this discussion post](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543853#3034987).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2Fb07dae4e95a61af232cd678be36d516e%2Fbackground.png?generation=1730726806041614&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2F63105d03f4122d7967b25ca8a546853c%2Fpixel.png?generation=1730726844414945&alt=media)\n\n## Initial Dip Estimation for Each Wavelength\n\n### Identifying the Transit Interval\n\nFor identifying the time location of the transit, the wavelength-averaged signal was used. After estimating the approximate time position using a rule-based algorithm, the exact time was identified with a fitting algorithm.\n\n### Gain Drift\n\nGain drift refers to the variation in the detector's gain over time and across different wavelengths. These drifts can introduce systematic errors in the observed signal and must be corrected to accurately estimate the transit dip.\n\nExoSim2 models the gain drift as:\n\n$$\n(1 + f(t) \\cdot g(\\lambda))\n$$\n\nwhere \\\\( f(t) \\\\) and \\\\( g(\\lambda) \\\\) are polynomial functions.\n\nIt is notable that the final function form is different from a two-variable polynomial \\\\( h(t, \\lambda) \\\\), because while the wavelength and time components can be separated in the former, they cannot in the latter.\n\nBy using this functional form directly in the fitting described below, we were able to achieve high performance in the dip estimation.\n\n### Dip Estimation with Gain Drift Fitting\n\nThe function used for fitting is as follows:\n\n$$\ny_{\\text{pred}} = I(\\lambda) \\times \\text{Box}(\\lambda) \\times (1 + f(t) \\cdot g(\\lambda))\n$$\n\nwhere \\\\( I(\\lambda) \\\\) is the spectrum of the star, and \\\\( \\text{Box}(\\lambda) \\\\) describes the dip; it is 1 outside the transit and \\\\( 1 - d_\\lambda \\\\) inside the transit.\n\nThe number of fitting parameters are as follows:\n\n- \\\\( f(t) \\\\): 5 parameters\n- \\\\( g(\\lambda) \\\\): 5 parameters\n\nIt is worth mentioning that the optimal \\\\( I(\\lambda) \\\\) and the optimal dip used in \\\\( \\text{Box}(\\lambda) \\\\), which minimize the mean squared error, can be found analytically if the other parameters are fixed, and they don't need to be considered by the fitting algorithm.\n\nTo stabilize the fitting process, a two-stage fitting was used. In the first stage, \\\\( I(\\lambda) \\\\) was directly estimated from the raw signal with temporal averaging and fixed during the fitting. In the second stage, \\\\( I(\\lambda) \\\\) was not fixed and was set optimally for each step during the fitting, while for the fitting parameters, the result from the first stage was used as the initial parameter. The error for each data point used in the fitting was estimated from the variation of the signal at each wavelength.\n\n### Dip Error Estimation with Bootstrapping\n\nThe errors in the dip estimation at each wavelength differ due to the varying signal-to-noise ratios associated with each wavelength. [Bootstrapping](https://en.wikipedia.org/wiki/Bootstrapping_(statistics)) was used to overcome this problem and to estimate the dip error of each wavelength.\n\nFurther details are available in our code.\n\n## Dip Estimation Considering Wavelength Correlations\n\nThree models were used: Gaussian Process Regression, AutoEncoder, and Non-negative Matrix Factorization (NMF). They were ensembled with the ratio 6:2:2.\n\n### Gaussian Process Regression\n\nNothing fancy, unlike the second-place solution.\n\nA simple kernel composed of RBF and Matern kernels was employed, and the errors calculated by bootstrapping were passed to `sklearn.gaussian_process.GaussianProcessRegressor` so that the model can consider the uncertainty of each data point.\n\n### AutoEncoder\n\nWe applied an autoencoder to capture relationships in the data.\n\nUnlike PCA, which captures linear relationships, autoencoders can model more complex and nonlinear patterns.\n\nAlso, we expected that by training the model with MSE loss, the autoencoder model could take noise into account and recover the \"optimal\" spectrum.\n\nImportant points were to normalize the data for each exoplanet and to take the moving median of the dip spectrum to smooth the input dip spectrum.\n\nThe model we used is as below, with 4 nodes in the hidden layer:\n\n```python\ninput_data = Input(shape=(input_dim,))\nencoded = Dense(encoding_dim, activation='relu')(input_data)\ndecoded = Dense(input_dim, activation='linear')(encoded)\n\nautoencoder = Model(input_data, decoded)\n```\n\n### NMF\n\nVery similar to the autoencoder, but we added it to enhance the diversity.\n\nUnlike with the autoencoder, the number of ranks was set to 5 for NMF.\n\n### Spectral Components Identified by NMF\n\nThe plot below shows the identified spectral components by NMF for the training data.\n\nIt can be seen that the main contributing gases of each component are:\n\n- **Component 1**: CO₂\n- **Component 2**: CH₄\n- **Component 3**: H₂O\n\nThis signifies NMF's ability to consider the correlation of the spectrum unsupervised.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2Fbc7f93b100cae6b768c0c17a04ff10a8%2Fnmf_spectral_components.png?generation=1730726867605917&alt=media)\n\n### Sigma\n\nWe took the weighted average of the following components:\n\n1. **Constant value** (planet and wavelength independent)\n2. **Standard deviation of the smoothed predicted dip spectrum** (planet dependent, wavelength independent)\n3. **Uncertainty predicted by Gaussian Process Regression** (planet and wavelength dependent)\n\nThe constant value played an important role in cases where the dip spectrum was almost constant but had some bias.\n\n## The Contribution of Each Component\n\nSince our solution incorporates a variety of ideas, we examined the contributions for those that seem important with late submissions.\n\n| Method                                                           | Public      | Private     | Public Loss      | Private Loss     |\n|------------------------------------------------------------------|-------------|-------------|------------------|------------------|\n| **Final Submission**                                             | 0.7330321   | 0.7420624   | 0.0000000        | 0.0000000        |\n| Only Gaussian Process Regression                                 | 0.7221480   | 0.7343485   | -0.0108841       | -0.0077139       |\n| Only AutoEncoder                                                 | 0.7078056   | 0.7181137   | -0.0252265       | -0.0239487       |\n| Only NMF                                                         | 0.7017943   | 0.7122631   | -0.0312378       | -0.0297993       |\n| With Hot Pixel Processing                                        | 0.7021653   | 0.7224989   | -0.0308668       | -0.0195635       |\n| Without Foreground Processing<br>with *1.008 to prediction       | 0.7225193   | 0.7298121   | -0.0105128       | -0.0122503       |\n\n## What Didn't Work\n\n### Absorption Spectra of Known Gases\n\nWe tried to use TauREx3 to fit the dip spectrum.\n\nAlthough it worked really well on the training data, the leaderboard score was horrible, presumably because of the shift in the composition of the atmosphere.\n\nWe did try to identify the gases in the test data, but without any success.\n\n### Denoising the Input Data with Machine Learning\n\nIt is quite hard to assume that taking the sum of the `[8:24]` channels is optimal.\n\nIn some cases, the optimal range could be `[9:23]` or `[10:22]`, or even taking a weighted sum could be optimal.\n\nTo resolve this problem, we tried several machine learning methods but didn't manage to beat `[8:24]`, probably due to the temporal spatial fluctuation of the signal.\n\n## What We Wanted to Do If We Had More Time\n\n- Utilize the FGS1 channel.\n\n# Final Comments\n\nDeep analysis of the simulation code led to critical ideas needed for our victory. The function for the gain drift or the foreground processing method couldn't have been found without observing the ExoSim2 code. Though we had believed that this is the destiny of competitions that use simulation-generated data, we were very surprised that some of the top teams were able to achieve high scores without using them, so again, a big applause to them.\n",
      "votes": 51
    },
    {
      "id": 3036357,
      "postDate": "2024-11-04T13:44:31.620Z",
      "content": "<p>Congratulations on your win! It seems I didn't approach the competition enough as 'reproduce the synthetic data generation'. Nonetheless you also seem to have done a better job of breaking down the spectra than my simple PCA, which I expect could also benefit the final mission.</p>\n<p>I'll also have to try a submission with the hot pixel processing disabled. This would be a bit disappointing if it matters much, since there were very few hot pixels in the training set (and so I never gave it much consideration).</p>",
      "rawMarkdown": "Congratulations on your win! It seems I didn't approach the competition enough as 'reproduce the synthetic data generation'. Nonetheless you also seem to have done a better job of breaking down the spectra than my simple PCA, which I expect could also benefit the final mission.\n\nI'll also have to try a submission with the hot pixel processing disabled. This would be a bit disappointing if it matters much, since there were very few hot pixels in the training set (and so I never gave it much consideration).",
      "votes": 6,
      "replies": [
        {
          "id": 3036370,
          "postDate": "2024-11-04T13:54:14.670Z",
          "content": "<p>I'm eager to see if your score jump to over 0.76 without the hot pixels, this would be amazing</p>",
          "rawMarkdown": "I'm eager to see if your score jump to over 0.76 without the hot pixels, this would be amazing",
          "votes": 2
        },
        {
          "id": 3036941,
          "postDate": "2024-11-05T04:45:19.040Z",
          "content": "<p>I tested this and did not see an improvement by disabling the hot pixel filter. <a href=\"https://www.kaggle.com/cnumber\" target=\"_blank\">@cnumber</a> , <a href=\"https://www.kaggle.com/daiwakun\" target=\"_blank\">@daiwakun</a> , do you apply any method to deal with invalid data? If not that might explain the difference. It's rather critical to deal with invalids, because jitter is only compensated when all rows are summed (as Jun Koda explains above). In fact you might gain more from some inpainting since it would also deal with dead pixels.</p>",
          "rawMarkdown": "I tested this and did not see an improvement by disabling the hot pixel filter. @cnumber , @daiwakun , do you apply any method to deal with invalid data? If not that might explain the difference. It's rather critical to deal with invalids, because jitter is only compensated when all rows are summed (as Jun Koda explains above). In fact you might gain more from some inpainting since it would also deal with dead pixels.",
          "votes": 3
        },
        {
          "id": 3037058,
          "postDate": "2024-11-05T09:40:28.540Z",
          "content": "<p>For the dead pixels, and originally hot pixels, we didn't do anything special. Just a simple np.nanmean over the x-axis to get the signal for every wavelength. We also agree that the PSF distortion over time and the xy-axis jitter (as <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> pointed out) were the reasons why the hot pixel process hurt our score so much.</p>\n<p>We saw that your code interpolates nan pixels with adjacent pixels, and are assuming that maybe this is the reason why you didn't gain from deleting the hot pixel process.</p>",
          "rawMarkdown": "For the dead pixels, and originally hot pixels, we didn't do anything special. Just a simple np.nanmean over the x-axis to get the signal for every wavelength. We also agree that the PSF distortion over time and the xy-axis jitter (as @junkoda pointed out) were the reasons why the hot pixel process hurt our score so much.\n\nWe saw that your code interpolates nan pixels with adjacent pixels, and are assuming that maybe this is the reason why you didn't gain from deleting the hot pixel process.",
          "votes": 1,
          "replies": [
            {
              "id": 3037059,
              "postDate": "2024-11-05T09:45:02.130Z",
              "content": "<p>Why would interpolating dead affect the hot pixels? They are s separated iirc</p>",
              "rawMarkdown": "Why would interpolating dead affect the hot pixels? They are s separated iirc"
            },
            {
              "id": 3037158,
              "postDate": "2024-11-05T12:27:24.707Z",
              "content": "<p>I interpolate both dead and hot pixels.</p>",
              "rawMarkdown": "I interpolate both dead and hot pixels.",
              "votes": 1
            },
            {
              "id": 3037276,
              "postDate": "2024-11-05T14:55:01.933Z",
              "content": "<p>We were curious of the impact of the interpolation on the score, so we added <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>'s interpolation code to our submission, and to our surprise found out that the interpolation had a huge impact on our score.</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Without hot pixel process</td>\n<td>0.7399976</td>\n<td>0.7450662</td>\n</tr>\n<tr>\n<td>With hot pixel process</td>\n<td>0.73814190</td>\n<td>0.7448809</td>\n</tr>\n</tbody>\n</table>",
              "rawMarkdown": "We were curious of the impact of the interpolation on the score, so we added @jeroencottaar's interpolation code to our submission, and to our surprise found out that the interpolation had a huge impact on our score.\n\n|Submission| Public | Private |\n|--- | --- | --- |\n|Without hot pixel process |0.7399976  | 0.7450662 |\n|With hot pixel process | 0.73814190|0.7448809 |",
              "votes": 2
            },
            {
              "id": 3037288,
              "postDate": "2024-11-05T15:07:43.690Z",
              "content": "<p>Nice! So hot pixels interpolation is the way</p>",
              "rawMarkdown": "Nice! So hot pixels interpolation is the way"
            }
          ]
        }
      ]
    },
    {
      "id": 3036704,
      "postDate": "2024-11-04T20:39:59.737Z",
      "content": "<p>I felt the dark current correction is too small (and speculated if the time 0.45 or 0.1 factor is really correct), so it makes scenes that keeping hot pixel is harmless and disadvantage in noise cancelling can be worse (but I didn't try).</p>\n<p>The PSF (point spread function) noise is highly correlated among spatial channels and the error is much less for the sum over 32 channels; not just 1/sqrt(N) reduction but noise is cancelling each other. I guessed this is subpixel pointing error - the star moves a little randomly during timestep so each spacial channel get a little different light but total is the same; or maybe the light get distorted for other reasons among pixels but within 32 channels. And the error cancellation become imperfect when pixel is removed, so not removing can be a benefit.</p>\n<p>Do you only disable the hot pixels from dark current and keep the dead pixel mask? I think the Auto Encoder is great for learning the correlated error and compensate for the masked signal.</p>\n<p>Congratulations for the 1st place!!!</p>",
      "rawMarkdown": "I felt the dark current correction is too small (and speculated if the time 0.45 or 0.1 factor is really correct), so it makes scenes that keeping hot pixel is harmless and disadvantage in noise cancelling can be worse (but I didn't try).\n\nThe PSF (point spread function) noise is highly correlated among spatial channels and the error is much less for the sum over 32 channels; not just 1/sqrt(N) reduction but noise is cancelling each other. I guessed this is subpixel pointing error - the star moves a little randomly during timestep so each spacial channel get a little different light but total is the same; or maybe the light get distorted for other reasons among pixels but within 32 channels. And the error cancellation become imperfect when pixel is removed, so not removing can be a benefit.\n\nDo you only disable the hot pixels from dark current and keep the dead pixel mask? I think the Auto Encoder is great for learning the correlated error and compensate for the masked signal.\n\nCongratulations for the 1st place!!!",
      "votes": 1,
      "replies": [
        {
          "id": 3037052,
          "postDate": "2024-11-05T09:23:55.470Z",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>Do you only disable the hot pixels from dark current and keep the dead pixel mask? <br>\n  Yes. We believe that the process of the dead pixel mask is reasonable, and it was only the hot pixel mask that was doing the wrong things.</p>\n</blockquote>",
          "rawMarkdown": "Thanks!\n\n>Do you only disable the hot pixels from dark current and keep the dead pixel mask? \nYes. We believe that the process of the dead pixel mask is reasonable, and it was only the hot pixel mask that was doing the wrong things.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3036376,
      "postDate": "2024-11-04T14:03:06.403Z",
      "content": "<p>The link to the notebook the not work</p>",
      "rawMarkdown": "The link to the notebook the not work",
      "votes": 1,
      "replies": [
        {
          "id": 3036377,
          "postDate": "2024-11-04T14:04:22.730Z",
          "content": "<p>Sorry, forgot to make it public.<br>\nHope it works now!</p>",
          "rawMarkdown": "Sorry, forgot to make it public.\nHope it works now!",
          "votes": 1,
          "replies": [
            {
              "id": 3036380,
              "postDate": "2024-11-04T14:09:25.647Z",
              "content": "<p>Thank you. Congratz on winning; outstanding work.</p>",
              "rawMarkdown": "Thank you. Congratz on winning; outstanding work."
            },
            {
              "id": 3036382,
              "postDate": "2024-11-04T14:11:53.547Z",
              "content": "<p>I see 'private datasource' attached to your notebook, are they needed for it to run? If so, please public them too…</p>",
              "rawMarkdown": "I see 'private datasource' attached to your notebook, are they needed for it to run? If so, please public them too..."
            },
            {
              "id": 3036386,
              "postDate": "2024-11-04T14:17:58.943Z",
              "content": "<p>It's for installing astropy (and taurex_cuda), so you won't need it if you pip install it like this.<br>\n<a href=\"https://www.kaggle.com/discussions/product-feedback/532336\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/532336</a></p>\n<p>It was after attaching the datasource that I realised Kaggle had introduced this nice feature :(</p>",
              "rawMarkdown": "It's for installing astropy (and taurex_cuda), so you won't need it if you pip install it like this.\nhttps://www.kaggle.com/discussions/product-feedback/532336\n\nIt was after attaching the datasource that I realised Kaggle had introduced this nice feature :(",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3036369,
      "postDate": "2024-11-04T13:52:53.213Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3036357,
      "author_name": "Jeroen Cottaar",
      "author_url": "",
      "post_date": "2024-11-04T13:44:31.620000",
      "content": "<p>Congratulations on your win! It seems I didn't approach the competition enough as 'reproduce the synthetic data generation'. Nonetheless you also seem to have done a better job of breaking down the spectra than my simple PCA, which I expect could also benefit the final mission.</p>\n<p>I'll also have to try a submission with the hot pixel processing disabled. This would be a bit disappointing if it matters much, since there were very few hot pixels in the training set (and so I never gave it much consideration).</p>",
      "votes": 6,
      "replies": [
        {
          "id": 3036370,
          "author_name": "greySnow",
          "author_url": "",
          "post_date": "2024-11-04T13:54:14.670000",
          "content": "<p>I'm eager to see if your score jump to over 0.76 without the hot pixels, this would be amazing</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3036941,
          "author_name": "Jeroen Cottaar",
          "author_url": "",
          "post_date": "2024-11-05T04:45:19.040000",
          "content": "<p>I tested this and did not see an improvement by disabling the hot pixel filter. <a href=\"https://www.kaggle.com/cnumber\" target=\"_blank\">@cnumber</a> , <a href=\"https://www.kaggle.com/daiwakun\" target=\"_blank\">@daiwakun</a> , do you apply any method to deal with invalid data? If not that might explain the difference. It's rather critical to deal with invalids, because jitter is only compensated when all rows are summed (as Jun Koda explains above). In fact you might gain more from some inpainting since it would also deal with dead pixels.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 3037058,
          "author_name": "c-number",
          "author_url": "",
          "post_date": "2024-11-05T09:40:28.540000",
          "content": "<p>For the dead pixels, and originally hot pixels, we didn't do anything special. Just a simple np.nanmean over the x-axis to get the signal for every wavelength. We also agree that the PSF distortion over time and the xy-axis jitter (as <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a> pointed out) were the reasons why the hot pixel process hurt our score so much.</p>\n<p>We saw that your code interpolates nan pixels with adjacent pixels, and are assuming that maybe this is the reason why you didn't gain from deleting the hot pixel process.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3037059,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-11-05T09:45:02.130000",
              "content": "<p>Why would interpolating dead affect the hot pixels? They are s separated iirc</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3037158,
              "author_name": "Jeroen Cottaar",
              "author_url": "",
              "post_date": "2024-11-05T12:27:24.707000",
              "content": "<p>I interpolate both dead and hot pixels.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3037276,
              "author_name": "c-number",
              "author_url": "",
              "post_date": "2024-11-05T14:55:01.933000",
              "content": "<p>We were curious of the impact of the interpolation on the score, so we added <a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a>'s interpolation code to our submission, and to our surprise found out that the interpolation had a huge impact on our score.</p>\n<table>\n<thead>\n<tr>\n<th>Submission</th>\n<th>Public</th>\n<th>Private</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>Without hot pixel process</td>\n<td>0.7399976</td>\n<td>0.7450662</td>\n</tr>\n<tr>\n<td>With hot pixel process</td>\n<td>0.73814190</td>\n<td>0.7448809</td>\n</tr>\n</tbody>\n</table>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3037288,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-11-05T15:07:43.690000",
              "content": "<p>Nice! So hot pixels interpolation is the way</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3036704,
      "author_name": "🐢 Jun Koda",
      "author_url": "",
      "post_date": "2024-11-04T20:39:59.737000",
      "content": "<p>I felt the dark current correction is too small (and speculated if the time 0.45 or 0.1 factor is really correct), so it makes scenes that keeping hot pixel is harmless and disadvantage in noise cancelling can be worse (but I didn't try).</p>\n<p>The PSF (point spread function) noise is highly correlated among spatial channels and the error is much less for the sum over 32 channels; not just 1/sqrt(N) reduction but noise is cancelling each other. I guessed this is subpixel pointing error - the star moves a little randomly during timestep so each spacial channel get a little different light but total is the same; or maybe the light get distorted for other reasons among pixels but within 32 channels. And the error cancellation become imperfect when pixel is removed, so not removing can be a benefit.</p>\n<p>Do you only disable the hot pixels from dark current and keep the dead pixel mask? I think the Auto Encoder is great for learning the correlated error and compensate for the masked signal.</p>\n<p>Congratulations for the 1st place!!!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3037052,
          "author_name": "c-number",
          "author_url": "",
          "post_date": "2024-11-05T09:23:55.470000",
          "content": "<p>Thanks!</p>\n<blockquote>\n  <p>Do you only disable the hot pixels from dark current and keep the dead pixel mask? <br>\n  Yes. We believe that the process of the dead pixel mask is reasonable, and it was only the hot pixel mask that was doing the wrong things.</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3036376,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-11-04T14:03:06.403000",
      "content": "<p>The link to the notebook the not work</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3036377,
          "author_name": "c-number",
          "author_url": "",
          "post_date": "2024-11-04T14:04:22.730000",
          "content": "<p>Sorry, forgot to make it public.<br>\nHope it works now!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3036380,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-11-04T14:09:25.647000",
              "content": "<p>Thank you. Congratz on winning; outstanding work.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3036382,
              "author_name": "greySnow",
              "author_url": "",
              "post_date": "2024-11-04T14:11:53.547000",
              "content": "<p>I see 'private datasource' attached to your notebook, are they needed for it to run? If so, please public them too…</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3036386,
              "author_name": "c-number",
              "author_url": "",
              "post_date": "2024-11-04T14:17:58.943000",
              "content": "<p>It's for installing astropy (and taurex_cuda), so you won't need it if you pip install it like this.<br>\n<a href=\"https://www.kaggle.com/discussions/product-feedback/532336\" target=\"_blank\">https://www.kaggle.com/discussions/product-feedback/532336</a></p>\n<p>It was after attaching the datasource that I realised Kaggle had introduced this nice feature :(</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3036369,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-04T13:52:53.213000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3036336": "[link to notebook](https://www.kaggle.com/code/cnumber/neurips-ariel-data-challenge-2024-final-submission)\n\n# Introduction\n\nWe ( @daiwakun and @cnumber) would like to thank the organizers for hosting this fascinating competition, especially @gordonyip and @lorenzomugnai, who worked tirelessly so that competitors could focus on the more \"interesting\" parts.\n\nAlso, a huge applause to @jeroencottaar for maintaining first place for almost the entire duration of the competition. Although we managed to grab victory at the end of the race, we are confident that we would have lost had the competition period been one day shorter or one day longer.\n\nMany ideas in our solution were found during a deep examination of the ExoSim2 and TauREx3 code used for data generation, including gain drift fitting and foreground processing. Although we didn't (or couldn't) find any leaks, we have to admit that our solution somewhat \"hacks\" the simulator. Nevertheless, we hope the hosts can make use of some aspects of our solution, and hopefully, it would be useful for the actual data.\n\nFor those who are curious about our significant score improvement during the last two days of the competition, we have decided to present what was going on in our team in [another post](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/544316)\n\n# Solution\n\n## Signal Preprocessing\n\nOnly the AIRS-CH0 channel was used because our solution heavily relies on the correlation between adjacent wavelengths, and also because we were unsure how to effectively utilize the FGS1 channel data due to issues with its distorted and fluctuating point spread function.\n\nThe public notebook (we wanted to put a link, but we couldn't find the notebook we used) was used with the following changes:\n\n1. **Disabling Hot Pixel Processing**: This led to a significant jump on the leaderboard. We assume that the hot pixel processing leads to unwanted information loss of the pixels in the middle of the sensor, making it almost impossible to correct the noise introduced by the time-dependent PSF distortion (see image below). Perhaps it is the sigma clip algorithm that is doing the wrong thing and that there exist algorithms that can handle hot pixels properly, but we weren't able to find them.\n\n2. **Foreground Processing**: Some teams have noticed that multiplying the final spectrum by a coefficient around 1.006~1.008 gives a huge boost. This is because a foreground is added to the signal during the ExoSim2 simulation [link to GitHub]. The correct way to handle this is to estimate the wavelength-dependent foreground signal and to subtract it from the signal in the central region. In our solution, we decided to estimate the foreground signal with the regions `[0:8]` and `[24:32]`, and subtracted it from the central region `[8:24]`. The effect of considering the foreground properly can be seen in [this discussion post](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543853#3034987).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2Fb07dae4e95a61af232cd678be36d516e%2Fbackground.png?generation=1730726806041614&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2F63105d03f4122d7967b25ca8a546853c%2Fpixel.png?generation=1730726844414945&alt=media)\n\n## Initial Dip Estimation for Each Wavelength\n\n### Identifying the Transit Interval\n\nFor identifying the time location of the transit, the wavelength-averaged signal was used. After estimating the approximate time position using a rule-based algorithm, the exact time was identified with a fitting algorithm.\n\n### Gain Drift\n\nGain drift refers to the variation in the detector's gain over time and across different wavelengths. These drifts can introduce systematic errors in the observed signal and must be corrected to accurately estimate the transit dip.\n\nExoSim2 models the gain drift as:\n\n$$\n(1 + f(t) \\cdot g(\\lambda))\n$$\n\nwhere \\\\( f(t) \\\\) and \\\\( g(\\lambda) \\\\) are polynomial functions.\n\nIt is notable that the final function form is different from a two-variable polynomial \\\\( h(t, \\lambda) \\\\), because while the wavelength and time components can be separated in the former, they cannot in the latter.\n\nBy using this functional form directly in the fitting described below, we were able to achieve high performance in the dip estimation.\n\n### Dip Estimation with Gain Drift Fitting\n\nThe function used for fitting is as follows:\n\n$$\ny_{\\text{pred}} = I(\\lambda) \\times \\text{Box}(\\lambda) \\times (1 + f(t) \\cdot g(\\lambda))\n$$\n\nwhere \\\\( I(\\lambda) \\\\) is the spectrum of the star, and \\\\( \\text{Box}(\\lambda) \\\\) describes the dip; it is 1 outside the transit and \\\\( 1 - d_\\lambda \\\\) inside the transit.\n\nThe number of fitting parameters are as follows:\n\n- \\\\( f(t) \\\\): 5 parameters\n- \\\\( g(\\lambda) \\\\): 5 parameters\n\nIt is worth mentioning that the optimal \\\\( I(\\lambda) \\\\) and the optimal dip used in \\\\( \\text{Box}(\\lambda) \\\\), which minimize the mean squared error, can be found analytically if the other parameters are fixed, and they don't need to be considered by the fitting algorithm.\n\nTo stabilize the fitting process, a two-stage fitting was used. In the first stage, \\\\( I(\\lambda) \\\\) was directly estimated from the raw signal with temporal averaging and fixed during the fitting. In the second stage, \\\\( I(\\lambda) \\\\) was not fixed and was set optimally for each step during the fitting, while for the fitting parameters, the result from the first stage was used as the initial parameter. The error for each data point used in the fitting was estimated from the variation of the signal at each wavelength.\n\n### Dip Error Estimation with Bootstrapping\n\nThe errors in the dip estimation at each wavelength differ due to the varying signal-to-noise ratios associated with each wavelength. [Bootstrapping](https://en.wikipedia.org/wiki/Bootstrapping_(statistics)) was used to overcome this problem and to estimate the dip error of each wavelength.\n\nFurther details are available in our code.\n\n## Dip Estimation Considering Wavelength Correlations\n\nThree models were used: Gaussian Process Regression, AutoEncoder, and Non-negative Matrix Factorization (NMF). They were ensembled with the ratio 6:2:2.\n\n### Gaussian Process Regression\n\nNothing fancy, unlike the second-place solution.\n\nA simple kernel composed of RBF and Matern kernels was employed, and the errors calculated by bootstrapping were passed to `sklearn.gaussian_process.GaussianProcessRegressor` so that the model can consider the uncertainty of each data point.\n\n### AutoEncoder\n\nWe applied an autoencoder to capture relationships in the data.\n\nUnlike PCA, which captures linear relationships, autoencoders can model more complex and nonlinear patterns.\n\nAlso, we expected that by training the model with MSE loss, the autoencoder model could take noise into account and recover the \"optimal\" spectrum.\n\nImportant points were to normalize the data for each exoplanet and to take the moving median of the dip spectrum to smooth the input dip spectrum.\n\nThe model we used is as below, with 4 nodes in the hidden layer:\n\n```python\ninput_data = Input(shape=(input_dim,))\nencoded = Dense(encoding_dim, activation='relu')(input_data)\ndecoded = Dense(input_dim, activation='linear')(encoded)\n\nautoencoder = Model(input_data, decoded)\n```\n\n### NMF\n\nVery similar to the autoencoder, but we added it to enhance the diversity.\n\nUnlike with the autoencoder, the number of ranks was set to 5 for NMF.\n\n### Spectral Components Identified by NMF\n\nThe plot below shows the identified spectral components by NMF for the training data.\n\nIt can be seen that the main contributing gases of each component are:\n\n- **Component 1**: CO₂\n- **Component 2**: CH₄\n- **Component 3**: H₂O\n\nThis signifies NMF's ability to consider the correlation of the spectrum unsupervised.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8458050%2Fbc7f93b100cae6b768c0c17a04ff10a8%2Fnmf_spectral_components.png?generation=1730726867605917&alt=media)\n\n### Sigma\n\nWe took the weighted average of the following components:\n\n1. **Constant value** (planet and wavelength independent)\n2. **Standard deviation of the smoothed predicted dip spectrum** (planet dependent, wavelength independent)\n3. **Uncertainty predicted by Gaussian Process Regression** (planet and wavelength dependent)\n\nThe constant value played an important role in cases where the dip spectrum was almost constant but had some bias.\n\n## The Contribution of Each Component\n\nSince our solution incorporates a variety of ideas, we examined the contributions for those that seem important with late submissions.\n\n| Method                                                           | Public      | Private     | Public Loss      | Private Loss     |\n|------------------------------------------------------------------|-------------|-------------|------------------|------------------|\n| **Final Submission**                                             | 0.7330321   | 0.7420624   | 0.0000000        | 0.0000000        |\n| Only Gaussian Process Regression                                 | 0.7221480   | 0.7343485   | -0.0108841       | -0.0077139       |\n| Only AutoEncoder                                                 | 0.7078056   | 0.7181137   | -0.0252265       | -0.0239487       |\n| Only NMF                                                         | 0.7017943   | 0.7122631   | -0.0312378       | -0.0297993       |\n| With Hot Pixel Processing                                        | 0.7021653   | 0.7224989   | -0.0308668       | -0.0195635       |\n| Without Foreground Processing<br>with *1.008 to prediction       | 0.7225193   | 0.7298121   | -0.0105128       | -0.0122503       |\n\n## What Didn't Work\n\n### Absorption Spectra of Known Gases\n\nWe tried to use TauREx3 to fit the dip spectrum.\n\nAlthough it worked really well on the training data, the leaderboard score was horrible, presumably because of the shift in the composition of the atmosphere.\n\nWe did try to identify the gases in the test data, but without any success.\n\n### Denoising the Input Data with Machine Learning\n\nIt is quite hard to assume that taking the sum of the `[8:24]` channels is optimal.\n\nIn some cases, the optimal range could be `[9:23]` or `[10:22]`, or even taking a weighted sum could be optimal.\n\nTo resolve this problem, we tried several machine learning methods but didn't manage to beat `[8:24]`, probably due to the temporal spatial fluctuation of the signal.\n\n## What We Wanted to Do If We Had More Time\n\n- Utilize the FGS1 channel.\n\n# Final Comments\n\nDeep analysis of the simulation code led to critical ideas needed for our victory. The function for the gain drift or the foreground processing method couldn't have been found without observing the ExoSim2 code. Though we had believed that this is the destiny of competitions that use simulation-generated data, we were very surprised that some of the top teams were able to achieve high scores without using them, so again, a big applause to them.\n",
    "3036357": "Congratulations on your win! It seems I didn't approach the competition enough as 'reproduce the synthetic data generation'. Nonetheless you also seem to have done a better job of breaking down the spectra than my simple PCA, which I expect could also benefit the final mission.\n\nI'll also have to try a submission with the hot pixel processing disabled. This would be a bit disappointing if it matters much, since there were very few hot pixels in the training set (and so I never gave it much consideration).",
    "3036704": "I felt the dark current correction is too small (and speculated if the time 0.45 or 0.1 factor is really correct), so it makes scenes that keeping hot pixel is harmless and disadvantage in noise cancelling can be worse (but I didn't try).\n\nThe PSF (point spread function) noise is highly correlated among spatial channels and the error is much less for the sum over 32 channels; not just 1/sqrt(N) reduction but noise is cancelling each other. I guessed this is subpixel pointing error - the star moves a little randomly during timestep so each spacial channel get a little different light but total is the same; or maybe the light get distorted for other reasons among pixels but within 32 channels. And the error cancellation become imperfect when pixel is removed, so not removing can be a benefit.\n\nDo you only disable the hot pixels from dark current and keep the dead pixel mask? I think the Auto Encoder is great for learning the correlated error and compensate for the masked signal.\n\nCongratulations for the 1st place!!!",
    "3036376": "The link to the notebook the not work",
    "3036369": ""
  }
}