{
  "id": 543760,
  "title": "5th place solution - NeurIPS ARIEL 2024",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543760",
  "author_name": "Pascal Pfeiffer",
  "post_date": "2024-11-01T10:22:42.417000",
  "votes": 36,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, I would like to thank the organizers for this amazing competition and my teammate Youri. It was a pleasure to work with you on this competition for the last two weeks. Congratulations on your Grandmaster title!</p>\n<h2>Tldr, give me the Code</h2>\n<p>Data Prep Notebook: <a href=\"https://www.kaggle.com/code/ilu000/neurips-ariel24-data-prep-5th-place-solution\" target=\"_blank\">https://www.kaggle.com/code/ilu000/neurips-ariel24-data-prep-5th-place-solution</a><br>\nData Augmentation Notebook: <a href=\"https://www.kaggle.com/code/ilu000/ariel24-data-augmentation-5th-place-solution\" target=\"_blank\">https://www.kaggle.com/code/ilu000/ariel24-data-augmentation-5th-place-solution</a><br>\nFinal Train Dataset: <a href=\"https://www.kaggle.com/datasets/ilu000/neurips-ariel24-5th-place-solution-data\" target=\"_blank\">https://www.kaggle.com/datasets/ilu000/neurips-ariel24-5th-place-solution-data</a><br>\nFinal Models: <a href=\"https://www.kaggle.com/datasets/ilu000/ariel24-models\" target=\"_blank\">https://www.kaggle.com/datasets/ilu000/ariel24-models</a><br>\nTrain Notebook: <a href=\"https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-train\" target=\"_blank\">https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-train</a><br>\nInference Notebook: <a href=\"https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-inference\" target=\"_blank\">https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-inference</a></p>\n<h2>Data preprocessing</h2>\n<p>As with most competitions with synthetic data, when starting this competition I focused on data exploration and modeling the aggregated light curves. Data preprocessing in our final solution was mainly done as described in <a href=\"https://www.kaggle.com/code/ilu000/ariel24-data-prep-fixed-dt\" target=\"_blank\">https://www.kaggle.com/code/ilu000/ariel24-data-prep-fixed-dt</a> and as shared by the hosts. Minor improvements to speed up the process by using GPU and parallel processing were helped to get the runtime down to a about 1h for the full test dataset. As test has different hot and dead pixels, using interpolation to fill the gaps was helpful to generalize to the distribution shift.</p>\n<h2>Modeling the light curve</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F996f94eb1f56e04591f655ae21fb2e4c%2Fpolyfit.png?generation=1730455263584439&amp;alt=media\" alt=\"\"></p>\n<p>As many have figured out, artifical drift over time as a polynomial of up to an order of 4 or 5 has been added to the data. After proper data preprocessing, detecting and removing this drift was the key to get to a decent score. Detecting transit points, fitting this polynomial and using directly the mean predictions with a proper sigma (tuned by star on LB score) gives an public LB score of around 0.546 and private of 0.578. Excluding the regions around the transit points makes this a bit more stable. By the end of the competition, this type of solution was well known and was the highest scoring public solution which would have ended in the bronze region with some tuning. Up to this point, local validation could be done by splitting by star and had great correlation with the public LB score.</p>\n<h2>Wavelength features</h2>\n<p>All planets in the training dataset featured weak to strong wavelength dependent absorption caused by the molecules in the atmosphere. Actually looking at train labels, we can overlay them with absorption spectra of known molecules and see that train is composed of a mixture of H2O, CO2 and CH4.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F5d6cb1bd93b2387824f3b3443a77dbc1%2Fco2_ch4.png?generation=1730456674856560&amp;alt=media\" alt=\"\"></p>\n<p>When we teamed up, we had two different ways to deal with this. I was using a model-less approach to extract absorption from a sliding window over the wavelength. With some smoothing and binning, public LB score that can be achieved with this approach and a constant sigma is about 0.590 and still correlates well with validation as it doesn't overfit on specific destributions. 2D Gaussian + Markov chain Monte Carlo (MCMC) as shared in the forums and what apparently is the state of the art when modelling the light curves in literature helped to get slightly better results, but the runtime for MCMC prevented us for ever using it in any kernel submission.</p>\n<p>When adding any kind of model on top of this, local validation score skyrockets by 0.100 easily as the model can fit to the three molecules in the training set. But public LB didn't follow, indicating that the test set has different molecules present.</p>\n<p>Youri did use a model-based approach and cleverly constrained it by modelling 5 gaussians that fitted approximately the absorption peaks of H2O, CO2 and CH4. Then, splitted the signal with these gaussians and used the aggregates as features for a Linear Regression model. He also subtracted the mean absorption first to further constrain the model. This approach worked very well on local validation while still able to carry over most to LB. From these results, it was quite clear that models are very powerful to fit the training data, but it was hard for us to fully generalize to the test set.</p>\n<p>As we expected and close to the end of the competition, the hosts confirmed that there are new molecules in the test set that were not present in the training set. We experimented with different splits of the training data and even experimented with clusters by dominant molecule, but public LB stayed to be the best way to validate.</p>\n<h2>Sigma estimation</h2>\n<p>Estimating the sigma was a crucial part of the competition. We tried different ways to estimate the sigma, and some were harder to validate locally. What did work amazingly locally and on LB was the strong correlation from pred.std(-1) (standard deviation over wavelength for each planet) vs. the error. Scaling the mean sigmas by this factor improved local validation and LB score by ~0.030. Additionally, we used a small fully connected NN to predict the sigma and blended these predictions to gain a few more points on LB.</p>\n<h2>Synthetic data</h2>\n<p>Generating more data to extend the training set to more molecules pushed our LB score significantly in the last days. With Taurex3, we generated absortion spectra for 9 more molecules (all_molecules = [(H2O, CH4, CO2), CO, NH3, HCN, C2H2, SO2, C2H4, H2S, PH3, TiO]). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F78858788777cf5745ef9b4122e32ec88%2Fabsorption.png?generation=1730455495079511&amp;alt=media\" alt=\"\"></p>\n<p>We had to make assumptions regarding planet temperature which we took from literature, but still we couldn't get ideal matches between Taurex3 and what we see in the training data for known molecules such as H2O. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F06073112b2d1b69ea967d713453eaced%2FBildschirmfoto%20vom%202024-11-01%2011-15-29.png?generation=1730456163601916&amp;alt=media\" alt=\"\"></p>\n<p>We used the training data and augmented it with these absorption spectra in two ways. First, we added it on top of the training labels and modified the signal accordingly (using the transit zone from our pipeline). Second, we used clean/mean signals from the training data and added the absorption spectra to it, again modifying the signal accordingly. Then, we retrained our models on this augmented data and improved our LB score by ~0.020 while local validation score slightly decreased as it is harder to fit more molecules. Interestingly, only adding a few molecules improves both, local and LB score, but adding more molecules only improves LB score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Fda04e8b055d836ef5897c7fa1cde22be%2Fsynthetic_data.png?generation=1730455583429604&amp;alt=media\" alt=\"\"></p>\n<p>I am confident that with more LB probing for the right molecules, we could have improved our solution even further. But, we are happy with the 5th place and are looking forward to read about the solutions of the top teams.</p>\n<h2>What did not work</h2>\n<ul>\n<li>Linear combinations of absorption spectra. Still not sure why this didn't work, probably never got the settings quite right and the models just learned it better.</li>\n<li>Scale sigma by absolute absorption. We saw a slight correlation in the training data, but it didn't work on the test data.</li>\n<li>Use a network to predict the absorption. It might work now with more training data that better fits what we have in test, but in the end time was running out to test it again.</li>\n<li>many more</li>\n</ul>",
  "messages": [
    {
      "id": 3033577,
      "postDate": "2024-11-01T10:22:42.417Z",
      "content": "<p>First of all, I would like to thank the organizers for this amazing competition and my teammate Youri. It was a pleasure to work with you on this competition for the last two weeks. Congratulations on your Grandmaster title!</p>\n<h2>Tldr, give me the Code</h2>\n<p>Data Prep Notebook: <a href=\"https://www.kaggle.com/code/ilu000/neurips-ariel24-data-prep-5th-place-solution\" target=\"_blank\">https://www.kaggle.com/code/ilu000/neurips-ariel24-data-prep-5th-place-solution</a><br>\nData Augmentation Notebook: <a href=\"https://www.kaggle.com/code/ilu000/ariel24-data-augmentation-5th-place-solution\" target=\"_blank\">https://www.kaggle.com/code/ilu000/ariel24-data-augmentation-5th-place-solution</a><br>\nFinal Train Dataset: <a href=\"https://www.kaggle.com/datasets/ilu000/neurips-ariel24-5th-place-solution-data\" target=\"_blank\">https://www.kaggle.com/datasets/ilu000/neurips-ariel24-5th-place-solution-data</a><br>\nFinal Models: <a href=\"https://www.kaggle.com/datasets/ilu000/ariel24-models\" target=\"_blank\">https://www.kaggle.com/datasets/ilu000/ariel24-models</a><br>\nTrain Notebook: <a href=\"https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-train\" target=\"_blank\">https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-train</a><br>\nInference Notebook: <a href=\"https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-inference\" target=\"_blank\">https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-inference</a></p>\n<h2>Data preprocessing</h2>\n<p>As with most competitions with synthetic data, when starting this competition I focused on data exploration and modeling the aggregated light curves. Data preprocessing in our final solution was mainly done as described in <a href=\"https://www.kaggle.com/code/ilu000/ariel24-data-prep-fixed-dt\" target=\"_blank\">https://www.kaggle.com/code/ilu000/ariel24-data-prep-fixed-dt</a> and as shared by the hosts. Minor improvements to speed up the process by using GPU and parallel processing were helped to get the runtime down to a about 1h for the full test dataset. As test has different hot and dead pixels, using interpolation to fill the gaps was helpful to generalize to the distribution shift.</p>\n<h2>Modeling the light curve</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F996f94eb1f56e04591f655ae21fb2e4c%2Fpolyfit.png?generation=1730455263584439&amp;alt=media\" alt=\"\"></p>\n<p>As many have figured out, artifical drift over time as a polynomial of up to an order of 4 or 5 has been added to the data. After proper data preprocessing, detecting and removing this drift was the key to get to a decent score. Detecting transit points, fitting this polynomial and using directly the mean predictions with a proper sigma (tuned by star on LB score) gives an public LB score of around 0.546 and private of 0.578. Excluding the regions around the transit points makes this a bit more stable. By the end of the competition, this type of solution was well known and was the highest scoring public solution which would have ended in the bronze region with some tuning. Up to this point, local validation could be done by splitting by star and had great correlation with the public LB score.</p>\n<h2>Wavelength features</h2>\n<p>All planets in the training dataset featured weak to strong wavelength dependent absorption caused by the molecules in the atmosphere. Actually looking at train labels, we can overlay them with absorption spectra of known molecules and see that train is composed of a mixture of H2O, CO2 and CH4.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F5d6cb1bd93b2387824f3b3443a77dbc1%2Fco2_ch4.png?generation=1730456674856560&amp;alt=media\" alt=\"\"></p>\n<p>When we teamed up, we had two different ways to deal with this. I was using a model-less approach to extract absorption from a sliding window over the wavelength. With some smoothing and binning, public LB score that can be achieved with this approach and a constant sigma is about 0.590 and still correlates well with validation as it doesn't overfit on specific destributions. 2D Gaussian + Markov chain Monte Carlo (MCMC) as shared in the forums and what apparently is the state of the art when modelling the light curves in literature helped to get slightly better results, but the runtime for MCMC prevented us for ever using it in any kernel submission.</p>\n<p>When adding any kind of model on top of this, local validation score skyrockets by 0.100 easily as the model can fit to the three molecules in the training set. But public LB didn't follow, indicating that the test set has different molecules present.</p>\n<p>Youri did use a model-based approach and cleverly constrained it by modelling 5 gaussians that fitted approximately the absorption peaks of H2O, CO2 and CH4. Then, splitted the signal with these gaussians and used the aggregates as features for a Linear Regression model. He also subtracted the mean absorption first to further constrain the model. This approach worked very well on local validation while still able to carry over most to LB. From these results, it was quite clear that models are very powerful to fit the training data, but it was hard for us to fully generalize to the test set.</p>\n<p>As we expected and close to the end of the competition, the hosts confirmed that there are new molecules in the test set that were not present in the training set. We experimented with different splits of the training data and even experimented with clusters by dominant molecule, but public LB stayed to be the best way to validate.</p>\n<h2>Sigma estimation</h2>\n<p>Estimating the sigma was a crucial part of the competition. We tried different ways to estimate the sigma, and some were harder to validate locally. What did work amazingly locally and on LB was the strong correlation from pred.std(-1) (standard deviation over wavelength for each planet) vs. the error. Scaling the mean sigmas by this factor improved local validation and LB score by ~0.030. Additionally, we used a small fully connected NN to predict the sigma and blended these predictions to gain a few more points on LB.</p>\n<h2>Synthetic data</h2>\n<p>Generating more data to extend the training set to more molecules pushed our LB score significantly in the last days. With Taurex3, we generated absortion spectra for 9 more molecules (all_molecules = [(H2O, CH4, CO2), CO, NH3, HCN, C2H2, SO2, C2H4, H2S, PH3, TiO]). </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F78858788777cf5745ef9b4122e32ec88%2Fabsorption.png?generation=1730455495079511&amp;alt=media\" alt=\"\"></p>\n<p>We had to make assumptions regarding planet temperature which we took from literature, but still we couldn't get ideal matches between Taurex3 and what we see in the training data for known molecules such as H2O. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F06073112b2d1b69ea967d713453eaced%2FBildschirmfoto%20vom%202024-11-01%2011-15-29.png?generation=1730456163601916&amp;alt=media\" alt=\"\"></p>\n<p>We used the training data and augmented it with these absorption spectra in two ways. First, we added it on top of the training labels and modified the signal accordingly (using the transit zone from our pipeline). Second, we used clean/mean signals from the training data and added the absorption spectra to it, again modifying the signal accordingly. Then, we retrained our models on this augmented data and improved our LB score by ~0.020 while local validation score slightly decreased as it is harder to fit more molecules. Interestingly, only adding a few molecules improves both, local and LB score, but adding more molecules only improves LB score.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Fda04e8b055d836ef5897c7fa1cde22be%2Fsynthetic_data.png?generation=1730455583429604&amp;alt=media\" alt=\"\"></p>\n<p>I am confident that with more LB probing for the right molecules, we could have improved our solution even further. But, we are happy with the 5th place and are looking forward to read about the solutions of the top teams.</p>\n<h2>What did not work</h2>\n<ul>\n<li>Linear combinations of absorption spectra. Still not sure why this didn't work, probably never got the settings quite right and the models just learned it better.</li>\n<li>Scale sigma by absolute absorption. We saw a slight correlation in the training data, but it didn't work on the test data.</li>\n<li>Use a network to predict the absorption. It might work now with more training data that better fits what we have in test, but in the end time was running out to test it again.</li>\n<li>many more</li>\n</ul>",
      "rawMarkdown": "First of all, I would like to thank the organizers for this amazing competition and my teammate Youri. It was a pleasure to work with you on this competition for the last two weeks. Congratulations on your Grandmaster title!\n\n## Tldr, give me the Code\nData Prep Notebook: https://www.kaggle.com/code/ilu000/neurips-ariel24-data-prep-5th-place-solution\nData Augmentation Notebook: https://www.kaggle.com/code/ilu000/ariel24-data-augmentation-5th-place-solution\nFinal Train Dataset: https://www.kaggle.com/datasets/ilu000/neurips-ariel24-5th-place-solution-data\nFinal Models: https://www.kaggle.com/datasets/ilu000/ariel24-models\nTrain Notebook: https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-train\nInference Notebook: https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-inference\n\n## Data preprocessing\n\nAs with most competitions with synthetic data, when starting this competition I focused on data exploration and modeling the aggregated light curves. Data preprocessing in our final solution was mainly done as described in https://www.kaggle.com/code/ilu000/ariel24-data-prep-fixed-dt and as shared by the hosts. Minor improvements to speed up the process by using GPU and parallel processing were helped to get the runtime down to a about 1h for the full test dataset. As test has different hot and dead pixels, using interpolation to fill the gaps was helpful to generalize to the distribution shift.\n\n## Modeling the light curve\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F996f94eb1f56e04591f655ae21fb2e4c%2Fpolyfit.png?generation=1730455263584439&alt=media)\n\nAs many have figured out, artifical drift over time as a polynomial of up to an order of 4 or 5 has been added to the data. After proper data preprocessing, detecting and removing this drift was the key to get to a decent score. Detecting transit points, fitting this polynomial and using directly the mean predictions with a proper sigma (tuned by star on LB score) gives an public LB score of around 0.546 and private of 0.578. Excluding the regions around the transit points makes this a bit more stable. By the end of the competition, this type of solution was well known and was the highest scoring public solution which would have ended in the bronze region with some tuning. Up to this point, local validation could be done by splitting by star and had great correlation with the public LB score.\n\n## Wavelength features\n\nAll planets in the training dataset featured weak to strong wavelength dependent absorption caused by the molecules in the atmosphere. Actually looking at train labels, we can overlay them with absorption spectra of known molecules and see that train is composed of a mixture of H2O, CO2 and CH4.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F5d6cb1bd93b2387824f3b3443a77dbc1%2Fco2_ch4.png?generation=1730456674856560&alt=media)\n\nWhen we teamed up, we had two different ways to deal with this. I was using a model-less approach to extract absorption from a sliding window over the wavelength. With some smoothing and binning, public LB score that can be achieved with this approach and a constant sigma is about 0.590 and still correlates well with validation as it doesn't overfit on specific destributions. 2D Gaussian + Markov chain Monte Carlo (MCMC) as shared in the forums and what apparently is the state of the art when modelling the light curves in literature helped to get slightly better results, but the runtime for MCMC prevented us for ever using it in any kernel submission.\n\nWhen adding any kind of model on top of this, local validation score skyrockets by 0.100 easily as the model can fit to the three molecules in the training set. But public LB didn't follow, indicating that the test set has different molecules present.\n\nYouri did use a model-based approach and cleverly constrained it by modelling 5 gaussians that fitted approximately the absorption peaks of H2O, CO2 and CH4. Then, splitted the signal with these gaussians and used the aggregates as features for a Linear Regression model. He also subtracted the mean absorption first to further constrain the model. This approach worked very well on local validation while still able to carry over most to LB. From these results, it was quite clear that models are very powerful to fit the training data, but it was hard for us to fully generalize to the test set.\n\nAs we expected and close to the end of the competition, the hosts confirmed that there are new molecules in the test set that were not present in the training set. We experimented with different splits of the training data and even experimented with clusters by dominant molecule, but public LB stayed to be the best way to validate.\n\n## Sigma estimation\n\nEstimating the sigma was a crucial part of the competition. We tried different ways to estimate the sigma, and some were harder to validate locally. What did work amazingly locally and on LB was the strong correlation from pred.std(-1) (standard deviation over wavelength for each planet) vs. the error. Scaling the mean sigmas by this factor improved local validation and LB score by ~0.030. Additionally, we used a small fully connected NN to predict the sigma and blended these predictions to gain a few more points on LB.\n\n## Synthetic data\n\nGenerating more data to extend the training set to more molecules pushed our LB score significantly in the last days. With Taurex3, we generated absortion spectra for 9 more molecules (all_molecules = [(H2O, CH4, CO2), CO, NH3, HCN, C2H2, SO2, C2H4, H2S, PH3, TiO]). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F78858788777cf5745ef9b4122e32ec88%2Fabsorption.png?generation=1730455495079511&alt=media)\n\nWe had to make assumptions regarding planet temperature which we took from literature, but still we couldn't get ideal matches between Taurex3 and what we see in the training data for known molecules such as H2O. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F06073112b2d1b69ea967d713453eaced%2FBildschirmfoto%20vom%202024-11-01%2011-15-29.png?generation=1730456163601916&alt=media)\n\nWe used the training data and augmented it with these absorption spectra in two ways. First, we added it on top of the training labels and modified the signal accordingly (using the transit zone from our pipeline). Second, we used clean/mean signals from the training data and added the absorption spectra to it, again modifying the signal accordingly. Then, we retrained our models on this augmented data and improved our LB score by ~0.020 while local validation score slightly decreased as it is harder to fit more molecules. Interestingly, only adding a few molecules improves both, local and LB score, but adding more molecules only improves LB score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Fda04e8b055d836ef5897c7fa1cde22be%2Fsynthetic_data.png?generation=1730455583429604&alt=media)\n\nI am confident that with more LB probing for the right molecules, we could have improved our solution even further. But, we are happy with the 5th place and are looking forward to read about the solutions of the top teams.\n\n## What did not work\n\n- Linear combinations of absorption spectra. Still not sure why this didn't work, probably never got the settings quite right and the models just learned it better.\n- Scale sigma by absolute absorption. We saw a slight correlation in the training data, but it didn't work on the test data.\n- Use a network to predict the absorption. It might work now with more training data that better fits what we have in test, but in the end time was running out to test it again.\n- many more\n",
      "votes": 35
    },
    {
      "id": 3033611,
      "postDate": "2024-11-01T10:58:51.793Z",
      "content": "<p>Very interesting approach to relay on the known absorption spectra for modelling/prediction. Congrats on 5th place and Youri's GM.</p>",
      "rawMarkdown": "Very interesting approach to relay on the known absorption spectra for modelling/prediction. Congrats on 5th place and Youri's GM.",
      "votes": 2,
      "replies": [
        {
          "id": 3033803,
          "postDate": "2024-11-01T14:25:25.203Z",
          "content": "<p>Thank you, looking forward to reading about your solution!</p>",
          "rawMarkdown": "Thank you, looking forward to reading about your solution!"
        },
        {
          "id": 3034269,
          "postDate": "2024-11-02T00:19:36.453Z",
          "content": "<p>Congratulations on your impressive 5th place finish, Pascal! Your approach to leveraging known absorption spectra for modeling is really insightful. It's great to see the collaboration with Youri and how you both tackled the challenges of the competition.<br>\nI’m looking forward to seeing your final code once it's cleaned up. It would be really helpful for those of us looking to learn from your methodologies. Thanks for sharing your detailed insights—it's always inspiring to see how different strategies can lead to success!</p>",
          "rawMarkdown": "Congratulations on your impressive 5th place finish, Pascal! Your approach to leveraging known absorption spectra for modeling is really insightful. It's great to see the collaboration with Youri and how you both tackled the challenges of the competition.\nI’m looking forward to seeing your final code once it's cleaned up. It would be really helpful for those of us looking to learn from your methodologies. Thanks for sharing your detailed insights—it's always inspiring to see how different strategies can lead to success!"
        }
      ]
    },
    {
      "id": 3035318,
      "postDate": "2024-11-03T09:17:27.910Z",
      "content": "<p>Thank you for your excellent approach using absorption spectra of known molecules. <br>\nYou mentioned that the molecules included in the training set and the test set are different. Were there any differences between star0 and star1 in the training set?</p>",
      "rawMarkdown": "Thank you for your excellent approach using absorption spectra of known molecules. \nYou mentioned that the molecules included in the training set and the test set are different. Were there any differences between star0 and star1 in the training set?"
    },
    {
      "id": 3035175,
      "postDate": "2024-11-03T04:57:07.480Z",
      "content": "<p>Congratulations Pascal and Youri. </p>\n<p>Seeing the molecules adsorption with wavelength gives true sense of what this competition is about. Thanks for sharing. </p>\n<p>Regarding 2D Gaussian + Markov chain Monte Carlo (MCMC), other than the runtime issue, how was the prediction accuracy? Did you try any approximate methods such as variational inference (VI)?  Thanks</p>",
      "rawMarkdown": "Congratulations Pascal and Youri. \n\nSeeing the molecules adsorption with wavelength gives true sense of what this competition is about. Thanks for sharing. \n\nRegarding 2D Gaussian + Markov chain Monte Carlo (MCMC), other than the runtime issue, how was the prediction accuracy? Did you try any approximate methods such as variational inference (VI)?  Thanks"
    },
    {
      "id": 3033836,
      "postDate": "2024-11-01T15:01:28.257Z",
      "content": "<p>Hi. Can you share the final code as well ?</p>",
      "rawMarkdown": "Hi. Can you share the final code as well ?",
      "replies": [
        {
          "id": 3033849,
          "postDate": "2024-11-01T15:10:28.020Z",
          "content": "<p>Yes, we will share the full solution and training code as required by the winner license. Just needs a bit of clean-up first.</p>",
          "rawMarkdown": "Yes, we will share the full solution and training code as required by the winner license. Just needs a bit of clean-up first.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3033611,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-11-01T10:58:51.793000",
      "content": "<p>Very interesting approach to relay on the known absorption spectra for modelling/prediction. Congrats on 5th place and Youri's GM.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3033803,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2024-11-01T14:25:25.203000",
          "content": "<p>Thank you, looking forward to reading about your solution!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3034269,
          "author_name": "yiyao6442",
          "author_url": "",
          "post_date": "2024-11-02T00:19:36.453000",
          "content": "<p>Congratulations on your impressive 5th place finish, Pascal! Your approach to leveraging known absorption spectra for modeling is really insightful. It's great to see the collaboration with Youri and how you both tackled the challenges of the competition.<br>\nI’m looking forward to seeing your final code once it's cleaned up. It would be really helpful for those of us looking to learn from your methodologies. Thanks for sharing your detailed insights—it's always inspiring to see how different strategies can lead to success!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3035318,
      "author_name": "Ty-Yuki",
      "author_url": "",
      "post_date": "2024-11-03T09:17:27.910000",
      "content": "<p>Thank you for your excellent approach using absorption spectra of known molecules. <br>\nYou mentioned that the molecules included in the training set and the test set are different. Were there any differences between star0 and star1 in the training set?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3035175,
      "author_name": "Viji",
      "author_url": "",
      "post_date": "2024-11-03T04:57:07.480000",
      "content": "<p>Congratulations Pascal and Youri. </p>\n<p>Seeing the molecules adsorption with wavelength gives true sense of what this competition is about. Thanks for sharing. </p>\n<p>Regarding 2D Gaussian + Markov chain Monte Carlo (MCMC), other than the runtime issue, how was the prediction accuracy? Did you try any approximate methods such as variational inference (VI)?  Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033836,
      "author_name": "JamshaidSohail",
      "author_url": "",
      "post_date": "2024-11-01T15:01:28.257000",
      "content": "<p>Hi. Can you share the final code as well ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3033849,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2024-11-01T15:10:28.020000",
          "content": "<p>Yes, we will share the full solution and training code as required by the winner license. Just needs a bit of clean-up first.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3033577": "First of all, I would like to thank the organizers for this amazing competition and my teammate Youri. It was a pleasure to work with you on this competition for the last two weeks. Congratulations on your Grandmaster title!\n\n## Tldr, give me the Code\nData Prep Notebook: https://www.kaggle.com/code/ilu000/neurips-ariel24-data-prep-5th-place-solution\nData Augmentation Notebook: https://www.kaggle.com/code/ilu000/ariel24-data-augmentation-5th-place-solution\nFinal Train Dataset: https://www.kaggle.com/datasets/ilu000/neurips-ariel24-5th-place-solution-data\nFinal Models: https://www.kaggle.com/datasets/ilu000/ariel24-models\nTrain Notebook: https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-train\nInference Notebook: https://www.kaggle.com/code/ilu000/neurips-ariel24-5th-place-solution-inference\n\n## Data preprocessing\n\nAs with most competitions with synthetic data, when starting this competition I focused on data exploration and modeling the aggregated light curves. Data preprocessing in our final solution was mainly done as described in https://www.kaggle.com/code/ilu000/ariel24-data-prep-fixed-dt and as shared by the hosts. Minor improvements to speed up the process by using GPU and parallel processing were helped to get the runtime down to a about 1h for the full test dataset. As test has different hot and dead pixels, using interpolation to fill the gaps was helpful to generalize to the distribution shift.\n\n## Modeling the light curve\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F996f94eb1f56e04591f655ae21fb2e4c%2Fpolyfit.png?generation=1730455263584439&alt=media)\n\nAs many have figured out, artifical drift over time as a polynomial of up to an order of 4 or 5 has been added to the data. After proper data preprocessing, detecting and removing this drift was the key to get to a decent score. Detecting transit points, fitting this polynomial and using directly the mean predictions with a proper sigma (tuned by star on LB score) gives an public LB score of around 0.546 and private of 0.578. Excluding the regions around the transit points makes this a bit more stable. By the end of the competition, this type of solution was well known and was the highest scoring public solution which would have ended in the bronze region with some tuning. Up to this point, local validation could be done by splitting by star and had great correlation with the public LB score.\n\n## Wavelength features\n\nAll planets in the training dataset featured weak to strong wavelength dependent absorption caused by the molecules in the atmosphere. Actually looking at train labels, we can overlay them with absorption spectra of known molecules and see that train is composed of a mixture of H2O, CO2 and CH4.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F5d6cb1bd93b2387824f3b3443a77dbc1%2Fco2_ch4.png?generation=1730456674856560&alt=media)\n\nWhen we teamed up, we had two different ways to deal with this. I was using a model-less approach to extract absorption from a sliding window over the wavelength. With some smoothing and binning, public LB score that can be achieved with this approach and a constant sigma is about 0.590 and still correlates well with validation as it doesn't overfit on specific destributions. 2D Gaussian + Markov chain Monte Carlo (MCMC) as shared in the forums and what apparently is the state of the art when modelling the light curves in literature helped to get slightly better results, but the runtime for MCMC prevented us for ever using it in any kernel submission.\n\nWhen adding any kind of model on top of this, local validation score skyrockets by 0.100 easily as the model can fit to the three molecules in the training set. But public LB didn't follow, indicating that the test set has different molecules present.\n\nYouri did use a model-based approach and cleverly constrained it by modelling 5 gaussians that fitted approximately the absorption peaks of H2O, CO2 and CH4. Then, splitted the signal with these gaussians and used the aggregates as features for a Linear Regression model. He also subtracted the mean absorption first to further constrain the model. This approach worked very well on local validation while still able to carry over most to LB. From these results, it was quite clear that models are very powerful to fit the training data, but it was hard for us to fully generalize to the test set.\n\nAs we expected and close to the end of the competition, the hosts confirmed that there are new molecules in the test set that were not present in the training set. We experimented with different splits of the training data and even experimented with clusters by dominant molecule, but public LB stayed to be the best way to validate.\n\n## Sigma estimation\n\nEstimating the sigma was a crucial part of the competition. We tried different ways to estimate the sigma, and some were harder to validate locally. What did work amazingly locally and on LB was the strong correlation from pred.std(-1) (standard deviation over wavelength for each planet) vs. the error. Scaling the mean sigmas by this factor improved local validation and LB score by ~0.030. Additionally, we used a small fully connected NN to predict the sigma and blended these predictions to gain a few more points on LB.\n\n## Synthetic data\n\nGenerating more data to extend the training set to more molecules pushed our LB score significantly in the last days. With Taurex3, we generated absortion spectra for 9 more molecules (all_molecules = [(H2O, CH4, CO2), CO, NH3, HCN, C2H2, SO2, C2H4, H2S, PH3, TiO]). \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F78858788777cf5745ef9b4122e32ec88%2Fabsorption.png?generation=1730455495079511&alt=media)\n\nWe had to make assumptions regarding planet temperature which we took from literature, but still we couldn't get ideal matches between Taurex3 and what we see in the training data for known molecules such as H2O. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F06073112b2d1b69ea967d713453eaced%2FBildschirmfoto%20vom%202024-11-01%2011-15-29.png?generation=1730456163601916&alt=media)\n\nWe used the training data and augmented it with these absorption spectra in two ways. First, we added it on top of the training labels and modified the signal accordingly (using the transit zone from our pipeline). Second, we used clean/mean signals from the training data and added the absorption spectra to it, again modifying the signal accordingly. Then, we retrained our models on this augmented data and improved our LB score by ~0.020 while local validation score slightly decreased as it is harder to fit more molecules. Interestingly, only adding a few molecules improves both, local and LB score, but adding more molecules only improves LB score.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Fda04e8b055d836ef5897c7fa1cde22be%2Fsynthetic_data.png?generation=1730455583429604&alt=media)\n\nI am confident that with more LB probing for the right molecules, we could have improved our solution even further. But, we are happy with the 5th place and are looking forward to read about the solutions of the top teams.\n\n## What did not work\n\n- Linear combinations of absorption spectra. Still not sure why this didn't work, probably never got the settings quite right and the models just learned it better.\n- Scale sigma by absolute absorption. We saw a slight correlation in the training data, but it didn't work on the test data.\n- Use a network to predict the absorption. It might work now with more training data that better fits what we have in test, but in the end time was running out to test it again.\n- many more\n",
    "3033611": "Very interesting approach to relay on the known absorption spectra for modelling/prediction. Congrats on 5th place and Youri's GM.",
    "3035318": "Thank you for your excellent approach using absorption spectra of known molecules. \nYou mentioned that the molecules included in the training set and the test set are different. Were there any differences between star0 and star1 in the training set?",
    "3035175": "Congratulations Pascal and Youri. \n\nSeeing the molecules adsorption with wavelength gives true sense of what this competition is about. Thanks for sharing. \n\nRegarding 2D Gaussian + Markov chain Monte Carlo (MCMC), other than the runtime issue, how was the prediction accuracy? Did you try any approximate methods such as variational inference (VI)?  Thanks",
    "3033836": "Hi. Can you share the final code as well ?"
  }
}