{
  "id": 543666,
  "title": "6th place solution",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543666",
  "author_name": "Sergei Fironov",
  "post_date": "2024-10-31T23:59:31.002000",
  "votes": 57,
  "comment_count": 23,
  "views": 0,
  "content": "<p>It was an excellent competition! For many years, I've been following astrophysics news with great interest. What else will they see in these blurry pixels? When will they finally find a habitable planet? And now I've been lucky enough to contribute to this myself.</p>\n<p>I want to thank the competition hosts, the Ariel mission, and specifically <a href=\"https://www.kaggle.com/GordonYip\" target=\"_blank\">@GordonYip</a> for answering important questions. I'd also like to thank all the authors of the notebook with the correct data calibration process. Without this, my participation would hardly have been meaningful. It's so difficult to find information about calibrations in articles! Thank <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> for the meme thread and funny comments 😀</p>\n<p>My solution isn't particularly sophisticated. Here's the breakdown:</p>\n<p><strong>Data Processing</strong>:<br>\nI noticed outliers in pixel data and smoothed them with a Gaussian in the frequency axis direction. The \"extra\" wavelengths not included in the final spectrum were somewhat useful here. Consistently improved by 0.002.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F71783%2Faf1f3b347ae2379d01f6be441ade29d5%2Fbang.jpg?generation=1730420235408838&amp;alt=media\" alt=\"\"><br>\n<strong>Smoothing</strong>:<br>\nNothing in the world is better than SG. It's fast enough and smooths everything. I even smoothed the predictions.</p>\n<p><strong>Transit Zone Prediction</strong>:<br>\nNothing is more painful than getting 0 in submission results. The main problem is incorrectly predicted transit boundaries. In the test, they're wider or narrower than those in the training set. Using SG smoothing in the time direction to find boundaries. Ensuring light flux decreases at the left boundary and increases at the right.</p>\n<p><strong>Features for Models</strong>:<br>\nThe feature construction process is similar to my baseline. We build a polynomial, find how to 'raise' the transit zone to align with the non-transit part. Used MAE and logpdf functions to evaluate how well we found the coefficient.<br>\nBuild a general polynomial for the frequency sum. For each frequency, try to find coefficients pulling the transit as B and external parts as A to this polynomial. 1 - A/B is the target feature for individual frequencies. Sums across frequency intervals. I know from theory that some zones correspond to absorption areas of basic gases CH4, H2O, CO2, CO, NH3. But the best frequency areas for building features turned out to be different. I used a mix of genetic algorithm and manual selection to find the best intervals.</p>\n<p><strong>Sigma Estimation</strong>:<br>\nI only estimated average sigma per spectrum. Build regression from the difference between maximum and minimum predictions for the planet.</p>\n<p><strong>Models</strong>:<br>\nTwo types: heuristic and CNN.</p>\n<p><strong>Heuristic</strong>:<br>\nSum of features with weights that look realistic. 0.666 on training set and 0.666 on public. Two different approaches for predictions with differences &lt;0.00018 and larger. For larger spectra, found coefficients can be used. For small ones, better not to.</p>\n<p><strong>CNN</strong>:<br>\nMain challenge was avoiding overfitting. Experimentally established maximum frequency count in window - 21. There's physical sense in this, close to distance between different gases' transmission bands. Important that model doesn't overfit on two bands. Two-headed model, one predicting 283 spectrum points, second head predicting average sigma. Loss function - GaussianLoss.<br>\n0.678 on training set, 0.684 on test.</p>\n<p><strong>Mixing</strong>:<br>\nWeight selection by best RMSE on training set.<br>\n0.681 on training, 0.692 on public.</p>\n<p><strong>Didn't work on public</strong>:<br>\nPredicting gas compositions and summing predicted TauREX gas spectra.</p>\n<p>Final notebook <a href=\"https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791\" target=\"_blank\">https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791</a></p>",
  "messages": [
    {
      "id": 3033229,
      "postDate": "2024-10-31T23:59:31.003Z",
      "content": "<p>It was an excellent competition! For many years, I've been following astrophysics news with great interest. What else will they see in these blurry pixels? When will they finally find a habitable planet? And now I've been lucky enough to contribute to this myself.</p>\n<p>I want to thank the competition hosts, the Ariel mission, and specifically <a href=\"https://www.kaggle.com/GordonYip\" target=\"_blank\">@GordonYip</a> for answering important questions. I'd also like to thank all the authors of the notebook with the correct data calibration process. Without this, my participation would hardly have been meaningful. It's so difficult to find information about calibrations in articles! Thank <a href=\"https://www.kaggle.com/shlomoron\" target=\"_blank\">@shlomoron</a> for the meme thread and funny comments 😀</p>\n<p>My solution isn't particularly sophisticated. Here's the breakdown:</p>\n<p><strong>Data Processing</strong>:<br>\nI noticed outliers in pixel data and smoothed them with a Gaussian in the frequency axis direction. The \"extra\" wavelengths not included in the final spectrum were somewhat useful here. Consistently improved by 0.002.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F71783%2Faf1f3b347ae2379d01f6be441ade29d5%2Fbang.jpg?generation=1730420235408838&amp;alt=media\" alt=\"\"><br>\n<strong>Smoothing</strong>:<br>\nNothing in the world is better than SG. It's fast enough and smooths everything. I even smoothed the predictions.</p>\n<p><strong>Transit Zone Prediction</strong>:<br>\nNothing is more painful than getting 0 in submission results. The main problem is incorrectly predicted transit boundaries. In the test, they're wider or narrower than those in the training set. Using SG smoothing in the time direction to find boundaries. Ensuring light flux decreases at the left boundary and increases at the right.</p>\n<p><strong>Features for Models</strong>:<br>\nThe feature construction process is similar to my baseline. We build a polynomial, find how to 'raise' the transit zone to align with the non-transit part. Used MAE and logpdf functions to evaluate how well we found the coefficient.<br>\nBuild a general polynomial for the frequency sum. For each frequency, try to find coefficients pulling the transit as B and external parts as A to this polynomial. 1 - A/B is the target feature for individual frequencies. Sums across frequency intervals. I know from theory that some zones correspond to absorption areas of basic gases CH4, H2O, CO2, CO, NH3. But the best frequency areas for building features turned out to be different. I used a mix of genetic algorithm and manual selection to find the best intervals.</p>\n<p><strong>Sigma Estimation</strong>:<br>\nI only estimated average sigma per spectrum. Build regression from the difference between maximum and minimum predictions for the planet.</p>\n<p><strong>Models</strong>:<br>\nTwo types: heuristic and CNN.</p>\n<p><strong>Heuristic</strong>:<br>\nSum of features with weights that look realistic. 0.666 on training set and 0.666 on public. Two different approaches for predictions with differences &lt;0.00018 and larger. For larger spectra, found coefficients can be used. For small ones, better not to.</p>\n<p><strong>CNN</strong>:<br>\nMain challenge was avoiding overfitting. Experimentally established maximum frequency count in window - 21. There's physical sense in this, close to distance between different gases' transmission bands. Important that model doesn't overfit on two bands. Two-headed model, one predicting 283 spectrum points, second head predicting average sigma. Loss function - GaussianLoss.<br>\n0.678 on training set, 0.684 on test.</p>\n<p><strong>Mixing</strong>:<br>\nWeight selection by best RMSE on training set.<br>\n0.681 on training, 0.692 on public.</p>\n<p><strong>Didn't work on public</strong>:<br>\nPredicting gas compositions and summing predicted TauREX gas spectra.</p>\n<p>Final notebook <a href=\"https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791\" target=\"_blank\">https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791</a></p>",
      "rawMarkdown": "It was an excellent competition! For many years, I've been following astrophysics news with great interest. What else will they see in these blurry pixels? When will they finally find a habitable planet? And now I've been lucky enough to contribute to this myself.\n\nI want to thank the competition hosts, the Ariel mission, and specifically @GordonYip for answering important questions. I'd also like to thank all the authors of the notebook with the correct data calibration process. Without this, my participation would hardly have been meaningful. It's so difficult to find information about calibrations in articles! Thank @shlomoron for the meme thread and funny comments 😀\n\nMy solution isn't particularly sophisticated. Here's the breakdown:\n\n**Data Processing**:\nI noticed outliers in pixel data and smoothed them with a Gaussian in the frequency axis direction. The \"extra\" wavelengths not included in the final spectrum were somewhat useful here. Consistently improved by 0.002.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F71783%2Faf1f3b347ae2379d01f6be441ade29d5%2Fbang.jpg?generation=1730420235408838&alt=media)\n**Smoothing**:\nNothing in the world is better than SG. It's fast enough and smooths everything. I even smoothed the predictions.\n\n**Transit Zone Prediction**:\nNothing is more painful than getting 0 in submission results. The main problem is incorrectly predicted transit boundaries. In the test, they're wider or narrower than those in the training set. Using SG smoothing in the time direction to find boundaries. Ensuring light flux decreases at the left boundary and increases at the right.\n\n**Features for Models**:\nThe feature construction process is similar to my baseline. We build a polynomial, find how to 'raise' the transit zone to align with the non-transit part. Used MAE and logpdf functions to evaluate how well we found the coefficient.\nBuild a general polynomial for the frequency sum. For each frequency, try to find coefficients pulling the transit as B and external parts as A to this polynomial. 1 - A/B is the target feature for individual frequencies. Sums across frequency intervals. I know from theory that some zones correspond to absorption areas of basic gases CH4, H2O, CO2, CO, NH3. But the best frequency areas for building features turned out to be different. I used a mix of genetic algorithm and manual selection to find the best intervals.\n\n**Sigma Estimation**:\nI only estimated average sigma per spectrum. Build regression from the difference between maximum and minimum predictions for the planet.\n\n**Models**:\nTwo types: heuristic and CNN.\n\n**Heuristic**:\nSum of features with weights that look realistic. 0.666 on training set and 0.666 on public. Two different approaches for predictions with differences <0.00018 and larger. For larger spectra, found coefficients can be used. For small ones, better not to.\n\n**CNN**:\nMain challenge was avoiding overfitting. Experimentally established maximum frequency count in window - 21. There's physical sense in this, close to distance between different gases' transmission bands. Important that model doesn't overfit on two bands. Two-headed model, one predicting 283 spectrum points, second head predicting average sigma. Loss function - GaussianLoss.\n0.678 on training set, 0.684 on test.\n\n**Mixing**:\nWeight selection by best RMSE on training set.\n0.681 on training, 0.692 on public.\n\n**Didn't work on public**:\nPredicting gas compositions and summing predicted TauREX gas spectra.\n\nFinal notebook https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791",
      "votes": 57
    },
    {
      "id": 3033233,
      "postDate": "2024-11-01T00:02:30.507Z",
      "content": "<p>Great job Sergei. Well deserved. When I joined this competition, I found your public notebook and build my solution off of that. Thanks for upping the level of the competition.</p>",
      "rawMarkdown": "Great job Sergei. Well deserved. When I joined this competition, I found your public notebook and build my solution off of that. Thanks for upping the level of the competition.",
      "votes": 6
    },
    {
      "id": 3033253,
      "postDate": "2024-11-01T00:29:28.180Z",
      "content": "<p>Congratz on 6th place! Thank you for being such a great partner for LOLz on the LB. I was rooting for you in the last couple of days to finish in the price range, and I'm glad you made it. Your public notebook was helpful for me, too. I remember that you linked something about 2dGP + MCMC, so why does this have nothing to do with your final solution? 🤣</p>",
      "rawMarkdown": "Congratz on 6th place! Thank you for being such a great partner for LOLz on the LB. I was rooting for you in the last couple of days to finish in the price range, and I'm glad you made it. Your public notebook was helpful for me, too. I remember that you linked something about 2dGP + MCMC, so why does this have nothing to do with your final solution? 🤣",
      "votes": 3,
      "replies": [
        {
          "id": 3033257,
          "postDate": "2024-11-01T00:33:04.963Z",
          "content": "<p>In this competition, I implemented and buried a huge number of different ideas. The MCMC idea worked poorly on the training dataset.</p>",
          "rawMarkdown": "In this competition, I implemented and buried a huge number of different ideas. The MCMC idea worked poorly on the training dataset.",
          "votes": 4
        },
        {
          "id": 3033258,
          "postDate": "2024-11-01T00:33:30.360Z",
          "content": "<p>Congrats on 4th place! Very impressive</p>",
          "rawMarkdown": "Congrats on 4th place! Very impressive",
          "votes": 1
        }
      ]
    },
    {
      "id": 3035670,
      "postDate": "2024-11-03T18:18:14.820Z",
      "content": "<p>This is a great way to think about the problem.🙂</p>",
      "rawMarkdown": "This is a great way to think about the problem.🙂"
    },
    {
      "id": 3033889,
      "postDate": "2024-11-01T15:42:30.717Z",
      "content": "<p>Congrats and thanks for sharing the solution! </p>",
      "rawMarkdown": "Congrats and thanks for sharing the solution! "
    },
    {
      "id": 3033764,
      "postDate": "2024-11-01T13:41:35.217Z",
      "content": "<p>A sincere congratulations from me as well for the gold medal. Well deserved spot! <br>\nYour solution is intriguing and I thank you for sharing the code. I learned many things from your discussions throughout the competition.</p>",
      "rawMarkdown": "A sincere congratulations from me as well for the gold medal. Well deserved spot! \nYour solution is intriguing and I thank you for sharing the code. I learned many things from your discussions throughout the competition."
    },
    {
      "id": 3033707,
      "postDate": "2024-11-01T12:51:27.590Z",
      "content": "<p>This is my first medal competition in which I participated and contributed. It was a bit difficult to understand and learn a whole new set of information. I ended up with 36.1 percent.<br>\nbut,<br>\nThanks for posting the solution. It would be great if other winners could also post their solutions!<br>\nAnd Ofcoure <a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a> man this code is inspiring and at the same time cool.</p>",
      "rawMarkdown": "This is my first medal competition in which I participated and contributed. It was a bit difficult to understand and learn a whole new set of information. I ended up with 36.1 percent.\nbut,\nThanks for posting the solution. It would be great if other winners could also post their solutions!\nAnd Ofcoure @sergeifironov man this code is inspiring and at the same time cool."
    },
    {
      "id": 3033602,
      "postDate": "2024-11-01T10:54:06.943Z",
      "content": "<p>Thanks for your sharing! May I ask what do you mean by</p>\n<blockquote>\n  <p>Experimentally established maximum frequency count in window - 21</p>\n</blockquote>",
      "rawMarkdown": "Thanks for your sharing! May I ask what do you mean by\n>Experimentally established maximum frequency count in window - 21",
      "replies": [
        {
          "id": 3033613,
          "postDate": "2024-11-01T10:59:11.780Z",
          "content": "<p>I used CNN where convolutions went along the wavelength axis. Each convolution with a kernel width of 2n+1 captures n points to the right and left of it. If we use two consecutive layers with kernels of 3 and 5, we capture +1+2 points on the right and left. If the sum of convolution kernels resulted in capturing more than 21 points, the public score dropped immediately. Up to 21 points, it correlated well with the training score.</p>",
          "rawMarkdown": "I used CNN where convolutions went along the wavelength axis. Each convolution with a kernel width of 2n+1 captures n points to the right and left of it. If we use two consecutive layers with kernels of 3 and 5, we capture +1+2 points on the right and left. If the sum of convolution kernels resulted in capturing more than 21 points, the public score dropped immediately. Up to 21 points, it correlated well with the training score.\n",
          "replies": [
            {
              "id": 3033618,
              "postDate": "2024-11-01T11:10:49.397Z",
              "content": "<p>So amazing! It is hard to imagine such a correlation and I believe you must have conducted numerous experiments to discover this. Greatly appreciate your sharing!</p>",
              "rawMarkdown": "So amazing! It is hard to imagine such a correlation and I believe you must have conducted numerous experiments to discover this. Greatly appreciate your sharing!"
            },
            {
              "id": 3033630,
              "postDate": "2024-11-01T11:21:57.427Z",
              "content": "<p>I believe this is related to the peaks of CH4, CO2, and H2O. As soon as two peaks from different gases fall into one window, it becomes too easy for the model to calculate their ratio and draw the corresponding peaks. However, this doesn't generalize well to the test set, where there are other molecules that the model can't capture in this way.</p>",
              "rawMarkdown": "I believe this is related to the peaks of CH4, CO2, and H2O. As soon as two peaks from different gases fall into one window, it becomes too easy for the model to calculate their ratio and draw the corresponding peaks. However, this doesn't generalize well to the test set, where there are other molecules that the model can't capture in this way.",
              "votes": 1
            },
            {
              "id": 3033635,
              "postDate": "2024-11-01T11:27:50.927Z",
              "content": "<p>Your interpretation is very convincing. I think this is the most interesting part of the competition.</p>",
              "rawMarkdown": "Your interpretation is very convincing. I think this is the most interesting part of the competition."
            }
          ]
        }
      ]
    },
    {
      "id": 3033279,
      "postDate": "2024-11-01T00:59:48.617Z",
      "content": "<p>Congratulations, Sergei. I'm also inspired greatly by your notebook. </p>\n<blockquote>\n  <p>CNN:<br>\n  Main challenge was avoiding overfitting. Experimentally established maximum frequency count in window - 21. There's physical sense in this, close to distance between different gases' transmission bands. Important that model doesn't overfit on two bands. Two-headed model, one predicting 283 spectrum points, second head predicting average sigma. Loss function - GaussianLoss.</p>\n</blockquote>\n<p>What is your CNN? 1d or 2d? like your public mobilenet? </p>",
      "rawMarkdown": "Congratulations, Sergei. I'm also inspired greatly by your notebook. \n\n>CNN:\nMain challenge was avoiding overfitting. Experimentally established maximum frequency count in window - 21. There's physical sense in this, close to distance between different gases' transmission bands. Important that model doesn't overfit on two bands. Two-headed model, one predicting 283 spectrum points, second head predicting average sigma. Loss function - GaussianLoss.\n\nWhat is your CNN? 1d or 2d? like your public mobilenet? ",
      "replies": [
        {
          "id": 3033284,
          "postDate": "2024-11-01T01:05:01.907Z",
          "content": "<p>1D. I have published the code. This model is already based on features obtained similarly to baseline. The main reason to use convolutions here is the high spread of features for neighboring wavelengths. Perhaps if there was more time and more submits per day, I could reproduce this solution for a 2d network as well. My work was close.</p>",
          "rawMarkdown": "1D. I have published the code. This model is already based on features obtained similarly to baseline. The main reason to use convolutions here is the high spread of features for neighboring wavelengths. Perhaps if there was more time and more submits per day, I could reproduce this solution for a 2d network as well. My work was close."
        }
      ]
    },
    {
      "id": 3033259,
      "postDate": "2024-11-01T00:33:39.073Z",
      "content": "<p>Thanks for the great notebook and congratulations for the gold!!! I wasn't able to generalize to the test data at all. Amazing that your CNN worked for the test data! </p>",
      "rawMarkdown": "Thanks for the great notebook and congratulations for the gold!!! I wasn't able to generalize to the test data at all. Amazing that your CNN worked for the test data! "
    },
    {
      "id": 3033243,
      "postDate": "2024-11-01T00:19:32.407Z",
      "content": "<p>congrats!<br>\nI'm still playing with overfitting lol. cv 0.667 lb 0.589, private lb 0.621, no idea what's the problem, trying to figured it out</p>",
      "rawMarkdown": "congrats!\nI'm still playing with overfitting lol. cv 0.667 lb 0.589, private lb 0.621, no idea what's the problem, trying to figured it out",
      "replies": [
        {
          "id": 3033245,
          "postDate": "2024-11-01T00:23:10.833Z",
          "content": "<p>btw, would you like to share your final notebook? really excited to know more details about your approach</p>",
          "rawMarkdown": "btw, would you like to share your final notebook? really excited to know more details about your approach",
          "replies": [
            {
              "id": 3033249,
              "postDate": "2024-11-01T00:26:40.920Z",
              "content": "<p><a href=\"https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791\" target=\"_blank\">https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791</a></p>\n<p>Yes, of course. But the code is quite difficult to understand, I code very quickly but terribly messy.</p>",
              "rawMarkdown": "https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791\n\nYes, of course. But the code is quite difficult to understand, I code very quickly but terribly messy.",
              "votes": 1
            },
            {
              "id": 3033256,
              "postDate": "2024-11-01T00:33:02.827Z",
              "content": "<p>Sure, thank you, will look into it. :)</p>",
              "rawMarkdown": "Sure, thank you, will look into it. :)"
            }
          ]
        }
      ]
    },
    {
      "id": 3033241,
      "postDate": "2024-11-01T00:15:48.697Z",
      "content": "<p>Congratulations! Well deserved! The solution you shared has greatly benefited everyone.</p>",
      "rawMarkdown": "Congratulations! Well deserved! The solution you shared has greatly benefited everyone."
    },
    {
      "id": 3033238,
      "postDate": "2024-11-01T00:12:29.183Z",
      "content": "<p>Great job. Learned a lot from your codes and discussions.</p>",
      "rawMarkdown": "Great job. Learned a lot from your codes and discussions."
    },
    {
      "id": 3033423,
      "postDate": "2024-11-01T05:09:24.377Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!"
    }
  ],
  "comments": [
    {
      "id": 3033233,
      "author_name": "JungleBeastDS",
      "author_url": "",
      "post_date": "2024-11-01T00:02:30.507000",
      "content": "<p>Great job Sergei. Well deserved. When I joined this competition, I found your public notebook and build my solution off of that. Thanks for upping the level of the competition.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 3033253,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-11-01T00:29:28.180000",
      "content": "<p>Congratz on 6th place! Thank you for being such a great partner for LOLz on the LB. I was rooting for you in the last couple of days to finish in the price range, and I'm glad you made it. Your public notebook was helpful for me, too. I remember that you linked something about 2dGP + MCMC, so why does this have nothing to do with your final solution? 🤣</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3033257,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-11-01T00:33:04.963000",
          "content": "<p>In this competition, I implemented and buried a huge number of different ideas. The MCMC idea worked poorly on the training dataset.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 3033258,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-11-01T00:33:30.360000",
          "content": "<p>Congrats on 4th place! Very impressive</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3035670,
      "author_name": "Data Warrior",
      "author_url": "",
      "post_date": "2024-11-03T18:18:14.820000",
      "content": "<p>This is a great way to think about the problem.🙂</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033889,
      "author_name": "SCHEN",
      "author_url": "",
      "post_date": "2024-11-01T15:42:30.717000",
      "content": "<p>Congrats and thanks for sharing the solution! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033764,
      "author_name": "Andrei Zamfir",
      "author_url": "",
      "post_date": "2024-11-01T13:41:35.217000",
      "content": "<p>A sincere congratulations from me as well for the gold medal. Well deserved spot! <br>\nYour solution is intriguing and I thank you for sharing the code. I learned many things from your discussions throughout the competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033707,
      "author_name": "_student",
      "author_url": "",
      "post_date": "2024-11-01T12:51:27.590000",
      "content": "<p>This is my first medal competition in which I participated and contributed. It was a bit difficult to understand and learn a whole new set of information. I ended up with 36.1 percent.<br>\nbut,<br>\nThanks for posting the solution. It would be great if other winners could also post their solutions!<br>\nAnd Ofcoure <a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a> man this code is inspiring and at the same time cool.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033602,
      "author_name": "Zhu Siqi",
      "author_url": "",
      "post_date": "2024-11-01T10:54:06.943000",
      "content": "<p>Thanks for your sharing! May I ask what do you mean by</p>\n<blockquote>\n  <p>Experimentally established maximum frequency count in window - 21</p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 3033613,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-11-01T10:59:11.780000",
          "content": "<p>I used CNN where convolutions went along the wavelength axis. Each convolution with a kernel width of 2n+1 captures n points to the right and left of it. If we use two consecutive layers with kernels of 3 and 5, we capture +1+2 points on the right and left. If the sum of convolution kernels resulted in capturing more than 21 points, the public score dropped immediately. Up to 21 points, it correlated well with the training score.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3033618,
              "author_name": "Zhu Siqi",
              "author_url": "",
              "post_date": "2024-11-01T11:10:49.397000",
              "content": "<p>So amazing! It is hard to imagine such a correlation and I believe you must have conducted numerous experiments to discover this. Greatly appreciate your sharing!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3033630,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-11-01T11:21:57.427000",
              "content": "<p>I believe this is related to the peaks of CH4, CO2, and H2O. As soon as two peaks from different gases fall into one window, it becomes too easy for the model to calculate their ratio and draw the corresponding peaks. However, this doesn't generalize well to the test set, where there are other molecules that the model can't capture in this way.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3033635,
              "author_name": "Zhu Siqi",
              "author_url": "",
              "post_date": "2024-11-01T11:27:50.927000",
              "content": "<p>Your interpretation is very convincing. I think this is the most interesting part of the competition.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3033279,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2024-11-01T00:59:48.617000",
      "content": "<p>Congratulations, Sergei. I'm also inspired greatly by your notebook. </p>\n<blockquote>\n  <p>CNN:<br>\n  Main challenge was avoiding overfitting. Experimentally established maximum frequency count in window - 21. There's physical sense in this, close to distance between different gases' transmission bands. Important that model doesn't overfit on two bands. Two-headed model, one predicting 283 spectrum points, second head predicting average sigma. Loss function - GaussianLoss.</p>\n</blockquote>\n<p>What is your CNN? 1d or 2d? like your public mobilenet? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3033284,
          "author_name": "Sergei Fironov",
          "author_url": "",
          "post_date": "2024-11-01T01:05:01.907000",
          "content": "<p>1D. I have published the code. This model is already based on features obtained similarly to baseline. The main reason to use convolutions here is the high spread of features for neighboring wavelengths. Perhaps if there was more time and more submits per day, I could reproduce this solution for a 2d network as well. My work was close.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3033259,
      "author_name": "🐢 Jun Koda",
      "author_url": "",
      "post_date": "2024-11-01T00:33:39.073000",
      "content": "<p>Thanks for the great notebook and congratulations for the gold!!! I wasn't able to generalize to the test data at all. Amazing that your CNN worked for the test data! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033243,
      "author_name": "ZHANG XINGHAN",
      "author_url": "",
      "post_date": "2024-11-01T00:19:32.407000",
      "content": "<p>congrats!<br>\nI'm still playing with overfitting lol. cv 0.667 lb 0.589, private lb 0.621, no idea what's the problem, trying to figured it out</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3033245,
          "author_name": "ZHANG XINGHAN",
          "author_url": "",
          "post_date": "2024-11-01T00:23:10.833000",
          "content": "<p>btw, would you like to share your final notebook? really excited to know more details about your approach</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3033249,
              "author_name": "Sergei Fironov",
              "author_url": "",
              "post_date": "2024-11-01T00:26:40.920000",
              "content": "<p><a href=\"https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791\" target=\"_blank\">https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791</a></p>\n<p>Yes, of course. But the code is quite difficult to understand, I code very quickly but terribly messy.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3033256,
              "author_name": "ZHANG XINGHAN",
              "author_url": "",
              "post_date": "2024-11-01T00:33:02.827000",
              "content": "<p>Sure, thank you, will look into it. :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3033241,
      "author_name": "Zhuang Jia",
      "author_url": "",
      "post_date": "2024-11-01T00:15:48.697000",
      "content": "<p>Congratulations! Well deserved! The solution you shared has greatly benefited everyone.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033238,
      "author_name": "Natan Labarrère",
      "author_url": "",
      "post_date": "2024-11-01T00:12:29.183000",
      "content": "<p>Great job. Learned a lot from your codes and discussions.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3033423,
      "author_name": "Donal",
      "author_url": "",
      "post_date": "2024-11-01T05:09:24.377000",
      "content": "<p>Congratulations!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3033229": "It was an excellent competition! For many years, I've been following astrophysics news with great interest. What else will they see in these blurry pixels? When will they finally find a habitable planet? And now I've been lucky enough to contribute to this myself.\n\nI want to thank the competition hosts, the Ariel mission, and specifically @GordonYip for answering important questions. I'd also like to thank all the authors of the notebook with the correct data calibration process. Without this, my participation would hardly have been meaningful. It's so difficult to find information about calibrations in articles! Thank @shlomoron for the meme thread and funny comments 😀\n\nMy solution isn't particularly sophisticated. Here's the breakdown:\n\n**Data Processing**:\nI noticed outliers in pixel data and smoothed them with a Gaussian in the frequency axis direction. The \"extra\" wavelengths not included in the final spectrum were somewhat useful here. Consistently improved by 0.002.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F71783%2Faf1f3b347ae2379d01f6be441ade29d5%2Fbang.jpg?generation=1730420235408838&alt=media)\n**Smoothing**:\nNothing in the world is better than SG. It's fast enough and smooths everything. I even smoothed the predictions.\n\n**Transit Zone Prediction**:\nNothing is more painful than getting 0 in submission results. The main problem is incorrectly predicted transit boundaries. In the test, they're wider or narrower than those in the training set. Using SG smoothing in the time direction to find boundaries. Ensuring light flux decreases at the left boundary and increases at the right.\n\n**Features for Models**:\nThe feature construction process is similar to my baseline. We build a polynomial, find how to 'raise' the transit zone to align with the non-transit part. Used MAE and logpdf functions to evaluate how well we found the coefficient.\nBuild a general polynomial for the frequency sum. For each frequency, try to find coefficients pulling the transit as B and external parts as A to this polynomial. 1 - A/B is the target feature for individual frequencies. Sums across frequency intervals. I know from theory that some zones correspond to absorption areas of basic gases CH4, H2O, CO2, CO, NH3. But the best frequency areas for building features turned out to be different. I used a mix of genetic algorithm and manual selection to find the best intervals.\n\n**Sigma Estimation**:\nI only estimated average sigma per spectrum. Build regression from the difference between maximum and minimum predictions for the planet.\n\n**Models**:\nTwo types: heuristic and CNN.\n\n**Heuristic**:\nSum of features with weights that look realistic. 0.666 on training set and 0.666 on public. Two different approaches for predictions with differences <0.00018 and larger. For larger spectra, found coefficients can be used. For small ones, better not to.\n\n**CNN**:\nMain challenge was avoiding overfitting. Experimentally established maximum frequency count in window - 21. There's physical sense in this, close to distance between different gases' transmission bands. Important that model doesn't overfit on two bands. Two-headed model, one predicting 283 spectrum points, second head predicting average sigma. Loss function - GaussianLoss.\n0.678 on training set, 0.684 on test.\n\n**Mixing**:\nWeight selection by best RMSE on training set.\n0.681 on training, 0.692 on public.\n\n**Didn't work on public**:\nPredicting gas compositions and summing predicted TauREX gas spectra.\n\nFinal notebook https://www.kaggle.com/code/sergeifironov/new-blend-pdf?scriptVersionId=204505791",
    "3033233": "Great job Sergei. Well deserved. When I joined this competition, I found your public notebook and build my solution off of that. Thanks for upping the level of the competition.",
    "3033253": "Congratz on 6th place! Thank you for being such a great partner for LOLz on the LB. I was rooting for you in the last couple of days to finish in the price range, and I'm glad you made it. Your public notebook was helpful for me, too. I remember that you linked something about 2dGP + MCMC, so why does this have nothing to do with your final solution? 🤣",
    "3035670": "This is a great way to think about the problem.🙂",
    "3033889": "Congrats and thanks for sharing the solution! ",
    "3033764": "A sincere congratulations from me as well for the gold medal. Well deserved spot! \nYour solution is intriguing and I thank you for sharing the code. I learned many things from your discussions throughout the competition.",
    "3033707": "This is my first medal competition in which I participated and contributed. It was a bit difficult to understand and learn a whole new set of information. I ended up with 36.1 percent.\nbut,\nThanks for posting the solution. It would be great if other winners could also post their solutions!\nAnd Ofcoure @sergeifironov man this code is inspiring and at the same time cool.",
    "3033602": "Thanks for your sharing! May I ask what do you mean by\n>Experimentally established maximum frequency count in window - 21",
    "3033279": "Congratulations, Sergei. I'm also inspired greatly by your notebook. \n\n>CNN:\nMain challenge was avoiding overfitting. Experimentally established maximum frequency count in window - 21. There's physical sense in this, close to distance between different gases' transmission bands. Important that model doesn't overfit on two bands. Two-headed model, one predicting 283 spectrum points, second head predicting average sigma. Loss function - GaussianLoss.\n\nWhat is your CNN? 1d or 2d? like your public mobilenet? ",
    "3033259": "Thanks for the great notebook and congratulations for the gold!!! I wasn't able to generalize to the test data at all. Amazing that your CNN worked for the test data! ",
    "3033243": "congrats!\nI'm still playing with overfitting lol. cv 0.667 lb 0.589, private lb 0.621, no idea what's the problem, trying to figured it out",
    "3033241": "Congratulations! Well deserved! The solution you shared has greatly benefited everyone.",
    "3033238": "Great job. Learned a lot from your codes and discussions.",
    "3033423": "Congratulations!"
  }
}