{
  "id": 543681,
  "title": "Quick Update on My Solution(15th)",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543681",
  "author_name": "takaito",
  "post_date": "2024-11-01T00:44:40.714000",
  "votes": 19,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I’d like to briefly share my solution here. <br>\nI hope to make time later to provide a more detailed write-up. <br>\nMy approach is as follows:<br>\n1.Use the baseline <a href=\"https://www.kaggle.com/code/sergeifironov/ariel-only-correlation\" target=\"_blank\">notebook</a> to determine the mean_s for each planet.<br>\n2.Use multiple deep learning models to predict the difference between the true value and 3.mean_s (with the input waveform transformed using mean_s).<br>\nCalculate sigma as the standard deviation of predictions from multiple models and adjust the scale.<br>\nAs mentioned in the Discussion, I thought it might be difficult to directly predict each planet’s wavelength. Given that I couldn’t gain enough domain knowledge in time, I decided to rely on the tips I had acquired on Kaggle instead of focusing on preprocessing. As a result, I chose to train a model to predict the difference between mean_s and each true wavelength. This approach has two significant advantages:<br>\nNo Need for Absolute Wavelength Prediction: Planetary wavelengths vary greatly, and attempting to predict them directly may not work well with the provided dataset. Predicting the difference, on the other hand, only requires predicting the deviation, making the task simpler. Also, since absolute-level prediction isn’t needed, we can normalize the input data to ensure consistency and reduce possible patterns in the input data, making this approach well-suited for a competition with limited samples and unknown planetary data in the test set.<br>\nThe Target Variable Has a Zero-Centered Distribution: This brings various benefits, which I plan to detail in another post.<br>\nFor predicting the difference, I transformed the input waveform data using the calculated mean_s. Specifically, I multiplied the interval where the planet passes by the star by the calculated (mean_s + 1). If (mean_s + 1) matches the true value, the transformed data should allow for a reasonable polynomial (of about the third degree). If it differs significantly from the true value, this transformation should result in a noticeable deviation, which the model can detect and predict.<br>\nTo enhance training, I added random noise to the input data and applied time-axis flipping for augmentation. For cross-validation, I used Group-2-fold, splitting based on star information.<br>\nI aimed to keep the model as simple as possible, and both BiLSTM and CNN worked well with my approach. Since the input was quite noisy, I found that adding a filter to average neighboring values significantly improved the cross-validation score.<br>\nOne of the biggest challenges was estimating sigma. I considered several sigma estimation methods, reviewing the Discussion and research papers like <a href=\"https://arxiv.org/abs/1612.01474\" target=\"_blank\">Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles</a>, and tested the following approaches:</p>\n<ol>\n<li>Directly Using the Loss as the Evaluation Metric: This involves using GaussianNLLLoss as the loss function. I ultimately abandoned this approach since the training was very unstable.</li>\n<li>Using Quantile Regression to Estimate the 15 and 85 Percentiles and Calculating the Difference: This method was mentioned in a Discussion post, <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/532740\" target=\"_blank\">Estimate sigma through quantile regression</a>, and initially boosted my score. Although I didn’t use it for the final sigma, I kept the quantile regression loss as part of the custom loss function.</li>\n<li>Using the Range Between the Maximum and Minimum Predictions of Multiple Models: I found this method effective with a moderate number of models, although with a larger ensemble, it became susceptible to outliers. Despite my reservations, this method did help to improve the score.</li>\n<li>Using the Standard Deviation of Predictions from Multiple Models: This method tends to perform better with a larger number of models, and I ultimately chose this approach.<br>\nHowever, regardless of the method, sigma estimates didn’t perform well on the test set, posing a significant challenge. Estimating sigma for unseen data proved difficult, and I spent considerable time working on this issue. Eventually, I settled on a target range based on the public leaderboard score.<br>\nI will post further details in a Discussion post later.</li>\n</ol>",
  "messages": [
    {
      "id": 3033268,
      "postDate": "2024-11-01T00:44:40.713Z",
      "content": "<p>I’d like to briefly share my solution here. <br>\nI hope to make time later to provide a more detailed write-up. <br>\nMy approach is as follows:<br>\n1.Use the baseline <a href=\"https://www.kaggle.com/code/sergeifironov/ariel-only-correlation\" target=\"_blank\">notebook</a> to determine the mean_s for each planet.<br>\n2.Use multiple deep learning models to predict the difference between the true value and 3.mean_s (with the input waveform transformed using mean_s).<br>\nCalculate sigma as the standard deviation of predictions from multiple models and adjust the scale.<br>\nAs mentioned in the Discussion, I thought it might be difficult to directly predict each planet’s wavelength. Given that I couldn’t gain enough domain knowledge in time, I decided to rely on the tips I had acquired on Kaggle instead of focusing on preprocessing. As a result, I chose to train a model to predict the difference between mean_s and each true wavelength. This approach has two significant advantages:<br>\nNo Need for Absolute Wavelength Prediction: Planetary wavelengths vary greatly, and attempting to predict them directly may not work well with the provided dataset. Predicting the difference, on the other hand, only requires predicting the deviation, making the task simpler. Also, since absolute-level prediction isn’t needed, we can normalize the input data to ensure consistency and reduce possible patterns in the input data, making this approach well-suited for a competition with limited samples and unknown planetary data in the test set.<br>\nThe Target Variable Has a Zero-Centered Distribution: This brings various benefits, which I plan to detail in another post.<br>\nFor predicting the difference, I transformed the input waveform data using the calculated mean_s. Specifically, I multiplied the interval where the planet passes by the star by the calculated (mean_s + 1). If (mean_s + 1) matches the true value, the transformed data should allow for a reasonable polynomial (of about the third degree). If it differs significantly from the true value, this transformation should result in a noticeable deviation, which the model can detect and predict.<br>\nTo enhance training, I added random noise to the input data and applied time-axis flipping for augmentation. For cross-validation, I used Group-2-fold, splitting based on star information.<br>\nI aimed to keep the model as simple as possible, and both BiLSTM and CNN worked well with my approach. Since the input was quite noisy, I found that adding a filter to average neighboring values significantly improved the cross-validation score.<br>\nOne of the biggest challenges was estimating sigma. I considered several sigma estimation methods, reviewing the Discussion and research papers like <a href=\"https://arxiv.org/abs/1612.01474\" target=\"_blank\">Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles</a>, and tested the following approaches:</p>\n<ol>\n<li>Directly Using the Loss as the Evaluation Metric: This involves using GaussianNLLLoss as the loss function. I ultimately abandoned this approach since the training was very unstable.</li>\n<li>Using Quantile Regression to Estimate the 15 and 85 Percentiles and Calculating the Difference: This method was mentioned in a Discussion post, <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/532740\" target=\"_blank\">Estimate sigma through quantile regression</a>, and initially boosted my score. Although I didn’t use it for the final sigma, I kept the quantile regression loss as part of the custom loss function.</li>\n<li>Using the Range Between the Maximum and Minimum Predictions of Multiple Models: I found this method effective with a moderate number of models, although with a larger ensemble, it became susceptible to outliers. Despite my reservations, this method did help to improve the score.</li>\n<li>Using the Standard Deviation of Predictions from Multiple Models: This method tends to perform better with a larger number of models, and I ultimately chose this approach.<br>\nHowever, regardless of the method, sigma estimates didn’t perform well on the test set, posing a significant challenge. Estimating sigma for unseen data proved difficult, and I spent considerable time working on this issue. Eventually, I settled on a target range based on the public leaderboard score.<br>\nI will post further details in a Discussion post later.</li>\n</ol>",
      "rawMarkdown": "\n\n\nI’d like to briefly share my solution here. \nI hope to make time later to provide a more detailed write-up. \nMy approach is as follows:\n1.Use the baseline [notebook](https://www.kaggle.com/code/sergeifironov/ariel-only-correlation) to determine the mean_s for each planet.\n2.Use multiple deep learning models to predict the difference between the true value and 3.mean_s (with the input waveform transformed using mean_s).\n\nCalculate sigma as the standard deviation of predictions from multiple models and adjust the scale.\nAs mentioned in the Discussion, I thought it might be difficult to directly predict each planet’s wavelength. Given that I couldn’t gain enough domain knowledge in time, I decided to rely on the tips I had acquired on Kaggle instead of focusing on preprocessing. As a result, I chose to train a model to predict the difference between mean_s and each true wavelength. This approach has two significant advantages:\n\nNo Need for Absolute Wavelength Prediction: Planetary wavelengths vary greatly, and attempting to predict them directly may not work well with the provided dataset. Predicting the difference, on the other hand, only requires predicting the deviation, making the task simpler. Also, since absolute-level prediction isn’t needed, we can normalize the input data to ensure consistency and reduce possible patterns in the input data, making this approach well-suited for a competition with limited samples and unknown planetary data in the test set.\n\nThe Target Variable Has a Zero-Centered Distribution: This brings various benefits, which I plan to detail in another post.\n\nFor predicting the difference, I transformed the input waveform data using the calculated mean_s. Specifically, I multiplied the interval where the planet passes by the star by the calculated (mean_s + 1). If (mean_s + 1) matches the true value, the transformed data should allow for a reasonable polynomial (of about the third degree). If it differs significantly from the true value, this transformation should result in a noticeable deviation, which the model can detect and predict.\n\nTo enhance training, I added random noise to the input data and applied time-axis flipping for augmentation. For cross-validation, I used Group-2-fold, splitting based on star information.\n\nI aimed to keep the model as simple as possible, and both BiLSTM and CNN worked well with my approach. Since the input was quite noisy, I found that adding a filter to average neighboring values significantly improved the cross-validation score.\n\nOne of the biggest challenges was estimating sigma. I considered several sigma estimation methods, reviewing the Discussion and research papers like [Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles](https://arxiv.org/abs/1612.01474), and tested the following approaches:\n\n1. Directly Using the Loss as the Evaluation Metric: This involves using GaussianNLLLoss as the loss function. I ultimately abandoned this approach since the training was very unstable.\n2. Using Quantile Regression to Estimate the 15 and 85 Percentiles and Calculating the Difference: This method was mentioned in a Discussion post, [Estimate sigma through quantile regression](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/532740), and initially boosted my score. Although I didn’t use it for the final sigma, I kept the quantile regression loss as part of the custom loss function.\n3. Using the Range Between the Maximum and Minimum Predictions of Multiple Models: I found this method effective with a moderate number of models, although with a larger ensemble, it became susceptible to outliers. Despite my reservations, this method did help to improve the score.\n4. Using the Standard Deviation of Predictions from Multiple Models: This method tends to perform better with a larger number of models, and I ultimately chose this approach.\nHowever, regardless of the method, sigma estimates didn’t perform well on the test set, posing a significant challenge. Estimating sigma for unseen data proved difficult, and I spent considerable time working on this issue. Eventually, I settled on a target range based on the public leaderboard score.\n\nI will post further details in a Discussion post later.\n\n",
      "votes": 19
    },
    {
      "id": 3037036,
      "postDate": "2024-11-05T08:55:42.443Z",
      "content": "<p>I appreciate you sharing that information.</p>",
      "rawMarkdown": "I appreciate you sharing that information."
    }
  ],
  "comments": [
    {
      "id": 3037036,
      "author_name": "Humayra Khanom Rime",
      "author_url": "",
      "post_date": "2024-11-05T08:55:42.443000",
      "content": "<p>I appreciate you sharing that information.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3033268": "\n\n\nI’d like to briefly share my solution here. \nI hope to make time later to provide a more detailed write-up. \nMy approach is as follows:\n1.Use the baseline [notebook](https://www.kaggle.com/code/sergeifironov/ariel-only-correlation) to determine the mean_s for each planet.\n2.Use multiple deep learning models to predict the difference between the true value and 3.mean_s (with the input waveform transformed using mean_s).\n\nCalculate sigma as the standard deviation of predictions from multiple models and adjust the scale.\nAs mentioned in the Discussion, I thought it might be difficult to directly predict each planet’s wavelength. Given that I couldn’t gain enough domain knowledge in time, I decided to rely on the tips I had acquired on Kaggle instead of focusing on preprocessing. As a result, I chose to train a model to predict the difference between mean_s and each true wavelength. This approach has two significant advantages:\n\nNo Need for Absolute Wavelength Prediction: Planetary wavelengths vary greatly, and attempting to predict them directly may not work well with the provided dataset. Predicting the difference, on the other hand, only requires predicting the deviation, making the task simpler. Also, since absolute-level prediction isn’t needed, we can normalize the input data to ensure consistency and reduce possible patterns in the input data, making this approach well-suited for a competition with limited samples and unknown planetary data in the test set.\n\nThe Target Variable Has a Zero-Centered Distribution: This brings various benefits, which I plan to detail in another post.\n\nFor predicting the difference, I transformed the input waveform data using the calculated mean_s. Specifically, I multiplied the interval where the planet passes by the star by the calculated (mean_s + 1). If (mean_s + 1) matches the true value, the transformed data should allow for a reasonable polynomial (of about the third degree). If it differs significantly from the true value, this transformation should result in a noticeable deviation, which the model can detect and predict.\n\nTo enhance training, I added random noise to the input data and applied time-axis flipping for augmentation. For cross-validation, I used Group-2-fold, splitting based on star information.\n\nI aimed to keep the model as simple as possible, and both BiLSTM and CNN worked well with my approach. Since the input was quite noisy, I found that adding a filter to average neighboring values significantly improved the cross-validation score.\n\nOne of the biggest challenges was estimating sigma. I considered several sigma estimation methods, reviewing the Discussion and research papers like [Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles](https://arxiv.org/abs/1612.01474), and tested the following approaches:\n\n1. Directly Using the Loss as the Evaluation Metric: This involves using GaussianNLLLoss as the loss function. I ultimately abandoned this approach since the training was very unstable.\n2. Using Quantile Regression to Estimate the 15 and 85 Percentiles and Calculating the Difference: This method was mentioned in a Discussion post, [Estimate sigma through quantile regression](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/532740), and initially boosted my score. Although I didn’t use it for the final sigma, I kept the quantile regression loss as part of the custom loss function.\n3. Using the Range Between the Maximum and Minimum Predictions of Multiple Models: I found this method effective with a moderate number of models, although with a larger ensemble, it became susceptible to outliers. Despite my reservations, this method did help to improve the score.\n4. Using the Standard Deviation of Predictions from Multiple Models: This method tends to perform better with a larger number of models, and I ultimately chose this approach.\nHowever, regardless of the method, sigma estimates didn’t perform well on the test set, posing a significant challenge. Estimating sigma for unseen data proved difficult, and I spent considerable time working on this issue. Eventually, I settled on a target range based on the public leaderboard score.\n\nI will post further details in a Discussion post later.\n\n",
    "3037036": "I appreciate you sharing that information."
  }
}