{
  "id": 609227,
  "title": "58th place solution",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609227",
  "author_name": "Kh0a",
  "post_date": "2025-09-25T03:52:33.846000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, I would like to express my gratitude to the organizers and fellow competitors for hosting such an incredible competition. I joined this challenge quite late and focused intensively on building my solution during the final 7 days. Initially, I explored public notebooks and came across <a href=\"https://www.kaggle.com/code/qaedtgyh/resnet-for-airs-0-360-lb\" target=\"_blank\">ResNet for AIRS 0.360 LB</a>, which inspired me to train my own version of the model. I implemented a 5-fold Cross-Validation approach and began optimizing the solution from there.</p>\n<h1>Solution Overview</h1>\n<p>My solution combined domain knowledge, feature engineering, and advanced machine learning techniques. The key components of the pipeline were:</p>\n<p><strong>Signal Preprocessing:</strong> Calibrating and preprocessing raw signals from the Ariel telescope.<br>\n<strong>Transit Depth Prediction:</strong> Using a ResNet-based model to predict transit depths.<br>\n<strong>Uncertainty Estimation:</strong> Training a SigmaNet model to predict uncertainties (sigma values) for both FGS and AIRS channels.<br>\n<strong>Post-Processing and Blending:</strong> Applying smoothing and blending techniques to improve predictions.<br>\n<strong>Ensemble Learning:</strong> Leveraging a 5-fold ensemble for robust predictions.</p>\n<h1>Data Preprocessing</h1>\n<h2>Signal Calibration</h2>\n<p>Signals from the FGS1 and AIRS-CH0 sensors were calibrated using dark, flat, and linear correction files.<br>\nA custom preprocessing pipeline was implemented to handle sensor-specific nuances, such as binning and clipping.</p>\n<h2>Feature Engineering</h2>\n<h4>Features for ResNet</h4>\n<p>Extracted astrophysical features such as <strong>transit_depth, Rs, Ms, Ts, Mp, e, P, sma,</strong> and <strong>i</strong>.</p>\n<h4>Features for SigmaNet</h4>\n<p>Generated additional features from the predicted transit depths (mu) and ensemble statistics:<br>\n<strong>Mean</strong> and <strong>standard deviation</strong> of AIRS predictions.<br>\nDispersion metrics across the 5-fold ensemble.</p>\n<h1>Modeling</h1>\n<h2>Transit Depth Prediction</h2>\n<p><strong>Model Architecture:</strong> A ResNet-based MLP (ResNetMLP2) with residual blocks, Squeeze-and-Excitation (SE) layers, and Gaussian noise for regularization. Can view model detail <a href=\"https://www.kaggle.com/code/llkh0a/5-folds-full-with-se-blocks#SE-structure\" target=\"_blank\">here</a><br>\n<strong>Hyperparameter Optimization:</strong> Used Optuna to tune key parameters such as learning rate, number of blocks, dropout rate, and hidden dimensions.<br>\n<strong>Ensemble:</strong> Trained 5 models using stratified 5-fold cross-validation. Predictions were averaged across folds for robustness.<br>\n<em>Total trainable params each fold:  1,851,035</em></p>\n<h2>Uncertainty Estimation</h2>\n<p><strong>SigmaNet:</strong> A custom neural network (SigmaNetMul) was trained to predict multiplicative residuals over baseline sigma values.<a href=\"https://www.kaggle.com/code/llkh0a/5-folds-full-with-se-blocks#Sigma-net\" target=\"_blank\">SigmaNet</a><br>\n<strong>Loss Function:</strong> Maximized the competition's Gaussian Log-Likelihood (GLL) score using a differentiable implementation in PyTorch.<br>\n<strong>Features:</strong> Combined astrophysical features, OOF predictions, and ensemble statistics to create a comprehensive feature set for sigma prediction.<br>\n<em>Total params: 807,195</em></p>\n<h2>Post-Processing</h2>\n<p>Applied <strong>Savitzky-Golay</strong> smoothing to AIRS predictions to reduce noise.<br>\nBlended raw and smoothed AIRS predictions using a weighted average (0.65 raw, 0.35 smoothed).</p>\n<h1>What Did Not Work</h1>\n<h2>Multitask Model</h2>\n<p>One of the approaches I explored was building a multitask model that simultaneously predicted both the transit depth (mu) and the uncertainty (sigma) in a single architecture. The idea was to share the feature extraction layers between the two tasks, leveraging the shared information to improve both predictions. However, this approach did not yield better results compared to the separate models for mu (ResNet) and sigma (SigmaNet). The multitask model struggled to balance the competing objectives of minimizing the Gaussian Log-Likelihood (GLL) loss for mu and accurately estimating sigma. This imbalance often led to suboptimal predictions for both tasks.</p>\n<h2>SpectralTransformer</h2>\n<p>One of the approaches I attempted was using a SpectralTransformer architecture to model the spectral data. The idea was to leverage the transformer's ability to capture long-range dependencies across wavelengths. However, this approach did not yield competitive results. The primary issue was that I fed raw features directly into the SpectralTransformer which make the model training process does not go well</p>",
  "messages": [
    {
      "id": 3293945,
      "postDate": "2025-09-25T03:52:33.847Z",
      "content": "<p>First, I would like to express my gratitude to the organizers and fellow competitors for hosting such an incredible competition. I joined this challenge quite late and focused intensively on building my solution during the final 7 days. Initially, I explored public notebooks and came across <a href=\"https://www.kaggle.com/code/qaedtgyh/resnet-for-airs-0-360-lb\" target=\"_blank\">ResNet for AIRS 0.360 LB</a>, which inspired me to train my own version of the model. I implemented a 5-fold Cross-Validation approach and began optimizing the solution from there.</p>\n<h1>Solution Overview</h1>\n<p>My solution combined domain knowledge, feature engineering, and advanced machine learning techniques. The key components of the pipeline were:</p>\n<p><strong>Signal Preprocessing:</strong> Calibrating and preprocessing raw signals from the Ariel telescope.<br>\n<strong>Transit Depth Prediction:</strong> Using a ResNet-based model to predict transit depths.<br>\n<strong>Uncertainty Estimation:</strong> Training a SigmaNet model to predict uncertainties (sigma values) for both FGS and AIRS channels.<br>\n<strong>Post-Processing and Blending:</strong> Applying smoothing and blending techniques to improve predictions.<br>\n<strong>Ensemble Learning:</strong> Leveraging a 5-fold ensemble for robust predictions.</p>\n<h1>Data Preprocessing</h1>\n<h2>Signal Calibration</h2>\n<p>Signals from the FGS1 and AIRS-CH0 sensors were calibrated using dark, flat, and linear correction files.<br>\nA custom preprocessing pipeline was implemented to handle sensor-specific nuances, such as binning and clipping.</p>\n<h2>Feature Engineering</h2>\n<h4>Features for ResNet</h4>\n<p>Extracted astrophysical features such as <strong>transit_depth, Rs, Ms, Ts, Mp, e, P, sma,</strong> and <strong>i</strong>.</p>\n<h4>Features for SigmaNet</h4>\n<p>Generated additional features from the predicted transit depths (mu) and ensemble statistics:<br>\n<strong>Mean</strong> and <strong>standard deviation</strong> of AIRS predictions.<br>\nDispersion metrics across the 5-fold ensemble.</p>\n<h1>Modeling</h1>\n<h2>Transit Depth Prediction</h2>\n<p><strong>Model Architecture:</strong> A ResNet-based MLP (ResNetMLP2) with residual blocks, Squeeze-and-Excitation (SE) layers, and Gaussian noise for regularization. Can view model detail <a href=\"https://www.kaggle.com/code/llkh0a/5-folds-full-with-se-blocks#SE-structure\" target=\"_blank\">here</a><br>\n<strong>Hyperparameter Optimization:</strong> Used Optuna to tune key parameters such as learning rate, number of blocks, dropout rate, and hidden dimensions.<br>\n<strong>Ensemble:</strong> Trained 5 models using stratified 5-fold cross-validation. Predictions were averaged across folds for robustness.<br>\n<em>Total trainable params each fold:  1,851,035</em></p>\n<h2>Uncertainty Estimation</h2>\n<p><strong>SigmaNet:</strong> A custom neural network (SigmaNetMul) was trained to predict multiplicative residuals over baseline sigma values.<a href=\"https://www.kaggle.com/code/llkh0a/5-folds-full-with-se-blocks#Sigma-net\" target=\"_blank\">SigmaNet</a><br>\n<strong>Loss Function:</strong> Maximized the competition's Gaussian Log-Likelihood (GLL) score using a differentiable implementation in PyTorch.<br>\n<strong>Features:</strong> Combined astrophysical features, OOF predictions, and ensemble statistics to create a comprehensive feature set for sigma prediction.<br>\n<em>Total params: 807,195</em></p>\n<h2>Post-Processing</h2>\n<p>Applied <strong>Savitzky-Golay</strong> smoothing to AIRS predictions to reduce noise.<br>\nBlended raw and smoothed AIRS predictions using a weighted average (0.65 raw, 0.35 smoothed).</p>\n<h1>What Did Not Work</h1>\n<h2>Multitask Model</h2>\n<p>One of the approaches I explored was building a multitask model that simultaneously predicted both the transit depth (mu) and the uncertainty (sigma) in a single architecture. The idea was to share the feature extraction layers between the two tasks, leveraging the shared information to improve both predictions. However, this approach did not yield better results compared to the separate models for mu (ResNet) and sigma (SigmaNet). The multitask model struggled to balance the competing objectives of minimizing the Gaussian Log-Likelihood (GLL) loss for mu and accurately estimating sigma. This imbalance often led to suboptimal predictions for both tasks.</p>\n<h2>SpectralTransformer</h2>\n<p>One of the approaches I attempted was using a SpectralTransformer architecture to model the spectral data. The idea was to leverage the transformer's ability to capture long-range dependencies across wavelengths. However, this approach did not yield competitive results. The primary issue was that I fed raw features directly into the SpectralTransformer which make the model training process does not go well</p>",
      "rawMarkdown": "First, I would like to express my gratitude to the organizers and fellow competitors for hosting such an incredible competition. I joined this challenge quite late and focused intensively on building my solution during the final 7 days. Initially, I explored public notebooks and came across [ResNet for AIRS 0.360 LB](https://www.kaggle.com/code/qaedtgyh/resnet-for-airs-0-360-lb), which inspired me to train my own version of the model. I implemented a 5-fold Cross-Validation approach and began optimizing the solution from there.\n# Solution Overview \nMy solution combined domain knowledge, feature engineering, and advanced machine learning techniques. The key components of the pipeline were:\n\n**Signal Preprocessing:** Calibrating and preprocessing raw signals from the Ariel telescope.\n**Transit Depth Prediction:** Using a ResNet-based model to predict transit depths.\n**Uncertainty Estimation:** Training a SigmaNet model to predict uncertainties (sigma values) for both FGS and AIRS channels.\n**Post-Processing and Blending:** Applying smoothing and blending techniques to improve predictions.\n**Ensemble Learning:** Leveraging a 5-fold ensemble for robust predictions.\n# Data Preprocessing\n## Signal Calibration\nSignals from the FGS1 and AIRS-CH0 sensors were calibrated using dark, flat, and linear correction files.\nA custom preprocessing pipeline was implemented to handle sensor-specific nuances, such as binning and clipping.\n## Feature Engineering\n#### Features for ResNet\nExtracted astrophysical features such as **transit_depth, Rs, Ms, Ts, Mp, e, P, sma,** and **i**.\n#### Features for SigmaNet\nGenerated additional features from the predicted transit depths (mu) and ensemble statistics:\n**Mean** and **standard deviation** of AIRS predictions.\nDispersion metrics across the 5-fold ensemble.\n# Modeling\n## Transit Depth Prediction\n**Model Architecture:** A ResNet-based MLP (ResNetMLP2) with residual blocks, Squeeze-and-Excitation (SE) layers, and Gaussian noise for regularization. Can view model detail [here](https://www.kaggle.com/code/llkh0a/5-folds-full-with-se-blocks#SE-structure)\n**Hyperparameter Optimization:** Used Optuna to tune key parameters such as learning rate, number of blocks, dropout rate, and hidden dimensions.\n**Ensemble:** Trained 5 models using stratified 5-fold cross-validation. Predictions were averaged across folds for robustness.\n*Total trainable params each fold:  1,851,035*\n## Uncertainty Estimation\n**SigmaNet:** A custom neural network (SigmaNetMul) was trained to predict multiplicative residuals over baseline sigma values.[SigmaNet](https://www.kaggle.com/code/llkh0a/5-folds-full-with-se-blocks#Sigma-net)\n**Loss Function:** Maximized the competition's Gaussian Log-Likelihood (GLL) score using a differentiable implementation in PyTorch.\n**Features:** Combined astrophysical features, OOF predictions, and ensemble statistics to create a comprehensive feature set for sigma prediction.\n*Total params: 807,195*\n## Post-Processing\nApplied **Savitzky-Golay** smoothing to AIRS predictions to reduce noise.\nBlended raw and smoothed AIRS predictions using a weighted average (0.65 raw, 0.35 smoothed).\n# What Did Not Work\n## Multitask Model\nOne of the approaches I explored was building a multitask model that simultaneously predicted both the transit depth (mu) and the uncertainty (sigma) in a single architecture. The idea was to share the feature extraction layers between the two tasks, leveraging the shared information to improve both predictions. However, this approach did not yield better results compared to the separate models for mu (ResNet) and sigma (SigmaNet). The multitask model struggled to balance the competing objectives of minimizing the Gaussian Log-Likelihood (GLL) loss for mu and accurately estimating sigma. This imbalance often led to suboptimal predictions for both tasks.\n## SpectralTransformer\nOne of the approaches I attempted was using a SpectralTransformer architecture to model the spectral data. The idea was to leverage the transformer's ability to capture long-range dependencies across wavelengths. However, this approach did not yield competitive results. The primary issue was that I fed raw features directly into the SpectralTransformer which make the model training process does not go well\n\n\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3293945": "First, I would like to express my gratitude to the organizers and fellow competitors for hosting such an incredible competition. I joined this challenge quite late and focused intensively on building my solution during the final 7 days. Initially, I explored public notebooks and came across [ResNet for AIRS 0.360 LB](https://www.kaggle.com/code/qaedtgyh/resnet-for-airs-0-360-lb), which inspired me to train my own version of the model. I implemented a 5-fold Cross-Validation approach and began optimizing the solution from there.\n# Solution Overview \nMy solution combined domain knowledge, feature engineering, and advanced machine learning techniques. The key components of the pipeline were:\n\n**Signal Preprocessing:** Calibrating and preprocessing raw signals from the Ariel telescope.\n**Transit Depth Prediction:** Using a ResNet-based model to predict transit depths.\n**Uncertainty Estimation:** Training a SigmaNet model to predict uncertainties (sigma values) for both FGS and AIRS channels.\n**Post-Processing and Blending:** Applying smoothing and blending techniques to improve predictions.\n**Ensemble Learning:** Leveraging a 5-fold ensemble for robust predictions.\n# Data Preprocessing\n## Signal Calibration\nSignals from the FGS1 and AIRS-CH0 sensors were calibrated using dark, flat, and linear correction files.\nA custom preprocessing pipeline was implemented to handle sensor-specific nuances, such as binning and clipping.\n## Feature Engineering\n#### Features for ResNet\nExtracted astrophysical features such as **transit_depth, Rs, Ms, Ts, Mp, e, P, sma,** and **i**.\n#### Features for SigmaNet\nGenerated additional features from the predicted transit depths (mu) and ensemble statistics:\n**Mean** and **standard deviation** of AIRS predictions.\nDispersion metrics across the 5-fold ensemble.\n# Modeling\n## Transit Depth Prediction\n**Model Architecture:** A ResNet-based MLP (ResNetMLP2) with residual blocks, Squeeze-and-Excitation (SE) layers, and Gaussian noise for regularization. Can view model detail [here](https://www.kaggle.com/code/llkh0a/5-folds-full-with-se-blocks#SE-structure)\n**Hyperparameter Optimization:** Used Optuna to tune key parameters such as learning rate, number of blocks, dropout rate, and hidden dimensions.\n**Ensemble:** Trained 5 models using stratified 5-fold cross-validation. Predictions were averaged across folds for robustness.\n*Total trainable params each fold:  1,851,035*\n## Uncertainty Estimation\n**SigmaNet:** A custom neural network (SigmaNetMul) was trained to predict multiplicative residuals over baseline sigma values.[SigmaNet](https://www.kaggle.com/code/llkh0a/5-folds-full-with-se-blocks#Sigma-net)\n**Loss Function:** Maximized the competition's Gaussian Log-Likelihood (GLL) score using a differentiable implementation in PyTorch.\n**Features:** Combined astrophysical features, OOF predictions, and ensemble statistics to create a comprehensive feature set for sigma prediction.\n*Total params: 807,195*\n## Post-Processing\nApplied **Savitzky-Golay** smoothing to AIRS predictions to reduce noise.\nBlended raw and smoothed AIRS predictions using a weighted average (0.65 raw, 0.35 smoothed).\n# What Did Not Work\n## Multitask Model\nOne of the approaches I explored was building a multitask model that simultaneously predicted both the transit depth (mu) and the uncertainty (sigma) in a single architecture. The idea was to share the feature extraction layers between the two tasks, leveraging the shared information to improve both predictions. However, this approach did not yield better results compared to the separate models for mu (ResNet) and sigma (SigmaNet). The multitask model struggled to balance the competing objectives of minimizing the Gaussian Log-Likelihood (GLL) loss for mu and accurately estimating sigma. This imbalance often led to suboptimal predictions for both tasks.\n## SpectralTransformer\nOne of the approaches I attempted was using a SpectralTransformer architecture to model the spectral data. The idea was to leverage the transformer's ability to capture long-range dependencies across wavelengths. However, this approach did not yield competitive results. The primary issue was that I fed raw features directly into the SpectralTransformer which make the model training process does not go well\n\n\n"
  }
}