{
  "id": 609210,
  "title": "7th Place Solution",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609210",
  "author_name": "Horikita Saku",
  "post_date": "2025-09-25T02:01:27.087000",
  "votes": 28,
  "comment_count": 6,
  "views": 0,
  "content": "<h1>Introduction</h1>\n<p>We ( <a href=\"https://www.kaggle.com/horikitasaku\" target=\"_blank\">@horikitasaku</a> and <a href=\"https://www.kaggle.com/takaito\" target=\"_blank\">@takaito</a>) would like to express our sincere gratitude to the organizers for such an outstanding competition.<br>\nSpecial thanks also for staff members like Sohier Dane <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> for all their hard work and support.</p>\n<p>This was truly a well-designed competition, and it clearly reflected the professionalism and scientific rigor of the organizing team (e.g., <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> and colleagues). Physicists have always been the group of people I respect the most.</p>\n<p><strong>Our approach can be summarized in the following stages:</strong></p>\n<ul>\n<li><strong>Step0:</strong> Preprocessing</li>\n<li><strong>Step1:</strong> Signal, transit, and feature extraction; calculation of the base <code>wl_preds</code> (Horikita)</li>\n<li><strong>Step2:</strong> Difference correction with NN and base sigma prediction (takaito)</li>\n<li><strong>Step3:</strong> Sigma scale adjustment using a gradient boosting model (Horikita)</li>\n<li><strong>Step4:</strong> Pseudo Labeling (takaito)</li>\n</ul>\n<h1>Step0: Preprocessing</h1>\n<p>Here we basically followed <a href=\"https://www.kaggle.com/code/ilu000/ariel25-quick-data-prep-improved\" target=\"_blank\">Pascal’s notebook</a> almost as it is, and processed both <code>obs0</code> and <code>obs1</code>.</p>\n<h1>Step1: Signal, Transit, Features – Calculation of the base <code>wl_preds</code> (CV: ~0.42, LB:-0.45)</h1>\n<h2>Step1.1: Phase Detector</h2>\n<p>The core idea of transit detection is <strong>using the extrema of the signal’s gradient</strong>.<br>\nI assumed that the transit boundaries would appear as characteristic features on the derivative curve. The overall process was as follows:</p>\n<ul>\n<li><p><strong>Smoothing</strong>: Smooth the signal and its gradient to remove high-frequency noise.</p></li>\n<li><p><strong>Outlier removal</strong>: Clip gradient values beyond 5σ, set the edges to zero, and then apply additional Gaussian smoothing.</p></li>\n<li><p><strong>Threshold scanning</strong>:</p>\n<ul>\n<li>Negative threshold (neg_thr) → <em>t1, t2</em> (ingress)</li>\n<li>Positive threshold (pos_thr) → <em>t3, t4</em> (egress)</li></ul></li>\n<li><p><strong>Local sampling</strong>: Sample around (t1, t2) and (t3, t4), and compare the baseline with the transit segment. If the ingress/egress depth is too shallow, shift the boundary toward the start or end of the sequence.</p></li>\n<li><p><strong>Post-processing</strong>: Fix abnormal samples. In the initial version this was not implemented, but after the dataset update I observed bugs from <code>Phase Detector</code> . Upon analysis, I found many abnormal signals, so this step was added.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F2b03340d9a21e83b7f313e3aaa710b6b%2F438737d84bb3b050faf0ae4f1adb8fd6.png?generation=1758761787946320&amp;alt=media\" alt=\"p1\"></p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F3d63b1fc0686b22ebb403229af6eb2d5%2F6d1506b7-957f-4bc2-8810-1992f705f8a8.png?generation=1758761822135400&amp;alt=media\" alt=\"\"></p>\n<h2>Step1.2: Physics-based modeling and feature extraction</h2>\n<p>Here I referenced <strong>two top solutions from Ariel 2024</strong>:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/xomox-10th-place-solution\" target=\"_blank\">xomox (10th place)</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/c-number-daiwakun-1st-place-solution\" target=\"_blank\">cnumber (1st place)</a></li>\n</ul>\n<p>Unfortunately, this part was almost useless for sigma prediction.<br>\nIn my solo experiments, using these sigmas often reduced the final score by about <strong>0.1 below the theoretical upper bound score</strong>. </p>\n<p>So the real value here was in generating the <strong>base <code>wl_preds</code></strong> and <strong>features derived from the physical process</strong>. We adopted two main patterns:</p>\n<ul>\n<li><p><strong>Pattern 1 (main)</strong>: Based on <em>xomox</em>’s pipeline, with several modifications:</p>\n<ul>\n<li>Restricted the polynomial degree of the baseline to 2–3 (higher orders caused instability in some signals).Reducing the order of the baseline polynomial resulted in a roughly 0.01 improvement for me.</li>\n<li>Adjusted binning and sampling settings.</li>\n<li>Introduced a more complex Gaussian Process (GP). In our case, the kernel design combined multiple components — RBF terms with different length scales, a periodic kernel to capture repeating structures, a Matérn kernel for local smoothness, a linear kernel tied to wavelength features, and a white noise kernel for robustness. What turned out to be very interesting is that although the GP itself was not directly effective for improving sigma prediction, it played an unexpectedly strong role in stabilizing all the downstream tasks, including the linear correction I tested independently and the NN model developed by takaito. <br>\nAs a result, we obtained stable <code>wl_preds</code>, reconstructed signals, ideal transit models, and physics-derived features.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2Faa1b5c478d7feedae89b9eab3929bd7f%2Fd718ac9a-1f1f-44da-90c4-7a78c229a52b.png?generation=1758762273752078&amp;alt=media\" alt=\"\"></li></ul></li>\n<li><p><strong>Pattern 2 (secondary)</strong>: Implemented <em>cnumber</em>’s approach with only minor changes. This provided an independent set of <code>wl_preds</code> and additional features.</p></li>\n</ul>\n<h1>Step2: Difference correction with NN and base sigma prediction (CV: ~0.56x, LB: ~0.55x)</h1>\n<p>Inspired by methods that worked well in ADC2024, instead of directly predicting <code>wl</code>, we trained deep learning models to predict the difference <code>wl_diff</code> between the roughly computed <code>wl_preds</code> from <strong>step1</strong> and the ground truth <code>wl_true</code>. In other words, the NN performs a <strong>residual correction</strong>.</p>\n<h2>Basic information</h2>\n<h3>Model inputs/outputs</h3>\n<ul>\n<li><code>wl_preds</code> from Step1 are used as the base predictions.</li>\n<li>The target variable is <code>wl_diff (= wl_true - wl_preds)</code>.</li>\n<li>Input features include signal, transit, mask, position, and various extracted features.</li>\n</ul>\n<h3>Model architectures</h3>\n<ul>\n<li>Multilayer Perceptron (MLP)</li>\n<li>BiGRU + CNN + MLP</li>\n<li>BiLSTM + CNN + MLP</li>\n<li>BiLSTM + CNN + BiLSTM</li>\n</ul>\n<h2>Key techniques</h2>\n<ul>\n<li>Used <strong>quantile regression</strong> so the model can also predict sigma.</li>\n<li>Applied <strong>Adversarial Weight Perturbation (AWP)</strong>.</li>\n<li>Performed data augmentation by flipping signal and transit sequences along the time axis.</li>\n<li>Added noise to the data for further augmentation.</li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>More complex model architectures did not bring improvements.</li>\n</ul>\n<h1>Step3: Sigma scale adjustment with Gradient Boosting (LB +~0.01)</h1>\n<p>Intuitively, Step2 already provided per-wavelength σ (<code>step2_sigma</code>), but there were still systematic shifts remaining at the sample level.<br>\nTo handle this, we summarized the shift into a single scalar <strong>scale factor <code>a</code></strong>, learned separately for FGS (wl0) and AIRS (all other bands). The idea was simple:</p>\n<ul>\n<li>Correct only the overall scale</li>\n<li>Keep the relative shape across wavelengths unchanged</li>\n</ul>\n<p>Concretely:</p>\n<ol>\n<li><strong>Boundary search</strong>: Used <code>minimize_scalar(method='bounded')</code> to find the optimal <code>a</code>.</li>\n<li><strong>Discretization</strong>: Rounded <code>a</code> to 0.01 increments to prevent overfitting to random noise in the evaluation metric.</li>\n</ol>\n<p>Finally, we trained a <strong>Gradient Boosting model</strong> to learn this <code>a</code> and applied it for scaling, which stabilized sigma predictions.</p>\n<pre><code>sigma_scaled[:, ] = (\n    np.array(mean_a_preds_fgs)\n    * step2_sigma[:, ]\n).clip(CFG.MIN_SIGMA)\n\n\nsigma_scaled[:, :] = (\n    (a_preds_airs).reshape(-, )\n    * step2_sigma[:, :]\n).clip(CFG.MIN_SIGMA)\n</code></pre>\n<h1>Step4: Pseudo Labeling (LB: ~+0.004)</h1>\n<p>To boost the leaderboard score, we retrained the model using the test data along with their predicted values.<br>\nTo avoid overfitting, we only used the predictions for <code>wl_diff</code> and did <strong>not</strong> use the predicted sigma values.</p>\n<p>(Step2 involved checking a large number of learning curves, and we observed that even when MSE overfit the training data, the validation performance rarely deteriorated. This gave us confidence to apply pseudo labeling here.)</p>\n<p>Below are the trends we observed:</p>\n<ul>\n<li><strong>MSE transition</strong>: Almost no folds showed clear overfitting.</li>\n<li><strong>Competition metric transition</strong>: Tended to overfit more easily.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F530edb9c4e96db81340ba3be9c937a2e%2F584f47bf-182c-43d8-8cb6-e3f223a3f6a8.png?generation=1758762783767691&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F4391760215c0f2ee849c24bdde20ad0c%2Fffd8de63-7407-4fba-bcee-86691107dbec.png?generation=1758762800260659&amp;alt=media\" alt=\"\"></li>\n</ul>",
  "messages": [
    {
      "id": 3293907,
      "postDate": "2025-09-25T02:01:27.087Z",
      "content": "<h1>Introduction</h1>\n<p>We ( <a href=\"https://www.kaggle.com/horikitasaku\" target=\"_blank\">@horikitasaku</a> and <a href=\"https://www.kaggle.com/takaito\" target=\"_blank\">@takaito</a>) would like to express our sincere gratitude to the organizers for such an outstanding competition.<br>\nSpecial thanks also for staff members like Sohier Dane <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> for all their hard work and support.</p>\n<p>This was truly a well-designed competition, and it clearly reflected the professionalism and scientific rigor of the organizing team (e.g., <a href=\"https://www.kaggle.com/gordonyip\" target=\"_blank\">@gordonyip</a> and colleagues). Physicists have always been the group of people I respect the most.</p>\n<p><strong>Our approach can be summarized in the following stages:</strong></p>\n<ul>\n<li><strong>Step0:</strong> Preprocessing</li>\n<li><strong>Step1:</strong> Signal, transit, and feature extraction; calculation of the base <code>wl_preds</code> (Horikita)</li>\n<li><strong>Step2:</strong> Difference correction with NN and base sigma prediction (takaito)</li>\n<li><strong>Step3:</strong> Sigma scale adjustment using a gradient boosting model (Horikita)</li>\n<li><strong>Step4:</strong> Pseudo Labeling (takaito)</li>\n</ul>\n<h1>Step0: Preprocessing</h1>\n<p>Here we basically followed <a href=\"https://www.kaggle.com/code/ilu000/ariel25-quick-data-prep-improved\" target=\"_blank\">Pascal’s notebook</a> almost as it is, and processed both <code>obs0</code> and <code>obs1</code>.</p>\n<h1>Step1: Signal, Transit, Features – Calculation of the base <code>wl_preds</code> (CV: ~0.42, LB:-0.45)</h1>\n<h2>Step1.1: Phase Detector</h2>\n<p>The core idea of transit detection is <strong>using the extrema of the signal’s gradient</strong>.<br>\nI assumed that the transit boundaries would appear as characteristic features on the derivative curve. The overall process was as follows:</p>\n<ul>\n<li><p><strong>Smoothing</strong>: Smooth the signal and its gradient to remove high-frequency noise.</p></li>\n<li><p><strong>Outlier removal</strong>: Clip gradient values beyond 5σ, set the edges to zero, and then apply additional Gaussian smoothing.</p></li>\n<li><p><strong>Threshold scanning</strong>:</p>\n<ul>\n<li>Negative threshold (neg_thr) → <em>t1, t2</em> (ingress)</li>\n<li>Positive threshold (pos_thr) → <em>t3, t4</em> (egress)</li></ul></li>\n<li><p><strong>Local sampling</strong>: Sample around (t1, t2) and (t3, t4), and compare the baseline with the transit segment. If the ingress/egress depth is too shallow, shift the boundary toward the start or end of the sequence.</p></li>\n<li><p><strong>Post-processing</strong>: Fix abnormal samples. In the initial version this was not implemented, but after the dataset update I observed bugs from <code>Phase Detector</code> . Upon analysis, I found many abnormal signals, so this step was added.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F2b03340d9a21e83b7f313e3aaa710b6b%2F438737d84bb3b050faf0ae4f1adb8fd6.png?generation=1758761787946320&amp;alt=media\" alt=\"p1\"></p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F3d63b1fc0686b22ebb403229af6eb2d5%2F6d1506b7-957f-4bc2-8810-1992f705f8a8.png?generation=1758761822135400&amp;alt=media\" alt=\"\"></p>\n<h2>Step1.2: Physics-based modeling and feature extraction</h2>\n<p>Here I referenced <strong>two top solutions from Ariel 2024</strong>:</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/xomox-10th-place-solution\" target=\"_blank\">xomox (10th place)</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/c-number-daiwakun-1st-place-solution\" target=\"_blank\">cnumber (1st place)</a></li>\n</ul>\n<p>Unfortunately, this part was almost useless for sigma prediction.<br>\nIn my solo experiments, using these sigmas often reduced the final score by about <strong>0.1 below the theoretical upper bound score</strong>. </p>\n<p>So the real value here was in generating the <strong>base <code>wl_preds</code></strong> and <strong>features derived from the physical process</strong>. We adopted two main patterns:</p>\n<ul>\n<li><p><strong>Pattern 1 (main)</strong>: Based on <em>xomox</em>’s pipeline, with several modifications:</p>\n<ul>\n<li>Restricted the polynomial degree of the baseline to 2–3 (higher orders caused instability in some signals).Reducing the order of the baseline polynomial resulted in a roughly 0.01 improvement for me.</li>\n<li>Adjusted binning and sampling settings.</li>\n<li>Introduced a more complex Gaussian Process (GP). In our case, the kernel design combined multiple components — RBF terms with different length scales, a periodic kernel to capture repeating structures, a Matérn kernel for local smoothness, a linear kernel tied to wavelength features, and a white noise kernel for robustness. What turned out to be very interesting is that although the GP itself was not directly effective for improving sigma prediction, it played an unexpectedly strong role in stabilizing all the downstream tasks, including the linear correction I tested independently and the NN model developed by takaito. <br>\nAs a result, we obtained stable <code>wl_preds</code>, reconstructed signals, ideal transit models, and physics-derived features.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2Faa1b5c478d7feedae89b9eab3929bd7f%2Fd718ac9a-1f1f-44da-90c4-7a78c229a52b.png?generation=1758762273752078&amp;alt=media\" alt=\"\"></li></ul></li>\n<li><p><strong>Pattern 2 (secondary)</strong>: Implemented <em>cnumber</em>’s approach with only minor changes. This provided an independent set of <code>wl_preds</code> and additional features.</p></li>\n</ul>\n<h1>Step2: Difference correction with NN and base sigma prediction (CV: ~0.56x, LB: ~0.55x)</h1>\n<p>Inspired by methods that worked well in ADC2024, instead of directly predicting <code>wl</code>, we trained deep learning models to predict the difference <code>wl_diff</code> between the roughly computed <code>wl_preds</code> from <strong>step1</strong> and the ground truth <code>wl_true</code>. In other words, the NN performs a <strong>residual correction</strong>.</p>\n<h2>Basic information</h2>\n<h3>Model inputs/outputs</h3>\n<ul>\n<li><code>wl_preds</code> from Step1 are used as the base predictions.</li>\n<li>The target variable is <code>wl_diff (= wl_true - wl_preds)</code>.</li>\n<li>Input features include signal, transit, mask, position, and various extracted features.</li>\n</ul>\n<h3>Model architectures</h3>\n<ul>\n<li>Multilayer Perceptron (MLP)</li>\n<li>BiGRU + CNN + MLP</li>\n<li>BiLSTM + CNN + MLP</li>\n<li>BiLSTM + CNN + BiLSTM</li>\n</ul>\n<h2>Key techniques</h2>\n<ul>\n<li>Used <strong>quantile regression</strong> so the model can also predict sigma.</li>\n<li>Applied <strong>Adversarial Weight Perturbation (AWP)</strong>.</li>\n<li>Performed data augmentation by flipping signal and transit sequences along the time axis.</li>\n<li>Added noise to the data for further augmentation.</li>\n</ul>\n<h2>What did not work</h2>\n<ul>\n<li>More complex model architectures did not bring improvements.</li>\n</ul>\n<h1>Step3: Sigma scale adjustment with Gradient Boosting (LB +~0.01)</h1>\n<p>Intuitively, Step2 already provided per-wavelength σ (<code>step2_sigma</code>), but there were still systematic shifts remaining at the sample level.<br>\nTo handle this, we summarized the shift into a single scalar <strong>scale factor <code>a</code></strong>, learned separately for FGS (wl0) and AIRS (all other bands). The idea was simple:</p>\n<ul>\n<li>Correct only the overall scale</li>\n<li>Keep the relative shape across wavelengths unchanged</li>\n</ul>\n<p>Concretely:</p>\n<ol>\n<li><strong>Boundary search</strong>: Used <code>minimize_scalar(method='bounded')</code> to find the optimal <code>a</code>.</li>\n<li><strong>Discretization</strong>: Rounded <code>a</code> to 0.01 increments to prevent overfitting to random noise in the evaluation metric.</li>\n</ol>\n<p>Finally, we trained a <strong>Gradient Boosting model</strong> to learn this <code>a</code> and applied it for scaling, which stabilized sigma predictions.</p>\n<pre><code>sigma_scaled[:, ] = (\n    np.array(mean_a_preds_fgs)\n    * step2_sigma[:, ]\n).clip(CFG.MIN_SIGMA)\n\n\nsigma_scaled[:, :] = (\n    (a_preds_airs).reshape(-, )\n    * step2_sigma[:, :]\n).clip(CFG.MIN_SIGMA)\n</code></pre>\n<h1>Step4: Pseudo Labeling (LB: ~+0.004)</h1>\n<p>To boost the leaderboard score, we retrained the model using the test data along with their predicted values.<br>\nTo avoid overfitting, we only used the predictions for <code>wl_diff</code> and did <strong>not</strong> use the predicted sigma values.</p>\n<p>(Step2 involved checking a large number of learning curves, and we observed that even when MSE overfit the training data, the validation performance rarely deteriorated. This gave us confidence to apply pseudo labeling here.)</p>\n<p>Below are the trends we observed:</p>\n<ul>\n<li><strong>MSE transition</strong>: Almost no folds showed clear overfitting.</li>\n<li><strong>Competition metric transition</strong>: Tended to overfit more easily.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F530edb9c4e96db81340ba3be9c937a2e%2F584f47bf-182c-43d8-8cb6-e3f223a3f6a8.png?generation=1758762783767691&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F4391760215c0f2ee849c24bdde20ad0c%2Fffd8de63-7407-4fba-bcee-86691107dbec.png?generation=1758762800260659&amp;alt=media\" alt=\"\"></li>\n</ul>",
      "rawMarkdown": "# Introduction\nWe ( @horikitasaku and @takaito) would like to express our sincere gratitude to the organizers for such an outstanding competition.\nSpecial thanks also for staff members like Sohier Dane @sohier for all their hard work and support.\n\nThis was truly a well-designed competition, and it clearly reflected the professionalism and scientific rigor of the organizing team (e.g., @gordonyip and colleagues). Physicists have always been the group of people I respect the most.\n\n\n**Our approach can be summarized in the following stages:**\n\n* **Step0:** Preprocessing\n* **Step1:** Signal, transit, and feature extraction; calculation of the base `wl_preds` (Horikita)\n* **Step2:** Difference correction with NN and base sigma prediction (takaito)\n* **Step3:** Sigma scale adjustment using a gradient boosting model (Horikita)\n* **Step4:** Pseudo Labeling (takaito)\n\n# Step0: Preprocessing\n\nHere we basically followed [Pascal’s notebook](https://www.kaggle.com/code/ilu000/ariel25-quick-data-prep-improved) almost as it is, and processed both `obs0` and `obs1`.\n\n# Step1: Signal, Transit, Features – Calculation of the base `wl_preds` (CV: \\~0.42, LB:-0.45)\n\n## Step1.1: Phase Detector\n\nThe core idea of transit detection is **using the extrema of the signal’s gradient**.\nI assumed that the transit boundaries would appear as characteristic features on the derivative curve. The overall process was as follows:\n\n* **Smoothing**: Smooth the signal and its gradient to remove high-frequency noise.\n* **Outlier removal**: Clip gradient values beyond 5σ, set the edges to zero, and then apply additional Gaussian smoothing.\n* **Threshold scanning**:\n\n  * Negative threshold (neg\\_thr) → *t1, t2* (ingress)\n  * Positive threshold (pos\\_thr) → *t3, t4* (egress)\n* **Local sampling**: Sample around (t1, t2) and (t3, t4), and compare the baseline with the transit segment. If the ingress/egress depth is too shallow, shift the boundary toward the start or end of the sequence.\n* **Post-processing**: Fix abnormal samples. In the initial version this was not implemented, but after the dataset update I observed bugs from `Phase Detector` . Upon analysis, I found many abnormal signals, so this step was added.\n![p1](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F2b03340d9a21e83b7f313e3aaa710b6b%2F438737d84bb3b050faf0ae4f1adb8fd6.png?generation=1758761787946320&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F3d63b1fc0686b22ebb403229af6eb2d5%2F6d1506b7-957f-4bc2-8810-1992f705f8a8.png?generation=1758761822135400&alt=media)\n\n## Step1.2: Physics-based modeling and feature extraction \n\nHere I referenced **two top solutions from Ariel 2024**:\n\n* [xomox (10th place)](https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/xomox-10th-place-solution)\n* [cnumber (1st place)](https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/c-number-daiwakun-1st-place-solution)\n\nUnfortunately, this part was almost useless for sigma prediction.\nIn my solo experiments, using these sigmas often reduced the final score by about **0.1 below the theoretical upper bound score**. \n\nSo the real value here was in generating the **base `wl_preds`** and **features derived from the physical process**. We adopted two main patterns:\n\n* **Pattern 1 (main)**: Based on *xomox*’s pipeline, with several modifications:\n\n  * Restricted the polynomial degree of the baseline to 2–3 (higher orders caused instability in some signals).Reducing the order of the baseline polynomial resulted in a roughly 0.01 improvement for me.\n  * Adjusted binning and sampling settings.\n  * Introduced a more complex Gaussian Process (GP). In our case, the kernel design combined multiple components — RBF terms with different length scales, a periodic kernel to capture repeating structures, a Matérn kernel for local smoothness, a linear kernel tied to wavelength features, and a white noise kernel for robustness. What turned out to be very interesting is that although the GP itself was not directly effective for improving sigma prediction, it played an unexpectedly strong role in stabilizing all the downstream tasks, including the linear correction I tested independently and the NN model developed by takaito. \n  As a result, we obtained stable `wl_preds`, reconstructed signals, ideal transit models, and physics-derived features.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2Faa1b5c478d7feedae89b9eab3929bd7f%2Fd718ac9a-1f1f-44da-90c4-7a78c229a52b.png?generation=1758762273752078&alt=media)\n\n* **Pattern 2 (secondary)**: Implemented *cnumber*’s approach with only minor changes. This provided an independent set of `wl_preds` and additional features.\n\n# Step2: Difference correction with NN and base sigma prediction (CV: \\~0.56x, LB: \\~0.55x)\n\nInspired by methods that worked well in ADC2024, instead of directly predicting `wl`, we trained deep learning models to predict the difference `wl_diff` between the roughly computed `wl_preds` from **step1** and the ground truth `wl_true`. In other words, the NN performs a **residual correction**.\n\n## Basic information\n\n### Model inputs/outputs\n\n* `wl_preds` from Step1 are used as the base predictions.\n* The target variable is `wl_diff (= wl_true - wl_preds)`.\n* Input features include signal, transit, mask, position, and various extracted features.\n\n### Model architectures\n\n* Multilayer Perceptron (MLP)\n* BiGRU + CNN + MLP\n* BiLSTM + CNN + MLP\n* BiLSTM + CNN + BiLSTM\n\n## Key techniques\n\n* Used **quantile regression** so the model can also predict sigma.\n* Applied **Adversarial Weight Perturbation (AWP)**.\n* Performed data augmentation by flipping signal and transit sequences along the time axis.\n* Added noise to the data for further augmentation.\n\n## What did not work\n\n* More complex model architectures did not bring improvements.\n\n# Step3: Sigma scale adjustment with Gradient Boosting (LB +\\~0.01)\n\nIntuitively, Step2 already provided per-wavelength σ (`step2_sigma`), but there were still systematic shifts remaining at the sample level.\nTo handle this, we summarized the shift into a single scalar **scale factor `a`**, learned separately for FGS (wl0) and AIRS (all other bands). The idea was simple:\n\n* Correct only the overall scale\n* Keep the relative shape across wavelengths unchanged\n\nConcretely:\n\n1. **Boundary search**: Used `minimize_scalar(method='bounded')` to find the optimal `a`.\n2. **Discretization**: Rounded `a` to 0.01 increments to prevent overfitting to random noise in the evaluation metric.\n\nFinally, we trained a **Gradient Boosting model** to learn this `a` and applied it for scaling, which stabilized sigma predictions.\n\n```python\nsigma_scaled[:, 0] = (\n    np.array(mean_a_preds_fgs)\n    * step2_sigma[:, 0]\n).clip(CFG.MIN_SIGMA)\n\n# AIRS\nsigma_scaled[:, 1:] = (\n    (a_preds_airs).reshape(-1, 1)\n    * step2_sigma[:, 1:]\n).clip(CFG.MIN_SIGMA)\n```\n\n# Step4: Pseudo Labeling (LB: \\~+0.004)\n\nTo boost the leaderboard score, we retrained the model using the test data along with their predicted values.\nTo avoid overfitting, we only used the predictions for `wl_diff` and did **not** use the predicted sigma values.\n\n(Step2 involved checking a large number of learning curves, and we observed that even when MSE overfit the training data, the validation performance rarely deteriorated. This gave us confidence to apply pseudo labeling here.)\n\nBelow are the trends we observed:\n\n* **MSE transition**: Almost no folds showed clear overfitting.\n* **Competition metric transition**: Tended to overfit more easily.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F530edb9c4e96db81340ba3be9c937a2e%2F584f47bf-182c-43d8-8cb6-e3f223a3f6a8.png?generation=1758762783767691&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F4391760215c0f2ee849c24bdde20ad0c%2Fffd8de63-7407-4fba-bcee-86691107dbec.png?generation=1758762800260659&alt=media)",
      "votes": 28
    },
    {
      "id": 3294155,
      "postDate": "2025-09-25T13:32:07.230Z",
      "content": "<p>Congratulations on achieving 7th place in the Ariel Challenge 2025! As a university student participating in my first NeurIPS competition, I’m inspired by your solution. I have a few questions about your approach:</p>\n<ol>\n<li><p>How did you come up with the idea to extract multiple features for the neural network? What inspired your feature engineering strategy? I saw some solutions only extract the based transit-depth and use that to tune the prediction so quite curious on that.</p></li>\n<li><p>Can you discuss more in details how you apply Gaussian Process in the Step 1?</p></li>\n<li><p>Did you incorporate star_info (e.g., stellar radius, mass, temperature, or planetary parameters) into your featuring engineering?</p></li>\n<li><p>How did you combine the extracted features with the time-series data ? Did you train models on these data separately and use an ensemble, or integrate them in a single model?</p></li>\n</ol>",
      "rawMarkdown": "Congratulations on achieving 7th place in the Ariel Challenge 2025! As a university student participating in my first NeurIPS competition, I’m inspired by your solution. I have a few questions about your approach:\n\n1. How did you come up with the idea to extract multiple features for the neural network? What inspired your feature engineering strategy? I saw some solutions only extract the based transit-depth and use that to tune the prediction so quite curious on that.\n\n2. Can you discuss more in details how you apply Gaussian Process in the Step 1?\n\n3. Did you incorporate star_info (e.g., stellar radius, mass, temperature, or planetary parameters) into your featuring engineering?\n\n4. How did you combine the extracted features with the time-series data ? Did you train models on these data separately and use an ensemble, or integrate them in a single model?",
      "replies": [
        {
          "id": 3294185,
          "postDate": "2025-09-25T14:52:40.650Z",
          "content": "<ol>\n<li>Extracting features was actually a rather straightforward idea, since we wanted to provide the neural network with information derived from the base <code>wl_pred</code>. However, using the raw predictions directly as inputs would be extremely high-dimensional and inefficient. The “based transit depth” is indeed one part of it, but I expanded this considerably by designing a much richer feature set.</li>\n<li>The inspiration for using GP came from cnumber’s 1st-place solution in the ADC2024. I extended their approach by introducing a more complex kernel design. </li>\n</ol>\n<pre><code>dip_ave = np.mean(dip)\n\ndata = signal_list[idx].copy()\nx = np.arange()\ny = dip.copy() - dip_ave\ns = err * \nX = x.reshape(-, )\n\nkernel1 = C(y.() - y.(), (, )) * RBF(, (, ))\nkernel2 = C(y.() - y.(), (, )) * Matern(\n    length_scale=, length_scale_bounds=(, ), nu=\n)\nkernel = kernel1 + kernel2\n\ngp = GaussianProcessRegressor(kernel=kernel, alpha=s**, n_restarts_optimizer=)\ngp.fit(X, y)\n\nx_pred = np.arange().reshape(-, )\ny_pred, y_std = gp.predict(x_pred, return_std=)\ny_pred = y_pred[::-] + dip_ave\n</code></pre>\n<ol>\n<li><p>Yes, we did some simple feature engineering using star_info. Among them, the most effective turned out to be features centered on planetary mass (Mp), including log-transformed versions.</p></li>\n<li><p>We did not train on them separately. Instead, we appended the extracted features and planetary parameters to every timestep of the binned time-series data (375 steps after binning). In other words, the statistical or planetary features were duplicated across the temporal dimension, so that the final input was <code>[time_len × (series_dim + feature_dim)]</code>. This way, the network could consume both sequential dynamics and global descriptors jointly in a single model.</p></li>\n</ol>",
          "rawMarkdown": "1. Extracting features was actually a rather straightforward idea, since we wanted to provide the neural network with information derived from the base `wl_pred`. However, using the raw predictions directly as inputs would be extremely high-dimensional and inefficient. The “based transit depth” is indeed one part of it, but I expanded this considerably by designing a much richer feature set.\n2. The inspiration for using GP came from cnumber’s 1st-place solution in the ADC2024. I extended their approach by introducing a more complex kernel design. \n```python\ndip_ave = np.mean(dip)\n\ndata = signal_list[idx].copy()\nx = np.arange(282)\ny = dip.copy() - dip_ave\ns = err * 1.6\nX = x.reshape(-1, 1)\n\nkernel1 = C(y.max() - y.min(), (1e-9, 1e3)) * RBF(10, (1, 1e5))\nkernel2 = C(y.max() - y.min(), (1e-9, 1e3)) * Matern(\n    length_scale=10, length_scale_bounds=(1, 1e5), nu=1.5\n)\nkernel = kernel1 + kernel2\n\ngp = GaussianProcessRegressor(kernel=kernel, alpha=s**2, n_restarts_optimizer=10)\ngp.fit(X, y)\n\nx_pred = np.arange(283).reshape(-1, 1)\ny_pred, y_std = gp.predict(x_pred, return_std=True)\ny_pred = y_pred[::-1] + dip_ave\n```\n3. Yes, we did some simple feature engineering using star_info. Among them, the most effective turned out to be features centered on planetary mass (Mp), including log-transformed versions.\n\n4. We did not train on them separately. Instead, we appended the extracted features and planetary parameters to every timestep of the binned time-series data (375 steps after binning). In other words, the statistical or planetary features were duplicated across the temporal dimension, so that the final input was `[time_len × (series_dim + feature_dim)]`. This way, the network could consume both sequential dynamics and global descriptors jointly in a single model."
        }
      ]
    },
    {
      "id": 3293975,
      "postDate": "2025-09-25T05:44:44.590Z",
      "content": "<p>Congrats in your gold medal, specially to your first gold medal <a href=\"https://www.kaggle.com/horikitasaku\" target=\"_blank\">@horikitasaku</a>! <br>\nI am curious about 2 points:</p>\n<ul>\n<li>What type of features have you used to detect/correct limb darkening? Your GRU model works along the wavelength dimension as seq_len or along the time dimension?</li>\n<li>What did you do to solve planets with no out of transit? </li>\n</ul>",
      "rawMarkdown": "Congrats in your gold medal, specially to your first gold medal @horikitasaku! \nI am curious about 2 points:\n- What type of features have you used to detect/correct limb darkening? Your GRU model works along the wavelength dimension as seq_len or along the time dimension?\n- What did you do to solve planets with no out of transit? \n",
      "replies": [
        {
          "id": 3294003,
          "postDate": "2025-09-25T07:17:23.640Z",
          "content": "<ol>\n<li><p>In fact, I didn’t handle these cases which has no out of transit directly. My approach is: when the phase detector flags this type of sample, I set the time points t1, t2, t3, t4 to the two sides (edges). The idea is that, for such samples, connecting the two endpoints on both sides to form a straight line is the most reasonable baseline. (There will still be loss, but I think this gets as close as we can.)At the same time, when fitting the baseline, use lower-degree polynomials (as I have written, such as a second-degree polynomial).<br>\nfor eg. Index 496<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F79740c40a3a9a630b292bda49b7c6420%2F3d4718a8-9655-4ca9-ac8b-b30b25c8881b.png?generation=1758784446169730&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F9131ea5c4ef0681d843b7131cbe8216f%2F0ec39ab5-8149-4990-9241-cbf028670bc3.png?generation=1758784468292314&amp;alt=media\" alt=\"\"></p></li>\n<li><p>The features are largely similar to those used by xomox (10th place) and others, but I extended them further.<br>\nMy features include dividing the full-band depth sequence all_d into bins at multiple resolutions (such as 283/100/50 and 200/150) and extracting per-segment statistics including mean, median, extrema, standard deviation, and quantiles, while also adding IQR, extended IQR, center of mass weighted by |d|, and gradient-based measures (mean, standard deviation, RMSE); at the same time, in the out-of-transit region (before t2 and after t3) I fit the normalized curve with a polynomial and record RMSE, MAE, MAPE, R², residual mean/standard deviation/range/extrema, and the cosine similarity between prediction and truth; I further include the first-order autocorrelation of the residuals as well as the mean and standard deviation of the first- and second-order differences of all_d (slope and curvature) to capture correlation structure and overall shape; finally, I perform an rFFT on the mean-subtracted all_d and take the leading frequency magnitudes as coarse spectral features. These features are unexpectedly very powerful.</p></li>\n</ol>",
          "rawMarkdown": "1. In fact, I didn’t handle these cases which has no out of transit directly. My approach is: when the phase detector flags this type of sample, I set the time points t1, t2, t3, t4 to the two sides (edges). The idea is that, for such samples, connecting the two endpoints on both sides to form a straight line is the most reasonable baseline. (There will still be loss, but I think this gets as close as we can.)At the same time, when fitting the baseline, use lower-degree polynomials (as I have written, such as a second-degree polynomial).\nfor eg. Index 496\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F79740c40a3a9a630b292bda49b7c6420%2F3d4718a8-9655-4ca9-ac8b-b30b25c8881b.png?generation=1758784446169730&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F9131ea5c4ef0681d843b7131cbe8216f%2F0ec39ab5-8149-4990-9241-cbf028670bc3.png?generation=1758784468292314&alt=media)\n\n2. The features are largely similar to those used by xomox (10th place) and others, but I extended them further.\nMy features include dividing the full-band depth sequence all_d into bins at multiple resolutions (such as 283/100/50 and 200/150) and extracting per-segment statistics including mean, median, extrema, standard deviation, and quantiles, while also adding IQR, extended IQR, center of mass weighted by |d|, and gradient-based measures (mean, standard deviation, RMSE); at the same time, in the out-of-transit region (before t2 and after t3) I fit the normalized curve with a polynomial and record RMSE, MAE, MAPE, R², residual mean/standard deviation/range/extrema, and the cosine similarity between prediction and truth; I further include the first-order autocorrelation of the residuals as well as the mean and standard deviation of the first- and second-order differences of all_d (slope and curvature) to capture correlation structure and overall shape; finally, I perform an rFFT on the mean-subtracted all_d and take the leading frequency magnitudes as coarse spectral features. These features are unexpectedly very powerful."
        },
        {
          "id": 3294141,
          "postDate": "2025-09-25T12:38:24.437Z",
          "content": "<p>Our gru model is along the time dimension. Features include ideal transit models, reconstructed baseline, and selected features)</p>",
          "rawMarkdown": "Our gru model is along the time dimension. Features include ideal transit models, reconstructed baseline, and selected features)",
          "replies": [
            {
              "id": 3294148,
              "postDate": "2025-09-25T13:09:00.943Z",
              "content": "<p>Thanks so much for the explanation. Congrats again in your gold!</p>",
              "rawMarkdown": "Thanks so much for the explanation. Congrats again in your gold!",
              "votes": 1
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3294155,
      "author_name": "Long2003",
      "author_url": "",
      "post_date": "2025-09-25T13:32:07.230000",
      "content": "<p>Congratulations on achieving 7th place in the Ariel Challenge 2025! As a university student participating in my first NeurIPS competition, I’m inspired by your solution. I have a few questions about your approach:</p>\n<ol>\n<li><p>How did you come up with the idea to extract multiple features for the neural network? What inspired your feature engineering strategy? I saw some solutions only extract the based transit-depth and use that to tune the prediction so quite curious on that.</p></li>\n<li><p>Can you discuss more in details how you apply Gaussian Process in the Step 1?</p></li>\n<li><p>Did you incorporate star_info (e.g., stellar radius, mass, temperature, or planetary parameters) into your featuring engineering?</p></li>\n<li><p>How did you combine the extracted features with the time-series data ? Did you train models on these data separately and use an ensemble, or integrate them in a single model?</p></li>\n</ol>",
      "votes": 0,
      "replies": [
        {
          "id": 3294185,
          "author_name": "Horikita Saku",
          "author_url": "",
          "post_date": "2025-09-25T14:52:40.650000",
          "content": "<ol>\n<li>Extracting features was actually a rather straightforward idea, since we wanted to provide the neural network with information derived from the base <code>wl_pred</code>. However, using the raw predictions directly as inputs would be extremely high-dimensional and inefficient. The “based transit depth” is indeed one part of it, but I expanded this considerably by designing a much richer feature set.</li>\n<li>The inspiration for using GP came from cnumber’s 1st-place solution in the ADC2024. I extended their approach by introducing a more complex kernel design. </li>\n</ol>\n<pre><code>dip_ave = np.mean(dip)\n\ndata = signal_list[idx].copy()\nx = np.arange()\ny = dip.copy() - dip_ave\ns = err * \nX = x.reshape(-, )\n\nkernel1 = C(y.() - y.(), (, )) * RBF(, (, ))\nkernel2 = C(y.() - y.(), (, )) * Matern(\n    length_scale=, length_scale_bounds=(, ), nu=\n)\nkernel = kernel1 + kernel2\n\ngp = GaussianProcessRegressor(kernel=kernel, alpha=s**, n_restarts_optimizer=)\ngp.fit(X, y)\n\nx_pred = np.arange().reshape(-, )\ny_pred, y_std = gp.predict(x_pred, return_std=)\ny_pred = y_pred[::-] + dip_ave\n</code></pre>\n<ol>\n<li><p>Yes, we did some simple feature engineering using star_info. Among them, the most effective turned out to be features centered on planetary mass (Mp), including log-transformed versions.</p></li>\n<li><p>We did not train on them separately. Instead, we appended the extracted features and planetary parameters to every timestep of the binned time-series data (375 steps after binning). In other words, the statistical or planetary features were duplicated across the temporal dimension, so that the final input was <code>[time_len × (series_dim + feature_dim)]</code>. This way, the network could consume both sequential dynamics and global descriptors jointly in a single model.</p></li>\n</ol>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3293975,
      "author_name": "Carlos Pérez Ricardo",
      "author_url": "",
      "post_date": "2025-09-25T05:44:44.590000",
      "content": "<p>Congrats in your gold medal, specially to your first gold medal <a href=\"https://www.kaggle.com/horikitasaku\" target=\"_blank\">@horikitasaku</a>! <br>\nI am curious about 2 points:</p>\n<ul>\n<li>What type of features have you used to detect/correct limb darkening? Your GRU model works along the wavelength dimension as seq_len or along the time dimension?</li>\n<li>What did you do to solve planets with no out of transit? </li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 3294003,
          "author_name": "Horikita Saku",
          "author_url": "",
          "post_date": "2025-09-25T07:17:23.640000",
          "content": "<ol>\n<li><p>In fact, I didn’t handle these cases which has no out of transit directly. My approach is: when the phase detector flags this type of sample, I set the time points t1, t2, t3, t4 to the two sides (edges). The idea is that, for such samples, connecting the two endpoints on both sides to form a straight line is the most reasonable baseline. (There will still be loss, but I think this gets as close as we can.)At the same time, when fitting the baseline, use lower-degree polynomials (as I have written, such as a second-degree polynomial).<br>\nfor eg. Index 496<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F79740c40a3a9a630b292bda49b7c6420%2F3d4718a8-9655-4ca9-ac8b-b30b25c8881b.png?generation=1758784446169730&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F9131ea5c4ef0681d843b7131cbe8216f%2F0ec39ab5-8149-4990-9241-cbf028670bc3.png?generation=1758784468292314&amp;alt=media\" alt=\"\"></p></li>\n<li><p>The features are largely similar to those used by xomox (10th place) and others, but I extended them further.<br>\nMy features include dividing the full-band depth sequence all_d into bins at multiple resolutions (such as 283/100/50 and 200/150) and extracting per-segment statistics including mean, median, extrema, standard deviation, and quantiles, while also adding IQR, extended IQR, center of mass weighted by |d|, and gradient-based measures (mean, standard deviation, RMSE); at the same time, in the out-of-transit region (before t2 and after t3) I fit the normalized curve with a polynomial and record RMSE, MAE, MAPE, R², residual mean/standard deviation/range/extrema, and the cosine similarity between prediction and truth; I further include the first-order autocorrelation of the residuals as well as the mean and standard deviation of the first- and second-order differences of all_d (slope and curvature) to capture correlation structure and overall shape; finally, I perform an rFFT on the mean-subtracted all_d and take the leading frequency magnitudes as coarse spectral features. These features are unexpectedly very powerful.</p></li>\n</ol>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3294141,
          "author_name": "Horikita Saku",
          "author_url": "",
          "post_date": "2025-09-25T12:38:24.437000",
          "content": "<p>Our gru model is along the time dimension. Features include ideal transit models, reconstructed baseline, and selected features)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3294148,
              "author_name": "Carlos Pérez Ricardo",
              "author_url": "",
              "post_date": "2025-09-25T13:09:00.943000",
              "content": "<p>Thanks so much for the explanation. Congrats again in your gold!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3293907": "# Introduction\nWe ( @horikitasaku and @takaito) would like to express our sincere gratitude to the organizers for such an outstanding competition.\nSpecial thanks also for staff members like Sohier Dane @sohier for all their hard work and support.\n\nThis was truly a well-designed competition, and it clearly reflected the professionalism and scientific rigor of the organizing team (e.g., @gordonyip and colleagues). Physicists have always been the group of people I respect the most.\n\n\n**Our approach can be summarized in the following stages:**\n\n* **Step0:** Preprocessing\n* **Step1:** Signal, transit, and feature extraction; calculation of the base `wl_preds` (Horikita)\n* **Step2:** Difference correction with NN and base sigma prediction (takaito)\n* **Step3:** Sigma scale adjustment using a gradient boosting model (Horikita)\n* **Step4:** Pseudo Labeling (takaito)\n\n# Step0: Preprocessing\n\nHere we basically followed [Pascal’s notebook](https://www.kaggle.com/code/ilu000/ariel25-quick-data-prep-improved) almost as it is, and processed both `obs0` and `obs1`.\n\n# Step1: Signal, Transit, Features – Calculation of the base `wl_preds` (CV: \\~0.42, LB:-0.45)\n\n## Step1.1: Phase Detector\n\nThe core idea of transit detection is **using the extrema of the signal’s gradient**.\nI assumed that the transit boundaries would appear as characteristic features on the derivative curve. The overall process was as follows:\n\n* **Smoothing**: Smooth the signal and its gradient to remove high-frequency noise.\n* **Outlier removal**: Clip gradient values beyond 5σ, set the edges to zero, and then apply additional Gaussian smoothing.\n* **Threshold scanning**:\n\n  * Negative threshold (neg\\_thr) → *t1, t2* (ingress)\n  * Positive threshold (pos\\_thr) → *t3, t4* (egress)\n* **Local sampling**: Sample around (t1, t2) and (t3, t4), and compare the baseline with the transit segment. If the ingress/egress depth is too shallow, shift the boundary toward the start or end of the sequence.\n* **Post-processing**: Fix abnormal samples. In the initial version this was not implemented, but after the dataset update I observed bugs from `Phase Detector` . Upon analysis, I found many abnormal signals, so this step was added.\n![p1](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F2b03340d9a21e83b7f313e3aaa710b6b%2F438737d84bb3b050faf0ae4f1adb8fd6.png?generation=1758761787946320&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F3d63b1fc0686b22ebb403229af6eb2d5%2F6d1506b7-957f-4bc2-8810-1992f705f8a8.png?generation=1758761822135400&alt=media)\n\n## Step1.2: Physics-based modeling and feature extraction \n\nHere I referenced **two top solutions from Ariel 2024**:\n\n* [xomox (10th place)](https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/xomox-10th-place-solution)\n* [cnumber (1st place)](https://www.kaggle.com/competitions/ariel-data-challenge-2024/writeups/c-number-daiwakun-1st-place-solution)\n\nUnfortunately, this part was almost useless for sigma prediction.\nIn my solo experiments, using these sigmas often reduced the final score by about **0.1 below the theoretical upper bound score**. \n\nSo the real value here was in generating the **base `wl_preds`** and **features derived from the physical process**. We adopted two main patterns:\n\n* **Pattern 1 (main)**: Based on *xomox*’s pipeline, with several modifications:\n\n  * Restricted the polynomial degree of the baseline to 2–3 (higher orders caused instability in some signals).Reducing the order of the baseline polynomial resulted in a roughly 0.01 improvement for me.\n  * Adjusted binning and sampling settings.\n  * Introduced a more complex Gaussian Process (GP). In our case, the kernel design combined multiple components — RBF terms with different length scales, a periodic kernel to capture repeating structures, a Matérn kernel for local smoothness, a linear kernel tied to wavelength features, and a white noise kernel for robustness. What turned out to be very interesting is that although the GP itself was not directly effective for improving sigma prediction, it played an unexpectedly strong role in stabilizing all the downstream tasks, including the linear correction I tested independently and the NN model developed by takaito. \n  As a result, we obtained stable `wl_preds`, reconstructed signals, ideal transit models, and physics-derived features.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2Faa1b5c478d7feedae89b9eab3929bd7f%2Fd718ac9a-1f1f-44da-90c4-7a78c229a52b.png?generation=1758762273752078&alt=media)\n\n* **Pattern 2 (secondary)**: Implemented *cnumber*’s approach with only minor changes. This provided an independent set of `wl_preds` and additional features.\n\n# Step2: Difference correction with NN and base sigma prediction (CV: \\~0.56x, LB: \\~0.55x)\n\nInspired by methods that worked well in ADC2024, instead of directly predicting `wl`, we trained deep learning models to predict the difference `wl_diff` between the roughly computed `wl_preds` from **step1** and the ground truth `wl_true`. In other words, the NN performs a **residual correction**.\n\n## Basic information\n\n### Model inputs/outputs\n\n* `wl_preds` from Step1 are used as the base predictions.\n* The target variable is `wl_diff (= wl_true - wl_preds)`.\n* Input features include signal, transit, mask, position, and various extracted features.\n\n### Model architectures\n\n* Multilayer Perceptron (MLP)\n* BiGRU + CNN + MLP\n* BiLSTM + CNN + MLP\n* BiLSTM + CNN + BiLSTM\n\n## Key techniques\n\n* Used **quantile regression** so the model can also predict sigma.\n* Applied **Adversarial Weight Perturbation (AWP)**.\n* Performed data augmentation by flipping signal and transit sequences along the time axis.\n* Added noise to the data for further augmentation.\n\n## What did not work\n\n* More complex model architectures did not bring improvements.\n\n# Step3: Sigma scale adjustment with Gradient Boosting (LB +\\~0.01)\n\nIntuitively, Step2 already provided per-wavelength σ (`step2_sigma`), but there were still systematic shifts remaining at the sample level.\nTo handle this, we summarized the shift into a single scalar **scale factor `a`**, learned separately for FGS (wl0) and AIRS (all other bands). The idea was simple:\n\n* Correct only the overall scale\n* Keep the relative shape across wavelengths unchanged\n\nConcretely:\n\n1. **Boundary search**: Used `minimize_scalar(method='bounded')` to find the optimal `a`.\n2. **Discretization**: Rounded `a` to 0.01 increments to prevent overfitting to random noise in the evaluation metric.\n\nFinally, we trained a **Gradient Boosting model** to learn this `a` and applied it for scaling, which stabilized sigma predictions.\n\n```python\nsigma_scaled[:, 0] = (\n    np.array(mean_a_preds_fgs)\n    * step2_sigma[:, 0]\n).clip(CFG.MIN_SIGMA)\n\n# AIRS\nsigma_scaled[:, 1:] = (\n    (a_preds_airs).reshape(-1, 1)\n    * step2_sigma[:, 1:]\n).clip(CFG.MIN_SIGMA)\n```\n\n# Step4: Pseudo Labeling (LB: \\~+0.004)\n\nTo boost the leaderboard score, we retrained the model using the test data along with their predicted values.\nTo avoid overfitting, we only used the predictions for `wl_diff` and did **not** use the predicted sigma values.\n\n(Step2 involved checking a large number of learning curves, and we observed that even when MSE overfit the training data, the validation performance rarely deteriorated. This gave us confidence to apply pseudo labeling here.)\n\nBelow are the trends we observed:\n\n* **MSE transition**: Almost no folds showed clear overfitting.\n* **Competition metric transition**: Tended to overfit more easily.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F530edb9c4e96db81340ba3be9c937a2e%2F584f47bf-182c-43d8-8cb6-e3f223a3f6a8.png?generation=1758762783767691&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F11676771%2F4391760215c0f2ee849c24bdde20ad0c%2Fffd8de63-7407-4fba-bcee-86691107dbec.png?generation=1758762800260659&alt=media)",
    "3294155": "Congratulations on achieving 7th place in the Ariel Challenge 2025! As a university student participating in my first NeurIPS competition, I’m inspired by your solution. I have a few questions about your approach:\n\n1. How did you come up with the idea to extract multiple features for the neural network? What inspired your feature engineering strategy? I saw some solutions only extract the based transit-depth and use that to tune the prediction so quite curious on that.\n\n2. Can you discuss more in details how you apply Gaussian Process in the Step 1?\n\n3. Did you incorporate star_info (e.g., stellar radius, mass, temperature, or planetary parameters) into your featuring engineering?\n\n4. How did you combine the extracted features with the time-series data ? Did you train models on these data separately and use an ensemble, or integrate them in a single model?",
    "3293975": "Congrats in your gold medal, specially to your first gold medal @horikitasaku! \nI am curious about 2 points:\n- What type of features have you used to detect/correct limb darkening? Your GRU model works along the wavelength dimension as seq_len or along the time dimension?\n- What did you do to solve planets with no out of transit? \n"
  }
}