{
  "id": 609574,
  "title": "29th Discovery and Solution",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609574",
  "author_name": "shanzhong8",
  "post_date": "2025-09-28T00:50:50.279000",
  "votes": 6,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I would like to express my sincere gratitude to Kaggle and University College London for hosting this competition. I learned a great deal, and it was truly an exciting and rewarding experience. I would also like to thank my teammates <a href=\"https://www.kaggle.com/ajinomoto132\" target=\"_blank\">@ajinomoto132</a> <a href=\"https://www.kaggle.com/chenzhenyuan\" target=\"_blank\">@chenzhenyuan</a> <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> <a href=\"https://www.kaggle.com/larrylin666\" target=\"_blank\">@larrylin666</a> —together, we proudly earned a silver medal in this competition.</p>\n<h1><strong>Summary</strong></h1>\n<ul>\n<li><strong>Preprocessing:</strong> Noise reduction</li>\n<li><strong>Stage 1:</strong> Transit time prediction and smoothing of noisy samples</li>\n<li><strong>Stage 2:</strong> Transit modeling and end-to-end neural network</li>\n<li><strong>Postprocessing:</strong> Simple refinements for the transit model and ML-based sigma prediction for the neural network</li>\n</ul>\n<h1><strong>Data Preparation</strong></h1>\n<p>Based on the official submission baseline, we made the following modifications: we adopted a binning factor of 5 to preserve richer temporal information and applied background noise removal to improve signal quality.</p>\n<h2>augmentation</h2>\n<p>We augment the dataset by pairing each signal with all available calibration files, rather than only its nominal pair, to increase sample diversity and improve model robustness.</p>\n<pre><code>signal0 → sample1 \nsignal1 → sample2 \nsignal0 → sample3 \nsignal1 → sample4\n</code></pre>\n<h2>smoothing and normalization</h2>\n<pre><code> ():\n    \n    batch_size, n_channels, n_wavelengths = train_signal.shape\n\n     ():\n        x = np.arange(-size// + , size// + )\n        g = np.exp(-(x**) / (*sigma**))\n         g / g.()\n    gauss_coefs = gaussian_kernel(size=, sigma=)\n\n    \n    q = train_signal[:, :, -win:+win]  \n\n    \n    q = q / q.mean(axis=, keepdims=)\n\n    \n    q_coef = q.mean(axis=(,))  \n\n    \n    t_smooth = train_signal[:, :, -win:+win].copy()\n\n    \n     l  (win, t_smooth.shape[]-win):\n        coefs = q_coef[l-win:l+win+] / q_coef[l]              \n        window = train_signal[:, :, -win+l-win:-win+l+win+]  \n\n        \n        t_smooth[:, :, l] = np.tensordot(window * coefs, gauss_coefs, axes=([],[]))\n\n    \n     win &gt; :\n        t_smooth = t_smooth[:, :, win:-win][:, :, ::-]\n    :\n        t_smooth = t_smooth[:, :, ::-]\n\n    \n    first_col = train_signal[:, :, :]  \n     np.concatenate([first_col, t_smooth], axis=)\n</code></pre>\n<p>Then we do normalization afterwards.</p>\n<h2><strong>Adjust pipeline from physical principle:</strong></h2>\n<p>Dark current represents an <strong>additive noise component</strong> in the measured signal. It should therefore be subtracted <strong>before nonlinearity correction</strong>, which is designed to operate on the <em>true physical signal</em>. Including dark current in the nonlinear transformation would violate this assumption and introduce systematic errors.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F1e1cce4adbabc8694a88b46e35bac821%2Fdown.png?generation=1759019549468977&amp;alt=media\" alt=\"formula\"></p>\n<p><strong>Benefit:</strong><br>\nApplying dark current subtraction before nonlinearity correction ensures that the polynomial correction operates on a cleaner signal, free from additive noise. This prevents the dark current component from being <strong>amplified or distorted by the nonlinear transformation</strong>, resulting in more accurate signal calibration and improved downstream performance.</p>\n<h2>modelling</h2>\n<p>Method 1 : NN with filter kernels, CNN backbone, and projection layer.</p>\n<p>Method 2: transit model and tree model stacking</p>\n<h1>Visualization</h1>\n<p>Shown below are examples of typical anomalous samples frequently observed in the dataset.</p>\n<h2>distribution outliers</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F0cabcf4965f208e652a4a89820dce432%2Fdistribution_outliers_examples.png?generation=1759017681133538&amp;alt=media\" alt=\"distribution_outliers_examples\"></p>\n<h2>extreme values</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F78fdbb98ddd89bdd88c58446a89616c5%2Fextreme_values_examples.png?generation=1759017692452285&amp;alt=media\" alt=\"extreme_values_examples\"></p>\n<h2>gradient anomalies</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F8535d073e9bc82c8379608a2cd771fd2%2Fgradient_anomalies_examples.png?generation=1759017704561057&amp;alt=media\" alt=\"gradient_anomalies_examples\"></p>\n<h2>phase anomalies</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2Fca4371c4c84d0452f006e9e0917a8468%2Fphase_anomalies_examples.png?generation=1759017717715771&amp;alt=media\" alt=\"phase_anomalies_examples\"></p>\n<h2>stability anomalies</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2Fcea8a39b93263b15bf1df0866668e641%2Fstability_anomalies_examples.png?generation=1759017733992433&amp;alt=media\" alt=\"stability_anomalies_examples\"></p>\n<h2>transit anomalies</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F75915953850ccb382e55c6d2cfb8d1a7%2Ftransit_anomalies_examples.png?generation=1759017778125801&amp;alt=media\" alt=\"transit_anomalies_examples\"></p>\n<h2>Postprocessing</h2>\n<p>We didn't figure out robust postprocessing.</p>\n<h1>Conclusion</h1>\n<p>In this work, we built upon the official baseline by introducing a series of targeted improvements across data preprocessing, augmentation, modeling, and calibration. This competition reinforced the importance of careful data preprocessing, the combination of domain knowledge with machine learning, and iterative experimentation. </p>\n<h2><strong>Cheers to all Kagglers, and really have fun!</strong></h2>",
  "messages": [
    {
      "id": 3295184,
      "postDate": "2025-09-28T00:50:50.280Z",
      "content": "<p>I would like to express my sincere gratitude to Kaggle and University College London for hosting this competition. I learned a great deal, and it was truly an exciting and rewarding experience. I would also like to thank my teammates <a href=\"https://www.kaggle.com/ajinomoto132\" target=\"_blank\">@ajinomoto132</a> <a href=\"https://www.kaggle.com/chenzhenyuan\" target=\"_blank\">@chenzhenyuan</a> <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> <a href=\"https://www.kaggle.com/larrylin666\" target=\"_blank\">@larrylin666</a> —together, we proudly earned a silver medal in this competition.</p>\n<h1><strong>Summary</strong></h1>\n<ul>\n<li><strong>Preprocessing:</strong> Noise reduction</li>\n<li><strong>Stage 1:</strong> Transit time prediction and smoothing of noisy samples</li>\n<li><strong>Stage 2:</strong> Transit modeling and end-to-end neural network</li>\n<li><strong>Postprocessing:</strong> Simple refinements for the transit model and ML-based sigma prediction for the neural network</li>\n</ul>\n<h1><strong>Data Preparation</strong></h1>\n<p>Based on the official submission baseline, we made the following modifications: we adopted a binning factor of 5 to preserve richer temporal information and applied background noise removal to improve signal quality.</p>\n<h2>augmentation</h2>\n<p>We augment the dataset by pairing each signal with all available calibration files, rather than only its nominal pair, to increase sample diversity and improve model robustness.</p>\n<pre><code>signal0 → sample1 \nsignal1 → sample2 \nsignal0 → sample3 \nsignal1 → sample4\n</code></pre>\n<h2>smoothing and normalization</h2>\n<pre><code> ():\n    \n    batch_size, n_channels, n_wavelengths = train_signal.shape\n\n     ():\n        x = np.arange(-size// + , size// + )\n        g = np.exp(-(x**) / (*sigma**))\n         g / g.()\n    gauss_coefs = gaussian_kernel(size=, sigma=)\n\n    \n    q = train_signal[:, :, -win:+win]  \n\n    \n    q = q / q.mean(axis=, keepdims=)\n\n    \n    q_coef = q.mean(axis=(,))  \n\n    \n    t_smooth = train_signal[:, :, -win:+win].copy()\n\n    \n     l  (win, t_smooth.shape[]-win):\n        coefs = q_coef[l-win:l+win+] / q_coef[l]              \n        window = train_signal[:, :, -win+l-win:-win+l+win+]  \n\n        \n        t_smooth[:, :, l] = np.tensordot(window * coefs, gauss_coefs, axes=([],[]))\n\n    \n     win &gt; :\n        t_smooth = t_smooth[:, :, win:-win][:, :, ::-]\n    :\n        t_smooth = t_smooth[:, :, ::-]\n\n    \n    first_col = train_signal[:, :, :]  \n     np.concatenate([first_col, t_smooth], axis=)\n</code></pre>\n<p>Then we do normalization afterwards.</p>\n<h2><strong>Adjust pipeline from physical principle:</strong></h2>\n<p>Dark current represents an <strong>additive noise component</strong> in the measured signal. It should therefore be subtracted <strong>before nonlinearity correction</strong>, which is designed to operate on the <em>true physical signal</em>. Including dark current in the nonlinear transformation would violate this assumption and introduce systematic errors.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F1e1cce4adbabc8694a88b46e35bac821%2Fdown.png?generation=1759019549468977&amp;alt=media\" alt=\"formula\"></p>\n<p><strong>Benefit:</strong><br>\nApplying dark current subtraction before nonlinearity correction ensures that the polynomial correction operates on a cleaner signal, free from additive noise. This prevents the dark current component from being <strong>amplified or distorted by the nonlinear transformation</strong>, resulting in more accurate signal calibration and improved downstream performance.</p>\n<h2>modelling</h2>\n<p>Method 1 : NN with filter kernels, CNN backbone, and projection layer.</p>\n<p>Method 2: transit model and tree model stacking</p>\n<h1>Visualization</h1>\n<p>Shown below are examples of typical anomalous samples frequently observed in the dataset.</p>\n<h2>distribution outliers</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F0cabcf4965f208e652a4a89820dce432%2Fdistribution_outliers_examples.png?generation=1759017681133538&amp;alt=media\" alt=\"distribution_outliers_examples\"></p>\n<h2>extreme values</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F78fdbb98ddd89bdd88c58446a89616c5%2Fextreme_values_examples.png?generation=1759017692452285&amp;alt=media\" alt=\"extreme_values_examples\"></p>\n<h2>gradient anomalies</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F8535d073e9bc82c8379608a2cd771fd2%2Fgradient_anomalies_examples.png?generation=1759017704561057&amp;alt=media\" alt=\"gradient_anomalies_examples\"></p>\n<h2>phase anomalies</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2Fca4371c4c84d0452f006e9e0917a8468%2Fphase_anomalies_examples.png?generation=1759017717715771&amp;alt=media\" alt=\"phase_anomalies_examples\"></p>\n<h2>stability anomalies</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2Fcea8a39b93263b15bf1df0866668e641%2Fstability_anomalies_examples.png?generation=1759017733992433&amp;alt=media\" alt=\"stability_anomalies_examples\"></p>\n<h2>transit anomalies</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F75915953850ccb382e55c6d2cfb8d1a7%2Ftransit_anomalies_examples.png?generation=1759017778125801&amp;alt=media\" alt=\"transit_anomalies_examples\"></p>\n<h2>Postprocessing</h2>\n<p>We didn't figure out robust postprocessing.</p>\n<h1>Conclusion</h1>\n<p>In this work, we built upon the official baseline by introducing a series of targeted improvements across data preprocessing, augmentation, modeling, and calibration. This competition reinforced the importance of careful data preprocessing, the combination of domain knowledge with machine learning, and iterative experimentation. </p>\n<h2><strong>Cheers to all Kagglers, and really have fun!</strong></h2>",
      "rawMarkdown": "I would like to express my sincere gratitude to Kaggle and University College London for hosting this competition. I learned a great deal, and it was truly an exciting and rewarding experience. I would also like to thank my teammates @ajinomoto132 @chenzhenyuan @atamazian @larrylin666 —together, we proudly earned a silver medal in this competition.\n\n\n# **Summary**\n\n* **Preprocessing:** Noise reduction\n* **Stage 1:** Transit time prediction and smoothing of noisy samples\n* **Stage 2:** Transit modeling and end-to-end neural network\n* **Postprocessing:** Simple refinements for the transit model and ML-based sigma prediction for the neural network\n\n\n# **Data Preparation**\nBased on the official submission baseline, we made the following modifications: we adopted a binning factor of 5 to preserve richer temporal information and applied background noise removal to improve signal quality.\n\n## augmentation\nWe augment the dataset by pairing each signal with all available calibration files, rather than only its nominal pair, to increase sample diversity and improve model robustness.\n```markdown\nsignal_0 + calibration_0 → sample1 \nsignal_0 + calibration_1 → sample2 \nsignal_1 + calibration_0 → sample3 \nsignal_1 + calibration_1 → sample4\n```\n\n## smoothing and normalization\n\n```python\ndef smooth_data_lambda_batch(train_signal, win=3):\n    \"\"\"\n    Smooth spectral data with Gaussian filter (batch version).\n    \n    Args:\n        train_signal: numpy array of shape (batch_size, n_channels, n_wavelengths)\n        win: window half-size (default=3)\n    \n    Returns:\n        Smoothed signal of shape (batch_size, n_channels, ?)\n    \"\"\"\n    batch_size, n_channels, n_wavelengths = train_signal.shape\n\n    def gaussian_kernel(size=7, sigma=1.0):\n        x = np.arange(-size//2 + 1, size//2 + 1)\n        g = np.exp(-(x**2) / (2*sigma**2))\n        return g / g.sum()\n    gauss_coefs = gaussian_kernel(size=7, sigma=1.0)\n\n    # Slice region of interest\n    q = train_signal[:, :, 40-win:322+win]  # (B, C, slice_len)\n\n    # Normalize each channel by its mean (per batch & channel)\n    q = q / q.mean(axis=2, keepdims=True)\n\n    # Reference spectrum: mean across batch & channels\n    q_coef = q.mean(axis=(0,1))  # shape (slice_len,)\n\n    # Copy ROI for smoothing\n    t_smooth = train_signal[:, :, 40-win:322+win].copy()\n\n    # Loop over wavelengths inside ROI\n    for l in range(win, t_smooth.shape[2]-win):\n        coefs = q_coef[l-win:l+win+1] / q_coef[l]              # (2*win+1,)\n        window = train_signal[:, :, 40-win+l-win:40-win+l+win+1]  # (B, C, 2*win+1)\n        \n        # Weighted Gaussian smoothing\n        t_smooth[:, :, l] = np.tensordot(window * coefs, gauss_coefs, axes=([2],[0]))\n\n    # Trim edges and reverse order\n    if win > 0:\n        t_smooth = t_smooth[:, :, win:-win][:, :, ::-1]\n    else:\n        t_smooth = t_smooth[:, :, ::-1]\n\n    # Concatenate first column (wavelength=0) back\n    first_col = train_signal[:, :, 0:1]  # (B, C, 1)\n    return np.concatenate([first_col, t_smooth], axis=2)\n```\nThen we do normalization afterwards.\n\n## **Adjust pipeline from physical principle:**\nDark current represents an **additive noise component** in the measured signal. It should therefore be subtracted **before nonlinearity correction**, which is designed to operate on the *true physical signal*. Including dark current in the nonlinear transformation would violate this assumption and introduce systematic errors.\n![formula](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F1e1cce4adbabc8694a88b46e35bac821%2Fdown.png?generation=1759019549468977&alt=media)\n\n**Benefit:**\nApplying dark current subtraction before nonlinearity correction ensures that the polynomial correction operates on a cleaner signal, free from additive noise. This prevents the dark current component from being **amplified or distorted by the nonlinear transformation**, resulting in more accurate signal calibration and improved downstream performance.\n\n\n\n## modelling\nMethod 1 : NN with filter kernels, CNN backbone, and projection layer.\n\nMethod 2: transit model and tree model stacking\n\n# Visualization\nShown below are examples of typical anomalous samples frequently observed in the dataset.\n\n## distribution outliers\n![distribution_outliers_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F0cabcf4965f208e652a4a89820dce432%2Fdistribution_outliers_examples.png?generation=1759017681133538&alt=media)\n\n## extreme values \n![extreme_values_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F78fdbb98ddd89bdd88c58446a89616c5%2Fextreme_values_examples.png?generation=1759017692452285&alt=media)\n\n## gradient anomalies \n![gradient_anomalies_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F8535d073e9bc82c8379608a2cd771fd2%2Fgradient_anomalies_examples.png?generation=1759017704561057&alt=media)\n\n## phase anomalies \n![phase_anomalies_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2Fca4371c4c84d0452f006e9e0917a8468%2Fphase_anomalies_examples.png?generation=1759017717715771&alt=media)\n\n## stability anomalies \n![stability_anomalies_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2Fcea8a39b93263b15bf1df0866668e641%2Fstability_anomalies_examples.png?generation=1759017733992433&alt=media)\n\n\n## transit anomalies \n![transit_anomalies_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F75915953850ccb382e55c6d2cfb8d1a7%2Ftransit_anomalies_examples.png?generation=1759017778125801&alt=media)\n\n\n\n## Postprocessing\nWe didn't figure out robust postprocessing.\n\n\n# Conclusion\n\nIn this work, we built upon the official baseline by introducing a series of targeted improvements across data preprocessing, augmentation, modeling, and calibration. This competition reinforced the importance of careful data preprocessing, the combination of domain knowledge with machine learning, and iterative experimentation. \n\n## **Cheers to all Kagglers, and really have fun!**\n\n",
      "votes": 6
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3295184": "I would like to express my sincere gratitude to Kaggle and University College London for hosting this competition. I learned a great deal, and it was truly an exciting and rewarding experience. I would also like to thank my teammates @ajinomoto132 @chenzhenyuan @atamazian @larrylin666 —together, we proudly earned a silver medal in this competition.\n\n\n# **Summary**\n\n* **Preprocessing:** Noise reduction\n* **Stage 1:** Transit time prediction and smoothing of noisy samples\n* **Stage 2:** Transit modeling and end-to-end neural network\n* **Postprocessing:** Simple refinements for the transit model and ML-based sigma prediction for the neural network\n\n\n# **Data Preparation**\nBased on the official submission baseline, we made the following modifications: we adopted a binning factor of 5 to preserve richer temporal information and applied background noise removal to improve signal quality.\n\n## augmentation\nWe augment the dataset by pairing each signal with all available calibration files, rather than only its nominal pair, to increase sample diversity and improve model robustness.\n```markdown\nsignal_0 + calibration_0 → sample1 \nsignal_0 + calibration_1 → sample2 \nsignal_1 + calibration_0 → sample3 \nsignal_1 + calibration_1 → sample4\n```\n\n## smoothing and normalization\n\n```python\ndef smooth_data_lambda_batch(train_signal, win=3):\n    \"\"\"\n    Smooth spectral data with Gaussian filter (batch version).\n    \n    Args:\n        train_signal: numpy array of shape (batch_size, n_channels, n_wavelengths)\n        win: window half-size (default=3)\n    \n    Returns:\n        Smoothed signal of shape (batch_size, n_channels, ?)\n    \"\"\"\n    batch_size, n_channels, n_wavelengths = train_signal.shape\n\n    def gaussian_kernel(size=7, sigma=1.0):\n        x = np.arange(-size//2 + 1, size//2 + 1)\n        g = np.exp(-(x**2) / (2*sigma**2))\n        return g / g.sum()\n    gauss_coefs = gaussian_kernel(size=7, sigma=1.0)\n\n    # Slice region of interest\n    q = train_signal[:, :, 40-win:322+win]  # (B, C, slice_len)\n\n    # Normalize each channel by its mean (per batch & channel)\n    q = q / q.mean(axis=2, keepdims=True)\n\n    # Reference spectrum: mean across batch & channels\n    q_coef = q.mean(axis=(0,1))  # shape (slice_len,)\n\n    # Copy ROI for smoothing\n    t_smooth = train_signal[:, :, 40-win:322+win].copy()\n\n    # Loop over wavelengths inside ROI\n    for l in range(win, t_smooth.shape[2]-win):\n        coefs = q_coef[l-win:l+win+1] / q_coef[l]              # (2*win+1,)\n        window = train_signal[:, :, 40-win+l-win:40-win+l+win+1]  # (B, C, 2*win+1)\n        \n        # Weighted Gaussian smoothing\n        t_smooth[:, :, l] = np.tensordot(window * coefs, gauss_coefs, axes=([2],[0]))\n\n    # Trim edges and reverse order\n    if win > 0:\n        t_smooth = t_smooth[:, :, win:-win][:, :, ::-1]\n    else:\n        t_smooth = t_smooth[:, :, ::-1]\n\n    # Concatenate first column (wavelength=0) back\n    first_col = train_signal[:, :, 0:1]  # (B, C, 1)\n    return np.concatenate([first_col, t_smooth], axis=2)\n```\nThen we do normalization afterwards.\n\n## **Adjust pipeline from physical principle:**\nDark current represents an **additive noise component** in the measured signal. It should therefore be subtracted **before nonlinearity correction**, which is designed to operate on the *true physical signal*. Including dark current in the nonlinear transformation would violate this assumption and introduce systematic errors.\n![formula](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F1e1cce4adbabc8694a88b46e35bac821%2Fdown.png?generation=1759019549468977&alt=media)\n\n**Benefit:**\nApplying dark current subtraction before nonlinearity correction ensures that the polynomial correction operates on a cleaner signal, free from additive noise. This prevents the dark current component from being **amplified or distorted by the nonlinear transformation**, resulting in more accurate signal calibration and improved downstream performance.\n\n\n\n## modelling\nMethod 1 : NN with filter kernels, CNN backbone, and projection layer.\n\nMethod 2: transit model and tree model stacking\n\n# Visualization\nShown below are examples of typical anomalous samples frequently observed in the dataset.\n\n## distribution outliers\n![distribution_outliers_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F0cabcf4965f208e652a4a89820dce432%2Fdistribution_outliers_examples.png?generation=1759017681133538&alt=media)\n\n## extreme values \n![extreme_values_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F78fdbb98ddd89bdd88c58446a89616c5%2Fextreme_values_examples.png?generation=1759017692452285&alt=media)\n\n## gradient anomalies \n![gradient_anomalies_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F8535d073e9bc82c8379608a2cd771fd2%2Fgradient_anomalies_examples.png?generation=1759017704561057&alt=media)\n\n## phase anomalies \n![phase_anomalies_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2Fca4371c4c84d0452f006e9e0917a8468%2Fphase_anomalies_examples.png?generation=1759017717715771&alt=media)\n\n## stability anomalies \n![stability_anomalies_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2Fcea8a39b93263b15bf1df0866668e641%2Fstability_anomalies_examples.png?generation=1759017733992433&alt=media)\n\n\n## transit anomalies \n![transit_anomalies_examples](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F22146004%2F75915953850ccb382e55c6d2cfb8d1a7%2Ftransit_anomalies_examples.png?generation=1759017778125801&alt=media)\n\n\n\n## Postprocessing\nWe didn't figure out robust postprocessing.\n\n\n# Conclusion\n\nIn this work, we built upon the official baseline by introducing a series of targeted improvements across data preprocessing, augmentation, modeling, and calibration. This competition reinforced the importance of careful data preprocessing, the combination of domain knowledge with machine learning, and iterative experimentation. \n\n## **Cheers to all Kagglers, and really have fun!**\n\n"
  }
}