{
  "id": 609375,
  "title": "12th Place Solution: E2E NN",
  "url": "/competitions/ariel-data-challenge-2025/discussion/609375",
  "author_name": "sroger",
  "post_date": "2025-09-26T09:43:43.491000",
  "votes": 8,
  "comment_count": 2,
  "views": 0,
  "content": "<p><strong>Preface</strong><br>\nThanks to the host and Kaggle team for organizing this competition, I learned a lot from this first contact with astronomy and I hope to participate again.<br>\nMy approach is quite simple and just a neural network end to end. I ended up ensembling 3 for my final submission, but a single model can reach 0.552.</p>\n<p><strong>Calibration</strong><br>\nI adapted this amazing <a href=\"https://www.kaggle.com/code/ilu000/ariel25-quick-data-prep-improved\" target=\"_blank\">notebook</a> by Pascal with minor changes, keeping the 15 binning.</p>\n<p><strong>Preprocessing</strong><br>\nSignals were clipped to remove outliers then group norm (over lambda) was applied.<br>\nOrbit parameters in star_info were not used.</p>\n<p><strong>Transits</strong><br>\nDuring inference the model internally outputs transit points t1 to t4, and they are trained in a semi-supervised manner with labels derived from the gradients of the savgol smoothed signal and filtered for validity.</p>\n<p><strong>Architecture</strong><br>\nThe model is a bert style transformer encoder and takes preprocessed signals as input and outputs mu and sigma (logvar).<br>\nThe intermediary transit points are fed into the final output head.</p>\n<p><strong>Training</strong><br>\nSince the model outputs both mu and sigma, I trained directly on the competition metric with gaussian NLL, and the (intermediary) transit head was trained with smooth L1 and an unsupervised symmetry loss (also smooth l1).<br>\nFor LR, cosine decay scheduling was used in addition to very high weight decay and a little dropout.</p>\n<p><strong>Augmentation</strong><br>\nFlip and aperture augmentation were applied to the signals during training as well as adding noise to the transit points.</p>\n<p><strong>Inference</strong><br>\nFlip TTA boosted LB by quite a bit and ensembling was done with inverse variance weighting.</p>\n<p><strong>Caveats</strong><br>\nTraining variance was high (without deterministic algorithms), the same seed can converge at quite different places, likely because of the sensitivity of GNLL. This made it quite difficult to tune hyperparameters reliably with my limited compute.<br>\nExtra signals didn't have a noticeable effect during training nor inference.</p>\n<p><strong>Todo</strong><br>\nUtilize orbit parameters and disentangle gain drift.</p>",
  "messages": [
    {
      "id": 3294507,
      "postDate": "2025-09-26T09:43:43.490Z",
      "content": "<p><strong>Preface</strong><br>\nThanks to the host and Kaggle team for organizing this competition, I learned a lot from this first contact with astronomy and I hope to participate again.<br>\nMy approach is quite simple and just a neural network end to end. I ended up ensembling 3 for my final submission, but a single model can reach 0.552.</p>\n<p><strong>Calibration</strong><br>\nI adapted this amazing <a href=\"https://www.kaggle.com/code/ilu000/ariel25-quick-data-prep-improved\" target=\"_blank\">notebook</a> by Pascal with minor changes, keeping the 15 binning.</p>\n<p><strong>Preprocessing</strong><br>\nSignals were clipped to remove outliers then group norm (over lambda) was applied.<br>\nOrbit parameters in star_info were not used.</p>\n<p><strong>Transits</strong><br>\nDuring inference the model internally outputs transit points t1 to t4, and they are trained in a semi-supervised manner with labels derived from the gradients of the savgol smoothed signal and filtered for validity.</p>\n<p><strong>Architecture</strong><br>\nThe model is a bert style transformer encoder and takes preprocessed signals as input and outputs mu and sigma (logvar).<br>\nThe intermediary transit points are fed into the final output head.</p>\n<p><strong>Training</strong><br>\nSince the model outputs both mu and sigma, I trained directly on the competition metric with gaussian NLL, and the (intermediary) transit head was trained with smooth L1 and an unsupervised symmetry loss (also smooth l1).<br>\nFor LR, cosine decay scheduling was used in addition to very high weight decay and a little dropout.</p>\n<p><strong>Augmentation</strong><br>\nFlip and aperture augmentation were applied to the signals during training as well as adding noise to the transit points.</p>\n<p><strong>Inference</strong><br>\nFlip TTA boosted LB by quite a bit and ensembling was done with inverse variance weighting.</p>\n<p><strong>Caveats</strong><br>\nTraining variance was high (without deterministic algorithms), the same seed can converge at quite different places, likely because of the sensitivity of GNLL. This made it quite difficult to tune hyperparameters reliably with my limited compute.<br>\nExtra signals didn't have a noticeable effect during training nor inference.</p>\n<p><strong>Todo</strong><br>\nUtilize orbit parameters and disentangle gain drift.</p>",
      "rawMarkdown": "**Preface**\nThanks to the host and Kaggle team for organizing this competition, I learned a lot from this first contact with astronomy and I hope to participate again.\nMy approach is quite simple and just a neural network end to end. I ended up ensembling 3 for my final submission, but a single model can reach 0.552.\n\n**Calibration**\nI adapted this amazing [notebook](https://www.kaggle.com/code/ilu000/ariel25-quick-data-prep-improved) by Pascal with minor changes, keeping the 15 binning.\n\n**Preprocessing**\nSignals were clipped to remove outliers then group norm (over lambda) was applied.\nOrbit parameters in star_info were not used.\n\n**Transits**\nDuring inference the model internally outputs transit points t1 to t4, and they are trained in a semi-supervised manner with labels derived from the gradients of the savgol smoothed signal and filtered for validity.\n\n**Architecture**\nThe model is a bert style transformer encoder and takes preprocessed signals as input and outputs mu and sigma (logvar).\nThe intermediary transit points are fed into the final output head.\n\n**Training**\nSince the model outputs both mu and sigma, I trained directly on the competition metric with gaussian NLL, and the (intermediary) transit head was trained with smooth L1 and an unsupervised symmetry loss (also smooth l1).\nFor LR, cosine decay scheduling was used in addition to very high weight decay and a little dropout.\n\n**Augmentation**\nFlip and aperture augmentation were applied to the signals during training as well as adding noise to the transit points.\n\n**Inference**\nFlip TTA boosted LB by quite a bit and ensembling was done with inverse variance weighting.\n\n**Caveats**\nTraining variance was high (without deterministic algorithms), the same seed can converge at quite different places, likely because of the sensitivity of GNLL. This made it quite difficult to tune hyperparameters reliably with my limited compute.\nExtra signals didn't have a noticeable effect during training nor inference.\n\n**Todo**\nUtilize orbit parameters and disentangle gain drift.",
      "votes": 8
    },
    {
      "id": 3294907,
      "postDate": "2025-09-27T04:48:16.123Z",
      "content": "<p>Congratulations! <br>\n'Training variance was high (without deterministic algorithms), the same seed can converge at quite different places, likely because of the sensitivity of GNLL'.</p>\n<p>I saw a bit of the same thing when working with NLL. I only did 3 seeds per fold, but I keep wondering if pushing that number up could’ve improved the score. Definitely something worth trying out.</p>",
      "rawMarkdown": "Congratulations! \n'Training variance was high (without deterministic algorithms), the same seed can converge at quite different places, likely because of the sensitivity of GNLL'.\n\nI saw a bit of the same thing when working with NLL. I only did 3 seeds per fold, but I keep wondering if pushing that number up could’ve improved the score. Definitely something worth trying out.",
      "votes": 1,
      "replies": [
        {
          "id": 3294959,
          "postDate": "2025-09-27T08:34:43.447Z",
          "content": "<p>Nice to see some confirmation! It's my first time working with NLL so this was quite unexpected.</p>",
          "rawMarkdown": "Nice to see some confirmation! It's my first time working with NLL so this was quite unexpected."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3294907,
      "author_name": "Viji",
      "author_url": "",
      "post_date": "2025-09-27T04:48:16.123000",
      "content": "<p>Congratulations! <br>\n'Training variance was high (without deterministic algorithms), the same seed can converge at quite different places, likely because of the sensitivity of GNLL'.</p>\n<p>I saw a bit of the same thing when working with NLL. I only did 3 seeds per fold, but I keep wondering if pushing that number up could’ve improved the score. Definitely something worth trying out.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3294959,
          "author_name": "sroger",
          "author_url": "",
          "post_date": "2025-09-27T08:34:43.447000",
          "content": "<p>Nice to see some confirmation! It's my first time working with NLL so this was quite unexpected.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3294507": "**Preface**\nThanks to the host and Kaggle team for organizing this competition, I learned a lot from this first contact with astronomy and I hope to participate again.\nMy approach is quite simple and just a neural network end to end. I ended up ensembling 3 for my final submission, but a single model can reach 0.552.\n\n**Calibration**\nI adapted this amazing [notebook](https://www.kaggle.com/code/ilu000/ariel25-quick-data-prep-improved) by Pascal with minor changes, keeping the 15 binning.\n\n**Preprocessing**\nSignals were clipped to remove outliers then group norm (over lambda) was applied.\nOrbit parameters in star_info were not used.\n\n**Transits**\nDuring inference the model internally outputs transit points t1 to t4, and they are trained in a semi-supervised manner with labels derived from the gradients of the savgol smoothed signal and filtered for validity.\n\n**Architecture**\nThe model is a bert style transformer encoder and takes preprocessed signals as input and outputs mu and sigma (logvar).\nThe intermediary transit points are fed into the final output head.\n\n**Training**\nSince the model outputs both mu and sigma, I trained directly on the competition metric with gaussian NLL, and the (intermediary) transit head was trained with smooth L1 and an unsupervised symmetry loss (also smooth l1).\nFor LR, cosine decay scheduling was used in addition to very high weight decay and a little dropout.\n\n**Augmentation**\nFlip and aperture augmentation were applied to the signals during training as well as adding noise to the transit points.\n\n**Inference**\nFlip TTA boosted LB by quite a bit and ensembling was done with inverse variance weighting.\n\n**Caveats**\nTraining variance was high (without deterministic algorithms), the same seed can converge at quite different places, likely because of the sensitivity of GNLL. This made it quite difficult to tune hyperparameters reliably with my limited compute.\nExtra signals didn't have a noticeable effect during training nor inference.\n\n**Todo**\nUtilize orbit parameters and disentangle gain drift.",
    "3294907": "Congratulations! \n'Training variance was high (without deterministic algorithms), the same seed can converge at quite different places, likely because of the sensitivity of GNLL'.\n\nI saw a bit of the same thing when working with NLL. I only did 3 seeds per fold, but I keep wondering if pushing that number up could’ve improved the score. Definitely something worth trying out."
  }
}