{
  "id": 544026,
  "title": "37th place solution",
  "url": "/competitions/ariel-data-challenge-2024/discussion/544026",
  "author_name": "Shamil Yagiyayev",
  "post_date": "2024-11-02T18:30:45.729000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to thank the hosts for the preparations, my teammates <a href=\"https://www.kaggle.com/imayushsaxena\" target=\"_blank\">@imayushsaxena</a> and <a href=\"https://www.kaggle.com/octaviograu\" target=\"_blank\">@octaviograu</a> and all the competitors who made it that challenging! </p>\n<h2>Our top solutions</h2>\n<p>The solution got </p>\n<table>\n<thead>\n<tr>\n<th>variant</th>\n<th>public</th>\n<th>private</th>\n<th>LB place</th>\n<th>Final place</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>10th</td>\n<td>0.6152848</td>\n<td>0.6198108</td>\n<td>53</td>\n<td>-</td>\n</tr>\n<tr>\n<td>11th</td>\n<td>0.6128403</td>\n<td><em>0.6293136</em></td>\n<td>-</td>\n<td>37</td>\n</tr>\n</tbody>\n</table>\n<h2>Summary of 11th</h2>\n<p>We used only AIRS data for the solution</p>\n<p>The basic flow of the solution:</p>\n<ol>\n<li>Standard correction, but without binning / shrinking the flux </li>\n<li>Normalize the signal</li>\n<li>Denoise the signal with <code>gaussian_filter1d</code></li>\n<li>Detect transit phases, including limb darkening</li>\n<li>Calculate features for the whole signal</li>\n<li>Calculate features for binned wavelengths by grouping nearest 6, 10, and 25 wavelengths.</li>\n<li>On top of the features 1000 models were trained </li>\n<li>The mean of the model's predictions gives the transit depth prediction</li>\n<li>Sigma is based on the std for the model's  predictions per wavelength</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/code/thegrey/ad24-eleventh?scriptVersionId=204479373\" target=\"_blank\">Link</a> to the 11th solution.</p>\n<h2>Backbone pipeline</h2>\n<p>We tried many denoising strategies, so we needed a simple way to iterate. We built a functional pipeline that's easy to configure:</p>\n<pre><code> functools  reduce\n\nfull_pipeline = [\n      **kwargs: bin_wavelengths(**kwargs),\n      **kwargs: normalize_signal(**kwargs),\n     **kwargs: filter_signal_with_gaussian(sigma=, **kwargs),\n     **kwargs: polynomial_detrend(**kwargs),\n     **kwargs: fit_transit_poly(**kwargs),\n     **kwargs: collect_features(**kwargs),\n]\n\ninitial_kwargs = {: , : , :, :}\nresult = reduce( kw, func: func(**kw), full_pipeline, initial_kwargs)\n</code></pre>\n<p>Then we can combine <code>result['features']</code> for different pipelines. </p>\n<p>Each step in the pipeline can be organized like this:</p>\n<pre><code> ():\n    signal = kwargs[]\n    filtered_signal = median_filter(signal, size=window)\n    kwargs[] = filtered_signal\n     kwargs\n</code></pre>\n<h2>Transit phase detection</h2>\n<p>After experimenting with the <a href=\"https://lkreidberg.github.io/batman/docs/html/index.html\" target=\"_blank\">BATMAN</a> library, we found that there are several possible structures of the transit flux signal, even when fully denoised and detrended. Limb darkening can influence the calculation of transit depth, affecting how we calculate the transit edges</p>\n<pre><code>  ():\n    signal = kwargs[] + \n\n    grad1 = np.gradient(signal)\n    mn = np.argmin(grad1)\n    mx = np.argmax(grad1)\n\n    peaks = find_peaks(grad1, width=)[]\n    ingress_left = peaks[(peaks &lt; mn) &amp; (mn - peaks &lt;= )][-]  np.((peaks &lt; mn) &amp; (mn - peaks &lt;= ))  mn-\n    ingress_right = peaks[(peaks &gt; mn) &amp; (peaks - mn &lt;= )][]  np.((peaks &gt; mn) &amp; (peaks - mn &lt;= ))  mn+\n\n    dips = find_peaks(-grad1, width=)[]\n    egress_left = dips[(dips &lt; mx) &amp; (mx - dips &lt;= )][-]  np.((dips &lt; mx) &amp; (mx - dips &lt;= ))  mx-\n    egress_right = dips[(dips &gt; mx) &amp; (dips - mx &lt;= )][]  np.((dips &gt; mx) &amp; (dips - mx &lt;= ))  mx+\n\n    kwargs[] = ingress_left\n    kwargs[] = ingress_right\n    kwargs[] = egress_left\n    kwargs[] = egress_right\n\n     kwargs\n</code></pre>\n<h3>Features calculations</h3>\n<p>Knowing the <em>ingress</em> and <em>egress</em> of the transit, we fit a polynomial into the flux signal (ignoring the transit) and then perform detrending.</p>\n<p>Then we can assume that once the plain detrended signal is on the level of constant 1, then we can derive a transit depth by the formula:</p>\n<pre><code> ():\n     /s - \n</code></pre>\n<p>Here <code>s</code> can be a mean flux in the transit, flux at the middle point, etc. We collect quite a few here. This was a huge speed improvement compared to the <code>minimize</code> method. </p>\n<p>Additionally, we collect generic features on the in-transit flux, like the standard deviation.</p>\n<h2>Models training</h2>\n<p>We have generated 500 Ridge models and 500 PLSRegression models that formed the ensemble. </p>\n<p>Alpha for Ridge was random:</p>\n<pre><code>alpha =  ** ( -  * np.random.rand())\n</code></pre>\n<p>The same as <code>n_components</code> for PLSRegression:</p>\n<pre><code>model = PLSRegression(n_components=n_components, tol=)\n</code></pre>\n<p>We used different subsets of the training data for each model:</p>\n<pre><code>subset_size = (X_train.shape[] * np.random.uniform(, ))\n</code></pre>\n<p>We sorted the model predictions for each wavelength, removed the lowest and highest 30%, and took the mean of the remaining predictions.</p>\n<h3>Sigma calculation</h3>\n<p>We calculated the standard deviation of the sorted predictions for each wavelength:  </p>\n<pre><code>std_preds = np.std(sorted_preds, axis=)\n</code></pre>\n<p>We then multiplied the standard deviation by a coefficient that worked best in LB:</p>\n<pre><code>COEFF_0 =  \nsigma_response = sigma * COEFF_0\n</code></pre>\n<h2>What worked well, though didn't make it to the final submission</h2>\n<ol>\n<li>Usage of wavelets (<code>pywt</code>) - our favorite one was <code>db29</code></li>\n<li><code>median_filter</code></li>\n<li><code>savgol_filter</code></li>\n</ol>\n<h2>What didn't work</h2>\n<ol>\n<li>Gaussian Process Regression - we tried different kernels here, but didn't see breakthroughs while a big increase of the processing time</li>\n<li>LinearRegresion, SVM, LGBM, CatBoost, RandomForest, XGBoost… </li>\n<li>BayesianRidge - worked well, though long time (and didn't fit into the memory for more features)</li>\n<li>We tried to train a neural network with a loss function for both sigma and transit depth (the competition score), but didn't have much luck due to the limited time</li>\n</ol>\n<h2>Improvements We Lacked Time For</h2>\n<ol>\n<li>Generate the train data with BATMAN  </li>\n<li>Use FGS data</li>\n<li>Use data for real molecules to generate additional training data</li>\n<li>Collect more meaningful features (and filter out those which don't help)</li>\n<li>Sigma approaches - we didn't use all the potential here</li>\n</ol>\n<p>Please feel free to explore our <a href=\"https://www.kaggle.com/code/thegrey/ad24-eleventh?scriptVersionId=204479373\" target=\"_blank\">best solution</a>.</p>",
  "messages": [
    {
      "id": 3034896,
      "postDate": "2024-11-02T18:30:45.730Z",
      "content": "<p>First of all, I would like to thank the hosts for the preparations, my teammates <a href=\"https://www.kaggle.com/imayushsaxena\" target=\"_blank\">@imayushsaxena</a> and <a href=\"https://www.kaggle.com/octaviograu\" target=\"_blank\">@octaviograu</a> and all the competitors who made it that challenging! </p>\n<h2>Our top solutions</h2>\n<p>The solution got </p>\n<table>\n<thead>\n<tr>\n<th>variant</th>\n<th>public</th>\n<th>private</th>\n<th>LB place</th>\n<th>Final place</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>10th</td>\n<td>0.6152848</td>\n<td>0.6198108</td>\n<td>53</td>\n<td>-</td>\n</tr>\n<tr>\n<td>11th</td>\n<td>0.6128403</td>\n<td><em>0.6293136</em></td>\n<td>-</td>\n<td>37</td>\n</tr>\n</tbody>\n</table>\n<h2>Summary of 11th</h2>\n<p>We used only AIRS data for the solution</p>\n<p>The basic flow of the solution:</p>\n<ol>\n<li>Standard correction, but without binning / shrinking the flux </li>\n<li>Normalize the signal</li>\n<li>Denoise the signal with <code>gaussian_filter1d</code></li>\n<li>Detect transit phases, including limb darkening</li>\n<li>Calculate features for the whole signal</li>\n<li>Calculate features for binned wavelengths by grouping nearest 6, 10, and 25 wavelengths.</li>\n<li>On top of the features 1000 models were trained </li>\n<li>The mean of the model's predictions gives the transit depth prediction</li>\n<li>Sigma is based on the std for the model's  predictions per wavelength</li>\n</ol>\n<p><a href=\"https://www.kaggle.com/code/thegrey/ad24-eleventh?scriptVersionId=204479373\" target=\"_blank\">Link</a> to the 11th solution.</p>\n<h2>Backbone pipeline</h2>\n<p>We tried many denoising strategies, so we needed a simple way to iterate. We built a functional pipeline that's easy to configure:</p>\n<pre><code> functools  reduce\n\nfull_pipeline = [\n      **kwargs: bin_wavelengths(**kwargs),\n      **kwargs: normalize_signal(**kwargs),\n     **kwargs: filter_signal_with_gaussian(sigma=, **kwargs),\n     **kwargs: polynomial_detrend(**kwargs),\n     **kwargs: fit_transit_poly(**kwargs),\n     **kwargs: collect_features(**kwargs),\n]\n\ninitial_kwargs = {: , : , :, :}\nresult = reduce( kw, func: func(**kw), full_pipeline, initial_kwargs)\n</code></pre>\n<p>Then we can combine <code>result['features']</code> for different pipelines. </p>\n<p>Each step in the pipeline can be organized like this:</p>\n<pre><code> ():\n    signal = kwargs[]\n    filtered_signal = median_filter(signal, size=window)\n    kwargs[] = filtered_signal\n     kwargs\n</code></pre>\n<h2>Transit phase detection</h2>\n<p>After experimenting with the <a href=\"https://lkreidberg.github.io/batman/docs/html/index.html\" target=\"_blank\">BATMAN</a> library, we found that there are several possible structures of the transit flux signal, even when fully denoised and detrended. Limb darkening can influence the calculation of transit depth, affecting how we calculate the transit edges</p>\n<pre><code>  ():\n    signal = kwargs[] + \n\n    grad1 = np.gradient(signal)\n    mn = np.argmin(grad1)\n    mx = np.argmax(grad1)\n\n    peaks = find_peaks(grad1, width=)[]\n    ingress_left = peaks[(peaks &lt; mn) &amp; (mn - peaks &lt;= )][-]  np.((peaks &lt; mn) &amp; (mn - peaks &lt;= ))  mn-\n    ingress_right = peaks[(peaks &gt; mn) &amp; (peaks - mn &lt;= )][]  np.((peaks &gt; mn) &amp; (peaks - mn &lt;= ))  mn+\n\n    dips = find_peaks(-grad1, width=)[]\n    egress_left = dips[(dips &lt; mx) &amp; (mx - dips &lt;= )][-]  np.((dips &lt; mx) &amp; (mx - dips &lt;= ))  mx-\n    egress_right = dips[(dips &gt; mx) &amp; (dips - mx &lt;= )][]  np.((dips &gt; mx) &amp; (dips - mx &lt;= ))  mx+\n\n    kwargs[] = ingress_left\n    kwargs[] = ingress_right\n    kwargs[] = egress_left\n    kwargs[] = egress_right\n\n     kwargs\n</code></pre>\n<h3>Features calculations</h3>\n<p>Knowing the <em>ingress</em> and <em>egress</em> of the transit, we fit a polynomial into the flux signal (ignoring the transit) and then perform detrending.</p>\n<p>Then we can assume that once the plain detrended signal is on the level of constant 1, then we can derive a transit depth by the formula:</p>\n<pre><code> ():\n     /s - \n</code></pre>\n<p>Here <code>s</code> can be a mean flux in the transit, flux at the middle point, etc. We collect quite a few here. This was a huge speed improvement compared to the <code>minimize</code> method. </p>\n<p>Additionally, we collect generic features on the in-transit flux, like the standard deviation.</p>\n<h2>Models training</h2>\n<p>We have generated 500 Ridge models and 500 PLSRegression models that formed the ensemble. </p>\n<p>Alpha for Ridge was random:</p>\n<pre><code>alpha =  ** ( -  * np.random.rand())\n</code></pre>\n<p>The same as <code>n_components</code> for PLSRegression:</p>\n<pre><code>model = PLSRegression(n_components=n_components, tol=)\n</code></pre>\n<p>We used different subsets of the training data for each model:</p>\n<pre><code>subset_size = (X_train.shape[] * np.random.uniform(, ))\n</code></pre>\n<p>We sorted the model predictions for each wavelength, removed the lowest and highest 30%, and took the mean of the remaining predictions.</p>\n<h3>Sigma calculation</h3>\n<p>We calculated the standard deviation of the sorted predictions for each wavelength:  </p>\n<pre><code>std_preds = np.std(sorted_preds, axis=)\n</code></pre>\n<p>We then multiplied the standard deviation by a coefficient that worked best in LB:</p>\n<pre><code>COEFF_0 =  \nsigma_response = sigma * COEFF_0\n</code></pre>\n<h2>What worked well, though didn't make it to the final submission</h2>\n<ol>\n<li>Usage of wavelets (<code>pywt</code>) - our favorite one was <code>db29</code></li>\n<li><code>median_filter</code></li>\n<li><code>savgol_filter</code></li>\n</ol>\n<h2>What didn't work</h2>\n<ol>\n<li>Gaussian Process Regression - we tried different kernels here, but didn't see breakthroughs while a big increase of the processing time</li>\n<li>LinearRegresion, SVM, LGBM, CatBoost, RandomForest, XGBoost… </li>\n<li>BayesianRidge - worked well, though long time (and didn't fit into the memory for more features)</li>\n<li>We tried to train a neural network with a loss function for both sigma and transit depth (the competition score), but didn't have much luck due to the limited time</li>\n</ol>\n<h2>Improvements We Lacked Time For</h2>\n<ol>\n<li>Generate the train data with BATMAN  </li>\n<li>Use FGS data</li>\n<li>Use data for real molecules to generate additional training data</li>\n<li>Collect more meaningful features (and filter out those which don't help)</li>\n<li>Sigma approaches - we didn't use all the potential here</li>\n</ol>\n<p>Please feel free to explore our <a href=\"https://www.kaggle.com/code/thegrey/ad24-eleventh?scriptVersionId=204479373\" target=\"_blank\">best solution</a>.</p>",
      "rawMarkdown": "First of all, I would like to thank the hosts for the preparations, my teammates @imayushsaxena and @octaviograu and all the competitors who made it that challenging! \n\n## Our top solutions\nThe solution got \n| variant | public | private | LB place | Final place |\n| --- | --- | --- | --- | --- |\n| 10th |  0.6152848 |  0.6198108 | 53 | - |\n| 11th |  0.6128403 |  *0.6293136* | - | 37 |\n\n## Summary of 11th\n\nWe used only AIRS data for the solution\n\nThe basic flow of the solution:\n1. Standard correction, but without binning / shrinking the flux \n2. Normalize the signal\n3. Denoise the signal with `gaussian_filter1d`\n4. Detect transit phases, including limb darkening\n5. Calculate features for the whole signal\n6. Calculate features for binned wavelengths by grouping nearest 6, 10, and 25 wavelengths.\n7. On top of the features 1000 models were trained \n8. The mean of the model's predictions gives the transit depth prediction\n9. Sigma is based on the std for the model's  predictions per wavelength\n\n[Link](https://www.kaggle.com/code/thegrey/ad24-eleventh?scriptVersionId=204479373) to the 11th solution.\n\n## Backbone pipeline \n\nWe tried many denoising strategies, so we needed a simple way to iterate. We built a functional pipeline that's easy to configure:\n\n```python\nfrom functools import reduce\n\nfull_pipeline = [\n    lambda  **kwargs: bin_wavelengths(**kwargs),\n    lambda  **kwargs: normalize_signal(**kwargs),\n    lambda **kwargs: filter_signal_with_gaussian(sigma=25, **kwargs),\n    lambda **kwargs: polynomial_detrend(**kwargs),\n    lambda **kwargs: fit_transit_poly(**kwargs),\n    lambda **kwargs: collect_features(**kwargs),\n]\n\ninitial_kwargs = {'signal': None, 'margin': 1, 'from_wavelength':39, 'to_wavelength':321}\nresult = reduce(lambda kw, func: func(**kw), full_pipeline, initial_kwargs)\n```\n\nThen we can combine `result['features']` for different pipelines. \n\nEach step in the pipeline can be organized like this:\n\n```python\ndef filter_signal_with_median(window=5, **kwargs):\n    signal = kwargs['signal']\n    filtered_signal = median_filter(signal, size=window)\n    kwargs['signal'] = filtered_signal\n    return kwargs\n```\n\n## Transit phase detection\n\nAfter experimenting with the [BATMAN](https://lkreidberg.github.io/batman/docs/html/index.html) library, we found that there are several possible structures of the transit flux signal, even when fully denoised and detrended. Limb darkening can influence the calculation of transit depth, affecting how we calculate the transit edges\n\n```python\n def get_transit_phases2(**kwargs):\n    signal = kwargs['signal'] + 1\n\n    grad1 = np.gradient(signal)\n    mn = np.argmin(grad1)\n    mx = np.argmax(grad1)\n    \n    peaks = find_peaks(grad1, width=25)[0]\n    ingress_left = peaks[(peaks < mn) & (mn - peaks <= 250)][-1] if np.any((peaks < mn) & (mn - peaks <= 250)) else mn-110\n    ingress_right = peaks[(peaks > mn) & (peaks - mn <= 250)][0] if np.any((peaks > mn) & (peaks - mn <= 250)) else mn+110\n    \n    dips = find_peaks(-grad1, width=25)[0]\n    egress_left = dips[(dips < mx) & (mx - dips <= 250)][-1] if np.any((dips < mx) & (mx - dips <= 250)) else mx-110\n    egress_right = dips[(dips > mx) & (dips - mx <= 250)][0] if np.any((dips > mx) & (dips - mx <= 250)) else mx+110\n\n    kwargs['ingress_left'] = ingress_left\n    kwargs['ingress_right'] = ingress_right\n    kwargs['egress_left'] = egress_left\n    kwargs['egress_right'] = egress_right\n\n    return kwargs\n```\n\n### Features calculations\n\nKnowing the *ingress* and *egress* of the transit, we fit a polynomial into the flux signal (ignoring the transit) and then perform detrending.\n\nThen we can assume that once the plain detrended signal is on the level of constant 1, then we can derive a transit depth by the formula:\n\n```python\ndef transit_formula(s):\n    return 1/s - 1\n```\nHere `s` can be a mean flux in the transit, flux at the middle point, etc. We collect quite a few here. This was a huge speed improvement compared to the `minimize` method. \n\nAdditionally, we collect generic features on the in-transit flux, like the standard deviation.\n\n## Models training\n\nWe have generated 500 Ridge models and 500 PLSRegression models that formed the ensemble. \n\nAlpha for Ridge was random:\n```python\nalpha = 10 ** (2 - 3 * np.random.rand())\n```\n\nThe same as `n_components` for PLSRegression:\n```python\nmodel = PLSRegression(n_components=n_components, tol=1e-8)\n```\n\nWe used different subsets of the training data for each model:\n```python\nsubset_size = int(X_train.shape[0] * np.random.uniform(0.5, 1))\n```\n\nWe sorted the model predictions for each wavelength, removed the lowest and highest 30%, and took the mean of the remaining predictions.\n\n### Sigma calculation\n\nWe calculated the standard deviation of the sorted predictions for each wavelength:  \n```python\nstd_preds = np.std(sorted_preds, axis=0)\n```\n\nWe then multiplied the standard deviation by a coefficient that worked best in LB:\n```python\nCOEFF_0 = 4 # This one worked the best\nsigma_response = sigma * COEFF_0\n```\n\n## What worked well, though didn't make it to the final submission\n1. Usage of wavelets (`pywt`) - our favorite one was `db29`\n2. `median_filter`\n3. `savgol_filter`\n\n## What didn't work\n1. Gaussian Process Regression - we tried different kernels here, but didn't see breakthroughs while a big increase of the processing time\n2. LinearRegresion, SVM, LGBM, CatBoost, RandomForest, XGBoost... \n3. BayesianRidge - worked well, though long time (and didn't fit into the memory for more features)\n4. We tried to train a neural network with a loss function for both sigma and transit depth (the competition score), but didn't have much luck due to the limited time\n\n## Improvements We Lacked Time For\n1. Generate the train data with BATMAN  \n2. Use FGS data\n3. Use data for real molecules to generate additional training data\n4. Collect more meaningful features (and filter out those which don't help)\n5. Sigma approaches - we didn't use all the potential here\n\nPlease feel free to explore our [best solution](https://www.kaggle.com/code/thegrey/ad24-eleventh?scriptVersionId=204479373).",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3034896": "First of all, I would like to thank the hosts for the preparations, my teammates @imayushsaxena and @octaviograu and all the competitors who made it that challenging! \n\n## Our top solutions\nThe solution got \n| variant | public | private | LB place | Final place |\n| --- | --- | --- | --- | --- |\n| 10th |  0.6152848 |  0.6198108 | 53 | - |\n| 11th |  0.6128403 |  *0.6293136* | - | 37 |\n\n## Summary of 11th\n\nWe used only AIRS data for the solution\n\nThe basic flow of the solution:\n1. Standard correction, but without binning / shrinking the flux \n2. Normalize the signal\n3. Denoise the signal with `gaussian_filter1d`\n4. Detect transit phases, including limb darkening\n5. Calculate features for the whole signal\n6. Calculate features for binned wavelengths by grouping nearest 6, 10, and 25 wavelengths.\n7. On top of the features 1000 models were trained \n8. The mean of the model's predictions gives the transit depth prediction\n9. Sigma is based on the std for the model's  predictions per wavelength\n\n[Link](https://www.kaggle.com/code/thegrey/ad24-eleventh?scriptVersionId=204479373) to the 11th solution.\n\n## Backbone pipeline \n\nWe tried many denoising strategies, so we needed a simple way to iterate. We built a functional pipeline that's easy to configure:\n\n```python\nfrom functools import reduce\n\nfull_pipeline = [\n    lambda  **kwargs: bin_wavelengths(**kwargs),\n    lambda  **kwargs: normalize_signal(**kwargs),\n    lambda **kwargs: filter_signal_with_gaussian(sigma=25, **kwargs),\n    lambda **kwargs: polynomial_detrend(**kwargs),\n    lambda **kwargs: fit_transit_poly(**kwargs),\n    lambda **kwargs: collect_features(**kwargs),\n]\n\ninitial_kwargs = {'signal': None, 'margin': 1, 'from_wavelength':39, 'to_wavelength':321}\nresult = reduce(lambda kw, func: func(**kw), full_pipeline, initial_kwargs)\n```\n\nThen we can combine `result['features']` for different pipelines. \n\nEach step in the pipeline can be organized like this:\n\n```python\ndef filter_signal_with_median(window=5, **kwargs):\n    signal = kwargs['signal']\n    filtered_signal = median_filter(signal, size=window)\n    kwargs['signal'] = filtered_signal\n    return kwargs\n```\n\n## Transit phase detection\n\nAfter experimenting with the [BATMAN](https://lkreidberg.github.io/batman/docs/html/index.html) library, we found that there are several possible structures of the transit flux signal, even when fully denoised and detrended. Limb darkening can influence the calculation of transit depth, affecting how we calculate the transit edges\n\n```python\n def get_transit_phases2(**kwargs):\n    signal = kwargs['signal'] + 1\n\n    grad1 = np.gradient(signal)\n    mn = np.argmin(grad1)\n    mx = np.argmax(grad1)\n    \n    peaks = find_peaks(grad1, width=25)[0]\n    ingress_left = peaks[(peaks < mn) & (mn - peaks <= 250)][-1] if np.any((peaks < mn) & (mn - peaks <= 250)) else mn-110\n    ingress_right = peaks[(peaks > mn) & (peaks - mn <= 250)][0] if np.any((peaks > mn) & (peaks - mn <= 250)) else mn+110\n    \n    dips = find_peaks(-grad1, width=25)[0]\n    egress_left = dips[(dips < mx) & (mx - dips <= 250)][-1] if np.any((dips < mx) & (mx - dips <= 250)) else mx-110\n    egress_right = dips[(dips > mx) & (dips - mx <= 250)][0] if np.any((dips > mx) & (dips - mx <= 250)) else mx+110\n\n    kwargs['ingress_left'] = ingress_left\n    kwargs['ingress_right'] = ingress_right\n    kwargs['egress_left'] = egress_left\n    kwargs['egress_right'] = egress_right\n\n    return kwargs\n```\n\n### Features calculations\n\nKnowing the *ingress* and *egress* of the transit, we fit a polynomial into the flux signal (ignoring the transit) and then perform detrending.\n\nThen we can assume that once the plain detrended signal is on the level of constant 1, then we can derive a transit depth by the formula:\n\n```python\ndef transit_formula(s):\n    return 1/s - 1\n```\nHere `s` can be a mean flux in the transit, flux at the middle point, etc. We collect quite a few here. This was a huge speed improvement compared to the `minimize` method. \n\nAdditionally, we collect generic features on the in-transit flux, like the standard deviation.\n\n## Models training\n\nWe have generated 500 Ridge models and 500 PLSRegression models that formed the ensemble. \n\nAlpha for Ridge was random:\n```python\nalpha = 10 ** (2 - 3 * np.random.rand())\n```\n\nThe same as `n_components` for PLSRegression:\n```python\nmodel = PLSRegression(n_components=n_components, tol=1e-8)\n```\n\nWe used different subsets of the training data for each model:\n```python\nsubset_size = int(X_train.shape[0] * np.random.uniform(0.5, 1))\n```\n\nWe sorted the model predictions for each wavelength, removed the lowest and highest 30%, and took the mean of the remaining predictions.\n\n### Sigma calculation\n\nWe calculated the standard deviation of the sorted predictions for each wavelength:  \n```python\nstd_preds = np.std(sorted_preds, axis=0)\n```\n\nWe then multiplied the standard deviation by a coefficient that worked best in LB:\n```python\nCOEFF_0 = 4 # This one worked the best\nsigma_response = sigma * COEFF_0\n```\n\n## What worked well, though didn't make it to the final submission\n1. Usage of wavelets (`pywt`) - our favorite one was `db29`\n2. `median_filter`\n3. `savgol_filter`\n\n## What didn't work\n1. Gaussian Process Regression - we tried different kernels here, but didn't see breakthroughs while a big increase of the processing time\n2. LinearRegresion, SVM, LGBM, CatBoost, RandomForest, XGBoost... \n3. BayesianRidge - worked well, though long time (and didn't fit into the memory for more features)\n4. We tried to train a neural network with a loss function for both sigma and transit depth (the competition score), but didn't have much luck due to the limited time\n\n## Improvements We Lacked Time For\n1. Generate the train data with BATMAN  \n2. Use FGS data\n3. Use data for real molecules to generate additional training data\n4. Collect more meaningful features (and filter out those which don't help)\n5. Sigma approaches - we didn't use all the potential here\n\nPlease feel free to explore our [best solution](https://www.kaggle.com/code/thegrey/ad24-eleventh?scriptVersionId=204479373)."
  }
}