{
  "id": 543983,
  "title": "9th place solution",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543983",
  "author_name": "Arseny Poyda",
  "post_date": "2024-11-02T16:19:55.587000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I want to thank the hosts and Kaggle team for this competition.</p>\n<h1>Data Preprocessing</h1>\n<ul>\n<li>Signal calibration is the same as the host's one, except binning and hot-pixels handling are omitted.</li>\n<li>Signal splitting is based on the first and second derivatives like in <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543944\" target=\"_blank\">3rd place solution</a></li>\n</ul>\n<h1>Spectrum (μ and σ) Prediction</h1>\n<p>The prediction procedure can be divided into 3 parts:</p>\n<ul>\n<li>coarse estimation of μ for different chunks of wavelengths;</li>\n<li>σ estimation;</li>\n<li>μ refinement based on σ.</li>\n</ul>\n<h3>Coarse Estimation of μ</h3>\n<ul>\n<li>μ is estimated as the relative drop of the signal during transit. I used two approaches to calculate μ. The first one is a slightly modified version of the polynomial fit method proposed by <a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a>. The second approach is to take the difference between the signals around the fall region and divide it by the higher one. Since the signal can include a low-frequency trend, it is important to consider only small regions (used 90 timestamps). The final μ is the weighted combination of these two approaches.</li>\n<li>The estimate of mean μ_1 is calulated over the entire 282 wavelengths (1 big chunk). The estimate of μ_47 for individual wavelengths is obtained by dividing the wavelengths into 47 chunks (each contains 6 pixels).</li>\n<li>In both cases the signal is filtered by Butterworth over time. Moreover, in the case of 47 chunks, the signal is additionally filtered by Hann window (1D convolution) over wavelengths.</li>\n</ul>\n<h3>Estimation of σ</h3>\n<ul>\n<li>The approach is increadibly simple. Calculate μ_4 (4 μs for 4 chunks, each with ~70 pixels) and just take the std() over these 4 values: σ = μ_4.std().</li>\n</ul>\n<h3>Refinement of μ</h3>\n<ul>\n<li>Despite the filtering over wavelengths, μ_47 still has extreme deviation. Thus it is essential to filter the obtained μs again .</li>\n<li>The final prediction of μ is the weighted combination of the mean μ_1 (1 big chunk) and μ_47 (47 chunks). The important find is the greater the TRUE σ, the better μ_47 approximates the TRUE μ. So the weight w_47 of μ_47 grows monotonically with estimated σ.</li>\n</ul>\n<p>In this discussion I have mentioned only the main details and omitted many small features.<br>\nThe inference code: <a href=\"https://www.kaggle.com/code/arsenypoyda/ariel-inference-9th-place\" target=\"_blank\">https://www.kaggle.com/code/arsenypoyda/ariel-inference-9th-place</a></p>",
  "messages": [
    {
      "id": 3034777,
      "postDate": "2024-11-02T16:19:55.587Z",
      "content": "<p>First of all, I want to thank the hosts and Kaggle team for this competition.</p>\n<h1>Data Preprocessing</h1>\n<ul>\n<li>Signal calibration is the same as the host's one, except binning and hot-pixels handling are omitted.</li>\n<li>Signal splitting is based on the first and second derivatives like in <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543944\" target=\"_blank\">3rd place solution</a></li>\n</ul>\n<h1>Spectrum (μ and σ) Prediction</h1>\n<p>The prediction procedure can be divided into 3 parts:</p>\n<ul>\n<li>coarse estimation of μ for different chunks of wavelengths;</li>\n<li>σ estimation;</li>\n<li>μ refinement based on σ.</li>\n</ul>\n<h3>Coarse Estimation of μ</h3>\n<ul>\n<li>μ is estimated as the relative drop of the signal during transit. I used two approaches to calculate μ. The first one is a slightly modified version of the polynomial fit method proposed by <a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a>. The second approach is to take the difference between the signals around the fall region and divide it by the higher one. Since the signal can include a low-frequency trend, it is important to consider only small regions (used 90 timestamps). The final μ is the weighted combination of these two approaches.</li>\n<li>The estimate of mean μ_1 is calulated over the entire 282 wavelengths (1 big chunk). The estimate of μ_47 for individual wavelengths is obtained by dividing the wavelengths into 47 chunks (each contains 6 pixels).</li>\n<li>In both cases the signal is filtered by Butterworth over time. Moreover, in the case of 47 chunks, the signal is additionally filtered by Hann window (1D convolution) over wavelengths.</li>\n</ul>\n<h3>Estimation of σ</h3>\n<ul>\n<li>The approach is increadibly simple. Calculate μ_4 (4 μs for 4 chunks, each with ~70 pixels) and just take the std() over these 4 values: σ = μ_4.std().</li>\n</ul>\n<h3>Refinement of μ</h3>\n<ul>\n<li>Despite the filtering over wavelengths, μ_47 still has extreme deviation. Thus it is essential to filter the obtained μs again .</li>\n<li>The final prediction of μ is the weighted combination of the mean μ_1 (1 big chunk) and μ_47 (47 chunks). The important find is the greater the TRUE σ, the better μ_47 approximates the TRUE μ. So the weight w_47 of μ_47 grows monotonically with estimated σ.</li>\n</ul>\n<p>In this discussion I have mentioned only the main details and omitted many small features.<br>\nThe inference code: <a href=\"https://www.kaggle.com/code/arsenypoyda/ariel-inference-9th-place\" target=\"_blank\">https://www.kaggle.com/code/arsenypoyda/ariel-inference-9th-place</a></p>",
      "rawMarkdown": "First of all, I want to thank the hosts and Kaggle team for this competition.\n\n# Data Preprocessing\n- Signal calibration is the same as the host's one, except binning and hot-pixels handling are omitted.\n- Signal splitting is based on the first and second derivatives like in [3rd place solution](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543944)\n\n# Spectrum (μ and σ) Prediction\nThe prediction procedure can be divided into 3 parts:\n- coarse estimation of μ for different chunks of wavelengths;\n- σ estimation;\n- μ refinement based on σ.\n\n### Coarse Estimation of μ\n- μ is estimated as the relative drop of the signal during transit. I used two approaches to calculate μ. The first one is a slightly modified version of the polynomial fit method proposed by @sergeifironov. The second approach is to take the difference between the signals around the fall region and divide it by the higher one. Since the signal can include a low-frequency trend, it is important to consider only small regions (used 90 timestamps). The final μ is the weighted combination of these two approaches.\n- The estimate of mean μ_1 is calulated over the entire 282 wavelengths (1 big chunk). The estimate of μ_47 for individual wavelengths is obtained by dividing the wavelengths into 47 chunks (each contains 6 pixels).\n- In both cases the signal is filtered by Butterworth over time. Moreover, in the case of 47 chunks, the signal is additionally filtered by Hann window (1D convolution) over wavelengths.\n\n### Estimation of σ\n- The approach is increadibly simple. Calculate μ_4 (4 μs for 4 chunks, each with ~70 pixels) and just take the std() over these 4 values: σ = μ_4.std().\n\n### Refinement of μ\n- Despite the filtering over wavelengths, μ_47 still has extreme deviation. Thus it is essential to filter the obtained μs again .\n- The final prediction of μ is the weighted combination of the mean μ_1 (1 big chunk) and μ_47 (47 chunks). The important find is the greater the TRUE σ, the better μ_47 approximates the TRUE μ. So the weight w_47 of μ_47 grows monotonically with estimated σ.\n\nIn this discussion I have mentioned only the main details and omitted many small features.\nThe inference code: https://www.kaggle.com/code/arsenypoyda/ariel-inference-9th-place",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3034777": "First of all, I want to thank the hosts and Kaggle team for this competition.\n\n# Data Preprocessing\n- Signal calibration is the same as the host's one, except binning and hot-pixels handling are omitted.\n- Signal splitting is based on the first and second derivatives like in [3rd place solution](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/543944)\n\n# Spectrum (μ and σ) Prediction\nThe prediction procedure can be divided into 3 parts:\n- coarse estimation of μ for different chunks of wavelengths;\n- σ estimation;\n- μ refinement based on σ.\n\n### Coarse Estimation of μ\n- μ is estimated as the relative drop of the signal during transit. I used two approaches to calculate μ. The first one is a slightly modified version of the polynomial fit method proposed by @sergeifironov. The second approach is to take the difference between the signals around the fall region and divide it by the higher one. Since the signal can include a low-frequency trend, it is important to consider only small regions (used 90 timestamps). The final μ is the weighted combination of these two approaches.\n- The estimate of mean μ_1 is calulated over the entire 282 wavelengths (1 big chunk). The estimate of μ_47 for individual wavelengths is obtained by dividing the wavelengths into 47 chunks (each contains 6 pixels).\n- In both cases the signal is filtered by Butterworth over time. Moreover, in the case of 47 chunks, the signal is additionally filtered by Hann window (1D convolution) over wavelengths.\n\n### Estimation of σ\n- The approach is increadibly simple. Calculate μ_4 (4 μs for 4 chunks, each with ~70 pixels) and just take the std() over these 4 values: σ = μ_4.std().\n\n### Refinement of μ\n- Despite the filtering over wavelengths, μ_47 still has extreme deviation. Thus it is essential to filter the obtained μs again .\n- The final prediction of μ is the weighted combination of the mean μ_1 (1 big chunk) and μ_47 (47 chunks). The important find is the greater the TRUE σ, the better μ_47 approximates the TRUE μ. So the weight w_47 of μ_47 grows monotonically with estimated σ.\n\nIn this discussion I have mentioned only the main details and omitted many small features.\nThe inference code: https://www.kaggle.com/code/arsenypoyda/ariel-inference-9th-place"
  }
}