{
  "id": 543823,
  "title": "74th place solution",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543823",
  "author_name": "FabienDaniel",
  "post_date": "2024-11-01T15:50:19.221000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First, thanks to the host for hosting such a great competition, and being so present in the forums ! And to all the paricipants who shared their insights and their code.</p>\n<p>My solution follows the line that others shared during this competition and largely beneficiated from the work shared by others. In no particular order, I principally made use of :</p>\n<ul>\n<li>the <a href=\"https://www.kaggle.com/code/lblhandsome/30min-faster-data-processing\" target=\"_blank\">cupy preprocessing</a> shared by <a href=\"https://www.kaggle.com/lblhandsome\" target=\"_blank\">@lblhandsome</a> </li>\n<li>the <a href=\"https://www.kaggle.com/code/sergeifironov/ariel-only-correlation\" target=\"_blank\">ratio estimate based on polynomials</a> developed by <a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a> </li>\n<li>the <a href=\"https://www.kaggle.com/code/gordonyip/host-starter-solution\" target=\"_blank\">2D CNN</a> shared by the host, that I trancsripted to pytorch</li>\n</ul>\n<p>I first adapted this work to develop a single mean estimate, with a <em>sigma</em> value constant for each star: <a href=\"https://www.kaggle.com/code/fabiendaniel/adc24-polyfit-inference\" target=\"_blank\">mean flux prediction kernel</a> ; where I aggregated multiple estimate for the mean value, based on linear regression models with different sets of input features.</p>\n<p>In a second time, I trained a 2D CNN that aimed at predicting the spectral signature of the exoplanet, normalised w.r.t. the mean flux: <a href=\"https://www.kaggle.com/code/fabiendaniel/adc24-polyfit-2d-cnn-inference\" target=\"_blank\">mean + spectral signature prediction</a>.</p>\n<p>Overall, the fact that I could not find a good way of estimating <em>sigma</em> stopped me, mainly because the best estimate I obtained on the trianing set were a factor ~3 different from what I could guess for the test set by LB probing. Since sigma had such a huge impact on the final score, I just stopped there without further insights. </p>",
  "messages": [
    {
      "id": 3033897,
      "postDate": "2024-11-01T15:50:19.220Z",
      "content": "<p>First, thanks to the host for hosting such a great competition, and being so present in the forums ! And to all the paricipants who shared their insights and their code.</p>\n<p>My solution follows the line that others shared during this competition and largely beneficiated from the work shared by others. In no particular order, I principally made use of :</p>\n<ul>\n<li>the <a href=\"https://www.kaggle.com/code/lblhandsome/30min-faster-data-processing\" target=\"_blank\">cupy preprocessing</a> shared by <a href=\"https://www.kaggle.com/lblhandsome\" target=\"_blank\">@lblhandsome</a> </li>\n<li>the <a href=\"https://www.kaggle.com/code/sergeifironov/ariel-only-correlation\" target=\"_blank\">ratio estimate based on polynomials</a> developed by <a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a> </li>\n<li>the <a href=\"https://www.kaggle.com/code/gordonyip/host-starter-solution\" target=\"_blank\">2D CNN</a> shared by the host, that I trancsripted to pytorch</li>\n</ul>\n<p>I first adapted this work to develop a single mean estimate, with a <em>sigma</em> value constant for each star: <a href=\"https://www.kaggle.com/code/fabiendaniel/adc24-polyfit-inference\" target=\"_blank\">mean flux prediction kernel</a> ; where I aggregated multiple estimate for the mean value, based on linear regression models with different sets of input features.</p>\n<p>In a second time, I trained a 2D CNN that aimed at predicting the spectral signature of the exoplanet, normalised w.r.t. the mean flux: <a href=\"https://www.kaggle.com/code/fabiendaniel/adc24-polyfit-2d-cnn-inference\" target=\"_blank\">mean + spectral signature prediction</a>.</p>\n<p>Overall, the fact that I could not find a good way of estimating <em>sigma</em> stopped me, mainly because the best estimate I obtained on the trianing set were a factor ~3 different from what I could guess for the test set by LB probing. Since sigma had such a huge impact on the final score, I just stopped there without further insights. </p>",
      "rawMarkdown": "First, thanks to the host for hosting such a great competition, and being so present in the forums ! And to all the paricipants who shared their insights and their code.\n\nMy solution follows the line that others shared during this competition and largely beneficiated from the work shared by others. In no particular order, I principally made use of :\n- the [cupy preprocessing](https://www.kaggle.com/code/lblhandsome/30min-faster-data-processing) shared by @lblhandsome \n- the [ratio estimate based on polynomials](https://www.kaggle.com/code/sergeifironov/ariel-only-correlation) developed by @sergeifironov \n- the [2D CNN](https://www.kaggle.com/code/gordonyip/host-starter-solution) shared by the host, that I trancsripted to pytorch\n\nI first adapted this work to develop a single mean estimate, with a *sigma* value constant for each star: [mean flux prediction kernel](https://www.kaggle.com/code/fabiendaniel/adc24-polyfit-inference) ; where I aggregated multiple estimate for the mean value, based on linear regression models with different sets of input features.\n\nIn a second time, I trained a 2D CNN that aimed at predicting the spectral signature of the exoplanet, normalised w.r.t. the mean flux: [mean + spectral signature prediction](https://www.kaggle.com/code/fabiendaniel/adc24-polyfit-2d-cnn-inference).\n\nOverall, the fact that I could not find a good way of estimating *sigma* stopped me, mainly because the best estimate I obtained on the trianing set were a factor ~3 different from what I could guess for the test set by LB probing. Since sigma had such a huge impact on the final score, I just stopped there without further insights. \n\n",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3033897": "First, thanks to the host for hosting such a great competition, and being so present in the forums ! And to all the paricipants who shared their insights and their code.\n\nMy solution follows the line that others shared during this competition and largely beneficiated from the work shared by others. In no particular order, I principally made use of :\n- the [cupy preprocessing](https://www.kaggle.com/code/lblhandsome/30min-faster-data-processing) shared by @lblhandsome \n- the [ratio estimate based on polynomials](https://www.kaggle.com/code/sergeifironov/ariel-only-correlation) developed by @sergeifironov \n- the [2D CNN](https://www.kaggle.com/code/gordonyip/host-starter-solution) shared by the host, that I trancsripted to pytorch\n\nI first adapted this work to develop a single mean estimate, with a *sigma* value constant for each star: [mean flux prediction kernel](https://www.kaggle.com/code/fabiendaniel/adc24-polyfit-inference) ; where I aggregated multiple estimate for the mean value, based on linear regression models with different sets of input features.\n\nIn a second time, I trained a 2D CNN that aimed at predicting the spectral signature of the exoplanet, normalised w.r.t. the mean flux: [mean + spectral signature prediction](https://www.kaggle.com/code/fabiendaniel/adc24-polyfit-2d-cnn-inference).\n\nOverall, the fact that I could not find a good way of estimating *sigma* stopped me, mainly because the best estimate I obtained on the trianing set were a factor ~3 different from what I could guess for the test set by LB probing. Since sigma had such a huge impact on the final score, I just stopped there without further insights. \n\n"
  }
}