{
  "id": 543682,
  "title": "Really Simple Approach to Get 39th Place",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543682",
  "author_name": "moto",
  "post_date": "2024-11-01T00:47:46.908000",
  "votes": 15,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I would like to thank the host for organizing such a wonderful competition, also to the person who shared a fantastic idea in the discussion, and my teammate, kyu999 and yuto.<br>\nWe had only one month to work on this competition, but we enjoyed this competition very much.</p>\n<p>Our solution is relatively simple.<br>\nOur solution consists of wl prediction part and sigma prediction part.</p>\n<h1>WL Prediction Part</h1>\n<p>WL Prediction Part is based on Sergei's method.<br>\nWe achieved LB 0.572 by this method.<br>\nWe expanded the Sergei's method to a multiple wavelength prediction version.<br>\nSimply predicting by the light curve of each wavelength in AIRS is not enough to expand this method due to the noise.<br>\nWe reduced noise by taking the sliding window mean along the wavelength axis in AIRS, which made possible to expand this method.<br>\nFGS1 data seems noisy, so we predicted wl_1 by the mean of light curves of FGS1 and AIRS.</p>\n<h1>Sigma Prediction Part</h1>\n<p>We utilized Random Forest to predict sigma.<br>\nThis method boosted LB 0.572 to 0.607.<br>\nWe wanted to find better approach, but we found this method 2 days before the deadline and we didn't have much time to improve it.:(<br>\nWe thought that models can overfit easily, so we chose Random Forest.<br>\nWe used the following features to predict sigma of i th wavelength.<br>\nNote that q means MAE/mean(signal during transit) whose shape is (num_planet, num_wavelength), s means wl prediction by the Sergei's method whose shape is (num_planet, num_wavelength), and data means light curve signals whose shape is (num_planet, num_time, num_wavelength).</p>\n<ul>\n<li>q[:, i:i+1]</li>\n<li>s[:, i:i+1]</li>\n<li>std(q, axis=1)</li>\n<li>mean(q, axis=1)</li>\n<li>entropy(abs(q) + 1e-8, axis=1)</li>\n<li>std(s, axis=1)</li>\n<li>mean(s, axis=1)</li>\n<li>skew(data[:, :, i], axis=1)</li>\n<li>kurtosis(data[:, :, i], axis=1)</li>\n</ul>\n<p>We simply used 5 fold Stratified KFold for cross validation.</p>\n<h1>Deep Learning Approach(Not Worked)</h1>\n<p>We could not construct good Deep Learning Model better than the LB score 0.531 using CNN.<br>\nValidation score fluctuated very much but making batch size large and adding 1e-8 to the sigma prediction made the training more stable.<br>\nI'm looking forward to the deep learning solution.</p>\n<p>↓our submission code<br>\n<a href=\"https://www.kaggle.com/yamashitamotokazu/39th-place-solution-submission-code\" target=\"_blank\">https://www.kaggle.com/yamashitamotokazu/39th-place-solution-submission-code</a></p>",
  "messages": [
    {
      "id": 3033270,
      "postDate": "2024-11-01T00:47:46.910Z",
      "content": "<p>I would like to thank the host for organizing such a wonderful competition, also to the person who shared a fantastic idea in the discussion, and my teammate, kyu999 and yuto.<br>\nWe had only one month to work on this competition, but we enjoyed this competition very much.</p>\n<p>Our solution is relatively simple.<br>\nOur solution consists of wl prediction part and sigma prediction part.</p>\n<h1>WL Prediction Part</h1>\n<p>WL Prediction Part is based on Sergei's method.<br>\nWe achieved LB 0.572 by this method.<br>\nWe expanded the Sergei's method to a multiple wavelength prediction version.<br>\nSimply predicting by the light curve of each wavelength in AIRS is not enough to expand this method due to the noise.<br>\nWe reduced noise by taking the sliding window mean along the wavelength axis in AIRS, which made possible to expand this method.<br>\nFGS1 data seems noisy, so we predicted wl_1 by the mean of light curves of FGS1 and AIRS.</p>\n<h1>Sigma Prediction Part</h1>\n<p>We utilized Random Forest to predict sigma.<br>\nThis method boosted LB 0.572 to 0.607.<br>\nWe wanted to find better approach, but we found this method 2 days before the deadline and we didn't have much time to improve it.:(<br>\nWe thought that models can overfit easily, so we chose Random Forest.<br>\nWe used the following features to predict sigma of i th wavelength.<br>\nNote that q means MAE/mean(signal during transit) whose shape is (num_planet, num_wavelength), s means wl prediction by the Sergei's method whose shape is (num_planet, num_wavelength), and data means light curve signals whose shape is (num_planet, num_time, num_wavelength).</p>\n<ul>\n<li>q[:, i:i+1]</li>\n<li>s[:, i:i+1]</li>\n<li>std(q, axis=1)</li>\n<li>mean(q, axis=1)</li>\n<li>entropy(abs(q) + 1e-8, axis=1)</li>\n<li>std(s, axis=1)</li>\n<li>mean(s, axis=1)</li>\n<li>skew(data[:, :, i], axis=1)</li>\n<li>kurtosis(data[:, :, i], axis=1)</li>\n</ul>\n<p>We simply used 5 fold Stratified KFold for cross validation.</p>\n<h1>Deep Learning Approach(Not Worked)</h1>\n<p>We could not construct good Deep Learning Model better than the LB score 0.531 using CNN.<br>\nValidation score fluctuated very much but making batch size large and adding 1e-8 to the sigma prediction made the training more stable.<br>\nI'm looking forward to the deep learning solution.</p>\n<p>↓our submission code<br>\n<a href=\"https://www.kaggle.com/yamashitamotokazu/39th-place-solution-submission-code\" target=\"_blank\">https://www.kaggle.com/yamashitamotokazu/39th-place-solution-submission-code</a></p>",
      "rawMarkdown": "I would like to thank the host for organizing such a wonderful competition, also to the person who shared a fantastic idea in the discussion, and my teammate, kyu999 and yuto.\nWe had only one month to work on this competition, but we enjoyed this competition very much.\n\nOur solution is relatively simple.\nOur solution consists of wl prediction part and sigma prediction part.\n\n# WL Prediction Part\nWL Prediction Part is based on Sergei's method.\nWe achieved LB 0.572 by this method.\nWe expanded the Sergei's method to a multiple wavelength prediction version.\nSimply predicting by the light curve of each wavelength in AIRS is not enough to expand this method due to the noise.\nWe reduced noise by taking the sliding window mean along the wavelength axis in AIRS, which made possible to expand this method.\nFGS1 data seems noisy, so we predicted wl_1 by the mean of light curves of FGS1 and AIRS.\n\n\n# Sigma Prediction Part\nWe utilized Random Forest to predict sigma.\nThis method boosted LB 0.572 to 0.607.\nWe wanted to find better approach, but we found this method 2 days before the deadline and we didn't have much time to improve it.:(\nWe thought that models can overfit easily, so we chose Random Forest.\nWe used the following features to predict sigma of i th wavelength.\nNote that q means MAE/mean(signal during transit) whose shape is (num_planet, num_wavelength), s means wl prediction by the Sergei's method whose shape is (num_planet, num_wavelength), and data means light curve signals whose shape is (num_planet, num_time, num_wavelength).\n- q[:, i:i+1]\n- s[:, i:i+1]\n- std(q, axis=1)\n- mean(q, axis=1)\n- entropy(abs(q) + 1e-8, axis=1)\n- std(s, axis=1)\n- mean(s, axis=1)\n- skew(data[:, :, i], axis=1)\n- kurtosis(data[:, :, i], axis=1)\n\nWe simply used 5 fold Stratified KFold for cross validation.\n\n\n# Deep Learning Approach(Not Worked)\nWe could not construct good Deep Learning Model better than the LB score 0.531 using CNN.\nValidation score fluctuated very much but making batch size large and adding 1e-8 to the sigma prediction made the training more stable.\nI'm looking forward to the deep learning solution.\n\n↓our submission code\nhttps://www.kaggle.com/yamashitamotokazu/39th-place-solution-submission-code",
      "votes": 15
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3033270": "I would like to thank the host for organizing such a wonderful competition, also to the person who shared a fantastic idea in the discussion, and my teammate, kyu999 and yuto.\nWe had only one month to work on this competition, but we enjoyed this competition very much.\n\nOur solution is relatively simple.\nOur solution consists of wl prediction part and sigma prediction part.\n\n# WL Prediction Part\nWL Prediction Part is based on Sergei's method.\nWe achieved LB 0.572 by this method.\nWe expanded the Sergei's method to a multiple wavelength prediction version.\nSimply predicting by the light curve of each wavelength in AIRS is not enough to expand this method due to the noise.\nWe reduced noise by taking the sliding window mean along the wavelength axis in AIRS, which made possible to expand this method.\nFGS1 data seems noisy, so we predicted wl_1 by the mean of light curves of FGS1 and AIRS.\n\n\n# Sigma Prediction Part\nWe utilized Random Forest to predict sigma.\nThis method boosted LB 0.572 to 0.607.\nWe wanted to find better approach, but we found this method 2 days before the deadline and we didn't have much time to improve it.:(\nWe thought that models can overfit easily, so we chose Random Forest.\nWe used the following features to predict sigma of i th wavelength.\nNote that q means MAE/mean(signal during transit) whose shape is (num_planet, num_wavelength), s means wl prediction by the Sergei's method whose shape is (num_planet, num_wavelength), and data means light curve signals whose shape is (num_planet, num_time, num_wavelength).\n- q[:, i:i+1]\n- s[:, i:i+1]\n- std(q, axis=1)\n- mean(q, axis=1)\n- entropy(abs(q) + 1e-8, axis=1)\n- std(s, axis=1)\n- mean(s, axis=1)\n- skew(data[:, :, i], axis=1)\n- kurtosis(data[:, :, i], axis=1)\n\nWe simply used 5 fold Stratified KFold for cross validation.\n\n\n# Deep Learning Approach(Not Worked)\nWe could not construct good Deep Learning Model better than the LB score 0.531 using CNN.\nValidation score fluctuated very much but making batch size large and adding 1e-8 to the sigma prediction made the training more stable.\nI'm looking forward to the deep learning solution.\n\n↓our submission code\nhttps://www.kaggle.com/yamashitamotokazu/39th-place-solution-submission-code"
  }
}