{
  "id": 543776,
  "title": "8th place solution",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543776",
  "author_name": "Ruby",
  "post_date": "2024-11-01T12:25:31.366000",
  "votes": 19,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Thanks Kaggle and competition host for this interesting competition and also specials thanks to sergeifironov for the great starter notebook, I basically follow the same way and add some details.</p>\n<p><strong>Data Processing</strong></p>\n<ul>\n<li>almost same as public notebooks, remove two more dimensions for last 60 wavelengths</li>\n</ul>\n<p><strong>Phase Split</strong></p>\n<ul>\n<li>detrend by polynomial fitting with boundary points</li>\n<li>find gap center with accumulated gradient</li>\n<li>find exact gap edge based on monotonic condition</li>\n</ul>\n<p><strong>model1</strong></p>\n<ul>\n<li>transit_depth=(oot-it)/oot, transit_depth is coefficient before oot term in linear equation, so this can be solved by OLS directly</li>\n<li>fit the average of all wavelengths</li>\n<li>fit by quantile regression first and remove points with large residual then fit by OLS. Same points will also be removed in model2 without running quantile regression again. This makes model robust to unclear phase edges.</li>\n</ul>\n<p><strong>model2</strong></p>\n<ul>\n<li>apply svd to all wavelengths and reconstruct them with principal components, different wavelengths are weighted based on SNR before svd</li>\n<li>exponential smoothing over nearby wavelengths, different wavelengths are further weighted based on SNR</li>\n<li>past two operations only apply to ch0</li>\n<li>fit all wavelengths independently with same way in model1</li>\n<li>coefficient covariance based on coefficient standard error of OLS and residual correlation</li>\n</ul>\n<p><strong>ensemble</strong></p>\n<ul>\n<li>basic models are run on data with bins 12*10, 10, then I average 10 results with shifted start index. When bin size becomes smaller the error in variables start to have significant impact</li>\n</ul>\n<p><strong>Post-Process</strong></p>\n<ul>\n<li>apply svd to model2 coefficient to get first k ( 3/4 is best for train/test set, since test molecules are super set of training collection) principal components, then reconstruct coefficients by GLS with covariance shrinkage, the covariance of model2 are transformed based on OLS since simple GLS covariance seems underestimate the actual variance.</li>\n<li>base_coef=0.6<em>coef1+0.4</em>coef2</li>\n<li>base_coef*mean(coef2)/mean(coef1) to deals with bias caused by imbalanced energy distribution across wavelengths</li>\n<li>replace base_coef with coef2 when significant difference exists between coef1 and coef2, this is done in planet-wavelength level and planet level.</li>\n<li>sigma for base_coef: max(|coef1-coef2|, coef_std2) Imagine we ’observe’ highly biased model1 from less biased model2. Sigma for coef1 is bounded by average gap between coef1 and coef2 globally, this works when accurate coef2 estimation is infeasible</li>\n<li>sigma for coef2: (2*coef_std+max coef gap between nearing wavelength), since model2 is still biased and there exists unknown bias in data. I extend coef_std2 based on some heuristic rules</li>\n<li>I add some rules to capture cases when the models might fail, and replace by default results, but seems they never work on training set and public test set, not sure if it worked in private part.</li>\n</ul>\n<p><strong>Reproduce</strong><br>\n<a href=\"https://www.kaggle.com/code/w5833946/adc-final-reproduce\" target=\"_blank\">https://www.kaggle.com/code/w5833946/adc-final-reproduce</a><br>\nI also run it with training set in ver1.</p>",
  "messages": [
    {
      "id": 3033677,
      "postDate": "2024-11-01T12:25:31.367Z",
      "content": "<p>Thanks Kaggle and competition host for this interesting competition and also specials thanks to sergeifironov for the great starter notebook, I basically follow the same way and add some details.</p>\n<p><strong>Data Processing</strong></p>\n<ul>\n<li>almost same as public notebooks, remove two more dimensions for last 60 wavelengths</li>\n</ul>\n<p><strong>Phase Split</strong></p>\n<ul>\n<li>detrend by polynomial fitting with boundary points</li>\n<li>find gap center with accumulated gradient</li>\n<li>find exact gap edge based on monotonic condition</li>\n</ul>\n<p><strong>model1</strong></p>\n<ul>\n<li>transit_depth=(oot-it)/oot, transit_depth is coefficient before oot term in linear equation, so this can be solved by OLS directly</li>\n<li>fit the average of all wavelengths</li>\n<li>fit by quantile regression first and remove points with large residual then fit by OLS. Same points will also be removed in model2 without running quantile regression again. This makes model robust to unclear phase edges.</li>\n</ul>\n<p><strong>model2</strong></p>\n<ul>\n<li>apply svd to all wavelengths and reconstruct them with principal components, different wavelengths are weighted based on SNR before svd</li>\n<li>exponential smoothing over nearby wavelengths, different wavelengths are further weighted based on SNR</li>\n<li>past two operations only apply to ch0</li>\n<li>fit all wavelengths independently with same way in model1</li>\n<li>coefficient covariance based on coefficient standard error of OLS and residual correlation</li>\n</ul>\n<p><strong>ensemble</strong></p>\n<ul>\n<li>basic models are run on data with bins 12*10, 10, then I average 10 results with shifted start index. When bin size becomes smaller the error in variables start to have significant impact</li>\n</ul>\n<p><strong>Post-Process</strong></p>\n<ul>\n<li>apply svd to model2 coefficient to get first k ( 3/4 is best for train/test set, since test molecules are super set of training collection) principal components, then reconstruct coefficients by GLS with covariance shrinkage, the covariance of model2 are transformed based on OLS since simple GLS covariance seems underestimate the actual variance.</li>\n<li>base_coef=0.6<em>coef1+0.4</em>coef2</li>\n<li>base_coef*mean(coef2)/mean(coef1) to deals with bias caused by imbalanced energy distribution across wavelengths</li>\n<li>replace base_coef with coef2 when significant difference exists between coef1 and coef2, this is done in planet-wavelength level and planet level.</li>\n<li>sigma for base_coef: max(|coef1-coef2|, coef_std2) Imagine we ’observe’ highly biased model1 from less biased model2. Sigma for coef1 is bounded by average gap between coef1 and coef2 globally, this works when accurate coef2 estimation is infeasible</li>\n<li>sigma for coef2: (2*coef_std+max coef gap between nearing wavelength), since model2 is still biased and there exists unknown bias in data. I extend coef_std2 based on some heuristic rules</li>\n<li>I add some rules to capture cases when the models might fail, and replace by default results, but seems they never work on training set and public test set, not sure if it worked in private part.</li>\n</ul>\n<p><strong>Reproduce</strong><br>\n<a href=\"https://www.kaggle.com/code/w5833946/adc-final-reproduce\" target=\"_blank\">https://www.kaggle.com/code/w5833946/adc-final-reproduce</a><br>\nI also run it with training set in ver1.</p>",
      "rawMarkdown": "Thanks Kaggle and competition host for this interesting competition and also specials thanks to sergeifironov for the great starter notebook, I basically follow the same way and add some details.\n\n**Data Processing**\n- almost same as public notebooks, remove two more dimensions for last 60 wavelengths\n\n**Phase Split**\n- detrend by polynomial fitting with boundary points\n- find gap center with accumulated gradient\n- find exact gap edge based on monotonic condition\n\n**model1**\n- transit_depth=(oot-it)/oot, transit_depth is coefficient before oot term in linear equation, so this can be solved by OLS directly\n- fit the average of all wavelengths\n- fit by quantile regression first and remove points with large residual then fit by OLS. Same points will also be removed in model2 without running quantile regression again. This makes model robust to unclear phase edges.\n\n**model2**\n- apply svd to all wavelengths and reconstruct them with principal components, different wavelengths are weighted based on SNR before svd\n- exponential smoothing over nearby wavelengths, different wavelengths are further weighted based on SNR\n- past two operations only apply to ch0\n- fit all wavelengths independently with same way in model1\n- coefficient covariance based on coefficient standard error of OLS and residual correlation\n\n**ensemble**\n- basic models are run on data with bins 12*10, 10, then I average 10 results with shifted start index. When bin size becomes smaller the error in variables start to have significant impact\n\n**Post-Process**\n- apply svd to model2 coefficient to get first k ( 3/4 is best for train/test set, since test molecules are super set of training collection) principal components, then reconstruct coefficients by GLS with covariance shrinkage, the covariance of model2 are transformed based on OLS since simple GLS covariance seems underestimate the actual variance.\n- base_coef=0.6*coef1+0.4*coef2\n- base_coef*mean(coef2)/mean(coef1) to deals with bias caused by imbalanced energy distribution across wavelengths\n- replace base_coef with coef2 when significant difference exists between coef1 and coef2, this is done in planet-wavelength level and planet level.\n- sigma for base_coef: max(|coef1-coef2|, coef_std2) Imagine we ’observe’ highly biased model1 from less biased model2. Sigma for coef1 is bounded by average gap between coef1 and coef2 globally, this works when accurate coef2 estimation is infeasible\n- sigma for coef2: (2*coef_std+max coef gap between nearing wavelength), since model2 is still biased and there exists unknown bias in data. I extend coef_std2 based on some heuristic rules\n- I add some rules to capture cases when the models might fail, and replace by default results, but seems they never work on training set and public test set, not sure if it worked in private part.\n\n**Reproduce**\nhttps://www.kaggle.com/code/w5833946/adc-final-reproduce\nI also run it with training set in ver1.\n",
      "votes": 19
    },
    {
      "id": 3034373,
      "postDate": "2024-11-02T04:41:53.043Z",
      "content": "<p>Congrats Brother 👍🏻</p>",
      "rawMarkdown": "Congrats Brother 👍🏻",
      "votes": 1
    },
    {
      "id": 3034270,
      "postDate": "2024-11-02T00:21:21.510Z",
      "content": "<p>Congratulations on your impressive 8th place finish, Ruby! Your solution showcases a solid understanding of data processing and modeling techniques. I find your use of SVD as a denoising method particularly interesting—it's a clever way to handle the noise while still extracting meaningful features.</p>\n<p>It’s fascinating that your post-processing step generated principal components that align with true labels. This highlights the potential for SVD not just as a denoising tool, but also for revealing underlying structures in the data.</p>\n<p>I'm curious, did you encounter any challenges when implementing the polynomial fitting for detrending? It would be great to hear more about how you fine-tuned that aspect, especially given its importance in your overall modeling approach!</p>",
      "rawMarkdown": "Congratulations on your impressive 8th place finish, Ruby! Your solution showcases a solid understanding of data processing and modeling techniques. I find your use of SVD as a denoising method particularly interesting—it's a clever way to handle the noise while still extracting meaningful features.\n\nIt’s fascinating that your post-processing step generated principal components that align with true labels. This highlights the potential for SVD not just as a denoising tool, but also for revealing underlying structures in the data.\n\nI'm curious, did you encounter any challenges when implementing the polynomial fitting for detrending? It would be great to hear more about how you fine-tuned that aspect, especially given its importance in your overall modeling approach!",
      "votes": 1,
      "replies": [
        {
          "id": 3034370,
          "postDate": "2024-11-02T04:37:54.207Z",
          "content": "<p>Do you mean the detrending in phase split part? It is not very sensitive, infact just gradient accumulation can work for most planets even without detrending, and the detrending is not applied to data used for later stage regression. I didn't spend time turing hyperparameters here.</p>",
          "rawMarkdown": "Do you mean the detrending in phase split part? It is not very sensitive, infact just gradient accumulation can work for most planets even without detrending, and the detrending is not applied to data used for later stage regression. I didn't spend time turing hyperparameters here."
        }
      ]
    },
    {
      "id": 3033856,
      "postDate": "2024-11-01T15:14:00.303Z",
      "content": "<p>Congratulation, Ruby. Some tricks are incredible, especially for your postprocess. </p>\n<blockquote>\n  <p>apply svd to all wavelengths and reconstruct them with principal components, different wavelengths are weighted based on SNR before svd</p>\n</blockquote>\n<p>You want to analyze the atmospheric components behind the wavelengths？</p>",
      "rawMarkdown": "Congratulation, Ruby. Some tricks are incredible, especially for your postprocess. \n\n>apply svd to all wavelengths and reconstruct them with principal components, different wavelengths are weighted based on SNR before svd\n\nYou want to analyze the atmospheric components behind the wavelengths？",
      "votes": 1,
      "replies": [
        {
          "id": 3033881,
          "postDate": "2024-11-01T15:35:40.873Z",
          "content": "<p>This step I just use svd as a general denoise method like what is done implicitly in ridge.But svd in postprocess step indeed generate main components that similar to those directly extracted from true labels.</p>",
          "rawMarkdown": "This step I just use svd as a general denoise method like what is done implicitly in ridge.But svd in postprocess step indeed generate main components that similar to those directly extracted from true labels."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3034373,
      "author_name": "Sumit_08",
      "author_url": "",
      "post_date": "2024-11-02T04:41:53.043000",
      "content": "<p>Congrats Brother 👍🏻</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3034270,
      "author_name": "yiyao6442",
      "author_url": "",
      "post_date": "2024-11-02T00:21:21.510000",
      "content": "<p>Congratulations on your impressive 8th place finish, Ruby! Your solution showcases a solid understanding of data processing and modeling techniques. I find your use of SVD as a denoising method particularly interesting—it's a clever way to handle the noise while still extracting meaningful features.</p>\n<p>It’s fascinating that your post-processing step generated principal components that align with true labels. This highlights the potential for SVD not just as a denoising tool, but also for revealing underlying structures in the data.</p>\n<p>I'm curious, did you encounter any challenges when implementing the polynomial fitting for detrending? It would be great to hear more about how you fine-tuned that aspect, especially given its importance in your overall modeling approach!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3034370,
          "author_name": "Ruby",
          "author_url": "",
          "post_date": "2024-11-02T04:37:54.207000",
          "content": "<p>Do you mean the detrending in phase split part? It is not very sensitive, infact just gradient accumulation can work for most planets even without detrending, and the detrending is not applied to data used for later stage regression. I didn't spend time turing hyperparameters here.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3033856,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2024-11-01T15:14:00.303000",
      "content": "<p>Congratulation, Ruby. Some tricks are incredible, especially for your postprocess. </p>\n<blockquote>\n  <p>apply svd to all wavelengths and reconstruct them with principal components, different wavelengths are weighted based on SNR before svd</p>\n</blockquote>\n<p>You want to analyze the atmospheric components behind the wavelengths？</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3033881,
          "author_name": "Ruby",
          "author_url": "",
          "post_date": "2024-11-01T15:35:40.873000",
          "content": "<p>This step I just use svd as a general denoise method like what is done implicitly in ridge.But svd in postprocess step indeed generate main components that similar to those directly extracted from true labels.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3033677": "Thanks Kaggle and competition host for this interesting competition and also specials thanks to sergeifironov for the great starter notebook, I basically follow the same way and add some details.\n\n**Data Processing**\n- almost same as public notebooks, remove two more dimensions for last 60 wavelengths\n\n**Phase Split**\n- detrend by polynomial fitting with boundary points\n- find gap center with accumulated gradient\n- find exact gap edge based on monotonic condition\n\n**model1**\n- transit_depth=(oot-it)/oot, transit_depth is coefficient before oot term in linear equation, so this can be solved by OLS directly\n- fit the average of all wavelengths\n- fit by quantile regression first and remove points with large residual then fit by OLS. Same points will also be removed in model2 without running quantile regression again. This makes model robust to unclear phase edges.\n\n**model2**\n- apply svd to all wavelengths and reconstruct them with principal components, different wavelengths are weighted based on SNR before svd\n- exponential smoothing over nearby wavelengths, different wavelengths are further weighted based on SNR\n- past two operations only apply to ch0\n- fit all wavelengths independently with same way in model1\n- coefficient covariance based on coefficient standard error of OLS and residual correlation\n\n**ensemble**\n- basic models are run on data with bins 12*10, 10, then I average 10 results with shifted start index. When bin size becomes smaller the error in variables start to have significant impact\n\n**Post-Process**\n- apply svd to model2 coefficient to get first k ( 3/4 is best for train/test set, since test molecules are super set of training collection) principal components, then reconstruct coefficients by GLS with covariance shrinkage, the covariance of model2 are transformed based on OLS since simple GLS covariance seems underestimate the actual variance.\n- base_coef=0.6*coef1+0.4*coef2\n- base_coef*mean(coef2)/mean(coef1) to deals with bias caused by imbalanced energy distribution across wavelengths\n- replace base_coef with coef2 when significant difference exists between coef1 and coef2, this is done in planet-wavelength level and planet level.\n- sigma for base_coef: max(|coef1-coef2|, coef_std2) Imagine we ’observe’ highly biased model1 from less biased model2. Sigma for coef1 is bounded by average gap between coef1 and coef2 globally, this works when accurate coef2 estimation is infeasible\n- sigma for coef2: (2*coef_std+max coef gap between nearing wavelength), since model2 is still biased and there exists unknown bias in data. I extend coef_std2 based on some heuristic rules\n- I add some rules to capture cases when the models might fail, and replace by default results, but seems they never work on training set and public test set, not sure if it worked in private part.\n\n**Reproduce**\nhttps://www.kaggle.com/code/w5833946/adc-final-reproduce\nI also run it with training set in ver1.\n",
    "3034373": "Congrats Brother 👍🏻",
    "3034270": "Congratulations on your impressive 8th place finish, Ruby! Your solution showcases a solid understanding of data processing and modeling techniques. I find your use of SVD as a denoising method particularly interesting—it's a clever way to handle the noise while still extracting meaningful features.\n\nIt’s fascinating that your post-processing step generated principal components that align with true labels. This highlights the potential for SVD not just as a denoising tool, but also for revealing underlying structures in the data.\n\nI'm curious, did you encounter any challenges when implementing the polynomial fitting for detrending? It would be great to hear more about how you fine-tuned that aspect, especially given its importance in your overall modeling approach!",
    "3033856": "Congratulation, Ruby. Some tricks are incredible, especially for your postprocess. \n\n>apply svd to all wavelengths and reconstruct them with principal components, different wavelengths are weighted based on SNR before svd\n\nYou want to analyze the atmospheric components behind the wavelengths？"
  }
}