{
  "id": 543812,
  "title": "Adding Two Features -> score: +0.015",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543812",
  "author_name": "Makoto Kine",
  "post_date": "2024-11-01T14:59:14.654000",
  "votes": 4,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I was personally satisfied with my ranking, but compared to other members who shared their ideas, mine was quite simple (I added two features to the data of size (673, 283) = (number of samples, number of wavelengths) and performed a simple Gaussian Process Regression with a linear kernel). <br>\nTherefore, I will focus on what two features I added.</p>\n<p>By simply adding these features, my score increased by +0.015, and since it utilized the FGS signal, I found it personally interesting.</p>\n<p>What I did was add as features the average light intensity average values of the four corners of the FGS1 images and the upper and lower edges of the AIRS-CH0 images for each sample. (yellow zone)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F302d3251623d6772c243fdd77c666dbb%2F2024-11-01%2023.52.04.png?generation=1730472870237724&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F4e835df30fe6390cfa978cba3e782417%2F2024-11-01%2023.52.14.png?generation=1730472885814664&amp;alt=media\" alt=\"\"></p>\n<p>The reason why this feature led to a score improvement is likely because, as shown below, this feature could very clearly cluster the stars to which the planets belong.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F90745bfe5f18ca291635f9fcc03d6683%2F2024-11-01%2023.58.33.png?generation=1730473134199393&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": 3033831,
      "postDate": "2024-11-01T14:59:14.653Z",
      "content": "<p>I was personally satisfied with my ranking, but compared to other members who shared their ideas, mine was quite simple (I added two features to the data of size (673, 283) = (number of samples, number of wavelengths) and performed a simple Gaussian Process Regression with a linear kernel). <br>\nTherefore, I will focus on what two features I added.</p>\n<p>By simply adding these features, my score increased by +0.015, and since it utilized the FGS signal, I found it personally interesting.</p>\n<p>What I did was add as features the average light intensity average values of the four corners of the FGS1 images and the upper and lower edges of the AIRS-CH0 images for each sample. (yellow zone)<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F302d3251623d6772c243fdd77c666dbb%2F2024-11-01%2023.52.04.png?generation=1730472870237724&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F4e835df30fe6390cfa978cba3e782417%2F2024-11-01%2023.52.14.png?generation=1730472885814664&amp;alt=media\" alt=\"\"></p>\n<p>The reason why this feature led to a score improvement is likely because, as shown below, this feature could very clearly cluster the stars to which the planets belong.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F90745bfe5f18ca291635f9fcc03d6683%2F2024-11-01%2023.58.33.png?generation=1730473134199393&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I was personally satisfied with my ranking, but compared to other members who shared their ideas, mine was quite simple (I added two features to the data of size (673, 283) = (number of samples, number of wavelengths) and performed a simple Gaussian Process Regression with a linear kernel). \nTherefore, I will focus on what two features I added.\n\nBy simply adding these features, my score increased by +0.015, and since it utilized the FGS signal, I found it personally interesting.\n\nWhat I did was add as features the average light intensity average values of the four corners of the FGS1 images and the upper and lower edges of the AIRS-CH0 images for each sample. (yellow zone)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F302d3251623d6772c243fdd77c666dbb%2F2024-11-01%2023.52.04.png?generation=1730472870237724&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F4e835df30fe6390cfa978cba3e782417%2F2024-11-01%2023.52.14.png?generation=1730472885814664&alt=media)\n\n\nThe reason why this feature led to a score improvement is likely because, as shown below, this feature could very clearly cluster the stars to which the planets belong.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F90745bfe5f18ca291635f9fcc03d6683%2F2024-11-01%2023.58.33.png?generation=1730473134199393&alt=media)\n\n\n\n",
      "votes": 4
    },
    {
      "id": 3033855,
      "postDate": "2024-11-01T15:13:58.873Z",
      "content": "<p>Interesting. Though the star was given as a primary feature even in test, wouldn't need such a proxy feature. </p>",
      "rawMarkdown": "Interesting. Though the star was given as a primary feature even in test, wouldn't need such a proxy feature. ",
      "votes": 1
    },
    {
      "id": 3034256,
      "postDate": "2024-11-01T23:42:25.310Z",
      "content": "<p>I subtracted the mean of your yellow zone from the signal. I think this is the background light not coming from the star, and because star luminosity is different, it adds differently to the total light for different stars. It is an offset not contributing to the light difference of transit, but contribute to the denominator and I think it reduces 0.5 - 1% of systematic error.</p>\n<p>1 - target_pred = (high - low) / (high - background)</p>\n<p>Nice post! Congratulations for the medal!</p>",
      "rawMarkdown": "I subtracted the mean of your yellow zone from the signal. I think this is the background light not coming from the star, and because star luminosity is different, it adds differently to the total light for different stars. It is an offset not contributing to the light difference of transit, but contribute to the denominator and I think it reduces 0.5 - 1% of systematic error.\n\n1 - target_pred = (high - low) / (high - background)\n\nNice post! Congratulations for the medal!"
    }
  ],
  "comments": [
    {
      "id": 3033855,
      "author_name": "Pascal Pfeiffer",
      "author_url": "",
      "post_date": "2024-11-01T15:13:58.873000",
      "content": "<p>Interesting. Though the star was given as a primary feature even in test, wouldn't need such a proxy feature. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3034256,
      "author_name": "🐢 Jun Koda",
      "author_url": "",
      "post_date": "2024-11-01T23:42:25.310000",
      "content": "<p>I subtracted the mean of your yellow zone from the signal. I think this is the background light not coming from the star, and because star luminosity is different, it adds differently to the total light for different stars. It is an offset not contributing to the light difference of transit, but contribute to the denominator and I think it reduces 0.5 - 1% of systematic error.</p>\n<p>1 - target_pred = (high - low) / (high - background)</p>\n<p>Nice post! Congratulations for the medal!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3033831": "I was personally satisfied with my ranking, but compared to other members who shared their ideas, mine was quite simple (I added two features to the data of size (673, 283) = (number of samples, number of wavelengths) and performed a simple Gaussian Process Regression with a linear kernel). \nTherefore, I will focus on what two features I added.\n\nBy simply adding these features, my score increased by +0.015, and since it utilized the FGS signal, I found it personally interesting.\n\nWhat I did was add as features the average light intensity average values of the four corners of the FGS1 images and the upper and lower edges of the AIRS-CH0 images for each sample. (yellow zone)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F302d3251623d6772c243fdd77c666dbb%2F2024-11-01%2023.52.04.png?generation=1730472870237724&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F4e835df30fe6390cfa978cba3e782417%2F2024-11-01%2023.52.14.png?generation=1730472885814664&alt=media)\n\n\nThe reason why this feature led to a score improvement is likely because, as shown below, this feature could very clearly cluster the stars to which the planets belong.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F16997855%2F90745bfe5f18ca291635f9fcc03d6683%2F2024-11-01%2023.58.33.png?generation=1730473134199393&alt=media)\n\n\n\n",
    "3033855": "Interesting. Though the star was given as a primary feature even in test, wouldn't need such a proxy feature. ",
    "3034256": "I subtracted the mean of your yellow zone from the signal. I think this is the background light not coming from the star, and because star luminosity is different, it adds differently to the total light for different stars. It is an offset not contributing to the light difference of transit, but contribute to the denominator and I think it reduces 0.5 - 1% of systematic error.\n\n1 - target_pred = (high - low) / (high - background)\n\nNice post! Congratulations for the medal!"
  }
}