{
  "id": 529533,
  "title": "Easily Find The Transit Zones [With Full Code Inside!]",
  "url": "/competitions/ariel-data-challenge-2024/discussion/529533",
  "author_name": "Pascal Pfeiffer",
  "post_date": "2024-08-21T13:23:48.805000",
  "votes": 40,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>The issue</h1>\n<p>As <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> has shown in <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527167\" target=\"_blank\">this discussion</a>, the transit time for the planets to pass the stars is not the same for each sample. Though, calculation and feature engineering on the light curves requires exact knowledge of these cutoffs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9b51760603a8f72d3604305ed76b8841%2Fsl.png?generation=1723307476416972&amp;alt=media\" alt=\"\"></p>\n<h1>The solution</h1>\n<p>Here, I am presenting a very robust solution to identify the transit zone that works on ALL train samples, with both sensor types. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F815d519fc2f09a15a96282223f432561%2F__results___6_0.png?generation=1724246295909588&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Fc8dcc7b6800036ce45a352d298c99249%2F__results___6_1.png?generation=1724246305869502&amp;alt=media\" alt=\"\"></p>\n<pre><code> scipy.signal  savgol_filter\n\n\n ():\n     savgol_filter(data, window_size, )  \n\n\n ():\n    best_breakpoint = initial_breakpoint\n    best_score = ()\n    midpoint = (data) // \n    smoothed_data = smooth_data(data, smooth_window)\n\n     i  (-window_size, window_size):\n        new_breakpoint = initial_breakpoint + i\n         new_breakpoint &gt; buffer_size  new_breakpoint &lt; midpoint - buffer_size:\n            region1 = data[: new_breakpoint - buffer_size]\n            region2 = data[\n                new_breakpoint\n                + buffer_size :  * midpoint\n                - new_breakpoint\n                - buffer_size\n            ]\n            region3 = data[ * midpoint - new_breakpoint + buffer_size :]\n\n            \n            breakpoint_region1 = smoothed_data[new_breakpoint - buffer_size: new_breakpoint + buffer_size]\n            breakpoint_region2 = smoothed_data[new_breakpoint - buffer_size: new_breakpoint + buffer_size]\n\n            mean_diff = (np.mean(region1) - np.mean(region2)) + (\n                np.mean(region2) - np.mean(region3)\n            )\n            var_sum = np.var(region1) + np.var(region2) + np.var(region3)\n            range_at_breakpoint1 = (np.(breakpoint_region1) - np.(breakpoint_region1))\n            range_at_breakpoint2 = (np.(breakpoint_region2) - np.(breakpoint_region2))\n\n            mean_range_at_breakpoint = (range_at_breakpoint1 + range_at_breakpoint2) / \n\n            score = mean_diff -  * var_sum + mean_range_at_breakpoint\n\n             score &gt; best_score:\n                best_score = score\n                best_breakpoint = new_breakpoint\n\n     best_breakpoint\n</code></pre>",
  "messages": [
    {
      "id": 2966012,
      "postDate": "2024-08-21T13:23:48.807Z",
      "content": "<h1>The issue</h1>\n<p>As <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> has shown in <a href=\"https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527167\" target=\"_blank\">this discussion</a>, the transit time for the planets to pass the stars is not the same for each sample. Though, calculation and feature engineering on the light curves requires exact knowledge of these cutoffs.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9b51760603a8f72d3604305ed76b8841%2Fsl.png?generation=1723307476416972&amp;alt=media\" alt=\"\"></p>\n<h1>The solution</h1>\n<p>Here, I am presenting a very robust solution to identify the transit zone that works on ALL train samples, with both sensor types. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F815d519fc2f09a15a96282223f432561%2F__results___6_0.png?generation=1724246295909588&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Fc8dcc7b6800036ce45a352d298c99249%2F__results___6_1.png?generation=1724246305869502&amp;alt=media\" alt=\"\"></p>\n<pre><code> scipy.signal  savgol_filter\n\n\n ():\n     savgol_filter(data, window_size, )  \n\n\n ():\n    best_breakpoint = initial_breakpoint\n    best_score = ()\n    midpoint = (data) // \n    smoothed_data = smooth_data(data, smooth_window)\n\n     i  (-window_size, window_size):\n        new_breakpoint = initial_breakpoint + i\n         new_breakpoint &gt; buffer_size  new_breakpoint &lt; midpoint - buffer_size:\n            region1 = data[: new_breakpoint - buffer_size]\n            region2 = data[\n                new_breakpoint\n                + buffer_size :  * midpoint\n                - new_breakpoint\n                - buffer_size\n            ]\n            region3 = data[ * midpoint - new_breakpoint + buffer_size :]\n\n            \n            breakpoint_region1 = smoothed_data[new_breakpoint - buffer_size: new_breakpoint + buffer_size]\n            breakpoint_region2 = smoothed_data[new_breakpoint - buffer_size: new_breakpoint + buffer_size]\n\n            mean_diff = (np.mean(region1) - np.mean(region2)) + (\n                np.mean(region2) - np.mean(region3)\n            )\n            var_sum = np.var(region1) + np.var(region2) + np.var(region3)\n            range_at_breakpoint1 = (np.(breakpoint_region1) - np.(breakpoint_region1))\n            range_at_breakpoint2 = (np.(breakpoint_region2) - np.(breakpoint_region2))\n\n            mean_range_at_breakpoint = (range_at_breakpoint1 + range_at_breakpoint2) / \n\n            score = mean_diff -  * var_sum + mean_range_at_breakpoint\n\n             score &gt; best_score:\n                best_score = score\n                best_breakpoint = new_breakpoint\n\n     best_breakpoint\n</code></pre>",
      "rawMarkdown": "# The issue\n\nAs @ambrosm has shown in [this discussion](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527167), the transit time for the planets to pass the stars is not the same for each sample. Though, calculation and feature engineering on the light curves requires exact knowledge of these cutoffs.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9b51760603a8f72d3604305ed76b8841%2Fsl.png?generation=1723307476416972&alt=media)\n\n# The solution\n\nHere, I am presenting a very robust solution to identify the transit zone that works on ALL train samples, with both sensor types. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F815d519fc2f09a15a96282223f432561%2F__results___6_0.png?generation=1724246295909588&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Fc8dcc7b6800036ce45a352d298c99249%2F__results___6_1.png?generation=1724246305869502&alt=media)\n\n```python\nfrom scipy.signal import savgol_filter\n\n\ndef smooth_data(data, window_size):\n    return savgol_filter(data, window_size, 3)  # window size 51, polynomial order 3\n\n\ndef optimize_breakpoint(data, initial_breakpoint, window_size=500, buffer_size=50, smooth_window=121):\n    best_breakpoint = initial_breakpoint\n    best_score = float(\"-inf\")\n    midpoint = len(data) // 2\n    smoothed_data = smooth_data(data, smooth_window)\n\n    for i in range(-window_size, window_size):\n        new_breakpoint = initial_breakpoint + i\n        if new_breakpoint > buffer_size and new_breakpoint < midpoint - buffer_size:\n            region1 = data[: new_breakpoint - buffer_size]\n            region2 = data[\n                new_breakpoint\n                + buffer_size : 2 * midpoint\n                - new_breakpoint\n                - buffer_size\n            ]\n            region3 = data[2 * midpoint - new_breakpoint + buffer_size :]\n\n            # calc on smoothed data\n            breakpoint_region1 = smoothed_data[new_breakpoint - buffer_size: new_breakpoint + buffer_size]\n            breakpoint_region2 = smoothed_data[new_breakpoint - buffer_size: new_breakpoint + buffer_size]\n\n            mean_diff = abs(np.mean(region1) - np.mean(region2)) + abs(\n                np.mean(region2) - np.mean(region3)\n            )\n            var_sum = np.var(region1) + np.var(region2) + np.var(region3)\n            range_at_breakpoint1 = (np.max(breakpoint_region1) - np.min(breakpoint_region1))\n            range_at_breakpoint2 = (np.max(breakpoint_region2) - np.min(breakpoint_region2))\n\n            mean_range_at_breakpoint = (range_at_breakpoint1 + range_at_breakpoint2) / 2\n\n            score = mean_diff - 0.5 * var_sum + mean_range_at_breakpoint\n\n            if score > best_score:\n                best_score = score\n                best_breakpoint = new_breakpoint\n\n    return best_breakpoint\n```\n",
      "votes": 40
    },
    {
      "id": 2966070,
      "postDate": "2024-08-21T14:07:16.297Z",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> -&gt; Thanks for sharing. is it similar to <strong>Box least square approach</strong>?  </p>\n</blockquote>",
      "rawMarkdown": "> @ilu000 -> Thanks for sharing. is it similar to **Box least square approach**?  ",
      "votes": 2,
      "replies": [
        {
          "id": 2966085,
          "postDate": "2024-08-21T14:23:08.100Z",
          "content": "<p>I haven't heard of \"Box least square approach\". This is an approach that I found empirically. It uses prior knowledge about the approximate position of the transit and prior knowledge that the transit is always centered and then optimizes a heuristic to find the best fit. </p>\n<ul>\n<li>high difference between the mean of the regions</li>\n<li>low variance in the regions</li>\n<li>large step at the breakpoints</li>\n</ul>\n<pre><code>score = mean_diff -  * var_sum + mean_range_at_breakpoint\n</code></pre>",
          "rawMarkdown": "I haven't heard of \"Box least square approach\". This is an approach that I found empirically. It uses prior knowledge about the approximate position of the transit and prior knowledge that the transit is always centered and then optimizes a heuristic to find the best fit. \n\n- high difference between the mean of the regions\n- low variance in the regions\n- large step at the breakpoints\n\n```python\nscore = mean_diff - 0.5 * var_sum + mean_range_at_breakpoint\n```",
          "votes": 3,
          "replies": [
            {
              "id": 2966768,
              "postDate": "2024-08-22T07:36:52.110Z",
              "content": "<p>The Box Least Squares (BLS) method is a statistical tool that astronomers often use to detect exoplanets that transit in front of their host stars. What it does is look for those periodic dips in a star's brightness that might indicate a planet passing by. Instead of getting too complex, BLS keeps it simple by fitting a straightforward box-shaped model to the light curve data. The \"box\" here is just a way to represent the time period when the star’s light dims because the planet is crossing in front of it. It’s not a perfect method, but it’s been quite useful in helping us find those small signals that might otherwise go unnoticed.</p>",
              "rawMarkdown": "The Box Least Squares (BLS) method is a statistical tool that astronomers often use to detect exoplanets that transit in front of their host stars. What it does is look for those periodic dips in a star's brightness that might indicate a planet passing by. Instead of getting too complex, BLS keeps it simple by fitting a straightforward box-shaped model to the light curve data. The \"box\" here is just a way to represent the time period when the star’s light dims because the planet is crossing in front of it. It’s not a perfect method, but it’s been quite useful in helping us find those small signals that might otherwise go unnoticed.",
              "votes": 7
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2966070,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2024-08-21T14:07:16.297000",
      "content": "<blockquote>\n  <p><a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> -&gt; Thanks for sharing. is it similar to <strong>Box least square approach</strong>?  </p>\n</blockquote>",
      "votes": 2,
      "replies": [
        {
          "id": 2966085,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2024-08-21T14:23:08.100000",
          "content": "<p>I haven't heard of \"Box least square approach\". This is an approach that I found empirically. It uses prior knowledge about the approximate position of the transit and prior knowledge that the transit is always centered and then optimizes a heuristic to find the best fit. </p>\n<ul>\n<li>high difference between the mean of the regions</li>\n<li>low variance in the regions</li>\n<li>large step at the breakpoints</li>\n</ul>\n<pre><code>score = mean_diff -  * var_sum + mean_range_at_breakpoint\n</code></pre>",
          "votes": 3,
          "replies": [
            {
              "id": 2966768,
              "author_name": "费文轩",
              "author_url": "",
              "post_date": "2024-08-22T07:36:52.110000",
              "content": "<p>The Box Least Squares (BLS) method is a statistical tool that astronomers often use to detect exoplanets that transit in front of their host stars. What it does is look for those periodic dips in a star's brightness that might indicate a planet passing by. Instead of getting too complex, BLS keeps it simple by fitting a straightforward box-shaped model to the light curve data. The \"box\" here is just a way to represent the time period when the star’s light dims because the planet is crossing in front of it. It’s not a perfect method, but it’s been quite useful in helping us find those small signals that might otherwise go unnoticed.</p>",
              "votes": 7,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2966012": "# The issue\n\nAs @ambrosm has shown in [this discussion](https://www.kaggle.com/competitions/ariel-data-challenge-2024/discussion/527167), the transit time for the planets to pass the stars is not the same for each sample. Though, calculation and feature engineering on the light curves requires exact knowledge of these cutoffs.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F7917824%2F9b51760603a8f72d3604305ed76b8841%2Fsl.png?generation=1723307476416972&alt=media)\n\n# The solution\n\nHere, I am presenting a very robust solution to identify the transit zone that works on ALL train samples, with both sensor types. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2F815d519fc2f09a15a96282223f432561%2F__results___6_0.png?generation=1724246295909588&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2675447%2Fc8dcc7b6800036ce45a352d298c99249%2F__results___6_1.png?generation=1724246305869502&alt=media)\n\n```python\nfrom scipy.signal import savgol_filter\n\n\ndef smooth_data(data, window_size):\n    return savgol_filter(data, window_size, 3)  # window size 51, polynomial order 3\n\n\ndef optimize_breakpoint(data, initial_breakpoint, window_size=500, buffer_size=50, smooth_window=121):\n    best_breakpoint = initial_breakpoint\n    best_score = float(\"-inf\")\n    midpoint = len(data) // 2\n    smoothed_data = smooth_data(data, smooth_window)\n\n    for i in range(-window_size, window_size):\n        new_breakpoint = initial_breakpoint + i\n        if new_breakpoint > buffer_size and new_breakpoint < midpoint - buffer_size:\n            region1 = data[: new_breakpoint - buffer_size]\n            region2 = data[\n                new_breakpoint\n                + buffer_size : 2 * midpoint\n                - new_breakpoint\n                - buffer_size\n            ]\n            region3 = data[2 * midpoint - new_breakpoint + buffer_size :]\n\n            # calc on smoothed data\n            breakpoint_region1 = smoothed_data[new_breakpoint - buffer_size: new_breakpoint + buffer_size]\n            breakpoint_region2 = smoothed_data[new_breakpoint - buffer_size: new_breakpoint + buffer_size]\n\n            mean_diff = abs(np.mean(region1) - np.mean(region2)) + abs(\n                np.mean(region2) - np.mean(region3)\n            )\n            var_sum = np.var(region1) + np.var(region2) + np.var(region3)\n            range_at_breakpoint1 = (np.max(breakpoint_region1) - np.min(breakpoint_region1))\n            range_at_breakpoint2 = (np.max(breakpoint_region2) - np.min(breakpoint_region2))\n\n            mean_range_at_breakpoint = (range_at_breakpoint1 + range_at_breakpoint2) / 2\n\n            score = mean_diff - 0.5 * var_sum + mean_range_at_breakpoint\n\n            if score > best_score:\n                best_score = score\n                best_breakpoint = new_breakpoint\n\n    return best_breakpoint\n```\n",
    "2966070": "> @ilu000 -> Thanks for sharing. is it similar to **Box least square approach**?  "
  }
}