{
  "id": 544418,
  "title": "26th place solution",
  "url": "/competitions/ariel-data-challenge-2024/discussion/544418",
  "author_name": "xaoyang",
  "post_date": "2024-11-05T03:47:43.424000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Thank you to the ARIEL team for hosting this competition and to everyone who shared insightful public notebooks, which were extremely helpful.</p>\n<p>In my opinion, our improvements on the public solutions are as follows:<br>\n(The code omits some details and highlights key parts only)</p>\n<h3>1.Sliding window</h3>\n<p>The sliding window helps to smooth data by calculating averages or applying filters within the window, reducing noise and making the signal more stable and continuous.</p>\n<pre><code>a_raw_train.shape\n</code></pre>\n<blockquote>\n  <p>(1, 187, 356)</p>\n</blockquote>\n<pre><code>window = \nwv_smooth_a_raw_train = []\n i  (cut_inf, cut_sup):\n    wv_smooth_a_raw_train.append( a_raw_train[:, :, (i-window, ):(i+window)].mean(axis=-) )\nwv_smooth_a_raw_train = np.stack(wv_smooth_a_raw_train, axis=-)\n</code></pre>\n<pre><code>wv_smooth_a_raw_train.shape\n</code></pre>\n<blockquote>\n  <p>(1, 187, 282)</p>\n</blockquote>\n<h3>2.Smooth a_wv_depth</h3>\n<p>We smooth the calculated transit depth for each band. It can suppress random fluctuations in the data, reduce noise, and make the signal or trend clearer, making it easier to identify primary patterns.</p>\n<pre><code>a_wv_depth.shape\n</code></pre>\n<blockquote>\n  <p>(1, 282)</p>\n</blockquote>\n<pre><code>plt.plot(a_wv_depth[])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2Fb437c333343dc37e6f709def8707dcc9%2F11.png?generation=1730528008347835&amp;alt=media\" alt=\"\"></p>\n<pre><code> scipy.signal  savgol_filter\n ():\n     savgol_filter(data, window_size, ) \n</code></pre>\n<pre><code>window_size = \nsmoothed = []\n pred  tqdm(a_wv_depth):\n    smooth_pred = smooth_data(pred, window_size=window_size)\n    smoothed.append(smooth_pred)\nsmoothed = np.row_stack(smoothed)\nsmoothed.shape\n</code></pre>\n<blockquote>\n  <p>(1, 282)</p>\n</blockquote>\n<pre><code>plt.plot(smoothed[])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2F24edee6b9011a0208e29d6036104bbe9%2F22.png?generation=1730528040529759&amp;alt=media\" alt=\"\"></p>\n<h3>3.Linear regression for sigma prediction</h3>\n<p>The sigma at planetary granularity is positively correlated with the differences in the depth of each band.</p>\n<pre><code>train_label_df = pd.read_csv(, index_col=)\ntrain_wv_smooth_pred = np.load()\n</code></pre>\n<pre><code>train_wv_smooth_pred.shape\n</code></pre>\n<blockquote>\n  <p>(673, 283)</p>\n</blockquote>\n<pre><code>train_label_df.shape\n</code></pre>\n<blockquote>\n  <p>(673, 283)</p>\n</blockquote>\n<pre><code>sigma_gt = np.(train_label_df.values - train_wv_smooth_pred)\nsigma_gt.shape\n</code></pre>\n<blockquote>\n  <p>(673, 283)</p>\n</blockquote>\n<pre><code>x = sigma_gt.mean(axis=-)\ny = train_wv_smooth_pred.std(axis=-)\nr = np.corrcoef(x, y)[, ]\nplt.scatter(x, y)\nplt.title()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2Fcae81420ef0c1189cd0f71b93c8e5776%2F33.png?generation=1730528081599637&amp;alt=media\" alt=\"\"></p>\n<blockquote>\n  <p>Text(0.5, 1.0, 'r=0.700985672809354')</p>\n</blockquote>\n<pre><code>star_model = LinearRegression()\nstar_model.fit(X, sigma_gt)\n</code></pre>\n<pre><code>sigma = []\n star, x  (adc_info.star, wv_smooth_pred.std(axis=-)):\n    out = star_model.predict(x.reshape(, -))\n    out = out.clip() * \n    sigma.append(out)\n\nsigma = np.row_stack(sigma)\nsigma.shape\n</code></pre>\n<blockquote>\n  <p>(1, 283)</p>\n</blockquote>\n<h2>4.Some other approaches will be updated afterwards.</h2>\n<p>Lastly, I want to thank every member of our team; each of them has been a tremendous help to me.</p>",
  "messages": [
    {
      "id": 3036914,
      "postDate": "2024-11-05T03:47:43.423Z",
      "content": "<p>Thank you to the ARIEL team for hosting this competition and to everyone who shared insightful public notebooks, which were extremely helpful.</p>\n<p>In my opinion, our improvements on the public solutions are as follows:<br>\n(The code omits some details and highlights key parts only)</p>\n<h3>1.Sliding window</h3>\n<p>The sliding window helps to smooth data by calculating averages or applying filters within the window, reducing noise and making the signal more stable and continuous.</p>\n<pre><code>a_raw_train.shape\n</code></pre>\n<blockquote>\n  <p>(1, 187, 356)</p>\n</blockquote>\n<pre><code>window = \nwv_smooth_a_raw_train = []\n i  (cut_inf, cut_sup):\n    wv_smooth_a_raw_train.append( a_raw_train[:, :, (i-window, ):(i+window)].mean(axis=-) )\nwv_smooth_a_raw_train = np.stack(wv_smooth_a_raw_train, axis=-)\n</code></pre>\n<pre><code>wv_smooth_a_raw_train.shape\n</code></pre>\n<blockquote>\n  <p>(1, 187, 282)</p>\n</blockquote>\n<h3>2.Smooth a_wv_depth</h3>\n<p>We smooth the calculated transit depth for each band. It can suppress random fluctuations in the data, reduce noise, and make the signal or trend clearer, making it easier to identify primary patterns.</p>\n<pre><code>a_wv_depth.shape\n</code></pre>\n<blockquote>\n  <p>(1, 282)</p>\n</blockquote>\n<pre><code>plt.plot(a_wv_depth[])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2Fb437c333343dc37e6f709def8707dcc9%2F11.png?generation=1730528008347835&amp;alt=media\" alt=\"\"></p>\n<pre><code> scipy.signal  savgol_filter\n ():\n     savgol_filter(data, window_size, ) \n</code></pre>\n<pre><code>window_size = \nsmoothed = []\n pred  tqdm(a_wv_depth):\n    smooth_pred = smooth_data(pred, window_size=window_size)\n    smoothed.append(smooth_pred)\nsmoothed = np.row_stack(smoothed)\nsmoothed.shape\n</code></pre>\n<blockquote>\n  <p>(1, 282)</p>\n</blockquote>\n<pre><code>plt.plot(smoothed[])\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2F24edee6b9011a0208e29d6036104bbe9%2F22.png?generation=1730528040529759&amp;alt=media\" alt=\"\"></p>\n<h3>3.Linear regression for sigma prediction</h3>\n<p>The sigma at planetary granularity is positively correlated with the differences in the depth of each band.</p>\n<pre><code>train_label_df = pd.read_csv(, index_col=)\ntrain_wv_smooth_pred = np.load()\n</code></pre>\n<pre><code>train_wv_smooth_pred.shape\n</code></pre>\n<blockquote>\n  <p>(673, 283)</p>\n</blockquote>\n<pre><code>train_label_df.shape\n</code></pre>\n<blockquote>\n  <p>(673, 283)</p>\n</blockquote>\n<pre><code>sigma_gt = np.(train_label_df.values - train_wv_smooth_pred)\nsigma_gt.shape\n</code></pre>\n<blockquote>\n  <p>(673, 283)</p>\n</blockquote>\n<pre><code>x = sigma_gt.mean(axis=-)\ny = train_wv_smooth_pred.std(axis=-)\nr = np.corrcoef(x, y)[, ]\nplt.scatter(x, y)\nplt.title()\n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2Fcae81420ef0c1189cd0f71b93c8e5776%2F33.png?generation=1730528081599637&amp;alt=media\" alt=\"\"></p>\n<blockquote>\n  <p>Text(0.5, 1.0, 'r=0.700985672809354')</p>\n</blockquote>\n<pre><code>star_model = LinearRegression()\nstar_model.fit(X, sigma_gt)\n</code></pre>\n<pre><code>sigma = []\n star, x  (adc_info.star, wv_smooth_pred.std(axis=-)):\n    out = star_model.predict(x.reshape(, -))\n    out = out.clip() * \n    sigma.append(out)\n\nsigma = np.row_stack(sigma)\nsigma.shape\n</code></pre>\n<blockquote>\n  <p>(1, 283)</p>\n</blockquote>\n<h2>4.Some other approaches will be updated afterwards.</h2>\n<p>Lastly, I want to thank every member of our team; each of them has been a tremendous help to me.</p>",
      "rawMarkdown": "Thank you to the ARIEL team for hosting this competition and to everyone who shared insightful public notebooks, which were extremely helpful.\n\nIn my opinion, our improvements on the public solutions are as follows:\n(The code omits some details and highlights key parts only)\n\n### 1.Sliding window\nThe sliding window helps to smooth data by calculating averages or applying filters within the window, reducing noise and making the signal more stable and continuous.\n```python\na_raw_train.shape\n```\n> (1, 187, 356)\n```python\nwindow = 37\nwv_smooth_a_raw_train = []\nfor i in range(cut_inf, cut_sup):\n    wv_smooth_a_raw_train.append( a_raw_train[:, :, max(i-window, 0):(i+window)].mean(axis=-1) )\nwv_smooth_a_raw_train = np.stack(wv_smooth_a_raw_train, axis=-1)\n```\n```python\nwv_smooth_a_raw_train.shape\n```\n> (1, 187, 282)\n\n### 2.Smooth a_wv_depth\nWe smooth the calculated transit depth for each band. It can suppress random fluctuations in the data, reduce noise, and make the signal or trend clearer, making it easier to identify primary patterns.\n```python\na_wv_depth.shape\n```\n> (1, 282)\n```python\nplt.plot(a_wv_depth[0])\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2Fb437c333343dc37e6f709def8707dcc9%2F11.png?generation=1730528008347835&alt=media)\n```python\nfrom scipy.signal import savgol_filter\ndef smooth_data(data, window_size):\n    return savgol_filter(data, window_size, 3) \n```\n```python\nwindow_size = 103\nsmoothed = []\nfor pred in tqdm(a_wv_depth):\n    smooth_pred = smooth_data(pred, window_size=window_size)\n    smoothed.append(smooth_pred)\nsmoothed = np.row_stack(smoothed)\nsmoothed.shape\n```\n> (1, 282)\n```python\nplt.plot(smoothed[0])\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2F24edee6b9011a0208e29d6036104bbe9%2F22.png?generation=1730528040529759&alt=media)\n\n### 3.Linear regression for sigma prediction \nThe sigma at planetary granularity is positively correlated with the differences in the depth of each band.\n```python\ntrain_label_df = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/train_labels.csv', index_col='planet_id')\ntrain_wv_smooth_pred = np.load('/kaggle/input/adc2024-train-wv-smooth-pred/train_wv_smooth_pred.npy')\n```\n```python\ntrain_wv_smooth_pred.shape\n```\n> (673, 283)\n```python\ntrain_label_df.shape\n```\n> (673, 283)\n```python\nsigma_gt = np.abs(train_label_df.values - train_wv_smooth_pred)\nsigma_gt.shape\n```\n> (673, 283)\n```python\nx = sigma_gt.mean(axis=-1)\ny = train_wv_smooth_pred.std(axis=-1)\nr = np.corrcoef(x, y)[1, 0]\nplt.scatter(x, y)\nplt.title(f'r={r}')\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2Fcae81420ef0c1189cd0f71b93c8e5776%2F33.png?generation=1730528081599637&alt=media)\n> Text(0.5, 1.0, 'r=0.700985672809354')\n```python\nstar_model = LinearRegression()\nstar_model.fit(X, sigma_gt)\n```\n```python\nsigma = []\nfor star, x in zip(adc_info.star, wv_smooth_pred.std(axis=-1)):\n    out = star_model.predict(x.reshape(1, -1))\n    out = out.clip(3e-5) * 1.5\n    sigma.append(out)\n    \nsigma = np.row_stack(sigma)\nsigma.shape\n```\n> (1, 283)\n\n## 4.Some other approaches will be updated afterwards.\n\nLastly, I want to thank every member of our team; each of them has been a tremendous help to me.",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "3036914": "Thank you to the ARIEL team for hosting this competition and to everyone who shared insightful public notebooks, which were extremely helpful.\n\nIn my opinion, our improvements on the public solutions are as follows:\n(The code omits some details and highlights key parts only)\n\n### 1.Sliding window\nThe sliding window helps to smooth data by calculating averages or applying filters within the window, reducing noise and making the signal more stable and continuous.\n```python\na_raw_train.shape\n```\n> (1, 187, 356)\n```python\nwindow = 37\nwv_smooth_a_raw_train = []\nfor i in range(cut_inf, cut_sup):\n    wv_smooth_a_raw_train.append( a_raw_train[:, :, max(i-window, 0):(i+window)].mean(axis=-1) )\nwv_smooth_a_raw_train = np.stack(wv_smooth_a_raw_train, axis=-1)\n```\n```python\nwv_smooth_a_raw_train.shape\n```\n> (1, 187, 282)\n\n### 2.Smooth a_wv_depth\nWe smooth the calculated transit depth for each band. It can suppress random fluctuations in the data, reduce noise, and make the signal or trend clearer, making it easier to identify primary patterns.\n```python\na_wv_depth.shape\n```\n> (1, 282)\n```python\nplt.plot(a_wv_depth[0])\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2Fb437c333343dc37e6f709def8707dcc9%2F11.png?generation=1730528008347835&alt=media)\n```python\nfrom scipy.signal import savgol_filter\ndef smooth_data(data, window_size):\n    return savgol_filter(data, window_size, 3) \n```\n```python\nwindow_size = 103\nsmoothed = []\nfor pred in tqdm(a_wv_depth):\n    smooth_pred = smooth_data(pred, window_size=window_size)\n    smoothed.append(smooth_pred)\nsmoothed = np.row_stack(smoothed)\nsmoothed.shape\n```\n> (1, 282)\n```python\nplt.plot(smoothed[0])\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2F24edee6b9011a0208e29d6036104bbe9%2F22.png?generation=1730528040529759&alt=media)\n\n### 3.Linear regression for sigma prediction \nThe sigma at planetary granularity is positively correlated with the differences in the depth of each band.\n```python\ntrain_label_df = pd.read_csv('/kaggle/input/ariel-data-challenge-2024/train_labels.csv', index_col='planet_id')\ntrain_wv_smooth_pred = np.load('/kaggle/input/adc2024-train-wv-smooth-pred/train_wv_smooth_pred.npy')\n```\n```python\ntrain_wv_smooth_pred.shape\n```\n> (673, 283)\n```python\ntrain_label_df.shape\n```\n> (673, 283)\n```python\nsigma_gt = np.abs(train_label_df.values - train_wv_smooth_pred)\nsigma_gt.shape\n```\n> (673, 283)\n```python\nx = sigma_gt.mean(axis=-1)\ny = train_wv_smooth_pred.std(axis=-1)\nr = np.corrcoef(x, y)[1, 0]\nplt.scatter(x, y)\nplt.title(f'r={r}')\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F12135819%2Fcae81420ef0c1189cd0f71b93c8e5776%2F33.png?generation=1730528081599637&alt=media)\n> Text(0.5, 1.0, 'r=0.700985672809354')\n```python\nstar_model = LinearRegression()\nstar_model.fit(X, sigma_gt)\n```\n```python\nsigma = []\nfor star, x in zip(adc_info.star, wv_smooth_pred.std(axis=-1)):\n    out = star_model.predict(x.reshape(1, -1))\n    out = out.clip(3e-5) * 1.5\n    sigma.append(out)\n    \nsigma = np.row_stack(sigma)\nsigma.shape\n```\n> (1, 283)\n\n## 4.Some other approaches will be updated afterwards.\n\nLastly, I want to thank every member of our team; each of them has been a tremendous help to me."
  }
}