{
  "id": 543857,
  "title": "Place 38: Improve public baseline by ~0.08 with ~20 lines of code 🚀",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543857",
  "author_name": "Benedikt Droste",
  "post_date": "2024-11-01T20:10:02.693000",
  "votes": 18,
  "comment_count": 4,
  "views": 0,
  "content": "<p>The challenge was very exciting. Unfortunately, <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> and I were only able to take part very late and we couldn't invest much time. Nevertheless, I would like to share a little trick with which the <a href=\"https://www.kaggle.com/code/vitalykudelya/neurips-ariel-data-correlation-parallel-scale\" target=\"_blank\">great public solution</a> (based on the <a href=\"https://www.kaggle.com/code/sergeifironov/ariel-only-correlation\" target=\"_blank\">impressive work</a> of <a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a> ) can be significantly improved with little effort.</p>\n<p>Many participants had problems finding suitable values for Sigma. At best, sigma should represent the deviation of the prediction from the ground truth. We assumed that the mean calculation from the public notebook was already quite good and concentrated on a better sigma prediction.</p>\n<p>We had the hypothesis that different wavelengths react differently to the measurable light intensity during the transit. Therefore, we wanted to cluster wavelengths with similar characteristics and then apply the public approach to these clusters. We then wanted to use the standard deviation of the cluster calculations as a sigma. This worked very well:</p>\n<table>\n<thead>\n<tr>\n<th>approach</th>\n<th>public score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>public baseline</td>\n<td>0.546</td>\n</tr>\n<tr>\n<td>public with improved sigma</td>\n<td>0.626</td>\n</tr>\n</tbody>\n</table>\n<p>We simply formed the clusters using the targets. If you look at a simple correlation matrix of the train labels, patterns are clearly visible:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fbe5b42cebfebf692e5f28ffe47e05a64%2Fcorr_matrix.png?generation=1730491076919086&amp;alt=media\" alt=\"\"></p>\n<p>We converted the matrix into a condensed distance matrix and finally divided it into hierarchical clusters. We determined the threshold by trial and error. In the end, we created 7 clusters:</p>\n<pre><code> ():\n    train_labels = train_labels.values[:,:]\n    correlation_matrix = np.corrcoef(train_labels, rowvar=)\n    distance_matrix =  - correlation_matrix\n\n    condensed_distance_matrix = squareform(distance_matrix, checks=)\n    Z = linkage(condensed_distance_matrix, method=)\n\n    clusters = fcluster(Z, threshold, criterion=)\n     clusters\n</code></pre>\n<p>We then made predictions for each cluster (analogous to the public notebook):</p>\n<pre><code>cluster_placeholder = np.zeros((adc_info.shape[], ))\n\n i  ((adc_info)):\n    base_signal = preprocessed_signal[i,:,:]\n     cluster_n  np.unique(clusters):\n        relevant_indizes = _get_cluster_indizes(cluster_n, clusters)\n        signal = base_signal[:,relevant_indizes]\n        s = predict_spectra(signal, gauss_sigma=)\n        cluster_placeholder[i, relevant_indizes] = s\n</code></pre>\n<p>We took the standard deviation of the clusters as the sigma:</p>\n<pre><code>cluster_sigmas = np.ones_like(cluster_placeholder) * cluster_placeholder.clip().std(axis=).reshape(cluster_placeholder.shape[],-)\n</code></pre>\n<p>Interestingly, the sigma matched better with the mean calculation than with the cluster predictions. That was already enough to score in the silver medal range.</p>",
  "messages": [
    {
      "id": 3034142,
      "postDate": "2024-11-01T20:10:02.693Z",
      "content": "<p>The challenge was very exciting. Unfortunately, <a href=\"https://www.kaggle.com/nischaydnk\" target=\"_blank\">@nischaydnk</a> and I were only able to take part very late and we couldn't invest much time. Nevertheless, I would like to share a little trick with which the <a href=\"https://www.kaggle.com/code/vitalykudelya/neurips-ariel-data-correlation-parallel-scale\" target=\"_blank\">great public solution</a> (based on the <a href=\"https://www.kaggle.com/code/sergeifironov/ariel-only-correlation\" target=\"_blank\">impressive work</a> of <a href=\"https://www.kaggle.com/sergeifironov\" target=\"_blank\">@sergeifironov</a> ) can be significantly improved with little effort.</p>\n<p>Many participants had problems finding suitable values for Sigma. At best, sigma should represent the deviation of the prediction from the ground truth. We assumed that the mean calculation from the public notebook was already quite good and concentrated on a better sigma prediction.</p>\n<p>We had the hypothesis that different wavelengths react differently to the measurable light intensity during the transit. Therefore, we wanted to cluster wavelengths with similar characteristics and then apply the public approach to these clusters. We then wanted to use the standard deviation of the cluster calculations as a sigma. This worked very well:</p>\n<table>\n<thead>\n<tr>\n<th>approach</th>\n<th>public score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>public baseline</td>\n<td>0.546</td>\n</tr>\n<tr>\n<td>public with improved sigma</td>\n<td>0.626</td>\n</tr>\n</tbody>\n</table>\n<p>We simply formed the clusters using the targets. If you look at a simple correlation matrix of the train labels, patterns are clearly visible:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fbe5b42cebfebf692e5f28ffe47e05a64%2Fcorr_matrix.png?generation=1730491076919086&amp;alt=media\" alt=\"\"></p>\n<p>We converted the matrix into a condensed distance matrix and finally divided it into hierarchical clusters. We determined the threshold by trial and error. In the end, we created 7 clusters:</p>\n<pre><code> ():\n    train_labels = train_labels.values[:,:]\n    correlation_matrix = np.corrcoef(train_labels, rowvar=)\n    distance_matrix =  - correlation_matrix\n\n    condensed_distance_matrix = squareform(distance_matrix, checks=)\n    Z = linkage(condensed_distance_matrix, method=)\n\n    clusters = fcluster(Z, threshold, criterion=)\n     clusters\n</code></pre>\n<p>We then made predictions for each cluster (analogous to the public notebook):</p>\n<pre><code>cluster_placeholder = np.zeros((adc_info.shape[], ))\n\n i  ((adc_info)):\n    base_signal = preprocessed_signal[i,:,:]\n     cluster_n  np.unique(clusters):\n        relevant_indizes = _get_cluster_indizes(cluster_n, clusters)\n        signal = base_signal[:,relevant_indizes]\n        s = predict_spectra(signal, gauss_sigma=)\n        cluster_placeholder[i, relevant_indizes] = s\n</code></pre>\n<p>We took the standard deviation of the clusters as the sigma:</p>\n<pre><code>cluster_sigmas = np.ones_like(cluster_placeholder) * cluster_placeholder.clip().std(axis=).reshape(cluster_placeholder.shape[],-)\n</code></pre>\n<p>Interestingly, the sigma matched better with the mean calculation than with the cluster predictions. That was already enough to score in the silver medal range.</p>",
      "rawMarkdown": "The challenge was very exciting. Unfortunately, @nischaydnk and I were only able to take part very late and we couldn't invest much time. Nevertheless, I would like to share a little trick with which the [great public solution](https://www.kaggle.com/code/vitalykudelya/neurips-ariel-data-correlation-parallel-scale) (based on the [impressive work](https://www.kaggle.com/code/sergeifironov/ariel-only-correlation) of @sergeifironov ) can be significantly improved with little effort.\n\nMany participants had problems finding suitable values for Sigma. At best, sigma should represent the deviation of the prediction from the ground truth. We assumed that the mean calculation from the public notebook was already quite good and concentrated on a better sigma prediction.\n\nWe had the hypothesis that different wavelengths react differently to the measurable light intensity during the transit. Therefore, we wanted to cluster wavelengths with similar characteristics and then apply the public approach to these clusters. We then wanted to use the standard deviation of the cluster calculations as a sigma. This worked very well:\n\n| approach |  public score |\n| --- | --- |\n| public baseline |  0.546 |\n| public with improved sigma| 0.626 |\n\nWe simply formed the clusters using the targets. If you look at a simple correlation matrix of the train labels, patterns are clearly visible:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fbe5b42cebfebf692e5f28ffe47e05a64%2Fcorr_matrix.png?generation=1730491076919086&alt=media)\n\nWe converted the matrix into a condensed distance matrix and finally divided it into hierarchical clusters. We determined the threshold by trial and error. In the end, we created 7 clusters:\n\n```python\ndef get_clusters(train_labels, threshold=0.002):\n    train_labels = train_labels.values[:,1:]\n    correlation_matrix = np.corrcoef(train_labels, rowvar=False)\n    distance_matrix = 1 - correlation_matrix\n\n    condensed_distance_matrix = squareform(distance_matrix, checks=False)\n    Z = linkage(condensed_distance_matrix, method='complete')\n\n    clusters = fcluster(Z, threshold, criterion='distance')\n    return clusters\n```\n\nWe then made predictions for each cluster (analogous to the public notebook):\n\n```python\ncluster_placeholder = np.zeros((adc_info.shape[0], 283))\n\nfor i in range(len(adc_info)):\n    base_signal = preprocessed_signal[i,:,:]\n    for cluster_n in np.unique(clusters):\n        relevant_indizes = _get_cluster_indizes(cluster_n, clusters)\n        signal = base_signal[:,relevant_indizes]\n        s = predict_spectra(signal, gauss_sigma=2.0)\n        cluster_placeholder[i, relevant_indizes] = s\n```\n\nWe took the standard deviation of the clusters as the sigma:\n```python\ncluster_sigmas = np.ones_like(cluster_placeholder) * cluster_placeholder.clip(0).std(axis=1).reshape(cluster_placeholder.shape[0],-1)\n```\nInterestingly, the sigma matched better with the mean calculation than with the cluster predictions. That was already enough to score in the silver medal range.",
      "votes": 18
    },
    {
      "id": 3034214,
      "postDate": "2024-11-01T22:00:57.047Z",
      "content": "<p>Nice! It's very elegant idea and code.</p>",
      "rawMarkdown": "Nice! It's very elegant idea and code.",
      "votes": 1,
      "replies": [
        {
          "id": 3034637,
          "postDate": "2024-11-02T12:30:03.633Z",
          "content": "<p>Thanks, the resulting sigma worked with your cnn pipeline as well. It was really strong out of the box.</p>\n<p>Btw: thank you for all your inputs and baseline notebooks during the competition :-)</p>",
          "rawMarkdown": "Thanks, the resulting sigma worked with your cnn pipeline as well. It was really strong out of the box.\n\nBtw: thank you for all your inputs and baseline notebooks during the competition :-)"
        }
      ]
    },
    {
      "id": 3034462,
      "postDate": "2024-11-02T06:56:52.623Z",
      "content": "<p>I shared the basic code here:<br>\n<a href=\"https://www.kaggle.com/code/benbla/neurips-ariel-data-correlation-parallel-sigma\" target=\"_blank\">https://www.kaggle.com/code/benbla/neurips-ariel-data-correlation-parallel-sigma</a></p>\n<p>It is also worth mentioning that we have smoothed the signal with a Gaussian filter. </p>",
      "rawMarkdown": "I shared the basic code here:\nhttps://www.kaggle.com/code/benbla/neurips-ariel-data-correlation-parallel-sigma\n\nIt is also worth mentioning that we have smoothed the signal with a Gaussian filter. ",
      "votes": 2,
      "replies": [
        {
          "id": 3034598,
          "postDate": "2024-11-02T11:12:12.427Z",
          "content": "<p>Thank you so much for the share <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> </p>",
          "rawMarkdown": "Thank you so much for the share @benbla ",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3034214,
      "author_name": "Sergei Fironov",
      "author_url": "",
      "post_date": "2024-11-01T22:00:57.047000",
      "content": "<p>Nice! It's very elegant idea and code.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3034637,
          "author_name": "Benedikt Droste",
          "author_url": "",
          "post_date": "2024-11-02T12:30:03.633000",
          "content": "<p>Thanks, the resulting sigma worked with your cnn pipeline as well. It was really strong out of the box.</p>\n<p>Btw: thank you for all your inputs and baseline notebooks during the competition :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3034462,
      "author_name": "Benedikt Droste",
      "author_url": "",
      "post_date": "2024-11-02T06:56:52.623000",
      "content": "<p>I shared the basic code here:<br>\n<a href=\"https://www.kaggle.com/code/benbla/neurips-ariel-data-correlation-parallel-sigma\" target=\"_blank\">https://www.kaggle.com/code/benbla/neurips-ariel-data-correlation-parallel-sigma</a></p>\n<p>It is also worth mentioning that we have smoothed the signal with a Gaussian filter. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 3034598,
          "author_name": "JamshaidSohail",
          "author_url": "",
          "post_date": "2024-11-02T11:12:12.427000",
          "content": "<p>Thank you so much for the share <a href=\"https://www.kaggle.com/benbla\" target=\"_blank\">@benbla</a> </p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3034142": "The challenge was very exciting. Unfortunately, @nischaydnk and I were only able to take part very late and we couldn't invest much time. Nevertheless, I would like to share a little trick with which the [great public solution](https://www.kaggle.com/code/vitalykudelya/neurips-ariel-data-correlation-parallel-scale) (based on the [impressive work](https://www.kaggle.com/code/sergeifironov/ariel-only-correlation) of @sergeifironov ) can be significantly improved with little effort.\n\nMany participants had problems finding suitable values for Sigma. At best, sigma should represent the deviation of the prediction from the ground truth. We assumed that the mean calculation from the public notebook was already quite good and concentrated on a better sigma prediction.\n\nWe had the hypothesis that different wavelengths react differently to the measurable light intensity during the transit. Therefore, we wanted to cluster wavelengths with similar characteristics and then apply the public approach to these clusters. We then wanted to use the standard deviation of the cluster calculations as a sigma. This worked very well:\n\n| approach |  public score |\n| --- | --- |\n| public baseline |  0.546 |\n| public with improved sigma| 0.626 |\n\nWe simply formed the clusters using the targets. If you look at a simple correlation matrix of the train labels, patterns are clearly visible:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2959383%2Fbe5b42cebfebf692e5f28ffe47e05a64%2Fcorr_matrix.png?generation=1730491076919086&alt=media)\n\nWe converted the matrix into a condensed distance matrix and finally divided it into hierarchical clusters. We determined the threshold by trial and error. In the end, we created 7 clusters:\n\n```python\ndef get_clusters(train_labels, threshold=0.002):\n    train_labels = train_labels.values[:,1:]\n    correlation_matrix = np.corrcoef(train_labels, rowvar=False)\n    distance_matrix = 1 - correlation_matrix\n\n    condensed_distance_matrix = squareform(distance_matrix, checks=False)\n    Z = linkage(condensed_distance_matrix, method='complete')\n\n    clusters = fcluster(Z, threshold, criterion='distance')\n    return clusters\n```\n\nWe then made predictions for each cluster (analogous to the public notebook):\n\n```python\ncluster_placeholder = np.zeros((adc_info.shape[0], 283))\n\nfor i in range(len(adc_info)):\n    base_signal = preprocessed_signal[i,:,:]\n    for cluster_n in np.unique(clusters):\n        relevant_indizes = _get_cluster_indizes(cluster_n, clusters)\n        signal = base_signal[:,relevant_indizes]\n        s = predict_spectra(signal, gauss_sigma=2.0)\n        cluster_placeholder[i, relevant_indizes] = s\n```\n\nWe took the standard deviation of the clusters as the sigma:\n```python\ncluster_sigmas = np.ones_like(cluster_placeholder) * cluster_placeholder.clip(0).std(axis=1).reshape(cluster_placeholder.shape[0],-1)\n```\nInterestingly, the sigma matched better with the mean calculation than with the cluster predictions. That was already enough to score in the silver medal range.",
    "3034214": "Nice! It's very elegant idea and code.",
    "3034462": "I shared the basic code here:\nhttps://www.kaggle.com/code/benbla/neurips-ariel-data-correlation-parallel-sigma\n\nIt is also worth mentioning that we have smoothed the signal with a Gaussian filter. "
  }
}