{
  "id": 543921,
  "title": "#60: Data augmentation explained: Generating more stars",
  "url": "/competitions/ariel-data-challenge-2024/discussion/543921",
  "author_name": "AmbrosM",
  "post_date": "2024-11-02T07:47:48.111000",
  "votes": 22,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm surprised that data augmentation hasn't been used more widely in this competition. Here's my approach.</p>\n<h2>Background: Dim stars make noisy pictures</h2>\n<p>If you've ever taken a picture with your mobile phone in low-light conditions, you know the effect: The camera adapts its shutter speed, but the picture will be noisier. In our dataset, the main feature of a star is its brightness, and the less bright a star is, the more noise we get in its picture: Star 1 is brighter than Star 0, and the unseen stars in the test set are less bright than Star 0. Furthermore, the stars emit less energy at higher wavelengths, which means that the right part of the spectra is noisier than the left one.</p>\n<p>The effect can be quantified: The variance of the signal is proportional to the signal itself (see <a href=\"https://en.wikipedia.org/wiki/Shot_noise\" target=\"_blank\">Wikipedia: Shot noise</a>). For the depth predictions, the variance contributed by the shot noise is inversely proportional to the signal strength.</p>\n<h2>Implementation</h2>\n<p>If we want to create data for a new star which is less bright than the existing ones, we can multiply all the data by a scaling factor &lt; 1. This multiplication reduces the variance by the square of the scaling factor so that the new variance is too low, and we need to compensate by injecting some additional variance.</p>\n<pre><code> ():\n    \n    old_brightness = data.mean(axis=).mean(axis=)\n    scaling = new_brightness / old_brightness.reshape(-, , )\n    new_data = data * scaling\n    time_binning =  \n    rng = np.random.default_rng()\n    new_data += rng.normal(scale=np.sqrt(( - scaling) * new_data / time_binning),\n                           size=new_data) \n     new_data\n</code></pre>\n<p>This implementation only works for downscaling but not for upscaling. Fortunately the stars in the test dataset are all less bright then Star 1 and we don't need to upscale the data.</p>\n<h2>Conclusion</h2>\n<p>Data augmentation has helped me tune the predictions for the sigma values.</p>",
  "messages": [
    {
      "id": 3034498,
      "postDate": "2024-11-02T07:47:48.110Z",
      "content": "<p>I'm surprised that data augmentation hasn't been used more widely in this competition. Here's my approach.</p>\n<h2>Background: Dim stars make noisy pictures</h2>\n<p>If you've ever taken a picture with your mobile phone in low-light conditions, you know the effect: The camera adapts its shutter speed, but the picture will be noisier. In our dataset, the main feature of a star is its brightness, and the less bright a star is, the more noise we get in its picture: Star 1 is brighter than Star 0, and the unseen stars in the test set are less bright than Star 0. Furthermore, the stars emit less energy at higher wavelengths, which means that the right part of the spectra is noisier than the left one.</p>\n<p>The effect can be quantified: The variance of the signal is proportional to the signal itself (see <a href=\"https://en.wikipedia.org/wiki/Shot_noise\" target=\"_blank\">Wikipedia: Shot noise</a>). For the depth predictions, the variance contributed by the shot noise is inversely proportional to the signal strength.</p>\n<h2>Implementation</h2>\n<p>If we want to create data for a new star which is less bright than the existing ones, we can multiply all the data by a scaling factor &lt; 1. This multiplication reduces the variance by the square of the scaling factor so that the new variance is too low, and we need to compensate by injecting some additional variance.</p>\n<pre><code> ():\n    \n    old_brightness = data.mean(axis=).mean(axis=)\n    scaling = new_brightness / old_brightness.reshape(-, , )\n    new_data = data * scaling\n    time_binning =  \n    rng = np.random.default_rng()\n    new_data += rng.normal(scale=np.sqrt(( - scaling) * new_data / time_binning),\n                           size=new_data) \n     new_data\n</code></pre>\n<p>This implementation only works for downscaling but not for upscaling. Fortunately the stars in the test dataset are all less bright then Star 1 and we don't need to upscale the data.</p>\n<h2>Conclusion</h2>\n<p>Data augmentation has helped me tune the predictions for the sigma values.</p>",
      "rawMarkdown": "I'm surprised that data augmentation hasn't been used more widely in this competition. Here's my approach.\n\n## Background: Dim stars make noisy pictures\n\nIf you've ever taken a picture with your mobile phone in low-light conditions, you know the effect: The camera adapts its shutter speed, but the picture will be noisier. In our dataset, the main feature of a star is its brightness, and the less bright a star is, the more noise we get in its picture: Star 1 is brighter than Star 0, and the unseen stars in the test set are less bright than Star 0. Furthermore, the stars emit less energy at higher wavelengths, which means that the right part of the spectra is noisier than the left one.\n\nThe effect can be quantified: The variance of the signal is proportional to the signal itself (see [Wikipedia: Shot noise](https://en.wikipedia.org/wiki/Shot_noise)). For the depth predictions, the variance contributed by the shot noise is inversely proportional to the signal strength.\n\n## Implementation\n\nIf we want to create data for a new star which is less bright than the existing ones, we can multiply all the data by a scaling factor < 1. This multiplication reduces the variance by the square of the scaling factor so that the new variance is too low, and we need to compensate by injecting some additional variance.\n\n```\ndef scale_planets(data, new_brightness):\n    \"\"\"Create the dataset for a new star with lower brightness.\n    \n    Parameters\n    data: array of shape (n_planets, n_timesteps, n_wavelengths)\n    new_brightness: float (desired brightness for the new star)\n    \n    Return value\n    new_data: array of shape (n_planets, n_timesteps, n_wavelengths)\n    \"\"\"\n    old_brightness = data.mean(axis=2).mean(axis=1)\n    scaling = new_brightness / old_brightness.reshape(-1, 1, 1)\n    new_data = data * scaling\n    time_binning = 25 # I downscaled the time dimension from 5625 to 225\n    rng = np.random.default_rng()\n    new_data += rng.normal(scale=np.sqrt((1 - scaling) * new_data / time_binning),\n                           size=new_data) # Increase the variance\n    return new_data\n```\n\nThis implementation only works for downscaling but not for upscaling. Fortunately the stars in the test dataset are all less bright then Star 1 and we don't need to upscale the data.\n\n## Conclusion\n\nData augmentation has helped me tune the predictions for the sigma values.\n    ",
      "votes": 22
    },
    {
      "id": 3034504,
      "postDate": "2024-11-02T07:53:43.673Z",
      "content": "<p>Interesting approach! I also wanted to take the opportunity here to say thank you for your threads and posts. You always share an incredible amount in the competitions in which you are active and thus help to raise the level in general. </p>\n<p>As far as I remember, you were the first to “publicly” point out that pure modeling can solve the problem to a large extent. Your training pipeline was certainly the first starting point for many participants. Thank you very much!</p>",
      "rawMarkdown": "Interesting approach! I also wanted to take the opportunity here to say thank you for your threads and posts. You always share an incredible amount in the competitions in which you are active and thus help to raise the level in general. \n\nAs far as I remember, you were the first to “publicly” point out that pure modeling can solve the problem to a large extent. Your training pipeline was certainly the first starting point for many participants. Thank you very much!",
      "votes": 3
    },
    {
      "id": 3036212,
      "postDate": "2024-11-04T11:04:09.043Z",
      "content": "<p>I tried to use augmentation for DL model. But, ultimately, classic methods were much more accurate, and for classic methods, augmentation is meaningless.</p>",
      "rawMarkdown": "I tried to use augmentation for DL model. But, ultimately, classic methods were much more accurate, and for classic methods, augmentation is meaningless.",
      "votes": 1,
      "replies": [
        {
          "id": 3248269,
          "postDate": "2025-07-14T08:30:43.930Z",
          "content": "<p>Is this why you shared how to augment data in this year comp? :D</p>",
          "rawMarkdown": "Is this why you shared how to augment data in this year comp? :D"
        }
      ]
    },
    {
      "id": 3034555,
      "postDate": "2024-11-02T09:32:56.587Z",
      "content": "<p>Thank you for sharing!<br>\nThats's simple, but smart augmentation approach, I didn't come up with at all…</p>",
      "rawMarkdown": "Thank you for sharing!\nThats's simple, but smart augmentation approach, I didn't come up with at all...",
      "votes": 1
    },
    {
      "id": 3035878,
      "postDate": "2024-11-04T00:49:55.833Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F286455%2F8454b2a4c7619c638061c63bd29cb31a%2FScreenshot%202024-11-03%20at%2016.48.06.png?generation=1730681350133755&amp;alt=media\" alt=\"\"></p>\n<p>Added <a href=\"https://explore.albumentations.ai/transform/ShotNoise\" target=\"_blank\">ShotNoise</a> to <a href=\"https://albumentations.ai/\" target=\"_blank\">Albumentations</a></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F286455%2F8454b2a4c7619c638061c63bd29cb31a%2FScreenshot%202024-11-03%20at%2016.48.06.png?generation=1730681350133755&alt=media)\n\nAdded [ShotNoise](https://explore.albumentations.ai/transform/ShotNoise) to [Albumentations](https://albumentations.ai/)",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 3034504,
      "author_name": "Benedikt Droste",
      "author_url": "",
      "post_date": "2024-11-02T07:53:43.673000",
      "content": "<p>Interesting approach! I also wanted to take the opportunity here to say thank you for your threads and posts. You always share an incredible amount in the competitions in which you are active and thus help to raise the level in general. </p>\n<p>As far as I remember, you were the first to “publicly” point out that pure modeling can solve the problem to a large extent. Your training pipeline was certainly the first starting point for many participants. Thank you very much!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3036212,
      "author_name": "greySnow",
      "author_url": "",
      "post_date": "2024-11-04T11:04:09.043000",
      "content": "<p>I tried to use augmentation for DL model. But, ultimately, classic methods were much more accurate, and for classic methods, augmentation is meaningless.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3248269,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2025-07-14T08:30:43.930000",
          "content": "<p>Is this why you shared how to augment data in this year comp? :D</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3034555,
      "author_name": "kyu999",
      "author_url": "",
      "post_date": "2024-11-02T09:32:56.587000",
      "content": "<p>Thank you for sharing!<br>\nThats's simple, but smart augmentation approach, I didn't come up with at all…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3035878,
      "author_name": "Vladimir Iglovikov",
      "author_url": "",
      "post_date": "2024-11-04T00:49:55.833000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F286455%2F8454b2a4c7619c638061c63bd29cb31a%2FScreenshot%202024-11-03%20at%2016.48.06.png?generation=1730681350133755&amp;alt=media\" alt=\"\"></p>\n<p>Added <a href=\"https://explore.albumentations.ai/transform/ShotNoise\" target=\"_blank\">ShotNoise</a> to <a href=\"https://albumentations.ai/\" target=\"_blank\">Albumentations</a></p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3034498": "I'm surprised that data augmentation hasn't been used more widely in this competition. Here's my approach.\n\n## Background: Dim stars make noisy pictures\n\nIf you've ever taken a picture with your mobile phone in low-light conditions, you know the effect: The camera adapts its shutter speed, but the picture will be noisier. In our dataset, the main feature of a star is its brightness, and the less bright a star is, the more noise we get in its picture: Star 1 is brighter than Star 0, and the unseen stars in the test set are less bright than Star 0. Furthermore, the stars emit less energy at higher wavelengths, which means that the right part of the spectra is noisier than the left one.\n\nThe effect can be quantified: The variance of the signal is proportional to the signal itself (see [Wikipedia: Shot noise](https://en.wikipedia.org/wiki/Shot_noise)). For the depth predictions, the variance contributed by the shot noise is inversely proportional to the signal strength.\n\n## Implementation\n\nIf we want to create data for a new star which is less bright than the existing ones, we can multiply all the data by a scaling factor < 1. This multiplication reduces the variance by the square of the scaling factor so that the new variance is too low, and we need to compensate by injecting some additional variance.\n\n```\ndef scale_planets(data, new_brightness):\n    \"\"\"Create the dataset for a new star with lower brightness.\n    \n    Parameters\n    data: array of shape (n_planets, n_timesteps, n_wavelengths)\n    new_brightness: float (desired brightness for the new star)\n    \n    Return value\n    new_data: array of shape (n_planets, n_timesteps, n_wavelengths)\n    \"\"\"\n    old_brightness = data.mean(axis=2).mean(axis=1)\n    scaling = new_brightness / old_brightness.reshape(-1, 1, 1)\n    new_data = data * scaling\n    time_binning = 25 # I downscaled the time dimension from 5625 to 225\n    rng = np.random.default_rng()\n    new_data += rng.normal(scale=np.sqrt((1 - scaling) * new_data / time_binning),\n                           size=new_data) # Increase the variance\n    return new_data\n```\n\nThis implementation only works for downscaling but not for upscaling. Fortunately the stars in the test dataset are all less bright then Star 1 and we don't need to upscale the data.\n\n## Conclusion\n\nData augmentation has helped me tune the predictions for the sigma values.\n    ",
    "3034504": "Interesting approach! I also wanted to take the opportunity here to say thank you for your threads and posts. You always share an incredible amount in the competitions in which you are active and thus help to raise the level in general. \n\nAs far as I remember, you were the first to “publicly” point out that pure modeling can solve the problem to a large extent. Your training pipeline was certainly the first starting point for many participants. Thank you very much!",
    "3036212": "I tried to use augmentation for DL model. But, ultimately, classic methods were much more accurate, and for classic methods, augmentation is meaningless.",
    "3034555": "Thank you for sharing!\nThats's simple, but smart augmentation approach, I didn't come up with at all...",
    "3035878": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F286455%2F8454b2a4c7619c638061c63bd29cb31a%2FScreenshot%202024-11-03%20at%2016.48.06.png?generation=1730681350133755&alt=media)\n\nAdded [ShotNoise](https://explore.albumentations.ai/transform/ShotNoise) to [Albumentations](https://albumentations.ai/)"
  }
}