{
  "id": 500816,
  "title": "The risk of underflow when converting to FP32",
  "url": "/competitions/leap-atmospheric-physics-ai-climsim/discussion/500816",
  "author_name": "Ryota",
  "post_date": "2024-05-07T03:42:43.129000",
  "votes": 32,
  "comment_count": 4,
  "views": 0,
  "content": "<p>In this competition, many participants are converting data to FP32 due to the immense size of the dataset. However, this poses a risk of underflow.</p>\n<p>Let's take a look at state_q0002_13, for example. It appears in the following table:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2F8f74fb5853c30e77af163728c37f3b8f%2F2024-05-07%2012.25.08.png?generation=1715052323136913&amp;alt=media\"></p>\n<p>The minimum value for FP32, even when including subnormal numbers, is <code>1.401298e-45</code>. If a number falls below this value, underflow occurs, and it is rounded to 0.<br>\nFurthermore, subnormal numbers experience a decrease in the number of significant digits as their values become smaller. Therefore, it is desirable to keep values within the range of normal numbers whenever possible, with the minimum normal number being <code>1.175494e-38</code>.</p>\n<p>According to my EDA, many other columns that also contain values smaller than <code>1.175494e-38</code>. For these columns, it would be advisable to scale up the values by multiplying them by 10^N or applying a similar scaling factor, although the extent to which this will impact the score is still uncertain.</p>",
  "messages": [
    {
      "id": 2797931,
      "postDate": "2024-05-07T03:42:43.130Z",
      "content": "<p>In this competition, many participants are converting data to FP32 due to the immense size of the dataset. However, this poses a risk of underflow.</p>\n<p>Let's take a look at state_q0002_13, for example. It appears in the following table:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2F8f74fb5853c30e77af163728c37f3b8f%2F2024-05-07%2012.25.08.png?generation=1715052323136913&amp;alt=media\"></p>\n<p>The minimum value for FP32, even when including subnormal numbers, is <code>1.401298e-45</code>. If a number falls below this value, underflow occurs, and it is rounded to 0.<br>\nFurthermore, subnormal numbers experience a decrease in the number of significant digits as their values become smaller. Therefore, it is desirable to keep values within the range of normal numbers whenever possible, with the minimum normal number being <code>1.175494e-38</code>.</p>\n<p>According to my EDA, many other columns that also contain values smaller than <code>1.175494e-38</code>. For these columns, it would be advisable to scale up the values by multiplying them by 10^N or applying a similar scaling factor, although the extent to which this will impact the score is still uncertain.</p>",
      "rawMarkdown": "In this competition, many participants are converting data to FP32 due to the immense size of the dataset. However, this poses a risk of underflow.\n\nLet's take a look at state_q0002_13, for example. It appears in the following table:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2F8f74fb5853c30e77af163728c37f3b8f%2F2024-05-07%2012.25.08.png?generation=1715052323136913&alt=media)\n\nThe minimum value for FP32, even when including subnormal numbers, is ``1.401298e-45``. If a number falls below this value, underflow occurs, and it is rounded to 0.\nFurthermore, subnormal numbers experience a decrease in the number of significant digits as their values become smaller. Therefore, it is desirable to keep values within the range of normal numbers whenever possible, with the minimum normal number being ``1.175494e-38``.\n\nAccording to my EDA, many other columns that also contain values smaller than ``1.175494e-38``. For these columns, it would be advisable to scale up the values by multiplying them by 10^N or applying a similar scaling factor, although the extent to which this will impact the score is still uncertain.",
      "votes": 31
    },
    {
      "id": 2798613,
      "postDate": "2024-05-07T10:03:16.703Z",
      "content": "<p>Maybe scaling the float64 data with StandardScaler and then storing it as float32 could solve this problem?</p>",
      "rawMarkdown": "Maybe scaling the float64 data with StandardScaler and then storing it as float32 could solve this problem?",
      "votes": 8,
      "replies": [
        {
          "id": 2798643,
          "postDate": "2024-05-07T10:20:15.147Z",
          "content": "<p>You're right, that approach does allow us to avoid this issue.</p>",
          "rawMarkdown": "You're right, that approach does allow us to avoid this issue.",
          "replies": [
            {
              "id": 2801666,
              "postDate": "2024-05-08T17:35:49.330Z",
              "content": "<p>I tried this approach and my public score dropped by 0.1.<br>\nI will continue investigating</p>",
              "rawMarkdown": "I tried this approach and my public score dropped by 0.1.\nI will continue investigating"
            },
            {
              "id": 2803070,
              "postDate": "2024-05-09T10:16:59.557Z",
              "content": "<p>I am confident that this is not a problem. Please check your code.</p>",
              "rawMarkdown": "I am confident that this is not a problem. Please check your code.",
              "votes": 2
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2798613,
      "author_name": "Zhuoqun Li",
      "author_url": "",
      "post_date": "2024-05-07T10:03:16.703000",
      "content": "<p>Maybe scaling the float64 data with StandardScaler and then storing it as float32 could solve this problem?</p>",
      "votes": 8,
      "replies": [
        {
          "id": 2798643,
          "author_name": "Ryota",
          "author_url": "",
          "post_date": "2024-05-07T10:20:15.147000",
          "content": "<p>You're right, that approach does allow us to avoid this issue.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2801666,
              "author_name": "Fernando Melo",
              "author_url": "",
              "post_date": "2024-05-08T17:35:49.330000",
              "content": "<p>I tried this approach and my public score dropped by 0.1.<br>\nI will continue investigating</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2803070,
              "author_name": "Zhuoqun Li",
              "author_url": "",
              "post_date": "2024-05-09T10:16:59.557000",
              "content": "<p>I am confident that this is not a problem. Please check your code.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2797931": "In this competition, many participants are converting data to FP32 due to the immense size of the dataset. However, this poses a risk of underflow.\n\nLet's take a look at state_q0002_13, for example. It appears in the following table:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3845682%2F8f74fb5853c30e77af163728c37f3b8f%2F2024-05-07%2012.25.08.png?generation=1715052323136913&alt=media)\n\nThe minimum value for FP32, even when including subnormal numbers, is ``1.401298e-45``. If a number falls below this value, underflow occurs, and it is rounded to 0.\nFurthermore, subnormal numbers experience a decrease in the number of significant digits as their values become smaller. Therefore, it is desirable to keep values within the range of normal numbers whenever possible, with the minimum normal number being ``1.175494e-38``.\n\nAccording to my EDA, many other columns that also contain values smaller than ``1.175494e-38``. For these columns, it would be advisable to scale up the values by multiplying them by 10^N or applying a similar scaling factor, although the extent to which this will impact the score is still uncertain.",
    "2798613": "Maybe scaling the float64 data with StandardScaler and then storing it as float32 could solve this problem?"
  }
}