{
  "id": 210438,
  "title": "Bounding box normalization ValueError",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/210438",
  "author_name": "InDSweTrust",
  "post_date": "2021-01-10T19:58:12.736000",
  "votes": 5,
  "comment_count": 9,
  "views": 0,
  "content": "<p>During bounding box normalization (with albumentations library), you may run into a <code>ValueError</code> caused by python floating point format which may cause out of range values &gt; 1.0, for example:</p>\n<p><code>ValueError: Expected y_max for bbox (0.029610556807209528, 0.32157104863221886, 0.12037335049887352, 1.0001213694697737, tensor(10)) to be in the range [0.0, 1.0], got 1.0001213694697737.</code></p>\n<p>I am using pre-processed resized  images and the bounding boxes are re-sized accordingly. This may occur when the bounding box is very close or right at the edge of the image boundary.</p>\n<p>The solution to this issue that I found so far is to modify the <code>def normalize_bbox</code> method in <code>bbox_utils.py</code> (albumentations) as follows:</p>\n<p><strong>Original</strong><br>\n<code>x_min, x_max = x_min / cols, x_max / cols</code><br>\n<code>y_min, y_max = y_min / rows, y_max / rows</code></p>\n<p><strong>Modified</strong><br>\n<code>x_min, x_max = min(x_min / cols, 1.0), min(x_max / cols, 1.0)</code><br>\n<code>y_min, y_max = min(y_min / rows, 1.0), min(y_max / rows, 1.0)</code></p>\n<p>As a result, floats such as 1.000121369…are converted to 1.0. <br>\nDid anyone else face this issue? If yes, maybe someone implemented an alternative solution that doesn't involve <code>bbox_utils.py</code> modification?</p>",
  "messages": [
    {
      "id": 1147918,
      "postDate": "2021-01-10T19:58:12.737Z",
      "content": "<p>During bounding box normalization (with albumentations library), you may run into a <code>ValueError</code> caused by python floating point format which may cause out of range values &gt; 1.0, for example:</p>\n<p><code>ValueError: Expected y_max for bbox (0.029610556807209528, 0.32157104863221886, 0.12037335049887352, 1.0001213694697737, tensor(10)) to be in the range [0.0, 1.0], got 1.0001213694697737.</code></p>\n<p>I am using pre-processed resized  images and the bounding boxes are re-sized accordingly. This may occur when the bounding box is very close or right at the edge of the image boundary.</p>\n<p>The solution to this issue that I found so far is to modify the <code>def normalize_bbox</code> method in <code>bbox_utils.py</code> (albumentations) as follows:</p>\n<p><strong>Original</strong><br>\n<code>x_min, x_max = x_min / cols, x_max / cols</code><br>\n<code>y_min, y_max = y_min / rows, y_max / rows</code></p>\n<p><strong>Modified</strong><br>\n<code>x_min, x_max = min(x_min / cols, 1.0), min(x_max / cols, 1.0)</code><br>\n<code>y_min, y_max = min(y_min / rows, 1.0), min(y_max / rows, 1.0)</code></p>\n<p>As a result, floats such as 1.000121369…are converted to 1.0. <br>\nDid anyone else face this issue? If yes, maybe someone implemented an alternative solution that doesn't involve <code>bbox_utils.py</code> modification?</p>",
      "rawMarkdown": "During bounding box normalization (with albumentations library), you may run into a `ValueError ` caused by python floating point format which may cause out of range values > 1.0, for example:\n\n`ValueError: Expected y_max for bbox (0.029610556807209528, 0.32157104863221886, 0.12037335049887352, 1.0001213694697737, tensor(10)) to be in the range [0.0, 1.0], got 1.0001213694697737.`\n\nI am using pre-processed resized  images and the bounding boxes are re-sized accordingly. This may occur when the bounding box is very close or right at the edge of the image boundary.\n\nThe solution to this issue that I found so far is to modify the `def normalize_bbox` method in `bbox_utils.py` (albumentations) as follows:\n\n**Original**\n`x_min, x_max = x_min / cols, x_max / cols`\n`y_min, y_max = y_min / rows, y_max / rows`\n\n**Modified**\n`x_min, x_max = min(x_min / cols, 1.0), min(x_max / cols, 1.0)`\n`y_min, y_max = min(y_min / rows, 1.0), min(y_max / rows, 1.0)`\n\nAs a result, floats such as 1.000121369...are converted to 1.0. \nDid anyone else face this issue? If yes, maybe someone implemented an alternative solution that doesn't involve `bbox_utils.py` modification?",
      "votes": 5
    },
    {
      "id": 1147960,
      "postDate": "2021-01-10T20:26:20.237Z",
      "content": "<p>I haven't checked the pre-processed, resized dataset yet, but the resized box should not be greater than one (or the resized width/height of the image). So, the problem is probably there. Can you share a link to the dataset?</p>",
      "rawMarkdown": "I haven't checked the pre-processed, resized dataset yet, but the resized box should not be greater than one (or the resized width/height of the image). So, the problem is probably there. Can you share a link to the dataset?\n",
      "votes": 1,
      "replies": [
        {
          "id": 1147983,
          "postDate": "2021-01-10T20:40:15.617Z",
          "content": "<p>I created the dataset myself - IMG size 512 with the original aspect ratio retained. But I created the dataset based on <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207955\" target=\"_blank\">this discussion</a> and I see the author uploaded a dataset of the same size <a href=\"https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-png-512px-original-ratio\" target=\"_blank\">here</a>. It should be the same as the dataset I used.</p>\n<p>That was my thought originally, but while checking the data, bounding box dimensions seemed OK…they do not seem to exceed image width or height. Have you tried working with pre-processed images or are you using the FastRCNN to resize?</p>",
          "rawMarkdown": "I created the dataset myself - IMG size 512 with the original aspect ratio retained. But I created the dataset based on [this discussion](https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207955) and I see the author uploaded a dataset of the same size [here](https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-png-512px-original-ratio). It should be the same as the dataset I used.\n\nThat was my thought originally, but while checking the data, bounding box dimensions seemed OK...they do not seem to exceed image width or height. Have you tried working with pre-processed images or are you using the FastRCNN to resize?",
          "votes": 1
        },
        {
          "id": 1148043,
          "postDate": "2021-01-10T21:17:45.510Z",
          "content": "<p>If the box dimensions are ok, then double-check the bbox_param format you pass to the Compose (pascal_voc should be in min/max, coco in xy/wh; default is normalized min/max). I usually mix these. Other than that, I'd do the same thing you did, but I'd fix the dimensions in the data loader, not in Albumentation.</p>\n<p>I use the dicom files, convert them to 800px with Albumentation <code>A.LongestMaxSize(max_size=800, p=1.0)</code> and I use the default FasterRCNN transforms (min-max: 800-1333; normalize with imagenet stats)</p>",
          "rawMarkdown": "If the box dimensions are ok, then double-check the bbox_param format you pass to the Compose (pascal_voc should be in min/max, coco in xy/wh; default is normalized min/max). I usually mix these. Other than that, I'd do the same thing you did, but I'd fix the dimensions in the data loader, not in Albumentation.\n\nI use the dicom files, convert them to 800px with Albumentation `A.LongestMaxSize(max_size=800, p=1.0)` and I use the default FasterRCNN transforms (min-max: 800-1333; normalize with imagenet stats)\n",
          "votes": 1
        },
        {
          "id": 1148065,
          "postDate": "2021-01-10T21:43:18.270Z",
          "content": "<p>Yes, bbox_param format is one of the first things I checked - currently passing pascal_voc, values are correctly min/max. I haven't tried other formats (or mixing), will do a bit later. <br>\nImplementing the max dim loader fix in dataloader is a good idea but I am not sure if its possible in this case…the bbox normalization (and the Value check) is done in Albumentations, so the error occurs before the normalized data gets back to the Dataset class. <br>\nI think the only way to avoid the mod in <code>bbox_utils.py</code> is to round down the upper bound bbox values in the data loader. But I find this method to be more dangerous ;)</p>\n<p>In my case resizing images with FasterRCNN was slower and bottlenecking my GPU…maybe I did something wrong (might try again later). But for now, using pre-processed images.</p>",
          "rawMarkdown": "Yes, bbox_param format is one of the first things I checked - currently passing pascal_voc, values are correctly min/max. I haven't tried other formats (or mixing), will do a bit later. \nImplementing the max dim loader fix in dataloader is a good idea but I am not sure if its possible in this case...the bbox normalization (and the Value check) is done in Albumentations, so the error occurs before the normalized data gets back to the Dataset class. \nI think the only way to avoid the mod in `bbox_utils.py` is to round down the upper bound bbox values in the data loader. But I find this method to be more dangerous ;)\n\nIn my case resizing images with FasterRCNN was slower and bottlenecking my GPU...maybe I did something wrong (might try again later). But for now, using pre-processed images.",
          "votes": 1
        },
        {
          "id": 1154139,
          "postDate": "2021-01-15T12:37:19.660Z",
          "content": "<p>is the problem solved? i have fixed this issue?</p>",
          "rawMarkdown": "is the problem solved? i have fixed this issue?"
        }
      ]
    },
    {
      "id": 1751011,
      "postDate": "2022-04-10T09:36:35.167Z",
      "content": "<p>Hi, have you solved this problem?</p>",
      "rawMarkdown": "Hi, have you solved this problem?\n\n"
    },
    {
      "id": 1521083,
      "postDate": "2021-09-22T21:23:11.790Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> , do you have a solution for this?</p>",
      "rawMarkdown": "Hi @morizin , do you have a solution for this?"
    },
    {
      "id": 1521082,
      "postDate": "2021-09-22T21:22:07.953Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> </p>\n<p>I am trying to train DETR on 1024 x 1024  dataset ( <a href=\"https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024\" target=\"_blank\">https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024</a> ) </p>\n<p>Its working fine for 512 x 512 but i am getting below error in case of 1024 x 1024:</p>\n<p><strong>ValueError: Expected x_max for bbox (0.5952381060633343, 0.578276682732394, 1.0590986667957623, 0.7091712844849098, 3) to be in the range [0.0, 1.0], got 1.0590986667957623.</strong></p>\n<p>Moreover i am scaling the bounding boxes according to the new resolution.</p>\n<p>Any idea what i might be missing?<br>\nThank you.</p>",
      "rawMarkdown": "Hi @indswetrust @pestipeti \n\nI am trying to train DETR on 1024 x 1024  dataset ( https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024 ) \n\nIts working fine for 512 x 512 but i am getting below error in case of 1024 x 1024:\n\n**ValueError: Expected x_max for bbox (0.5952381060633343, 0.578276682732394, 1.0590986667957623, 0.7091712844849098, 3) to be in the range [0.0, 1.0], got 1.0590986667957623.**\n\nMoreover i am scaling the bounding boxes according to the new resolution.\n\nAny idea what i might be missing?\nThank you."
    },
    {
      "id": 1254993,
      "postDate": "2021-03-28T10:14:27.397Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1147960,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2021-01-10T20:26:20.237000",
      "content": "<p>I haven't checked the pre-processed, resized dataset yet, but the resized box should not be greater than one (or the resized width/height of the image). So, the problem is probably there. Can you share a link to the dataset?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1147983,
          "author_name": "InDSweTrust",
          "author_url": "",
          "post_date": "2021-01-10T20:40:15.617000",
          "content": "<p>I created the dataset myself - IMG size 512 with the original aspect ratio retained. But I created the dataset based on <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/207955\" target=\"_blank\">this discussion</a> and I see the author uploaded a dataset of the same size <a href=\"https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-png-512px-original-ratio\" target=\"_blank\">here</a>. It should be the same as the dataset I used.</p>\n<p>That was my thought originally, but while checking the data, bounding box dimensions seemed OK…they do not seem to exceed image width or height. Have you tried working with pre-processed images or are you using the FastRCNN to resize?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1148043,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2021-01-10T21:17:45.510000",
          "content": "<p>If the box dimensions are ok, then double-check the bbox_param format you pass to the Compose (pascal_voc should be in min/max, coco in xy/wh; default is normalized min/max). I usually mix these. Other than that, I'd do the same thing you did, but I'd fix the dimensions in the data loader, not in Albumentation.</p>\n<p>I use the dicom files, convert them to 800px with Albumentation <code>A.LongestMaxSize(max_size=800, p=1.0)</code> and I use the default FasterRCNN transforms (min-max: 800-1333; normalize with imagenet stats)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1148065,
          "author_name": "InDSweTrust",
          "author_url": "",
          "post_date": "2021-01-10T21:43:18.270000",
          "content": "<p>Yes, bbox_param format is one of the first things I checked - currently passing pascal_voc, values are correctly min/max. I haven't tried other formats (or mixing), will do a bit later. <br>\nImplementing the max dim loader fix in dataloader is a good idea but I am not sure if its possible in this case…the bbox normalization (and the Value check) is done in Albumentations, so the error occurs before the normalized data gets back to the Dataset class. <br>\nI think the only way to avoid the mod in <code>bbox_utils.py</code> is to round down the upper bound bbox values in the data loader. But I find this method to be more dangerous ;)</p>\n<p>In my case resizing images with FasterRCNN was slower and bottlenecking my GPU…maybe I did something wrong (might try again later). But for now, using pre-processed images.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1154139,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-01-15T12:37:19.660000",
          "content": "<p>is the problem solved? i have fixed this issue?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1751011,
      "author_name": "Josekinkin",
      "author_url": "",
      "post_date": "2022-04-10T09:36:35.167000",
      "content": "<p>Hi, have you solved this problem?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1521083,
      "author_name": "Yasir Irfan",
      "author_url": "",
      "post_date": "2021-09-22T21:23:11.790000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> , do you have a solution for this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1521082,
      "author_name": "Yasir Irfan",
      "author_url": "",
      "post_date": "2021-09-22T21:22:07.953000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/indswetrust\" target=\"_blank\">@indswetrust</a> <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a> </p>\n<p>I am trying to train DETR on 1024 x 1024  dataset ( <a href=\"https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024\" target=\"_blank\">https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024</a> ) </p>\n<p>Its working fine for 512 x 512 but i am getting below error in case of 1024 x 1024:</p>\n<p><strong>ValueError: Expected x_max for bbox (0.5952381060633343, 0.578276682732394, 1.0590986667957623, 0.7091712844849098, 3) to be in the range [0.0, 1.0], got 1.0590986667957623.</strong></p>\n<p>Moreover i am scaling the bounding boxes according to the new resolution.</p>\n<p>Any idea what i might be missing?<br>\nThank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1254993,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-28T10:14:27.397000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1147918": "During bounding box normalization (with albumentations library), you may run into a `ValueError ` caused by python floating point format which may cause out of range values > 1.0, for example:\n\n`ValueError: Expected y_max for bbox (0.029610556807209528, 0.32157104863221886, 0.12037335049887352, 1.0001213694697737, tensor(10)) to be in the range [0.0, 1.0], got 1.0001213694697737.`\n\nI am using pre-processed resized  images and the bounding boxes are re-sized accordingly. This may occur when the bounding box is very close or right at the edge of the image boundary.\n\nThe solution to this issue that I found so far is to modify the `def normalize_bbox` method in `bbox_utils.py` (albumentations) as follows:\n\n**Original**\n`x_min, x_max = x_min / cols, x_max / cols`\n`y_min, y_max = y_min / rows, y_max / rows`\n\n**Modified**\n`x_min, x_max = min(x_min / cols, 1.0), min(x_max / cols, 1.0)`\n`y_min, y_max = min(y_min / rows, 1.0), min(y_max / rows, 1.0)`\n\nAs a result, floats such as 1.000121369...are converted to 1.0. \nDid anyone else face this issue? If yes, maybe someone implemented an alternative solution that doesn't involve `bbox_utils.py` modification?",
    "1147960": "I haven't checked the pre-processed, resized dataset yet, but the resized box should not be greater than one (or the resized width/height of the image). So, the problem is probably there. Can you share a link to the dataset?\n",
    "1751011": "Hi, have you solved this problem?\n\n",
    "1521083": "Hi @morizin , do you have a solution for this?",
    "1521082": "Hi @indswetrust @pestipeti \n\nI am trying to train DETR on 1024 x 1024  dataset ( https://www.kaggle.com/xhlulu/vinbigdata-chest-xray-resized-png-1024x1024 ) \n\nIts working fine for 512 x 512 but i am getting below error in case of 1024 x 1024:\n\n**ValueError: Expected x_max for bbox (0.5952381060633343, 0.578276682732394, 1.0590986667957623, 0.7091712844849098, 3) to be in the range [0.0, 1.0], got 1.0590986667957623.**\n\nMoreover i am scaling the bounding boxes according to the new resolution.\n\nAny idea what i might be missing?\nThank you.",
    "1254993": ""
  }
}