{
  "id": 208781,
  "title": "How to present the images that have multiple classes",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/208781",
  "author_name": "Weicong",
  "post_date": "2021-01-05T02:28:37.504000",
  "votes": 10,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We know the submission format is like \"14 1 0 0 1 1\" if no found in the image. However, we can find a lot of training images have multiple classes (different diseases) in the single x-ray image. In this case, how to present it in our submission? Should we use multiple rows?</p>",
  "messages": [
    {
      "id": 1138845,
      "postDate": "2021-01-05T02:28:37.503Z",
      "content": "<p>We know the submission format is like \"14 1 0 0 1 1\" if no found in the image. However, we can find a lot of training images have multiple classes (different diseases) in the single x-ray image. In this case, how to present it in our submission? Should we use multiple rows?</p>",
      "rawMarkdown": "We know the submission format is like \"14 1 0 0 1 1\" if no found in the image. However, we can find a lot of training images have multiple classes (different diseases) in the single x-ray image. In this case, how to present it in our submission? Should we use multiple rows?",
      "votes": 10
    },
    {
      "id": 1138903,
      "postDate": "2021-01-05T03:52:47.673Z",
      "content": "<p>Good question. From the competition description:</p>\n<blockquote>\n  <p>Images in the test set may contain more than one object. For each object in a given test image, you must predict a class ID, confidence score, and bounding box in format xmin ymin xmax ymax.</p>\n</blockquote>\n<p>Underneath that they have an example of an image with multiple objects, which looks like this in the submission file:</p>\n<pre><code>ID,TARGET\n004f33259ee4aef671c2b95d54e4be69,11 0.5 100 100 200 200 13 0.7 10 10 20 20\n</code></pre>\n<p>So this means that there are two predictions for image id <code>004f33259ee4aef671c2b95d54e4be69</code>:</p>\n<ol>\n<li>Object of class <code>11</code> (pleural thickening) with a confidence of <code>0.5</code> with a bounding box of <code>100</code> (xmin), <code>100</code> (ymin), <code>200</code> (xmax), and <code>200</code> (ymax).</li>\n<li>Object of class <code>13</code> (pulmonary fibrosis) with a confidence of <code>0.7</code> with a bounding box of <code>10</code> (xmin), <code>10</code> (ymin), <code>20</code> (xmax), and <code>20</code> (ymax).</li>\n</ol>\n<p>It looks like that second column <code>TARGET</code> is just a string with those numbers in it (class id, confidence, xmin, ymin, xmax, ymax) separated by a space. Multiple object predictions are just concatenated to that string with a space between each.</p>",
      "rawMarkdown": "Good question. From the competition description:\n\n> Images in the test set may contain more than one object. For each object in a given test image, you must predict a class ID, confidence score, and bounding box in format xmin ymin xmax ymax.\n\nUnderneath that they have an example of an image with multiple objects, which looks like this in the submission file:\n\n    ID,TARGET\n    004f33259ee4aef671c2b95d54e4be69,11 0.5 100 100 200 200 13 0.7 10 10 20 20\n\nSo this means that there are two predictions for image id `004f33259ee4aef671c2b95d54e4be69`:\n\n1. Object of class `11` (pleural thickening) with a confidence of `0.5` with a bounding box of `100` (xmin), `100` (ymin), `200` (xmax), and `200` (ymax).\n2. Object of class `13` (pulmonary fibrosis) with a confidence of `0.7` with a bounding box of `10` (xmin), `10` (ymin), `20` (xmax), and `20` (ymax).\n\nIt looks like that second column `TARGET` is just a string with those numbers in it (class id, confidence, xmin, ymin, xmax, ymax) separated by a space. Multiple object predictions are just concatenated to that string with a space between each.",
      "votes": 5,
      "replies": [
        {
          "id": 1139736,
          "postDate": "2021-01-05T15:33:31.090Z",
          "content": "<p>Thanks, Thomas! You gave a detailed explanation. Another question: do you know the directions of x_min, y_min? </p>\n<p>I find the shape of the image is (vertical axis, horizontal axis), something like the matrix we use in python, rather than a math chart. Meanwhile, checking the image, we can find the origin (0,0) is at the left top. Do you know the meaning of x_min, y_min, x_max, y_max? For example, x_min = 300. Should I count it from the left top down 300 pixes, or go right 300 pixes?</p>\n<p>Thanks for your patient and sharing!</p>\n<p>Best,</p>\n<p>Weicong</p>",
          "rawMarkdown": "Thanks, Thomas! You gave a detailed explanation. Another question: do you know the directions of x_min, y_min? \n\nI find the shape of the image is (vertical axis, horizontal axis), something like the matrix we use in python, rather than a math chart. Meanwhile, checking the image, we can find the origin (0,0) is at the left top. Do you know the meaning of x_min, y_min, x_max, y_max? For example, x_min = 300. Should I count it from the left top down 300 pixes, or go right 300 pixes?\n\nThanks for your patient and sharing!\n\n\nBest,\n\nWeicong",
          "votes": 1
        },
        {
          "id": 1140210,
          "postDate": "2021-01-05T21:17:55.263Z",
          "content": "<p><a href=\"https://www.kaggle.com/weicongf\" target=\"_blank\">@weicongf</a> here's my understanding of <code>x_min</code> and <code>y_min</code>. As you stated, the x-ray images have an <code>(x, y)</code> origin of <code>(0, 0)</code> - meaning the top left hand corner is the origin of the image. With that in mind:</p>\n<ul>\n<li>The combination of <code>(x_min, y_min)</code> represent the upper left hand corner of the bounding box.</li>\n<li>The combination of <code>(x_max, y_max)</code> represent the lower right hand corner of the bounding box.</li>\n<li>Both <code>x_min</code> and <code>x_max</code> are measured from the left hand side of the image. So <code>x_min</code> of 300 is 300 pixels away from the left hand side of the image, and <code>x_max</code> of 400 is 400 pixels away from the left hand side of the image.</li>\n<li>Both <code>y_min</code> and <code>y_max</code> are measured from the top of the image. So <code>y_min</code> of 100 is 100 pixels down from the top of the image, and <code>y_max</code> of 250 is 250 pixels down from the top of the image.</li>\n</ul>\n<p>To put this into more context, if you have:</p>\n<ul>\n<li><code>x_min</code> = 300</li>\n<li><code>x_max</code> = 400</li>\n<li><code>y_min</code> = 100</li>\n<li><code>y_max</code> = 250</li>\n</ul>\n<p>You would have a bounding box that has an upper left hand coordinate of <code>(300, 100)</code>, a lower right coordinate of <code>(400, 200)</code>, and the bounding box itself would be 100 pixels wide by 150 pixels tall. Hope this helps!</p>",
          "rawMarkdown": "@weicongf here's my understanding of `x_min` and `y_min`. As you stated, the x-ray images have an `(x, y)` origin of `(0, 0)` - meaning the top left hand corner is the origin of the image. With that in mind:\n\n* The combination of `(x_min, y_min)` represent the upper left hand corner of the bounding box.\n* The combination of `(x_max, y_max)` represent the lower right hand corner of the bounding box.\n* Both `x_min` and `x_max` are measured from the left hand side of the image. So `x_min` of 300 is 300 pixels away from the left hand side of the image, and `x_max` of 400 is 400 pixels away from the left hand side of the image.\n* Both `y_min` and `y_max` are measured from the top of the image. So `y_min` of 100 is 100 pixels down from the top of the image, and `y_max` of 250 is 250 pixels down from the top of the image.\n\nTo put this into more context, if you have:\n\n* `x_min` = 300\n* `x_max` = 400\n* `y_min` = 100\n* `y_max` = 250\n\nYou would have a bounding box that has an upper left hand coordinate of `(300, 100)`, a lower right coordinate of `(400, 200)`, and the bounding box itself would be 100 pixels wide by 150 pixels tall. Hope this helps!",
          "votes": 3
        },
        {
          "id": 1140248,
          "postDate": "2021-01-05T21:59:03.407Z",
          "content": "<p>Appreciate your detailed expanation, Craig!</p>",
          "rawMarkdown": "Appreciate your detailed expanation, Craig!"
        }
      ]
    },
    {
      "id": 1148393,
      "postDate": "2021-01-11T05:24:30.840Z",
      "content": "<p>you will get better understanding from this image.</p>\n<p>The bounding box has the following (x, y) coordinates of its corners: top-left is (x_min, y_min) or (98px, 345px), top-right is (x_max, y_min) or (420px, 345px), bottom-left is (x_min, y_max) or (98px, 462px), bottom-right is (x_max, y_max) or (420px, 462px). As you see, coordinates of the bounding box's corners are calculated with respect to the top-left corner of the image which has (x, y) coordinates (0, 0).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3800337%2F319592810894409010e781a1feb91b50%2Fwinbig.png?generation=1610342604014303&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://albumentations.ai/docs/getting_started/bounding_boxes_augmentation/\" target=\"_blank\"></a></p>",
      "rawMarkdown": "you will get better understanding from this image.\n\nThe bounding box has the following (x, y) coordinates of its corners: top-left is (x_min, y_min) or (98px, 345px), top-right is (x_max, y_min) or (420px, 345px), bottom-left is (x_min, y_max) or (98px, 462px), bottom-right is (x_max, y_max) or (420px, 462px). As you see, coordinates of the bounding box's corners are calculated with respect to the top-left corner of the image which has (x, y) coordinates (0, 0).\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3800337%2F319592810894409010e781a1feb91b50%2Fwinbig.png?generation=1610342604014303&alt=media)\n\n\n[](https://albumentations.ai/docs/getting_started/bounding_boxes_augmentation/)",
      "votes": 1
    },
    {
      "id": 1145421,
      "postDate": "2021-01-09T06:10:53.300Z",
      "content": "<p>I was also curious about the submission, thank you for a good explanation.</p>",
      "rawMarkdown": "I was also curious about the submission, thank you for a good explanation."
    }
  ],
  "comments": [
    {
      "id": 1138903,
      "author_name": "Craig Thomas",
      "author_url": "",
      "post_date": "2021-01-05T03:52:47.673000",
      "content": "<p>Good question. From the competition description:</p>\n<blockquote>\n  <p>Images in the test set may contain more than one object. For each object in a given test image, you must predict a class ID, confidence score, and bounding box in format xmin ymin xmax ymax.</p>\n</blockquote>\n<p>Underneath that they have an example of an image with multiple objects, which looks like this in the submission file:</p>\n<pre><code>ID,TARGET\n004f33259ee4aef671c2b95d54e4be69,11 0.5 100 100 200 200 13 0.7 10 10 20 20\n</code></pre>\n<p>So this means that there are two predictions for image id <code>004f33259ee4aef671c2b95d54e4be69</code>:</p>\n<ol>\n<li>Object of class <code>11</code> (pleural thickening) with a confidence of <code>0.5</code> with a bounding box of <code>100</code> (xmin), <code>100</code> (ymin), <code>200</code> (xmax), and <code>200</code> (ymax).</li>\n<li>Object of class <code>13</code> (pulmonary fibrosis) with a confidence of <code>0.7</code> with a bounding box of <code>10</code> (xmin), <code>10</code> (ymin), <code>20</code> (xmax), and <code>20</code> (ymax).</li>\n</ol>\n<p>It looks like that second column <code>TARGET</code> is just a string with those numbers in it (class id, confidence, xmin, ymin, xmax, ymax) separated by a space. Multiple object predictions are just concatenated to that string with a space between each.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1139736,
          "author_name": "Weicong",
          "author_url": "",
          "post_date": "2021-01-05T15:33:31.090000",
          "content": "<p>Thanks, Thomas! You gave a detailed explanation. Another question: do you know the directions of x_min, y_min? </p>\n<p>I find the shape of the image is (vertical axis, horizontal axis), something like the matrix we use in python, rather than a math chart. Meanwhile, checking the image, we can find the origin (0,0) is at the left top. Do you know the meaning of x_min, y_min, x_max, y_max? For example, x_min = 300. Should I count it from the left top down 300 pixes, or go right 300 pixes?</p>\n<p>Thanks for your patient and sharing!</p>\n<p>Best,</p>\n<p>Weicong</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1140210,
          "author_name": "Craig Thomas",
          "author_url": "",
          "post_date": "2021-01-05T21:17:55.263000",
          "content": "<p><a href=\"https://www.kaggle.com/weicongf\" target=\"_blank\">@weicongf</a> here's my understanding of <code>x_min</code> and <code>y_min</code>. As you stated, the x-ray images have an <code>(x, y)</code> origin of <code>(0, 0)</code> - meaning the top left hand corner is the origin of the image. With that in mind:</p>\n<ul>\n<li>The combination of <code>(x_min, y_min)</code> represent the upper left hand corner of the bounding box.</li>\n<li>The combination of <code>(x_max, y_max)</code> represent the lower right hand corner of the bounding box.</li>\n<li>Both <code>x_min</code> and <code>x_max</code> are measured from the left hand side of the image. So <code>x_min</code> of 300 is 300 pixels away from the left hand side of the image, and <code>x_max</code> of 400 is 400 pixels away from the left hand side of the image.</li>\n<li>Both <code>y_min</code> and <code>y_max</code> are measured from the top of the image. So <code>y_min</code> of 100 is 100 pixels down from the top of the image, and <code>y_max</code> of 250 is 250 pixels down from the top of the image.</li>\n</ul>\n<p>To put this into more context, if you have:</p>\n<ul>\n<li><code>x_min</code> = 300</li>\n<li><code>x_max</code> = 400</li>\n<li><code>y_min</code> = 100</li>\n<li><code>y_max</code> = 250</li>\n</ul>\n<p>You would have a bounding box that has an upper left hand coordinate of <code>(300, 100)</code>, a lower right coordinate of <code>(400, 200)</code>, and the bounding box itself would be 100 pixels wide by 150 pixels tall. Hope this helps!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1140248,
          "author_name": "Weicong",
          "author_url": "",
          "post_date": "2021-01-05T21:59:03.407000",
          "content": "<p>Appreciate your detailed expanation, Craig!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1148393,
      "author_name": "akash desai",
      "author_url": "",
      "post_date": "2021-01-11T05:24:30.840000",
      "content": "<p>you will get better understanding from this image.</p>\n<p>The bounding box has the following (x, y) coordinates of its corners: top-left is (x_min, y_min) or (98px, 345px), top-right is (x_max, y_min) or (420px, 345px), bottom-left is (x_min, y_max) or (98px, 462px), bottom-right is (x_max, y_max) or (420px, 462px). As you see, coordinates of the bounding box's corners are calculated with respect to the top-left corner of the image which has (x, y) coordinates (0, 0).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3800337%2F319592810894409010e781a1feb91b50%2Fwinbig.png?generation=1610342604014303&amp;alt=media\" alt=\"\"></p>\n<p><a href=\"https://albumentations.ai/docs/getting_started/bounding_boxes_augmentation/\" target=\"_blank\"></a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1145421,
      "author_name": "stonebell",
      "author_url": "",
      "post_date": "2021-01-09T06:10:53.300000",
      "content": "<p>I was also curious about the submission, thank you for a good explanation.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1138845": "We know the submission format is like \"14 1 0 0 1 1\" if no found in the image. However, we can find a lot of training images have multiple classes (different diseases) in the single x-ray image. In this case, how to present it in our submission? Should we use multiple rows?",
    "1138903": "Good question. From the competition description:\n\n> Images in the test set may contain more than one object. For each object in a given test image, you must predict a class ID, confidence score, and bounding box in format xmin ymin xmax ymax.\n\nUnderneath that they have an example of an image with multiple objects, which looks like this in the submission file:\n\n    ID,TARGET\n    004f33259ee4aef671c2b95d54e4be69,11 0.5 100 100 200 200 13 0.7 10 10 20 20\n\nSo this means that there are two predictions for image id `004f33259ee4aef671c2b95d54e4be69`:\n\n1. Object of class `11` (pleural thickening) with a confidence of `0.5` with a bounding box of `100` (xmin), `100` (ymin), `200` (xmax), and `200` (ymax).\n2. Object of class `13` (pulmonary fibrosis) with a confidence of `0.7` with a bounding box of `10` (xmin), `10` (ymin), `20` (xmax), and `20` (ymax).\n\nIt looks like that second column `TARGET` is just a string with those numbers in it (class id, confidence, xmin, ymin, xmax, ymax) separated by a space. Multiple object predictions are just concatenated to that string with a space between each.",
    "1148393": "you will get better understanding from this image.\n\nThe bounding box has the following (x, y) coordinates of its corners: top-left is (x_min, y_min) or (98px, 345px), top-right is (x_max, y_min) or (420px, 345px), bottom-left is (x_min, y_max) or (98px, 462px), bottom-right is (x_max, y_max) or (420px, 462px). As you see, coordinates of the bounding box's corners are calculated with respect to the top-left corner of the image which has (x, y) coordinates (0, 0).\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3800337%2F319592810894409010e781a1feb91b50%2Fwinbig.png?generation=1610342604014303&alt=media)\n\n\n[](https://albumentations.ai/docs/getting_started/bounding_boxes_augmentation/)",
    "1145421": "I was also curious about the submission, thank you for a good explanation."
  }
}