{
  "id": 227034,
  "title": "mmdetection inference output",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/227034",
  "author_name": "Daniel Shan",
  "post_date": "2021-03-18T17:13:19.849000",
  "votes": 3,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I have some questions about mmdetection inference output, and I couldn't find any documentation to answer my questions :(</p>\n<p>So in my test pipeline I do image scaling:</p>\n<pre><code>test_pipeline = [\n        ...\n        img_scale=[(1333, 480), (1333, 960)],\n        flip=False,\n        transforms=[\n            dict(type='Resize', keep_ratio=True),\n</code></pre>\n<p>The output per box looks like this: <code>[579, 279, 640, 347, .85]</code>.  My question is …</p>\n<ol>\n<li><p>Are the output coords for the <strong>resized</strong> image?  Or the <strong>original</strong> test image?  i.e. do I need to convert the output coords to the correct scale for the original test image?</p></li>\n<li><p>Are the output coords in (xmin, ymin, xmax, ymax) format?  It kind of looks like it but wanted to make sure it wasn't actually something else like (ymin, xmin, ymax, xmax).</p></li>\n</ol>",
  "messages": [
    {
      "id": 1244048,
      "postDate": "2021-03-18T17:13:19.850Z",
      "content": "<p>I have some questions about mmdetection inference output, and I couldn't find any documentation to answer my questions :(</p>\n<p>So in my test pipeline I do image scaling:</p>\n<pre><code>test_pipeline = [\n        ...\n        img_scale=[(1333, 480), (1333, 960)],\n        flip=False,\n        transforms=[\n            dict(type='Resize', keep_ratio=True),\n</code></pre>\n<p>The output per box looks like this: <code>[579, 279, 640, 347, .85]</code>.  My question is …</p>\n<ol>\n<li><p>Are the output coords for the <strong>resized</strong> image?  Or the <strong>original</strong> test image?  i.e. do I need to convert the output coords to the correct scale for the original test image?</p></li>\n<li><p>Are the output coords in (xmin, ymin, xmax, ymax) format?  It kind of looks like it but wanted to make sure it wasn't actually something else like (ymin, xmin, ymax, xmax).</p></li>\n</ol>",
      "rawMarkdown": "I have some questions about mmdetection inference output, and I couldn't find any documentation to answer my questions :(\n\nSo in my test pipeline I do image scaling:\n```\ntest_pipeline = [\n        ...\n        img_scale=[(1333, 480), (1333, 960)],\n        flip=False,\n        transforms=[\n            dict(type='Resize', keep_ratio=True),\n```\n\nThe output per box looks like this: `[579, 279, 640, 347, .85]`.  My question is ...\n\n1. Are the output coords for the **resized** image?  Or the **original** test image?  i.e. do I need to convert the output coords to the correct scale for the original test image?\n\n2. Are the output coords in (xmin, ymin, xmax, ymax) format?  It kind of looks like it but wanted to make sure it wasn't actually something else like (ymin, xmin, ymax, xmax).",
      "votes": 3
    },
    {
      "id": 1244125,
      "postDate": "2021-03-18T17:52:45.953Z",
      "content": "<ol>\n<li>They are for the size of your input images. If you changed the size from the original test images you need to convert them back.</li>\n<li>Yes as you thought.</li>\n</ol>\n<p>Btw, it is quite amazing that they don't provide any documentation on the output, their <a href=\"https://github.com/open-mmlab/mmdetection/issues/2463\" target=\"_blank\">response</a> to a question in this regard was not particularly helpful.. 😏</p>",
      "rawMarkdown": "1. They are for the size of your input images. If you changed the size from the original test images you need to convert them back.\n2. Yes as you thought.\n\nBtw, it is quite amazing that they don't provide any documentation on the output, their [response](https://github.com/open-mmlab/mmdetection/issues/2463) to a question in this regard was not particularly helpful.. 😏",
      "votes": 4,
      "replies": [
        {
          "id": 1254390,
          "postDate": "2021-03-27T16:21:24.307Z",
          "content": "<p><a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> <br>\nCan you answer the same question please</p>\n<ul>\n<li><p>I am forwarding original image size to the network but i am resizing to 1024 in the config. Is the boxes (which comes as output) is for coordinates in image size 1024 or original image</p></li>\n<li><p>can you say what the output from <code>output = inference_detector(model,image)</code>?</p></li>\n</ul>",
          "rawMarkdown": "@hannes82 \nCan you answer the same question please\n- I am forwarding original image size to the network but i am resizing to 1024 in the config. Is the boxes (which comes as output) is for coordinates in image size 1024 or original image\n\n- can you say what the output from `output = inference_detector(model,image)`?"
        },
        {
          "id": 1254476,
          "postDate": "2021-03-27T17:45:08.993Z",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> The coordinates are for the original image size.</p>\n<p>I don't use inference_detector, just results.pkl as output of test.py. So I don't know.</p>",
          "rawMarkdown": "@morizin The coordinates are for the original image size.\n\nI don't use inference_detector, just results.pkl as output of test.py. So I don't know.",
          "votes": 1
        },
        {
          "id": 1254488,
          "postDate": "2021-03-27T17:58:09.573Z",
          "content": "<p>Thanks for your reply</p>",
          "rawMarkdown": "Thanks for your reply",
          "votes": 1
        },
        {
          "id": 1254492,
          "postDate": "2021-03-27T18:04:03.910Z",
          "content": "<p>can you please clarify these also please<br>\n(ymin, xmin, ymax, xmax)<br>\nor <br>\n(xmin, ymin, xmax, ymax)</p>",
          "rawMarkdown": "can you please clarify these also please\n(ymin, xmin, ymax, xmax)\nor \n(xmin, ymin, xmax, ymax)"
        },
        {
          "id": 1254502,
          "postDate": "2021-03-27T18:15:42.723Z",
          "content": "<p>you are welcome. the latter: x_min,y_min,x_max, y_max</p>",
          "rawMarkdown": "you are welcome. the latter: x_min,y_min,x_max, y_max",
          "votes": 1
        }
      ]
    },
    {
      "id": 1255319,
      "postDate": "2021-03-28T17:15:36.710Z",
      "content": "<p><a href=\"https://www.kaggle.com/zekunn\" target=\"_blank\">@zekunn</a> my model LB is also different than CV but that seems to be the case for many people :) If it's a LOT worse, then maybe you're processing the output incorrectly?  That's the issue I had, thank you <a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> for your response!</p>",
      "rawMarkdown": "@zekunn my model LB is also different than CV but that seems to be the case for many people :) If it's a LOT worse, then maybe you're processing the output incorrectly?  That's the issue I had, thank you @hannes82 for your response!",
      "votes": 1,
      "replies": [
        {
          "id": 1255598,
          "postDate": "2021-03-29T00:26:38.547Z",
          "content": "<p>Hello,which process did you choose?I change the position of x and y ,but get worse.</p>",
          "rawMarkdown": "Hello,which process did you choose?I change the position of x and y ,but get worse."
        }
      ]
    },
    {
      "id": 1255307,
      "postDate": "2021-03-28T17:04:22.757Z",
      "content": "<p><a href=\"https://www.kaggle.com/kuanzhang\" target=\"_blank\">@kuanzhang</a> the img_scale values are just the size to which the images are resized to in your pipeline.  In the example above I'm using multiscaling which you can read about here: <a href=\"https://github.com/open-mmlab/mmdetection/issues/78#issuecomment-432971302\" target=\"_blank\">https://github.com/open-mmlab/mmdetection/issues/78#issuecomment-432971302</a></p>\n<p>As for training on downsampled data and testing on original data, I haven't tried that, I've always used the same dataset for both training and testing.  I don't think there's a right answer for what img_scale should* be although obviously the scale should make sense for the size of the images you're using.</p>",
      "rawMarkdown": "@kuanzhang the img_scale values are just the size to which the images are resized to in your pipeline.  In the example above I'm using multiscaling which you can read about here: https://github.com/open-mmlab/mmdetection/issues/78#issuecomment-432971302\n\nAs for training on downsampled data and testing on original data, I haven't tried that, I've always used the same dataset for both training and testing.  I don't think there's a right answer for what img_scale should* be although obviously the scale should make sense for the size of the images you're using.",
      "replies": [
        {
          "id": 1255344,
          "postDate": "2021-03-28T17:44:44.203Z",
          "content": "<p>Got it. <a href=\"https://www.kaggle.com/danshan\" target=\"_blank\">@danshan</a> That is great help! So if I train and test both on 1024x1024 datasize as I previously did on YOLO models, I may still keep img_scale = (1333, 800) in cfg_file, because this value (1333,800) here is not for the input or output data size. it is just the size assigned to scale the data and used in the model. Is my understanding correct?</p>\n<p>Besides, how do you prepare the json file for 1500 test dataset? I guess I already have partitions in YOLO format and only need to prepare the associated json files. And the json files for test should be different from training/testing ---- no \"annotations\" for testing. Do you have recommended one to follow? I notice this link: <a href=\"https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset\" target=\"_blank\">https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset</a> </p>\n<p>Thanks!</p>",
          "rawMarkdown": "Got it. @danshan That is great help! So if I train and test both on 1024x1024 datasize as I previously did on YOLO models, I may still keep img_scale = (1333, 800) in cfg_file, because this value (1333,800) here is not for the input or output data size. it is just the size assigned to scale the data and used in the model. Is my understanding correct?\n\nBesides, how do you prepare the json file for 1500 test dataset? I guess I already have partitions in YOLO format and only need to prepare the associated json files. And the json files for test should be different from training/testing ---- no \"annotations\" for testing. Do you have recommended one to follow? I notice this link: https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset \n\nThanks!"
        },
        {
          "id": 1255528,
          "postDate": "2021-03-28T21:56:23.440Z",
          "content": "<p>Your understanding is correct :)</p>\n<p>I prepared the json files with the notebook you linked.  For the test dataset, you can modify the notebook to produce a json with just the image names.  The annotations section will be empty.</p>",
          "rawMarkdown": "Your understanding is correct :)\n\nI prepared the json files with the notebook you linked.  For the test dataset, you can modify the notebook to produce a json with just the image names.  The annotations section will be empty."
        },
        {
          "id": 1255597,
          "postDate": "2021-03-29T00:20:40.900Z",
          "content": "<p>Thanks, Daniel! I finally made my first vfnet single model trained and submitted. (0.16 and 0.25 before and after the filtering).</p>\n<p>Do you know how to do ensemble inference in MMDetection? Is there a similar way to YOLO that we can do \"python detect.py --weights 1.pt 2.pt 3.pt…\"?   <a href=\"https://www.kaggle.com/danshan\" target=\"_blank\">@danshan</a> 👍</p>",
          "rawMarkdown": "Thanks, Daniel! I finally made my first vfnet single model trained and submitted. (0.16 and 0.25 before and after the filtering).\n\nDo you know how to do ensemble inference in MMDetection? Is there a similar way to YOLO that we can do \"python detect.py --weights 1.pt 2.pt 3.pt...\"?   @danshan 👍"
        },
        {
          "id": 1255600,
          "postDate": "2021-03-29T00:29:13.543Z",
          "content": "<p>Not that I know of unfortunately :( at least not natively within mmdet itself.</p>",
          "rawMarkdown": "Not that I know of unfortunately :( at least not natively within mmdet itself."
        }
      ]
    },
    {
      "id": 1255246,
      "postDate": "2021-03-28T16:05:47.953Z",
      "content": "<p><a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a>, <a href=\"https://www.kaggle.com/danshan\" target=\"_blank\">@danshan</a>, how shall we understand the img_scale values, e.g.,  1333, 480 and 960?</p>\n<p>Can we train on the 3x downsampled training data, and inference on the original-size test data? What would be the img_scale accordingly? Thanks!</p>",
      "rawMarkdown": "@hannes82, @danshan, how shall we understand the img_scale values, e.g.,  1333, 480 and 960?\n\nCan we train on the 3x downsampled training data, and inference on the original-size test data? What would be the img_scale accordingly? Thanks!\n",
      "replies": [
        {
          "id": 1255312,
          "postDate": "2021-03-28T17:10:00.190Z",
          "content": "<p>If I understand it correctly, the the short edge of the image will be randomly sampled from [480,960].</p>\n<p>I suspect that significantly changing image size at inference time compared to training time is not a good idea, but I might be wrong here.</p>\n<p>Setting an image scale at inference time (i.e. TTA) may hurt or improve the outcome.</p>",
          "rawMarkdown": "If I understand it correctly, the the short edge of the image will be randomly sampled from [480,960].\n\nI suspect that significantly changing image size at inference time compared to training time is not a good idea, but I might be wrong here.\n\nSetting an image scale at inference time (i.e. TTA) may hurt or improve the outcome."
        },
        {
          "id": 1255323,
          "postDate": "2021-03-28T17:17:35.700Z",
          "content": "<p><code>If I understand it correctly, the the short edge of the image will be randomly sampled from [480,960].</code></p>\n<p>Yup.</p>\n<p>Also for clarity: this was an experiment I was doing and it actually did not work very well XD.  I've had better results with fixed scale at test time.</p>",
          "rawMarkdown": "`If I understand it correctly, the the short edge of the image will be randomly sampled from [480,960].`\n\nYup.\n\nAlso for clarity: this was an experiment I was doing and it actually did not work very well XD.  I've had better results with fixed scale at test time.\n\n"
        },
        {
          "id": 1255345,
          "postDate": "2021-03-28T17:45:39.347Z",
          "content": "<p>Got it. Thanks, <a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> and <a href=\"https://www.kaggle.com/danshan\" target=\"_blank\">@danshan</a> 🙂</p>",
          "rawMarkdown": "Got it. Thanks, @hannes82 and @danshan 🙂"
        },
        {
          "id": 1255675,
          "postDate": "2021-03-29T04:00:07.360Z",
          "content": "<p>how to decode the boxes when I use that?</p>",
          "rawMarkdown": "how to decode the boxes when I use that?"
        }
      ]
    },
    {
      "id": 1254725,
      "postDate": "2021-03-28T02:56:28.783Z",
      "content": "<p>Hello,my mmdetection models lb is much different from cv….Is there something wrong?</p>",
      "rawMarkdown": "Hello,my mmdetection models lb is much different from cv....Is there something wrong?"
    }
  ],
  "comments": [
    {
      "id": 1244125,
      "author_name": "Hannes Öhler",
      "author_url": "",
      "post_date": "2021-03-18T17:52:45.953000",
      "content": "<ol>\n<li>They are for the size of your input images. If you changed the size from the original test images you need to convert them back.</li>\n<li>Yes as you thought.</li>\n</ol>\n<p>Btw, it is quite amazing that they don't provide any documentation on the output, their <a href=\"https://github.com/open-mmlab/mmdetection/issues/2463\" target=\"_blank\">response</a> to a question in this regard was not particularly helpful.. 😏</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1254390,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-27T16:21:24.307000",
          "content": "<p><a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> <br>\nCan you answer the same question please</p>\n<ul>\n<li><p>I am forwarding original image size to the network but i am resizing to 1024 in the config. Is the boxes (which comes as output) is for coordinates in image size 1024 or original image</p></li>\n<li><p>can you say what the output from <code>output = inference_detector(model,image)</code>?</p></li>\n</ul>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254476,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-27T17:45:08.993000",
          "content": "<p><a href=\"https://www.kaggle.com/morizin\" target=\"_blank\">@morizin</a> The coordinates are for the original image size.</p>\n<p>I don't use inference_detector, just results.pkl as output of test.py. So I don't know.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1254488,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-27T17:58:09.573000",
          "content": "<p>Thanks for your reply</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1254492,
          "author_name": "Mohammed Rizin V K",
          "author_url": "",
          "post_date": "2021-03-27T18:04:03.910000",
          "content": "<p>can you please clarify these also please<br>\n(ymin, xmin, ymax, xmax)<br>\nor <br>\n(xmin, ymin, xmax, ymax)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1254502,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-27T18:15:42.723000",
          "content": "<p>you are welcome. the latter: x_min,y_min,x_max, y_max</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1255319,
      "author_name": "Daniel Shan",
      "author_url": "",
      "post_date": "2021-03-28T17:15:36.710000",
      "content": "<p><a href=\"https://www.kaggle.com/zekunn\" target=\"_blank\">@zekunn</a> my model LB is also different than CV but that seems to be the case for many people :) If it's a LOT worse, then maybe you're processing the output incorrectly?  That's the issue I had, thank you <a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> for your response!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1255598,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-03-29T00:26:38.547000",
          "content": "<p>Hello,which process did you choose?I change the position of x and y ,but get worse.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1255307,
      "author_name": "Daniel Shan",
      "author_url": "",
      "post_date": "2021-03-28T17:04:22.757000",
      "content": "<p><a href=\"https://www.kaggle.com/kuanzhang\" target=\"_blank\">@kuanzhang</a> the img_scale values are just the size to which the images are resized to in your pipeline.  In the example above I'm using multiscaling which you can read about here: <a href=\"https://github.com/open-mmlab/mmdetection/issues/78#issuecomment-432971302\" target=\"_blank\">https://github.com/open-mmlab/mmdetection/issues/78#issuecomment-432971302</a></p>\n<p>As for training on downsampled data and testing on original data, I haven't tried that, I've always used the same dataset for both training and testing.  I don't think there's a right answer for what img_scale should* be although obviously the scale should make sense for the size of the images you're using.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1255344,
          "author_name": "Kuan Zhang",
          "author_url": "",
          "post_date": "2021-03-28T17:44:44.203000",
          "content": "<p>Got it. <a href=\"https://www.kaggle.com/danshan\" target=\"_blank\">@danshan</a> That is great help! So if I train and test both on 1024x1024 datasize as I previously did on YOLO models, I may still keep img_scale = (1333, 800) in cfg_file, because this value (1333,800) here is not for the input or output data size. it is just the size assigned to scale the data and used in the model. Is my understanding correct?</p>\n<p>Besides, how do you prepare the json file for 1500 test dataset? I guess I already have partitions in YOLO format and only need to prepare the associated json files. And the json files for test should be different from training/testing ---- no \"annotations\" for testing. Do you have recommended one to follow? I notice this link: <a href=\"https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset\" target=\"_blank\">https://www.kaggle.com/sreevishnudamodaran/vinbigdata-fusing-bboxes-coco-dataset</a> </p>\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1255528,
          "author_name": "Daniel Shan",
          "author_url": "",
          "post_date": "2021-03-28T21:56:23.440000",
          "content": "<p>Your understanding is correct :)</p>\n<p>I prepared the json files with the notebook you linked.  For the test dataset, you can modify the notebook to produce a json with just the image names.  The annotations section will be empty.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1255597,
          "author_name": "Kuan Zhang",
          "author_url": "",
          "post_date": "2021-03-29T00:20:40.900000",
          "content": "<p>Thanks, Daniel! I finally made my first vfnet single model trained and submitted. (0.16 and 0.25 before and after the filtering).</p>\n<p>Do you know how to do ensemble inference in MMDetection? Is there a similar way to YOLO that we can do \"python detect.py --weights 1.pt 2.pt 3.pt…\"?   <a href=\"https://www.kaggle.com/danshan\" target=\"_blank\">@danshan</a> 👍</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1255600,
          "author_name": "Daniel Shan",
          "author_url": "",
          "post_date": "2021-03-29T00:29:13.543000",
          "content": "<p>Not that I know of unfortunately :( at least not natively within mmdet itself.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1255246,
      "author_name": "Kuan Zhang",
      "author_url": "",
      "post_date": "2021-03-28T16:05:47.953000",
      "content": "<p><a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a>, <a href=\"https://www.kaggle.com/danshan\" target=\"_blank\">@danshan</a>, how shall we understand the img_scale values, e.g.,  1333, 480 and 960?</p>\n<p>Can we train on the 3x downsampled training data, and inference on the original-size test data? What would be the img_scale accordingly? Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1255312,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-03-28T17:10:00.190000",
          "content": "<p>If I understand it correctly, the the short edge of the image will be randomly sampled from [480,960].</p>\n<p>I suspect that significantly changing image size at inference time compared to training time is not a good idea, but I might be wrong here.</p>\n<p>Setting an image scale at inference time (i.e. TTA) may hurt or improve the outcome.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1255323,
          "author_name": "Daniel Shan",
          "author_url": "",
          "post_date": "2021-03-28T17:17:35.700000",
          "content": "<p><code>If I understand it correctly, the the short edge of the image will be randomly sampled from [480,960].</code></p>\n<p>Yup.</p>\n<p>Also for clarity: this was an experiment I was doing and it actually did not work very well XD.  I've had better results with fixed scale at test time.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1255345,
          "author_name": "Kuan Zhang",
          "author_url": "",
          "post_date": "2021-03-28T17:45:39.347000",
          "content": "<p>Got it. Thanks, <a href=\"https://www.kaggle.com/hannes82\" target=\"_blank\">@hannes82</a> and <a href=\"https://www.kaggle.com/danshan\" target=\"_blank\">@danshan</a> 🙂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1255675,
          "author_name": "Zekun",
          "author_url": "",
          "post_date": "2021-03-29T04:00:07.360000",
          "content": "<p>how to decode the boxes when I use that?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1254725,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-03-28T02:56:28.783000",
      "content": "<p>Hello,my mmdetection models lb is much different from cv….Is there something wrong?</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1244048": "I have some questions about mmdetection inference output, and I couldn't find any documentation to answer my questions :(\n\nSo in my test pipeline I do image scaling:\n```\ntest_pipeline = [\n        ...\n        img_scale=[(1333, 480), (1333, 960)],\n        flip=False,\n        transforms=[\n            dict(type='Resize', keep_ratio=True),\n```\n\nThe output per box looks like this: `[579, 279, 640, 347, .85]`.  My question is ...\n\n1. Are the output coords for the **resized** image?  Or the **original** test image?  i.e. do I need to convert the output coords to the correct scale for the original test image?\n\n2. Are the output coords in (xmin, ymin, xmax, ymax) format?  It kind of looks like it but wanted to make sure it wasn't actually something else like (ymin, xmin, ymax, xmax).",
    "1244125": "1. They are for the size of your input images. If you changed the size from the original test images you need to convert them back.\n2. Yes as you thought.\n\nBtw, it is quite amazing that they don't provide any documentation on the output, their [response](https://github.com/open-mmlab/mmdetection/issues/2463) to a question in this regard was not particularly helpful.. 😏",
    "1255319": "@zekunn my model LB is also different than CV but that seems to be the case for many people :) If it's a LOT worse, then maybe you're processing the output incorrectly?  That's the issue I had, thank you @hannes82 for your response!",
    "1255307": "@kuanzhang the img_scale values are just the size to which the images are resized to in your pipeline.  In the example above I'm using multiscaling which you can read about here: https://github.com/open-mmlab/mmdetection/issues/78#issuecomment-432971302\n\nAs for training on downsampled data and testing on original data, I haven't tried that, I've always used the same dataset for both training and testing.  I don't think there's a right answer for what img_scale should* be although obviously the scale should make sense for the size of the images you're using.",
    "1255246": "@hannes82, @danshan, how shall we understand the img_scale values, e.g.,  1333, 480 and 960?\n\nCan we train on the 3x downsampled training data, and inference on the original-size test data? What would be the img_scale accordingly? Thanks!\n",
    "1254725": "Hello,my mmdetection models lb is much different from cv....Is there something wrong?"
  }
}