{
  "id": 227891,
  "title": "MMdetection questions",
  "url": "/competitions/vinbigdata-chest-xray-abnormalities-detection/discussion/227891",
  "author_name": "Mostafa Ibrahim",
  "post_date": "2021-03-22T16:47:05.426000",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello everyone, I  have built a YoloV5 model and now I was thinking about trying mmdetection. I have checked their documentation and this notebook <a href=\"https://www.kaggle.com/gauravsingh1/mmdet-eval-fasterrcnn-with-wbf/data\" target=\"_blank\">https://www.kaggle.com/gauravsingh1/mmdet-eval-fasterrcnn-with-wbf/data</a><br>\nBut, I couldn't find answers to my questions. This notebook is using this coco converted dataset<br>\n<a href=\"https://www.kaggle.com/sreevishnudamodaran/vinbigdata-coco-dataset-with-wbf-3x-downscaled\" target=\"_blank\">https://www.kaggle.com/sreevishnudamodaran/vinbigdata-coco-dataset-with-wbf-3x-downscaled</a><br>\nMy main question is that shouldn't the bounding boxes be normalized? I copied the notebook, ran it several times, made some changes and the score is always somewhere between 0.010 and 0.030 which is quite low for a FasterRCNN. The default YoloV5  config gives 0.12x…. so I am pretty sure something is wrong, I am just trying to figure out what it is. There aren't that many good mmdetection tutorials out there. (Please don't tell me to check their documentation as I have already done that).</p>\n<p>In summary that notebook / my one is using:<br>\nFasterRCNN with pretrained weights<br>\ntrained for around 30 epochs on 3x Downscaled images (so around 700-800 x 700-800)</p>",
  "messages": [
    {
      "id": 1254745,
      "postDate": "2021-03-28T04:24:08.230Z",
      "content": "<p>I have the same question.When I visualize the boxes , it seems right.</p>",
      "rawMarkdown": "I have the same question.When I visualize the boxes , it seems right.",
      "votes": 1
    },
    {
      "id": 1248519,
      "postDate": "2021-03-22T16:47:05.427Z",
      "content": "<p>Hello everyone, I  have built a YoloV5 model and now I was thinking about trying mmdetection. I have checked their documentation and this notebook <a href=\"https://www.kaggle.com/gauravsingh1/mmdet-eval-fasterrcnn-with-wbf/data\" target=\"_blank\">https://www.kaggle.com/gauravsingh1/mmdet-eval-fasterrcnn-with-wbf/data</a><br>\nBut, I couldn't find answers to my questions. This notebook is using this coco converted dataset<br>\n<a href=\"https://www.kaggle.com/sreevishnudamodaran/vinbigdata-coco-dataset-with-wbf-3x-downscaled\" target=\"_blank\">https://www.kaggle.com/sreevishnudamodaran/vinbigdata-coco-dataset-with-wbf-3x-downscaled</a><br>\nMy main question is that shouldn't the bounding boxes be normalized? I copied the notebook, ran it several times, made some changes and the score is always somewhere between 0.010 and 0.030 which is quite low for a FasterRCNN. The default YoloV5  config gives 0.12x…. so I am pretty sure something is wrong, I am just trying to figure out what it is. There aren't that many good mmdetection tutorials out there. (Please don't tell me to check their documentation as I have already done that).</p>\n<p>In summary that notebook / my one is using:<br>\nFasterRCNN with pretrained weights<br>\ntrained for around 30 epochs on 3x Downscaled images (so around 700-800 x 700-800)</p>",
      "rawMarkdown": "Hello everyone, I  have built a YoloV5 model and now I was thinking about trying mmdetection. I have checked their documentation and this notebook https://www.kaggle.com/gauravsingh1/mmdet-eval-fasterrcnn-with-wbf/data\nBut, I couldn't find answers to my questions. This notebook is using this coco converted dataset\nhttps://www.kaggle.com/sreevishnudamodaran/vinbigdata-coco-dataset-with-wbf-3x-downscaled\nMy main question is that shouldn't the bounding boxes be normalized? I copied the notebook, ran it several times, made some changes and the score is always somewhere between 0.010 and 0.030 which is quite low for a FasterRCNN. The default YoloV5  config gives 0.12x.... so I am pretty sure something is wrong, I am just trying to figure out what it is. There aren't that many good mmdetection tutorials out there. (Please don't tell me to check their documentation as I have already done that).\n\nIn summary that notebook / my one is using:\nFasterRCNN with pretrained weights\ntrained for around 30 epochs on 3x Downscaled images (so around 700-800 x 700-800)",
      "votes": 1
    },
    {
      "id": 1248670,
      "postDate": "2021-03-22T18:34:25.423Z",
      "content": "<p>This code works for square dataset (256x256 for example. In this case dim = 256), you could try to adapt it. The DataFrame is the original one, with original coordinates. Tell me if you have any questions.</p>\n<pre><code>def to_coco(df_path, name) :\n\n    df = pd.read_csv(df_path) \n    annotations = []\n    images = []\n    obj_count = 0\n    for i, image_id in enumerate(df.image_id.unique()) :\n        images.append(dict(file_name = f'{image_id}.png', height=dim, width=dim, id=i))\n\n        for index, row in df.query(\"image_id == @image_id\").iterrows():\n            h_ratio = dim / row['height']\n            w_ratio = dim / row['width']\n            data_anno = dict(\n                image_id=i,\n                iscrowd=0,\n                id=obj_count,\n                category_id=row['class_id'],\n                bbox=[\n                float(row[\"x_min\"]) * w_ratio,\n                float(row[\"y_min\"]) * h_ratio,\n                float(row[\"x_max\"]) * w_ratio - float(row[\"x_min\"]) * w_ratio,\n                float(row[\"y_max\"]) * h_ratio - float(row[\"y_min\"]) * h_ratio],\n                area=(float(row[\"x_max\"]) * w_ratio - float(row[\"x_min\"]) * w_ratio) * (float(row[\"y_max\"]) * h_ratio - float(row[\"y_min\"]) * h_ratio))\n            annotations.append(data_anno)\n            obj_count += 1\n    coco_format_json = dict(\n        images=images,\n        annotations=annotations,\n        categories=cat2label)\n    mmcv.dump(coco_format_json, f'{name}')\n</code></pre>",
      "rawMarkdown": "This code works for square dataset (256x256 for example. In this case dim = 256), you could try to adapt it. The DataFrame is the original one, with original coordinates. Tell me if you have any questions.\n```\n\ndef to_coco(df_path, name) :\n\n    df = pd.read_csv(df_path) \n    annotations = []\n    images = []\n    obj_count = 0\n    for i, image_id in enumerate(df.image_id.unique()) :\n        images.append(dict(file_name = f'{image_id}.png', height=dim, width=dim, id=i))\n\n        for index, row in df.query(\"image_id == @image_id\").iterrows():\n            h_ratio = dim / row['height']\n            w_ratio = dim / row['width']\n            data_anno = dict(\n                image_id=i,\n                iscrowd=0,\n                id=obj_count,\n                category_id=row['class_id'],\n                bbox=[\n                float(row[\"x_min\"]) * w_ratio,\n                float(row[\"y_min\"]) * h_ratio,\n                float(row[\"x_max\"]) * w_ratio - float(row[\"x_min\"]) * w_ratio,\n                float(row[\"y_max\"]) * h_ratio - float(row[\"y_min\"]) * h_ratio],\n                area=(float(row[\"x_max\"]) * w_ratio - float(row[\"x_min\"]) * w_ratio) * (float(row[\"y_max\"]) * h_ratio - float(row[\"y_min\"]) * h_ratio))\n            annotations.append(data_anno)\n            obj_count += 1\n    coco_format_json = dict(\n        images=images,\n        annotations=annotations,\n        categories=cat2label)\n    mmcv.dump(coco_format_json, f'{name}')\n```",
      "votes": 2,
      "replies": [
        {
          "id": 1248742,
          "postDate": "2021-03-22T20:08:35.607Z",
          "content": "<p>Awesome, thanks very much! I will try it and let you know. This isn't writing normalised coordinates to the json file though, so I assume you don't need to normalise them?</p>",
          "rawMarkdown": "Awesome, thanks very much! I will try it and let you know. This isn't writing normalised coordinates to the json file though, so I assume you don't need to normalise them?",
          "votes": 1
        },
        {
          "id": 1249833,
          "postDate": "2021-03-23T14:56:50.240Z",
          "content": "<p><code>float(row[\"x_min\"]) * w_ratio</code> You just need to adapt the new coordinates to the new image dimension </p>",
          "rawMarkdown": "``float(row[\"x_min\"]) * w_ratio`` You just need to adapt the new coordinates to the new image dimension "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1254745,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-03-28T04:24:08.230000",
      "content": "<p>I have the same question.When I visualize the boxes , it seems right.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1248670,
      "author_name": "Matthieu Planté",
      "author_url": "",
      "post_date": "2021-03-22T18:34:25.423000",
      "content": "<p>This code works for square dataset (256x256 for example. In this case dim = 256), you could try to adapt it. The DataFrame is the original one, with original coordinates. Tell me if you have any questions.</p>\n<pre><code>def to_coco(df_path, name) :\n\n    df = pd.read_csv(df_path) \n    annotations = []\n    images = []\n    obj_count = 0\n    for i, image_id in enumerate(df.image_id.unique()) :\n        images.append(dict(file_name = f'{image_id}.png', height=dim, width=dim, id=i))\n\n        for index, row in df.query(\"image_id == @image_id\").iterrows():\n            h_ratio = dim / row['height']\n            w_ratio = dim / row['width']\n            data_anno = dict(\n                image_id=i,\n                iscrowd=0,\n                id=obj_count,\n                category_id=row['class_id'],\n                bbox=[\n                float(row[\"x_min\"]) * w_ratio,\n                float(row[\"y_min\"]) * h_ratio,\n                float(row[\"x_max\"]) * w_ratio - float(row[\"x_min\"]) * w_ratio,\n                float(row[\"y_max\"]) * h_ratio - float(row[\"y_min\"]) * h_ratio],\n                area=(float(row[\"x_max\"]) * w_ratio - float(row[\"x_min\"]) * w_ratio) * (float(row[\"y_max\"]) * h_ratio - float(row[\"y_min\"]) * h_ratio))\n            annotations.append(data_anno)\n            obj_count += 1\n    coco_format_json = dict(\n        images=images,\n        annotations=annotations,\n        categories=cat2label)\n    mmcv.dump(coco_format_json, f'{name}')\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 1248742,
          "author_name": "Mostafa Ibrahim",
          "author_url": "",
          "post_date": "2021-03-22T20:08:35.607000",
          "content": "<p>Awesome, thanks very much! I will try it and let you know. This isn't writing normalised coordinates to the json file though, so I assume you don't need to normalise them?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1249833,
          "author_name": "Matthieu Planté",
          "author_url": "",
          "post_date": "2021-03-23T14:56:50.240000",
          "content": "<p><code>float(row[\"x_min\"]) * w_ratio</code> You just need to adapt the new coordinates to the new image dimension </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1254745": "I have the same question.When I visualize the boxes , it seems right.",
    "1248519": "Hello everyone, I  have built a YoloV5 model and now I was thinking about trying mmdetection. I have checked their documentation and this notebook https://www.kaggle.com/gauravsingh1/mmdet-eval-fasterrcnn-with-wbf/data\nBut, I couldn't find answers to my questions. This notebook is using this coco converted dataset\nhttps://www.kaggle.com/sreevishnudamodaran/vinbigdata-coco-dataset-with-wbf-3x-downscaled\nMy main question is that shouldn't the bounding boxes be normalized? I copied the notebook, ran it several times, made some changes and the score is always somewhere between 0.010 and 0.030 which is quite low for a FasterRCNN. The default YoloV5  config gives 0.12x.... so I am pretty sure something is wrong, I am just trying to figure out what it is. There aren't that many good mmdetection tutorials out there. (Please don't tell me to check their documentation as I have already done that).\n\nIn summary that notebook / my one is using:\nFasterRCNN with pretrained weights\ntrained for around 30 epochs on 3x Downscaled images (so around 700-800 x 700-800)",
    "1248670": "This code works for square dataset (256x256 for example. In this case dim = 256), you could try to adapt it. The DataFrame is the original one, with original coordinates. Tell me if you have any questions.\n```\n\ndef to_coco(df_path, name) :\n\n    df = pd.read_csv(df_path) \n    annotations = []\n    images = []\n    obj_count = 0\n    for i, image_id in enumerate(df.image_id.unique()) :\n        images.append(dict(file_name = f'{image_id}.png', height=dim, width=dim, id=i))\n\n        for index, row in df.query(\"image_id == @image_id\").iterrows():\n            h_ratio = dim / row['height']\n            w_ratio = dim / row['width']\n            data_anno = dict(\n                image_id=i,\n                iscrowd=0,\n                id=obj_count,\n                category_id=row['class_id'],\n                bbox=[\n                float(row[\"x_min\"]) * w_ratio,\n                float(row[\"y_min\"]) * h_ratio,\n                float(row[\"x_max\"]) * w_ratio - float(row[\"x_min\"]) * w_ratio,\n                float(row[\"y_max\"]) * h_ratio - float(row[\"y_min\"]) * h_ratio],\n                area=(float(row[\"x_max\"]) * w_ratio - float(row[\"x_min\"]) * w_ratio) * (float(row[\"y_max\"]) * h_ratio - float(row[\"y_min\"]) * h_ratio))\n            annotations.append(data_anno)\n            obj_count += 1\n    coco_format_json = dict(\n        images=images,\n        annotations=annotations,\n        categories=cat2label)\n    mmcv.dump(coco_format_json, f'{name}')\n```"
  }
}